Multi-source land cover data fusion method based on conflict index and self-supervised reconciliation

By constructing a teacher-student model architecture and a self-supervised pseudo-label generation mechanism, and using the conflict index to identify high-conflict areas, the problem of insufficient accuracy and consistency in multi-source land cover data fusion is solved, and efficient land cover data fusion is achieved.

CN120726441BActive Publication Date: 2025-11-04JILIN AGRICULTURAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511148859.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-04
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing multi-source land cover data fusion methods struggle to balance accuracy and consistency when faced with classification inconsistencies and uncertainties, especially in high-conflict areas and ambiguous boundaries, and they are highly dependent on annotation in situations where data is scarce.

Method used

We employ a multi-source land cover data fusion method based on conflict index and self-supervised reconciliation. By constructing a teacher-student model architecture, we use the conflict index to identify high-conflict areas and combine it with a self-supervised pseudo-label generation mechanism to improve the model's classification ability in high-conflict areas, thereby enhancing spatial consistency and robustness.

Benefits of technology

It improves the accuracy and consistency of land cover data, reduces the dependence on labeled data, and enhances the mIoU, OA, and kappa coefficients of the classification results, achieving a fusion effect of regional perception and adaptation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726441B_ABST
    Figure CN120726441B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of remote sensing land cover mapping, and is especially a multi-source land cover data fusion method based on conflict index and self-supervised reconciliation. The method comprises the following steps: step one, data preparation and data preprocessing; step two, feature construction; step three, patch data construction and enhancement; step four, model construction and training strategy; step five, whole map reasoning and fusion result output. The application can improve the precision of land cover data, enhance the spatial consistency and the discrimination and repair ability of high conflict areas, realize the regional perception and adaptability of the fusion model, reduce the dependence on a large amount of labeled data, and is based on the fusion idea of deep learning+conflict perception+self-supervised optimization, and a teacher-student double model architecture is constructed to guide the model to seek balance between global consistency and local difference.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing land cover mapping, in particular to a multi-source land cover data fusion method based on conflict index and self-supervised reconciliation. BACKGROUND

[0002] Land cover refers to the physical / biological cover types formed naturally or artificially on the earth's surface, such as forests, farmland, water bodies, cities, etc. Land cover data is one of the basic data sets in geographic information science, which can accurately reflect the type distribution and spatio-temporal variation characteristics of the earth's surface, and is widely used in many fields such as water resource evolution analysis, geological disaster monitoring, land degradation and desertification assessment, environmental quality assessment, and ecological security pattern research.

[0003] With the development of remote sensing technology, more and more land cover products have been released. In the face of the coexistence of multi-source land cover products, single land cover data often cannot balance accuracy and universality. The accuracy of data can be improved through fusion methods. Existing methods developed for multi-source land cover data fusion mainly include rule decision method, probability modeling method and machine learning method.

[0004] A commonly used method is the majority voting method, which takes the class with the highest frequency in multiple data as the final result. However, this method is highly sensitive to weights, and for areas with large errors, it is easy to amplify the influence of abnormal values.

[0005] To address the inconsistency and uncertainty between multi-source data, the Dempster-Shafer evidence theory method is applied to land cover data fusion. However, when facing serious conflicts in the classification of multiple land cover data, the traditional Dempster-Shafer evidence theory method is easily affected by the design of the belief function and is unstable, requiring more prior knowledge to design rules and models, which may be difficult to achieve in the case of data scarcity or lack of domain knowledge.

[0006] The multi-source land use and land cover data fusion method considering spatial correlation considers spatial correlation, but does not consider the processing of classification conflict areas of multi-source data. For areas with low consistency, improved Dempster-Shafer evidence theory is used for fusion processing, which determines the land cover type by integrating the reliability of multiple land cover products, effectively processing uncertain areas and improving the accuracy of classification. However, this consistency analysis only focuses on surface consistency, ignoring the spatial context consistency and fuzzy boundaries between pixels.

[0007] Therefore, we propose a multi-source land cover data fusion method based on conflict index and self-supervised reconciliation to solve the above problems. SUMMARY

[0008] This section is intended to summarize some aspects of the embodiments of the present application and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of the specification of the present application in order to avoid obscuring the purpose of this section, the abstract and the title of the specification. Such simplifications or omissions cannot be used to limit the scope of the present application.

[0009] To solve the above technical problems, according to one aspect of the present application, the present application provides the following technical solutions:

[0010] A multi-source land cover data fusion method based on conflict index and self-supervised reconciliation, comprising the following steps:

[0011] Step 1: Data preparation and data preprocessing:

[0012] Download land cover data and multiple geographic environment factors, reclassify and process the source land cover data;

[0013] Step 2: Feature construction:

[0014] One-hot encoding is performed on each category of land cover data, each stack of geographic environment factors is stacked into a channel, a conflict index is constructed and used as a separate channel;

[0015] Step 3: patch data construction and enhancement:

[0016] The study area is extracted by sliding window, and each patch contains multiple channel input features and a corresponding label map;

[0017] Step 4: Model construction and training strategy:

[0018] A pair of deep learning semantic segmentation models with the same structure and complementary functions, teacher model and student model, are constructed;

[0019] The teacher model includes one-hot encoding of multiple land cover data categories, multiple normalized geographic environment factors, and a conflict index;

[0020] Self-supervised uncertainty guidance module:

[0021] The output results of the teacher model and the conflict index are used to automatically generate high-confidence pseudo labels;

[0022] The training data for the student model includes raw patches from real labels and high-confidence pseudo-label patches selected based on the teacher's softmax output and conflict index.

[0023] Step 5: Output of whole-image reasoning and fusion results:

[0024] The reasoning results of teachers and students are merged and saved.

[0025] As a preferred embodiment of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation described in this invention, in step two, the function formula of the conflict index is:

[0026] ;

[0027] in, Represents a pixel The conflict index value reflects the degree of inconsistency in classification results across multiple source layers. Represents a pixel The mode frequency, Represents a pixel Consistency with the classification of its neighboring pixels, For pixels The entropy of classification information.

[0028] As a preferred embodiment of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation described in this invention, wherein the... The formula is:

[0029] ;

[0030] in, This represents the number of times the cell appears in the category that appears most frequently across all input layers. The mode frequency represents the total number of data points. The higher the mode frequency, the more consistent the multi-source classification results are at that position, and the lower the degree of conflict.

[0031] As a preferred embodiment of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation described in this invention, wherein the... The formula is:

[0032] ;

[0033] for The neighborhood set, For indicator functions, if is 1, otherwise 0, which measures the local smoothness of the pixel in space, and if the pixel and its neighborhood are low in classification consistency, it means that it may be in the boundary or conflict transition zone.

[0034] As a preferred scheme of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation provided by the application, wherein the teacher model training data are all from visual interpretation real labels, and a cross-entropy loss function with class weight is used for optimization, and the loss function is defined as: The formula is:

[0035] ;

[0036] Among them, represents the probability that the pixel is marked as k in multiple input layers, k is the total number of categories, and the information entropy reflects the distribution disorder degree of the classification result at the position, and the higher the entropy, the more uncertain and unstable the classification of the pixel is.

[0037] As a preferred scheme of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation provided by the application, wherein the teacher model training data are all from visual interpretation real labels, and a cross-entropy loss function with class weight is used for optimization, and the loss function is defined as:

[0038] ;

[0039] Among them, represents the total number of categories; represents 1 when the label is class c, otherwise 0; represents the prediction probability of the model for class c ; represents the weight of class c ; represents the overall loss function, which is weighted and summed according to the difference in the number of class samples, and the class weight is obtained by normalizing the reciprocal of the frequency of each class pixel.

[0040] As a preferred scheme of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation provided by the application, wherein the teacher model and the student model both use U-Net network as the basic framework, which is composed of symmetrical encoder and decoder, multi-channel input is constructed by introducing multi-source land cover data and geographical environmental factors, and differential design is made in training strategy and data source.

[0041] As a preferred scheme of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation, the teacher model relies on real labels to provide stable and reliable full map classification results, and is suitable for processing regions with strong consistency; the student model has stronger adaptability and repairability on the basis of introducing pseudo labels, and can improve the classification performance in regions with significant conflicts and fuzzy boundaries, and finally the advantages of the two are fused through the conflict perception guided decision mechanism, realizing a land cover fusion optimization strategy with stability and flexibility.

[0042] As a preferred scheme of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation, the first step of the self-supervised uncertainty guidance module is Softmax entropy calculation, the information entropy is calculated pixel by pixel through the softmax probability map output by the teacher model, so as to measure the prediction uncertainty of the model at each pixel point, and the higher the entropy value is, the lower the classification confidence of the model at the pixel is; on the contrary, the low-entropy pixel can be regarded as a high-confidence prediction result.

[0043] The second step is to read the conflict index, which is calculated from the inconsistency of the classification results of multi-source land cover products, and can reflect the conflict degree of different data sources at a certain pixel. By reading the conflict index map and jointly analyzing the entropy value, the high-conflict, high-uncertainty region and the ordinary region can be distinguished, providing a basis for subsequent pseudo label screening.

[0044] The third step is to apply double thresholds, which sets two screening standards of entropy value threshold and conflict index threshold.

[0045] The fourth step is to screen the pseudo label, and the class with the highest probability is selected as the pseudo label from the softmax output of the teacher model for the pixel meeting the condition.

[0046] The last step is to mix with the real label, in order to maintain the diversity and stability of the training data, the generated pseudo label patch and the real label patch obtained by visual interpretation are mixed in a ratio of 1:1 to construct the training set of the student model.

[0047] Compared with the prior art, the present application has the beneficial effects that: ①improve the accuracy of land cover data, by introducing the teacher-student model cooperative mechanism and the conflict index driven partition decision strategy, the teacher model maintains high accuracy in the low conflict stable area, and the student model focuses on repairing the high conflict and fuzzy boundary area, improving the precision indicators such as mIoU, OA and kappa coefficient of the overall classification.

[0048] ②Enhance the spatial consistency and the discrimination and repair ability of high conflict area, the method of the application accurately identifies the high conflict area through the construction of conflict index and guides the student model to focus on the repair learning of these areas. The network architecture strengthens the identification and optimization ability of the model to the inconsistent area of different source data classification, effectively improves the classification confusion of the classification result in the conflict area, and improves the consistency and semantic stability of the fused data.

[0049] ③Realize the region perception and adaptability of the fusion model, unlike the traditional method of unified classification decision of the whole image, the application can make adaptive decision according to the conflict index of each pixel, realize the pixel-level fusion control, and the fusion mechanism significantly improves the flexibility and generalization ability of the model.

[0050] ④Reduce the dependence on a large amount of labeled data, generate pseudo-labels with the help of the output of the teacher model and the double constraint screening mechanism to construct the student model training set, effectively expand the training sample size, and still achieve good effect in the case of data annotation scarcity, have the advantages of self-supervised learning, save manpower resources and time cost.

[0051] ⑤The method has strong universality and is easy to expand, the method is suitable for land cover product fusion scenes of multiple spatial resolutions, classification systems and data sources, and the input of geographical environmental factors has good expansibility, and multiple geographical environmental factors can be input into the model as needed. The overall framework has modular design and standardized process, and is convenient for migration to other regions or fusion tasks of other data. BRIEF DESCRIPTION OF DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the application, the application will be described in detail below in combination with the drawings and detailed embodiments. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor. Among them:

[0053] Figure 1 It is a flow chart of the application of a multi-source land cover data fusion method based on conflict index and self-supervised reconciliation;

[0054] Figure 2 It is a model architecture diagram of the application of a multi-source land cover data fusion method based on conflict index and self-supervised reconciliation;

[0055] Figure 3 It is a visualization diagram of the conflict index (CI) in the research area in embodiment 1 of the application of a multi-source land cover data fusion method based on conflict index and self-supervised reconciliation;

[0056] Figure 4 A visualization chart of the research area in the conflict index index - information entropy (Entropy) of Embodiment 1 of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation of the application;

[0057] Figure 5 A visualization chart of the research area in the conflict index index - Model Frequency (Model Frequency) of Embodiment 1 of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation of the application;

[0058] Figure 6 A visualization chart of the research area in the conflict index index - Local Consistency (Local Consistency) of Embodiment 1 of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation of the application;

[0059] Figure 7 A comparison chart of the same position of the real object and the source land cover map and the fusion land cover map in Embodiment 1 of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation of the application;

[0060] Figure 8 ESRI 2020 Land Cover land cover map used in Embodiment 1 of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation of the application;

[0061] Figure 9 MODIS MCD12Q1 LCCS2 land cover map used in Embodiment 1 of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation of the application;

[0062] Figure 10 MODIS MCD12Q1 IGBP land cover map used in Embodiment 1 of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation of the application;

[0063] Figure 11 ESA CCI-LC land cover map used in Embodiment 1 of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation of the application;

[0064] Figure 12 GLC-FCS30 land cover map used in Embodiment 1 of the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation of the application;

[0065] Figure 13The land cover map fused and generated by using the traditional U-Net architecture in the land cover data fusion method based on conflict index and self-supervised reconciliation of multiple sources according to the embodiment 1 of the application;

[0066] Figure 14 The land cover map fused and generated by using the self-repairing U-Net network in the land cover data fusion method based on conflict index and self-supervised reconciliation of multiple sources according to the embodiment 1 of the application. DETAILED DESCRIPTION

[0067] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the specific embodiments of the application will be described in detail below with reference to the accompanying drawings.

[0068] Secondly, the application is described in detail in combination with the schematic diagram. In the detailed description of the embodiments of the application, the cross-sectional view of the device structure is locally enlarged without the general proportion for the convenience of description, and the schematic diagram is only an example, which should not limit the scope of protection of the application here. In addition, the three-dimensional spatial dimensions of length, width and depth should be included in the actual manufacture.

[0069] In order to make the purposes, technical solutions and advantages of the application more clear, the embodiments of the application will be further described in detail below with reference to the accompanying drawings.

[0070] The application aims to solve the classification conflict problem caused by the inconsistency of classification systems, different spatial resolutions, different classification methods and other problems in the fusion of multiple land cover data, so as to construct land cover data with higher precision. The working principle is based on the fusion idea of deep learning + conflict perception + self-supervised optimization. By constructing a teacher-student double model architecture, the model is guided to seek balance between global consistency and local difference.

[0071] A land cover data fusion method based on conflict index and self-supervised reconciliation of multiple sources is provided to solve some problems in the existing method and its discrimination and processing of classification inconsistent regions between different data. The core idea is to identify the classification inconsistent region by constructing a conflict index, combine the uncertainty estimation of model prediction, guide the construction of self-supervised pseudo label, realize the key repair of high conflict area and the fusion optimization of overall classification result. This method not only enhances the fusion expression ability of the model to multiple source data, but also effectively improves the classification accuracy and spatial consistency of high conflict area.

[0072] To identify high conflict areas, the invention constructs a quantifiable indicator of inconsistency in the classification of multiple land cover data - conflict index CI, which comprehensively considers the category mode frequency, information entropy and local consistency, accurately measures the classification differences between multiple source data, and provides data support for the identification of high uncertainty classification areas and the adjustment of fusion strategy. In order to ensure that the influence of the three is balanced when evaluating the conflict degree of multi-source classification results, an equal weight allocation strategy is adopted, that is, the weight of each indicator is 1 / 3. This allocation method not only makes the contribution of each dimension feature equal, but also meets a basic constraint condition: the sum of the weights is equal to 1. This average weighting method helps to comprehensively reflect the classification conflict of the pixel from the three dimensions of category consistency, category confusion degree and spatial continuity, so as to construct a more objective and stable conflict index. The function formula of the conflict index is:

[0073]

[0074] Wherein, represents the conflict index value of the pixel , reflecting the inconsistency degree of its classification results in multiple source layers. The formula contains the following three sub-formulas:

[0075]

[0076] Wherein, represents the mode frequency of the pixel , is the number of occurrences of the most common category of the pixel in all input layers, is the total number of data. The higher the mode frequency, the more consistent the multi-source classification results at this position, and the lower the conflict degree.

[0077]

[0078] Wherein, represents the classification consistency of the pixel and the pixels in its neighborhood; is the neighborhood set (3x3 neighborhood window) of , is an indicator function (1 if , otherwise 0). This term is used to measure the local smoothness of the pixel in space. If the classification consistency of the pixel and its neighborhood is low, it means that it may be in the boundary or conflict transition zone.

[0079]

[0080] Wherein, is the classification information entropy of the pixel , indicates that the pixel is labeled as k the probability of the class, k is the total number of classes. The information entropy reflects the degree of distribution disorder of the classification results of the position, and the higher the entropy, the more uncertain and unstable the classification of the pixel is.

[0081] The conflict index is an index for measuring the inconsistency of the classification results of multi-source land cover data, which can quantify the classification conflict degree of each pixel in different products. The higher the CI value, the greater the classification difference of the pixel in multiple data and the stronger the uncertainty. The conflict index is not only used to find the problem area, but also guides the subsequent classification optimization and model reconstruction task, which is the core scheduling mechanism in the present application.

[0082] The present application constructs an uncertainty perception type fusion model - self-repairing U-Net network which is fused with the self-supervised learning idea based on the core architecture of U-Net. The model adopts a teacher-student collaborative mechanism and is composed of three main modules: a teacher model (Teacher U-Net), a self-supervised uncertainty guiding module (Uncertainty-Guided Module) and a student model (Student U-Net). The three cooperate with each other to improve the classification ability of the model in the high conflict area and enhance the overall spatial consistency and robustness.

[0083] The deep learning model constructed by the application adopts a classic U-Net encoder-decoder structure, and the network architecture is consistent in the teacher model (Teacher) and the student model (Student), but there are differences in input sources and training strategies. In the encoder part, the model successively passes through four down-sampling stages, each stage containing two 3*3 convolution layers (Conv2d + BatchNorm + ReLU) and a 2*2 maximum pooling operation (MaxPooling). The channel number is expanded from 64, 128, 256 to 512 after each down-sampling, and the feature map size is gradually reduced to one fourth of the original input. The bottleneck part in the middle uses two 1024-channel convolution layers for deep feature extraction. In the decoder part, each stage restores the feature map to a larger spatial resolution through 2*2 up-sampling (transposed convolution), and splices (skip connection) with the feature map of the corresponding encoding stage, followed by two convolution layers for fusion. The channel number of the decoding stage is reduced from 512, 256, 128 to 64, and finally an 1*1 convolution is used to map the output to the prediction map of the target class. The teacher model is trained using real labels, responsible for generating high-quality softmax classification probability maps; the student model is trained by combining real label patches and pseudo label patches selected based on conflict index and softmax entropy, and the training strategy (loss function, early stopping, learning rate scheduling, etc.) is consistent with the teacher model. The output is still the softmax classification map consistent with the teacher model.

[0084] The self-supervised uncertainty guidance module is located between the teacher and student models, mainly used to generate pseudo labels and construct the training set of the student model. The module first calculates the information entropy based on the softmax probability map output by the teacher model to measure the classification uncertainty of each pixel; then reads the conflict index map to comprehensively judge whether the pixel belongs to an unreliable area. The module selects high-confidence pixels as pseudo labels by setting double thresholds (entropy and conflict index). Then the pseudo label patches and real label patches are mixed in a 1:1 ratio to form the training data set of the student model.

[0085] The model training part is optimized by using a cross-entropy loss function. Considering the problem of unbalanced class proportion in land cover data, a weighted cross-entropy loss function (Weighted Cross Entropy Loss) is used to give different weights to different classes. The calculation method of class weight is the reciprocal of the frequency of each class pixel and is normalized, and the weight of the less class is higher, so as to prevent the model from being biased to the main class. The out-of-area or unlabeled pixels are set as class 0, and ignore_index = 0 is used in the loss function to exclude this part of the area, so as to avoid affecting the training.

[0086] As for the optimizer, the Adam optimizer is selected to speed up the convergence, and the initial learning rate is set to 0.001, and the weight decay parameter 1e-5 is set to suppress overfitting. In order to improve the training stability, the learning rate scheduling strategy StepLR is introduced, and the learning rate is automatically decayed to 0.1 times of the original every 10 epochs. In the training process, the EarlyStopping strategy is introduced, and the validation set mIoU is set to continuously improve for 15 epochs without improvement, and the training is terminated in advance to prevent the model from overfitting to the training set. The batch size is set to 8, and a maximum of 100 epochs is trained each time. In most cases, the training can be terminated in advance within 30-50 epochs.

[0087] Finally, the teacher model provides stable full-image prediction results, while the student model focuses on repairing and optimizing the conflict significant area. In the fusion stage, the spatial distribution of the conflict index map is used to realize regional perception decision: the student model output is used in the high conflict area, and the teacher model result is retained in the remaining area, so as to realize the output of high-quality land cover classification map with stability and flexibility. The overall parameter size of the model is reasonable, which not only ensures strong expression ability, but also can run efficiently under acceptable computing resources.

[0088] As Figures 1-2 shown, the multi-source land cover data fusion method based on conflict index and self-supervised reconciliation provided by the application comprises the following steps:

[0089] Step 1: Data preparation and data preprocessing

[0090] Download the land cover data and multiple geographic environmental factors that need to be fused.

[0091] Use ArcGIS Pro software to unify the spatial resolution of all data through resampling, and also unify the projection coordinate system of all data.

[0092] According to the research purpose, the target classification needed is selected.

[0093] A classification mapping table between the classification system of the multiple source land cover products and the target classification system is constructed according to the semantic affinity score. All source land cover data are reclassified to realize the unification of the classification system using ArcGIS Pro software.

[0094] A certain number of random points in the study area are generated as validation point data, and the land cover type of each point is obtained through visual interpretation of each point by high-definition satellite remote sensing images.

[0095] The study area data and the interpreted validation point data are used as the basis for effective pixel screening and accuracy evaluation, and after rasterization, spatial resolution and projection coordinate system are also unified.

[0096] All geographic environmental factor data are normalized.

[0097] After all data processing is completed, it is converted into Numpy format.

[0098] Step two: feature construction

[0099] Each category of land cover data is one-hot encoded, and each category of land cover data is treated as a channel. The number of corresponding channels is the product of the number of land cover data and the number of target categories.

[0100] Each geographic environmental factor is stacked into a channel.

[0101] A conflict index is constructed and used as a separate channel.

[0102] Step three: patch data construction and enhancement

[0103] The study area is extracted by sliding window (patch size = 256, stride = 128). Each patch contains multiple channel input features and a corresponding label map. The label comes from the visual interpretation of the validation point. The area outside or without label pixels is set to 0 and ignored in training as ignore index. To improve the generalization ability of the model, data augmentation is also performed, including horizontal flip, vertical flip and 90° rotation. Each original patch generates 3 augmented samples.

[0104] Step four: model construction and training strategy

[0105] The application constructs a pair of deep learning semantic segmentation models with the same structure and complementary functions: a teacher model (Teacher) and a student model (Student). Both of them use U-Net network as the basic framework, which is composed of symmetrical encoder and decoder, and has good feature extraction and spatial positioning ability. By introducing multi-source land cover data and geographical environmental factors to construct multi-channel input, and differentiating the design of training strategy and data source, the discrimination and optimization of the classification conflict area are realized, and the spatial consistency of the classification result is enhanced.

[0106] Teacher Model

[0107] The teacher model is the core classifier of the first stage of the entire network architecture, which is responsible for learning the mapping relationship from the fused features to the standard land cover classification results under the supervised condition. The input of the model is a multi-channel feature map, including one-hot encoding of multiple land cover data, multiple normalized geographical environmental factors and a conflict index. The output of the model is multi-channel (target class number plus 1) prediction data, corresponding to all target valid categories and 1 invalid category (category 0 is set as ignore index, which does not participate in loss calculation).

[0108] The training data of the teacher model is all from the real label of visual interpretation, and the weighted cross entropy loss function (Weighted Cross Entropy Loss) is used for optimization, which is defined as:

[0109]

[0110] Among them, represents the total number of categories; represents that the value is 1 when the label is category c, otherwise it is 0 (i.e. the first c component of one-hot encoding); represents the prediction probability of the model for category c (softmax output); represents the weight of category c , which is usually calculated according to the inverse of the category frequency (see the formula below); represents the overall loss function, which is weighted and summed according to the difference of the number of category samples. This formula is used to solve the class imbalance problem to prevent small classes from being ignored.

[0111] The category weight is obtained by normalizing the inverse of the frequency of each category, and the specific calculation method is as follows:

[0112]

[0113] Among them, represents the frequency of class c in the training set (i.e. the proportion of the total number of pixels); represents the weight of class c; the denominator is the sum of the reciprocals of all class frequencies, used to normalize the sum of weights to 1. This formula ensures that rare classes have large weights and common classes have small weights, which is the basis of weighted cross-entropy.

[0114] Self-Supervised Uncertainty-Guided Module

[0115] This module is the key link between the teacher and student models in the teacher-student architecture, using the output results and conflict indexes of the teacher model to automatically generate high-confidence pseudo labels, thereby constructing supplementary training data to guide the student model to better learn the classification features of high-conflict areas.

[0116] The first step of the module is Compute Softmax Entropy.

[0117] Through the softmax probability map output by the teacher model, the information entropy is calculated pixel by pixel to measure the prediction uncertainty of the model at each pixel. The higher the entropy value, the lower the classification confidence of the model at that pixel; on the contrary, low-entropy pixels can be considered as high-confidence prediction results. This process helps us identify high-uncertainty areas that need to be focused on.

[0118] The second step is Read Conflict Index.

[0119] The conflict index is calculated from the inconsistency of the classification results of multiple land cover products, and can reflect the conflict degree of different data sources at a certain pixel. By reading the conflict index map, we can distinguish high-conflict, high-uncertainty areas from ordinary areas, providing a basis for subsequent pseudo label selection.

[0120] The third step is Apply Dual Thresholds.

[0121] This step sets two screening criteria for entropy value threshold and conflict index threshold. By setting dual thresholds, we can effectively eliminate low-confidence and high-conflict noise pixels, ensuring the reliability of pseudo labels.

[0122] The fourth step is Select Pseudo Labels.

[0123] For pixels that meet the conditions, the class with the highest probability from the softmax output of the teacher model is selected as the pseudo label. This way we can quickly get a set of relatively reliable label data, covering the classification deficiencies of the teacher model in complex areas.

[0124] The last step is Mix with Real Labels (1:1).

[0125] To maintain the diversity and stability of the training data, the module mixes the generated pseudo-label patches with the real label patches obtained through visual interpretation at a ratio of 1:1 to construct the training set for the student model. This mixing strategy can simultaneously utilize the high accuracy of real labels and the high coverage of pseudo labels, improving the student model's discriminative ability in high-conflict areas.

[0126] Student Model

[0127] The student model, as a supplementary classifier in the second stage, has a structure similar to the teacher model and is also a U-Net network based on multi-channel input. However, its training data consists of two parts: one part is the original patch from the real label, and the other part is the high-confidence pseudo-label patch selected based on the teacher's softmax output and conflict index. Specifically, by calculating the softmax entropy value of each pixel and setting the mask condition jointly with the conflict index, the reliable area is selected and the pseudo-label is generated to construct the self-supervised training set. In the verification stage, only real label data is used for evaluation to maintain the consistency and reliability of the evaluation index. The training strategy (loss function, early stopping, learning rate scheduling) is consistent with the teacher model.

[0128] Although the teacher model and the student model have consistent network structures based on the U-Net framework with multi-channel input, they have a division of labor in function. The teacher model relies on real labels to provide stable and reliable full-image classification results, which is suitable for processing areas with strong consistency. The student model, based on the introduction of pseudo labels, has stronger adaptability and repair ability, and can improve the classification performance in areas with significant conflicts and fuzzy boundaries. Finally, through the conflict-aware guided decision mechanism, the advantages of the two are integrated to achieve a land cover fusion optimization strategy that balances stability and flexibility.

[0129] Step Five: Whole-image Reasoning and Fusion Result Output

[0130] After the teacher model is trained, it is used for whole-image sliding window reasoning to output a softmax probability map, and the initial classification map is obtained through the argmax operation, providing a probability basis for subsequent self-supervised pseudo label generation.

[0131] After the student model is trained, the student model is used for whole-image reasoning to also generate a softmax probability map, and the classification map based on the conflict mask is obtained through the argmax operation.

[0132] Output fusion optimization, fuse the reasoning results of teachers and students: for high conflict areas above the conflict index threshold, use the classification results of the student, and the remaining areas retain the teacher output. The pixels outside the region are set to 0, and the final fusion result is saved.

[0133] Embodiment 1

[0134] See Figures 3-14 wherein Figure 7 , a1 b2 represents two high-resolution remote sensing satellite images of real objects, a2 b2 represents the land cover map fused using the self-repairing U-Net network, a3 b3 represents the land cover map fused using the traditional U-Net architecture, a4 b4 represents the land cover map of ESRI 2020 Land Cover, a5 b5 represents the land cover map of MODIS MCD12Q1 LCCS2, a6 b6 represents the land cover map of MODIS MCD12Q1 IGBP, a7 b7 represents the land cover map of ESA CCI-LC, and a8 b8 represents the land cover map of GLC-FCS30.

[0135] The Sanjiang Plain region in China and the Russian Primorsky Krai bordering the Sanjiang Plain region are selected as the experimental region. Both regions are geographically adjacent and located in the East Asian mid-high latitude monsoon climate zone, with similarities in geography and ecology, while there are differences in landform structure and land use methods, belonging to a typical ecological transition zone. This type of region has complex class boundaries, heterogeneous information overlap, and significant classification conflicts, making it an ideal sample area for testing the applicability and stability of multi-source land cover fusion methods.

[0136] Five mainstream land cover data products in 2020 are selected in this embodiment, including: ESA CCI-LC, GLC-FCS30, MODIS MCD12Q1 IGBP, MODIS MCD12Q1 LCCS2, and ESRI 2020 Land Cover. The goal of the fusion experiment is to improve the classification accuracy and spatial consistency of cultivated land, forest land, and grassland in the study area, and to generate a cultivated land, forest land, and grassland cover map for the study area in 2020. Auxiliary input data includes five geographic environmental factors: digital elevation model (DEM), slope, soil moisture, normalized vegetation index (NDVI), and land surface temperature (LST). The data is uniformly resampled to 30-meter resolution, and the projection coordinate system is kept consistent (WGS84 projection coordinate system). All land cover products are reclassified into four categories after semantic comparison and expert revision: cultivated land, forest land, grassland, and others. All geographic environmental factors are normalized.

[0137] 3686 points were evenly distributed in the study area as validation point data, and Google high-resolution remote sensing images were used to visually interpret each point. After interpretation, the validation point data was converted to raster format, and the projection and resolution were unified.

[0138] All data were then converted to NumPy format, with five land cover data encoded as one-hot (4 channels for each product, a total of 20 channels), geographical factors stacked by quantity as 5 channels, and conflict index constructed as the 26th channel. The final input feature tensor is 26 channels, with a size of (H, W, 26).

[0139] The conflict index is constructed based on three sub-indices: majority class frequency, local consistency, and information entropy, which measure the consistency of multi-source classification, neighborhood structure consistency, and classification uncertainty, respectively. The three sub-indices are assigned equal weights, and the sum of the weights is 1. The constructed conflict index map has a value range of [0, 3.46875], which is divided into eight intervals according to the numerical value: 0.0, (0.0, 0.04444], (0.04444, 0.08889], (0.08889, 0.13333], (0.13333, 0.17778], (0.17778, 0.6], (0.6, 1.0], >1.0, where the area with pixel value >1 is defined as the high conflict area.

[0140] Patch segmentation is then performed using a sliding window strategy (patch size = 256, stride = 128) to slide the study area and extract the original patch. Data augmentation operations (90° rotation, horizontal flip, vertical flip) are performed on each patch, and each original patch is augmented to 4 samples to improve the model's generalization ability.

[0141] This study uses a teacher-student model architecture, both based on the standard U-Net network structure, with encoder-decoder symmetric structure, skip connection mechanism, and adaptation to multi-channel input. In the first stage, the teacher model is supervised trained with real visual interpretation labels, and the optimization target is the weighted cross-entropy loss function, with class weights normalized by the inverse of the sample frequency. EarlyStopping (monitoring validation set mIoU, patience = 15) and StepLR learning rate decay strategy are used in training, with an initial learning rate of 0.001.

[0142] The teacher model outputs a softmax probability map after inference, which is used for pseudo-label construction of the second-stage student model. The pseudo-label screening adopts a joint threshold strategy: pixels with softmax entropy < 0.8 and conflict index < 1 are considered as high-confidence pseudo-labels and participate in the training of the student model. The pseudo-label patches and the real label patches are mixed in a 1:1 ratio to form the training set of the student model.

[0143] The student model has the same structure as the teacher model, but the training data contains pseudo-label regions, so it has stronger boundary recognition and region repair capabilities. The training strategy is consistent with that of the teacher model.

[0144] In the final fusion stage, a conflict-aware guided decision mechanism is adopted: the prediction results of the student model are used for regions with a conflict index greater than 1, and the remaining regions retain the output of the teacher model. The final land cover map after fusion is saved, and the pixels outside the region are set to invalid values (0).

[0145]

[0146] Although the present application has been described with reference to the embodiments above, various improvements can be made thereto and components thereof can be substituted with equivalents without departing from the scope of the present application. In particular, features in the embodiments disclosed by the present application can be combined with each other in any manner as long as there is no structural conflict, and the combinations are not exhaustively described in the present specification only for the purpose of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for fusing multi-source land cover data based on conflict index and self-supervised reconciliation, characterized in that, Includes the following steps: Step 1: Data Preparation and Preprocessing Download land cover data and multiple geographic environmental factors, and reclassify and process the source land cover data; Step 2: Feature Construction One-hot encoding is performed on each category of land cover data. Each geographic environmental factor is stacked into a channel, and a conflict index is constructed as a separate channel. The function formula for the conflict index is as follows: ; in, Represents a pixel The conflict index value reflects the degree of inconsistency in classification results across multiple source layers. Represents a pixel The mode frequency, Represents a pixel Consistency with the classification of its neighboring pixels, For pixels The entropy of classification information; Step 3: Patch Data Construction and Enhancement: The study area is extracted using a sliding window method, with each patch containing multiple channel input features and a corresponding label map. Step 4: Model Building and Training Strategies Construct a pair of deep learning semantic segmentation models with identical structures and complementary functions: a teacher model and a student model; The teacher model includes one-hot encoding of multiple land cover data categories, multiple normalized geographic environmental factors, and a conflict index; Self-supervised uncertainty guidance module: The output of the teacher model and the conflict index are used to automatically generate pseudo-labels with high confidence. The training data for the student model includes raw patches from real labels and high-confidence pseudo-label patches selected based on the teacher's softmax output and conflict index. Step 5: Output of whole-image reasoning and fusion results: The reasoning results of teachers and students are merged and saved.

2. The multi-source land cover data fusion method based on conflict index and self-supervised reconciliation as described in claim 1, characterized in that, The The formula is: ; in, This represents the number of times the cell appears in the category that appears most frequently across all input layers. The mode frequency represents the total number of data points. The higher the mode frequency, the more consistent the multi-source classification results are at that position, and the lower the degree of conflict.

3. The multi-source land cover data fusion method based on conflict index and self-supervised reconciliation according to claim 1, characterized in that, The The formula is: ; for The neighborhood set, For indicator functions, if If the value is 1, it is 1; otherwise, it is 0. This item is used to measure the local smoothness of a cell in space. If the cell has low consistency with its neighborhood classification, it indicates that it may be located in a boundary or conflict transition zone.

4. The multi-source land cover data fusion method based on conflict index and self-supervised reconciliation as described in claim 1, characterized in that, The The formula is: ; in, This indicates that the cell is marked as [label] in multiple input layers. k The probability of the category, k The total number of categories is represented by the information entropy, which reflects the degree of disorder in the distribution of classification results at that location. The higher the entropy, the more uncertain and unstable the classification of that pixel is.

5. The multi-source land cover data fusion method based on conflict index and self-supervised reconciliation as described in claim 1, characterized in that, The teacher model training data all comes from real labels in visual interpretation, and is optimized using a cross-entropy loss function with class weights. This loss function is defined as: ; in, Indicates the total number of categories; The value is 1 if the label is category c, and 0 otherwise. The model represents the categories c The predicted probability; Indicates category c The weight, The overall loss function is a weighted sum based on the differences in the number of samples in each category. The category weights are obtained by normalizing the inverse of the frequency of each type of pixel.

6. The multi-source land cover data fusion method based on conflict index and self-supervised reconciliation according to claim 1, characterized in that, Both the teacher and student models use the U-Net network as the basic framework, consisting of a symmetrical encoder and decoder. Multi-channel inputs are constructed by introducing multi-source land cover data and geographical environmental factors, and the training strategies and data sources are designed differently.

7. The multi-source land cover data fusion method based on conflict index and self-supervised reconciliation according to claim 1, characterized in that, The teacher model, relying on real labels, provides stable and reliable full-map classification results, suitable for processing areas with strong consistency. The student model, based on the introduction of pseudo-labels, has stronger adaptability and repair capabilities, and can improve classification performance in areas with significant conflicts and blurred boundaries. Finally, through a conflict perception-guided decision-making mechanism, the advantages of both models are integrated to achieve a land cover fusion optimization strategy that balances stability and flexibility.

8. The multi-source land cover data fusion method based on conflict index and self-supervised reconciliation according to claim 1, characterized in that, The first step of the self-supervised uncertainty guidance module is the calculation of Softmax entropy. The information entropy is calculated pixel by pixel using the softmax probability map output by the teacher model. This measures the prediction uncertainty of the model at each pixel. The higher the entropy value, the lower the classification confidence of the model at that pixel. Conversely, low-entropy pixels can be regarded as high-confidence prediction results. The second step is to read the conflict index. The conflict index is calculated from the inconsistency of the classification results of multi-source land cover products. It can reflect the degree of conflict of different data sources on a certain pixel. By reading the conflict index map and analyzing it together with the entropy value, high conflict and high uncertainty areas can be distinguished from ordinary areas, providing a basis for subsequent pseudo-label screening. The third step is to apply dual thresholds, which sets two screening criteria: an entropy threshold and a conflict index threshold. The fourth step is to filter pseudo-labels. For pixels that meet the criteria, the category with the highest probability is directly selected from the softmax output of the teacher model as the pseudo-label. The final step is to mix the generated pseudo-label patches with the real labels obtained from visual interpretation in a 1:1 ratio to maintain the diversity and stability of the training data, thus constructing the training set for the student model.

Citation Information

Patent Citations

  • Urban mining waste land utilization optimization method and system based on cooperation of space conflict and ecological barrier

    CN119168305A

  • Type identification, causal analysis and simulation coping method for territorial space conflicts

    CN119939420A