Map-image semantic change detection method and system based on multitask-single label learning

By using an optimized framework of multi-task-single-label learning, and by mapping semantic category probabilities to a binary change probability space using a distribution transformation function, the problem of high training complexity of semantic change detection models in the remote sensing field is solved, and efficient semantic change detection is achieved.

CN120808138APending Publication Date: 2025-10-17WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510841036.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing semantic change detection methods in the remote sensing field suffer from high model training complexity due to the difficulty in obtaining labels, requiring a lot of time, manpower and financial resources, and the multi-task framework is heavily dependent on pixel-level labels.

Method used

We design an optimized framework for multi-task-single-label learning. By extracting features from maps and images, we use a distribution transformation function to map semantic category probabilities to a binary change probability space, thereby achieving joint learning of change localization and semantic segmentation and reducing the dependence on multi-task supervised labels.

Benefits of technology

It effectively reduces the training cost of semantic change detection models, improves model performance, enables efficient learning with only semantic or change labels, and reduces annotation difficulty and data collection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808138A_ABST
    Figure CN120808138A_ABST
Patent Text Reader

Abstract

The invention discloses a map-image semantic change detection method and system based on multitask-single label learning. According to the method, efficient learning of semantic change detection of a previous land cover map and a newest optical image is realized through semantic change detection and distribution transformation. The method comprises the following steps: firstly, extracting a change region between map data and an optical image and a semantic category of image pixels by using a multi-task model, and then converting a semantic category probability into a binary change probability through a distribution transformation function; through the design, on one hand, the latest and best semantic change detection model can be integrated; and on the other hand, learning of the semantic change detection model can be efficiently driven only by using a binary change tag or a semantic segmentation tag, and finally, high-performance and low-cost semantic change detection is realized. According to the method, modeling of map-image pair change area positioning and pixel category identification can be realized by virtue of an existing optical remote sensing image semantic change detection method based on mature research.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of computer vision, and particularly relates to a map-image semantic change detection method and system based on multi-task-single-label learning. BACKGROUND

[0002] Semantic change detection is an important intelligent interpretation task in the field of remote sensing, aiming to analyze the change information of the earth's surface from double-time series of earth observation data, and to provide information support for urban planning, natural resource monitoring and disaster assessment. The current research work mainly extracts the changed area and type of the earth's surface by comparing the historical remote sensing image with the latest image. Due to cloud cover, incomplete file records and sensor differences, the early image information of the earth's surface faces the "three no" dilemma of unavailability, unavailability and incomparability, which hinders the semantic change detection. In the long history of earth observation, people have collected global land cover information through field investigation and image interpretation and drawn it into a land cover map. With the deepening of the idea of cooperation and sharing and the development of Internet technology, map data has become readily available and well maintained. Therefore, using easily accessible map data as a reference object for the latest earth image can solve the problem of historical data gaps, facilitating the analysis of the dynamic of the earth's surface. However, the inter-class diversity of ground objects significantly enriches the semantic change pattern, increasing the difficulty of extracting the change information of the earth's surface.

[0003] In order to solve this problem, the existing semantic change detection method usually decomposes semantic change detection into two sub-tasks of change localization and semantic segmentation, and then jointly learns them by means of multi-task framework. A large number of research work has shown that the strategy of task decoupling and joint learning can effectively improve the accuracy of semantic change detection. However, the training of multi-task framework requires pixel-level change localization and semantic segmentation labels. In the field of remote sensing, it requires more time, manpower and financial resources to obtain semantic dense labels, which requires the annotators to have professional knowledge of remote sensing and frequent annotation operations. Therefore, it is urgent to design an efficient multi-task learning framework to reduce the difficulty of obtaining high-performance semantic change detection model. SUMMARY

[0004] In order to solve the problem of high learning complexity of semantic change detection model, the present application designs a multi-task-single-label learning optimization framework. First, the multi-task model is used to extract the features of ground objects in the map and image, model the difference representation of the features of ground objects, and detect the changed area of the earth's surface. At the same time, the association between the image features of ground objects and the semantic categories is modeled to identify the class attributes of the pixels. Then, through a distribution transformation function, the semantic category probability is mapped into the binary change probability space, so as to optimize the parameters of the multi-task model with the supervision signal of the binary change label or the supervision signal of the semantic label.

[0005] The technical scheme adopted by the system of the application is: a map-image semantic change detection method based on multi-task-single-label learning, sequentially comprising the following stages: A semantic change detection stage: extracting features of ground objects in the map and the image, detecting changed ground surface areas and ground object types; A distribution transformation stage: using semantic knowledge of the map to unify the probability spaces of the binary change and the semantic categories, facilitating learning of change positioning and semantic segmentation driven by a single task label signal.

[0006] Specifically, the specific content of the semantic change detection stage comprises: An arbitrary semantic change detection model based on a multi-task learning framework is used as a benchmark model, which generally includes a double-branch feature extraction module, a change positioning module and a semantic segmentation module. In order to avoid interference between heterogeneous map and image features, the double-branch feature extraction module uses two feature extractors that are not associated with each other and do not share parameters to process map and image data respectively. The extracted map and image features are input into the change positioning module to obtain the detection result of the change area, and the image features are input into the semantic segmentation module to obtain the recognition result of the pixel semantics.

[0007] Specifically, the specific content of the distribution transformation stage comprises: The semantic knowledge of the map is used to align the probability spaces of the semantic categories and the binary change, breaking the gap between the multi-task label signals. When only the semantic category label is available, the semantic information is used to generate a binary change label to supervise the learning of the change positioning module; and when only the binary change label is available, the category probability output by the semantic segmentation module is converted into a binary change probability to facilitate the supervision of the semantic segmentation module by the binary change label signal.

[0008] Further, the double-branch feature extraction module comprises a map encoder and an image encoder; The semantic information of the map is represented by one-hot encoding Then, the map encoder is used to extract the map features of the ground objects:

[0009] The image encoder is used to process the optical image data of the ground surface , extract the image features of the ground objects:

[0010] The map encoder and the image encoder do not share model parameters with each other and independently extract map and image information, ensuring the discriminability of the ground object features.

[0011] Further, the change positioning module, i.e. a change decoder Constructing the time sequence information of the map and image data, generating the difference feature representation of the feature changes and predicting the change area:

[0012] wherein, respectively represent the map feature and the image feature, represent the change probability of the map-image pair.

[0013] Further, through the semantic segmentation module, i.e. the semantic decoder Classify the visual features of the feature, and obtain the pixel-level class attribute:

[0014] wherein, represent the image feature, represent the semantic probability of the image.

[0015] Further, in the distribution transformation stage, first, a derivable distribution transformation process is constructed, which is related to the semantic class and the probability space of binary changes; given a class probability matrix before the change and a class probability matrix after the change , wherein each probability vector and obeys the class distribution; in the map-image data pair, because the semantic information before the change is known, each probability vector is in discrete, one-hot form, i.e. and ; in order to generate the binary change probability, the argmax operation is used to discretize the probability distribution , generate one-hot encoding, and then compare it with ; in view of the feature class before the change is determined, i.e. , then the change probability of the corresponding position is equivalent to the probability that the corresponding pixel does not belong to the class , i.e. , b=1 represents that the feature change occurs at the corresponding position; similarly, the non-change probability at the position is equivalent to the probability that the corresponding pixel does not belong to the class , i.e. , b=0 represents that the feature change does not occur at the corresponding position; therefore, the distribution transformation function about the semantic probability is defined as:

[0016] wherein, represent the number of feature classes, The distribution transformation function converts the loose semantic distribution by multiplication operation, making it derivable.

[0017] Further, when only semantic labels are available for model training, the semantic labels are converted into category probability matrix using one-hot encoding operation and pseudo change labels are generated by distribution transformation function The loss function of the learning for supervised change localization task is as follows:

[0018] wherein, denotes the change probability of the map-image pair, denotes the semantic probability of the image, and denote the loss functions of change localization and semantic segmentation respectively, both of which are realized by cross-entropy loss function.

[0019] Further, when only change labels are available for model training, the predicted semantic probability is converted into binary change probability by distribution transformation function, so as to be supervised by change labels; a local information entropy penalty term is designed to increase the certainty of semantic probability by minimizing the local information entropy; the definition of the local information entropy function is as follows:

[0020]

[0021] wherein, H and W respectively denote height and width, K denotes the category of the ground object, denotes the change label, 0 denotes that the ground object has no change, and 1 denotes that the ground object has change.

[0022] Further, the final loss function of the trained model is as follows:

[0023] wherein, denotes the change probability of the map-image pair, denotes the semantic probability of the image, and denote the loss functions of change localization and semantic segmentation respectively, both of which are realized by cross-entropy loss function.

[0024] Further, it also includes the evaluation index SCS to compare the semantic change detection performance under different training label conditions.

[0025] The application further provides a map-image semantic change detection system based on multi-task-single-label learning, comprising: A processor and a memory, the memory is used for storing program instructions, and the processor is used for calling the stored instructions in the memory to execute the map-image semantic change detection method based on multi-task-single-label learning.

[0026] In summary, the map-image semantic change detection method based on multi-task-single-label learning can realize cross-task supervision of heterogeneous labels by utilizing the causal relationship between change positioning and semantic segmentation labels, and facilitate the acquisition of high-performance semantic change detection models. The application can reduce the dependence of the semantic change detection model based on the multi-task learning framework on multi-task supervision labels and reduce the training cost of the semantic change detection model. The final experimental verification also proves that the method described in the application exceeds the performance of the post-classification strategy based on SSG2 under the same number of labels. Therefore, the application can be effectively applied to the map-image semantic change detection task. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 The map-image semantic change detection method based on multi-task-single-label learning of the embodiment of the application is shown in the flowchart. DETAILED DESCRIPTION

[0028] In order to facilitate those skilled in the art to understand and implement the application, the application will be further described in detail below in combination with the drawings and examples. It should be understood that the implementation examples described herein are only used to illustrate and explain the application, and are not used to limit the application.

[0029] As shown in Figure 1 The application is directed to the map-image semantic change detection task, and proposes an efficient multi-task learning framework. The method mainly includes two steps, which are the semantic change detection phase and the distribution transformation phase.

[0030] First, in the semantic change detection phase, as Figure 1As shown in the left half of FIG. 1, the proposed method uses a semantic change detection model based on a multi-task learning framework (generally including a dual-branch feature extraction network, a change decoder, and a semantic decoder, etc.) to extract the change information of the map-image pair. Because the proposed method aligns the probability distribution of the change localization and semantic segmentation results in the output space of the model, it can be compatible with any semantic change detection model based on a multi-task learning framework. In this embodiment, TBFFNet is used as the semantic change detection model of the present application. The dual-branch feature extraction network of TBFFNet uniformly uses ResNet34 as the encoder of the two branches and shares the parameters of the two encoders. In order to adapt to the map-image-based semantic change detection, the present application uses the encoder of the time series 1 branch as the map encoder to extract the map features, and uses the encoder of the time series 2 branch as the image encoder to extract the image features. In addition, in view of the modal difference between the map and image data, the present application removes the parameter sharing mechanism of the dual-branch feature extraction network to avoid the mutual interference of the map and image features. The change decoder of TBFFNet uses three change decoding modules based on convolutional neural layers to gradually extract the difference features of the map and image and realize change localization. The semantic decoder of TBFFNet uses two levels of stacked semantic decoding modules, each of which sequentially analyzes the feature of the ground object by two convolutional neural modules based on residual connection and one neural module based on extended convolution to realize image segmentation. The specific steps of predicting the semantic change detection result are as follows: Step S11: using one-hot encoding to represent the semantic information of the map , wherein represents the height of the map, represents the width of the map, and represents the number of categories of the ground object. Then, the map encoder extracts the map features of the ground object:

[0031] Step S12: processing the optical image data of the ground surface by the image encoder to extract the image features of the ground object:

[0032] Step S13: after extracting the features of the ground object, the change decoder constructs the time series information of the map and image data, generates the difference feature representation of the ground object change, and predicts the change area:

[0033] Step S14: in view of the common characteristics of the change localization and semantic segmentation tasks, the image features of both are shared and the semantic decoder ​Classify the visual features of the ground objects to obtain the class attribute at the pixel level:

[0034] After the above steps, the change probability of the map-image pair and the semantic probability of the image .

[0035] After completing the semantic change detection phase, enter the distribution transformation phase. This phase utilizes the causal relationship between change localization and semantic segmentation tasks, enabling the semantic change detection model based on the multi-task learning framework to learn with only change labels or semantic labels, as shown in the right half of Figure 1 . The specific transformation steps are as follows: Step S21, construct a derivable distribution transformation process to associate the semantic class and the binary change probability space. Given a class probability matrix before change and a class probability matrix after change , where each probability vector and obeys the class distribution. In the map-image data pair, because the semantic information before the change is known, each probability vector is discrete and one-hot, i.e. and . To generate the binary change probability, the argmax operation is generally used to discretize the probability distribution of , generate one-hot encoding, and then compare it with . However, the non-derivable nature of the argmax and comparison operations hinders the transmission of gradient signals from the binary change probability to the semantic class probability. Given that the ground object class before the change is determined, i.e. , the change probability at the corresponding position is equivalent to the probability that the pixel does not belong to class , i.e. , b=1 indicates that the position has a ground object change. Similarly, the non-change probability at position is equivalent to the probability that the pixel belongs to class , i.e. , b=0 indicates that the position has no ground object change. Therefore, the distribution transformation function for semantic probability can be defined as:

[0036] where . This distribution transformation function loosens the conversion process of the semantic distribution through multiplication operations, making it derivable.

[0037] Step S22, when only semantic labels are available for model training, use one-hot encoding operation to replace semantic labels with category probability matrix and generate pseudo change labels through distribution transformation function Learning for supervised change localization task. The loss function of the trained model is as follows:

[0038] wherein, and respectively represent the loss functions of change localization and semantic segmentation, both of which are realized by cross-entropy loss function.

[0039] Step 23, when only change labels are available for model training, the predicted semantic probability is converted to binary change probability by distribution transformation function, so as to be supervised by change labels. In the change region, the change label provides the information that the pixels in the region do not belong to a certain category, and this signal is weaker than the semantic label in supervising the semantic segmentation module. In order to further constrain the learning of the semantic segmentation module, a local information entropy penalty term is designed to increase the certainty of semantic probability by minimizing the local information entropy. The definition of local information entropy function is as follows:

[0040]

[0041] wherein, represents change label, 0 represents that the ground object has not changed, and 1 represents that the ground object has changed. Finally, the loss function of the trained model is as follows:

[0042] In order to verify the effective performance of the application in the case of only semantic segmentation or change training label, the HiUCD dataset containing map-image pairs and their change labels and image semantic segmentation labels is selected for experiment. The trained semantic change detection model is evaluated on the test set of the HiUCD dataset. In the experimental verification, the commonly used semantic change segmentation (SCS) evaluation index SCS is used to compare the semantic change detection performance of the model under different training label conditions. At the same time, the post-classification strategy based on the advanced semantic segmentation model SSG2 (for details, please refer to the literature: Diakogiannis F I, Furby S, Caccetta P, et al. Ssg2: A new modeling paradigm for semantic segmentation [J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2024, 215: 44-61.) is selected for comparison. This strategy obtains the change area and type of the map-image pair through image semantic segmentation and label comparison, so it only needs to train the segmentation model of the image pixel semantic label. This method can also obtain semantic change detection in the case of only single-task training label. The experimental results are shown in the following table:

[0043] In summary, the map-image semantic change detection method based on multi-task-single-label learning provided by the application aligns the output space of the change localization and image semantic segmentation decoder through distribution transformation, which facilitates cross-task supervision of heterogeneous labels. The application can train a semantic change detection model based on a multi-task learning framework in the case of only semantic segmentation and change training labels, reducing the difficulty of collecting and labeling time-series data and the cost of model training. The final experimental verification shows that the application can still achieve good semantic change performance in the case of only single-task training labels, and significantly outperforms the post-classification method based on SSG2. The application can be effectively applied to the semantic change detection task based on map-image pairs to achieve efficient identification of change areas and types.

[0044] On the other hand, the embodiment of the application also provides a map-image semantic change detection system based on multi-task-single-label learning, comprising: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the map-image semantic change detection method based on multi-task-single-label learning as described in the above technical solution.

[0045] It should be understood that parts of the specification not specifically described in detail are part of the prior art.

[0046] It should be understood that the above description of the preferred embodiments is merely a detailed explanation and is not considered as a limitation to the scope of patent protection of the present application. Any modification or alternation made by those skilled in the art without departing from the scope of the present application as defined in the claims shall fall within the scope of the present application. The scope of patent protection of the present application shall be subject to the appended claims.

Claims

1. A map-image semantic change detection method based on multi-task-single-label learning, characterized by: include: Semantic change detection stage: Use any semantic change detection model based on a multi-task learning framework as a baseline model to extract change information of the map-image pair. The baseline model includes a dual-branch feature extraction module, a change localization module, and a semantic segmentation module. The dual-branch feature extraction module uses two feature extractors to process the map and image data respectively. The extracted map and image features are passed through the change localization module to obtain the detection results of the changed area and the change probability of the map-image pair. Simultaneously, the image features are input into the semantic segmentation module to obtain the pixel semantic recognition results and the semantic probability of the image. Distribution transformation stage: Utilizing the causal relationship between change localization and semantic segmentation tasks, the semantic change detection model based on the multi-task learning framework can be learned with only change labels or semantic labels. When there are only semantic category labels, the semantic information is used to generate binary change labels to supervise the learning of the change localization module. When there are only binary change labels, the category probabilities output by the semantic segmentation module are converted into binary change probabilities to facilitate the supervision of the semantic segmentation module by the binary change label signal.

2. The map-image semantic change detection method based on multi-task-single-label learning according to claim 1, characterized in that: The dual-branch feature extraction module includes a map encoder and an image encoder; Use one-hot encoding to represent the semantic information of the map , then, using the map encoder Extract map features of objects: Through the image encoder Processing optical image data of the earth's surface , extract the image features of the object: The map encoder and the image encoder do not share model parameters with each other and independently extract map and image information, thereby ensuring the discriminability of ground feature characteristics.

3. The map-image semantic change detection method based on multi-task-single-label learning according to claim 1, characterized in that: By changing the positioning module, that is, changing the decoder Construct time series information of map and image data, generate differential feature representations of land feature changes, and predict change areas: in, Represent map features and image features respectively, represents the probability of change of a map-image pair.

4. The map-image semantic change detection method based on multi-task-single-label learning according to claim 1, characterized in that: Through the semantic segmentation module, namely the semantic decoder Classify the visual features of the ground objects and obtain pixel-level category attributes: in, Represents image features, Represents the semantic probability of an image.

5. The map-image semantic change detection method based on multi-task-single-label learning according to claim 1, characterized in that: In the distribution transformation phase, we first construct a differentiable distribution transformation process to associate the semantic categories with the probability space of binary changes; given a category probability matrix before the change and a changed class probability matrix , where each probability vector and All obey the category distribution; in the map-image data pair, since the semantic information before the change is known, each probability vector Discrete, one-hot form, i.e. and ; To generate binary change probabilities, use the argmax operation to discretize The probability distribution of , generates a unique hot encoding, and then compares it with For comparison; given that the land feature category before the change is certain, that is, , then the probability of change of the corresponding position This is equivalent to the corresponding pixel not belonging to the category The probability of , b=1 means that the corresponding position has changed; similarly, at position The unchanged probability is equivalent to the corresponding pixel not belonging to the category The probability of , b = 0 means that there is no ground feature change at the corresponding location; therefore, the distribution transformation function of semantic probability is defined as: in, Indicates the number of categories of features, , the distribution transformation function loosens the transformation process of semantic distribution through multiplication operation, making it differentiable.

6. The map-image semantic change detection method based on multi-task-single-label learning according to claim 5, characterized in that: When only semantic labels are available for model training, a one-hot encoding operation is used to convert the semantic labels into a class probability matrix. , and generate pseudo change labels through the distribution transformation function For supervised learning of the change localization task, the loss function of the training model is as follows: in, represents the probability of change of the map-image pair, represents the semantic probability of the image, and They represent the loss functions for change localization and semantic segmentation, respectively, both of which are implemented by the cross entropy loss function.

7. The map-image semantic change detection method based on multi-task-single-label learning according to claim 5, characterized in that: When only the change labels are available for model training, the predicted semantic probabilities are transformed through the distribution transformation function. Convert to binary change probability , which is convenient for changing labels to supervise it; Design a local information entropy penalty term to increase the certainty of semantic probability by minimizing the local information entropy; the local information entropy function The definition of is as follows: Among them, H and W represent height and width respectively, K represents the category of the object, Indicates the change label, 0 means the feature has not changed, and 1 means the feature has changed.

8. The map-image semantic change detection method based on multi-task-single-label learning according to claim 7, characterized in that: The final loss function of the trained model is as follows: in, represents the probability of change of the map-image pair, represents the semantic probability of the image, and They represent the loss functions for change localization and semantic segmentation, respectively, both of which are implemented by the cross entropy loss function.

9. The map-image semantic change detection method based on multi-task-single-label learning according to claim 1, characterized in that: It also includes the use of the evaluation index SCS to compare the semantic change detection performance under different training label situations.

10. A map-image semantic change detection system based on multi-task-single-label learning, characterized by: include: A processor and a memory, the memory being used to store program instructions, and the processor being used to call the stored instructions in the memory to execute a map-image semantic change detection method based on multi-task-single-label learning as described in any one of claims 1 to 9.