A method and system for three-dimensional and semantic change detection
By integrating knowledge graphs and generative contrastive adversarial networks, positive and negative generators and semantic detectors are constructed, solving the problems of limited training samples and weak deception of negative samples in existing technologies. This enables intelligent, automated, and refined mining of simultaneous 3D and semantic changes, and can detect changes in multimodal data under small sample conditions.
Patent Information
- Application Number
- CN202311257853.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-26
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2043-09-26
AI Technical Summary
Existing technologies for detecting simultaneous changes in 3D and semantic data suffer from problems such as limited training samples, weak deceptiveness of negative samples, and difficulty in effective adversarial analysis. This makes it difficult to achieve intelligent, efficient, and interpretable analysis of multimodal data and small-sample, intelligent, automated, and refined data mining.
By employing a method that integrates knowledge graphs and generative adversarial networks, positive and negative generators and semantic detectors are constructed. By combining contrastive learning and generative adversarial networks, multi-branch, multi-task 3D and semantic synchronous change detection is achieved. Large models are used to assist semantic segmentation to obtain small sample labels, preventing negative samples from being deceptive, and self-supervised learning is performed.
It enables simultaneous 3D and semantic change detection under small sample conditions, and can detect whether and how the 3D and semantics change. The detection results are refined and robust, and it can handle intelligent, efficient and interpretable analysis of multimodal data.
Smart Images

Figure CN117274811B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of remote sensing and artificial intelligence, and relates to a three-dimensional and semantic synchronous change detection method and system, in particular to a three-dimensional and semantic synchronous change detection method and system fusing a knowledge graph and a generative contrastive adversarial. BACKGROUND
[0002] The change detection dataset and the change detection algorithm are important foundations for three-dimensional and semantic synchronous change detection, and the current research status thereof is as follows:
[0003] The data acquisition methods, different angle descriptions, and the like result in a multi-modal problem. How to use machine learning methods to process multi-modal data has become a research focus in the field of three-dimensional and semantic synchronous change detection. A knowledge graph is a knowledge base for describing entities, concepts and their relationships in the objective world, and is widely used in the fields of remote sensing image scene classification, semantic segmentation and the like. With the development of computer technology, machine learning technology is widely used in the field of three-dimensional and semantic synchronous change detection. The current literatures on change detection based on a generative adversarial network (GAN) include:
[0004] (1) GAN-based remote sensing image change detection; Kousuke et al. use a generative adversarial network for cross-domain change detection (document 1), Alvarez et al. propose a self-supervised conditional generative adversarial network to solve the binary change detection task (document 2), Li et al. and Chen et al. use a generative adversarial network to detect changes in buildings (documents 3 and 4), Wu et al. use a generative adversarial network for regional supervision change detection (document 5), and Hou et al. use a W-Net as a generator for generative adversarial training (document 6);
[0005] (2) GAN-based remote sensing three-dimensional change detection; Niu et al. use a conditional generative adversarial network to perform three-dimensional change detection on heterogeneous synthetic aperture radar (document 7), Nagy et al. use a generative adversarial network to perform three-dimensional change detection on coarsely registered point clouds of street view data (document 8), and Yang Xiaodong et al. use a generative adversarial network to learn three-dimensional features and perform change detection (document 9);
[0006] The current literatures on change detection based on contrastive learning include:
[0007] (1) Contrastive learning-based remote sensing image change detection; Wang et al. train an encoder based on a supervised contrastive learning method for change detection (document 10); Ou et al. use contrastive learning to perform change detection on hyperspectral images (document 11); Zhang et al. use the idea of contrastive learning to perform change detection on buildings (document 12);
[0008] (2) Contrastive learning based remote sensing 3D change detection; de Gélis et al. proposed a 3D point cloud change detection method based on unsupervised learning of deep clustering and contrastive learning (document 13);
[0009] The above prior art has the following shortcomings:
[0010] Kousuke et al. used limited training samples (document 1), Alvarez et al. used a self-supervised method to only perform 2D change detection (document 2), Li et al. and Chen et al. used generative adversarial networks to only detect changes in buildings (documents 3, 4), Wu et al. used generative adversarial networks to perform a game process of generator, segmenter and discriminator (document 5), which is difficult to solve the problem that the negative sample is not strong in deception and difficult to effectively antagonize; Hou et al. used W-Net as a generator to only generate a change label result (document 6), Niu et al. and Nagy et al. can only perform 3D change detection (documents 7, 8), Yang Xiaodong et al. need to use artificial labeling method to obtain training samples (document 9), Wang et al. use contrastive learning to perform supervised change detection on high-resolution remote sensing images (document 10), Ou et al. use Gaussian noise data augmentation strategy (document 11), Zhang can only detect building changes or unchanged conditions (document 12), and de Gélis et al. mainly focus on whether the feature type of change detection changes (document 13).
[0011] Document source:
[0012] Document 1: Kousuke Y, Kanji T, Takuma S. Use of Generative Adversarial Network for Cross-Domain Change Detection [J]. arXiv preprint arXiv:1712.08868, 2017.
[0013] Document 2: Alvarez J L H, Ravanbakhsh M, Demir B. S2-cGAN: Self-supervised adversarial representation learning for binary change detection in multispectral images [C]. IGARSS 2020-2020 IEEE International Geoscience and Remote Sensing Symposium, 2020: 2515-2518. Document 3: Li Y, Chen H, Dong S, et al. Multi-Temporal Sample Pair Generation for Building Change Detection Promotion in Optical Remote Sensing Domain Based on Generative Adversarial Network [J]. Remote Sensing, 2023, 15(9): 2470.
[0014] Document 4: Chen H, Li W, Shi Z. Adversarial Instance Augmentation for Building Change Detection in Remote Sensing Images [J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 1-16.
[0015] Document 5: Wu C, Du B, Zhang L. Fully Convolutional Change Detection Framework with Generative Adversarial Network for Unsupervised, Weakly Supervised and Regional Supervised Change Detection [J], 2022.
[0016] Document 6: Hou B, Liu Q, Wang H, et al. From W-Net to CDGAN: Bitemporal change detection via deep learning techniques [J]. IEEE Transactions on Geoscience and Remote Sensing, 2019, 58(3): 1790-1802.
[0017] Document 7: Niu X, Gong M, Zhan T, et al. A Conditional Adversarial Network for Change Detection in Heterogeneous Images [J]. IEEE Geoscience and Remote Sensing Letters, 2019, 16(1): 45-49. Document 8: Nagy B, Kovacs L, Benedek C. ChangeGAN: A deep network for change detection in coarsely registered point clouds [J]. IEEE Robotics and Automation Letters, 2021, 6(4): 8277-8284. Document 9: Yang Xiaodong, Yan Hua. Fusion of DOM and DSM data three-dimensional change detection algorithm [J]. Scientific and Technological Innovation, 2022(028): 000.
[0018] Document 10: Wang J, Zhong Y, Zhang L. Change Detection Based on Supervised Contrastive Learning for High-Resolution Remote Sensing Imagery [J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 1-16.
[0019] Document 11: Ou X, Liu L, Tan S, et al. A Hyperspectral Image Change Detection Framework With Self-Supervised Contrastive Learning Pretrained Model[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2022, 15: 7724-7740.
[0020] Document 12: Zhang M, Li Q, Yuan Y, et al. Edge Neighborhood Contrastive Learning for Building Change Detection[J]. IEEE Geoscience and Remote Sensing Letters, 2023, 20: 1-5. Document 13: De Gélis I, Saha S, Shahzad M, et al. Deep Unsupervised Learning for 3D ALS Point Clouds Change Detection[J]. arXiv preprint arXiv: 2305.03529, 2023. SUMMARY
[0021] In order to realize intelligent and efficient explainable analysis of multi-dimensional multi-modal data, and achieve small sample, intelligent, automatic and fine mining of three-dimensional and semantic synchronous change information, the present application proposes a three-dimensional and semantic synchronous change detection method and system combining knowledge graph and generative contrastive adversarial.
[0022] The technical scheme adopted by the method of the present application is: a three-dimensional and semantic synchronous change detection method, characterized by comprising the following steps:
[0023] Step 1: Obtain a multi-modal three-dimensional and semantic change detection small sample data set, including three-dimensional geometric data, attribute data and text data of two time phases;
[0024] Step 2: Preprocess the three-dimensional geometric data, attribute data and text data of the two time phases respectively to obtain geometric data sets, attribute data sets and text data sets of the two time phases;
[0025] Step 3: Based on the two-phase geometric data set and attribute data set, a three-dimensional and semantic synchronous change detection data set is constructed, including semantic change, three-dimensional geometric change data, semantic change, three-dimensional geometric invariant data, semantic invariant, three-dimensional geometric change data and semantic invariant, three-dimensional geometric invariant data;
[0026] Step 4: Based on the knowledge graph, a positive and negative generator is constructed, and a large model is used to assist in semantic segmentation of the three-dimensional and semantic synchronous change detection data set to obtain positive and negative samples of three-dimensional and semantic change detection. Then, the three-dimensional and semantic synchronous change detection data set is expanded to obtain three-dimensional geometric change labels and semantic change labels; including geometric real change labels, geometric false alarm error labels and geometric missed alarm error labels, semantic real change labels, semantic false alarm error labels and semantic missed alarm error labels;
[0027] Step 5: Based on the idea of contrast learning, a three-dimensional and semantic momentum contrast encoder is constructed, and the positive and negative samples of the two-phase three-dimensional data and attribute data are input to pretrain the three-dimensional and semantic momentum contrast encoder;
[0028] Step 6: Based on contrast learning, a semantic detector is constructed, and a three-dimensional momentum contrast encoder, a semantic momentum contrast encoder and a semantic encoder are established respectively to encode the two-phase three-dimensional data, attribute data and semantic text data. Then, the decoder is decoded, the loss is calculated, and the gradient is updated;
[0029] Step 7: Based on the loss function of the adversarial network, the game of the positive and negative generator and the semantic detector is realized, and then the self-supervised three-dimensional and semantic synchronous change detection under the condition of small sample is completed; by judging whether the generated adversarial game meets the iteration termination condition, if the iteration termination condition is met, step 8 is executed; otherwise, step 4 is executed;
[0030] Step 8: Three-dimensional and semantic synchronous change detection information mining, three-dimensional momentum contrast encoder, semantic momentum contrast encoder and semantic encoder obtain three-dimensional geometric change, two-phase semantic segmentation result and semantic change result respectively;
[0031] Step 9: Analysis of three-dimensional and semantic synchronous change detection results, further analysis of semantic information mined by three-dimensional and semantic synchronous change detection, and results of semantic change, three-dimensional geometric change, semantic change, three-dimensional geometric invariance, semantic invariance, three-dimensional geometric change and semantic invariance, three-dimensional geometric invariance.
[0032] As preferred, in step 2, the geometric correction and radiometric correction are performed on the two-phase three-dimensional point cloud data and digital surface model data DSM to obtain a geometric data set;
[0033] The geometric correction is from image to map correction, first selecting a reference plane, then automatically matching to obtain the same name points, eliminating the mis-matching points, then performing orthographic correction, and finally generating a digital orthographic map DOM;
[0034] The radiation correction is completed by an empirical model method, an atmospheric correction method or a ground surface reflection correction method.
[0035] As preferred, in step 2, the text data set is based on the geoscience knowledge graph to obtain the semantic and topological relationship of the ground object information of two phases required for three-dimensional and geometric synchronous change detection, and is stored in the text data set;
[0036] The geoscience knowledge graph is a tree structure, and the leaf nodes represent the topological and spatial relationship between the ground objects. The construction of the knowledge graph includes the following steps:
[0037] (1) Geoscience knowledge extraction, extracting the semantic information of geographical scenes and geographical features contained in the geometric and attribute data, obtaining entity, attribute and relationship elements;
[0038] (2) Geoscience knowledge fusion, extracting and analyzing the semantic information of spatial pattern, evolution process, interaction mechanism and spatio-temporal distribution contained in the geometric and attribute data to generate a graph unit;
[0039] (3) Geoscience knowledge reasoning, performing knowledge reasoning from the perspectives of three-dimensional geometric features, spectral features, texture features or topological features to form a small sample text data set.
[0040] As preferred, in step 4, the positive and negative generator uses the semantic and topological information of different ground object types provided by the geoscience knowledge graph to establish the positive and negative generator, respectively establishes intelligent interpretation labels for the geometric and semantic information of multi-modal data, and generates positive samples and negative samples; the positive samples include geometric real change labels, geometric real invariable labels, semantic real change labels and semantic real invariable labels; the negative samples include geometric false detection error labels, geometric missed detection error labels, semantic false detection error labels and semantic missed detection error labels;
[0041] As preferred, the positive and negative generator based on the geoscience knowledge graph generates positive samples, and the specific implementation steps are as follows:
[0042] Step 4.1.1: Geometric transformation, the three-dimensional and semantic synchronous change detection data set contains three-dimensional geometric data of phase 1 and phase 2, image rotation and image mirror processing are performed on the two-phase point cloud and image, and geometric sample enhancement is performed on the three-dimensional point cloud and two-phase images;
[0043] Step 4.1.2: Radiation enhancement, three-dimensional and semantic synchronous change detection dataset contains attribute data of phase 1 and phase 2, histogram matching, histogram stretching, mixed object addition processing are performed on the two-phase images, and sample enhancement of radiation is performed;
[0044] The positive and negative generator based on the geosciences knowledge graph generates negative samples, and the specific implementation steps are as follows:
[0045] Step 4.2.1: Generate geometric correction negative samples;
[0046] (1) Generate single-pixel image registration error; registration error is generated by registering the positive samples generated by geometric transformation;
[0047] (2) Generate multi-pixel image registration error; based on the geosciences knowledge graph, the misregistration error caused by the overall shift of the image edge between the two images is simulated through morphological processing based on the positive samples generated by geometric transformation;
[0048] Step 4.2.2: Generate radiation correction negative samples; add radiation noise to the positive samples generated by radiation enhancement to simulate the pseudo change caused by shadows.
[0049] As preferred, the specific implementation of step 5 includes the following sub-steps:
[0050] Step 5.1: Establish a momentum contrast encoder based on the idea of contrast learning to realize the pre-training process of the downstream task semantic detector construction;
[0051] Step 5.1.1: Positive samples and negative samples are input respectively; the positive samples and negative samples are input into the DSM-based momentum contrast encoder and the image-based momentum contrast encoder respectively, and the three-dimensional momentum contrast encoder and the semantic momentum contrast encoder are obtained by training respectively;
[0052] Among them, the input of the three-dimensional momentum contrast encoder is the positive samples and negative samples of the multi-modal geometric height of the two phases; the input of the semantic momentum contrast encoder is the positive samples and negative samples of the multi-modal remote sensing image of the two phases;
[0053] Step 5.1.2: initialize three-dimensional momentum contrast encoder 3DEncoder, semantic momentum encoder ImageEncoder; during the pre-training process of the three-dimensional and semantic-based momentum contrast encoder, 3DEncoder and ImageEncoder correspond to a momentum encoder 3DMomentumEncoder, ImageMomentumEncoder respectively, and the parameters of 3DEncoder and 3DMomentumEncoder are kept consistent, and the parameters of ImageEncoder and ImageMomentumEncoder are kept consistent during the training process;
[0054] Step 5.1.3: repeatedly train the three-dimensional momentum contrast encoder and the semantic momentum contrast encoder;
[0055] (1) randomly sample samples 3Dx and Imagex from positive and negative samples input into the three-dimensional and semantic momentum contrast encoders, the randomly sampled samples 3Dx and Imagex are called anchor points, and the remaining positive and negative samples except the anchor points are called queues;
[0056] (2) randomly data augment the anchor points, and input the anchor points and the data augmented data into encoders 3DEncoder and ImageEncoder respectively, and calculate feature vectors 3Dx q and Imagex q respectively;
[0057] (3) input samples into 3DMomentumEncoder and ImageMomentumEncoder and calculate feature vectors 3Dx k and Imagex k ; the samples input into the 3DMomentumEncoder and ImageMomentumEncoder are samples in the queues and samples after another data augmentation of the anchor points; the samples after another data augmentation of the anchor points are positive samples, and the samples after data augmentation of the anchor points form a positive sample pair with the anchor points;
[0058] (4) calculate the inner product of the anchor points and the corresponding positive samples as the similarity, wherein the inner product is the corresponding multiplication of the feature vectors 3Dx q and 3Dx k , Imagex q and Imagex k ;
[0059] (5) calculate the inner product of the anchor points and all negative samples;
[0060] (6) combine the inner product of the anchor point and the positive sample and the inner product of the anchor point and the negative sample together;
[0061] (7) calculate the loss of the anchor point and the positive sample, and then distinguish the correct sample and the error sample; the loss can be divided into supervised loss function calculation and unsupervised loss function calculation according to whether the label data of the positive and negative samples is input; the more similar the anchor point and the positive sample are, the smaller the loss value is; the closer the anchor point and the negative sample are, the larger the loss value is;
[0062] (8) update the parameters of the encoder 3DEncoder, ImageEncoder, 3DMomentumEncoder and ImageMomentumEncoder;
[0063] Step 5.1.4: Momentum Contrast Encoder Result Output, output the three-dimensional momentum contrast encoder and semantic momentum encoder trained by the momentum contrast encoder to subsequent tasks for use.
[0064] As preferred, the specific implementation of step 6 includes the following sub-steps:
[0065] Step 6.1: Constructing a contrast learning semantic decoder, decoding the three-dimensional and semantic feature vectors 3Dx q and Imagex q into original three-dimensional and semantic samples 3Dx and Imagex, training the semantic decoder to extract the three-dimensional and semantic features obtained by the three-dimensional momentum contrast encoder and the semantic momentum encoder;
[0066] Step 6.2: Under the condition of small samples, combining the three-dimensional and semantic encoders obtained by the momentum contrast encoder and the contrast learning semantic decoder to construct a contrast learning semantic detector;
[0067] Step 6.2.1: Data input, including remote sensing image data of the previous phase, geometric elevation data of the previous phase, text information data of the previous phase, remote sensing image data of the later phase, geometric elevation data of the later phase and text information data of the later phase; the remote sensing image data of the previous and later phases is used to obtain the remote sensing image segmentation results of the previous and later phases, the text information data of the previous and later phases is used to obtain the semantic change results, and the geometric elevation data of the previous and later phases is used to obtain the three-dimensional geometric change results, and the four semantic detector identification results are obtained at the same time;
[0068] Step 6.2.2: calling the pre-trained encoders 3DEncoder, ImageEncoder;
[0069] Step 6.2.3: decoding by the contrast learning semantic decoder;
[0070] Step 6.2.4: input sample label and calculate loss value, gradient update;
[0071] Step 6.2.5: semantic detector result analysis, the identification result of the semantic detector includes the semantic segmentation result of the pre-phase remote sensing image, the semantic segmentation result of the post-phase remote sensing image, the semantic change result and the three-dimensional geometric change result.
[0072] As a preferred, the specific implementation of step 8 includes the following sub-steps:
[0073] Step 8.1: semantic detector based on contrastive learning detects three-dimensional change information, and three-dimensional change information is mined through small sample self-supervised learning;
[0074] Step 8.2: semantic information of geometric and semantic change is obtained through adversarial game, loss function of generative contrastive adversarial network is calculated, and then whether three-dimensional geometric and semantic information changes and how it changes are obtained; the result of three-dimensional change detection includes: semantic segmentation result of pre-phase remote sensing image, semantic segmentation result of post-phase remote sensing image, semantic change result and three-dimensional geometric change result; the three-dimensional change detection result is updated to the three-dimensional change detection dataset;
[0075] Step 8.3: dynamic update of three-dimensional and semantic synchronous change detection dataset;
[0076] Step 8.3.1: sample label update, small sample label dataset based on knowledge graph, positive and negative samples generated by positive and negative generator;
[0077] Step 8.3.2: geometric and semantic detection results of three-dimensional change detection, while generating detection results of three-dimensional change detection, real-time dynamic update of dataset is realized.
[0078] As a preferred, the specific implementation of step 9 includes the following sub-steps:
[0079] Step 9.1: three-dimensional change detection network performance analysis, the convergence and robustness performance of the generative contrastive adversarial network are judged by calculating the network loss value of three-dimensional and semantic synchronous change detection;
[0080] Step 9.2: quantitative evaluation of three-dimensional and semantic synchronous change detection;
[0081] Step 9.3: qualitative evaluation of three-dimensional and semantic synchronous change detection;
[0082] The results of three-dimensional and semantic synchronous change detection are used to analyze the semantic segmentation results of the remote sensing images of the pre-phase and post-phase, to obtain how the semantics change; the semantic change results and the three-dimensional geometric change results are analyzed to obtain the following types of changes: semantic change, three-dimensional geometric change, semantic change, three-dimensional geometric invariance, semantic invariance, three-dimensional geometric change, semantic invariance, and three-dimensional geometric invariance.
[0083] The technical scheme adopted by the system of the application is: a three-dimensional and semantic synchronous change detection system, characterized by comprising:
[0084] One or more processors;
[0085] A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional and semantic synchronous change detection method.
[0086] The innovation points of the application include:
[0087] (1) A multi-task three-dimensional and semantic synchronous change detection method is used to realize three-dimensional change, geometric change, and semantic change synchronously;
[0088] (2) Only positive and negative generators and semantic detector adversarial detection methods are used;
[0089] (3) A generative adversarial network and contrastive learning are combined to realize a multi-branch and multi-input method for three-dimensional and semantic synchronous change detection;
[0090] (4) A knowledge graph is used to prevent the negative sample labels generated by the generator from being deceptive, i.e., the pseudo change is not realistic enough, and the problem of ineffective confrontation cannot be solved;
[0091] (5) A large model is used to assist semantic segmentation to obtain small sample labels for three-dimensional change detection, and small sample three-dimensional and semantic change detection is realized.
[0092] Compared with the prior art, the beneficial effects of the application include:
[0093] (1) GAN-based change detection:
[0094] Kousuke et al. used limited training samples (document 1), while the application not only uses a large model to assist semantic segmentation to obtain small sample labels for three-dimensional change detection, but also constructs a three-dimensional change detection dataset based on a knowledge graph, and realizes small sample three-dimensional and semantic synchronous change detection, and realizes the detection of whether three-dimensional and semantic change and how to change;
[0095] Alvarez et al. used a self-supervised method for two-dimensional change detection (document 2), while the positive and negative generator based on the knowledge graph of the present application can generate intelligent labels for both three-dimensional data and attribute data under the condition of small samples, realizing three-dimensional and semantic change detection under the condition of small samples; in addition, the present application realizes three-dimensional and attribute synchronous change detection through the intelligent game process of the positive and negative generator and the semantic detector, and completes the multi-branch input of three-dimensional and semantic synchronous change detection, multi-task implementation through the contrastive learning and the generative adversarial network;
[0096] Li et al. and Chen et al. used generative adversarial networks for building change detection (documents 3, 4), while the present application can realize three-dimensional and semantic synchronous change detection, not only detecting whether the building and other feature types change, but also detecting the change of different feature types and three-dimensional terrain; in addition, the present application constructs a positive and negative generator based on the idea of knowledge graph, and constructs a semantic detector based on the idea of contrastive learning, and realizes three-dimensional and semantic synchronous change detection under the condition of small samples through the game process of the positive and negative generator and the semantic detector;
[0097] Wu et al. used a game process of generator, segmenter and discriminator using generative adversarial networks (document 5), while the present application constructs a generative contrastive adversarial network using only the game of positive and negative generator and semantic detector, realizes three-dimensional and semantic synchronous change detection, and detects the semantic information of three-dimensional and attribute change, three-dimensional change, attribute change, etc.; and the generator based on the knowledge graph of the present application generates small sample labels to effectively prevent the problem that the negative sample is not strong in cheating and is difficult to effectively counter;
[0098] Hou et al. used W-Net as a generator to only generate change label results (document 6), while the positive and negative generator of the present application generates labels and images of positive samples and negative samples, and the negative sample will cheat the semantic detector to game, in addition, the present application uses knowledge graph to construct the positive and negative generator, and uses contrastive learning to construct the semantic detector to cooperate with the positive and negative generator to conduct three-dimensional and semantic synchronous change detection;
[0099] Niu et al. and Nagy et al. used synthetic aperture radar and coarse registration point cloud for three-dimensional change detection (documents 7, 8), while the multi-modal three-dimensional change detection dataset constructed by the present application not only contains three-dimensional geometric information and spectral information, but also provides fine and high-quality small sample interpretation labels for neural network modeling through the positive and negative generator based on the full use of knowledge graph; in addition, the present application designs a semantic detector based on contrastive learning to realize multi-branch input of three-dimensional and semantic synchronous change detection; through the game process of generative adversarial, it realizes the detection of whether the three-dimensional attribute changes, how to change, etc.
[0100] Yang Xiaodong et al. use artificial labeling method to obtain training samples (document 9), while the present invention generates a small sample labeled dataset for three-dimensional and semantic synchronous change detection based on the positive and negative generator of the knowledge graph, and performs three-dimensional and semantic synchronous change detection under the condition of small sample; in addition, the present invention constructs a positive and negative generator and a semantic detector based on the knowledge graph and contrastive learning, and through the game of the two, a multi-branch, multi-task generative adversarial network is constructed, and then three-dimensional and semantic synchronous change detection is realized;
[0101] (2) Change detection based on contrastive learning;
[0102] Wang et al. use contrastive learning for supervised change detection of high-resolution remote sensing images (document 10), while the present invention generates positive and negative samples based on the positive and negative generator of the knowledge graph, thereby realizing three-dimensional and semantic synchronous change detection under the condition of small sample; in addition, the present invention realizes three-dimensional and semantic synchronous change detection through the game of the positive and negative generator and the semantic detector, which not only can detect whether the high-resolution image changes, but also can perform three-dimensional change detection, and simultaneously realize the detection problems of whether it changes and how it changes;
[0103] Ou et al. use a data augmentation strategy of Gaussian noise (document 11), while the present invention uses a large model assisted semantic segmentation based on a knowledge graph to obtain a small sample label for three-dimensional change detection; the positive and negative generator designed by the present invention can not only generate change labels, but also generate image corresponding to the labels; in addition, the present invention fuses contrastive learning and generative adversarial network, and realizes three-dimensional and semantic synchronous change detection;
[0104] Zhang et al. use the idea of contrastive learning to detect building changes or invariable conditions (document 12), while the present invention not only detects whether the feature type changes, but also detects three-dimensional changes, how the feature type changes, etc.; the generative adversarial process of the positive and negative generator and the semantic detector designed by the present invention realizes multi-task training of three-dimensional and semantic synchronous change detection;
[0105] de Gélis et al. mainly focus on whether the feature type changes in change detection (document 13), while the three-dimensional and semantic synchronous change detection method proposed by the present invention is based on the fusion of three-dimensional information, and designs a positive and negative generator and a semantic detector, which realizes intelligent semantic information mining of three-dimensional change detection through a generative adversarial network, and analyzes four three-dimensional change detection results, i.e. whether the three-dimensional geometry and semantic information changes, how to change. BRIEF DESCRIPTION OF DRAWINGS
[0106] The technical solutions herein are further described below using examples and specific embodiments. In addition, some drawings are also used in the description of the technical solutions. Other drawings and the intent of the present application can also be obtained by those skilled in the art without creative labor based on these drawings.
[0107] Figure 1 is a three-dimensional and semantic synchronous change detection network flowchart of an embodiment of the present application;
[0108] Figure 2 is a momentum contrast encoder flowchart of an embodiment of the present application;
[0109] Figure 3 is a semantic detector flowchart of an embodiment of the present application. DETAILED DESCRIPTION
[0110] In order to facilitate those skilled in the art to understand and implement the present application, the present application is further described in detail below in combination with the drawings and examples. It should be understood that the implementation examples described herein are only used to illustrate and explain the present application, and are not used to limit the present application.
[0111] See Figure 1 The three-dimensional and semantic synchronous change detection method provided in the embodiment includes the following steps:
[0112] Step 1: Obtain a multi-modal three-dimensional and semantic change detection small sample data set, including three-dimensional geometric data, attribute data and text data of two time phases;
[0113] In an embodiment, the three-dimensional geometric data includes three-dimensional point cloud data and digital surface model data DSM of two time phases; the attribute data includes digital orthophoto map data DOM and land use and land cover data LULC of two time phases; and the text data is text data TXT of two time phases generated based on a geoscience knowledge graph.
[0114] Step 2: Preprocess the three-dimensional geometric data, attribute data and text data of the two time phases respectively to obtain geometric data sets, attribute data sets and text data sets of the two time phases;
[0115] In an embodiment, the three-dimensional point cloud data and digital surface model data DSM of the two time phases are subjected to geometric correction and radiation correction to obtain the geometric data sets.
[0116] The geometric correction is from image to map correction, first selecting a reference plane, then automatically matching to obtain the same name point by using an automatic scheme, using the RANSAC algorithm to remove the mis-matching points, then using cloud control photogrammetry technology for ortho correction, and finally generating a digital ortho map DOM (Digital Ortho Map);
[0117] The radiation correction is the basis of change detection, and is completed by an empirical model method, an atmospheric correction method or a surface reflection correction method. Through the radiation correction, the difference between the two time-phase images is reduced as much as possible.
[0118] In an embodiment, the text data set is used to obtain the semantic and topological relationship of the ground object information of two time phases required for three-dimensional and geometric synchronous change detection based on a geoscience knowledge graph, and is stored in the text data set;
[0119] The geoscience knowledge graph is a tree structure, and the leaf nodes represent the topological and spatial relationship between the ground objects. The construction of the knowledge graph includes the following steps:
[0120] (1) Geoscience knowledge extraction, extracting the semantic information of the geographic scene and geographic features contained in the geometric and attribute data, obtaining entity, attribute and relationship elements;
[0121] (2) Geoscience knowledge fusion, extracting and analyzing the semantic information of the spatial pattern, evolution process, interaction mechanism and spatio-temporal distribution contained in the geometric and attribute data to generate a graph unit;
[0122] (3) Geoscience knowledge reasoning, knowledge reasoning from the perspective of three-dimensional geometric features, spectral features, texture features or topological features to form a small sample text data set.
[0123] Step 3: Based on the geometric data set and the attribute data set of two time phases, a three-dimensional and semantic synchronous change detection data set is constructed, including semantic change, three-dimensional geometric change data, semantic change, three-dimensional geometric invariant data, semantic invariant, three-dimensional geometric change data and semantic invariant, three-dimensional geometric invariant data;
[0124] Step 4: Based on the knowledge graph, a positive and negative generator is constructed, a large model is used to assist in semantic segmentation of the three-dimensional and semantic synchronous change detection data set to obtain positive samples and negative samples of three-dimensional and semantic change detection, and then the three-dimensional and semantic synchronous change detection data set is expanded to obtain three-dimensional geometric change labels and semantic change labels; including geometric real change labels, geometric false alarm error labels and geometric missed detection error labels, semantic real change labels, semantic false alarm error labels and semantic missed detection error labels;
[0125] In an embodiment, the positive-negative generator uses the semantic and topological information of different ground object types provided by the geoscience knowledge graph to establish the positive-negative generator, and respectively establishes intelligent interpretation labels for the geometric and semantic information of multi-modal data to generate positive samples and negative samples; the positive samples include geometric real change labels, geometric real invariance labels, semantic real change labels, and semantic real invariance labels; and the negative samples include geometric false detection error labels, geometric missed detection error labels, semantic false detection error labels, and semantic missed detection error labels.
[0126] In an embodiment, the positive-negative generator based on the geoscience knowledge graph generates positive samples, and the specific implementation steps are as follows:
[0127] Step 4.1.1: geometric transformation, three-dimensional and semantic synchronous change detection data set contains three-dimensional geometric data of time phase 1 and time phase 2, and image rotation, image mirror processing is performed on two time phase point clouds and images, and geometric sample enhancement is performed on three-dimensional point clouds and two time phase images;
[0128] Step 4.1.2: radiation enhancement, the three-dimensional and semantic synchronous change detection data set contains attribute data of time phase 1 and time phase 2, and histogram matching, histogram stretching, and mixed ground object addition processing are performed on two time phase images to perform radiation sample enhancement;
[0129] In an embodiment, the positive-negative generator based on the geoscience knowledge graph generates negative samples, and the specific implementation steps are as follows:
[0130] Step 4.2.1: generate geometric correction negative samples;
[0131] (1) generate single pixel image registration error; generate registration error by registration on the positive samples generated by geometric transformation;
[0132] (2) generate multi-pixel image registration error; based on the positive samples generated by geometric transformation, the geoscience knowledge graph simulates the mismatching error caused by the overall shift of the image edge of two images, such as the deformation of the mountain caused by the image edge misregistration, by morphological processing such as dilation and erosion; the morphological processing operation refers to a collection of image processing operations based on image shape, and the dilation and erosion are common means in morphological processing;
[0133] 1) the result of image A after dilation by structure element B is as follows:
[0134]
[0135] wherein, B x ={x+b|b∈B} represents the geometry of the pixels after the structure element is translated by x, and b is the coordinate of the pixel in image B;
[0136] 2) The result of the erosion of image A by structure element B is as follows:
[0137]
[0138] where the definitions of the variables remain the same as in equation (1);
[0139] Step 4.2.2: Generate radiation correction negative samples; add radiation noise to the positive samples generated by radiation enhancement to simulate the pseudo changes caused by shadows.
[0140] In an embodiment, the positive and negative generator based on the geosciences knowledge graph performs sample expansion, and the specific implementation steps are as follows: the positive samples and negative samples are respectively processed by image rotation, image mirroring, etc. to perform sample expansion, and then generate geometric correct samples, geometric error samples, semantic correct samples, and semantic error samples;
[0141] Step 5: Construct a three-dimensional and semantic momentum contrast encoder based on the idea of contrast learning, input the positive and negative samples of the three-dimensional data and attribute data of two time phases, and pre-train the three-dimensional and semantic momentum contrast encoder;
[0142] See Figure 2 In an embodiment, the specific implementation of step 5 includes the following sub-steps:
[0143] Step 5.1: Establish a momentum contrast encoder based on the idea of contrast learning to realize the pre-training process of the downstream task semantic detector construction;
[0144] Step 5.1.1: The positive samples and negative samples are respectively input, and the correct samples and error samples are input under the condition of sufficient guarantee of balance; the positive samples and negative samples are respectively input into the DSM-based momentum contrast encoder and the image-based momentum contrast encoder, and are respectively trained to obtain the three-dimensional momentum contrast encoder and the semantic momentum contrast encoder; the input of the positive and negative samples can divide the momentum contrast encoder into a supervised momentum contrast encoder and an unsupervised momentum contrast encoder according to whether the three-dimensional and semantic label data are input;
[0145] Among them, the input of the three-dimensional momentum contrast encoder is the positive and negative samples of the multi-modal geometric height of two time phases; the input of the semantic momentum contrast encoder is the positive and negative samples of the multi-modal remote sensing image of two time phases;
[0146] Step 5.1.2: initialize the three-dimensional momentum contrast encoder and the semantic momentum encoder, initialize the three-dimensional momentum contrast encoder 3DEncoder, the semantic momentum encoder ImageEncoder; in the pre-training process of the three-dimensional and semantic-based momentum contrast encoder, 3DEncoder and ImageEncoder correspond to a momentum encoder 3DMomentumEncoder, ImageMomentumEncoder respectively, and the parameters of 3DEncoder and 3DMomentumEncoder are kept consistent, and the parameters of ImageEncoder and ImageMomentumEncoder are kept consistent during the training process;
[0147] Step 5.1.3: repeatedly train the three-dimensional momentum contrast encoder and the semantic momentum contrast encoder;
[0148] (1) randomly sample samples 3Dx and Imagex from positive and negative samples input into the three-dimensional and semantic momentum contrast encoder, the randomly sampled samples 3Dx and Imagex are called anchor points, and the remaining positive and negative samples except the anchor points are called queues;
[0149] (2) randomly data augment the anchor points, and input the anchor points and the data augmented data into the encoders 3DEncoder and ImageEncoder respectively, and calculate the feature vectors 3Dx q and Imagex q respectively;
[0150] (3) input the samples into 3DMomentumEncoder and ImageMomentumEncoder and calculate the feature vectors 3Dx k and Imagex k ; the samples input into the 3DMomentumEncoder and ImageMomentumEncoder are samples in the queues and samples after another data augmentation of the anchor points; the samples after another data augmentation of the anchor points are used as positive samples, and the samples after data augmentation of the anchor points form a positive sample pair; the samples in the queues are used as negative samples; the momentum used when updating the parameters of the momentum encoders 3D Momentum Encoder f_k_3D and Image Momentum Encoder f_k_Image is as follows:
[0151] θ k ←γθ k +(1-γ)θ q (3)
[0152] wherein, θ kθ represents the parameters to be updated of 3D Momentum Encoder f_k_3D and Image Momentum Encoder f_k_Image q θ represents the parameters of 3D Encoder f_q_3D and Image Encoder f_q_Image q γ (γ ∈ [0, 1]) is a momentum parameter, which is updated by back propagation.
[0153] (4) Calculate the inner product of the anchor point and the corresponding positive sample as the similarity, and the inner product is the multiplication of the feature vectors 3Dx q and 3Dx k , Image x q and Image x k ;
[0154] (5) Calculate the inner product of the anchor point and all negative samples;
[0155] (6) Merge the inner product of the anchor point and the positive sample and the inner product of the negative sample together, and the merging uses the concat function;
[0156] (7) Calculate the loss of the anchor point and the positive sample, and then distinguish the correct sample and the error sample; the loss can be divided into supervised loss function calculation and unsupervised loss function calculation according to whether the label data of the positive and negative samples is input; the more similar the anchor point and the positive sample are, the smaller the loss value is; the closer the anchor point and the negative sample are, the larger the loss value is; the basic principle of the momentum contrast encoder loss definition is as follows:
[0157]
[0158] Where τ is a temperature hyperparameter, q represents the anchor point, k + represents the positive sample, and k i (i = 0, 1, …, K) represents the negative sample, and the momentum contrast encoder loss is the log loss based on the softmax classifier.
[0159] (8) Update the parameters of the encoders 3D Encoder, Image Encoder, 3D Momentum Encoder and
[0160] Image Momentum Encoder;
[0161] Step 5.1.4: Momentum contrast encoder result output, output the three-dimensional momentum contrast encoder and semantic momentum encoder trained by the momentum contrast encoder to the subsequent tasks for use.
[0162] Step 6: Constructing a semantic detector based on contrastive learning, respectively establishing a three-dimensional momentum contrast encoder, a semantic momentum contrast encoder, and a semantic encoder, respectively encoding the three-dimensional data, attribute data, and semantic text data of the two time phases, and then performing decoder decoding, loss calculation, and gradient update operations;
[0163] See Figure 2 In an embodiment, the specific implementation of step 6 includes the following sub-steps:
[0164] Step 6.1: Constructing a contrastive learning semantic decoder, combining the three-dimensional and semantic feature vectors 3Dx q and Imagex q decoded into original three-dimensional and semantic samples 3Dx and Imagex, training the semantic decoder to extract three-dimensional and semantic features obtained by the three-dimensional momentum contrast encoder and the semantic momentum encoder;
[0165] Step 6.2: Under the condition of small samples, combining the three-dimensional and semantic encoders obtained by the momentum contrast encoder and the contrastive learning semantic decoder to construct a contrastive learning semantic detector;
[0166] Step 6.2.1: Data input, including remote sensing image data of the previous time phase, geometric elevation data of the previous time phase, text information data of the previous time phase, remote sensing image data of the later time phase, geometric elevation data of the later time phase, and text information data of the later time phase; the remote sensing image data of the previous and later time phases are used to obtain remote sensing image segmentation results of the previous and later time phases, the text information data of the previous and later time phases are used to obtain semantic change results, and the geometric elevation data of the previous and later time phases are used to obtain three-dimensional geometric change results, and the four semantic detector identification results are obtained simultaneously;
[0167] Step 6.2.2: Call the pre-trained encoder 3DEncoder, ImageEncoder in the upstream task;
[0168] Step 6.2.3: Contrastive learning semantic decoder decoding;
[0169] Step 6.2.4: Input sample labels and calculate loss values, and perform gradient updates;
[0170] The loss functions used by the semantic detector are defined as follows:
[0171] Loss s1 :
[0172] Loss s1 is the loss function of the semantic segmentation result of time phase 1, using the LossDice function, defined as follows:
[0173]
[0174] wherein, X1 is the real pixel label of time phase 1, Y1 is the segmentation pixel label of time phase 1; |X1∩Y1| represents the result of the dot product of the real label pixels of time phase 1 and the segmentation image pixels and then summing; |X1| and |Y1| respectively represent the sum of the corresponding pixels of the real pixels and the segmentation pixels of time phase 1; y 1(i) is the label of three-dimensional geometric sample i (i = 1, …, m), the changed pixel is 1, and the unchanged pixel is 0; m is the number of input samples; p 1(i) is the probability that the three-dimensional geometric sample i (i = 1, …, m) is predicted as a changed pixel; and is an adaptive weight, δ s1 , δ s1(1) and δ s1(2) are noise parameters; Loss s2 :
[0175] Loss s2 is a loss function of time phase 2 semantic segmentation result, using LossDice function, defined as follows:
[0176]
[0177] wherein, X2 is the real pixel label of time phase 2, Y2 is the segmentation pixel label of time phase 2; |X2∩Y2| represents the result of the dot product of the real label pixels of time phase 2 and the segmentation image pixels and then summing; |X2| and |Y2| respectively represent the sum of the corresponding pixels of the real pixels and the segmentation pixels of time phase 2; y 2(i) is the label of three-dimensional geometric sample i (i = 1, …, m), the changed pixel is 1, and the unchanged pixel is 0; m is the number of input samples; p 2(i) is the probability that the three-dimensional geometric sample i (i = 1, …, m) is predicted as a changed pixel; and is an adaptive weight, δ s2 , δ s2(1) and δ s2(2) are noise parameters; Loss 3D :
[0178] Loss 3D is a loss function for judging whether the three-dimensional geometry changes or not, using cross-entropy loss function, defined as follows:
[0179]
[0180] wherein, y 3D(i) is the label of three-dimensional geometric sample i (i = 1, …, m), the changed pixel is 1, and the unchanged pixel is 0; m is the number of input samples; p 3D(i)P (i) is the probability of the three-dimensional geometric sample i (i = 1, …, m) being predicted as a changed pixel;
[0181] Loss Img :
[0182] Loss Img The loss function for judging whether the semantics change or not uses a cross-entropy loss function, and is defined as follows:
[0183]
[0184] Wherein, y Img(i) is the label of the semantic sample i (i = 1, …, m), and the changed pixel is 1 and the unchanged pixel is 0; m is the number of input samples; p Img(i) is the probability of the semantic sample i (i = 1, …, m) being predicted as a changed pixel;
[0185] Loss Joint :
[0186] Loss Joint The semantic detector joint loss function is defined as follows:
[0187] Loss Jiont = w1Loss s1 + w2Loss s2 + w3Loss 3D + w4Loss Img + δ T (9)
[0188] Wherein, δ T is a noise parameter, w1, w2, w3 and w4 are adaptive weights of the joint loss function, and are defined as follows:
[0189]
[0190] Wherein, δ1, δ2, δ3 and δ4 are noise parameters corresponding to w1, w2, w3 and w4 respectively;
[0191] Step 6.2.5: semantic detector result analysis, the identification result of the semantic detector includes the semantic segmentation result of the previous phase remote sensing image, the semantic segmentation result of the later phase remote sensing image, the semantic change result and the three-dimensional geometric change result. The semantic change includes whether the attribute of the ground object changes, and what type of ground object the original ground object changes into.
[0192] Step 7: Realize the game of positive and negative generators and semantic detectors based on the loss function of the adversarial network, and then complete the self-supervised three-dimensional and semantic synchronous change detection under the condition of small samples; through judging whether the generated adversarial game meets the judgment condition of iteration termination, if the iteration termination condition is met, steps 8 and 9 are executed; otherwise, steps 4, 5 and 6 are executed again.
[0193] The three-dimensional and semantic synchronous change detection semantic information intelligent mining of the embodiment is realized through the game process of the positive and negative generators G and the semantic detector D of the generative adversarial network, and further semantic information of the geometric and semantic synchronous three-dimensional change detection is mined. For each iteration process in the training, the following steps are executed:
[0194] (1) For k steps in one iteration, the following steps are executed:
[0195] 1) Three-dimensional geometric and semantic data input: input the three-dimensional geometric and semantic data corrected by geometric radiation, from the original data distribution p data (x) Sample a small batch of m samples {x (1) ,x (2) ,……,x (m)};
[0196] 2) Positive and negative generator G generates correct and incorrect samples: the positive and negative generator generates correct and incorrect samples based on the knowledge graph, from the prior noise distribution p g (z) Sample a small batch of m noise samples {z (1) ,z (2) ,……,z (m)};
[0197] 3) Semantic detector D detects three-dimensional change information based on contrast learning: realize the mining of three-dimensional change information through small sample self-supervised learning; update the semantic detector by ascending random gradient, and the basic principle is as follows:
[0198]
[0199] Where θ d represents the parameters of the three-dimensional or semantic momentum contrast encoder, D is the semantic detector, G is the positive and negative generator, x (i) is the original label, and G(z (i) ) is the generated sample.
[0200] (2) Positive and negative generator G generates correct and incorrect samples, from the prior noise distribution p g (z) Sample a small batch of m noise samples {z (1)z (2) ,……,z (m)};
[0201] (3) Update the positive and negative generator by descending stochastic gradient, the basic principle is as follows:
[0202]
[0203] Wherein, θ g represents the parameters of the positive and negative generator, D is the semantic detector, G(z (i) ) is the generated sample;
[0204] The semantic information of geometric and semantic changes is obtained by the adversarial game, and the loss function of the generated contrastive adversarial network is calculated in the adversarial game process;
[0205] The semantic detector needs to distinguish the original sample x and the generated sample G(z) to the maximum, so it needs to satisfy D(x) as large as possible and D(G(z)) as small as possible, therefore, the loss function of the generated contrastive adversarial network needs to satisfy the following formula:
[0206]
[0207] Wherein, L(D) represents the semantic detector loss; L(G) represents the positive and negative generator loss; x is the original sample, which is subject to the distribution p data (x); G(z) is the sample generated by the positive and negative generator G, which is subject to the distribution p g (z);
[0208] Through the game process of the positive and negative generator G and the semantic detector D of the generated contrastive adversarial network, whether the three-dimensional geometric and semantic information changes and how it changes are obtained; the result of the three-dimensional change detection has four kinds:
[0209] (1) The semantic segmentation result of the front time phase remote sensing image;
[0210] (2) The semantic segmentation result of the remote sensing image of the later time phase;
[0211] (3) The semantic change result; the semantic change includes whether the attribute of the ground object changes, and what type of ground object the original ground object changes into;
[0212] (4) The three-dimensional geometric change result.
[0213] Step 8: three-dimensional and semantic synchronous change detection information mining, three-dimensional momentum contrast encoder, semantic momentum contrast encoder and semantic encoder respectively obtain the results of whether the three-dimensional geometry changes, the semantic segmentation results of two time phases, and whether the semantic changes;
[0214] In an embodiment, the specific implementation of step 8 includes the following sub-steps:
[0215] Step 8.1: Detecting three-dimensional change information based on a contrast learning semantic detector, and mining three-dimensional change information through small sample self-supervised learning;
[0216] Step 8.2: Obtain semantic information of geometric and semantic changes through adversarial game, calculate the loss function of the generated contrast adversarial network, and then obtain whether the three-dimensional geometric and semantic information changes and how it changes; the result of the three-dimensional change detection includes: semantic segmentation result of the previous phase remote sensing image, semantic segmentation result of the later phase remote sensing image, semantic change result and three-dimensional geometric change result; the semantic change includes whether the attribute of the ground object changes, and the original ground object changes into what type of ground object; update the three-dimensional change detection result to the three-dimensional change detection dataset;
[0217] Step 8.3: Dynamic update of three-dimensional and semantic synchronous change detection dataset;
[0218] Step 8.3.1: Update of sample label, generation of positive and negative samples based on small sample label dataset of knowledge graph and positive and negative generators;
[0219] Step 8.3.2: Geometric and semantic detection results of three-dimensional change detection, real-time dynamic update of dataset while generating detection results of three-dimensional change detection.
[0220] Step 9: Analysis of three-dimensional and semantic synchronous change detection results, further analysis of semantic information mined by three-dimensional and semantic synchronous change detection, and obtain results of semantic change, three-dimensional geometric change, semantic change, three-dimensional geometric invariance, semantic invariance, three-dimensional geometric change and semantic invariance, three-dimensional geometric invariance.
[0221] In an embodiment, the specific implementation of step 9 includes the following sub-steps:
[0222] Step 9.1: Analysis of three-dimensional change detection network performance, judge the convergence and robustness of the generated contrast adversarial network by calculating the network loss value of three-dimensional and semantic synchronous change detection;
[0223] Step 9.2: Quantitative evaluation of three-dimensional and semantic synchronous change detection, the present application uses a confusion matrix and intersection over Union (IOU) to jointly detect the three-dimensional change detection result. The overall accuracy (OA), kappa coefficient, producer accuracy (PA), user accuracy (UA), and other traditional accuracy evaluation indexes are calculated through the confusion matrix. The IOU is defined as the proportion of the area of the intersection of the correctly detected geometric and semantic change boundary and the true geometric and semantic change boundary to the union set.
[0224] Step 9.3: Qualitative evaluation of three-dimensional and semantic synchronous change detection;
[0225] The results of three-dimensional and semantic synchronous change detection, the semantic segmentation results of the remote sensing images of the pre-phase and post-phase are analyzed to obtain how the semantics change; the semantic change result and the three-dimensional geometric change result are analyzed to obtain the following types of changes: semantic change, three-dimensional geometric change, semantic change, three-dimensional geometric invariance, semantic invariance, three-dimensional geometric change, semantic invariance, three-dimensional geometric invariance.
[0226] The embodiment also provides a three-dimensional and semantic synchronous change detection system, comprising:
[0227] One or more processors;
[0228] A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional and semantic synchronous change detection method.
[0229] Because the fusion knowledge graph realizes the construction of a three-dimensional change detection data set of multi-modal data and serves the three-dimensional change detection of a neural network, it has necessary practical significance; on the basis of a generative adversarial network and contrastive learning, a positive and negative generator and a semantic detector are designed to extract three-dimensional change semantic information, and through the process of game, the three-dimensional change detection result is intelligently optimized, three-dimensional change detection under a small sample condition is realized, and it has important theoretical and application significance, the three-dimensional and semantic synchronous change detection network of the fusion knowledge graph and the generative contrastive adversarial network is first proposed in the present application, and the functions are as follows:
[0230] (1) Construct a multi-modal three-dimensional change detection data set, construct a three-dimensional change detection data set based on multi-dimensional multi-modal data using a knowledge graph, provide sufficient data support for three-dimensional and semantic synchronous change detection based on a neural network, and provide an accurate small sample three-dimensional and semantic synchronous change detection data set;
[0231] (2) Establish a three-dimensional and semantic change detection generation contrast adversarial network, realize self-supervised three-dimensional and semantic synchronous change detection under small sample conditions, based on the idea of contrast learning, realize multi-branch input, multi-task network construction, construct momentum contrast encoder to extract geometric semantic information of change detection; based on the design of adversarial loss function of generative adversarial network, the semantic information of accurate three-dimensional feature change or not, how to change is mined through the game process, and then the automatic and intelligent three-dimensional and semantic synchronous change detection is realized;
[0232] (3) Realize self-supervised three-dimensional and semantic change detection under small sample conditions, based on knowledge graph, use deep learning large model to assist semantic segmentation to generate small sample three-dimensional and semantic synchronous change detection interpretation label, and realize intelligent optimization of label through the game process of generative adversarial network;
[0233] It should be understood that the above description of the preferred embodiments is more detailed, and therefore cannot be considered as a limitation on the scope of patent protection of the present application. Those skilled in the art can make substitutions or modifications without departing from the scope of the present application, which falls within the scope of protection of the present application. The scope of protection of the present application shall be subject to the appended claims.
Claims
1. A method for detecting changes in three dimensions and semantics, characterized in that, The method comprises the following steps: Step 1: obtaining a multi-modal three-dimensional and semantic change detection small sample data set, including three-dimensional geometric data, attribute data and text data of two time phases; Step 2: respectively pre-processing the three-dimensional geometric data, attribute data and text data of the two time phases to obtain geometric data sets, attribute data sets and text data sets of the two time phases; Step 3: based on the geometric data sets and attribute data sets of the two time phases, constructing a three-dimensional and semantic synchronous change detection data set, including semantic change, three-dimensional geometric change data, semantic change, three-dimensional geometric invariant data, semantic invariant, three-dimensional geometric change data and semantic invariant, three-dimensional geometric invariant data; Step 4: based on a knowledge graph, constructing a positive and negative generator, using a large model to assist in semantic segmentation of the three-dimensional and semantic synchronous change detection data set, obtaining positive samples and negative samples of three-dimensional and semantic change detection, and then expanding the three-dimensional and semantic synchronous change detection data set to obtain three-dimensional geometric change labels and semantic change labels; including geometric real change labels, geometric false alarm error labels and geometric missed detection error labels, semantic real change labels, semantic false alarm error labels and semantic missed detection error labels; Step 5: based on the idea of contrast learning, constructing a three-dimensional and semantic momentum contrast encoder, inputting positive samples and negative samples of three-dimensional data and attribute data of two time phases, and pre-training the three-dimensional and semantic momentum contrast encoder; Step 6: based on contrast learning, constructing a semantic detector, respectively establishing a three-dimensional momentum contrast encoder, a semantic momentum contrast encoder and a semantic encoder, respectively encoding three-dimensional data, attribute data and semantic text data of two time phases, and then performing decoder decoding, loss calculation and gradient update operations; Step 7: based on the loss function of the adversarial network, realizing the game of the positive and negative generator and the semantic detector, and then completing the self-supervised three-dimensional and semantic synchronous change detection under the condition of small samples; whether the generated adversarial game meets the judgment condition of iteration termination is judged, if the iteration termination condition is met, step 8 is executed; otherwise, step 4 is executed; Step 8: three-dimensional and semantic synchronous change detection information mining, the three-dimensional momentum contrast encoder, the semantic momentum contrast encoder and the semantic encoder respectively obtain three-dimensional geometric change, two-time-phase semantic segmentation results and semantic change results; Step 9: three-dimensional and semantic synchronous change detection result analysis, further analyzing the semantic information mined by the three-dimensional and semantic synchronous change detection to obtain semantic change, three-dimensional geometric change, semantic change, three-dimensional geometric invariance, semantic invariance, three-dimensional geometric change and semantic invariance, three-dimensional geometric invariance results.
2. The method of claim 1, wherein: In step 2, the three-dimensional point cloud data and digital surface model data DSM of the two time phases are geometrically corrected and radiometrically corrected to obtain a geometric data set; The geometric correction is from image to map correction, first selecting a reference plane, then automatically matching to obtain corresponding points, eliminating mis-matching points, then performing orthographic correction, and finally generating a digital orthographic map DOM; The radiometric correction is completed by an empirical model method, an atmospheric correction method or a ground surface reflection correction method.
3. The method of claim 1, wherein: In step 2, the text dataset is based on the geosciences knowledge graph to obtain the semantic and topological relationship of the ground object information of two time phases required for three-dimensional and geometric synchronous change detection, and is stored in the text dataset; The geosciences knowledge graph is a tree structure, and the leaf nodes represent the topological and spatial relationship between ground objects. The construction of the knowledge graph includes the following steps: (1) Geosciences knowledge extraction, extracting the semantic information of geographical scenes and geographical features contained in geometric and attribute data, obtaining entity, attribute, and relationship elements; (2) Geosciences knowledge fusion, extracting and analyzing the semantic information of spatial patterns, evolution processes, interaction mechanisms, and temporal and spatial distribution contained in geometric and attribute data to generate graph units; (3) Geosciences knowledge reasoning, knowledge reasoning from the perspectives of three-dimensional geometric features, spectral features, texture features, or topological features to form a small sample text dataset.
4. The method of claim 1, wherein: In step 4, the positive and negative generator uses the semantic and topological information of different ground object types provided by the geosciences knowledge graph to establish a positive and negative generator, and establishes intelligent interpretation labels for the geometric and semantic information of multi-modal data to generate positive samples and negative samples; the positive samples include geometric real change labels, geometric real invariant labels, semantic real change labels, and semantic real invariant labels; The negative samples include geometric false positive error labels, geometric false negative error labels, semantic false positive error labels, and semantic false negative error labels.
5. The method of claim 4, wherein: The positive and negative generator based on the geosciences knowledge graph generates positive samples, and the specific implementation steps are as follows: Step 4.1.1: geometric transformation, the three-dimensional and semantic synchronous change detection dataset contains three-dimensional geometric data of time phase 1 and time phase 2, image rotation and image mirroring processing are performed on the two time phase point clouds and images, and geometric sample enhancement is performed on the three-dimensional point cloud and the two time phase images; Step 4.1.2: radiation enhancement, the three-dimensional and semantic synchronous change detection dataset contains attribute data of time phase 1 and time phase 2, histogram matching, histogram stretching, and mixed ground object addition processing are performed on the two time phase images to perform radiation sample enhancement; The positive and negative generator based on the geosciences knowledge graph generates negative samples, and the specific implementation steps are as follows: Step 4.2.1: generate geometric correction negative samples; (1) generate single pixel image registration error; generate registration error by registering the positive samples generated by geometric transformation; (2) generate multi-pixel image registration error; based on the geosciences knowledge graph, the mis-matching error caused by the overall shift of the image edge of the two images is simulated through morphological processing based on the positive samples generated by geometric transformation; Step 4.2.2: generate radiation correction negative samples; add radiation noise to the positive samples generated by radiation enhancement to simulate pseudo changes caused by shadows.
6. The method of claim 1, wherein: The specific implementation of step 5 includes the following sub-steps: Step 5.1: establish a momentum contrast encoder based on the idea of contrast learning to realize the pre-training process of the downstream task semantic detector construction; Step 5.1.1: the positive sample and the negative sample are respectively input; the positive sample and the negative sample are respectively input into the DSM-based momentum contrast encoder and the image-based momentum contrast encoder, and are respectively trained to obtain the three-dimensional momentum contrast encoder and the semantic momentum contrast encoder; wherein the input of the three-dimensional momentum contrast encoder is the positive sample and the negative sample of the multi-modal geometric height of two time phases; the input of the semantic momentum contrast encoder is the positive sample and the negative sample of the multi-modal remote sensing image of two time phases; Step 5.1.2: initialize the three-dimensional momentum contrast encoder 3DEncoder and the semantic momentum encoder ImageEncoder; in the pre-training process of the three-dimensional and semantic momentum contrast encoder, 3DEncoder and ImageEncoder correspond to a momentum encoder 3DMomentumEncoder and ImageMomentumEncoder respectively, and the parameters of 3DEncoder and 3DMomentumEncoder are consistent, and the parameters of ImageEncoder and ImageMomentumEncoder are consistent during the training process; Step 5.1.3: repeatedly train the three-dimensional momentum contrast encoder and the semantic momentum contrast encoder; (1) randomly sample samples 3Dx and Imagex from the positive and negative samples input into the three-dimensional and semantic momentum contrast encoder, the randomly sampled samples 3Dx and Imagex are called anchor points, and the remaining positive and negative samples except the anchor points are called queues; (2) Random data augmentation is performed on the anchor points, and the anchor points and the data after data augmentation are input into the encoder 3D Encoder and ImageEncoder, respectively, and the feature vectors 3Dx q and Imagex q are calculated, respectively. (3) 3D Momentum Encoder, Image Momentum Encoder input samples and calculate feature vectors 3D x k and Image x k ; the samples input by the 3D Momentum Encoder, Image Momentum Encoder are samples after another data augmentation of anchor point data and samples in a queue; the samples after another data augmentation of the anchor point are used as positive samples, and form a positive sample pair with the samples after data augmentation of the anchor point; the samples in the queue are used as negative samples; (4) Compute the inner product of the anchor and the corresponding positive sample as the similarity, the inner product is the feature vector 3Dx q and 3Dx k corresponding multiplication, Imagex q and Imagex k corresponding multiplication; (5) calculate the inner product of the anchor point and all negative samples; (6) combine the inner product of the anchor point and the positive sample and the inner product of the negative sample; (7) calculate the loss of the anchor point and the positive sample to distinguish the correct sample and the error sample; the loss can be divided into supervised loss function calculation and unsupervised loss function calculation according to whether the label data of the positive and negative samples is input; the more similar the anchor point and the positive sample are, the smaller the loss value is; the closer the anchor point and the negative sample are, the larger the loss value is; (8) update the parameters of the encoders 3DEncoder, ImageEncoder, 3DMomentumEncoder and ImageMomentumEncoder; Step 5.1.4: momentum contrast encoder result output, output the three-dimensional momentum contrast encoder and the semantic momentum encoder trained by the momentum contrast encoder to the subsequent task.
7. The method of claim 1, wherein: The specific implementation of step 6 includes the following sub-steps: Step 6.1: Constructing the contrastive learning semantic decoder, the three-dimensional and semantic feature vectors 3Dx q and Imagex q decoded into the original three-dimensional and semantic samples 3Dx and Imagex, the semantic decoder is trained to extract the three-dimensional and semantic features obtained by the three-dimensional momentum contrast encoder and the semantic momentum encoder; Step 6.2: under the condition of small sample, combine the three-dimensional and semantic encoders and the contrast learning semantic decoder obtained by the momentum contrast encoder to construct a contrast learning semantic detector. Step 6.2.1: Data input, including remote sensing image data of the previous phase, geometric elevation data of the previous phase, text information data of the previous phase, remote sensing image data of the later phase, geometric elevation data of the later phase, and text information data of the later phase; the remote sensing image data of the previous and later phases is used to obtain remote sensing image segmentation results of the previous and later phases, the text information data of the previous and later phases is used to obtain semantic change results, and the geometric elevation data of the previous and later phases is used to obtain three-dimensional geometric change results, and the four semantic detectors simultaneously obtain the identification results; Step 6.2.2: Call the pre-trained encoder 3DEncoder and ImageEncoder; Step 6.2.3: Contrastive learning semantic decoder decoding; Step 6.2.4: Input sample labels and calculate loss value, and perform gradient update; Step 6.2.5: Semantic detector result analysis, the identification results of the semantic detector include the semantic segmentation results of the remote sensing images of the previous and later phases, the semantic change results, and the three-dimensional geometric change results.
8. The method of claim 1, wherein: The specific implementation of step 8 includes the following sub-steps: Step 8.1: Detecting three-dimensional change information based on a contrastive learning semantic detector, and mining three-dimensional change information through small sample self-supervised learning; Step 8.2: Obtaining semantic information of geometric and semantic changes through adversarial game, calculating a loss function of a generated contrastive adversarial network, and then obtaining whether three-dimensional geometric and semantic information changes and how it changes; The results of the three-dimensional change detection include: semantic segmentation results of the remote sensing images of the previous phase, semantic segmentation results of the remote sensing images of the later phase, semantic change results, and three-dimensional geometric change results; and the three-dimensional change detection results are updated to a three-dimensional change detection dataset; Step 8.3: Dynamic updating of the three-dimensional and semantic synchronous change detection dataset; Step 8.3.1: Sample label updating, small sample label dataset based on a knowledge graph, and positive and negative samples generated by a positive and negative generator; Step 8.3.2: Geometric and semantic detection results of three-dimensional change detection, while generating detection results of three-dimensional change detection, realizing real-time dynamic updating of the dataset.
9. The method of claim 1, wherein: The specific implementation of step 9 includes the following sub-steps: Step 9.1: Three-dimensional change detection network performance analysis, judging the convergence and robustness of the generated contrastive adversarial network by calculating the network loss value of the three-dimensional and semantic synchronous change detection; Step 9.2: Quantitative evaluation of three-dimensional and semantic synchronous change detection; Step 9.3: Qualitative evaluation of three-dimensional and semantic synchronous change detection; The results of the three-dimensional and semantic synchronous change detection, analyzing the semantic segmentation results of the remote sensing images of the previous and later phases to obtain how the semantics change, and analyzing the semantic change results and the three-dimensional geometric change results to obtain the following types of changes: semantic change, three-dimensional geometric change, semantic change, three-dimensional geometric invariance, semantic invariance, three-dimensional geometric change, and semantic invariance, three-dimensional geometric invariance.
10. A three-dimensional and semantic synchronous change detection system, characterized in that: Comprise: One or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the three-dimensional and semantic change detection method of any one of claims 1 to 9.
Citation Information
Patent Citations
High-resolution remote sensing image weak supervision building change detection method guided by prior semantic knowledge
CN113936217A
High-resolution remote sensing image semantic change detection method based on binary change detection contrast learning
CN116524346A