A method for urban scene change detection based on twin graph neural network
By applying twin graph neural network and feature interaction module in remote sensing image scene-level change detection, combining the node attribute characteristics and graph structure of the two-time phase, deep structural features are extracted and graph similarity modeled, the problem of difficulty in using image structure information in the existing technology is solved, and the accuracy and accuracy of change detection are improved.
Patent Information
- Application Number
- CN202311196705.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-16
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2043-09-16
AI Technical Summary
The existing remote sensing image scene-level change detection methods are difficult to make full use of image structure information, resulting in missed and missed detection in urban land use change detection.
The urban market scene change detection method based on twin graph neural network is adopted. Through the twin graph neural network and feature interaction module, the node attribute characteristics and graph structure of the two-time phase are combined, and the deep structural features of the scene are extracted, and the graph similarity modeling and node feature interaction are improved to improve the accuracy of change detection.
It effectively overcomes the problem that traditional methods cannot fully utilize image structure information, improves the accuracy of scene-level change detection, reduces missed and missed detection, and provides technical support for efficient acquisition of urban land use change information.
Smart Images

Figure CN117173575B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image scene-level change detection, and particularly to an urban scene change detection method based on a twin graph neural network. Background Art
[0002] The changes in the surface environment are closely related to human activities. Based on remote sensing images, information on land cover and land use changes can be obtained quickly and accurately, which is of great significance for aspects such as urban and rural development planning, natural disaster assessment, environmental dynamic monitoring, and natural resource management. Generally speaking, according to the granularity of the basic analysis unit, change detection methods can be divided into pixel-level, object-level, and scene-level. Usually, the first two change detection methods can only detect changes in the landscape, such as the change from vegetation to buildings, and cannot identify semantic or functional changes at the scene level, such as the transformation of an industrial area into a residential area. However, most of the existing change detection methods focus on the pixel and object levels. With the advancement of high-quality urbanization, it is very important to extract urban land use change information at the scene level for explaining urban functional zoning, grasping the urbanization trend, and making sustainable development decisions.
[0003] There is relatively little research on the existing remote sensing image scene-level change detection, and most of them focus on using traditional methods for analysis, or comparing feature distances after using convolutional neural networks to extract deep features. These methods mainly focus on processing the attribute information of the scene, while ignoring the extraction and utilization of the topological structure information among the various elements within the scene. The graph neural network is a deep learning model constructed for structural data, suitable for processing non-Euclidean space data, and has excellent structural feature abstraction ability. It can not only effectively mine the deep structural information of graph data, but also automatically fuse the rich attribute information of graph node elements. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide an urban scene change detection method based on a twin graph neural network, which overcomes the problem that traditional methods and convolutional neural networks cannot fully utilize the image structure information, and provides technical support for efficiently obtaining urban land use change information in remote sensing images.
[0005] To achieve the above purpose, the present invention adopts the following technical scheme: An urban scene change detection method based on a twin graph neural network, based on a twin graph neural network and a feature interaction module, includes the following steps:
[0006] Step S1: Obtain high-resolution remote sensing images of the study area at two time phases, and perform preprocessing on the images, including radiometric correction, orthorectification, atmospheric correction, image fusion, image registration, and image resampling operations;
[0007] Step S2: Perform sliding window cropping on the preprocessed image to obtain scene samples of a specified size; select an image segmentation algorithm to perform superpixel segmentation on the scene samples to obtain the scene superpixel segmentation result; construct a graph structure for the scene samples based on the superpixel segmentation result;
[0008] Step S3: Based on the graph structure constructed in Step S2, extract the attribute features of each node by combining the original scene image; use the attribute features and graph structure of the nodes in two time phases as inputs, and utilize a siamese graph neural network to fuse the attribute information of the nodes to extract the deep structure features of the scene;
[0009] Step S4: Perform graph similarity modeling and node feature interaction based on the structure features of the graph nodes extracted in Step S3; specifically, for graph similarity modeling, first perform graph pooling on the node features through an attention mechanism based on the inner product to output a graph vector representation, and use a Neural Tensor Network to model the similarity of the graph vector representations in two time phases to obtain the graph-level interaction result; for the node interaction operation, construct a node feature interaction matrix by calculating the inner product of the node features in two time phases for each node, and statistically obtain the node interaction result by calculating its histogram features; concatenate the node-level interaction result and the graph-level interaction result for change detection;
[0010] Step S5: Train the model by combining the real scene sample labels, and use the binary cross-entropy loss function to calculate the difference between the model prediction result and the label to guide the tuning of the model parameters;
[0011] Step S6: Use the model weights obtained from training for change detection of the scene samples obtained in Step S2, and concatenate the detection results to complete the detection of the entire region.
[0012] In a preferred embodiment, Step S2 specifically includes the following steps:
[0013] Step S21: Use sliding window cropping to crop the remote sensing image of the study area into scene pictures of a fixed size, select an image segmentation algorithm to perform superpixel segmentation on the scene pictures, and obtain an ideal segmentation result through parameter adjustment;
[0014] Step S22: Based on the superpixel segmentation result, regard each superpixel object in the scene as a graph node, and construct a regional adjacency graph. Specifically, for superpixel objects that are spatially adjacent, their adjacency relationship is defined as 1, and the adjacency relationship between non-adjacent objects is defined as 0. The calculation formula for the adjacency matrix is expressed as follows:
[0015]
[0016] where A ij represents the adjacency relationship between object i and object j.
[0017] In a preferred embodiment, step S3 specifically includes the following steps:
[0018] Step S31: Combine the original images of the dual-temporal scene, and extract the attribute features of the graph nodes as the feature input of the Siamese graph neural network; The attribute feature extraction can be carried out by two methods: feature engineering and deep learning. Among them, the feature engineering method refers to manually designing image features using expert knowledge, such as spectral, texture, geometric and other features, as well as low-level visual features such as Scale-Invariant Feature Transform (SIFT) and Histogram of Oriented Gradients (HOG); The deep learning method mainly automatically extracts high-level semantic features based on a pre-trained deep neural network.
[0019] Step S32: Use the Graph Sampling and Aggregation Network (GraphSAGE) as the backbone network for information transmission update and structural feature extraction, and construct a structural feature encoder with two shared-weight branches. Each encoder stacks multiple graph neural structures, and input the attribute features extracted in step S31 into the structural feature encoder to extract the deep structural features of the scene.
[0020] In a preferred embodiment, step S4 specifically includes the following steps:
[0021] Step S41: The node interaction module first calculates the product of two matrices, and performs a Sigmoid mapping on the calculation result to obtain a node feature interaction matrix. Each element represents the similarity of the corresponding node. After specifying the number of intervals, its histogram feature is statistically calculated as the node feature interaction result.
[0022] Step S42: First, use the inner product-based attention mechanism to perform graph pooling on the graph node features to output the graph vector representation. Specifically, calculate the mean of the dual-temporal node features respectively, and calculate the inner product with the feature mean for each node, and perform a Sigmoid mapping on the calculation result. Use the mapping result to represent the similarity between each node feature and the mean; The closer the mapping result is to 1, the higher the correlation, and vice versa.
[0023] Step S43: Use the neural tensor network to perform correlation modeling on the graph vector representation obtained in step S42. The calculation formula of the neural tensor network is:
[0024]
[0025] In the formula, e 1 and e 2 respectively represent the dual-temporal graph vector representations, g represents the similarity modeling result of the dual-temporal graph vector representations, f is the activation function, W is the tensor slice, used to model the similarity between e 1 and e 2 , K is a hyperparameter used to set the number of tensor slices, V is the weight matrix, and b is the bias vector.
[0026] Step S44: Concatenate the interaction results of Step S41 and Step S43 for subsequent change detection.
[0027] In a preferred embodiment, Step S5 specifically includes the following steps:
[0028] Step S51: Based on the image data preprocessed in Step S1, perform label making. The common label making method is manual interpretation; collect scene change sample labels, and divide the data into a training set, a validation set, and a test set according to a certain proportion;
[0029] Step S52: Set hyperparameters, set the initial learning rate Lr, the optimizer type O, the training batch N of the model, and the number of model iterations R;
[0030] Step S53: Use the binary cross-entropy loss function (BCE Loss) to calculate and evaluate the difference between the model prediction result and the true label. The definition of BCE Loss is as follows:
[0031] L BCE = -p(x)·log((q(x))+(1 - p(x))·log(1 - q(x))
[0032] By minimizing the loss function, the model can adjust its own parameters to improve the prediction accuracy; the goal of training is to continuously optimize the loss function to make it reach the minimum value.
[0033] In a preferred embodiment, Step S6 specifically includes the following steps:
[0034] Step S61: After the parameter setting is completed, use the dataset obtained in Step S5 to train the Siamese multi-graph neural scene change detection model and save the optimal model weights;
[0035] Step S62: Use the weight file obtained in Step S61 to perform change detection on the scene samples obtained in Step S1, and concatenate the detection results to complete the detection of the entire area.
[0036] Compared with the prior art, the present invention has the following beneficial effects: (1) Utilize the superior structure feature abstraction ability of the Siamese graph neural network to mine the deep structure information of the scene and fuse the attribute information of the graph node elements, solving the problem that traditional methods cannot fully utilize the image structure information; (2) Integrate the graph interaction module and the node interaction module, supplementing fine-grained node-level information while modeling the similarity of the two-temporal scenes, overcoming the problems of missed detection and false detection prone to traditional methods, further improving the accuracy of scene-level change detection, and providing technical support for efficiently obtaining urban land use change information in remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 Schematic diagram of the method flow of the preferred embodiment of the present invention;
[0038] Figure 2 Structural diagram of the twin graph neural network of the preferred embodiment of the present invention;
[0039] Figure 3 Detection result diagram of the scene change in the research area of the preferred embodiment of the present invention, where (a) binary scene change, (b) change area, and (c) detailed result display. Detailed implementation manners
[0040] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0041] It should be noted that the following detailed description is exemplary and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0042] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0043] As Figure 1 shown, this embodiment provides a method for detecting urban scene changes based on a twin graph neural network, including the following steps:
[0044] Step S1: Obtain dual-temporal high-resolution remote sensing images of the research area, and perform preprocessing on the images, including operations such as radiometric correction, orthorectification, atmospheric correction, image fusion, image registration, and image resampling;
[0045] Step S2: Perform sliding window cropping on the preprocessed images to obtain scene samples of a specified size; select an image segmentation algorithm to perform superpixel segmentation on the scene samples to obtain the scene superpixel segmentation result; construct a graph structure for the scene samples based on the superpixel segmentation result;
[0046] Step S3: Based on the graph structure constructed in Step S2, extract the attribute features of each node by combining the original scene images; use the attribute features and graph structure of the dual-temporal nodes as inputs, and use the twin graph neural network to fuse the attribute information of the nodes to extract the deep structure features of the scene;
[0047] Step S4: Perform graph similarity modeling and node feature interaction based on the structural features of the graph nodes extracted in Step S3; specifically, for graph similarity modeling, first perform graph pooling on the node features through an attention mechanism based on the inner product to output a graph vector representation, and use a Neural Tensor Network to model the similarity of the dual-temporal graph vector representations to obtain a graph-level interaction result; for node interaction operations, construct a node feature interaction matrix by calculating the inner product of the dual-temporal node features for each node, and statistically obtain the node interaction result by calculating its histogram features; concatenate the node-level interaction result and the graph-level interaction result for change detection;
[0048] Step S5: Train the model in combination with the real-scene sample labels, and use the binary cross-entropy loss function to calculate the difference between the model prediction result and the label to guide the tuning of the model parameters;
[0049] Step S6: Use the model weights obtained from training for the change detection of the scene samples obtained in Step S2, and concatenate the detection results to complete the detection of the entire region.
[0050] In this embodiment, Step S2 specifically includes the following steps:
[0051] Step S21: Use a sliding window to crop the remote sensing image of the study area into scene pictures of 256×256 pixels, and select eCognition multi-scale segmentation to perform superpixel segmentation on the scene pictures. Through parameter tuning, an ideal segmentation result is obtained when the shape factor is 0.2, the compactness is 0.4, and the segmentation scale is 100;
[0052] Step S22: Based on the superpixel segmentation result, regard each superpixel object in the scene as a graph node, and construct a regional adjacency graph. Specifically, for spatially adjacent superpixel objects, their adjacency relationship is defined as 1, and the adjacency relationship between non-adjacent objects is defined as 0. The calculation formula of the adjacency matrix is expressed as follows:
[0053]
[0054] where A ij represents the adjacency relationship between object i and object j.
[0055] In this embodiment, Step S3 specifically includes the following steps:
[0056] Step S31: Combine the original images of the dual-temporal scene and extract the attribute features of the graph nodes as the feature input of the siamese graph neural network. The extraction of attribute features can be achieved through two methods: feature engineering and deep learning. Feature engineering refers to the manual design of image features using expert knowledge, such as spectral, texture, geometric features, as well as low-level visual features like Scale-Invariant Feature Transform (SIFT) and Histogram of Oriented Gradients (HOG). Deep learning mainly involves automatically extracting high-level semantic features based on a (pre)-trained deep neural network.
[0057] Step S32: Use Graph Sampling and Aggregation Network (GraphSAGE) as the backbone network for information passing update and structural feature extraction to construct a structural feature encoder with two shared-weight branches. Other common graph neural networks such as Graph Convolutional Network (GCN), Graph Isomorphism Network (GIN), and Graph Attention Network (GAT) can also be used to construct the structural feature encoder. In this example, GraphSAGE is selected as the backbone network for information passing update and structural feature extraction, and three identical graph neural structures are stacked in each encoder to input the features extracted in Step S31 into the network to mine deep structural information.
[0058] In this embodiment, Step S4 specifically includes the following steps:
[0059] Step S41: The node interaction module first calculates the product of two matrices (the feature matrix of the latter temporal phase needs to be transposed before calculating the product), and performs a Sigmoid mapping on the calculation result to obtain the node feature interaction matrix, where each element represents the similarity of the corresponding node. The specified number of intervals is 16, and its histogram feature is counted as the node feature interaction result.
[0060] Step S42: First, use the inner-product-based attention mechanism to perform graph pooling on the graph node features to output the graph vector representation. Specifically, calculate the mean of the dual-temporal node features respectively, and calculate the inner product with the feature mean for each node, and perform a Sigmoid mapping on the calculation result. Use the mapping result to represent the similarity between each node feature and the mean. The closer the mapping result is to 1, the higher the correlation, and vice versa.
[0061] Step S43: Use the neural tensor network to perform correlation modeling on the graph vector representation obtained in Step S42. The calculation formula of the neural tensor network is:
[0062]
[0063] In the formula, e 1 and e 2 respectively represent the dual-temporal graph vector representations, g represents the similarity modeling result of the dual-temporal graph vector representations, f is the activation function, and W is the tensor slice used to model e 1and e 2 The similarity of, where K is a hyperparameter used to set the number of tensor slices, V is the weight matrix, and b is the bias vector. In this example, the value of K is set to 64.
[0064] Step S44: Concatenate the interaction results of Step S41 and Step S43 for subsequent change detection.
[0065] In this embodiment, Step S5 specifically includes the following steps:
[0066] Step S51: Based on the image data preprocessed in Step S1, perform label making. A common label making method is manual interpretation. Collect scene change sample labels and divide the data into a training set, a validation set, and a test set in a ratio of 6:2:2;
[0067] Step S52: Set the initial learning rate to 0.001, use the Adam optimizer as the optimizer of the model, set the training batch of the model to 16, and set the number of model iterations to 100;
[0068] Step S53: Use the binary cross-entropy loss function (BCE Loss) to calculate and evaluate the difference between the model prediction result and the true label. The BCE Loss is defined as follows:
[0069] L BCE = -p(x)·log((q(x))+(1 - p(x))·log(1 - q(x))
[0070] By minimizing the loss function, the model can adjust its own parameters to improve the accuracy of prediction. The goal of training is to continuously optimize the loss function to make it reach the minimum value.
[0071] In this embodiment, Step S6 specifically includes the following steps:
[0072] Step S61: After the parameter settings are completed, use the dataset obtained in Step S5 to train the Siamese multi-graph neural scene change detection model and save the optimal model weights;
[0073] Step S62: Use the weight file obtained in Step S61 to perform change detection on the scene samples obtained in Step S1, and concatenate the detection results to complete the detection of the entire region.
[0074] In this embodiment, a partial urban area of Fuzhou City, Fujian Province is used as the study area, and two-phase GF-2 high-resolution remote sensing image data of the study area in 2017 and 2020 are collected. After preprocessing, its spatial resolution is standardized to 1m, and the red, green, and blue three bands are used. After manual interpretation, this embodiment has obtained a total of 1003 pairs of scene samples and corresponding label data with a size of 256×256 pixels, including 789 pairs of unchanged samples and 214 pairs of changed samples, which are divided into training set, validation set, and test set according to the ratio of 6:2:2 for model training, validation, and testing.
[0075] Figure 2 The structural diagram of the Siamese graph neural network constructed in this embodiment is shown. It can be seen from the figure that the model is mainly composed of a structural feature extractor and a feature interaction module. The input of the structural feature extractor is the original node attribute feature matrix, which is composed of three layers of repeated graph neural networks and activation functions. After the node feature matrix is aggregated and updated by the graph structure feature extractor, the graph structure information is incorporated, and it is used as the input of the node interaction module and the aggregated output is the graph feature vector. Specifically, the node interaction module first calculates the product of two matrices (the feature matrix of the latter time phase needs to be transposed before calculating the product), and performs a Sigmoid mapping on the calculation result to obtain the node feature interaction matrix, where each element represents the similarity of the corresponding node; the specified number of intervals is 16, and the histogram feature of the node interaction matrix is counted as the node feature interaction result; on the other hand, the node features output by the graph neural network are aggregated through the inner product-based attention mechanism module to output the graph feature vector of the scene, and the graph feature vectors of the two time phases are output as the graph-level similarity vector after passing through the graph feature interaction module. Specifically, 64 weight slices are set to model the similarity of the graph feature vectors of the two time phases, and 1 conventional single-layer perceptron is combined to achieve graph-level feature interaction. The node-level interaction result and the graph-level interaction result are spliced for detection and classification.
[0076] Figure 3 The partial experimental result diagram of this embodiment on the Fuzhou GF-2 image is shown. Figure 3 (a) and Figure 3 (b) respectively show the binary change result and the change area extraction result of the study area; Figure 3 (c) shows the detailed information of the result. It can be seen from the figure that whether it is the scene change in a large area or the scene change in a small area, the detection results obtained by the proposed method have a high consistency with the ground truth, and the phenomena of missed detection and false detection are less. The experimental results show the effectiveness of the proposed method for urban scene change detection.
[0077] The twin graph neural scene change detection model proposed by the present invention combines the graph neural network algorithm and the feature interaction module, overcoming the problem that it is difficult for traditional methods to fully extract the image structure information; by integrating the graph interaction module and the node interaction module, the model can supplement fine-grained node-level information while modeling the similarity of the dual-temporal scene graph-level features, overcoming the problems of missed detection and false detection prone to traditional methods, and further improving the accuracy of scene-level change detection, providing technical support for efficiently obtaining urban land use change information in remote sensing images.
[0078] The above are only the preferred embodiments of the present invention, and all equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.
Claims
1. A method for detecting urban scene changes based on twin graph neural networks, characterized in that Based on the twin graph neural network and feature interaction module, the following steps are included: Step S1: Obtain dual-temporal high-resolution remote sensing images of the study area and preprocess the images, including radiation correction, orthorectification, atmospheric correction, image fusion, image registration, and image resampling operations; Step S2: Perform sliding window cropping on the preprocessed image to obtain a scene sample of a specified size; select an image segmentation algorithm to perform superpixel segmentation on the scene sample to obtain a scene superpixel segmentation result; and construct a graph structure for the scene sample based on the superpixel segmentation result; Step S3: Based on the graph structure constructed in step S2, the attribute features of each node are extracted in combination with the original scene image; the attribute features of the nodes in the dual time phase and the graph structure are used as input, and the attribute information of the nodes is fused using the twin graph neural network to extract the deep structural features of the scene; Step S4: Graph similarity modeling and node feature interaction are performed based on the structural features of the graph nodes extracted in step S3; specifically, the graph similarity modeling first performs graph pooling on the node features through an inner product-based attention mechanism to output a graph vector representation, and uses a neural tensor network to model the similarity of the bi-phase graph vector representation to obtain a graph-level interaction result; the node interaction operation constructs a node feature interaction matrix by calculating the inner product of the bi-phase node features node by node, and obtains the node interaction result by counting its histogram features; the node-level interaction results and the graph-level interaction results are spliced to detect changes; Step S5: training the model in combination with real scene sample labels, and using the binary cross entropy loss function to calculate the difference between the model prediction result and the label to guide model parameter tuning; Step S6: Using the model weights obtained through training to detect changes in the scene samples obtained in step S2, and splicing the detection results to complete the detection of the entire area; Step S2 specifically includes the following steps: Step S21: using sliding window cropping to crop the remote sensing image of the study area into a scene image of a fixed size, selecting an image segmentation algorithm to perform superpixel segmentation on the scene image, and obtaining a relatively ideal segmentation result by adjusting parameters; Step S22: Based on the superpixel segmentation result, each superpixel object in the scene is regarded as a graph node, and a regional adjacency graph is constructed. Specifically, for spatially adjacent superpixel objects, their adjacency relationship is defined as 1, and the adjacency relationship between non-adjacent objects is defined as 0. The adjacency matrix calculation formula is expressed as follows: In the formula, A ij Represents the adjacency relationship between object i and object j; Step S3 specifically includes the following steps: Step S31: Combine the original image of the dual-phase scene and extract the attribute features of the graph nodes as the feature input of the twin graph neural network; the attribute feature extraction is carried out through two methods: feature engineering and deep learning. The feature engineering method refers to the use of expert knowledge to manually design image features and HOG low-level visual features. The expert knowledge manually designed image features include spectrum, texture, and geometric features. The HOG low-level visual features include scale-invariant features SIFT and directional gradient histogram features. The deep learning method is mainly based on a pre-trained deep neural network to automatically extract high-level semantic features. Step S32: Using the graph sampling and aggregation network GraphSAGE as the backbone network for information transmission update and structural feature extraction, a structural feature encoder with two shared weight branches is constructed, each encoder is stacked with a multi-layer graph neural structure, and the attribute features extracted in step S31 are input into the structural feature encoder to extract the deep structural features of the scene; Step S4 specifically includes the following steps: Step S41: The node interaction module first calculates the product of the two matrices, and performs Sigmoid mapping on the calculation result to obtain a node feature interaction matrix, in which each element represents the similarity of the corresponding node. After specifying the number of intervals, its histogram features are counted as the node feature interaction result; Step S42: First, use the inner product-based attention mechanism to perform graph pooling on the graph node features and output a graph vector representation. Specifically, calculate the mean of the bi-phase node features respectively, and calculate the inner product with the feature mean node by node, perform Sigmoid mapping on the calculation results, and use the mapping results to represent the similarity between each node feature and the mean; the closer the mapping result is to 1, the higher the correlation, and vice versa; Step S43: Use a neural tensor network to perform correlation modeling on the graph vector representation obtained in step S42. The calculation formula of the neural tensor network is: Wherein, e1 and e2 represent the vector representation of the two-phase image, g represents the similarity modeling result of the vector representation of the two-phase image, f is the activation function, W is the tensor slice used to model the similarity between e1 and e2, K is a hyperparameter used to set the number of tensor slices, V is the weight matrix, and b is the bias vector; Step S44: concatenate the interactive results of step S41 and step S43 for subsequent change detection.
2. According to claim 1, a method for detecting urban scene changes based on twin graph neural network is characterized in that: Step S5 specifically includes the following steps: Step S51: labeling is performed based on the image data preprocessed in step S1. The commonly used labeling method is manual interpretation. Sample labels of scene changes are collected, and the data are divided into a training set, a validation set, and a test set according to a proportion. Step S52: Hyperparameter setting, setting the initial learning rate Lr, optimizer type O, model training batch N and model iteration number R; Step S53: The binary cross entropy loss function (BCE Loss) is used to calculate and evaluate the difference between the model prediction result and the true label. The BCE Loss is defined as follows: L BCE =-p(x)·log((q(x))+(1-p(x))·log(1-q(x)) By minimizing the loss function, the model adjusts its own parameters to improve the accuracy of the prediction; the goal of training is to continuously optimize the loss function to its minimum value.
3. According to claim 1, a method for detecting urban scene changes based on twin graph neural network is characterized in that: Step S6 specifically includes the following steps: Step S61: After the parameter setting is completed, the data set obtained in step S5 is used to train the twin multi-image neural scene change detection model and save the optimal model weights; Step S62: Use the weight file obtained in step S61 to perform change detection on the scene samples obtained in step S1, and splice the detection results to complete the detection of the entire area.