Scene Situation Generation Method Based on Remote Sensing Image Change Detection

By adopting the dual-branch U-net change detection network and graph convolutional neural network based on Swin Transformer in the remote sensing change detection technology, the problems of pseudo-change and missed detection in the existing technology are solved, and the rapid, accurate and comprehensive situation generation of remote sensing scenarios is achieved.

CN118736431BActive Publication Date: 2025-06-13BEIJING SATELLITE INFORMATION ENG RES INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410739940.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2025-06-13
Estimated Expiration
2044-06-07

AI Technical Summary

Technical Problem

Existing remote sensing change detection technology is difficult to effectively solve the problem of mist detection caused by cloud occlusion, pseudo-changes caused by seasonal changes, and the missed detection problems caused by changing objects occupying fewer pixels in large-scale remote sensing scenarios. It relies on a single machine vision understanding and is not very practical.

Method used

A dual-branch U-net change detection network based on Swin Transformer is used to detect changes on high-resolution satellite remote sensing images of different phases. A spatial relationship model is constructed through a graph convolution neural network, edge sets and adjacency matrix are generated, multi-task synchronous training is performed, and the situation of the changing scene is output.

Benefits of technology

It realizes rapid, accurate and comprehensive situation generation of remote sensing scenes with changing land objects over a period of time, improves the situation awareness level of remote sensing changing scenes, and can effectively reduce pseudo-change and missed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118736431B_ABST
    Figure CN118736431B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for generating a scene situation based on remote sensing image change detection, including: acquiring two high-resolution satellite remote sensing images of the same area at different times and performing preprocessing; constructing a dual-branch U-net change detection network based on Swin Transformer to perform change detection on the two high-resolution satellite remote sensing images at different times; modeling the spatial relationship according to the boundary information of the changed ground objects output by the change detection network, constructing a graph convolutional neural network, and generating an edge set and an adjacency matrix; using a manually annotated remote sensing change detection data set to train the graph convolutional neural network to obtain a scene situation generation model based on remote sensing image change detection; using the trained scene situation generation model based on remote sensing image change detection to test the data in the test set to obtain the situation of the remote sensing change image. The present invention makes full use of the rich semantic information of the dual-temporal remote sensing images to realize the automatic generation of the changed scene situation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a method for generating a scene situation based on remote sensing image change detection. Background Art

[0002] Remote sensing image change detection refers to detecting the differences of another phase relative to a reference image with a remote sensing image of one phase as the reference. With the rapid development of earth observation technology, the time resolution of remote sensing imaging has been rapidly improved, providing a rich data source for remote sensing ground object change monitoring. However, there are still problems in the current field of remote sensing change detection, such as pseudo-changes of observed objects caused by cloud cover, seasonal changes, etc., and missed detections caused by the fact that the changed objects occupy fewer pixels in a large-scale remote sensing scene.

[0003] Remote sensing scene situation generation refers to analyzing the differences in ground object changes based on remote sensing images obtained by observations at different times through methods such as change detection, so as to generate the development trend of the scene, which is conducive to quickly extracting information on ground object changes in a large-scale remote sensing scene, completing the transmission of ground object change information, and playing a role in explaining the content of the changed scene, aiming to provide guidance for scientific decision-making in fields such as agriculture, environment, and transportation. With the development of remote sensing technology, the number of remote sensing satellites increases several times every year, and both the time resolution and spatial resolution of remote sensing images have been greatly increased, providing a richer and more reliable data source for scene situation generation. However, the current remote sensing scene situation generation technology still mainly relies on single machine vision understanding, with poor practicability, and there is an urgent need for a fast, accurate, and comprehensive situation generation method based on change detection to improve the situation recognition level of remote sensing changed scenes. Summary of the Invention

[0004] To solve the above technical problems existing in the prior art, the purpose of the present invention is to provide a method for generating a scene situation based on remote sensing image change detection, which can generate a situation for a remote sensing scene with changed ground objects within a period of time.

[0005] To achieve the above invention purpose, the present invention provides a method for generating a scene situation based on remote sensing image change detection, including the following steps:

[0006] Step S1: Obtain two high-resolution satellite remote sensing images of the same area at different times and perform preprocessing;

[0007] Step S2: Construct a dual-branch U-net change detection network based on Swin Transformer to perform change detection on the two high-resolution satellite remote sensing images of different times;

[0008] Step S3: Model the spatial relationship based on the boundary information of the changed ground objects output by the change detection network, construct a graph convolutional neural network, and generate an edge set and an adjacency matrix;

[0009] Step S4: Use the manually annotated remote sensing change detection dataset to train the graph convolutional neural network to obtain a scene situation generation model for remote sensing image change detection;

[0010] Step S5: Use the trained scene situation generation model for remote sensing image change detection to test the data in the test set to obtain the situation of the remote sensing change image.

[0011] According to a technical solution of the present invention, in the step S1, after obtaining two high-resolution satellite remote sensing images of the same area at different times, the images are divided into blocks to form images with a size of 256×256, and preprocessing is performed on the two high-resolution satellite remote sensing images at different times.

[0012] The preprocessing at least includes data augmentation, cloud simulation, season simulation, and environmental noise simulation.

[0013] According to a technical solution of the present invention, in the step S2, it specifically includes:

[0014] Step S21: Perform patch segmentation and linear embedding on the preprocessed images respectively;

[0015] Step S22: Construct a dual-branch U-net change detection network based on Swin Transformer as the backbone network of the change detection part;

[0016] Step S23: Use the dual-branch three-level encoder part of the backbone network to perform hierarchical feature extraction on the dual-temporal remote sensing images respectively and convert them into feature tokens;

[0017] Step S24: Input the highest-level features extracted by the dual-branch encoder into the feature fusion module to fuse the dual-temporal feature maps;

[0018] Step S25: Use the three-level decoder part of the backbone network to fuse the hierarchical features extracted by the encoder in turn, gradually restore the change information, and obtain the change information with both high-level features and low-level features, where the change information includes change categories, change types, and change positions;

[0019] Step S26: Perform the last upsampling, insert a channel attention module to reduce information loss, and output the change detection information.

[0020] According to a technical solution of the present invention, in the step S3, it specifically includes:

[0021] Step S31: Consider the change objects of various ground features as nodes V = [v 1 , v 2 , v 3 , … v n . Take the change type and change status of the ground features as node features, and define the spatial distance between each pair of nodes as edges E = [e 1 , e 2 , e 3 , … e n . Obtain the node set V and the edge set E, and construct the graph neural network G = {V, E};

[0022] Step S32: Introduce the boundary position information box of the changed ground features. box i = [y 1 , x 1 , y 2 , x 2 to model the spatial relationship and normalize the position information of each changed ground feature, generating the edge set E and the adjacency matrix A. Then we have:

[0023]

[0024] e ij = ||v i - v j ||,

[0025]

[0026] where w and h are the width and height of the remote sensing image, e ij is the edge between the changed ground features i and j, and a ij is the element of the adjacency matrix A;

[0027] Step S33: Construct the degree matrix d and the Laplacian matrix L based on the adjacency matrix:

[0028]

[0029] Construct the weight matrix W, and update the node features by combining the Laplacian matrix and the adjacency matrix. The formula for graph convolution is as follows:

[0030]

[0031] where σ is the non-linear activation function, the parameters of W are learned by gradient descent and randomly initialized, and Z is the output;

[0032] Step S34: Introduce the attention mechanism. Assume that the feature vector corresponding to any node v i in the l-th layer of the constructed remote sensing image scene content graph is h i, after the feature aggregation of the representation information of the edges of adjacent nodes with the attention mechanism as the core, the new representation information is h' i ,

[0033] Set the central node as v i , then assume the neighbor nodes v j to v i weight coefficients, then there are:

[0034] e ij = LeakyReLu(α T [Wh i ||Wh j )

[0035]

[0036] Among them, W and α are the weight matrix and the parameter matrix to be learned respectively. The output of the entire layer of attention network is as follows:

[0037]

[0038] According to a technical solution of the present invention, in the step S4, it specifically includes:

[0039] Step S41: Divide the dual-temporal remote sensing change detection dataset made by manual annotation into a training set, a validation set, and a test set according to a preset ratio;

[0040] Step S42: Preprocess the remote sensing images in the training set to obtain preprocessed images;

[0041] Step S43: Select a weighted binary cross-entropy loss function for the change detection part, then there is:

[0042]

[0043] Among them, y i is the true category of the input instance x i , p i is the probability that the predicted input instance x i belongs to 1, and β is the weight parameter;

[0044] Step S44: Select a cross-entropy loss function for the graph convolutional neural network, then there is:

[0045]

[0046] The overall network loss function is the weighted average of the two loss functions, and the hyperparameter γ is used to balance the weights of different tasks. The calculation formula is as follows:

[0047] L = (1 - γ)loss class(Y, P) + γ_loss scene ;

[0048] Step S45: Use the training dataset for model training and the validation dataset for model validation. The validation metrics are precision, recall, F1-score, and OA, and the calculation formulas are as follows:

[0049] precision = TP / (TP + FP)

[0050] recall = TP / (TP + FN)

[0051]

[0052] OA = (TP + TN) / (TP + TN + FN + FP)

[0053] Among them, TP is the positive example correctly classified, TN is the negative example correctly classified, FP is the positive example misclassified, and FN is the negative example misclassified;

[0054] Step S46: Obtain the verified scene situation generation model for remote sensing image change detection.

[0055] According to one technical solution of the present invention, the dual-branch U-net change detection network based on Swin Transformer includes a multi-level encoder, a fusion module, and a multi-level decoder.

[0056] According to one technical solution of the present invention, in step S23, it specifically includes:

[0057] Cascade and linearly project the third-level output of the dual-branch decoder, and perform further feature aggregation through a two-layer Swin Transformer block.

[0058] According to one aspect of the present invention, there is provided an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes a method for generating a scene situation for remote sensing image change detection as described in any one of the above technical solutions.

[0059] According to one aspect of the present invention, there is provided a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, a method for generating a scene situation for remote sensing image change detection as described in any one of the above technical solutions is implemented.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] The present invention proposes a method for generating scene situation based on remote sensing image change detection, which makes full use of the rich semantic information of dual-temporal remote sensing images to realize the automatic generation of change scene situations. Using a dual-branch U-net based on Swin Transformer as the backbone network, multi-level encoders are used to extract multi-level features in the input dual-temporal remote sensing images; multi-level decoders are used to gradually restore the image resolution and generate change detection results; graph neural networks are constructed using change information, and graph convolutional networks are used to aggregate node information, and multi-task synchronous training is performed on the change detection network and the graph neural network to output the situation of the change scene. The present invention uses Swin Transformer as the basic unit to extract global information in remote sensing images, utilizes the rich semantic information in change scenes, and uses graph neural networks to efficiently model the semantic relationship between ground object changes and scene changes, and can be widely applied to fields such as agriculture, environment, transportation, and urban planning to provide information guidance for scientific decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the following described drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0063] Figure 1 Schematically showing a flow diagram of generating a scene situation based on remote sensing image change detection according to an embodiment of the present invention;

[0064] Figure 2 Schematically showing a flow diagram of generating a scene situation based on remote sensing image change detection according to another embodiment of the present invention;

[0065] Figure 3 Schematically showing a structural diagram of a change detection network based on Swin Transformer according to an embodiment of the present invention;

[0066] Figure 4 Schematically showing a method for modeling the semantic relationship between ground object changes and scene changes based on a graph neural network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The description of the embodiments of this specification should be combined with the corresponding drawings, which should be part of the complete specification. In the drawings, the shape or thickness of the embodiments may be enlarged, and simplified or convenient markings may be used. Furthermore, the parts of each structure in the drawings will be described separately. It should be noted that the elements not shown or described in words in the drawings are in forms known to those of ordinary skill in the art.

[0068] Any reference to directions and orientations in the description of the embodiments herein is for convenience of description only and should not be construed as any limitation on the scope of protection of the present invention. The following description of the preferred embodiments involves combinations of features that may exist independently or in combination. The present invention is not particularly limited to the preferred embodiments. The scope of the present invention is defined by the claims.

[0069] As Figure 1 and Figure 2 shown, a method for generating a scene situation based on remote sensing image change detection of the present invention includes the following steps:

[0070] Step S1, obtain two high-resolution satellite remote sensing images of the same area at different times and perform preprocessing;

[0071] In some embodiments of the present invention, in step S1, two high-resolution satellite remote sensing images of the same area at different times are obtained, and the images are divided into blocks of size 256×256. The images are preprocessed, including data enhancement, cloud simulation, season simulation, and environmental noise simulation, to eliminate the pseudo-changes in the remote sensing images and serve as the input of the change detection model.

[0072] As Figure 3 shown, in some embodiments of the present invention, step S2, construct a dual-branch U-net change detection network based on Swin Transformer to perform change detection on the two high-resolution satellite remote sensing images of different times, specifically including:

[0073] Step S21, perform patch segmentation and linear embedding on the preprocessed images respectively. Specifically, the input images can be divided into (4, 4) small patches, and each small patch can be regarded as a token. Subsequently, the token is linearly mapped to 96 dimensions using convolutional operations;

[0074] Step S22, construct a dual-branch U-net change detection network based on Swin Transformer as the backbone network of the change detection part;

[0075] Step S23, use the dual-branch three-level encoder part of the backbone network to perform hierarchical feature extraction on the dual-temporal remote sensing images respectively and convert them into feature tokens;

[0076] Step S24: Input the highest-level features extracted by the dual-branch encoder into the feature fusion module to fuse the dual-temporal feature maps;

[0077] Step S25: Use the three-level decoder part of the backbone network to fuse the features at all levels extracted by the encoder in sequence, gradually recover the change information, and obtain the change information with both high-level features and low-level features, where the change information includes the change category, change type, and change location;

[0078] Step S26: Perform the last upsampling, insert the channel attention module to reduce information loss, and output the change detection information.

[0079] Use the three-level decoder of the dual-branch of the change detection backbone network to gradually recover the change information from the three-level features generated by the encoder and fuse the high-level semantic features generated by the fusion module into the multi-level features. This includes: further fusing the multi-level features through the channel attention mechanism, recovering the image resolution through progressive upsampling, and reducing the computational amount through the patch shaping operation without affecting the feature information.

[0080] In some embodiments of the present invention, in step S3, based on the boundary information of the changed ground objects output by the change detection network, model the spatial relationship, construct a graph convolutional neural network, and generate an edge set and an adjacency matrix, specifically including:

[0081] Step S31: Regard the changed objects of various ground objects as nodes V = [v 1 , v 2 , v 3 , … v n , use the change type and change state of the ground object as node features, and define the spatial distance between each node as an edge E = [e 1 , e 2 , e 3 , … e n , obtain the node set V and the edge set E, and construct the graph neural network G = {V, E};

[0082] Construct a graph neural network, use the changed objects of the ground object as nodes, the change type and change state of the ground object as node features, and the spatial relationship between nodes as edges to generate graph data, and further cluster through the graph convolutional neural network to model the semantic relationship between the change of the ground object and the change of the scene.

[0083] Step S32: Introduce the boundary position information box of the changed ground object, box i = [y 1 , x 1 , y 2 , x 2Model the spatial relationship, normalize the position information of each changed object, generate the edge set E and the adjacency matrix A, then we have:

[0084]

[0085] e ij = ||v i - v j ||,

[0086]

[0087] where w and h are the width and height of the remote sensing image, e ij is the edge between changed objects i and j, and a ij is the element of the adjacency matrix A;

[0088] Step S33: Construct the degree matrix D and the Laplacian matrix L based on the adjacency matrix:

[0089]

[0090] Construct the weight matrix W, combine the Laplacian matrix and the adjacency matrix to update the node features, and the formula of graph convolution is as follows:

[0091]

[0092] where σ is the non-linear activation function, the parameters of W are learned through gradient descent and randomly initialized, and Z is represented as the output;

[0093] Step S34: As Figure 4 shown, introduce the attention mechanism, use the attention mechanism to perform weighted summation on the feature information of adjacent neighbor node representations, realize the adaptive coefficient allocation of different neighbor weights, and thus can greatly improve the representation ability of graph convolution. Assume that the feature vector corresponding to any node v i in the l-th layer of the constructed remote sensing image scene content graph is h i , after the feature aggregation of the representation information of the edges of adjacent nodes with the attention mechanism as the core, the new representation information is h i i ,

[0094] Set the central node as v i , and then assume the weight coefficients of the neighbor nodes v j to v i , then we have:

[0095] e ij = LeakyReLU(α T [Wh i || Wh j )

[0096]

[0097] Among them, W and α are the shared weight matrix and the parameter matrix to be learned respectively. Then, the output of the entire layer of attention network is as follows:

[0098] In some embodiments of the present invention, in step S4, the encoder, fusion module, decoder of the above change detection part, and the graph convolutional neural network for modeling the semantic relationship between ground object changes and scene changes are synchronously trained using the manually labeled remote sensing change detection dataset to obtain a scene situation generation model based on remote sensing image change detection, which specifically includes:

[0099] Step S41: Divide the dual-temporal remote sensing change detection dataset made by manual annotation into a training set, a validation set, and a test set according to the ratio of 7:1:2, and ensure that different parts have the same sample distribution in the way of stratified sampling;

[0100] Step S42: Preprocess the remote sensing images in the training set, including data augmentation, cloud simulation, season simulation, and environmental noise simulation, to eliminate the pseudo-changes in the remote sensing images and obtain the preprocessed images;

[0101] Step S43: Select a weighted binary cross-entropy loss function for the change detection part, then there is:

[0102]

[0103] where y i is the true category of the input instance x i , p i is the probability that the predicted input instance x i belongs to 1, and β is the weight parameter, set to 0.3, which is used to balance the representation capabilities of local features and global features and can be adjusted according to the usage scenario;

[0104] Step S44: Select a cross-entropy loss function for the graph convolutional neural network, then there is:

[0105]

[0106] The overall network loss function is the weighted average of the two loss functions, and the hyperparameter γ is used to balance the weights of different tasks. The calculation formula is as follows:

[0107] L = (1 - γ)loss class (Y, P) + γloss scene ;

[0108] Step S45: Use the training dataset to train the model and the validation dataset to validate the model. The validation metrics are precision, recall, F1-score, and OA, and the calculation formulas are as follows:

[0109] precision = TP / (TP + FP)

[0110] recall = TP / (TP + FN)

[0111]

[0112] OA = (TP + TN) / (TP + TN + FN + FP)

[0113] Among them, TP is the positive example correctly classified, TN is the negative example correctly classified, FP is the positive example misclassified, and FN is the negative example misclassified;

[0114] Step S46: Obtain the validated scene situation generation model for remote sensing image change detection.

[0115] In some embodiments of the present invention, in step S5, use the trained scene situation generation model for remote sensing image change detection to test the data in the test set and obtain the situation of the remote sensing change image.

[0116] In some embodiments of the present invention, the dual-branch U-net change detection network based on Swin Transformer includes a multi-level encoder, a fusion module, and a multi-level decoder, and a graph convolutional neural network for modeling the relationship between ground object changes and scene change semantics. The image tokens are further processed by the encoder in three stages. Each stage of the encoder consists of a Swin Transformer layer and a patch merging layer. The Swin Transformer layer is used to learn effective features, and the patch merging layer connects four feature maps in the channel dimension to implement the downsampling operation. The outputs of each level of the encoder will be input to the corresponding decoder and used together with the decoder features to restore the change information. In addition, the output of the highest level of the encoder is input to the fusion module.

[0117] The multi-level decoder is used to predict the change information stage by stage. Corresponding to the encoder, the decoder also consists of an upsampling, a patch connection, and a Swin Transformer layer. The encoder will output a feature map of the same size as the original image.

[0118] The fusion module is used to merge the high-level features generated by the encoder and consists of a concatenation layer, a linear projection layer, and a Swin Transformer layer, which are responsible for connecting the feature maps, reducing the feature dimension, and further fusing the features respectively.

[0119] In some embodiments of the present invention, in step S23, it specifically includes:

[0120] Cascade and linearly project the output of the third stage of the dual-branch decoder, and perform further feature aggregation through a two-layer SwinTransformer block.

[0121] According to one aspect of the present invention, there is provided an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes a method for generating a scene situation based on remote sensing image change detection according to any one of the above technical solutions.

[0122] According to one aspect of the present invention, there is provided a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, a method for generating a scene situation based on remote sensing image change detection according to any one of the above technical solutions is implemented.

[0123] The computer-readable storage medium may include any medium capable of storing or transmitting information. Examples of computer-readable storage media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, optical fiber media, radio frequency (RF) links, and so on. Code segments may be downloaded via a computer network such as the Internet, intranet, etc.

[0124] A method for generating a scene situation based on remote sensing image change detection of the present invention uses a dual-branch U-net based on SwinTransformer as the backbone network, extracts multi-level features in the input dual-temporal remote sensing images through a multi-level encoder; adopts a multi-level decoder to gradually restore the image resolution and generate a change detection result; constructs a graph neural network using the change information, aggregates node information using a graph convolutional network, and performs multi-task synchronous training on the change detection network and the graph neural network to output the situation of the change scene. The present invention uses the Swin Transformer as the basic unit to extract global information in the remote sensing image, utilizes the rich semantic information in the change scene, and uses the graph neural network to efficiently model the semantic relationship between ground object changes and scene changes, and can be widely applied to fields such as agriculture, environment, transportation, urban planning, etc., providing information guidance for scientific decision-making.

[0125] In addition, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0126] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0128] It should also be noted that in this document, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or terminal device including the said element.

[0129] Finally, it should be noted that the above description is the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those skilled in the art of this technology, once they know the basic creative concept of the present invention, without departing from the principle described in the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

Claims

1. A scene situation generation method based on remote sensing image change detection, characterized in that: The following steps are involved: Step S1, obtaining two high-resolution satellite remote sensing images of the same area at different phases and performing preprocessing; Step S2, constructing a dual-branch U-net change detection network based on Swin Transformer to perform change detection on the two high-resolution satellite remote sensing images in different phases; Step S3: Modeling the spatial relationship based on the boundary information of the changed objects output by the change detection network, constructing a graph convolutional neural network, and generating edge sets and adjacency matrices, specifically including: Step S31: Treat the changed objects of various types of ground features as nodes V = [v1, v2, v3, ... v n ], the change type and change state of the features are used as node features, and the spatial distance between each node is defined as the edge E = [e1, e2, e3, ... e n ], get the node set V and edge set E, and construct the graph neural network G = {V, E}; Step S32: Introduce the boundary position information box of the changed object, box i =[y1,x1,y2,x2] To model the spatial relationship, normalize the location information of each variable feature, generate the edge set E and the adjacency matrix A, then: Among them, w and h are the width and height of the remote sensing image, e ij is the edge between the changing features i and j, a ij is an element of the adjacency matrix A; Step S33: construct the degree matrix D and the Laplace matrix L based on the adjacency matrix: Construct the weight matrix W, combine the Laplacian matrix and the adjacency matrix to update the node features, and the formula for graph convolution is as follows: Where σ is a nonlinear activation function, the parameters of W are learned by gradient descent and randomly initialized, and Z is represented as the output; Step S34: Introduce the attention mechanism, assuming that any node v in the constructed remote sensing image scene content graph i The corresponding feature vector at layer l is h i After the representation information feature aggregation of the edges of adjacent nodes with the attention mechanism as the core, the new representation information is h i ' , Set the center node to v i , then assume that the neighbor node v j to v i The weight coefficient is: and ij =LeakyReLU(α T [Wh i ||Wh j ]) Among them, W and α are the weight matrix and the parameter matrix to be learned respectively. ij is an element in α, then the output of the entire attention network is as follows: Step S4: Use the manually annotated remote sensing change detection dataset to train the graph convolutional neural network to obtain a scene situation generation model based on remote sensing image change detection; Step S5: Using the trained scene situation generation model based on remote sensing image change detection, the data in the test set is tested to obtain the situation of the remote sensing change image.

2. The scene situation generation method based on remote sensing image change detection according to claim 1 is characterized in that: In step S1, after acquiring two high-resolution satellite remote sensing images of the same area at different phases, the images are divided into blocks and processed into images of size 256×256, and the two high-resolution satellite remote sensing images at different phases are preprocessed. The preprocessing includes at least data enhancement, cloud and fog simulation, season simulation, and environmental noise simulation.

3. The scene situation generation method based on remote sensing image change detection according to claim 1 is characterized in that: The step S2 specifically includes: Step S21, performing spot segmentation and linear embedding on the preprocessed image respectively; Step S22, constructing a dual-branch U-net change detection network based on Swin Transformer as the backbone network of the change detection part; Step S23, using the dual-branch three-stage encoder part of the backbone network to extract features of the dual-temporal remote sensing images step by step and convert them into feature tokens; Step S24, inputting the highest-level features extracted by the dual-branch encoder into a feature fusion module to fuse the dual-phase feature graph; Step S25, using the three-level decoder part of the backbone network to sequentially fuse the features of each level extracted by the encoder, gradually restore the change information, and obtain change information with both high-level features and low-level features, wherein the change information includes change category, change type and change position; Step S26: perform the last upsampling, insert the channel attention module to reduce information loss, and output change detection information.

4. The scene situation generation method based on remote sensing image change detection according to claim 3 is characterized in that: The step S4 specifically includes: Step S41, dividing the manually annotated dual-temporal remote sensing change detection dataset into a training set, a validation set, and a test set according to a preset ratio; Step S42, preprocessing the remote sensing images in the training set to obtain preprocessed images; Step S43: Select a weighted binary cross entropy loss function for the change detection part, then: Among them, y i For the input instance x i The true category, p i Input instance x for prediction i The probability of belonging to 1, β is the weight parameter; Step S44: Select the cross entropy loss function for the graph convolutional neural network, then: The overall network loss function is the weighted average of the two loss functions, using the hyperparameter γ to balance the weights of different tasks. The calculation formula is as follows: L=(1-γ)loss class (Y,P)+gloss scene ; Step S45: Use the training data set to train the model, and use the validation data set to validate the model. The validation indicators are precision, recall, F1-score, and OA. The calculation formula is as follows: precision=TP / (TP+FP) recall=TP / (TP+FN) OA=(TP+TN) / (TP+TN+FN+FP) Among them, TP is a correctly classified positive example, TN is a correctly classified negative example, FP is a misclassified positive example, and FN is a misclassified negative example; Step S46: Obtain a verified scene situation generation model based on remote sensing image change detection.

5. The scene situation generation method based on remote sensing image change detection according to claim 3 is characterized in that: The dual-branch U-net change detection network based on Swin Transformer includes a multi-stage encoder, a fusion module, and a multi-stage decoder.

6. The scene situation generation method based on remote sensing image change detection according to claim 5 is characterized in that: The step S23 specifically includes: The third-stage output of the multi-stage encoder is concatenated and linearly projected, and passed through a two-layer SwinTransformer block for further feature aggregation.

7. An electronic device, characterized in that: include: One or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the scene situation generation method based on remote sensing image change detection as described in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, implement the scene situation generation method based on remote sensing image change detection as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Remote sensing image ground object change detection method and device

    CN116385881A

  • Convolutional network and graph neural network hybrid-based change detection network and method

    CN116778317A

  • Aerial photo ground feature classification and change detection method based on graph neural network

    CN117437234A