Method, device and computer program product for identifying changing patterns
Through the change detection method combining feature extraction and multi-network detection, the problems of false changes and missed detection in the existing technology are solved, and higher-precision change detection is achieved.
Patent Information
- Application Number
- CN202511030606.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing change detection technologies have shortcomings in identifying false changes and missed detections, making it difficult to accurately identify changes in land objects, especially small plots and land objects with similar colors.
The trained feature extraction network, the first change detection network, the second change detection network and the third change detection network are used to perform feature extraction and change detection on the dual-phase images. Combined with the change classification model, the problems of missed detection and false detection are focused on respectively, and the detection accuracy is improved through feature fusion and classification recognition.
It improves the accuracy of change detection, reduces the misidentification of false changes, and enhances the detection ability of small plots and objects with similar colors.
Smart Images

Figure CN120526241B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing technology, and in particular to a method, device and computer program product for identifying change spots. Background Art
[0002] Change detection is a hot topic in remote sensing research. It identifies surface changes by comparing remote sensing images at different points in time. It has widespread applications in urban planning, environmental monitoring, disaster assessment, agricultural management, and other fields, providing crucial data support for decision makers. However, existing change detection technologies face numerous challenges in practical applications, with spurious changes and missed detections being urgent challenges that need to be addressed.
[0003] Pseudo-change patches are unchanging areas that are mistakenly identified as changing areas, such as the seasonal changes in farmland (green in spring, yellow in autumn). There are also pseudo-changes caused by natural phenomena: quicksand changes shape due to wind, which can easily be misidentified as ground changes. Conventional change detection methods are prone to missing small plots or features with a similar color to the background. Given these scenarios and the shortcomings of existing technologies, a more accurate change detection method is urgently needed to address the problem. Summary of the Invention
[0004] In order to solve the problems existing in the prior art, the present invention provides a method for identifying a change pattern, which includes:
[0005] Acquire a dual-phase image of the area to be studied, wherein the dual-phase image includes a front-phase remote sensing image and a back-phase remote sensing image; use a trained feature extraction network to extract features from the dual-phase image to obtain a feature map; use a trained first change detection network, a second change detection network, and a third change detection network to perform change detection on the feature map respectively to obtain a first prediction mask map, a second prediction mask map, and a third prediction mask map, wherein the second change detection network is a network that focuses on global change detection, the first change detection network is a network that focuses on missed detection problems, and the third change detection network is a network that focuses on false detection problems; use a trained change classification model to perform change classification and recognition processing on the dual-phase image, the first prediction mask map, the second prediction mask map, and the third prediction mask map to obtain a change classification and recognition result, wherein the change classification and recognition result includes: pseudo-change spots and real change spots.
[0006] In another aspect, the present invention provides a device for identifying a change pattern, comprising:
[0007] An image acquisition module is used to acquire a dual-phase image of the area to be studied, wherein the dual-phase image includes a front-phase remote sensing image and a back-phase remote sensing image; a feature extraction module is used to use a trained feature extraction network to perform feature extraction on the dual-phase image to obtain a feature map; a change detection module is used to use a trained first change detection network, a second change detection network and a third change detection network to perform change detection on the feature map respectively to obtain a first prediction mask map, a second prediction mask map and a third prediction mask map, wherein the second change detection network is a network focusing on global change detection, the first change detection network is a network focusing on missed detection problems, and the third change detection network is a network focusing on false detection problems; a classification and recognition module is used to use a trained change classification model to perform change classification and recognition processing on the dual-phase image, the first prediction mask map, the second prediction mask map and the third prediction mask map to obtain a change classification and recognition result, wherein the change classification and recognition result includes: pseudo-change spots and real change spots.
[0008] In yet another aspect of the present invention, an electronic device is provided, comprising: a processor and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the above-mentioned method for identifying change patterns.
[0009] On the other hand, the present invention provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the above-mentioned method for identifying the changing pattern.
[0010] In yet another aspect, the present invention provides a computer program product comprising instructions or programs, and when the computer program product is executed by a computer, the computer program product implements the above-mentioned method for identifying change patterns.
[0011] The technical solution provided by the embodiment of the present invention has the following beneficial effects:
[0012] By acquiring dual-phase images of the area to be studied, the dual-phase images include front-phase remote sensing images and back-phase remote sensing images; using the trained feature extraction network to extract features from the dual-phase images to obtain feature maps; using the trained first change detection network, second change detection network and third change detection network to perform change detection on the feature maps respectively to obtain the first prediction mask map, the second prediction mask map and the third prediction mask map, the second change detection network is a network that focuses on global change detection, the first change detection network is a network that focuses on missed detection problems, and the third change detection network is a network that focuses on false detection problems; using the trained change classification model to perform change classification and recognition processing on the dual-phase images, the first prediction mask map, the second prediction mask map and the third prediction mask map to obtain change classification and recognition results, thereby improving the accuracy of change detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 1 is a flow chart of a method for identifying a change pattern provided by an embodiment of the present invention;
[0014] Figure 2 is a schematic structural diagram of a first change detection network provided by an embodiment of the present invention;
[0015] Figure 3 1 is a schematic diagram of a framework of another method for identifying a change pattern provided by an embodiment of the present invention;
[0016] Figure 4 Schematic diagram of a framework of a feature extraction network and three change detection networks provided by an embodiment of the present invention;
[0017] Figure 5 Schematic diagram of a framework of a semi-supervised classification model provided by an embodiment of the present invention;
[0018] Figure 6 This is a front-phase remote sensing image provided by an embodiment of the present invention;
[0019] Figure 7 This is a post-temporal remote sensing image provided by an embodiment of the present invention;
[0020] Figure 8 This is a schematic diagram of a missed detection label provided by an embodiment of the present invention;
[0021] Figure 9 This is a schematic diagram of a real label provided by an embodiment of the present invention;
[0022] Figure 10 This is a schematic diagram of a misdetected label provided by an embodiment of the present invention;
[0023] Figure 11 is a fused prediction graph provided by an embodiment of the present invention;
[0024] Figure 12 This is a manually annotated fused prediction graph provided by an embodiment of the present invention;
[0025] Figure 13 It is a posterior phase remote sensing image in unlabeled data during training of a variational autoencoder provided by an embodiment of the present invention;
[0026] Figure 14 It is a front-phase remote sensing image in unlabeled data during training of a variational autoencoder provided by an embodiment of the present invention;
[0027] Figure 15 It is a label map of unlabeled data during training of a variational autoencoder provided by an embodiment of the present invention;
[0028] Figure 16 A variational autoencoder provided by an embodiment of the present invention is trained with a front-phase remote sensing image in labeled data;
[0029] Figure 17 A variational autoencoder provided by an embodiment of the present invention is trained with a posterior phase remote sensing image in labeled data;
[0030] Figure 18 It is a label graph in labeled data during training of a variational autoencoder provided by an embodiment of the present invention;
[0031] Figure 19 1 is a schematic diagram of the structure of a variational autoencoder provided by an embodiment of the present invention;
[0032] Figure 20 It is a structural diagram of a device for identifying changing patterns provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0034] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings.
[0035] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0036] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if monitoring (stated condition or event)" may be interpreted as "when determining" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0037] See also Figure 1 , an embodiment of the present invention provides a method for identifying a change pattern, which includes the following steps:
[0038] Step 101: Acquire a dual-phase image of the area to be studied.
[0039] Acquire a remote sensing image (or remote sensing image) of the area to be studied. The area to be studied is usually a surface area containing a variety of ground objects, such as various buildings, trees, rivers, and various terrains. Accordingly, a variety of ground objects will be reflected in the remote sensing image. There are many ways to acquire remote sensing images, which are not limited by the present invention. This recognition method processes a dual-phase image, which includes: a front-phase remote sensing image x1 and a back-phase remote sensing image x2, which can be seen in Figure 1 、 Figure 6 and Figure 7 , Figure 6 for Figure 1 A magnified schematic diagram of the mid-front phase remote sensing image. Figure 7 for Figure 1 A magnified schematic diagram of the mid- to late-phase remote sensing image.
[0040] Step S102: Using the trained feature extraction network to extract features from the dual-phase image to obtain a feature map.
[0041] Before use, an initial feature extraction network, a first change detection network, a second change detection network, and a third change detection network are constructed, and then trained to obtain a trained feature extraction network, a first change detection network, a second change detection network, and a third change detection network.
[0042] Regarding the feature extraction network, for example, the feature extraction network may include: an encoder, such as a twin network encoder. The input is a dual-phase image, and two encoders are configured accordingly, namely encoder1 (or e1) and encoder2 (or e2). The two encoders may have exactly the same network architecture, but their parameters are independent of each other, and are used to encode the front-phase remote sensing image and the back-phase remote sensing image, respectively, to obtain the front-phase features and the back-phase features. For example: when the input size is When H represents the height of the image, W represents the width of the image, and the output size is , the corresponding formula is expressed as:
[0043] Pre-phase characteristics: ,
[0044] Post-phase characteristics: .
[0045] The specific network architecture of the encoder is as follows:
[0046]
[0047] The residual block is used to indicate that two residual operations are performed. Taking the second stage as an example, the operation is explained as follows: first, a residual operation with a step size of 2 is performed, and then a residual operation with a step size of 1 is performed.
[0048] The feature extraction network can also include: Transformer feature fusion (or T), expressed as , that is, the feature extraction network includes: encoder and Transformer feature fusion. The Transformer feature fusion is used to fuse the front-phase features and the back-phase features output by the two encoders. The formula is expressed as:
[0049]
[0050] This embodiment does not limit the specific steps of feature fusion, which can be as follows:
[0051] First, perform spatial flattening to convert the 2D feature map into a 1D sequence. The formula is:
[0052]
[0053]
[0054]
[0055] In the formula, Flatten() means flattening.
[0056] Then add sinusoidal position encoding to get enhanced features and For example, when calculating the position encoding, sin is used for even dimensions and cos is used for odd dimensions. Pos represents the position index in the sequence, i represents the dimension index in the embedding vector, and d represents the total dimension of the embedding vector. The formula is expressed as:
[0057]
[0058]
[0059]
[0060]
[0061]
[0062] Then, a bidirectional cross-attention mechanism is used to fuse the foreground and background information using the cross-operation Query (or Q), Key (or K) and Value (or V) and find the difference between them. , , Represents the weight matrix of query, key, and value respectively. The weight matrix is a learning parameter used to map the input vector to different subspaces. This can be seen as information extracted from the original input for comparison with other outputs. Key is obtained by multiplying the original input with the weight matrix Multiplying is used to match the query to determine the degree of association between the inputs. The value is obtained by multiplying the original input with the weight matrix. It is obtained by multiplication, which contains the actual information or content and is weighted and summarized based on the similarity between the query and the key.
[0063] In this method, the query is generated using the features of the previous phase, the key / value is generated using the features of the next phase, and then the attention output is calculated. The relevant formula is expressed as:
[0064]
[0065]
[0066]
[0067]
[0068] Then use the later phase features to generate Query, use the earlier phase features to generate Key / Value, and then calculate , the relevant formula is expressed as:
[0069]
[0070]
[0071]
[0072]
[0073] Finally, the feedforward neural network is used for feature fusion, and the relevant formula is expressed as:
[0074]
[0075]
[0076]
[0077] Where b1 and b2 are parameters in the network, which are used to represent the threshold.
[0078] Then output the reconstructed features
[0079]
[0080] It should be noted that: LayerNorm() represents layer normalization, ReLu() represents activation function, Reshape() represents reconstruction, and Attention() represents attention.
[0081] Regarding the first change detection network (or head1 or d1), the second change detection network (or head2 or d2) and the third change detection network (or head3 or d3), the first change detection network is a network that focuses on missed detection issues, such as small plots and small buildings; the second change detection network is a network that focuses on global change detection; the third change detection network is a network that focuses on false detection issues, such as pseudo-change areas such as quicksand deformation and shadows. Exemplarily: the three change detection networks include decoders, which have the same network architecture, but the parameters of the three are independent of each other, and the output results are: the first prediction mask map (or missed detection prediction map or ), the second prediction mask map (also called conventional prediction map or ) and the third prediction mask map (also called false detection prediction map or ), it should be noted that, for the sake of distinction, the output of the trained change detection network can be called a prediction mask map, such as the first prediction mask map, the second prediction mask map, and the third prediction mask map. During training, the output of the change detection network can be called a prediction map, such as a missed detection prediction map, a regular prediction map, and a false detection prediction map. The relevant formula is expressed as:
[0082]
[0083]
[0084]
[0085] The specific network architecture of the decoder can be found in Figure 2 ,exist Figure 2 In the example, taking input as an example, the corresponding meaning is: (H / 32)*(W / 32)*512. Taking "Upsampling 1" as an example, the corresponding meaning is upsampling its input, and the upsampling number is 1.
[0086] For loss functions and training strategies, see Figure 4 , illustratively:
[0087] Generate a missed detection weight map, which is used to , so that the missed detection area has The loss value is times that of missed detection, and the relevant formula is expressed as:
[0088]
[0089] Generate a false positive weight map, which is used to , so that the false detection area has a The loss value is times that of the original value to focus on the problem of false detection. The relevant formula is expressed as:
[0090]
[0091] In the formula, y represents the true label, that is, the manually marked change area. It should be noted that in the application, and As a parameter. The value of can be 3, The value of can be 4.
[0092] Then, the three loss functions corresponding to the first change detection network, the second change detection network, and the third change detection network are calculated respectively: the first loss function L1, the second loss function L2, and the third loss function L3:
[0093]
[0094]
[0095]
[0096] It should be noted that: () represents the binary cross entropy loss function, i represents the i-th pixel, HW represents the total number of pixels in the image, and As an example, the meaning of the superscript (i) is represented by the value of the i-th pixel in the true label.
[0097] During training, parameters are isolated to prevent the currently optimized object from affecting other objects. Specifically, the feature extraction network and the second change detection network form the backbone network, and the backbone network is trained to obtain the trained second change detection network and feature extraction network. The first change detection network is trained separately to obtain the trained first change detection network. The third change detection network is trained separately to obtain the trained third change detection network. That is, according to the updated parameters, other networks / modules are temporarily frozen at this time; according to the updated parameters, other networks / modules are temporarily frozen at this time; and finally, according to the updated parameters, other networks / modules are temporarily frozen at this time. During backbone network training, the samples in the second training set include: dual-phase images and true labels, such as Figure 4 and Figure 9 As shown, Figure 9 for Figure 4 The enlarged diagram of the true label in the first training set during the training of the first change detection network includes: the output of the feature extraction network of the backbone network and the missed detection label. The missed detection label is the missed detection comparison result between the missed detection prediction map and the true label. The missed detection comparison result is used to indicate the area that exists in the true label but not in the missed detection prediction map, such as Figure 4 and Figure 8 The middle red part, Figure 8 for Figure 4 The samples in the third training set during the training of the third change detection network include: the output of the feature extraction network of the backbone network and the false positive label. The false positive label is the false positive comparison result between the false positive prediction map and the true label. The false positive comparison result is used to indicate the area that is not in the true label but is in the false positive prediction map, such as Figure 4 and Figure 10 The middle green part, Figure 10 for Figure 4 A magnified diagram of misdetected labels.
[0098] During training, the first change detection network performs change detection on the feature map to obtain a prediction map, which is called a missed detection prediction map; the second change detection network performs change detection on the feature map to obtain a prediction map, which is called a regular prediction map; the third change detection network performs change detection on the feature map to obtain a prediction map, which is called a false detection prediction map. Feature fusion is performed on the missed detection prediction map, the regular prediction map, and the false detection prediction map to obtain a fused prediction map (which can be called the fourth prediction mask map after training is completed). Figure 11 As shown. The specific fusion method can be found in the fusion method in step 104 below. The fused prediction map will also be manually annotated, such as marking real changes and pseudo changes such as snow cover, vegetation season, and quicksand deformation, to obtain a manually annotated fused prediction map, such as Figure 12 As shown, it can be used to train a variation classification model.
[0099] The trained feature extraction network is used to extract features from the dual-phase image to obtain a feature map.
[0100] Step 103: Use the trained first change detection network, the second change detection network, and the third change detection network to perform change detection on the feature map respectively to obtain a first prediction mask map, a second prediction mask map, and a third prediction mask map.
[0101] Regarding how the first change detection network, the second change detection network, and the third change detection network are trained, please refer to the description of step 102 above, which will not be described in detail here. Figure 3 , use the trained first change detection network to perform change detection on the feature map to obtain a first prediction mask map; use the trained second change detection network to perform change detection on the feature map to obtain a second prediction mask map; use the trained third change detection network to perform change detection on the feature map to obtain a third prediction mask map.
[0102] Step 104: Using the trained change classification model, perform change classification and recognition processing on the dual-phase image, the first prediction mask image, the second prediction mask image, and the third prediction mask image to obtain change classification and recognition results. The change classification and recognition results include: pseudo change spots and real change spots.
[0103] This step can be implemented as follows:
[0104] First, feature fusion is performed on the first prediction mask map, the second prediction mask map, and the third prediction mask map to obtain a fourth prediction mask map.
[0105] Exemplarily, the fourth prediction mask image is obtained The formula is as follows:
[0106]
[0107] in, represents the threshold for missed detection, represents the threshold for false detection, Represents pixel coordinates.
[0108] Then, the trained change classification model is used to perform change classification and recognition processing on the dual-phase image and the fourth prediction mask image to obtain a change classification and recognition result.
[0109] See also Figure 6 When training the change classification model, the samples in the training set include: dual-phase images, mask images corresponding to the dual-phase images, and classification labels. The classification labels mark pseudo-change spots and real change spots. The changes in the images corresponding to the pseudo-change spots can be caused by snow cover, vegetation season, or quicksand deformation. This embodiment does not specifically limit this.
[0110] When applied, first perform connected domain analysis on the fourth prediction mask map to extract all the change spots (or change spots or change areas). For each change spot , calculate its minimum circumscribed square, and extend the side length to the preset threshold, for example, the value is: , if the minimum circumscribed square side length is greater than , the change patch can be divided into multiple sub-change patches for processing. Then, according to the coordinate correspondence, the change patch is processed from the previous time to the remote sensing image. , post-phase remote sensing images Intercept the corresponding area and , thus obtaining 、 and Then, a patch classification dataset is constructed, and change classification and recognition processing is performed on each element in the patch classification dataset to obtain the corresponding change patch classification results.
[0111] The change classification model is preferably a semi-supervised classification model. Specifically, a small amount of labeled data and a large amount of unlabeled data are used to train the change classification model (or true / false change classification model), which greatly reduces the workload of manual labeling, removes false change spots, and improves the generalization ability and accuracy of the model.
[0112] Specifically, the change classification model consists of a feature extractor and a classifier connected in sequence. The feature extractor extracts features from the fourth prediction mask, the preceding remote sensing image, and the following remote sensing image. The classifier classifies the feature extractor's output to produce a change patch classification result. Change patch classification results include true change and false change. False change can be of various types, such as snow cover, vegetation season, and quicksand deformation. The change area type (or change patch type) can be coded separately, for example, 0 for true change, 1 for snow cover, 2 for vegetation season, and 3 for quicksand deformation.
[0113] The feature extractor is an encoder in the autoencoder, and the autoencoder is preferably a VAE (Variational Autoencoder), which has an encoder f e and decoder d e , the decoder d of the autoencoder e It can be ResNet (Residual Neural Network). VAE reconstructs the loss function and combines KL divergence (Kullback-Leibler Divergence) as a new loss function, making the encoder f of the autoencoder e It can extract discriminative latent features. VAE can perform cross-modal feature fusion, where the convolution kernel acts on spatiotemporal-spectral-variation information simultaneously. By automatically learning feature correlations, it can capture complex patterns such as "spectral anomalies in changing areas". VAE can be pre-trained using large-scale unlabeled image data (or unlabeled data). The structure of VAE can be seen in Figure 19 As shown, Figure 13 It indicates Figure 19 Post-temporal remote sensing images in unlabeled data; Figure 14 It indicates Figure 19 Front-phase remote sensing images in unlabeled data; Figure 15 It indicates Figure 19 Label graph in unlabeled data; Figure 16 The diagram shows the remote sensing image of the previous phase in the labeled data; Figure 17 The illustrated image is a later-phase remote sensing image in the labeled data; Figure 18 The diagram shows the label graph in the labeled data. The classifier includes convolutional pooling layers and linear layers.
[0114] First, a connected domain analysis is performed on the fourth predicted mask feature map. Specifically, for the fourth predicted mask feature map, Change Area , calculate its minimum circumscribed square, if the side length of the minimum circumscribed square is greater than the preset threshold , then the change area can be divided into multiple parts so that the minimum circumscribed square of each divided change area is no larger than . Preset threshold It can be 128, the unit is the number of pixels. Then according to the coordinates, the remote sensing image of the previous phase , post-phase remote sensing images Intercept the corresponding area and Then construct the spot classification dataset D2, the formula is expressed as:
[0115]
[0116] Where N represents the number of change patches.
[0117] Then the triple Spliced into a 7-channel tensor , expressed as:
[0118]
[0119] Among them, concat() means splicing, Change patches (or change areas) The corresponding area in the previous phase remote sensing image, Change pattern The corresponding area in the later phase remote sensing image.
[0120] Finally, a labeled dataset is constructed ,in, express
[0121] The input data is spliced together with Consistent, Represents the manually annotated label. The unlabeled sample data constitutes the unlabeled sample dataset, which is represented as .
[0122] The triplets are then concatenated into a 7-channel tensor Through the encoder and decoder The reconstruction result can be obtained . Use binary cross entropy to calculate the reconstruction loss , added to the KL divergence loss to get the total loss :
[0123]
[0124] in, is the feature extractor loss function, For the encoder, are input data and latent variables respectively, are encoder parameters, For the decoder, are decoder parameters, is the KL divergence loss function, is the prior distribution of latent variables, is the KL divergence weight coefficient.
[0125] Regarding the classifier, the encoder As a feature extractor, access the fully connected classification head , in the labeled dataset Fine-tune training on unlabeled data. At the same time, a dynamic proxy label generation mechanism is constructed: Perform multiple rounds of predictions. If the prediction category remains consistent for c consecutive periods (c = 5) and the confidence exceeds the threshold, When , the high confidence prediction results are converted into proxy labels , we get the proxy labels for the unlabeled dataset , and then calculate the loss functions separately and perform iterative collaborative optimization.
[0126] Each interval ( ) epochs to update the proxy labels, ensuring the model has sufficient time to learn the proxy labels. By alternately optimizing the feature extraction network and classifier parameters, a positive feedback loop of feature learning, proxy label generation, and model optimization is formed. Training is terminated when the maximum number of epochs is reached or the loss function stops decreasing within k epochs.
[0127] Specifically, the unlabeled image data is processed by the preliminarily trained change classification model to obtain the image data labeled with proxy labels, where the proxy labels are annotations of the position, size, and type of the image change area.
[0128] The unlabeled image data is processed according to the initially trained change classification model to obtain the classification prediction results of the change area.
[0129] If the classification prediction results are consistent for N consecutive times and the confidence level exceeds the preset confidence threshold , then set a proxy label for the changed area to obtain the image data annotated with the proxy label. Where N can be 5, the confidence threshold It can be determined according to actual conditions, and the present invention does not limit this.
[0130] The image data annotated by the proxy label is represented as ,in ,in is the feature extractor, For the classifier.
[0131] The model is trained based on the labeled image dataset combined with the image data annotated with the proxy label, and the aforementioned process is repeated to obtain a trained change classification model.
[0132] The loss function during training can be expressed as: ,in is the loss function when training labeled image data, is the loss function for training the image data annotated with the proxy label. In the actual training process, you can Update the image data annotated by the proxy label once every epoch to ensure that the change classification model is fully trained. It can be 3. By optimizing the parameters of the feature extractor and classifier, a positive feedback loop of feature learning-proxy label generation-model optimization is formed. When the preset number of epochs m is reached or the loss function does not decrease within n epochs, the training is stopped and the trained change classification model is obtained. Among them, m and n can be determined according to the actual situation, and the present invention does not limit this. The structure of the change classification model can be as follows Figure 5 As shown in the figure, in the application, the change classification model can remove pseudo-change spots and perform data set feedback correction. When the proportion of pseudo-change spots is greater than the preset threshold, the change data is added or corrected, the change detection training set is adjusted, and the network can be trained again.
[0133] By acquiring dual-phase images of the area to be studied, the dual-phase images include front-phase remote sensing images and back-phase remote sensing images; using the trained feature extraction network to extract features from the dual-phase images to obtain feature maps; using the trained first change detection network, second change detection network and third change detection network to perform change detection on the feature maps respectively to obtain the first prediction mask map, the second prediction mask map and the third prediction mask map, the second change detection network is a network that focuses on global change detection, the first change detection network is a network that focuses on missed detection problems, and the third change detection network is a network that focuses on false detection problems; using the trained change classification model to perform change classification and recognition processing on the dual-phase images, the first prediction mask map, the second prediction mask map and the third prediction mask map to obtain change classification and recognition results, thereby improving the accuracy of change detection.
[0134] See also Figure 20 An embodiment of the present invention provides a device for identifying change spots, which is used to execute a method for identifying change spots provided in the above embodiment. The identification device includes: an image acquisition module 201, a feature extraction module 202, a change detection module 203 and a classification and recognition module 204.
[0135] The image acquisition module 201 is used to acquire dual-phase images of the area to be studied, and the dual-phase images include a front-phase remote sensing image and a back-phase remote sensing image.
[0136] The feature extraction module 202 is used to extract features from the bi-temporal image using the trained feature extraction network to obtain a feature map.
[0137] The change detection module 203 is used to use the trained first change detection network, second change detection network and third change detection network to perform change detection on the feature map respectively to obtain the first prediction mask map, the second prediction mask map and the third prediction mask map. The second change detection network is a network that focuses on global change detection, the first change detection network is a network that focuses on missed detection problems, and the third change detection network is a network that focuses on false detection problems.
[0138] The classification and recognition module 204 is used to use the trained change classification model to perform change classification and recognition processing on the dual-phase image, the first prediction mask image, the second prediction mask image and the third prediction mask image to obtain change classification and recognition results. The change classification and recognition results include: pseudo change spots and real change spots.
[0139] Regarding the implementation of the specific functions of the above modules, please refer to the relevant content of the change pattern recognition method in the above embodiment, which will not be described in detail here.
[0140] It should be noted that the aforementioned embodiments of the apparatus for identifying changing patterns only illustrate the division of the aforementioned functional modules when identifying changing patterns. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus for identifying changing patterns provided in the aforementioned embodiments and the embodiment of the method for identifying changing patterns are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0141] An embodiment of the present invention provides an electronic device comprising: a memory and a processor. The processor is connected to the memory and is configured to execute the above-described method for identifying change patterns based on instructions stored in the memory. The number of processors may be one or more, and the processor may be single-core or multi-core. The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip. The memory may be an example of a computer-readable medium as described below.
[0142] An embodiment of the present invention provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the above-described method for identifying change patterns. Computer-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented using any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc-read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0143] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any method described in the aforementioned embodiment of the method for identifying change patterns.
[0144] It is understood from common technical knowledge that the present invention may be implemented by other embodiments that do not depart from its spirit or essential features. Therefore, the embodiments disclosed above are, in all respects, merely illustrative and not exclusive. All modifications within the scope of the present invention or equivalent to the scope of the present invention are intended to be encompassed by the present invention.
Claims
1. A method for identifying a change pattern, characterized in that: The identification method comprises: Acquire a dual-phase image of the area to be studied, wherein the dual-phase image includes a front-phase remote sensing image and a back-phase remote sensing image; Extracting features from the dual-phase image using a trained feature extraction network to obtain a feature map; Using the trained first change detection network, second change detection network, and third change detection network to perform change detection on the feature map, respectively, to obtain a first prediction mask map, a second prediction mask map, and a third prediction mask map, wherein the second change detection network is a network focusing on global change detection, the first change detection network is a network focusing on missed detection problems, and the third change detection network is a network focusing on false detection problems; Using the trained change classification model, the bi-temporal image, the first prediction mask image, the second prediction mask image, and the third prediction mask image are subjected to change classification and recognition processing to obtain a change classification and recognition result, wherein the change classification and recognition result includes: pseudo change spots and real change spots; The trained change classification model is used to perform change classification and recognition processing on the bi-temporal image, the first prediction mask image, the second prediction mask image, and the third prediction mask image to obtain a change classification and recognition result, including: performing feature fusion on the first prediction mask map, the second prediction mask map, and the third prediction mask map to obtain a fourth prediction mask map; Using the trained change classification model, the bi-temporal image and the fourth prediction mask image are subjected to change classification and recognition processing to obtain a change classification and recognition result; The identification method further includes: Constructing an initial feature extraction network, a first change detection network, a second change detection network, and a third change detection network; Performing network training on the feature extraction network and the second change detection network together using a second training set to obtain trained second change detection network and feature extraction network, wherein samples in the second training set include bi-temporal images and true labels, and the true labels are manually annotated change regions; Performing separate network training on the first change detection network using a first training set to obtain a trained first change detection network, wherein the samples in the first training set include outputs of a feature extraction network and missed detection labels, wherein the missed detection labels are missed detection comparison results of a missed detection prediction map corresponding to the output of the feature extraction network and the true label, wherein the missed detection comparison results are used to indicate regions that are present in the true label but not in the missed detection prediction map; The third change detection network is individually trained using a third training set to obtain the trained third change detection network. The samples in the third training set include the output of the feature extraction network and the false detection label. The false detection label is a false detection comparison result between the false detection prediction map corresponding to the output of the feature extraction network and the true label. The false detection comparison result is used to represent an area that is not in the true label but is in the false detection prediction map.
2. The identification method according to claim 1, characterized in that The identification method further includes: Obtain the first loss function, the second loss function, and the third loss function; The expression of the first loss function L1 is: The expression of the second loss function L2 is: The expression of the third loss function L3 is: Where H represents the image height in the bi-temporal image, W represents the image width in the bi-temporal image, i represents the i-th pixel, represents the binary cross entropy loss function, and are parameters, y represents the true label, pred1 represents the missed detection prediction map, pred2 represents the regular prediction map, and pred3 represents the false detection prediction map; Accordingly, the performing separate network training on the first change detection network using the first training set to obtain the trained first change detection network includes: Performing separate network training on the first change detection network using the first training set and the first loss function to obtain a trained first change detection network; Accordingly, the performing network training on the feature extraction network and the second change detection network together using the second training set to obtain the trained second change detection network and the feature extraction network includes: Performing network training on the feature extraction network and the second change detection network using the second training set and the second loss function to obtain trained second change detection network and feature extraction network; Accordingly, the method of performing separate network training on the third change detection network using the third training set to obtain the trained third change detection network includes: The third change detection network is individually trained using the third training set and the third loss function to obtain the trained third change detection network.
3. The identification method according to claim 1, characterized in that The feature extraction network includes: two encoders and a Transformer feature fusion device, the two encoders are used to correspond to the two-phase images one by one, and the Transformer feature fusion device is used to perform feature fusion on the outputs of the two encoders.
4. The identification method according to claim 3, characterized in that The first change detection network, the second change detection network and the third change detection network all include decoders, and have the same network architecture but independent parameters.
5. The identification method according to claim 1, characterized in that The change classification model is a semi-supervised classification model.
6. A device for identifying a changing pattern, characterized in that: The identification device comprises: An image acquisition module is used to acquire a dual-phase image of the area to be studied, wherein the dual-phase image includes a front-phase remote sensing image and a back-phase remote sensing image; A feature extraction module, configured to extract features from the dual-phase image using a trained feature extraction network to obtain a feature map; a change detection module, configured to perform change detection on the feature map using the trained first change detection network, the second change detection network, and the third change detection network, respectively, to obtain a first prediction mask map, a second prediction mask map, and a third prediction mask map, wherein the second change detection network focuses on global change detection, the first change detection network focuses on missed detection, and the third change detection network focuses on false detection; a classification and recognition module, configured to perform change classification and recognition processing on the bi-phase image, the first prediction mask image, the second prediction mask image, and the third prediction mask image using the trained change classification model to obtain a change classification and recognition result, wherein the change classification and recognition result includes: pseudo change spots and real change spots; The classification and recognition module is specifically configured to: perform feature fusion on the first prediction mask image, the second prediction mask image, and the third prediction mask image to obtain a fourth prediction mask image; perform change classification and recognition processing on the bi-temporal image and the fourth prediction mask image using the trained change classification model to obtain a change classification and recognition result; The change pattern recognition device further includes a training device for constructing an initial feature extraction network, a first change detection network, a second change detection network, and a third change detection network; Performing network training on the feature extraction network and the second change detection network together using a second training set to obtain trained second change detection network and feature extraction network, wherein samples in the second training set include bi-temporal images and true labels, and the true labels are manually annotated change regions; Performing separate network training on the first change detection network using a first training set to obtain a trained first change detection network, wherein the samples in the first training set include outputs of a feature extraction network and missed detection labels, wherein the missed detection labels are missed detection comparison results of a missed detection prediction map corresponding to the output of the feature extraction network and the true label, wherein the missed detection comparison results are used to indicate regions that are present in the true label but not in the missed detection prediction map; The third change detection network is individually trained using a third training set to obtain the trained third change detection network. The samples in the third training set include the output of the feature extraction network and the false detection label. The false detection label is a false detection comparison result between the false detection prediction map corresponding to the output of the feature extraction network and the true label. The false detection comparison result is used to represent an area that is not in the true label but is in the false detection prediction map.
7. A computer program product comprising instructions, characterized in that When the computer program product is run on a computer, the method according to any one of claims 1 to 5 is executed by the computer.
Citation Information
Patent Citations
Remote sensing change detection method based on edge enhanced cross attention and multi-dimensional loss
CN119671934A
Image processing method and apparatus, and electronic device and storage medium
WO2022160753A1