Electric power facility change detection method, device and equipment based on dual-time-phase remote sensing data

By obtaining remote sensing images before and after disasters in the power facility area, preprocessing and feature extraction are performed, and using multi-head self-attention mechanism and pyramid multi-scale feature extraction module, the problem of sparse images before and after disasters and the influence of unrelated factors is solved, and the accuracy of power facility change detection is improved.

CN120339657APending Publication Date: 2025-07-18FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510485322.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing change detection methods, due to uncertain time of disaster occurrence and inconsistent imaging conditions, the number of images before and after the disaster and the excessive unrelated change factors are affected, which affects the accuracy of power facilities' change detection.

Method used

By obtaining remote sensing image pairs of power facilities at the moment before and after disasters, pre-processing and extracting alignment feature pairs of different feature scales, a multi-head self-attention mechanism and a pyramid multi-scale feature extraction module are used to fuse the feature change map for classification prediction, and determine whether there are changes in the power facilities.

Benefits of technology

The accuracy of judging changes in disaster situations in power facilities in the area is improved, and the robustness of the model and the accuracy of change detection are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339657A_ABST
    Figure CN120339657A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power facility change detection method, device and equipment based on dual-time-phase remote sensing data, and the method comprises the steps: obtaining and preprocessing remote sensing image pairs of an electric power facility region before and after a disaster, and generating preprocessed remote sensing image pairs; iteratively extracting alignment feature pairs corresponding to different feature scales from the preprocessed remote sensing image pair; respectively fusing the aligned feature pairs according to each feature scale to obtain difference features; performing feature fusion on each difference feature to obtain a feature change graph; and performing up-sampling on the feature change image according to the size of the remote sensing image pair, then performing classification prediction, and judging whether the electric power facility area is changed or not, thereby effectively improving the accuracy of judging the change of the disaster condition of the electric power facility area in combination with the remote sensing image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of change detection, and particularly to a method, device and equipment for power facility change detection based on dual-temporal remote sensing data. Background Art

[0002] With the progress of technology and the development of society, the demand for monitoring and analyzing the earth's surface has become increasingly urgent. As a key technology, remote sensing image change detection strongly supports the in-depth understanding of the dynamic changes on the earth's surface, and its significance lies not only in the scientific research field, but also plays an important role in daily life decisions such as urban planning, environmental monitoring, and resource management.

[0003] In the power field, due to the large-scale geographical layout of transmission lines, power facilities are extremely vulnerable to natural disasters such as typhoons and floods. Typhoons may cause transmission lines to collapse, and floods will inundate substations, seriously threatening the safety and stability of power facilities and affecting the continuity and reliability of power supply. Therefore, regularly monitoring the surface changes around the transmission line corridors is crucial for promptly detecting potential risks and taking preventive measures. By using remote sensing change detection technology, the surface conditions around transmission lines can be monitored more efficiently, the impact of natural disasters on power facilities can be evaluated, and then the layout and maintenance strategies of power facilities can be optimized, the disaster resistance and recovery capabilities of the power system can be improved, and the safety and stability of power supply can be ensured.

[0004] However, most current change detection methods focus on building change detection, and there is a severe lack of change detection algorithms for power disaster investigation. The existing disaster change detection methods mainly perform change detection based on the disaster scenes in remote sensing images before and after the disaster. Due to the uncertain occurrence time of disasters, the scarce number of images before and after disasters caused by inconsistent imaging conditions of the devices used to capture disaster scenes, and too many irrelevant change factors, the change detection accuracy is affected. Summary of the Invention

[0005] The present invention provides a method, device and equipment for power facility change detection based on dual-temporal remote sensing data, which solves the technical problem that the scarce number of images before and after disasters and too many irrelevant change factors caused by the uncertain occurrence time of disasters and inconsistent imaging conditions of the devices used to capture disaster scenes affect the change detection accuracy.

[0006] A method for power facility change detection based on dual-temporal remote sensing data provided by the present invention includes:

[0007] Obtaining a pair of remote sensing images of the power facility area at different times before and after a disaster and preprocessing them to generate a pair of preprocessed remote sensing images;

[0008] Iteratively extracting corresponding aligned feature pairs corresponding to different feature scales from the pair of preprocessed remote sensing images;

[0009] Fuse the aligned feature pairs according to each of the described feature scales to obtain differential features;

[0010] Perform feature fusion on each of the differential features to obtain a feature change map;

[0011] Upsample the feature change map according to the size of the remote sensing image pair and then perform classification prediction to determine whether there is a change in the power facility area.

[0012] Optionally, the obtaining of the remote sensing image pair of the power facility area before and after the disaster and preprocessing to generate a preprocessed remote sensing image pair includes:

[0013] Obtain the remote sensing image pair of the power facility area before and after the disaster;

[0014] Perform data augmentation on the remote sensing image pair to obtain an augmented image pair;

[0015] Crop the augmented image pair according to a preset cropping size and area ratio to obtain a cropped image pair;

[0016] Perform temporal adjustment on the cropped image pair according to a preset probability value and adjust the image attributes within the cropped data to obtain an image pair to be converted;

[0017] Annotate the image pair to be converted and convert it to the data type required by the model to obtain a preprocessed remote sensing image pair.

[0018] Optionally, the iteratively extracting the aligned feature pairs corresponding to different feature scales from the preprocessed remote sensing image pair includes:

[0019] Call a feature extraction model to extract multiple general feature pairs of different feature scales from the preprocessed remote sensing image pair;

[0020] Call a change detection model to extract remote sensing domain feature pairs of the current feature scale from the preprocessed remote sensing image pair;

[0021] Align the general feature pairs and the remote sensing domain feature pairs with the current feature scale as the benchmark to obtain aligned feature pairs;

[0022] Call the change detection model to downsample and reduce the dimension of the aligned feature pairs, and combine the multi-head self-attention mechanism to extract new remote sensing domain feature pairs corresponding to the current feature scale;

[0023] Jump to execute the step of aligning the general feature pairs and the remote sensing domain feature pairs with the current feature scale as the benchmark to obtain aligned feature pairs until all the general feature pairs are aligned to obtain multiple aligned feature pairs.

[0024] Optionally, the feature extraction model includes multiple encoder layers, and each encoder is replaced with a rotation-variable size window attention mechanism; the step of invoking the feature extraction model to extract general feature pairs of multiple different feature scales from the preprocessed remote sensing image pair includes:

[0025] Perform image chunking on the preprocessed remote sensing image pair to obtain two sets of image chunk sequences, and each set of image chunk sequences includes multiple image chunks;

[0026] Flatten each image chunk into a one-dimensional vector and map it to a preset-dimensional feature space to obtain a projection vector; position encoding is set for each projection vector;

[0027] After normalizing the projection vectors through the encoder layer, extract the attention features of each projection vector in combination with the rotation-variable size window attention mechanism;

[0028] Extract paired attention features output by the encoder layers at multiple preset target layers to obtain general feature pairs of multiple different feature scales.

[0029] Optionally, the step of invoking the change detection model to extract remote sensing domain feature pairs of the current feature scale from the preprocessed remote sensing image pair includes:

[0030] Invoke the change detection model to perform convolutional downsampling on the preprocessed remote sensing image pair to obtain remote sensing feature pairs;

[0031] Multiply the remote sensing feature pairs by a preset weight matrix through the change detection model, and perform reshaping and linear transformation in sequence to obtain two sets of query vectors, key vectors, and value vectors;

[0032] Obtain remote sensing domain feature pairs through the change detection model according to the multi-head attention mechanism and each group of the query vectors, the key vectors, and the value vectors.

[0033] Optionally, the step of aligning the general feature pairs with the remote sensing domain feature pairs based on the current feature scale to obtain aligned feature pairs includes:

[0034] Select general feature pairs of the corresponding feature scale based on the current feature scale;

[0035] Perform layer normalization on the general feature pairs and project them to the same feature dimension as the remote sensing domain feature pairs to obtain intermediate feature pairs;

[0036] Calculate the correlation matrix between the intermediate feature pairs and the remote sensing domain feature pairs and calculate the attention weights in combination with the number of channels of the remote sensing domain features;

[0037] Calculate the dot product of the attention weights and the intermediate feature pairs to obtain weighted feature pairs;

[0038] Perform bilinear interpolation upsampling on the intermediate feature pairs to obtain upsampled feature pairs;

[0039] Use residual connections to stack the general feature pairs, the weighted feature pairs, and the upsampled feature pairs to obtain aligned feature pairs.

[0040] Optionally, fusing the aligned feature pairs according to each feature scale to obtain difference features includes:

[0041] Concatenate the pre-disaster fusion features and the post-disaster fusion features within the aligned feature pairs according to each feature scale to obtain concatenated features;

[0042] Perform two-dimensional convolution on the concatenated features and then perform linear activation to obtain activated features;

[0043] Normalize the activated features to obtain difference features.

[0044] Optionally, performing feature fusion on each of the difference features to obtain a feature change map includes:

[0045] After converting the number of channels of each of the difference features to a target number of channels, perform bilinear interpolation upsampling on each of the difference features to obtain multiple updated features of the same dimension;

[0046] Connect each of the updated features in the channel dimension and convert to the target number of channels to obtain a feature change map.

[0047] The second aspect of the present invention provides a power facility change detection device based on dual-temporal remote sensing data, including:

[0048] An image acquisition and preprocessing module for acquiring a pair of remote sensing images of a power facility area before and after a disaster and preprocessing them to generate a pair of preprocessed remote sensing images;

[0049] A feature pair iterative extraction module for iteratively extracting aligned feature pairs corresponding to different feature scales from the pair of preprocessed remote sensing images;

[0050] A difference feature fusion module for fusing the aligned feature pairs according to each feature scale to obtain difference features;

[0051] A feature change map generation module for performing feature fusion on each of the difference features to obtain a feature change map;

[0052] A change prediction module for upsampling the feature change map according to the size of the pair of remote sensing images and then performing classification prediction to determine whether there is a change in the power facility area.

[0053] In a third aspect of the present invention, an electronic device is provided, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor is caused to execute the steps of the method for detecting changes in power facilities based on dual-temporal remote sensing data according to any one of the first aspects of the present invention.

[0054] As can be seen from the above technical solutions, the present invention has the following advantages:

[0055] The present invention preprocesses a pair of remote sensing images of a power facility area before and after a disaster to generate a pair of preprocessed remote sensing images; iteratively extracts aligned feature pairs corresponding to different feature scales from the pair of preprocessed remote sensing images; fuses the aligned feature pairs according to each feature scale to obtain differential features; performs feature fusion on each differential feature to obtain a feature change map; performs upsampling on the feature change map according to the size of the pair of remote sensing images and then performs classification prediction to determine whether there are changes in the power facility area, thereby effectively improving the accuracy of judging the change in the disaster-affected situation of the power facility area by combining remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0057] Figure 1 It is a flowchart of the steps of a method for detecting changes in power facilities based on dual-temporal remote sensing data provided by an embodiment of the present invention;

[0058] Figure 2 It is a schematic diagram of the integrated model structure of a method for detecting changes in power facilities based on dual-temporal remote sensing data provided by an embodiment of the present invention;

[0059] Figure 3 It is a block diagram of the structure of a device for detecting changes in power facilities based on dual-temporal remote sensing data provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] In specific implementations, the use of pre-trained models in change detection tasks is increasing. Pre-trained models are trained on massive datasets and are widely applied to various downstream tasks in computer vision. Compared with starting training with randomly initialized model parameters, using pre-trained models can capture a large amount of general knowledge and effectively reduce the requirement for data volume. However, current pre-trained models are all pre-trained on natural images, and there are significant differences between natural images and remote sensing images: Remote sensing images are usually from an aerial perspective, lacking the color characteristics of natural images, and have a relatively low spatial resolution, making it difficult to directly use feature extraction models in change detection.

[0061] According to this characteristic of dual-temporal image input in change detection, existing change detection methods can be divided into two categories: image-level based and feature-level based change detection. The former directly concatenates the dual-temporal images along the channel dimension and then feeds them into a single segmentation network such as FCN (Fully Convolutional Networks) for change detection. Feature-level methods use a pair of twin networks with shared weights to independently obtain the features of each single-temporal image, fuse the features of different temporal phases extracted in the previous step through connection or other means, and then use a simple prediction head to predict the output.

[0062] In the feature extraction stage, existing models use convolutional modules because convolutional models can extract discriminative features of dual-temporal images. However, this module is limited by the size of the convolutional kernel and is difficult to focus on and extract large-scale features of the images before and after changes within the spatial and temporal ranges. Models based on Transformer and attention modules effectively solve the problems exposed by convolutional networks. Transformer captures global image information through the self-attention mechanism, is not restricted by local perception, and is suitable for dealing with large-scale changes in remote sensing images. The self-attention mechanism can perform parallel computing, making the Transformer model easier to train on large-scale data; the attention-based network model dynamically adjusts the information of parts such as channels and spaces of the input data by learning weights. In the change detection of remote sensing images, this model can better capture the key information of the changed areas. In the power scenario, there is currently no such change detection method. Using this method can capture the key information of the changed areas in the power scenario before and after disasters, which helps to improve the effect of change detection in the power scenario.

[0063] To this end, the embodiments of the present invention provide a method, device, and equipment for detecting changes in power facilities based on dual-temporal remote sensing data. By sequentially extracting the features of power remote sensing images in two temporal phases, the temporal features of the images are effectively extracted. At the same time, a feature extraction model for multiple remote sensing tasks is used to extract general knowledge in the remote sensing field and embed it into the change detection features, effectively solving the problem of different domains between natural images and remote sensing images. On this basis, the model uses a pyramid multi-scale feature extraction module to better retain the details and global structural features of the images to adapt to the pixel-level classification task of change detection. Finally, the model uses an attention-based alignment module to efficiently fuse features of different sizes, automatically extract the feature information useful for change detection in the feature extraction model, and efficiently fuse them for feature extraction. Thus, the technical problem that the number of images before and after a disaster is scarce and there are too many irrelevant change factors due to the uncertain occurrence time of the disaster and the inconsistent imaging conditions of the devices used to capture the disaster scene, which in turn affects the accuracy of change detection, is solved.

[0064] In order to make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0065] Please refer to Figure 1 , Figure 1 which is a flowchart of the steps of a method for detecting changes in power facilities based on dual-temporal remote sensing data provided by the embodiments of the present invention.

[0066] A method for detecting changes in power facilities based on dual-temporal remote sensing data provided by the present invention includes:

[0067] Step 101: Obtain a pair of remote sensing images of the power facility area before and after a disaster and preprocess them to generate a pair of preprocessed remote sensing images;

[0068] The pair of remote sensing images includes a pre-disaster remote sensing image and a post-disaster remote sensing image, which refers to the images of the area where the power facilities are located obtained by remote sensing technology, that is, without directly contacting the object, at a certain distance, using a sensor to receive the electromagnetic wave information reflected and radiated by the target ground object, and processing and analyzing it to reveal the characteristics, properties, and changes of the target ground object.

[0069] In an embodiment of the present invention, remote sensing image pairs of the power facility area at pre-disaster and post-disaster moments can be obtained. Each group of remote sensing image pairs includes a pre-disaster remote sensing image and a post-disaster remote sensing image, and each is respectively labeled with a corresponding label. Considering the high resolution of the remote sensing images and the fact that the sizes of each pair of images are not necessarily the same. For each single image of each pair, it needs to be pre-cut into non-overlapping images of 256×256 pixels in size, ensuring that the same position of the same pair of images is cut. Then, data augmentation is performed before the images are input into the model to obtain pre-processed remote sensing image pairs.

[0070] In an example of the present invention, step 101 may include the following sub-steps:

[0071] Obtain remote sensing image pairs of the power facility area at pre-disaster and post-disaster moments;

[0072] Perform data augmentation on the remote sensing image pairs to obtain augmented image pairs;

[0073] Crop the augmented image pairs according to a preset cropping size and area occupancy ratio to obtain cropped image pairs;

[0074] Perform temporal adjustment on the cropped image pairs according to a preset probability value and adjust the image attributes within the cropped data to obtain image pairs to be converted;

[0075] Label the image pairs to be converted and convert them into the data type required by the model to obtain pre-processed remote sensing image pairs.

[0076] In this embodiment, remote sensing image pairs of the power facility area at pre-disaster and post-disaster moments are obtained, and data augmentation is performed on the remote sensing image pairs to obtain augmented image pairs, that is, randomly rotate and translate the images. Specifically, there is a 50% probability of randomly rotating the image, and the rotation angle range is between -20 degrees and 20 degrees; there is also a 50% probability of horizontal or vertical flipping, so as to increase the diversity of training data and reduce the risk of model overfitting.

[0077] Furthermore, according to a preset cropping size and area occupancy ratio, such as the cropping size is 256×256, and the maximum proportion of the target category in the cropped area cannot exceed 0.75, the augmented image pairs are cropped to obtain cropped image pairs. This ensures that the cropped area contains both the target object and the background, which is beneficial for the model to learn more comprehensive features.

[0078] After the cropping of the image pair is completed, the time series of the cropped image pair is adjusted according to a preset probability value, that is, the images before and after the disaster change are exchanged, and the time order in the image sequence is randomly exchanged with a probability of 50%. This effectively enhances the time series data and helps the model learn time-independent features. In addition, the image attributes within the cropped data are adjusted, such as adjusting the brightness of the image (randomly increasing or decreasing by 10 on the current brightness), contrast (adjusted to between 0.8 and 1.2), saturation (in the same range as the contrast), and hue (randomly changing by 10 units), to obtain the image pair to be transformed, thereby simulating images under different lighting conditions and enhancing the robustness of the model.

[0079] Finally, the image pair to be transformed is labeled and converted into the data type required by the model to obtain the preprocessed remote sensing image pair, which includes converting the image data into a specific data type (such as Tensor) and organizing the labeled data to match the input requirements of the model to ensure the smooth progress of the training process.

[0080] Step 102: Iteratively extract the aligned feature pairs corresponding to different feature scales from the preprocessed remote sensing image pair;

[0081] In this embodiment, by calling the feature extraction model, the general feature pairs of different feature scales are extracted from the preprocessed remote sensing image pair. At the same time, the remote sensing domain feature pairs extracted from the preprocessed remote sensing image pair by the change detection model are iteratively aligned to fuse and obtain the aligned feature pairs of different feature scales.

[0082] In an example of the present invention, step 102 may include the following sub-steps S11-S15:

[0083] S11: Call the feature extraction model to extract multiple general feature pairs of different feature scales from the preprocessed remote sensing image pair;

[0084] The general feature pairs include the pre-disaster general features and the post-disaster general features, which are used to obtain various image information in the preprocessed remote sensing image and wait for further screening later.

[0085] Considering the variable directions of remote sensing targets due to the bird's-eye view, a learnable rotation angle factor is additionally introduced into the feature extraction model to expand the variable-size window attention in the vision Transformer, generating adaptive scaling, translation, and rotation windows, so that the feature extraction model can extract multiple general feature pairs of different feature scales from the preprocessed remote sensing image pair.

[0086] Furthermore, the feature extraction model includes multiple encoder layers, and each encoder is replaced with a rotation-variable size window attention mechanism; S11 may include the following sub-steps:

[0087] The preprocessed remote sensing image pair is subjected to image chunking to obtain two sets of image chunk sequences, and each set of image chunk sequences includes multiple image chunks;

[0088] Each image chunk is flattened into a one-dimensional vector and mapped to a preset dimensional feature space to obtain a projection vector; position encodings are set for all the projection vectors;

[0089] After normalizing the projection vectors through the encoder layer, the attention features of each projection vector are extracted by combining the rotation-variable-sized window attention mechanism;

[0090] The paired attention features output by the encoder layer at multiple preset target layers are extracted to obtain multiple general feature pairs with different feature scales.

[0091] In this embodiment, the preprocessed remote sensing image pair is subjected to image chunking to obtain two sets of image chunk sequences, and each set of image chunk sequences includes multiple image chunks. For each image chunk within each set of image chunk sequences, each image chunk is respectively flattened into a one-dimensional vector and mapped to a preset dimensional feature space to obtain a projection vector, and position encodings are set for all the projection vectors. The feature extraction model can be pre-trained and can be a VIT (Vision Transformer) model, which is a model that applies the Transformer architecture to the field of images.

[0092] For each projection vector, taking it as the input feature (where C, H, and W are the number of channels, height, and width in I respectively) it is evenly divided into different windows, and the feature of each window can be expressed as (p is the window size), and a total of windows are obtained. Then, three linear layers are used to generate the query feature and the initial key and value features, which are respectively expressed as and , and then is used to predict the change of the window:

[0093]

[0094] where GAP represents global average pooling, is the scaling factor on the X-axis and Y-axis, is the offset on the X-axis and Y-axis, and in addition is the rotation angle. Taking the corner points of the window as an example:

[0095]

[0096] where are the coordinates of the upper left corner and the lower right corner of the initial window, is the coordinate of the center point of the window. , , are the distances between the corner point and the center point in the horizontal and vertical directions respectively. The transformation of the window can be achieved by using the obtained scaling, translation, and rotation factors:

[0097]

[0098] where are the coordinates of the corner points of the window after transformation. New key vectors and value vectors and are sampled from the obtained window, and the self-attention operation is performed through the following formula:

[0099]

[0100] is the feature output by the self-attention operation of a window, , where h is the number of self-attention (SA). The shape of the final output feature in the attention of the rotation-variable-size window is restored by concatenating the features of different self-attention in the channel dimension and merging the features of different windows along the spatial dimension to obtain the attention feature of each projection vector. The attention of the rotation-variable-size window is used to replace the multi-head full attention in the encoder layer of the original VIT model.

[0101] Extract the paired attention features output by the encoder layers at multiple preset target levels. For example, if the entire feature extraction model contains 24 layers, here the 7th, 11th, 15th, and 23rd layers are used as the preset target levels for extracting the encoder layer output. The deeper the layer, the more abstract the extracted features, and the shallower the layer, the finer the extracted features, so as to obtain multiple general feature pairs with different feature scales.

[0102] S12. Call the change detection model to extract the remote sensing field feature pairs at the current feature scale from the preprocessed remote sensing image pair;

[0103] In this embodiment, while the feature extraction model extracts the general feature pairs, the change detection model is called to extract the remote sensing field feature pairs at different feature scales from the preprocessed remote sensing image pair. Since the change detection model needs to iteratively extract at different feature scales, therefore, at the first extraction, the shallow feature scale is used as the current feature scale for extracting the remote sensing field feature pairs.

[0104] Further, S12 may include the following sub-steps:

[0105] Call the change detection model to perform convolutional downsampling on the preprocessed remote sensing image pair to obtain the remote sensing feature pairs;

[0106] The change detection model multiplies the remote sensing feature pairs using a preset weight matrix, and then performs reshaping and linear transformation in sequence to obtain two sets of query vectors, key vectors, and value vectors;

[0107] The change detection model obtains the remote sensing domain feature pairs according to the multi-head attention mechanism and each group of query vectors, key vectors, and value vectors.

[0108] In this embodiment, for each preprocessed remote sensing image in the preprocessed remote sensing image pair, a multi-level Transformer Block encoder is also used to extract multi-level features. The shallow features have high resolution and retain the detailed features of the original image, while the deep features have low resolution and extract the abstract features of the overall image.

[0109] Specifically, each preprocessed remote sensing image in the preprocessed remote sensing image pair is respectively input into the change detection model as input features (3 indicates that the image is in RGB format), and each Block will output features as tensors of dimension, where l represents the number of layers (layer), and the value of l is {1, 2, 3, 4}. Four types of Blocks are used in this part to extract four layers of features. After passing through each Block, downsampling is used and the number of channels is increased, which means .

[0110] Processing the features of the l-th layer , first downsample it to get . Here, a convolutional module is used to implement downsampling. A convolution with a 7×7 kernel size, a stride of 4, and a padding of 3 is used as the initial downsampling. For the subsequent three downsamplings, convolutions with a kernel size of 3×3, a stride of 2, and a padding of 1 are used, represents the downsampling process.

[0111]

[0112] The main building block of each Block is the self-attention module. The attention method adopted in the original Transformer is as follows:

[0113]

[0114] Among them , and respectively represent Query, Key, and Value. The dimensions of these three features are the same, all , and formula (9) has a high computational complexity , before using the attention mechanism, first perform dimensionality reduction on it , see formulas (10)-(11).

[0115]

[0116]

[0117] For the original feature , The operation means to convert into a tensor with a feature of , where r represents a hyperparameter, i.e., the scaling ratio. Then use the linear layer to remap into a feature of . For , a dimension-reduced feature is generated. For and it is the same. The computational complexity is reduced to .

[0118] Since in the attention mechanism, the dimension is expanded into one dimension, to fuse the position information, two multi-layer perceptrons and convolutions are used to add the position information:

[0119]

[0120] where is the feature of self-attention, GELU represents the Gaussian error linear unit activation. Using this position encoding is different from the fixed position encoding in ViT and can adapt to images with different resolutions.

[0121] The method for inputting each Block into the next Block is as follows:

[0122] First use downsampling and then generate the corresponding , where j represents the number of multi-head attention heads. Then perform dimensionality reduction on these three features and use multi-head self-attention to obtain the output . Finally, add the position information and combine the general knowledge to obtain the input of the next Block:

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130] For each obtained in each step will be used as the in step S13, and then is fused to obtain the output which is used as the input of the next Block Use to represent the process of step S13. The processed original image passes through four Blocks to obtain four different-scale remote sensing field feature pairs.

[0131] S13: Align the general feature pair and the remote sensing field feature pair based on the current feature scale to obtain the aligned feature pair;

[0132] Furthermore, S13 may include the following sub-steps:

[0133] Select the general feature pair corresponding to the current feature scale based on the current feature scale;

[0134] Perform layer normalization on the general feature pair and project it to the same feature dimension as the remote sensing field feature pair to obtain the intermediate feature pair;

[0135] Calculate the correlation matrix between the intermediate feature pair and the remote sensing field feature pair and calculate the attention weight in combination with the number of channels of the remote sensing field feature;

[0136] Calculate the dot product of the attention weight and the intermediate feature pair to obtain the weighted feature pair;

[0137] Perform bilinear interpolation upsampling on the intermediate feature pair to obtain the upsampled feature pair;

[0138] Adopt residual connection to stack the general feature pair, the weighted feature pair and the upsampled feature pair to obtain the aligned feature pair.

[0139] The features generated by pre-training and the feature sizes of the change detection model are often not necessarily the same, and not all pre-trained knowledge is necessarily effective. Specifically, two problems need to be considered: First, not all general knowledge is necessary, and it is necessary to determine how to filter out useless information and retain valuable general feature pairs; Second, in the feature extraction model, there are a total of 24 layers, and the features in the basic model are reduced by at least 10 times, and the obtained resolution is extremely low. It is necessary to fuse the highly abstract features with the change detection module.

[0140] In this embodiment, for the general features within the general feature pair , use as the remote sensing field features within the features extracted by the change detection model.

[0141] First, use a layer normalization LN to adjust the distribution of each general feature within the general feature pair to make it consistent with the general feature, and then use a linear layer to project to , obtaining the intermediate features within the intermediate feature pair :

[0142]

[0143] To effectively distinguish valuable features in general knowledge, use cosine metric to calculate and 's correlation matrix , and then divide by and use the softmax method to obtain the attention weights , finally use and to perform a dot product operation to obtain the weighted features :

[0144]

[0145] Convert back into a tensor with a dimension of , extract both the useful part of the features and solve the problem of scale misalignment, and finally use a simple residual connection to obtain the fused features at this level .

[0146]

[0147] Among them, represents using bilinear interpolation upsampling to obtain the upsampled features.

[0148] S14. Call the change detection model to downsample and reduce the dimension of the aligned feature pair, and combine the multi-head self-attention mechanism to extract the remote sensing field feature pair corresponding to the new current feature scale;

[0149] S15. Jump to execute the step of aligning the general feature pair and the remote sensing field feature pair based on the current feature scale until all general feature pairs are aligned, obtaining multiple aligned feature pairs.

[0150] In this embodiment, for each obtained in each step, it will be used as in step S13, and then is fused to obtain the output , use As the input for the next Block . Use to represent the process of step S13. The processed original image passes through four Blocks to obtain four different-scale remote sensing field feature pairs.

[0151] Step 103: Fuse and align the feature pairs according to each feature scale to obtain differential features;

[0152] In this embodiment, the aligned feature pairs include pre-disaster fusion features and post-disaster fusion features. For each different feature scale, fuse the pre-disaster fusion features and post-disaster fusion features at each feature scale respectively to obtain differential features at different feature scales, so as to determine the temporal change of the power facility area.

[0153] In an example of the present invention, step 103 may include the following sub-steps:

[0154] Stitch the pre-disaster fusion features and post-disaster fusion features in the aligned feature pairs according to each feature scale to obtain stitched features;

[0155] Perform two-dimensional convolution on the stitched features and then perform linear activation to obtain activated features;

[0156] Normalize the activated features to obtain differential features.

[0157] In this embodiment, four different-sized pre-disaster fusion features and post-disaster fusion features are obtained for the two time phases before and after the change, and the differential features are fused in the following manner , :

[0158]

[0159] Among them, Concat represents concatenating tensors in the channel dimension, and concatenating the pre-disaster fusion features and post-disaster fusion features in the channel dimension. In this way, the feature information of the two time phases is integrated together to prepare for subsequent extraction of differential information. Perform two-dimensional convolution operation on the stitched feature map, aiming to extract the spatial features and local patterns in the fused features. Apply the ReLU activation function to the result of the convolution to introduce non-linear factors, which helps the model learn more complex feature representations, enhances the expression ability of the model, and at the same time alleviates the problem of gradient disappearance, making the network training more efficient. Perform batch normalization processing on the features after ReLU activation for layers to adjust the distribution of the features to a standard distribution with a mean close to 0 and a variance close to 1.

[0160] Step 104: Perform feature fusion on each differential feature to obtain a feature change map;

[0161] In this embodiment, by performing feature fusion on each differential feature, the differential features of different feature scales are fused, and a feature change map for describing disaster change prediction is generated.

[0162] In an example of the present invention, step 104 may include the following sub-steps:

[0163] After converting the number of channels of each differential feature to the target number of channels, perform bilinear interpolation upsampling on each differential feature to obtain multiple updated features of the same dimension;

[0164] Connect each updated feature in the channel dimension and convert it to the target number of channels to obtain a feature change map.

[0165] In this embodiment, first use a 1×1 convolutional layer to convert the number of channels of each feature into a unified number of channels 128, then use bilinear interpolation to upsample the four features to a unified dimension, and finally connect the four features in the channel dimension and use a 1×1 convolution to map the number of channels back to 128 to obtain a feature change map :

[0166]

[0167]

[0168]

[0169] Step 105: After upsampling the feature change map according to the size of the remote sensing image pair, perform classification prediction to determine whether there is a change in the power facility area.

[0170] In this embodiment, after the generation of the feature change map is completed, the feature change map can be upsampled according to the size of any remote sensing image in the remote sensing image pair to obtain a feature change map of the same size. Then use a multi-layer perceptron to process according to the predicted category to generate a predicted change map CM, where the upsampling uses a transposed convolution with a stride of 2 and a convolution kernel size of 3. The predicted categories are two types, namely change and unchanged, so as to determine whether there is a change in the power facility area before and after the disaster.

[0171] In addition, as Figure 2 shown, the entire process between steps 101-105 in the embodiment of the present invention can be implemented in the form of an integrated model. For the training and evaluation of the integrated model, pixel-level cross-entropy loss can be used for evaluation, and the loss function is as follows:

[0172]

[0173] Where H and W respectively predict the height and width of the output of the change map, and respectively represent the label and the predicted change map the true label and the predicted value of the position.

[0174] The experimental results of its model evaluation are shown in Table 1 below:

[0175] Table 1

[0176]

[0177] Class: Represents the category to be detected.

[0178] "unchanged": Refers to the area in the image that has not changed.

[0179] "changed": Refers to the area in the image that has changed.

[0180] Fscore, also known as the F1 score, is the harmonic mean of precision and recall, and is used to measure the performance of a classification model. The higher the F1 score, the better the performance of the model.

[0181] Precision: Precision, which represents the proportion of samples correctly identified as this category by the model among all samples identified as this category by the model.

[0182] Recall: Recall, also known as the true positive rate or sensitivity, represents the proportion of all samples that are actually of this category and are correctly identified by the model.

[0183] IoU: Intersection over Union, is an indicator to measure the overlap degree of two sets, and is usually used in image segmentation tasks. In this context, it may be used to measure the overlap degree between the changed area identified by the model and the actual changed area.

[0184] Acc: Accuracy, represents the proportion of all samples that are correctly classified.

[0185] It can be seen from the table that:

[0186] The Fscore of the "changed" category is 88.85%, the Precision is 89.33%, the Recall is 88.37%, the IoU is 79.93%, and the accuracy is 88.37%. This indicates that the model performs well in identifying the changed area and can accurately detect the changes caused by disasters.

[0187] In an embodiment of the present invention, by obtaining and preprocessing a pair of remote sensing images of a power facility area before and after a disaster, a pair of preprocessed remote sensing images is generated; aligned feature pairs corresponding to different feature scales are iteratively extracted from the pair of preprocessed remote sensing images; the aligned feature pairs are fused according to each feature scale to obtain differential features; the differential features are feature-fused to obtain a feature change map; the feature change map is upsampled according to the size of the pair of remote sensing images and then classified and predicted to determine whether there is a change in the power facility area, thereby effectively improving the accuracy of judging the change in the disaster-affected situation of the power facility area in combination with the remote sensing images.

[0188] Please refer to Figure 3 , Figure 3 which is a structural block diagram of a power facility change detection device based on dual-temporal remote sensing data provided by an embodiment of the present invention.

[0189] An embodiment of the present invention provides a power facility change detection device based on dual-temporal remote sensing data, including:

[0190] An image acquisition and preprocessing module 301, configured to obtain and preprocess a pair of remote sensing images of a power facility area before and after a disaster, and generate a pair of preprocessed remote sensing images;

[0191] A feature pair iterative extraction module 302, configured to iteratively extract aligned feature pairs corresponding to different feature scales from the pair of preprocessed remote sensing images;

[0192] A differential feature fusion module 303, configured to fuse the aligned feature pairs according to each feature scale to obtain differential features;

[0193] A feature change map generation module 304, configured to perform feature fusion on each differential feature to obtain a feature change map;

[0194] A change prediction module 305, configured to perform upsampling on the feature change map according to the size of the pair of remote sensing images and then perform classification and prediction to determine whether there is a change in the power facility area.

[0195] Optionally, the image acquisition and preprocessing module 301 is specifically configured to:

[0196] Obtain a pair of remote sensing images of a power facility area before and after a disaster;

[0197] Perform data enhancement on the pair of remote sensing images to obtain an enhanced pair of images;

[0198] Crop the enhanced pair of images according to a preset cropping size and regional occupancy ratio to obtain a cropped pair of images;

[0199] Perform temporal adjustment on the cropped pair of images according to a preset probability value and adjust the image attributes within the cropped data to obtain a pair of images to be transformed;

[0200] Label the image pairs to be converted and convert them into the model requirement data type to obtain the preprocessed remote sensing image pairs.

[0201] Optionally, the feature pair iterative extraction module 302 includes:

[0202] A general feature pair extraction sub-module for calling a feature extraction model to extract general feature pairs of multiple different feature scales from the preprocessed remote sensing image pairs;

[0203] A remote sensing field feature pair extraction sub-module for calling a change detection model to extract remote sensing field feature pairs of the current feature scale from the preprocessed remote sensing image pairs;

[0204] An aligned feature pair generation sub-module for aligning the general feature pairs and the remote sensing field feature pairs based on the current feature scale to obtain aligned feature pairs;

[0205] A remote sensing field feature pair update sub-module for calling a change detection model to downsample and reduce the dimension of the aligned feature pairs, and combining the multi-head self-attention mechanism to extract new remote sensing field feature pairs corresponding to the current feature scale;

[0206] An iterative extraction sub-module for jumping to execute the step of aligning the general feature pairs and the remote sensing field feature pairs based on the current feature scale to obtain aligned feature pairs until all the general feature pairs are aligned, and obtaining multiple aligned feature pairs.

[0207] Optionally, the feature extraction model includes multiple encoder layers, and each encoder is replaced with a rotary variable-sized window attention mechanism; the general feature pair extraction sub-module is specifically used for:

[0208] Perform image chunking on the preprocessed remote sensing image pairs to obtain two groups of image chunk sequences, and each group of image chunk sequences includes multiple image chunks;

[0209] Flatten each image chunk into a one-dimensional vector and map it to a preset-dimensional feature space to obtain a projection vector; position encoding is set for each projection vector;

[0210] After normalizing the projection vectors through the encoder layers, combine the rotary variable-sized window attention mechanism to extract the attention features of each projection vector;

[0211] Extract the paired attention features output by the encoder layers at multiple preset target layers to obtain general feature pairs of multiple different feature scales.

[0212] Optionally, the remote sensing field feature pair extraction sub-module is specifically used for:

[0213] Call a change detection model to perform convolutional downsampling on the preprocessed remote sensing image pairs to obtain remote sensing feature pairs;

[0214] The change detection model multiplies the remote sensing feature pairs using a preset weight matrix, and then performs reshaping and linear transformation in sequence to obtain two sets of query vectors, key vectors, and value vectors;

[0215] The change detection model obtains the remote sensing domain feature pairs according to the multi-head attention mechanism and each set of query vectors, key vectors, and value vectors.

[0216] Optionally, the alignment feature pair generation sub-module is specifically used for:

[0217] Select the general feature pairs corresponding to the current feature scale based on the current feature scale;

[0218] Perform layer normalization on the general feature pairs and project them to the same feature dimension as the remote sensing domain feature pairs to obtain intermediate feature pairs;

[0219] Calculate the correlation matrix between the intermediate feature pairs and the remote sensing domain feature pairs and calculate the attention weights in combination with the number of channels of the remote sensing domain features;

[0220] Calculate the dot product of the attention weights and the intermediate feature pairs to obtain weighted feature pairs;

[0221] Perform bilinear interpolation upsampling on the intermediate feature pairs to obtain upsampled feature pairs;

[0222] Adopt residual connection to stack the general feature pairs, weighted feature pairs, and upsampled feature pairs to obtain aligned feature pairs.

[0223] Optionally, the difference feature fusion module 303 is specifically used for:

[0224] Concatenate the pre-disaster fusion features and post-disaster fusion features in the aligned feature pairs according to each feature scale respectively to obtain concatenated features;

[0225] Perform two-dimensional convolution on the concatenated features and then perform linear activation to obtain activated features;

[0226] Normalize the activated features to obtain difference features.

[0227] Optionally, the feature change map generation module 304 is specifically used for:

[0228] After converting the number of channels of each difference feature to the target number of channels, perform bilinear interpolation upsampling on each difference feature to obtain multiple updated features of the same dimension;

[0229] Connect each updated feature in the channel dimension and convert it to the target number of channels to obtain a feature change map.

[0230] An embodiment of the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the power facility change detection method based on dual-temporal remote sensing data as described in any embodiment of the present invention.

[0231] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, modules, and sub-modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0232] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0233] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0234] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0235] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting changes in power facilities based on dual-temporal remote sensing data, characterized in that Including: Obtain remote sensing image pairs of the power facility area before and after a disaster and perform preprocessing to generate preprocessed remote sensing image pairs; Iteratively extract aligned feature pairs corresponding to different feature scales from the preprocessed remote sensing image pairs; Fuse the aligned feature pairs according to each of the feature scales to obtain differential features; Perform feature fusion on each of the differential features to obtain a feature change map; Upsample the feature change map according to the size of the remote sensing image pair and then perform classification prediction to determine whether there are changes in the power facility area.

2. The method according to claim 1, characterized in that, The obtaining remote sensing image pairs of the power facility area before and after a disaster and performing preprocessing to generate preprocessed remote sensing image pairs includes: Obtain remote sensing image pairs of the power facility area before and after a disaster; Perform data augmentation on the remote sensing image pairs to obtain augmented image pairs; Crop the augmented image pairs according to a preset cropping size and regional occupancy ratio to obtain cropped image pairs; Perform temporal adjustment on the cropped image pairs according to a preset probability value and adjust the image attributes within the cropped data to obtain image pairs to be converted; Label the image pairs to be converted and convert them into the data type required by the model to obtain preprocessed remote sensing image pairs.

3. The method according to claim 1, wherein The iteratively extracting aligned feature pairs corresponding to different feature scales from the preprocessed remote sensing image pairs includes: Call a feature extraction model to extract a plurality of general feature pairs of different feature scales from the preprocessed remote sensing image pairs; Call a change detection model to extract remote sensing domain feature pairs of the current feature scale from the preprocessed remote sensing image pairs; Align the general feature pairs and the remote sensing domain feature pairs based on the current feature scale to obtain aligned feature pairs; Call the change detection model to downsample and reduce the dimension of the aligned feature pairs, and combine the multi-head self-attention mechanism to extract new remote sensing domain feature pairs corresponding to the current feature scale; Jump to execute the step of aligning the general feature pairs and the remote sensing domain feature pairs based on the current feature scale to obtain aligned feature pairs until all the general feature pairs are aligned to obtain a plurality of aligned feature pairs.

4. The method according to claim 3, characterized in that, The feature extraction model includes a plurality of encoder layers, and each encoder is replaced with a rotation-variable size window attention mechanism; the calling the feature extraction model to extract a plurality of general feature pairs of different feature scales from the preprocessed remote sensing image pairs includes: Perform image chunking on the preprocessed remote sensing image pairs to obtain two groups of image chunk sequences, and each group of the image chunk sequences includes a plurality of image chunks; Flatten each image chunk into a one-dimensional vector and map it to a preset-dimensional feature space to obtain a projection vector; position encoding is provided for each projection vector; After normalizing the projection vectors through the encoder layers, combine the rotation-variable size window attention mechanism to extract the attention features of each projection vector; Extract paired attention features output by the encoder layers at a plurality of preset target layers to obtain a plurality of general feature pairs of different feature scales.

5. The method according to claim 3, characterized in that, The calling the change detection model to extract remote sensing domain feature pairs of the current feature scale from the preprocessed remote sensing image pairs includes: Call the change detection model to perform convolutional downsampling on the preprocessed remote sensing image pair to obtain a remote sensing feature pair; Multiply the remote sensing feature pair by the preset weight matrix through the change detection model, and perform reshaping and linear transformation in sequence to obtain two groups of query vectors, key vectors, and value vectors; According to the multi-head attention mechanism and each group of the query vectors, the key vectors, and the value vectors, the change detection model obtains a remote sensing domain feature pair.

6. The method according to claim 3, characterized in that, Aligning the general feature pair and the remote sensing domain feature pair based on the current feature scale to obtain an aligned feature pair, including: Select the general feature pair corresponding to the corresponding feature scale based on the current feature scale; Perform layer normalization on the general feature pair and project it to the same feature dimension as the remote sensing domain feature pair to obtain an intermediate feature pair; Calculate the correlation matrix between the intermediate feature pair and the remote sensing domain feature pair and calculate the attention weight in combination with the number of channels of the remote sensing domain feature; Calculate the dot product of the attention weight and the intermediate feature pair to obtain a weighted feature pair; Perform bilinear interpolation upsampling on the intermediate feature pair to obtain an upsampled feature pair; Use residual connection to stack the general feature pair, the weighted feature pair, and the upsampled feature pair to obtain an aligned feature pair.

7. The method according to claim 1, wherein Fusing the aligned feature pairs according to each feature scale respectively to obtain a difference feature, including: Stitch the pre-disaster fusion feature and the post-disaster fusion feature in the aligned feature pair according to each feature scale respectively to obtain a stitched feature; Perform two-dimensional convolution on the stitched feature and then perform linear activation to obtain an activated feature; Normalize the activated feature to obtain a difference feature.

8. The method according to claim 1, wherein Perform feature fusion on each of the difference features to obtain a feature change map, including: After converting the number of channels of each difference feature to the target number of channels, perform bilinear interpolation upsampling on each difference feature to obtain multiple updated features of the same dimension; Connect each of the updated features in the channel dimension and convert it to the target number of channels to obtain a feature change map.

9. A power facility change detection device based on dual-temporal remote sensing data, characterized in that, Including: An image acquisition and preprocessing module, configured to acquire and preprocess remote sensing image pairs of the power facility area at pre-disaster and post-disaster moments to generate preprocessed remote sensing image pairs; A feature pair iterative extraction module, configured to iteratively extract aligned feature pairs corresponding to different feature scales from the preprocessed remote sensing image pairs; A difference feature fusion module, configured to fuse the aligned feature pairs according to each feature scale respectively to obtain a difference feature; A feature change map generation module, configured to perform feature fusion on each of the difference features to obtain a feature change map; A change prediction module, configured to perform upsampling on the feature change map according to the size of the remote sensing image pair and then perform classification prediction to determine whether there is a change in the power facility area.

10. An electronic device, characterized in that, Including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for detecting power facility changes based on dual-temporal remote sensing data according to any one of claims 1-8.