Urban change detection method and device based on global guidance-local feature alignment

By adopting a global guidance-local feature alignment method, the problem of single feature fusion and alignment strategies in existing technologies is solved, achieving high accuracy and robustness in urban change detection and adapting to change detection in complex urban scenarios.

CN122049694BActive Publication Date: 2026-07-07XIDIAN UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2026-04-14
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Existing urban change detection methods lack adaptability in feature fusion strategies and feature alignment strategies, resulting in inaccurate detection results and difficulty in adapting to complex and diverse urban change scenarios.

Method used

A method based on global guidance and local feature alignment is adopted. By acquiring a multi-scale feature pyramid, combining global guidance information and local window search for feature alignment and fusion, and using a gating mechanism for adaptive fusion, the change detection results are finally generated.

Benefits of technology

It improves the accuracy and generalization ability of change detection, enhances the stability and robustness of feature alignment, and adapts to change detection in complex urban scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049694B_ABST
    Figure CN122049694B_ABST
Patent Text Reader

Abstract

This invention discloses a method and device for urban change detection based on global guidance and local feature alignment, relating to the field of remote sensing monitoring technology. The invention acquires first and second period imagery of the area to be detected and extracts its multi-scale features and guidance features, obtaining first and second features at I levels and an initial guidance feature. At each level, feature projection is performed on the guidance feature, the first and second features, and global guidance information for the current level is generated based on the projected guidance features. Local feature alignment is performed on the projected first and second features based on local window search and the global guidance information. Adaptive feature fusion is performed based on the local feature alignment results and a gating mechanism, and the fused features of the current level are used as the guidance features for the next level. Decoding is performed based on the fused features of all levels. A change detection result map is generated based on the decoded features. This invention can effectively improve the accuracy and robustness of change detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing monitoring technology, specifically relating to a method and device for detecting urban changes based on global guidance and local feature alignment. Background Technology

[0002] With the development of high-resolution remote sensing satellites, various sensors mounted on platforms such as satellites, aircraft, and drones can acquire abundant remote sensing images. Among numerous remote sensing applications, change detection aims to identify and extract changes in land cover by analyzing remote sensing images from different time periods. Related technologies for change detection are also widely used in areas such as dynamic monitoring of land use, environmental protection, and natural disaster assessment.

[0003] In recent years, the global urbanization process has accelerated, urban space has continued to expand, and infrastructure construction has become increasingly frequent, posing enormous challenges to urban development, the ecological environment, and the socio-economic development. Quickly and accurately grasping information on urban changes is crucial for scientifically formulating urban plans, optimizing land resource allocation, monitoring environmental quality, and responding to sudden disasters. Compared to traditional ground surveys, remote sensing technology plays an irreplaceable role in monitoring dynamic urban changes due to its significant advantages such as wide coverage, short acquisition cycle, and rich information dimensions. By analyzing remote sensing images from different periods, various urban change patterns, such as urban expansion, building construction and demolition, changes in road networks, and changes in vegetation cover, can be effectively identified, providing strong decision support for urban managers. However, some current monitoring methods use fixed fusion rules for feature fusion, resulting in a single feature fusion strategy and a lack of adaptive selection capabilities, making it difficult to adapt to the complex and diverse change patterns in urban scenarios. Furthermore, some monitoring methods employ strict hard alignment strategies, either globally or locally, which cannot dynamically determine whether alignment should be performed based on the confidence level of image content (such as occlusion or shadow areas), easily leading to incorrect matches and inaccurate detection results. Summary of the Invention

[0004] To address the aforementioned problems in the existing technology, this invention provides a method and device for detecting urban changes based on global guidance and local feature alignment.

[0005] The technical problem to be solved by this invention is achieved through the following technical solution:

[0006] This invention provides a method for detecting urban changes based on global guidance and local feature alignment, comprising:

[0007] Obtain the first-period and second-period image images of the area to be detected;

[0008] Multi-scale features are extracted from the first period image and the second period image respectively to obtain a first feature pyramid and a second feature pyramid. The first feature pyramid contains I different levels of first features, and the second feature pyramid contains I different levels of second features, where I is a positive integer greater than or equal to 1.

[0009] Feature projection is performed on the current level's guiding features, the current level's first feature, and the current level's second feature. Based on the projected guiding features of the current level, global guiding information for the current level is generated. Based on local window search and the current level's global guiding information, local feature alignment is performed on the projected first feature and the projected second feature of the current level. Based on the local feature alignment result and gating mechanism, adaptive feature fusion is performed, and the resulting fused features of the current level are used as the guiding features of the next level. The guiding features of the first level are obtained by stitching together the first period image and the second period image and then extracting features.

[0010] Based on the guiding features of all levels, a bottom-up decoding process is performed to obtain the final decoded features;

[0011] Based on the final decoded features, a change detection result map of the region to be detected is generated.

[0012] The present invention also provides an urban change detection device based on global guidance-local feature alignment, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus;

[0013] The memory is used to store computer programs;

[0014] When the processor executes the program stored in the memory, it implements the steps of the above-described urban change detection method based on global guidance-local feature alignment.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0016] 1) This invention guides the alignment and fusion of the first and second features using global guidance information at the current level, and generates guidance information for the next level. This design improves the stability and robustness of feature alignment between the two periods, and through this adaptive feature fusion mechanism, enhances the change detection accuracy and generalization ability in complex urban scenarios.

[0017] 2) The global guidance-local alignment fusion module proposed in this invention is easy to integrate into existing change detection networks and has good engineering application value.

[0018] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating an urban change detection method based on global guidance and local feature alignment provided in an embodiment of the present invention.

[0020] Figure 2 This is a schematic diagram of a network architecture of a non-twin global guidance-local alignment change detection network provided in an embodiment of the present invention;

[0021] Figure 3 This is a schematic diagram of the workflow of the global guidance-local alignment fusion module provided in an embodiment of the present invention. Detailed Implementation

[0022] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0023] Figure 1 This is a flowchart illustrating a city change detection method based on global guidance and local feature alignment provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:

[0024] S101. Obtain the first-period image and the second-period image of the area to be detected.

[0025] For example, the area to be detected could be some areas of the city to be detected. The first period and the second period refer to two time points with a certain time interval, and the present invention does not limit the specific length of the interval.

[0026] S102. Extract multi-scale features from the first-period image and the second-period image respectively to obtain the first feature pyramid and the second feature pyramid. The first feature pyramid contains I different levels of first features, and the second feature pyramid contains I different levels of second features, where I is a positive integer greater than or equal to 1.

[0027] For example, I=5, based on which, the first feature pyramid includes: a first feature at the first level, a first feature at the second level, a first feature at the third level, a first feature at the fourth level, and a first feature at the fifth level; and the second feature pyramid includes: a second feature at the first level, a second feature at the second level, a second feature at the third level, a second feature at the fourth level, and a second feature at the fifth level.

[0028] S103. Project the current level's guiding features, the current level's first feature, and the current level's second feature respectively. Generate global guiding information for the current level based on the projected guiding features. Based on local window search and the current level's global guiding information, perform local feature alignment on the projected first feature and the projected second feature of the current level. Perform adaptive feature fusion based on the local feature alignment result and gating mechanism. Use the resulting fused features of the current level as the guiding features of the next level. The guiding features of the first level are obtained by stitching together the first-period image and the second-period image and then extracting features.

[0029] Specifically, the first-level guiding feature is a feature obtained by stitching together the first-period image and the second-period image, followed by convolution processing. The kernel size of the convolutional layer performing the convolution processing is... The step size and padding are 1.

[0030] Specifically, feature projection operations can be performed on the guiding feature, the first feature, and the second feature of the current layer using three independent convolutional layers. The kernel size of each convolutional layer is [size missing]. With a step size and padding of 1, we obtain the projected guided features, the first projected features, and the second projected features of the current level, so as to unify the number of channels of these three features to the output dimension of the current level.

[0031] S104. Perform bottom-up decoding based on the guiding features of all levels to obtain the final decoded features.

[0032] S105. Generate a change detection result map of the region to be detected based on the final decoding features.

[0033] The final decoded features can be converted into category probabilities. Then, threshold segmentation is performed on the probabilities to obtain the change detection result map of the region to be detected. Since the principle of threshold segmentation is a well-known technique, it will not be described in detail in this invention.

[0034] In some embodiments, the generation of global guidance information for the current level based on the projected guidance features of the current level in S103 is achieved through S10~S12:

[0035] S10. Generate a global guidance temperature map for the current level based on the guided features projected from the current level. The global guidance temperature map is used to control the sharpness of the distribution of the activation function Softmax.

[0036] For example, the global guided temperature map of the current level. The expression is as follows:

[0037] ;

[0038] in, This represents the guided features projected onto the current level. This represents the first double convolution operation, which is implemented by two convolutional layers. The kernel size of the first convolutional layer is [value missing]. The kernel size of the second convolutional layer is , This represents the hyperbolic tangent activation function. This represents the initial temperature heatmap. It should be noted that the current level's global guiding temperature map includes... The temperature thermal values ​​of each pixel in the image are used to generate higher temperature thermal values ​​in areas with clear textures to achieve accurate matching, and lower temperature thermal values ​​in smooth areas to smooth noise. In other words, the temperature map is used to enhance the matching peak in areas with rich textures and smooth the matching distribution in areas with weak textures.

[0039] S11. Generate the global guidance confidence of the current level based on the guided features after projection of the current level. The global guidance confidence is used to evaluate whether the pixel position is suitable for alignment operation.

[0040] For example, the global boot confidence level at the current level The expression is as follows:

[0041] ;

[0042] in, This represents the first double convolution operation. This indicates the global guidance confidence level at the current level. This represents the Sigmoid activation function. The global guided temperature map for the current level contains... The confidence score of each pixel is used to evaluate whether the pixel is suitable for feature alignment.

[0043] S12. Use the global guidance temperature map and the global guidance confidence of the current level as the global guidance information of the current level.

[0044] In some embodiments, the local feature alignment of the first projected feature and the second projected feature of the current level based on local window search and global guidance information of the current level in S103 is implemented through S13~S16:

[0045] S13, The first feature after projection onto the current level The second feature after projection of the current level Normalization is performed separately to obtain the normalized first feature of the current level. and the normalized second feature of the current level .

[0046] For example, the normalized first feature of the current level and the normalized second feature of the current level The expressions are as follows:

[0047] ;

[0048] in, It is a coefficient of size 1e-6, used to prevent division by zero. This indicates the calculation of the L2 norm.

[0049] S14. For the normalized first feature of the current level Each pixel in From the normalized second feature of the current level Determining the pixel A pixel at the same position is obtained as a pixel. .

[0050] S15, Normalized second feature at the current level The middle is determined by pixels Local search window centered .

[0051] Specifically, for the normalized first feature of the current level Each pixel in (denoted as) ), from the normalized second feature of the current level Determining the pixel A pixel at the same position (denoted as ) ), then, in the current level of normalized second feature Define a pixel Centered on, size is The window, thus obtaining the pixels. A corresponding local search window , It is a positive integer greater than 1, and The value of can be set according to actual needs, and this invention does not limit it.

[0052] S16, Pixel-based Local search window Global guidance temperature map at the current level Global guidance confidence at the current level and the second feature projected from the current level. Perform local feature alignment to obtain aligned features. .

[0053] Specifically, calculate each pixel. exist eigenvectors in With local search window Pixels at each location within exist eigenvectors in The dot product between them yields There are 1 similarity score, among which... Represents a local search window The number of pixels within, It is a positive integer, and The value ranges from 1 to Based on pixels Global guided temperature map at the current level Temperature thermodynamic value ,right Each similarity score is scaled, and then the pixel values ​​are calculated using Softmax. With pixels Normalized matching weights between Based on normalized matching weights and pixels eigenvectors Calculate pre-alignment features ;Utilize the global guidance confidence of the current level In pre-aligned features and Adaptive interpolation is performed between them to obtain alignment features. .

[0054] For example, pixels With pixels Similarity between The expression is as follows:

[0055] ;

[0056] in, It is the transpose symbol. express The transpose of .

[0057] For example, The expression is as follows:

[0058] ;

[0059] in, This represents the natural exponential function. It is a positive integer, and The value ranges from 1 to , Represents pixels With local search window Inner Similarity between pixels Indicates As variables, and for All corresponding Sum.

[0060] Here, based on the normalized matching weights and pixels eigenvectors You can get the pixels. pre-aligned feature vectors Thus, the normalized first feature of the current level The pre-aligned feature vectors of all pixels constitute the pre-aligned features. For example, The expression is as follows:

[0061] ;

[0062] in, Indicates As variables, and for All corresponding Sum.

[0063] For example, The expression is as follows:

[0064] .

[0065] The formula utilizes confidence levels. Construct a soft gating mechanism, that is, add gating in areas with high confidence. The weights are adjusted to correct feature misalignment caused by parallax, increasing the weights in regions with low confidence. The weights are determined to preserve the true features of ground changes or to avoid the introduction of noise.

[0066] In some embodiments, the adaptive fusion of features based on local feature alignment results and gating mechanism in S103 above, and the resulting fused features of the current level as guiding features for the next level, are implemented through S17~S20:

[0067] S17, Align Features The first feature after projection of the current level And the guided features after projection of the current level After concatenation along the channel dimension, a second double convolution operation is performed to obtain the alignment features. The first feature after projection of the current level And the guided features after projection of the current level Each has its corresponding routing weight.

[0068] It should be noted that the second double convolution operation is implemented by two convolutional layers, wherein the kernel size of the first convolutional layer is [missing information]. The kernel size of the second convolutional layer is The above S17 allows for dynamic weight allocation based on feature validity using a three-way gating mechanism. For example, aligning features... The first feature after projection of the current level And the guided features after projection of the current level The expressions for the corresponding route weights are as follows:

[0069] ;

[0070] ;

[0071] in, This indicates the second double convolution operation. For activation function, express The vector formed They represent The routing weight, and In reality, they are all spatial weighted graphs.

[0072] S18. Based on routing weight Alignment features The first feature after projection of the current level And the guided features after projection of the current level Perform a weighted summation to obtain the mixed features. .

[0073] For example, The expression is as follows:

[0074] .

[0075] S19. Based on alignment features The first feature after projection of the current level Determine the characteristics of the difference .

[0076] For example, The expression is as follows:

[0077] ;

[0078] in, This indicates that the absolute value is being calculated.

[0079] S20, Mixing features Sum and difference characteristics After concatenation, residual block processing is performed to obtain the fusion feature of the current level, and the fusion feature of the current level is used as the guiding feature of the next level. .

[0080] For example, the next level of guiding features The expression is as follows:

[0081] ;

[0082] in, This indicates a splicing operation. This represents the residual connection operation. Residual connections are an important computational technique in deep learning, primarily used to address the vanishing gradient problem during the training of deep neural networks, thereby improving the training efficiency and performance of the model. Residual connections skip certain computational layers of the input signal and add it to the calculated output to form the final output.

[0083] In some embodiments, the above-mentioned S104 is implemented through steps S1041 to S1043:

[0084] S1041, Fusion features of Level I After upsampling, decoding is performed to obtain the layer I decoding features. .

[0085] S1042, Decoding features of layer I Upsampling is then performed before fusing the features with the (I-1)th level. The concatenation is performed along the channel dimension, and the concatenated features are then decoded to obtain the (I-1)th layer decoded features. .

[0086] S1043, Decoding features of the (I-1)th layer Upsampling is then performed before fusing the features with the (I-2)th level. The features are concatenated along the channel dimension, and then decoded to obtain the decoded features of layer I-2. This process is repeated iteratively until the first layer of decoded features is obtained. .

[0087] For example, when I=5, then first... After upsampling, decoding is performed to obtain... After that, After upsampling, then with Concatenation is performed along the channel dimension, and the concatenated features are then decoded to obtain... Next, regarding After upsampling, then with Concatenation is performed along the channel dimension, and the concatenated features are then decoded to obtain... ;right After upsampling, then with Concatenation is performed along the channel dimension, and the concatenated features are then decoded to obtain... Next, regarding After upsampling, then with Concatenation is performed along the channel dimension, and the concatenated features are then decoded to obtain... .

[0088] In this invention, the final decoding feature is the decoding feature output by the trained non-twin global guidance-local alignment change detection network after inputting the first-period image map, the second-period image map, and the first-level guidance feature into the trained non-twin global guidance-local alignment change detection network. That is, the processing procedures corresponding to S102 to S104 are implemented by the trained non-twin global guidance-local alignment change detection network.

[0089] Specifically, the non-twin global-guided-local-alignment change detection network includes: a non-twin feature extraction network, a change encoder, and a change decoder; the non-twin feature extraction network includes two feature encoders with non-shared weights (referred to as two non-shared feature encoders or two non-twin feature encoders), each containing I convolutional blocks; the change encoder includes I global-guided-local-alignment fusion modules; the change decoder includes I deconvolutional decoding blocks; the first layer of each non-shared feature encoder... The output of the convolutional block at the first level and the output of the second level The input connections of the convolutional blocks at each level are such that the input features of the first level convolutional block of a non-shared feature encoder are the first-time image map, and the input features of the first level convolutional block of another non-shared feature encoder are the second-time image map. It is a positive integer and The value of is from 1 to I; the first The output of the first global guidance-local alignment fusion module and the first The first level of guidance features is also used as an input feature of the first global guidance-local alignment fusion module; The output of the deconvolution decode block and the first The input connection of the i-th deconvolutional decoding block, and the output of the i-th global-guided-local-aligned fusion module is connected to the input of the i-th deconvolutional decoding block, the i-th The global guidance-local alignment blending module also works with the first Skip connections between deconvolutional decoding blocks, It is a positive integer and The value of is from 1 to I-1.

[0090] Specifically, in the I-level convolutional blocks, each level performs a standard double convolution operation. Each level of convolutional block consists of two sets of convolutional layers, a normalization function, and an activation function. The kernel size of each set of convolutional layers is [missing information]. The stride and padding are both 1. Each of the I deconvolutional decoding blocks consists of an upsampling layer, a feature concatenation module, and a deconvolutional block. This deconvolutional block comprises a deconvolutional layer, a normalization function, an activation function, and a Dropout layer. The deconvolutional kernel size of the deconvolutional layer is... The step size and padding are 1.

[0091] The aforementioned non-twin global-guided-local-aligned change detection network may also include a linear mapping and normalization module, the input of which is connected to the output of the first deconvolutional decoding block, and the linear mapping and normalization module is used to convert the decoded features output by the first deconvolutional decoding block into a change detection result map.

[0092] For example, Figure 2 This is a schematic diagram of a non-twin global guidance-local alignment change detection network architecture. Figure 2 In this context, two non-shared feature encoders are referred to as two non-twin feature encoders. Each non-twin feature encoder contains five levels of convolutional blocks. For example, the five yellow cuboids in a non-twin feature encoder, represented by the yellow area, represent the five levels of convolutional blocks in that non-twin feature encoder, and the five green cuboids in a non-twin feature encoder, represented by the green area, represent the five levels of convolutional blocks in that non-twin feature encoder. It should be noted that... Figure 2 The document does not show the connections between convolutional blocks in adjacent layers across the five convolutional layers. Please refer to [link / reference needed]. Figure 2The change encoder includes five global guidance-local alignment fusion modules: Global Guidance-Local Alignment Fusion Module 1, Global Guidance-Local Alignment Fusion Module 2, Global Guidance-Local Alignment Fusion Module 3, Global Guidance-Local Alignment Fusion Module 4, and Global Guidance-Local Alignment Fusion Module 5. Two black arrows pointing to each global guidance-local alignment fusion module represent its two inputs, and the black arrow pointed to by that module represents its one output. (Continue to refer to...) Figure 2 The transformation decoder includes five deconvolution decoder blocks: deconvolution decoder block 1, deconvolution decoder block 2, deconvolution decoder block 3, deconvolution decoder block 4, and deconvolution decoder block 5. A black arrow pointing to each deconvolution decoder block indicates its input, while the black arrow pointed to by that block indicates its output. (Continue to refer to...) Figure 2 Each yellow dashed arrow represents a skip connection. For example, the yellow arrow pointing from the global guidance-local alignment fusion module 1 to the deconvolution decoding block 1 represents a skip connection between the global guidance-local alignment fusion module 1 and the deconvolution decoding block 1. The rest are similar and will not be elaborated here.

[0093] If we take the first period image map Second period image and the first-level guiding features Enter the above Figure 2 Following the non-twin global-guided-local alignment change detection network shown, the network's processing flow is as follows:

[0094] The first-period image was extracted using these two non-twin feature encoders. Second period image The multi-scale features yield five levels of first features, represented as follows: And the second feature of the five levels, uniformly represented as , and This represents the two non-twin feature encoders;

[0095] Global guidance - local alignment and fusion module For the first The first feature of the hierarchy and the Second feature of the hierarchy and the Hierarchical guidance features Perform feature projections separately, based on the first... The first feature after projection of the hierarchy and the Second feature after projection of hierarchy and the Hierarchical projection guidance features Generate the first Hierarchical fusion features and will As the first The hierarchical guiding features are used until the fusion features of the 5th level are obtained. Then input deconvolution decoding block 5;

[0096] 5 pairs of deconvolution decoding blocks Upsampling is performed before decoding to obtain the fifth layer of decoded features. The input is then fed into deconvolution decoding block 4; deconvolution decoding block 4... Upsampling is then performed before fusing with the features from the 4th level. The features are concatenated along the channel dimension, and then decoded to obtain the fourth layer of decoded features. Then, deconvolution decoding block 3 is input. Deconvolution decoding blocks 3 through 1 are executed sequentially according to this principle until the first layer of decoded features is output by deconvolution decoding block 1. ;

[0097] Finally, the linear mapping and normalization module according to Generate a change detection result image.

[0098] For example, Figure 3 This is a flowchart illustrating the workflow of the global guidance-local alignment and fusion module. For example... Figure 3 As shown, the guiding features of the current level will be... The first feature of the current level and the second feature of the current level After inputting a corresponding global guidance-local alignment fusion module, the module first projects the three features respectively to obtain the projected guidance features of the current level. The first feature after projection of the current level The second feature after projection of the current level ; after that, according to Generate a global guided temperature map for the current level. ,like Figure 3 As shown, this can be specifically achieved through... To generate [the product], convolution, normalization, ReLU activation, and convolution operations are performed sequentially. At the same time, according to Generate global bootstrapping confidence for the current level. ,like Figure 3 As shown, this can be specifically achieved through... The global bootstrap confidence of the current layer is generated by sequentially performing convolution, normalization, ReLU activation, convolution operation, and sigmoid activation operation. After obtaining and Then, by sequentially... , Normalization and expansion are performed, and correlation calculation (i.e., similarity calculation and utilization) are also performed. Similarity scaling), Softmax calculation, and aggregation calculation (i.e., calculating pre-aligned features). and utilize exist and Adaptive interpolation is performed between them to generate alignment features. Then, using a three-way gating mechanism (in... Figure 3 (represented as Router) calculates separately , and weight ,use right , and Perform weighted fusion, based on the features of the weighted fusion, and Generate fusion features for the current level. .

[0099] To accelerate network convergence and improve the discriminative power of intermediate layer features, the aforementioned non-twin global-guided-local alignment change detection network can be trained using a multi-scale deep supervision strategy. Specifically, during each training iteration, as described above... Figure 2 As shown, the decoded features output by the first deconvolutional decoding block Convert to change detection result image Furthermore, an auxiliary branch is introduced at the intermediate level of decoding, that is, the decoded features output by the second and third deconvolutional decoding blocks are... and Convert them into change detection result images respectively and This yields prediction results at different scales. Then, bilinear interpolation is used to restore the low-resolution prediction map to the original image size, facilitating loss calculation. Therefore, during network training, the total loss function... The expression is as follows:

[0100] ;

[0101] in, This represents the total loss (also known as the total loss function). Represents the cross-entropy loss function. Indicates the truth value label of the change. , and This represents the loss weight parameter.

[0102] It should be noted that during each training iteration, after calculating the total loss using the above formula, the network parameters of the convolutional layers in the two non-shared feature encoders, the change encoder, and the change decoder are adjusted simultaneously during backpropagation. Training can be stopped after a preset number of iterations, resulting in the trained two non-shared feature encoders, the trained change encoder, and the trained change decoder; thus, a trained non-Twin global guided-local aligned change detection network is obtained.

[0103] The present invention also provides an urban change detection device based on global guidance-local feature alignment, including a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory communicate with each other through the communication bus. The memory is used to store computer programs. When the processor executes the program stored in the memory, it implements the steps of the above-mentioned urban change detection method based on global guidance-local feature alignment.

[0104] To further verify the technical effect of the present invention, building datasets from two different cities were used to conduct simulation experiments on the present invention and some existing methods, respectively, and specific quantitative evaluation results were obtained. The specific quantitative evaluation results are shown in Table 1 below:

[0105] Table 1. Quantitative evaluation results of the present invention and some existing methods

[0106]

[0107] As shown in Table 1, the present invention can effectively suppress spurious changes caused by misalignment and significantly improve detection accuracy. Experimental results on two city change detection datasets verify the superior performance of the method of the present invention in change detection.

[0108] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0109] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0110] In this specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. While different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce a good effect.

[0111] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A method for detecting urban changes based on global guidance and local feature alignment, characterized in that, include: Obtain the first-period and second-period image images of the area to be detected; Multi-scale features are extracted from the first period image and the second period image respectively to obtain a first feature pyramid and a second feature pyramid. The first feature pyramid contains I different levels of first features, and the second feature pyramid contains I different levels of second features, where I is a positive integer greater than or equal to 3. Feature projection is performed on the current level's guiding features, the current level's first feature, and the current level's second feature. Based on the projected guiding features of the current level, global guiding information for the current level is generated. Based on local window search and the current level's global guiding information, local feature alignment is performed on the projected first feature and the projected second feature of the current level. Based on the local feature alignment result and gating mechanism, adaptive feature fusion is performed, and the resulting fused features of the current level are used as the guiding features of the next level. The guiding features of the first level are obtained by stitching together the first period image and the second period image and then extracting features. Based on the guiding features of all levels, a bottom-up decoding process is performed to obtain the final decoded features; Based on the final decoding features, a change detection result map of the region to be detected is generated; The step of generating global guidance information for the current level based on the projected guidance features of the current level includes: A global guidance temperature map for the current level is generated based on the projected guidance features of the current level, wherein the global guidance temperature map is used to control the distribution sharpness of the activation function Softmax; The global guidance confidence of the current level is generated based on the projected guidance features of the current level, wherein the global guidance confidence is used to evaluate whether the pixel position is suitable for feature alignment operation; The global guidance temperature map and the global guidance confidence level of the current level are used as the global guidance information of the current level. The step of performing local feature alignment on the projected first feature and the projected second feature of the current level based on local window search and the global guidance information of the current level includes: The first feature after projection of the current level and the second feature after projection of the current level are normalized respectively to obtain the normalized first feature and the normalized second feature of the current level. For each pixel in the normalized first feature of the current level Determine the pixel from the normalized second feature of the current level. A pixel at the same position is obtained as a pixel. ; The pixel is determined in the normalized second feature of the current level. Local search window centered ; Based on the pixel points The local search window The global guided temperature map of the current level, the global guided confidence of the current level, and the projected second feature of the current level are aligned locally to obtain aligned features.

2. The urban change detection method based on global guidance-local feature alignment according to claim 1, characterized in that, The expressions for the global guidance temperature map and the global guidance confidence level of the current level are as follows: ; ; in, This represents the guided features projected onto the current level. This represents the first double convolution operation. This represents the hyperbolic tangent activation function. This represents the initial temperature thermogram. This represents the global guidance temperature map for the current level. This represents the global guidance confidence level at the current level. This represents the Sigmoid activation function.

3. The urban change detection method based on global guidance-local feature alignment according to claim 1, characterized in that, Based on the pixel points The local search window The aligned features are obtained by aligning the global guided temperature map of the current level, the global guided confidence of the current level, and the projected second feature of the current level with local features, including: Calculate the pixel points Feature vectors and the local search window The dot product of the feature vectors of pixels at each location within the matrix yields multiple similarity scores. According to the pixel points Temperature thermodynamic values ​​in the global guided temperature map at the current level The similarity is scaled for each value, and then the pixel is calculated. With the local search window Normalized matching weights between pixels at each location within the array; Based on the normalized matching weights and the local search window Calculate the pre-aligned features from the feature vectors of pixels at each position within the matrix. Using the global guiding confidence of the current level, adaptive interpolation is performed between the pre-aligned feature and the second feature projected from the current level to obtain the aligned feature.

4. The urban change detection method based on global guidance-local feature alignment according to claim 1, characterized in that, The adaptive fusion of features based on local feature alignment results and gating mechanisms, and the use of the fused features of the current level as guiding features for the next level, includes: After concatenating the alignment feature, the first feature projected from the current level, and the guiding feature projected from the current level along the channel dimension, and then performing a second double convolution operation, the routing weights corresponding to the alignment feature, the first feature projected from the current level, and the guiding feature projected from the current level are obtained respectively. Based on the routing weight, the alignment feature, the first feature projected from the current level, and the guiding feature projected from the current level are weighted and summed to obtain the hybrid feature; The difference feature is determined based on the alignment feature and the first feature after projection of the current level; The fusion feature and the difference feature are concatenated and then processed by residual block to obtain the fusion feature of the current level, and the fusion feature of the current level is used as the guiding feature of the next level.

5. The urban change detection method based on global guidance-local feature alignment according to claim 1, characterized in that, The bottom-up decoding process based on the guiding features of all levels yields the final decoded features, including: The fused features of the I-th layer are upsampled and then decoded to obtain the decoded features of the I-th layer. The layer I decoding features are upsampled and then concatenated with the fusion features of layer I-1 in the channel dimension. The concatenated features are then decoded to obtain the layer I-1 decoding features. The decoding features of the (I-1)th layer are upsampled and then concatenated with the fusion features of the (I-2)th layer in the channel dimension. The concatenated features are then decoded to obtain the decoding features of the (I-2)th layer. This process is repeated until the decoding features of the first layer are obtained. The decoding features of the first layer are the final decoding features.

6. The urban change detection method based on global guidance-local feature alignment according to claim 1, characterized in that, The final decoding feature is the decoding feature generated by the trained non-twin global guidance-local alignment change detection network after inputting the first period image map, the second period image map, and the first level guidance feature into the trained non-twin global guidance-local alignment change detection network. The non-twin global guided-local alignment change detection network includes: a non-twin feature extraction network, a change encoder, and a change decoder; the non-twin feature extraction network includes two non-shared feature encoders, each containing I layers of convolutional blocks; the change encoder includes I global guided-local alignment fusion modules; and the change decoder includes I deconvolutional decoding blocks. The first non-shared feature encoder The output of the convolutional block at the first level and the output of the second level The input connections of the convolutional blocks at each level are as follows: the input features of the first level convolutional block of a non-shared feature encoder are the first period image map, and the input features of the first level convolutional block of another non-shared feature encoder are the second period image map. It is a positive integer and The value of is from 1 to 1; No. The output of the first global guidance-local alignment fusion module and the first The first level of guidance features is also used as an input feature of the first global guidance-local alignment fusion module. No. The output of the deconvolution decode block and the first The input connection of the i-th deconvolutional decoding block, and the output of the i-th global-guided-local-aligned fusion module is connected to the input of the i-th deconvolutional decoding block, the i-th The global guidance-local alignment blending module also works with the first Skip connections between deconvolutional decoding blocks, It is a positive integer and The value of is from 1 to I-1.

7. The urban change detection method based on global guidance-local feature alignment according to claim 6, characterized in that, The loss function of the non-twin global guidance-local alignment change detection network is expressed as follows: ; in, Indicates the total loss. Represents the cross-entropy loss function. This represents the change detection result map generated based on the features output by the first deconvolutional decoding block. This represents the change detection result map generated based on the features output by the second deconvolution decoding block. This represents the change detection result map generated based on the features output by the third deconvolution decoding block. Indicates the truth value label of the change. , and This represents the loss weight parameter.

8. A city change detection device based on global guidance-local feature alignment, comprising a processor, a communication interface, a memory, and a communication bus, characterized in that, The processor, the communication interface, and the memory communicate with each other via the communication bus; The memory is used to store computer programs; When the processor executes a program stored in the memory, it implements the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Lightweight sensing system for complex environment change and detection capability enhancing method

    CN121685991A