A remote sensing detection method for changes in urban building types

Through the deep learning method based on Swim Transformer, multi-scale bi-time phase image features are extracted and comparative loss calculations are performed, and the problem of difficult to identify building type changes and pseudo-change interference in the prior art is solved, and accurate detection of building type changes and pseudo-change mitigation is achieved.

CN117173573BActive Publication Date: 2025-05-16FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311195257.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-16
Publication Date
2025-05-16
Estimated Expiration
2043-09-16

AI Technical Summary

Technical Problem

The prior art is difficult to identify changes in the area and type of building at the same time, and there is a problem of pseudo-change interference.

Method used

Using the deep learning method based on Swim Transformer, a semantic segmentation and change detection task module is constructed through multi-scale bi-time phase image feature extraction and comparison loss calculation to identify building type changes and mitigate the impact of pseudo-change.

Benefits of technology

It realizes automatic and accurate detection of changes in building types, reduces pseudo-change interference, and improves the efficiency and accuracy of change detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117173573B_ABST
    Figure CN117173573B_ABST
Patent Text Reader

Abstract

The present invention proposes a remote sensing detection method for changes in urban building types. A model for detecting changes in urban building types based on high-resolution remote sensing images is constructed, which is specifically composed of two modules: semantic segmentation and change detection. The semantic segmentation module uses Swim Transformer as an encoder to improve the global perception ability in the feature extraction process. The building type information and boundary information are learned simultaneously through the parameter sharing mechanism of different tasks to solve the problem of detecting changes in building types, and multi-scale feature comparison learning is integrated into the feature extraction stage to improve the network's perception of changing objects. In conjunction with the change detection module, this method can simultaneously identify the boundaries and type change information of buildings. The present invention integrates an end-to-end deep neural network to realize the automatic detection of changes in urban building types using high-resolution remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of remote sensing detection methods for building type changes, in particular to a remote sensing detection method for urban building type changes. Background Art

[0002] In urban areas, building changes are an important manifestation of the urbanization process. While changing the surface cover, they directly affect the way people use the land. Obtaining accurate information on urban building changes, including regional changes and type changes, is basic information for understanding the current status of urban land use and is also the basic data support for urban sustainable development planning. In addition, automatically identifying building change information at the individual building scale will help carry out building change statistics in the geographic national conditions monitoring project. Traditional building type change information can only be obtained through manual surveying or manual visual interpretation, which is time-consuming and labor-intensive.

[0003] With the continuous development of remote sensing technology, the spatial resolution of remote sensing images has been continuously improved. Remote sensing images with high spatial resolution have brought opportunities for refined building change recognition. The current building change detection method mainly extracts image features directly, and then uses complex feature learning methods to identify change information. It is inevitably affected by pseudo-change interference caused by shooting perspective, image offset, etc., and pseudo-change information is mistakenly identified as real building changes. In addition, the existing automatic detection method of building changes in remote sensing images only focuses on the change information of building areas, and it is difficult to effectively identify the type change information of buildings. Therefore, how to simultaneously identify building area changes and type changes, and effectively reduce the impact of pseudo-changes and other noise, is a problem that needs to be solved urgently. This method uses deep learning technology to achieve automatic and accurate detection of building type changes. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a remote sensing detection method for urban building type changes, which overcomes the problem that traditional methods cannot identify building category change information and the identification results have obvious false changes.

[0005] To achieve the above object, the present invention adopts the following technical solution: a remote sensing detection method for urban building type changes, comprising the following steps:

[0006] Step S1: Obtain dual-phase high-resolution remote sensing images of the study area and perform preprocessing operations on the images, including radiation correction, orthorectification, image fusion, image resampling, etc.;

[0007] Step S2: Construct a semantic segmentation task module based on Swim Transformer, extract multi-scale dual-temporal image features, and learn building feature information; specifically, use multi-scale learning to alleviate the problem of weak representation capabilities of small-scale geometric features and large-scale semantic features in single-scale learning, and improve the accuracy of semantic segmentation tasks; and perform contrast loss calculation on the same-scale features of dual-temporal images, and through contrast loss optimization, reduce the similarity between real change features and expand the similarity between pseudo-change features; combine multi-scale to alleviate the problem of not being able to take into account the representation capabilities of semantic features and geometric features at the same time in a single scale, and highlight the real change information;

[0008] Step S3: Construct a change detection task module based on Swim Transformer to extract image change features and learn the change information between images; based on the multi-task concept, construct a dual-temporal multi-category building type change map and a binary change change map decoder respectively;

[0009] Step S4: Design semantic boundary loss and binary change loss for building changes in dual-temporal images; specifically, weighted loss is used in multi-classification semantic segmentation tasks to alleviate the problem of imbalanced number of building types during training;

[0010] Step S5: Create a multi-category building change sample set, divide the buildings into unchanged buildings, discrete buildings, simple single-purpose buildings, complex multi-purpose buildings and residential buildings; perform sliding cropping and sample enhancement operations on the dual-phase images to construct training and verification sample sets;

[0011] Step S6: Model training and prediction, finally obtaining a building type change map and a binary map of the change area of ​​the dual-temporal image.

[0012] In a preferred embodiment, step S1 specifically includes the following steps:

[0013] Step S11: Select two high-resolution remote sensing images of different time phases of the area to be studied, and preprocess the dual-time phase images, including atmospheric and radiation correction, panchromatic-multispectral fusion and other operations;

[0014] Step S12: Perform operations such as geo-registration and style transfer based on the dual-temporal images obtained in S11, and resample images of different resolutions to the same resolution to improve image processing efficiency.

[0015] In a preferred embodiment, step S2 specifically includes the following steps:

[0016] Step S21: Select Swim Transformer as the basic network, which mainly includes image fusion, downsampling module, mask module and multi-head moving window attention calculation module;

[0017] Step S22: Based on the twin multi-task network learning framework, two modules of semantic segmentation and change detection are constructed, and two weight-sharing multi-layer attention networks are used to encode the features of the dual-phase images, while learning the building boundaries and their type information in the images;

[0018] Step S23: Based on the multi-scale features extracted in step S22, a multi-scale contrast loss is constructed; specifically, compared with a single scale, a multi-scale can simultaneously take into account the representation capabilities of image building geometry features and texture features; the similarity between real change features is reduced and the similarity between pseudo change features is expanded through contrast loss calculation to alleviate the influence of pseudo changes in the image. The main methods are as follows:

[0019] Cosine similarity is used to calculate the multi-scale similarity based on the channel direction. The calculation formula is as follows:

[0020]

[0021] In the formula, A and B are the feature information sets of the same scale at T1 and T2 respectively. i , B i Respectively represent the feature vectors of the same channel at T1 and T2; S(A, B) is the vector similarity, and the value is between [-1, +1]. The similarity S(A, B) is converted into the distance D, and the calculation formula is as follows:

[0022]

[0023] Construct a feature distance D difference threshold to generate pseudo labels. After experiments, the feature D difference is between 0.1 and 0.4 when the same network is used to extract the same type of buildings, and the feature D difference is between 0.6 and 0.9 when different types of buildings are extracted. Finally, the D difference threshold is composed of a basic threshold and a learnable threshold. The basic threshold is 0.5, and the learnable threshold is learned from the label of the binary change map. The pseudo label reflects the feature change in the channel direction.

[0024] Based on the above steps, we can get the pseudo-label of whether the building information corresponding to the feature has changed; build a multi-scale feature loss calculation function based on the pseudo-label; use the multi-scale feature difference to expand the extracted change features and shrink the pseudo-change or unchanged features, thereby highlighting the change information and unchanged information; the loss calculation tool is as follows:

[0025]

[0026] l dis =l1+l2+…+lm

[0027] Where D is the calculation result in S32, l i Represents the feature difference loss of the i-th scale, label = 0 indicates that the feature corresponding to this channel has not changed, l = 1 indicates that the feature corresponding to this channel has changed; max is the maximum distance difference, when max is equal to 2, it means that the features are completely different; optimize the loss function, reduce the feature similarity of the unchanged targets in the two time phases, increase the feature similarity of the changed targets, and form the total contrast loss by adding and calculating the integrated multi-scale contrast loss.

[0028] In a preferred embodiment, step S3 specifically includes the following steps:

[0029] Step S31: Based on the multi-scale features of the dual-phase images extracted in step S2, a change feature encoder based on SwimTransformer is constructed to obtain the change features of the phase images;

[0030] Step S32: Based on the idea of ​​twin multi-task learning, a dual-phase building multi-classification decoder based on multi-scale building features is constructed, and a binary mask solution change map decoder based on image change features is constructed; it mainly includes a multi-scale feature fusion module, which fuses multi-scale information during the decoding process to avoid information loss.

[0031] In a preferred embodiment, step S4 specifically includes the following steps:

[0032] Step S41: Based on the twin multi-task network model constructed in steps S2 and S3, a multi-task loss function is designed to optimize the network weights to reduce the prediction error; specifically, the multi-classification cross entropy loss is used to calculate the loss of multi-classification semantic segmentation of dual-phase buildings, and the calculation formula is as follows:

[0033]

[0034] Where l m represents the bi-temporal multi-category semantic segmentation loss, represents the predicted probability of the jth category, cls represents the true category, and N represents the total number of categories; through the back propagation of the loss, the predicted category of the pixel position is made closer to the true category;

[0035] Step S42: using weighted contrast loss to calculate binary change mask loss; the calculation formula is as follows:

[0036] l c =-y c ·log(p c )-(1-y c )·log(1-p c )

[0037]

[0038]

[0039] l ct =W p ·l cp +W n ·l cn

[0040] Where: l c represents the change binary mask loss, y c Indicates the actual label of the pixel position c, whether it changes or not, p c Represents the prediction result of the c pixel position; l cp represents the loss of the actual change in the position of pixel c, l cn represents the loss of pixel position c which is actually unchanged; y cp The mask representing the actual change in the position of pixel c, y cn represents the real unchanged mask of pixel position c; n p The total number of pixels representing the change in the true label, n n The total number of pixels representing the change in the true label; W p is the calculation weight of the change loss, W n is the calculation weight of the unchanged loss, and the total loss of the changed mask is obtained by weighted addition. ct ;

[0041] Step S43: Use weighted cosine similarity loss to further highlight the change information; the calculation formula is as follows:

[0042]

[0043] Where: l cos represents the cosine similarity loss, x i Represents the classification probability vector arranged on a certain pixel position channel of the dual-phase image, cos(x1, x2) represents the cosine similarity of the probability vector, and max represents the boundary value; y represents the real label of whether the corresponding position has changed. When y=-1, it means that the x position has changed. At this time, the calculated cosine similarity should be small in theory. The cosine similarity of the changed position is optimized to 0 by optimizing the loss; y=1 means that the x position has not changed. At this time, the cosine similarity should be infinitely close to 1 in theory. The loss can be optimized to make the loss of the unchanged position tend to 1; the label information y is obtained by converting the real label;

[0044] Step S44: Based on the multiple loss functions constructed in steps S41, S42, and S43, the total loss of the network is obtained by weighted addition:

[0045] l t =λ1l dis +λ2l m +λ3l ct +λ4l cos

[0046] Where: l t represents the total loss of the network, λ i It represents the weights corresponding to different losses, and the entropy method is used to determine the weight ratios of different losses.

[0047] In a preferred embodiment, step S5 specifically includes the following steps:

[0048] Step S51: Based on the image data preprocessed in step S1, a commonly used labeling method is selected, such as the built-in surface vector construction method of ArcGIS software, and the semantic boundaries of the buildings with dual-temporal changes are manually visually interpreted, and field information is added to a single building, and different types of buildings are represented by different pixel values, including unchanged buildings, discrete buildings, simple single-purpose buildings, complex multi-purpose buildings, and residential buildings;

[0049] Step S52: Based on step S51, the vector label data is exported as single-channel raster data with projection and affine information, so as to vectorize the prediction result;

[0050] Step S53: Based on the sample data set constructed in S51 and S52, the data is cropped into a size of S×S pixels by using a sliding window, and the cropped images with a ratio of 20% to 80% of the changed buildings are selected; the training set, the validation set, and the test set are divided according to a certain ratio; in order to enhance the generalization ability of the model, the training sample data set is expanded by synchronously performing horizontal and vertical flipping and affine transformation on the dual-phase images;

[0051] Step S54: using the SGD optimizer and setting the initial learning rate lr as the optimizer for the segmentation of the changing building and the prediction of the binary changing mask;

[0052] Step S55: Use the modules in the attention model pre-trained on the large dataset ImageNet to initialize the network model weights to speed up the model convergence, improve the model training efficiency, and obtain a better model weight file.

[0053] In a preferred embodiment, step S6 specifically includes the following steps:

[0054] Step S61: Based on the sample set obtained in step S5, set reasonable hyperparameters such as number of iterations, learning rate, optimizer parameters, etc. to train the model, and store the weight file with the best effect in the validation set;

[0055] Step S62: Based on the weight file obtained in step S61, the images in the test set are predicted, and the prediction results of the dual-phase buildings are cropped using the binary mask image; and the color index is used to convert it into an RGB image, and the affine information and projection information are added to obtain the final prediction results.

[0056] Compared with the prior art, the present invention has the following beneficial effects: it can not only detect changes in building areas, but also detect changes in building types. The Swim Transformer based on the twin neural network is constructed to learn image building features and change features. Compared with the convolutional neural network, the global perception ability in the feature extraction process is improved; the decoding method using multi-scale feature fusion can effectively alleviate the feature loss in the transmission process of the multi-level network; in addition, in order to make the real change feature information in the dual-phase image extracted in the twin network have a large spacing, and the unchanged or pseudo-change feature information has a small spacing, multi-scale contrast loss is used for optimization, so that the network can better distinguish between real change information and pseudo-change or unchanged information; in addition, adding weighted contrast loss further optimizes the network's extraction of change information, and using a joint loss function to comprehensively optimize the network. Finally, an end-to-end network is formed to improve the efficiency of change detection, and provide technical support for remote sensing change detection of urban buildings in the process of rapid urban development. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A schematic diagram of a method flow chart of a preferred embodiment of the present invention;

[0058] Figure 2 Reference diagram (I) of the study area change of the preferred embodiment of the present invention;

[0059] Figure 3 Reference Figure (II) is a diagram showing changes in the study area of ​​a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0060] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0061] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0062] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.

[0063] A remote sensing detection method for urban building type changes based on twin multi-task network and multi-scale feature contrast learning, reference Figure 1-3 , including the following steps:

[0064] Step S1: Obtain dual-temporal high-resolution remote sensing images of the study area and perform preprocessing operations on the images, including radiation correction, orthorectification, image fusion, image resampling, etc.;

[0065] Step S2: Construct a semantic segmentation task module based on Swim Transformer, extract multi-scale dual-temporal image features, and learn building feature information. Specifically, use multi-scale learning to alleviate the problem of weak representation of small-scale geometric features and large-scale semantic features in single-scale learning, and improve the accuracy of semantic segmentation tasks. And calculate the contrast loss of the same-scale features of dual-temporal images. Through contrast loss optimization, reduce the similarity between real change features, expand the similarity between pseudo-change features, and combine multi-scale to alleviate the problem of not being able to take into account the representation capabilities of semantic features and geometric features at the same time in a single scale. Highlight real change information.

[0066] Step S3: Construct a change detection task module based on Swim Transformer. Extract image change features and learn the change information between images. Based on the multi-task idea, construct a dual-temporal multi-category building type change map and a binary change change map decoder respectively.

[0067] Step S4: Design semantic boundary loss and binary change loss for building changes in dual-temporal images. Specifically, weighted loss is used in multi-classification semantic segmentation tasks to alleviate the problem of imbalanced number of building types during training.

[0068] Step S5: Create a multi-category building change sample set, and divide the buildings into unchanged buildings, discrete buildings, simple single-purpose buildings, complex multi-purpose buildings, and residential buildings. Perform sliding cropping and sample enhancement operations on the dual-phase images to construct training and validation sample sets;

[0069] Step S6: Model training and prediction, finally obtaining a building type change map and a binary map of the change area of ​​the dual-temporal image.

[0070] In this embodiment, step S1 specifically includes the following steps:

[0071] Step S11: Select two high-resolution remote sensing images of different time phases of the area to be studied, and preprocess the dual-time phase images, including atmospheric and radiation correction, panchromatic-multispectral fusion and other operations;

[0072] Step S12: Perform operations such as geo-registration and style transfer based on the dual-temporal images obtained in S11, and resample images of different resolutions to the same resolution to improve image processing efficiency.

[0073] In this embodiment, step S2 specifically further includes the following steps:

[0074] Step S21: Select Swim Transformer as the basic network, which mainly includes image fusion, downsampling module, mask module and multi-head moving window attention calculation module.

[0075] Step S22: Based on the twin multi-task network learning framework, two modules of semantic segmentation and change detection are constructed. Two multi-layer attention networks with shared weights are used to encode the dual-phase image features, and the building boundaries and their type information in the image are learned at the same time.

[0076] Step S23: Based on the multi-scale features extracted in step S22, a multi-scale contrast loss is constructed. Specifically, compared with a single scale, multi-scale can take into account the representation capabilities of both the geometric features and the texture features of the image building. The similarity between the real change features is reduced by contrast loss calculation, and the similarity between the pseudo-change features is expanded to alleviate the impact of pseudo-changes in the image. The main methods are as follows:

[0077] Cosine similarity is used to calculate the multi-scale similarity based on the channel direction. The calculation formula is as follows:

[0078]

[0079] In the formula, A and B are the feature information sets of the same scale at T1 and T2 respectively. i , B i They represent the feature vectors of the same channel at T1 and T2 respectively. S(A, B) is the vector similarity, and its value is between [-1, +1]. The similarity S(A, B) is converted into the distance D, and the calculation formula is as follows:

[0080]

[0081] Construct a feature distance D difference threshold to generate pseudo labels. After testing, the feature D difference is between 0.1 and 0.4 when the same network is used to extract the same type of buildings, and the feature D difference is between 0.6 and 0.9 when different types of buildings are extracted. Finally, the D difference threshold is composed of a basic threshold and a learnable threshold. The basic threshold is 0.5, and the learnable threshold is learned from the label of the binary change map. The pseudo label reflects the feature change in the channel direction.

[0082] Based on the above steps, we can get the pseudo-label of whether the building information corresponding to the feature has changed. Based on the pseudo-label, we build a multi-scale feature loss calculation function. We use the multi-scale feature difference to expand the extracted change features and shrink the pseudo-change or unchanged features, thereby highlighting the change information and unchanged information. The loss calculation tool is as follows:

[0083]

[0084] l dis =l1+l2+…+l m

[0085] Where D is the calculation result in S32, l i Indicates the feature difference loss of the i-th scale, label = 0 means that the feature corresponding to this channel has not changed, l = 1 means that the feature corresponding to this channel has changed. max is the maximum distance difference, when max is equal to 2, it means that the features are completely different. Optimize the loss function, reduce the feature similarity of the unchanged targets in the two time phases, increase the feature similarity of the changed targets, and form the total contrast loss by adding and calculating the integrated multi-scale contrast loss.

[0086] In this embodiment, step S3 specifically further includes the following steps:

[0087] Step S31: Based on the multi-scale features of the dual-phase images extracted in step S2, a change feature encoder based on SwimTransformer is constructed to obtain the change features of the phase images.

[0088] Step S32: Based on the twin multi-task learning idea, a dual-phase building multi-classification decoder based on multi-scale building features is constructed, and a binary mask solution change map decoder based on image change features is constructed. It mainly includes a multi-scale feature fusion module, which fuses multi-scale information during the decoding process to avoid information loss.

[0089] In this embodiment, step S4 specifically includes the following steps:

[0090] Step S41: Based on the twin multi-task network model constructed in steps S2 and S3, a multi-task loss function is designed to optimize the network weights to reduce the prediction error. Specifically, the multi-classification cross entropy loss is used to calculate the loss of multi-classification semantic segmentation of dual-temporal buildings. The calculation formula is as follows:

[0091]

[0092] Where l m represents the bi-temporal multi-category semantic segmentation loss, Represents the predicted probability of the jth category, cls represents the true category, and N represents the total number of categories. Through the back propagation of the loss, the predicted category of the pixel position is made closer to the true category.

[0093] Step S42: Calculate the binary change mask loss using weighted contrast loss. The calculation formula is as follows:

[0094] l c =-y c ·log(p c )-(1-y c )·log(1-p c )

[0095]

[0096]

[0097] l ct =W p ·l cp +W n ·l cn

[0098] Where: l c represents the change binary mask loss, y c Indicates the actual label of the pixel position c, whether it changes or not, p c Represents the prediction result of the c pixel position. cp represents the loss of the actual change in the position of pixel c, l cn represents the loss of pixel position c which is not actually changed. cp The mask representing the actual change in the position of pixel c, y cn Represents the true unchanged mask of pixel position c. p The total number of pixels representing the change in the true label, n n The total number of pixels that represent changes in the true label. p is the calculation weight of the change loss, W n is the calculation weight of the unchanged loss, and the total loss of the changed mask is obtained by weighted addition. ct .

[0099] Step S43: Use weighted cosine similarity loss to further highlight the change information. The calculation formula is as follows:

[0100]

[0101] Where: l cos represents the cosine similarity loss, x iRepresents the classification probability vector arranged on a certain pixel position channel of the dual-phase image, cos(x1, x2) represents the cosine similarity of the probability vector, and max represents the boundary value. y represents the real label of whether the corresponding position has changed. When y=-1, it means that the x position has changed. At this time, the calculated cosine similarity should be small in theory. The cosine similarity of the changed position is optimized to 0 by optimizing the loss. y=1 means that the x position has not changed. At this time, the cosine similarity should be infinitely close to 1 in theory. The loss can be optimized to make the loss of the unchanged position tend to 1. The label information y is obtained by conversion from the real label.

[0102] Step S44: Based on the multiple loss functions constructed in steps S41, S42, and S43, the total loss of the network is obtained by weighted addition:

[0103] l t =λ1l dis +λ2l m +λ3l ct +λ4l cos

[0104] Where: l t represents the total loss of the network, λ i It represents the weights corresponding to different losses, and the entropy method is used to determine the weight ratios of different losses.

[0105] In this embodiment, step S5 specifically includes the following steps:

[0106] Step S51: Based on the image data preprocessed in step S1, a commonly used labeling method is selected, such as the built-in surface vector construction method of ArcGIS software, and the semantic boundaries of the buildings with dual-phase changes are manually visually interpreted. Field information is added to individual buildings, and different types of buildings are represented by different pixel values, including unchanged buildings, discrete buildings, simple single-purpose buildings, complex multi-purpose buildings, and residential buildings.

[0107] Step S52: Based on step S51, the vector label data is exported as single-channel raster data with projection and affine information, so as to vectorize the prediction result;

[0108] Step S53: Based on the sample data set constructed in S51 and S52, the data is cropped to S×S pixel size using a sliding window method, and the cropped images with a ratio of 20% to 80% of the changed buildings are selected. The training set, validation set, and test set are divided according to a certain ratio. In order to enhance the generalization ability of the model, the training sample data set is expanded by synchronously flipping the dual-phase images horizontally, vertically, and affinely changing;

[0109] Step S54: using the SGD optimizer and setting the initial learning rate lr as the optimizer for the segmentation of the changing building and the prediction of the binary changing mask;

[0110] Step S55: Use the modules in the attention model pre-trained on the large dataset ImageNet to initialize the network model weights to speed up the model convergence, improve the model training efficiency, and obtain a better model weight file.

[0111] In this embodiment, step S6 specifically includes the following steps:

[0112] Step S61: Based on the sample set obtained in step S5, set the number of epochs to 50, set reasonable learning rate, optimizer parameters and other hyperparameters to train the model, and store the weight file with the best effect on the validation set.

[0113] Step S62: Based on the weight file obtained in step S61, the images in the test set are predicted, and the prediction results of the dual-phase buildings are cropped using the binary mask image. The color index is used to convert it into an RGB image, and the affine information and projection information are added to obtain the final prediction result.

[0114] In this real example, the main urban area of ​​Quanzhou City, Fujian Province is used as the research object, and high-resolution remote sensing image data of the study area GF-2 in 2017 and 2020 are used to detect building changes in key disturbance areas in the study area, using red, green and blue bands. This example uses 2380 images with a pixel size of 256×256 and corresponding annotation data for model training, which are divided into 5 types of buildings, including unchanged buildings, discrete buildings, simple single-purpose buildings, complex multi-purpose buildings and residential apartments. Figure 2 For detailed comparison results, Figure 2 It contains high-resolution images of T1 and T2, as well as the corresponding actual labels and model prediction results. Combining the model prediction graph and labels, it can be seen that this model can not only detect boundary changes but also obtain the type of change.

[0115] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0116] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0117] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0119] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

Claims

1. A remote sensing detection method for changes in urban building types, characterized in that: The following steps are involved: Step S1: Obtain dual-temporal high-resolution remote sensing images of the study area and perform preprocessing operations on the images, including radiation correction, orthorectification, image fusion, and image resampling operations; Step S2: Construct a semantic segmentation task module based on Swim Transformer to extract multi-scale dual-temporal image features and learn building feature information; perform contrast loss calculation on the same-scale features of dual-temporal images; Step S3: Construct a change detection task module based on Swim Transformer to extract image change features and learn the change information between images; based on the multi-task concept, construct a dual-temporal multi-category building type change map and a binary change change map decoder respectively; Step S4: Design semantic boundary loss and binary change loss for building changes in dual-temporal images; specifically, weighted loss is used in multi-classification semantic segmentation tasks to alleviate the problem of imbalanced number of building types during training; Step S5: Create a multi-category building change sample set, divide the buildings into unchanged buildings, discrete buildings, simple single-purpose buildings, complex multi-purpose buildings and residential buildings; perform sliding cropping and sample enhancement operations on the dual-phase images to construct training and verification sample sets; Step S6: Model training and prediction, finally obtaining a building type change map and a binary map of the change area of ​​the dual-temporal image; Step S2 specifically includes the following steps: Step S21: Select Swim Transformer as the basic network, which mainly includes image fusion, downsampling module, mask module and multi-head moving window attention calculation module; Step S22: Based on the twin multi-task network learning framework, two modules of semantic segmentation and change detection are constructed, and two weight-sharing multi-layer attention networks are used to encode the features of the dual-phase images, while learning the building boundaries and their type information in the images; Step S23: Based on the multi-scale features extracted in step S22, a multi-scale contrast loss is constructed as follows: Cosine similarity is used to calculate the multi-scale similarity based on the channel direction. The calculation formula is as follows: In the formula, A and B are feature information sets of the same scale at T1 and T2 respectively. i ,B i Respectively represent the feature vectors of the same channel at T1 and T2; S(A,B) is the vector similarity, and the value is between [-1, +1]. The vector similarity S(A,B) is converted into the distance D. The calculation formula is as follows: Construct a feature distance D difference threshold to generate a pseudo label, and the pseudo label reflects the feature change in the channel direction; Based on the above steps, a pseudo label of whether the building information corresponding to the feature has changed can be obtained; a multi-scale feature loss calculation function is constructed based on the pseudo label; the extracted change features are expanded using multi-scale feature differences, and pseudo-change or unchanged features are reduced; the multi-scale feature loss calculation function is as follows: l dis =l1+l2+…+l m Where D is the calculation result in S32, l i Represents the feature difference loss of the i-th scale, label = 0 means that the feature corresponding to this channel has not changed, label = 1 means that the feature corresponding to this channel has changed; max is the maximum distance difference, when max is equal to 2, it means that the features are completely different; optimize the loss function, and integrate the multi-scale contrast loss by adding and calculating to form the total contrast loss; Step S3 specifically includes the following steps: Step S31: Based on the multi-scale features of the dual-phase images extracted in step S2, a change feature encoder based on Swim Transformer is constructed to obtain the change features of the phase images; Step S32: Based on the twin multi-task learning idea, a dual-phase building multi-classification decoder based on multi-scale building features is constructed, and a binary mask solution change map decoder based on image change features is constructed; a multi-scale feature fusion module is included to fuse multi-scale information in the decoding process; Step S4 specifically includes the following steps: Step S41: Based on the twin multi-task network model constructed in steps S2 and S3, a multi-task loss function is designed to optimize the network weights; specifically, the multi-classification cross entropy loss is used to calculate the loss of multi-classification semantic segmentation of dual-phase buildings, and the calculation formula is as follows: Where l m represents the bi-temporal multi-category semantic segmentation loss, represents the predicted probability of the jth category, cls represents the true category, and N represents the total number of categories; through the back propagation of the loss, the predicted category of the pixel position is made closer to the true category; Step S42: using weighted contrast loss to calculate binary change mask loss; the calculation formula is as follows: l c =-y c ·log(p c )-(1-)·log(1-p c ) l ct =W p ·l cp +W n ·l cn Where: l c represents the change binary mask loss, y c represents the actual label at the pixel position c, p c Represents the prediction result of the c pixel position; l cp represents the loss of the actual change in the position of pixel c, l cn represents the loss of pixel position c which is actually unchanged; y cp The mask representing the actual change in the position of pixel c, y cn represents the mask of the actual unchanged pixel position c; n p The total number of pixels representing the change in the true label, n n The total number of pixels representing the change in the true label; W p is the calculation weight of the change loss, W n is the calculation weight of the unchanged loss, and the total loss of the changed mask is obtained by weighted addition. ct ; Step S43: Use weighted cosine similarity loss to further highlight the change information; the calculation formula is as follows: Where: l cos represents the cosine similarity loss, x i Represents the classification probability vector arranged on a certain pixel position channel of the dual-phase image, cos(x1,x2) represents the cosine similarity of the probability vector, and max represents the boundary value; y represents the real label of whether the corresponding position has changed. When y=-1, it means that the x position has changed, and the cosine similarity of the changed position is optimized to 0 by optimizing the loss; y=1 means that the x position has not changed, and the loss of the unchanged position can be made to tend to 1 by optimizing the loss; the label information y is obtained by converting the real label; Step S44: Based on the multiple loss functions constructed in steps S23, S41, S42, and S43, the total loss of the network is obtained by weighted addition: l t =λ1l dis +λ2l m +λ3l ct +λ4l cos Where: l t represents the total loss of the network, λ i It represents the weights corresponding to different losses, and the entropy method is used to determine the weight ratios of different losses.

2. A remote sensing detection method for urban building type changes according to claim 1, characterized in that: Step S1 specifically includes the following steps: Step S11: Select high-resolution remote sensing images of two different time phases of the area to be studied, and preprocess the dual-time phase images, including atmosphere, radiation correction, and panchromatic-multispectral fusion operations; Step S12: Perform geo-registration and style transfer operations based on the dual-temporal images obtained in S11, and resample images of different resolutions to the same resolution.

3. A remote sensing detection method for urban building type changes according to claim 1, characterized in that: Step S5 specifically includes the following steps: Step S51: Based on the image data preprocessed in step S1, a commonly used labeling method is selected, such as the built-in surface vector construction method of ArcGIS software, and the semantic boundaries of the buildings with dual-temporal changes are manually visually interpreted, and field information is added to a single building, and different types of buildings are represented by different pixel values, including unchanged buildings, discrete buildings, simple single-purpose buildings, complex multi-purpose buildings, and residential buildings; Step S52: Based on step S51, the vector label data is exported as single-channel raster data with projection and affine information; Step S53: Based on the sample data set constructed in S51 and S52, the data is cropped into a size of S′×S′ pixels by using a sliding window method, and the cropped images with a ratio of 20% to 80% of the changed buildings are selected; the training set, the validation set, and the test set are divided according to a certain ratio; the training sample data set is expanded by synchronously using horizontal, vertical flipping, and affine change methods for the dual-phase images; Step S54: using the SGD optimizer and setting the initial learning rate lr as the optimizer for the segmentation of the changing building and the prediction of the binary changing mask; Step S55: Initialize the network model weights using the modules in the attention model pre-trained on the large dataset ImageNet.

4. A remote sensing detection method for urban building type changes according to claim 1, characterized in that: Step S6 specifically includes the following steps: Step S61: Based on the sample set obtained in step S5, set a reasonable number of iterations, learning rate, optimizer parameter hyperparameter training model, and store the best weight file in the validation set; Step S62: Based on the weight file obtained in step S61, the images in the test set are predicted, and the prediction results of the dual-phase buildings are cropped using the binary mask image; and the color index is used to convert it into an RGB image, and the affine information and projection information are added to obtain the final prediction results.

Citation Information

Patent Citations

  • Urban building change remote sensing detection method based on twin multitask network

    CN114821354A

  • Image processing method and apparatus, and electronic device and storage medium

    WO2022160753A1