New building identification method, device and equipment and medium
By converting synthetic aperture radar imagery into cloud-free multispectral imagery and combining it with a Siamese neural network model, the problem of low detection accuracy of optical multispectral imagery under low visibility or obstruction conditions is solved by extracting depth feature information and spatial dependencies, thus achieving high-precision identification of newly added buildings.
Patent Information
- Application Number
- CN202510977121.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies using single optical multispectral imaging for building change detection suffer from image information loss due to low visibility at night and obstructions such as clouds and haze, affecting detection accuracy.
A pre-trained adversarial generative network model is used to convert synthetic aperture radar images into cloud-free multispectral images. The Siamese neural network model is used to output the identification results of newly added buildings. Deep feature information and spatial dependencies are extracted through a dual-branch high-efficiency encoder and a dual-attention mechanism decoder to achieve high-precision building change detection.
It improves the accuracy of building change detection, solves the problem of missing image information due to low visibility or obstructions, and ensures accurate identification of new buildings under various weather conditions.
Smart Images

Figure CN120976735A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of remote sensing, and particularly relates to a newly-added building identification method, device, equipment and medium. BACKGROUND
[0002] With the development of remote sensing technology and the increase in the number of satellite constellations, remote sensing images have become an important data source for supporting natural resource investigation and monitoring work. The identification of newly-added buildings is an important basis for subsequent illegal building monitoring, map updating and other work, and the identification accuracy of newly-added buildings affects the quality of subsequent work.
[0003] When using single optical multispectral for building change detection, the resolution is high, and the image has spectral characteristics and a large amount of information. The defects are low night visibility and the existence of cloud cover and haze, which leads to image information loss and seriously affects the detection accuracy. SUMMARY
[0004] To solve the above problems, the present application provides a newly-added building identification method, device, equipment and medium, which can solve the problem of image information loss caused by low visibility or shielding, and improve the image detection accuracy.
[0005] The present application provides a newly-added building identification method, which comprises:
[0006] The pre-trained generative adversarial network model is used to convert the obtained synthetic aperture radar image into a cloud-free multispectral image;
[0007] The dual-time-phase cloud-free multispectral image is input into the pre-trained twin neural network model, and the identification result of the newly-added building in the image is output.
[0008] Preferably, the twin neural network model comprises a dual-branch efficient encoder and a dual-attention mechanism decoder.
[0009] The dual-time-phase cloud-free multispectral image is input into the pre-trained twin neural network model, and the identification result of the newly-added building in the image is output.
[0010] The pre-set dual-branch efficient encoder is used to extract deep feature information from the cloud-free multispectral image;
[0011] The spatial dependency context information and the channel dependency information in the deep feature information are captured according to the pre-set dual-attention mechanism decoder, the mask result of each pixel is output through feature mapping, and the identification result of the newly-added building is output according to the mask result of each pixel.
[0012] Preferably, the training process of the generative adversarial network model comprises:
[0013] obtaining synthetic aperture radar images registered with the cloud-free multispectral images as training data;
[0014] inputting the training data into a preset generator to generate fitted cloud-free multispectral images;
[0015] inputting the fitted cloud-free multispectral images and the registered cloud-free multispectral images into a preset discriminator to calculate an adversarial loss;
[0016] optimizing a mapping function in the generator according to the adversarial loss;
[0017] using the trained generator and discriminator as the adversarial generative network model.
[0018] As a preferred solution, the training process of the twin neural network model comprises:
[0019] obtaining registered double-time-phase images and corresponding change labels;
[0020] extracting deep feature information from the double-time-phase images using a double-branch efficient encoder;
[0021] capturing spatial dependency context information and channel dependency information in the deep feature information according to a double-attention mechanism decoder to obtain a result output;
[0022] calculating a loss value between the obtained result output and the change label using a preset loss function;
[0023] optimizing parameters of the double-branch efficient encoder and the double-attention mechanism decoder according to the calculated loss value as a training target;
[0024] using the trained double-branch efficient encoder and double-attention mechanism decoder as the twin neural network model.
[0025] Preferably, the loss function is:
[0026] wherein L is a loss value, p and y respectively represent the i-th predicted feature map and the i-th ground truth, Z represents a batch size, and a and β are hyperparameters. loss i i
[0027] Preferably, the double-branch efficient encoder comprises two single-branch EfficientNet B4 networks.
[0028] Each single-branch EfficientNet B4 network is a layer of convolutional layer structure, seven layers of repeatedly stacked MBConv structure, and a layer of convolutional layer structure.
[0029] Preferably, the dual attention mechanism decoder comprises a position attention module and a multi-scale fusion attention module.
[0030] The position attention module respectively performs feature mapping on the input features through three convolutional layers, resamples and transposes the multiplication of two mapped features, and uses a softmax layer to calculate to obtain a spatial attention map; another mapped feature is transposed and multiplied with the spatial attention map and resampled to obtain a spatial feature.
[0031] The multi-scale fusion attention module performs convolutional layer processing on high-level features and low-level features in the spatial feature to obtain feature mapping; compresses the obtained feature mapping through a global average pooling layer and generates channel statistical information to obtain pixel representation of each pixel; and activates channel relationship through a fully connected layer and an activation function to obtain output features.
[0032] The embodiment of the present application also provides a new building identification device, the device comprises:
[0033] The conversion module is configured to convert the acquired synthetic aperture radar image into a cloud-free multispectral image using a pre-trained generative adversarial network model.
[0034] The identification module is configured to input the dual-phase cloud-free multispectral image into a pre-trained twin neural network model to output an identification result of the new building in the image.
[0035] Another embodiment of the present application provides a terminal device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the new building identification method of any one of the above embodiments.
[0036] Another embodiment of the present application provides a computer-readable storage medium, comprising a stored computer program, wherein the computer program controls the device where the computer-readable storage medium is located to execute the new building identification method of any one of the above embodiments when the computer program is running.
[0037] This invention provides a method, apparatus, device, and medium for identifying newly added buildings. It employs a pre-trained generative adversarial network model to convert acquired synthetic aperture radar (SAR) images into cloud-free multispectral images. The dual-temporal cloud-free multispectral images are then input into a pre-trained Siamese neural network model, which outputs the identification result of newly added buildings in the images. This solution addresses the problem of inaccurate image information due to low visibility or obstructions, thereby improving image detection accuracy. Attached Figure Description
[0038] Figure 1 This is a flowchart illustrating a new building identification method provided in an embodiment of the present invention;
[0039] Figure 2 This is another flowchart illustrating the new building identification method provided in this embodiment of the invention;
[0040] Figure 3 This is a schematic diagram of the training process of the adversarial generative network model provided in an embodiment of the present invention;
[0041] Figure 4 This is a schematic diagram of the structure of the single-branch EfficientNet B4 network provided in an embodiment of the present invention;
[0042] Figure 5 This is a schematic diagram of the dual attention mechanism decoder provided in an embodiment of the present invention;
[0043] Figure 6 This is a structural schematic diagram of a newly added building identification device provided in an embodiment of the present invention;
[0044] Figure 7 This is a schematic diagram of the structure of a terminal device provided in an embodiment of the present invention. Detailed Implementation
[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0046] This application provides a new building identification method, see [link to relevant documentation] Figure 1 This is a flowchart illustrating a method for identifying newly added buildings provided in an embodiment of the present invention. The method includes steps S1 to S2:
[0047] Step S1: The acquired synthetic aperture radar imagery is converted into cloud-free multispectral imagery using a pre-trained adversarial generative network model.
[0048] Step S2: Input the cloudless multispectral images of the two time phases into the pre-trained twin neural network model, and output the recognition results of newly added buildings in the image.
[0049] In this specific implementation, the Generative Adversarial Network (GAN) model consists of a generator and a discriminator, which are continuously optimized through adversarial training. In the process of converting Synthetic Aperture Radar (SAR) imagery into cloudless multispectral imagery, the generator aims to learn the mapping relationship between SAR imagery and cloudless multispectral imagery, generating the most realistic cloudless multispectral imagery possible; the discriminator is responsible for distinguishing between real cloudless multispectral imagery and imagery generated by the generator. Through continuous adversarial training, the generator gradually becomes able to convert SAR imagery into high-quality cloudless multispectral imagery. Since SAR imagery has the ability to acquire data in all weather conditions and at all times, while multispectral imagery is more conducive to the identification of ground features, using a GAN to achieve the conversion between the two can fully leverage the advantages of both types of imagery.
[0050] The Siamese neural network consists of two subnetworks with identical structures and shared weights. Dual-temporal, cloud-free multispectral images are input into these two subnetworks, which then extract features from the input images. By calculating the similarity between the two feature vectors, the network determines whether changes exist in the images, thereby identifying newly added buildings. Through pre-training, this network learns the characteristic patterns and change patterns of buildings in images from different periods. When new dual-temporal images are input, it can quickly and accurately determine the location and extent of newly added buildings.
[0051] In practice, synthetic aperture radar (SAR) images to be processed are acquired and converted into cloudless multispectral images using a pre-trained adversarial generative network (GAN). Cloudless multispectral images from another period corresponding to these images are then selected to form a dual-temporal imagery.
[0052] The cloudless multispectral images of the two time phases are input into a pre-trained twin neural network model. The model outputs the recognition results of newly added buildings in the images, including information such as the location and extent of the newly added buildings.
[0053] See Figure 2 This is another flowchart illustrating the new building identification method provided in this embodiment of the invention.
[0054] First, a GAN model is used to convert SAR images into optical images, and the model is trained to improve its accuracy. Next, change detection is performed on the generated multispectral images to extract newly added buildings from the range of changes. The trained GAN model is then used to convert the SAR images to obtain cloud-free multispectral images, and a Siam-EMNet neural network model is used for change detection. The output is a binary change mask image where each pixel has an output value of 0 or 1, representing "no change" or "new building," respectively.
[0055] It should be noted that the model needs to undergo a pre-training process. A large amount of synthetic aperture radar imagery data and corresponding cloud-free multispectral imagery data are collected. The collected data undergoes preprocessing, including geometric correction, radiometric correction, cropping, and normalization, to ensure the data meets the requirements for model training.
[0056] The data is divided into training, validation, and test sets according to a certain ratio for the training, validation, and evaluation of generative adversarial network models and Siamese neural network models.
[0057] The Generative Adversarial Network (GAN) model is trained by inputting SAR imagery from the training set into the generator to produce preliminary cloud-free multispectral images. These generated images, along with real cloud-free multispectral images, are then input into a discriminator, which outputs its judgment. Based on the discriminator's feedback, the parameters of both the generator and discriminator are adjusted, and this process is repeated until the generator produces images of high quality that the discriminator can hardly distinguish between the generated and real images. During training, a validation set is used to evaluate the model and prevent overfitting.
[0058] To train a Siamese neural network model, cloudless multispectral imagery from two time phases (the training set) is input into two sub-networks of the Siamese neural network. The similarity between the two feature vectors is calculated and compared with the actual change labels (whether new buildings have been added), resulting in a loss function. The network parameters are then adjusted using backpropagation to minimize the loss function. Similarly, a validation set is used to evaluate and optimize the model during training.
[0059] This invention addresses the problem of severe cloud obstruction in optical imagery during the detection of suspected illegal buildings by proposing a detection method that fuses multispectral and SAR multimodal data. SAR can provide high-resolution imaging under various weather conditions, supplementing optical remote sensing images with additional information. It also overcomes the limitations of SAR images, which suffer from significant noise and are difficult to interpret, and multispectral images, which are susceptible to weather and lighting conditions.
[0060] Regarding the choice of fusion method, image-to-image translation (I2I) based on deep generative adversarial networks (GANs) shows great potential in mapping images from two different domains while preserving the main content, such as style transfer and super-resolution. Therefore, this invention uses a GAN model, trains it to convert SAR imagery into multispectral imagery, evaluates the accuracy of the model, and uses the de-clouded multispectral imagery it converts to detect suspected illegal buildings.
[0061] In another embodiment provided by the present invention, the Siamese neural network model includes a dual-branch high-efficiency encoder and a dual-attention mechanism decoder;
[0062] The step of inputting dual-temporal passive multispectral images into a pre-trained Siamese neural network model and outputting the recognition results of pixel changes in the image includes:
[0063] Depth feature information is extracted from the passive multispectral image using a preset dual-branch high-efficiency encoder;
[0064] The decoder captures the spatial dependency context information and channel dependency information in the depth feature information according to the preset dual attention mechanism, outputs the mask result of each pixel through feature mapping, and outputs the recognition result of the newly added building based on the mask result of each pixel.
[0065] In this specific implementation, the Siamese neural network model includes a dual-branch high-efficiency encoder and a dual-attention mechanism decoder.
[0066] This study uses a dual-branch high-efficiency encoder to explore the impact of different combinations of network width, depth, and image resolution on experimental accuracy. After receiving passive multispectral images from two temporal phases, the dual-branch high-efficiency encoder can rapidly extract depth feature information from the images using an efficient structure and algorithm. These depth features contain key information about buildings and other ground features.
[0067] The dual-attention mechanism decoder processes deep feature information from two dimensions. The spatial dependency context information capture module analyzes the spatial relationships between pixels to understand the spatial structure and layout of features; the channel dependency information capture module uncovers potential connections between different spectral channels, highlighting spectral features valuable for building identification. Through feature mapping, the final output is a mask for each pixel, thereby accurately determining the location and extent of newly added buildings.
[0068] A dual-temporal passive multispectral image training set is input into a dual-branch high-efficiency encoder to extract depth feature information. These features are then fed into a dual-attention mechanism decoder, which captures spatial and channel dependency information and performs feature mapping, outputting pixel mask results. By comparing with the true change labels, the loss function is calculated, and the network parameters are optimized using the backpropagation algorithm. Similarly, a validation set is used for model evaluation and optimization during the training process.
[0069] The dual-temporal images are input into the trained Siamese neural network model. The dual-branch high-efficiency encoder extracts depth features, and the dual-attention mechanism decoder processes the features and outputs pixel mask results. Based on this, the identification results of the newly added buildings are obtained, and the specific location and range of the newly added buildings are determined.
[0070] In another embodiment provided by the present invention, the adversarial generative network model training process includes:
[0071] Acquire synthetic aperture radar images registered with cloudless multispectral images as training data;
[0072] The training data is input into a preset generator to generate a fitted cloudless multispectral image.
[0073] The fitted cloudless multispectral image and the registered cloudless multispectral image are input into a preset discriminator to calculate the adversarial loss.
[0074] The mapping function in the generator is optimized based on the adversarial loss;
[0075] The trained generator and discriminator are used as the adversarial generative network model.
[0076] In this specific implementation, the goal of GAN model training is to learn how to synthesize cloudless multispectral images from SAR images through a generative adversarial network (GAN), thereby providing higher quality image input for subsequent building detection.
[0077] The training data input consists of SAR images with the same 3m resolution for both registration and multispectral imaging, and the output is the corresponding cloud-free multispectral image. In the network structure, the generator G converts the SAR image into a fitted multispectral image, and the discriminator D... y The generator G determines whether the generated image is indistinguishable from the real multispectral image. After training, the generator G can convert any SAR image into a pseudo-multispectral image.
[0078] As a general training framework for approximating generative models, GAN-based models typically consist of a generator G and a discriminator D. y Composition, trained in an adversarial manner, see Figure 3 This is a schematic diagram illustrating the training process of the adversarial generative network model provided in an embodiment of the present invention. Figure 3 As shown, given training images from the source domain X and the target domain Y, I2I aims to learn a mapping function G such that, given any unseen image in domain X, it can synthesize a fake image indistinguishable from a real image in domain Y, while preserving semantic content. The generator and discriminator are alternately optimized to compete with each other. The generator's goal is to generate fake images that can fool the discriminator, while the discriminator is trained to distinguish between fake and real images. The overall training objective of the GAN is called the adversarial loss, denoted as:
[0079]
[0080] Since the only monitoring signal comes from discriminator D ySince it makes predictions based on high-level features, the detailed reconstruction performance of the generator G is not optimal. Therefore, recent GAN-based I2I algorithms have utilized additional loss terms to further enhance the original GAN, which is optimized with adversarial loss. Depending on the availability of training data, GAN-based I2I algorithms can be categorized into two types: unsupervised and supervised methods. For unsupervised methods, images from both the source and target domains are required. For supervised methods, training images from both domains need to be paired. Pix2Pix is a representative supervised I2I algorithm that adds an additional loss term, guiding the generator G not only to deceive the discriminator but also to produce near-ground-truth fake images in the image space. The training objective is:
[0081]
[0082] L L1 (G)=E x~X,y~Y [||G(x)-y||1;
[0083] In addition to combating the loss term L GAN In addition, a pixel-level reconstruction loss with weight coefficients λ and L1 in the form of L1 norm is used to force the fake image to be similar to the real image in the target domain.
[0084] CycleGAN is a representative unsupervised I2I algorithm that introduces the transitivity concept into I2I tasks. It uses two generators (G and F) and a discriminator (D). y and D x ) is used to learn two mapping functions, namely G and D. y Represent X→Y and F, D x Let Y → X be the representation. CycleGAN uses a cycle consistency loss to regularize the mapping function, rather than minimizing the distance between real and fake images. CycleGAN assumes that the two mapping functions should be inverses of each other, ensuring that the translated images maintain consistency with themselves: F(G(x))≈x and G(F(y))≈y represent x∈X and y∈Y, respectively.
[0085] A new loss term, with weighting coefficient λL cyc The combined losses from the two adversarial approaches constitute the ultimate training objective of CycleGAN:
[0086]
[0087] L cyc (G, F) = E x~X [||xF(G(x))||1]+E y~Y [||yG(F(y))||1;
[0088] Among them, L GAN (G,D y ) represents the generator G and the discriminator D. y Adversarial loss from x to y, L GAN (F,D X ) represents the generator F and the discriminator D. x The adversarial loss from y to x; The cyclic consistency loss includes reconstruction errors in two directions: x→G(x)→F(G(x))≈x, y→F(y)→G(F(y))≈y, and λL. cyc : The weighting coefficient of the cycle consistency loss, used to control its influence on the total loss.
[0089] Cyclic consistency loss function L cyc (G, F), E x~X Let E represent the expectation of a sample x from the source domain X. y~Y This represents the expectation of the sample y from the target domain Y. Specifically, the first transformation path x→G(x)→F(G(x)) ensures that the image originating from the source domain X can be restored back to itself, and the second transformation path is y→F(y)→G(F(y)), which ensures that the image originating from the target domain Y can be restored back to itself.
[0090] The mapping function in the generator is optimized based on the adversarial loss;
[0091] The trained generator and discriminator are used as the adversarial generative network model.
[0092] In another embodiment of the present invention, the training process of the Siamese neural network model includes:
[0093] Obtain the registered dual-temporal images and their corresponding change labels;
[0094] A dual-branch high-efficiency encoder is used to extract depth feature information from the dual-temporal images;
[0095] The decoder captures the spatial dependency context information and channel dependency information in the deep feature information according to the dual attention mechanism, and then outputs the result.
[0096] The loss value between the result calculated using a preset loss function and the change label is calculated.
[0097] The calculated loss value is used as the training target to optimize the parameters of the dual-branch efficient encoder and the dual-attention mechanism decoder.
[0098] The trained dual-branch high-efficiency encoder and dual-attention mechanism decoder are used as the Siamese neural network model.
[0099] In this specific implementation, the Siamese neural network model training uses registered bi-temporal images and their corresponding change labels to provide learning samples for the model. The bi-temporal images record the state of the same region at different times, while the change labels indicate the actual changes that occurred; together, they constitute the supervisory information for model training.
[0100] The dual-branch high-efficiency encoder is based on a convolutional neural network (CNN) architecture, with two branches having identical structures and sharing weights. When two temporal images are input into the two branches respectively, the encoder extracts local features of the image step by step through multiple convolutional operations, and reduces data dimensionality and expands the receptive field through pooling operations, ultimately extracting deep features containing semantic and structural information of the image. Its efficiency is reflected in the optimization of the network structure, reducing redundant computation, and improving computational efficiency while ensuring the quality of feature extraction.
[0101] The dual-attention decoder comprises spatial attention and channel attention mechanisms. Spatial attention focuses on the spatial relationships between pixels in an image, highlighting regions relevant to target changes through weighted processing of feature maps. Channel attention, on the other hand, focuses on the importance differences between different spectral channels or feature channels, enhancing channel information valuable for target identification while suppressing irrelevant information. Working together, these two mechanisms deeply mine spatial and channel dependencies within deep features, enabling the model to understand image features from different perspectives and laying the foundation for accurate pixel change output.
[0102] Pre-defined loss functions (such as cross-entropy loss function, Dice loss function, etc.) are used to quantify the difference between the model's predicted results and the actual change labels. By calculating the loss value, the degree of deviation in the model's current prediction is determined. Based on gradient descent algorithms (such as stochastic gradient descent, Adam optimization algorithm, etc.), with the loss value as the optimization objective, the loss gradient is passed from the decoder to the encoder using the backpropagation algorithm to update the weights and bias parameters in the network, so that the model's predicted results continuously approach the actual change labels, gradually improving the model's recognition ability.
[0103] During training, dual-temporal image data from different regions, seasons, and weather conditions are collected from various sources, including satellite remote sensing and aerial photography. Simultaneously, precise change labels are created through manual annotation and integration with Geographic Information System (GIS) data to ensure that the labels accurately reflect changes in new buildings and land features within the images.
[0104] Feature point matching algorithms (such as SIFT and SURF) or deep learning-based registration methods (such as Deformable Convolution Networks) are used to register bi-temporal images. By finding corresponding points in the two images, a transformation model (such as affine transformation and perspective transformation) is established to adjust the images to the same coordinate system, ensuring that the positions of the same features correspond in the two images.
[0105] The registered images are normalized to grayscale, mapping pixel values to a uniform range (e.g., [0,1] or [-1,1]); images are cropped to remove useless edge regions and to standardize image size; one-hot encoding is performed on varying labels to adapt them to the model's input requirements. The processed data is then divided into training, validation, and test sets, typically in a 7:1:2 ratio.
[0106] The dual-temporal images of the training set are input into a dual-branch high-efficiency encoder. The two branches extract features from the images and output the corresponding depth feature vectors.
[0107] The deep feature vector output by the encoder is input into the dual attention mechanism decoder. The spatial attention mechanism and the channel attention mechanism process the features in turn. After upsampling, convolution and other operations, the mask result of each pixel is output, that is, the predicted pixel change.
[0108] The masking result output by the decoder and the corresponding real change label are substituted into the preset loss function to calculate the loss value and evaluate the accuracy of the model's current prediction.
[0109] Using the backpropagation algorithm, the gradients of each parameter are calculated based on the loss value. An optimizer (such as Adam) is then used to update the parameters of the efficient dual-branch encoder and the dual-attention decoder to reduce the loss value. The above steps of feature extraction, information processing, loss calculation, and parameter optimization are repeated continuously. During training, the model performance is periodically evaluated using a validation set to prevent overfitting.
[0110] Training is stopped when the loss value stabilizes and becomes low on both the training and validation sets, and the model's accuracy and recall on the validation set meet the expected requirements. The finally trained model is then independently evaluated using a test set. If the model maintains good performance on the test set, the dual-branch efficient encoder and dual-attention decoder are selected as the final Siamese neural network model for use in actual new building recognition tasks.
[0111] Through the synergistic effect of a dual-branch efficient encoder and a dual-attention mechanism decoder, the model can accurately capture subtle feature changes of newly added buildings in dual-temporal images, achieving pixel-level recognition accuracy. The dual-branch efficient encoder reduces the computational complexity of the model and accelerates the feature extraction process; the dual-attention mechanism decoder specifically enhances key information, improving the model's efficiency in utilizing effective features. During the training phase, the model converges faster, shortening the training time.
[0112] In another embodiment provided by the present invention, the loss function is:
[0113] Among them, L loss p is the loss value. i and y i Let Z and β represent the predicted i-th feature map and i-th ground truth, respectively, where Z represents the batch size, and α and β are hyperparameters.
[0114] In this specific implementation, the dice loss function is a good choice for scenarios with an imbalance between changed and unchanged samples. This function focuses more on mining information from changed regions during training, but it is unstable for training small target buildings and can lead to gradient overfitting in extreme cases. The cross-entropy loss function measures the distance between the actual output and the predicted output. The smaller the entropy value, the closer the two probability distributions are, and the entropy value also determines the experimental training gradient. The smaller the value, the slower the gradient update; the larger the gradient, the faster the parameter update. Parameter update iterations can be compensated for unstable results using a single dice loss function. Therefore, a hybrid loss function, a weighted combination of dice loss and cross-entropy loss, is used to optimize Siam-EMNet. It is defined as:
[0115]
[0116] Where, p i and y i Let β and α represent the predicted feature map and ground truth, respectively, and Z represent the batch size. Furthermore, two hyperparameters (0 < β < 5 and 0 < α < 5) are used to control the effect of the loss function. During training, the optimal training results are obtained when β = 0.5 and α = 2 by adjusting the parameters.
[0117] In another embodiment provided by the present invention, the dual-branch high-efficiency encoder includes two single-branch EfficientNet B4 networks;
[0118] Each single-branch EfficientNet B4 network consists of a single convolutional layer, seven stacked MBConv layers, and a single convolutional layer.
[0119] In this specific implementation, a Siam-EMNet network model is proposed to detect changes in buildings. This model uses bi-temporal images as input and output binary detection maps. The encoder consists of a Siamese dual-branch EfficientNet B4, with the two branches sharing weights. To improve the model's accuracy, this structure employs a composite scaling method for the network's width, depth, and resolution. The PAB module in the decoder is used to obtain spatial dependencies between pixels, while the MFAB module obtains channel dependencies between arbitrary feature maps by fusing high and low semantic features. Combined with a dual attention mechanism, multi-scale feature information in the image can be effectively utilized to improve accuracy.
[0120] Encoder-decoder based networks exhibit high feature extraction accuracy. The Siam-EMNet network can comprehensively and efficiently extract bi-temporal feature information from VHR remote sensing images, achieving efficient fusion of multi-level information. The encoder uses ImageNet pre-trained EfficientNet B4 as the backbone to extract deep features, improving overall accuracy. Furthermore, skip connections transfer encoder features to the decoder, enabling more effective fusion of deep and shallow features.
[0121] Highly efficient B4 encoders with two branches typically optimize model accuracy by increasing network depth, width, and resolution. AlexNet was the first to combine dropout, ReLU, and LRN techniques with CNNs, expanding the width and depth of CNNs. VGGNet, ResNet, and InceptionNet optimize models by increasing network depth and width, respectively. Huang et al. improved model performance by increasing image resolution. Most previous studies only adjusted width, depth, or resolution. EfficientNet is a family of CNNs proposed by Tan et al. based on existing CNNs. This study extends different combinations of network width, depth, and image resolution, exploring the impact of different combinations on experimental accuracy, and proposes eight versions from EfficientNet B0 to EfficientNet B7. Taking B0 as an example, each network is divided into nine blocks, Blocks 1 to 9. See [link to relevant documentation] Figure 4 This is a schematic diagram of the structure of the single-branch EfficientNet B4 network provided in an embodiment of the present invention.
[0122] The dual-branch efficient network B4 extracts feature information, with the two branches sharing weights, effectively utilizing feature information. Figure 4The structure of a single-branch EfficientNet B4 is shown. The first block is a convolutional layer with a kernel size of 3×3 and a stride of 2. Blocks 2 through 8 repeat the stacked MBConv structure, and Block 9 is a normal 1×1 convolutional layer.
[0123] EfficientNet optimizes the network in three dimensions: width, depth, and resolution. It utilizes Neural Architecture Search (NAS) to obtain its composite parameters, as described below:
[0124] max d,w,r Accuracy(Nas(s,d,r));
[0125]
[0126] Memory(Nas)≤target_memory;
[0127] FLOPS(Nas) ≤ target_flops;
[0128] in, and For the predefined parameters of the basic network, w, d, and r are coefficients representing the network width, depth, and resolution, respectively; st represents the constraints, ⊙ i=1,…s This represents a series multiplication operation. This indicates that Fi is repeated in the data block. Next, d represents depth scaling. represents the feature matrix of the input data block; target_memory and target_flops are restricted to Memory(Nas) and FLOPS(Nas) respectively, and the Memory(Nas) and FLOPS(Nas) of each model are the optimal values less than or equal to the restriction values.
[0129] In another embodiment provided by the present invention, the dual attention mechanism decoder includes a positional attention module and a multi-scale fusion attention module;
[0130] The location attention module performs feature mapping on the input features through three convolutional layers. It resamples and transposes two of the mapped features and multiplies them, then uses a softmax layer to calculate the spatial attention map. The other mapped feature is then transposed and multiplied with the spatial attention map and resampled to obtain the spatial features.
[0131] The multi-scale fusion attention module processes the high-level and low-level features in the spatial features through convolutional layers to obtain feature maps; it then compresses the obtained feature maps and generates channel statistics through a global average pooling layer to obtain the pixel representation of each pixel; finally, it activates the channel relationships through a fully connected layer and an activation function to obtain the output features.
[0132] In this specific implementation, the local feature information acquired by traditional CNNs may lead to false detections of targets. The decoder of the Siam-EMNet network model learns the MANet decoding structure, which has been maturely used for semantic segmentation of medical images. To establish a rich contextual connectivity model based on local features, dual attention modules (PAB and MFAB) are designed to capture the spatial and channel information of the decoding part. Inspired by the literature, PAB and MFAB are introduced. The PAB module utilizes rich spatial contextual information for modeling, enhancing its representational power. MFAB captures feature-channel relationships by combining high-level and low-level feature maps. Feature maps are enhanced and suppressed according to their importance in the building segmentation task. See [link to relevant documentation] Figure 5 This is a schematic diagram of the dual-attention mechanism decoder provided in an embodiment of the present invention. The structures of the two modules are as follows: Figure 5 As shown, Figure 5 (a) in the diagram represents the position attention module PAB. Figure 5 (b) in the text refers to b. Multiscale Fusion Attention Module (MFAB).
[0133] exist Figure 5 In a, firstly, The input with local features is used in the convolutional layer to generate two new feature maps, B and D, respectively. B and D were resampled as (N is the total number of pixels, N = H × W), then transpose and multiply, and use a softmax layer to calculate the spatial attention map.
[0134]
[0135] Here, Qji represents the influence of position i on position j. Simultaneously, another feature map generated by A is... Resampling
[0136] Then, multiply the transposes of matrices E and Q, and resample the result as follows: Finally, by applying the scale parameter α, pixel-by-pixel summation is performed on feature A to obtain the final output.
[0137]
[0138] Here, it is initialized to 0, and more new weights are obtained through training. P selectively aggregates context based on spatial attention graphs, improving intra-class compactness and semantic consistency.
[0139] Figure 5 b shows the structure of the MFAB module, and its operation steps are as follows.
[0140] Advanced Features DH* in The results were obtained by feeding them into 1×1 and 3×3 convolutional layers. DH in and DL in They have the same number of channels. The output feature map u = [u1, u2, ..., u... k The following can be calculated:
[0141]
[0142] Among them, D in =[d 1 ,d 2 ,…,d k ],D in ∈(DL in or DH in ), D represents the channel corresponding to the convolution kernel. in , and * denote convolution.
[0143] Global average pooling is used to compress feature U and generate channel statistics, denoted as S1 and S2 respectively. Its k-th pixel can be represented as...
[0144]
[0145] Where W and H represent width and height respectively, u k This represents the feature map for each channel.
[0146] A bottleneck layer is used to limit the complexity of the model and obtain channel relationship information x1 and x2.
[0147] x1=F LS (S1,R)=θ1(R1θ2(R2θ1));
[0148] x2=F HS (S2,R)=θ1(R1θ2(R2θ2));
[0149] Where R1 and R2 represent fully connected layers, and θ1 and θ2 represent sigmoid and ReLU, respectively. Then F is used. add The function combines the channel-by-channel outputs of low-level feature x1 and high-level feature x2.
[0150] x = F add (·)=x1+x2;
[0151] Activate the channel relationship information x, rescale the feature U, and obtain DH. out .
[0152]
[0153] exist It is an extended property of the information set U, V k This is information about channel relationship settings. and F base (T k V k ) is to obtain V k Multiplying them channel by channel yields the feature map. In addition, DH out DL in Connect them together to get the final output DH* out .
[0154] Compared to methods that rely solely on multispectral or SAR imagery for building change detection, this method trains a GAN model to convert SAR imagery into multispectral imagery, evaluates its accuracy, and applies it to scenarios with abundant cloud cover. Due to the imaging principle of SAR imagery, it is less affected by weather conditions, thus the multispectral imagery converted from SAR imagery undergoes cloud removal processing. Subsequently, the Siam-EMNet network model is used to detect suspected illegal buildings on the multispectral imagery converted from SAR imagery. This avoids the significant noise interference present in SAR imagery and overcomes the limitations of multispectral imagery, which is susceptible to weather and lighting conditions.
[0155] While most EO-SAR image conversion models are still trained using images at different resolutions, this invention uses data at a uniform 3m resolution. This will improve the accuracy of model training to some extent, making the conversion from SAR images to optical images more precise.
[0156] This application utilizes the fusion of multispectral and SAR multimodal imaging for the detection of suspected illegal buildings. Previous methods either used multispectral or SAR imagery for detection, never fusing the two. In this invention, a model for converting SAR imagery into multispectral imagery is trained based on a GAN model, its accuracy is evaluated, and it is applied to scenes with abundant cloud cover. Due to the imaging principle of SAR imagery, it is less affected by weather, and the multispectral imagery converted from SAR imagery undergoes cloud removal processing. This method avoids the significant noise interference present in SAR imagery and overcomes the limitations of multispectral imagery, which is susceptible to weather and lighting conditions.
[0157] Another embodiment of the present invention provides a new building identification device, see [link to previous document]. Figure 6 This is a structural schematic diagram of a newly added building identification device provided in an embodiment of the present invention. The device includes:
[0158] The conversion module is used to convert the acquired synthetic aperture radar imagery into cloud-free multispectral imagery using a pre-trained adversarial generative network model.
[0159] The recognition module is used to input cloudless multispectral images from two time phases into a pre-trained twin neural network model and output the recognition results of newly added buildings in the images.
[0160] The newly added building identification device provided in this embodiment can perform all the steps and functions of the newly added building identification method provided in any of the above embodiments. The specific functions of the device will not be described in detail here.
[0161] See Figure 7 This is a schematic diagram of the structure of a terminal device provided in an embodiment of the present invention. The terminal device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a new building recognition program. When the processor executes the computer program, it implements the steps in the various embodiments of the new building recognition method described above, for example... Figure 1 The steps S1 to S2 are shown. Alternatively, when the processor executes the computer program, it implements the functions of each module in the above-described device embodiments.
[0162] For example, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the terminal device. For example, the computer program can be divided into a detection module, an output power control module, and a window control module. The specific functions of each module have been described in detail in the above embodiment of a new building recognition method, and the specific functions of this device will not be repeated here.
[0163] The terminal device described can be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the schematic diagram is merely an example of a terminal device and does not constitute a limitation on any terminal device. It may include more or fewer components than illustrated, or combine certain components, or use different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.
[0164] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0165] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0166] If the module integrated into the terminal device is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0167] It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications are also considered to be within the scope of protection of this invention.
Claims
1. A new building identification method, characterized in that, The method includes: A pre-trained adversarial generative network model is used to convert the acquired synthetic aperture radar imagery into cloud-free multispectral imagery. The cloudless multispectral images of two time phases are input into a pre-trained twin neural network model, and the output is the recognition result of newly added buildings in the image.
2. The method for identifying newly added buildings according to claim 1, characterized in that, The twin neural network model includes a dual-branch high-efficiency encoder and a dual-attention mechanism decoder; The step of inputting dual-temporal passive multispectral images into a pre-trained Siamese neural network model and outputting the recognition results of pixel changes in the image includes: Depth feature information is extracted from the passive multispectral image using a preset dual-branch high-efficiency encoder; The decoder captures the spatial dependency context information and channel dependency information in the depth feature information according to the preset dual attention mechanism, outputs the mask result of each pixel through feature mapping, and outputs the recognition result of the newly added building based on the mask result of each pixel.
3. The method for identifying newly added buildings according to claim 1, characterized in that, The training process of the adversarial generative network model includes: Acquire synthetic aperture radar images registered with cloudless multispectral images as training data; The training data is input into a preset generator to generate a fitted cloudless multispectral image. The fitted cloudless multispectral image and the registered cloudless multispectral image are input into a preset discriminator to calculate the adversarial loss. The mapping function in the generator is optimized based on the adversarial loss; The trained generator and discriminator are used as the adversarial generative network model.
4. The method for identifying newly added buildings according to claim 1, characterized in that, The training process of the twin neural network model includes: Obtain the registered dual-temporal images and their corresponding change labels; A dual-branch high-efficiency encoder is used to extract depth feature information from the dual-temporal images; The decoder captures the spatial dependency context information and channel dependency information in the deep feature information according to the dual attention mechanism, and then outputs the result. The loss value between the result calculated using a preset loss function and the change label is calculated. The calculated loss value is used as the training target to optimize the parameters of the dual-branch efficient encoder and the dual-attention mechanism decoder. The trained dual-branch high-efficiency encoder and dual-attention mechanism decoder are used as the Siamese neural network model.
5. The method for identifying newly added buildings according to claim 1, characterized in that, The loss function is: Among them, L loss p is the loss value. i and y i Let Z represent the predicted i-th feature map and the i-th ground truth, respectively, where Z represents the batch size, and α and β are hyperparameters.
6. The method for identifying newly added buildings according to claim 4, characterized in that, The dual-branch high-efficiency encoder comprises two single-branch EfficientNet B4 networks; Each single-branch EfficientNet B4 network consists of a single convolutional layer, seven stacked MBConv layers, and a single convolutional layer.
7. The method for identifying newly added buildings according to claim 4, characterized in that, The dual-attention mechanism decoder includes a positional attention module and a multi-scale fusion attention module; The location attention module performs feature mapping on the input features through three convolutional layers. It resamples and transposes two of the mapped features and multiplies them, then uses a softmax layer to calculate the spatial attention map. The other mapped feature is then transposed and multiplied with the spatial attention map and resampled to obtain the spatial features. The multi-scale fusion attention module processes the high-level and low-level features in the spatial features through convolutional layers to obtain feature maps; The obtained feature map is compressed and channel statistics are generated by a global average pooling layer to obtain the pixel representation of each pixel; Output features are obtained by activating channel relationships through fully connected layers and activation functions.
8. A new building identification device, characterized in that, The device includes: The conversion module is used to convert the acquired synthetic aperture radar imagery into cloud-free multispectral imagery using a pre-trained adversarial generative network model. The recognition module is used to input cloudless multispectral images from two time phases into a pre-trained twin neural network model and output the recognition results of newly added buildings in the images.
9. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the new building identification method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the new building identification method as described in any one of claims 1 to 7.