Building roof recognition method, device, electronic device and storage medium

By using the building segmentation model built with the Transformer network and the improved attention mechanism, the problem of low building roof recognition accuracy was solved, fast and high-precision roof recognition was achieved, the cost of manual recognition was reduced, and the user experience was improved.

CN116188974BActive Publication Date: 2025-10-03GUANGDONG POWER GRID CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211686806.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-10-03
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

The existing technology has low accuracy in identifying building roofs, resulting in high manual recognition costs and poor user experience.

Method used

By acquiring the remote sensing image to be identified, the preset building segmentation model is used to perform building segmentation, and the hybrid framework structure model built based on the Transformer network is used to identify the building roof. Combined with the improved attention mechanism and the cross-shaped window context interaction module, global and local context information is captured.

Benefits of technology

It achieves high-precision and rapid recognition of building roofs, reduces manual recognition costs, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188974B_ABST
    Figure CN116188974B_ABST
Patent Text Reader

Abstract

The present invention discloses a building roof recognition method, device, electronic device, and storage medium. The building roof recognition method includes: acquiring a remote sensing image to be recognized; performing building segmentation on the remote sensing image to be recognized based on a preset building segmentation model; and recognizing the building roof based on the segmented remote sensing image to be recognized. By acquiring a remote sensing image to be recognized, performing building segmentation on the remote sensing image to be recognized based on a preset building segmentation model, and recognizing the building roof based on the segmented remote sensing image to be recognized, rapid building roof recognition is achieved, reducing the cost of manual building roof recognition, improving the accuracy of building roof recognition, and enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technology, and in particular to a building roof recognition method, device, electronic equipment and storage medium. Background Art

[0002] Renewable energy is a crucial component of sustainable development. Solar energy, a widely recognized renewable energy source, significantly reduces environmental pollution. Solar photovoltaic power generation is simple to maintain and is not restricted by geographic location. Rooftop solar photovoltaic power generation is a form of distributed solar photovoltaic power generation. Solar panels are typically installed on building rooftops, eliminating the need for significant land use. Consuming the electricity generated by rooftop photovoltaic panels locally or directly connecting it to a nearby power grid reduces carbon emissions and helps users save on electricity costs. The size of a building's rooftop and the surrounding environment both affect the amount of power generated by rooftop photovoltaic power generation. Therefore, high-precision identification of building rooftop information is crucial for distributed solar photovoltaic power generation.

[0003] Currently, the information of building roofs can be determined through manual identification or through neural network models, but the accuracy is low. Therefore, a convenient and accurate method for identifying building roofs has become an urgent problem to be solved. Summary of the Invention

[0004] The present invention provides a building roof recognition method, device, electronic device and storage medium to achieve high-precision recognition of building roofs, reduce the cost of manual recognition of building roofs, and improve user experience.

[0005] According to one aspect of the present invention, a method for identifying a building roof is provided, wherein the method comprises:

[0006] Acquire remote sensing images to be identified;

[0007] Perform building segmentation on the remote sensing image to be identified based on the preset building segmentation model;

[0008] Identify building roofs based on the segmented remote sensing image to be identified.

[0009] According to another aspect of the present invention, a device for identifying a building roof is provided, wherein the device comprises:

[0010] An image acquisition module is used to acquire remote sensing images to be identified;

[0011] A building segmentation module, used to perform building segmentation on the remote sensing image to be identified based on a preset building segmentation model;

[0012] The roof recognition module is used to identify the roof of the building based on the segmented remote sensing image to be identified.

[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the building roof recognition method according to any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, which are used to enable a processor to implement the building roof recognition method according to any embodiment of the present invention when executed.

[0018] The technical solution of the present invention obtains a remote sensing image to be identified, performs building segmentation on the image based on a preset building segmentation model, and then identifies building roofs based on the segmented image. This enables rapid roof recognition, improves roof recognition accuracy, reduces the cost of manual roof recognition, and enhances the user experience.

[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 This is a flow chart of a building roof recognition method provided according to the first embodiment of the present invention;

[0022] Figure 2 This is a flowchart of a method for training a preset building segmentation model according to a second embodiment of the present invention;

[0023] Figure 3 Flowchart of a method for training a preset building segmentation model according to a third embodiment of the present invention;

[0024] Figure 4is a flow chart of a building roof recognition method provided according to a fourth embodiment of the present invention;

[0025] Figure 5 2 is a schematic diagram of the structure of a preset building segmentation model provided according to the fourth embodiment of the present invention;

[0026] Figure 6 2 is a schematic diagram of the structure of the Transformer module provided in accordance with the fourth embodiment of the present invention;

[0027] Figure 7 2 is a schematic diagram of the structure of the improved attention mechanism provided in accordance with the fourth embodiment of the present invention;

[0028] Figure 8 is a schematic diagram of a cross-shaped window context interaction module provided according to a fourth embodiment of the present invention;

[0029] Figure 9 This is a schematic structural diagram of a building roof recognition device provided according to a third embodiment of the present invention;

[0030] Figure 10 The present invention is a schematic structural diagram of an electronic device for implementing a building roof recognition method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] Example 1

[0034] Figure 1This is a flow chart of a building roof recognition method provided according to the first embodiment of the present invention. This embodiment is applicable to the case of identifying building roofs. The method can be executed by a building roof recognition device. The building roof recognition device can be implemented in the form of hardware and / or software. The building roof recognition device can be configured in an electronic device. Figure 1 As shown, the method includes:

[0035] S110: Acquire a remote sensing image to be identified.

[0036] The remote sensing image to be identified may refer to a high-resolution image waiting for identification, and the remote sensing image to be identified is an image that has not been labeled.

[0037] In an embodiment of the invention, the remote sensing image to be identified can be collected by a high-resolution image sensor, or can be obtained from a high-resolution image data source. For example, the remote sensing image to be identified can be obtained from a map service provider.

[0038] S120 , performing building segmentation on the remote sensing image to be identified based on a preset building segmentation model.

[0039] The preset building segmentation model may refer to a pre-trained building segmentation model that can be used to perform building segmentation on remote sensing images. The preset building segmentation model may be a hybrid framework structure model based on a Transformer network.

[0040] In an embodiment of the present invention, a trained preset building segmentation model can be stored locally on the electronic device and retrieved locally from the electronic device. A remote sensing image to be identified can be input into the preset building segmentation model, and the preset building segmentation model can be used to perform building segmentation on the remote sensing image to be identified. In one embodiment, the preset building segmentation model can generate labels for buildings in the remote sensing image to be identified, with different labels being generated for different building areas.

[0041] S130: Identify the roof of the building according to the segmented remote sensing image to be identified.

[0042] In an embodiment of the invention, after the remote sensing image to be identified is segmented into buildings, the labels corresponding to the buildings in the segmented remote sensing image to be identified can be extracted, and the labels corresponding to the roofs of the buildings can be searched to identify the roofs of the buildings.

[0043] In this embodiment of the present invention, a remote sensing image to be identified is acquired, and then segmented based on a preset building segmentation model. Building roofs are then identified based on the segmented image. This allows for rapid roof recognition, reduces the cost of manual roof recognition, improves roof recognition accuracy, and enhances the user experience.

[0044] Example 2

[0045] Figure 2 FIG. 1 is a flow chart of a method for training a preset building segmentation model according to a second embodiment of the present invention. This embodiment is an explanation of a method for training a preset building segmentation model according to the above embodiment. Figure 2 As shown, the method includes:

[0046] S210. Generate a model dataset based on the high-resolution image, wherein the model dataset includes at least a training set, a validation set, and a test set.

[0047] Among them, high-resolution images can refer to images with high pixel density, which can be collected by high-resolution image sensors, or obtained through high-resolution image data sources. Since high-resolution images have a high pixel density, they are easier to extract features; model datasets can refer to datasets for training models, and the preset building segmentation model can be trained through the model datasets.

[0048] In an embodiment of the invention, after obtaining a high-resolution image, the buildings in the high-resolution image can be labeled, the roof area information of each building in the high-resolution image can be labeled, and the labeled high-resolution image can be used as a model data set. In actual operation, after obtaining a high-resolution image, the high-resolution image can be imported into a deep learning image annotation tool, and the roofs of the buildings in the high-resolution image can be labeled by the deep learning image annotation tool. Exemplarily, the high-resolution image annotation tool can include but is not limited to labelme software and RectLabel software. After the high-resolution image is labeled, each labeled high-resolution image can be used as a model data set, wherein the model data set includes at least a training set, a validation set and a test set. In one embodiment, the data in the data set can be divided into a training set, a validation set and a test set according to a ratio, which can include 8:1:1, 7:2:1, etc. Since a large amount of data is required in the model data set, the existing labeled high-resolution image can be enhanced to expand the data set. In one embodiment, the labeled high-resolution image can be cropped, toned, rotated, etc. to expand the data in the data set.

[0049] S220: Input the training set into a pre-built preset building segmentation model for training, wherein the preset building segmentation model is built based on a Transformer network.

[0050] The Transformer network can be a hybrid framework based on the attention mechanism or an encoder-decoder neural network architecture. A preset building segmentation model can be built using the Transformer network.

[0051] In an embodiment of the invention, a training set can be input into a pre-built preset building segmentation model, and the pre-built preset building segmentation model can be trained based on the training set. The preset building segmentation model can be built based on a Transformer network. In one embodiment, the network loss function of the preset building segmentation model can include, but is not limited to, a cross-entropy loss function. The optimizer can select Adam and set a learning rate. A local optimal solution can be obtained through a backpropagation algorithm and reverse iteration. Based on the local optimal solution, the convolution kernel parameters and bias size and other parameters of the preset building segmentation model can be adjusted to minimize the loss of the preset building segmentation model, thereby obtaining the optimal preset building segmentation model.

[0052] In an embodiment of the present invention, a model dataset is generated based on high-resolution images, and the training set is input into a pre-built preset building segmentation model for training, thereby implementing the training of the preset building segmentation model, and then implementing building segmentation of remote sensing images through the preset building segmentation model, thereby improving the user experience.

[0053] In one embodiment, a hybrid framework model encoder of a preset building segmentation model is composed of four Transformer modules connected in series, wherein the image processing sizes of the four Transformer modules are different, and wherein the Transformer module includes an improved attention mechanism module, a multi-layer perceptron, a convolutional layer, and an activation function.

[0054] Among them, the Transformer module may include an improved attention mechanism module, a multi-layer perceptron, a convolution layer, and an activation function. Among them, the improved attention mechanism module may refer to an attention mechanism module with a dual-branch structure, having a global attention branch and a convolution local branch to capture global and local context for visual perception. The multi-layer perceptron is an artificial neural network with a forward structure, comprising an input layer, an output layer, and multiple hidden layers, and the output of each hidden layer needs to be converted by an activation function. The convolution layer can be used to obtain position information instead of position encoding to solve the problem of decreased accuracy caused by inconsistent image resolution in the test set and the training set. The activation function can be used to change the linear relationship of the previous data so that the feature map is output correctly.

[0055] In an embodiment of the invention, a hybrid framework model encoder of a preset building segmentation model can downsample the features in the training set. The hybrid framework model encoder of the preset building segmentation model is composed of 4 Transformer modules connected in series. The image processing sizes of the 4 Transformer modules can be the same or different. In actual operation, when the image processing size of each Transformer module is double-downsampled, the size of each output feature map is half of the input feature map. The Transformer module includes an improved attention mechanism module, a multi-layer perceptron, a convolutional layer, and an activation function. In one embodiment, after the training set is input into the Transformer module, it can first pass through the improved attention mechanism module, and then pass through the multi-layer perceptron, the convolutional layer, and the activation function in sequence.

[0056] In one embodiment, the improved attention mechanism module includes a global branch and a local branch, wherein the global branch includes a window-based multi-head self-attention module and a cross-window context interaction module, and the local branch includes two parallel convolutional layers.

[0057] Among them, the multi-head self-attention module can be used to capture global context information; the cross-shaped window context interaction module can be composed of two parallel branches, which can be used to capture the relationship between windows and establish the dependency relationship between windows. The cross-shaped window context interaction module fuses the two feature maps generated by the horizontal average pooling layer and the vertical average pooling layer to capture the global context. Among them, the horizontal average pooling layer can establish the horizontal relationship between windows, and the vertical average pooling layer can establish the vertical relationship between windows. Among them, the kernel size of the convolution layer in the local branch can be unrestricted, and can include but is not limited to 1, 3, 4, etc. The kernel sizes of the two parallel convolution layers can be different.

[0058] In an embodiment of the invention, the improved attention mechanism module may include a global branch and a local branch. In the global branch, the multi-head self-attention module may obtain and capture global context information. In actual operation, the training set may be divided into several windows through the window partition module. The multi-head self-attention module may calculate the Q, K, and V matrices of all heads, where Q represents Query, K represents Key, and V represents Value. Then, by changing the shape and exchanging the dimensions, the Q, K, and V of multiple heads are placed in the same batch for the same calculation as the single-head attention. Finally, the attention vectors of multiple heads are spliced ​​together to calculate the attention value. Since there is no overlapping part between each window, the relationship between windows can be determined by the cross-shaped window context interaction module. In the local branch, the training set can be input into two parallel convolutional layers to obtain local context information.

[0059] In one embodiment, the dependency relationship between windows of the cross-shaped window context interaction module is at least one of the following:

[0060]

[0061] P1 (m+i,n) =D i (P1 (m,n) );

[0062]

[0063]

[0064] Among them, P1(m,n) represents any point in window 1, P2(m+w,n) represents any point in window 2, w represents the window size, D represents the attention calculation, and the dependency between windows can be determined through the cross-shaped window context interaction module.

[0065] In one embodiment, the decoder of the preset building segmentation model includes a multi-scale feature fusion module and four upsampling blocks.

[0066] The multi-scale feature fusion module can be used to fuse multi-scale features, concatenating feature maps of the same size to obtain features of different scales. The upsampling block can be used to convert the downsampled image to its original size. The upsampling factor of each upsampling block can be different, and the size of the downsampled image can correspond to upsampling blocks of different sizes.

[0067] In an embodiment of the invention, since the hybrid framework model encoder of the preset building segmentation model is composed of four transformer modules connected in series, when the image processing size of each transformer module is the same, the downsampling size of the feature map output by each transformer module is doubled. Each transformer module can be connected to an upsampling block. In one embodiment, all transformer modules can be upsampled and feature fusion can be performed through a multi-scale feature fusion module; alternatively, the feature maps of the second transformer module, the third transformer module, and the fourth transformer module can be input into the upsampling module and upsampled to feature maps of the processing size of the first transformer module. Feature fusion is performed through the multi-scale feature fusion module to obtain multi-scale context information and a fused feature map. The fused feature map is then fused with the feature map of the first transformer module and input into the upsampling block to restore the original image size. Exemplarily, when the image processing size of each Transformer module is double downsampled, the output feature map size of the first Transformer module is half of the input feature map, the output feature map size of the second Transformer module is one-quarter of the input feature map; the output feature map size of the third Transformer module is one-eighth of the input feature map; and the output feature map size of the fourth Transformer module is one-sixteenth of the input feature map. In this case, the upsampling block connected to the fourth Transformer module can be 8x upsampling, the upsampling block connected to the third Transformer module can be 4x upsampling, and the upsampling block connected to the second Transformer module can be 2x upsampling. The feature maps output by the second Transformer module, the third Transformer module, and the fourth Transformer module are fused to obtain multi-scale context information and a fused feature map. The fused feature map is then fused with the feature map of the first Transformer module and input into the upsampling block to restore the original image size.

[0068] In this embodiment, the improved attention mechanism module in the Transformer module enables the network to obtain local information features while obtaining global context information. By adding a cross-shaped window context interaction module, a dependency relationship between windows is established to capture the relationship between windows and improve computing efficiency.

[0069] Example 3

[0070] Figure 31 is a flowchart of a method for training a preset building segmentation model according to a third embodiment of the present invention. This embodiment further illustrates a method for training a preset building segmentation model based on the above embodiment. Figure 3 As shown, the method includes:

[0071] S310. Generate a model dataset based on the high-resolution image.

[0072] S320: Determine a cross entropy loss function as a network loss function of the preset building segmentation model.

[0073] The cross-entropy loss function is a smooth function. The network loss function can be used to calculate the difference between the forward calculation result of each iteration of the neural network and the true value, thereby guiding the next step of training in the right direction. The network loss function can be used to measure the distance between the predicted information of the neural network model and the expected information (label). The closer the predicted information is to the expected information, the smaller the loss function value.

[0074] In an embodiment of the invention, a network loss function of a preset building segmentation model may be determined, and the network loss function of the preset building segmentation model may include but is not limited to a cross entropy loss function. In one embodiment, the cross entropy loss function may be used as the network loss function of the preset building segmentation model, and the cross entropy loss function may include Among them, y i Represents the label value, y′ i Represents the predicted value.

[0075] S330. Determine a local optimal solution of the preset building segmentation model through a back propagation algorithm based on an Adam optimizer with a set learning rate.

[0076] The Adam optimizer is a stochastic optimizer that can calculate adaptive learning rates for different parameters. In one embodiment, the learning rate of the Adam optimizer can be set by business personnel based on experience. The backpropagation algorithm can be used to train a preset building segmentation model. A local optimal solution can refer to a solution for building segmentation that is optimal within a certain range or region.

[0077] In an embodiment of the invention, an Adam optimizer with a learning rate can be configured to determine a local optimal solution for a preset building segmentation model through reverse iteration using a back-propagation algorithm. In actual operation, the local optimal solution can be determined using an evaluation metric. In one embodiment, Mean IOU can be selected as the evaluation metric, and its expression is as follows: Where k represents the number of categories, p ii Indicates the correct pixel predicted, p ij represents the pixel whose category is j and is predicted to be i, p jiIt represents the pixel whose category is i and predicted to be j. When the evaluation index value reaches the maximum, it can be considered as the local optimal solution of the building segmentation model.

[0078] S340. Adjust the convolution kernel parameters and the offset size of the preset building segmentation model based on the local optimal solution.

[0079] In an embodiment of the invention, after determining the local optimal solution, the convolution kernel parameters and bias size of the preset building segmentation model can be adjusted according to the further optimal solution to minimize the model loss and obtain the optimal preset building segmentation model.

[0080] In an embodiment of the present invention, a model data set is generated through high-resolution images, a cross-entropy loss function is determined as the network loss function of a preset building segmentation model, a local optimal solution of the preset building segmentation model is determined through a back-propagation algorithm based on an Adam optimizer set with a learning rate, and the convolution kernel parameters and bias size of the preset building segmentation model are adjusted based on the local optimal solution to implement the preset building segmentation model training, so that the model loss is minimized, thereby obtaining the optimal preset building segmentation model.

[0081] Example 4

[0082] Figure 4 1 is a flow chart of a building roof recognition method provided according to a fourth embodiment of the present invention. This embodiment further illustrates a building roof recognition method based on the above embodiment. Figure 4 As shown, the method includes:

[0083] S410: Obtain a high-resolution image generation dataset, wherein the model dataset includes at least a training set, a validation set, and a test set.

[0084] In an embodiment, high-resolution images can be obtained from a map service provider, manually annotated as a dataset, and data augmentation can be used to expand the dataset size. The dataset is then divided into a training set, a validation set, and a test set in an 8:1:1 ratio. Data augmentation can include random cropping, color adjustment, rotation, etc., to ensure that the image data in the training set covers all possible situations.

[0085] S420: Input the training set into a preset building segmentation model for training.

[0086] In one embodiment, Figure 5 3 is a schematic structural diagram of a preset building segmentation model provided according to the fourth embodiment of the present invention.

[0087] Among them, the encoder of the preset building segmentation model can be composed of 4 Transformer modules in series, which can be recorded as Transformer block1, Transformer block2, Transformer block3 and Transformerblock4; the decoder consists of 4 upsampling blocks and a multi-scale feature fusion module.

[0088] In one embodiment, Figure 6 2 is a schematic diagram of the structure of the Transformer module provided according to the fourth embodiment of the present invention.

[0089] The Transformer module can be composed of an improved attention mechanism module, a multi-layer perceptron, a convolutional layer, and an activation function. The output feature maps of each Transformer module can be denoted as O1, O2, O3, and O4, and the size of each output feature map is half of the input feature map.

[0090] In one embodiment, Figure 7 2 is a schematic diagram of the structure of the improved attention mechanism provided according to the fourth embodiment of the present invention.

[0091] Among them, the improved attention mechanism module can adopt a dual-branch structure, with a global attention branch and a convolution local branch to capture global and local context for visual perception. The global branch includes a multi-head self-attention module and a cross-window context interaction module; the local branch can include two parallel convolution layers (with kernel sizes of 3 and 1 respectively) to extract local context, and then the two results are normalized. In the global branch, the image can be divided into several small images through a window partition module, and the attention value is calculated. In one embodiment, the training set can be divided into several windows, and the multi-head self-attention module can calculate the Q, K, and V matrices of all heads, calculate the similarity of each pixel, and calculate the attention value. Since there is no overlap between each window, the relationship between windows can be determined by the cross-window context interaction module to obtain the relationship between the windows, the global branch context relationship and the local branch context relationship are spliced ​​in the channel dimension, and the number of channels is adjusted using 1×1 convolution to output the result.

[0092] In one embodiment, Figure 8 2 is a schematic diagram of a cross-shaped window context interaction module provided according to a fourth embodiment of the present invention.

[0093] Among them, the cross-shaped interaction module consists of two parallel branches, upper and lower, which can establish dependencies between windows to capture the relationship between windows. The cross-shaped window context interaction module fuses the two feature maps generated by the horizontal average pooling layer and the vertical average pooling layer to capture the global context. In one embodiment, the horizontal average pooling layer establishes a horizontal relationship between windows, and the vertical average pooling layer establishes a vertical relationship between windows. For any point P1(m,n) in window 1, its dependency relationship with P2(m+w,n) in window 2 can include:

[0094]

[0095] P1 (m+i,n) =D i (P1 (m,n) );

[0096]

[0097]

[0098] Among them, w represents the window size and D represents the attention calculation.

[0099] The multi-scale feature fusion module adopts a four-parallel branch structure, and performs dilated convolution on the O1, O2, O3, and O4 feature maps after unified size and splicing operation at different expansion rates to obtain different scale features of the feature maps.

[0100] In one embodiment, the encoding process involves sequentially feeding an input image into Transformer block 1, Transformer block 2, Transformer block 3, and Transformer block 4 to obtain feature maps O1, O2, O3, and O4 containing global and local context information. Within the Transformer modules, the input first passes through an improved attention mechanism module, then through a multi-layer perceptron, a convolutional layer, and an activation function. The convolutional layer can be used to obtain positional information, replacing positional encoding and addressing the accuracy degradation caused by inconsistent resolution between test and training images.

[0101] In one embodiment, the decoder of the preset building segmentation model includes a multi-scale feature fusion module and four upsampling blocks.

[0102] In one embodiment, the decoding process may include upsampling the feature map output by Transformer block 4 by a factor of 8, upsampling the feature map output by Transformer block 3 by a factor of 4, and then concatenating and adjusting the number of channels with the feature map output by Transformer block 2. This operation fuses high-level language information with low-level local information to produce feature map F1. Feature map F1 is input into a multi-scale feature fusion module to obtain multi-scale context information and produce feature map F2. Feature map F2 is combined with the output of Transformer block 1 to fuse high-level semantics with low-level features. The feature map is then input into an upsampling module and upsampled by a factor of 2 to restore it to its original size.

[0103] The cross entropy loss function is selected as the loss function of the preset building segmentation model, and its expression is as follows:

[0104]

[0105] where y i Represents the label value, y′ i Represents the predicted value.

[0106] Mean IOU is selected as the evaluation indicator, and its expression is as follows:

[0107]

[0108] Where k represents the number of categories, p ii Indicates the correct pixel predicted, p ij represents the pixel whose category is j and is predicted to be i, p ji Represents pixels whose category is i and predicted to be j.

[0109] The optimizer selects Adam, sets the learning rate, and uses the back-propagation algorithm and reverse iteration to obtain the local optimal solution. The convolution kernel parameters and hyperparameters such as the bias size are adjusted to minimize the model loss and obtain the optimal preset building segmentation model.

[0110] S430: Segment the remote sensing image to be identified based on a preset building segmentation model, and identify the roofs of the buildings.

[0111] In an embodiment, a remote sensing image to be identified may be obtained, and buildings may be segmented in an unlabeled remote sensing image using a trained preset building segmentation model to identify building roofs.

[0112] In an embodiment of the present invention, building segments are performed on remote sensing images to be identified based on a preset building segmentation model, and building roofs are identified, thereby achieving rapid identification of building roofs and reducing the cost of manual identification of building roofs. Through an improved attention mechanism module and the included cross-window context interaction module, global context and local context information can be efficiently captured, and computational complexity can be greatly reduced.

[0113] Example 5

[0114] Figure 9 FIG. 1 is a schematic diagram of a structure of a building roof recognition device according to the third embodiment of the present invention. Figure 9 As shown, the device includes: an image acquisition module 51, a building segmentation module 52 and a roof recognition module 53.

[0115] The image acquisition module 51 is used to acquire the remote sensing image to be identified.

[0116] Building segmentation module 52, for performing building segmentation on the remote sensing image to be identified based on a preset building segmentation model.

[0117] The roof recognition module 53 is used to recognize the roof of the building according to the segmented remote sensing image to be recognized.

[0118] In this embodiment of the present invention, the image acquisition module acquires a remote sensing image to be identified. The building segmentation module segments the image based on a preset building segmentation model. The roof recognition module then identifies building roofs based on the segmented image. This enables rapid roof recognition, reduces the cost of manual roof recognition, improves roof recognition accuracy, and enhances the user experience.

[0119] In one embodiment, a building roof recognition device further includes:

[0120] A data set generation module is used to generate a model data set based on high-resolution images, wherein the model data set includes at least a training set, a validation set, and a test set;

[0121] The model training module is used to input the training set into a pre-built preset building segmentation model for training, wherein the preset building segmentation model is built based on the Transformer network.

[0122] In one embodiment, the model training module includes:

[0123] The loss function determining unit is used to determine a cross entropy loss function as a network loss function of a preset building segmentation model.

[0124] The optimal solution determination unit is used to determine the local optimal solution of the preset building segmentation model through a back propagation algorithm based on an Adam optimizer set with a learning rate.

[0125] The model adjustment module is used to adjust the convolution kernel parameters and bias size of the preset building segmentation model based on the local optimal solution.

[0126] In one embodiment, the hybrid framework model encoder of the preset building segmentation model in the building segmentation module 52 is composed of four Transformer modules connected in series, wherein the image processing sizes of the four Transformer modules are different, and the Transformer module includes an improved attention mechanism module, a multi-layer perceptron, a convolutional layer and an activation function.

[0127] In one embodiment, the decoder of the preset building segmentation model in the building segmentation module 52 includes a multi-scale feature fusion module and four upsampling blocks.

[0128] In one embodiment, the global branch and the local branch of the attention mechanism module are improved, wherein the global branch includes a window-based multi-head self-attention module and a cross-window context interaction module, and the local branch includes two parallel convolutional layers.

[0129] In one embodiment, the dependency relationship between windows of the cross-shaped window context interaction module is at least one of the following:

[0130]

[0131] P1 (m+i,n) =D i (P1 (m,n) );

[0132]

[0133]

[0134] Among them, P1(m,n) represents any point in window 1, P2(m+w,n) represents any point in window 2, w represents the window size, and D represents the attention calculation.

[0135] A building roof recognition device provided by an embodiment of the present invention can execute a building roof recognition method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects of the execution method.

[0136] Example 10

[0137] Figure 10Schematic diagram of an electronic device for implementing a building roof recognition method according to an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0138] like Figure 10 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0139] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0140] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a method for identifying building rooftops.

[0141] In some embodiments, a building rooftop identification method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the building rooftop identification method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform a building rooftop identification method in any other suitable manner (e.g., via firmware).

[0142] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0143] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0144] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0145] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0146] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0147] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0148] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0149] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A building roof recognition method, characterized in that: The method comprises: Acquire remote sensing images to be identified; Performing building segmentation on the remote sensing image to be identified based on a preset building segmentation model; Identify the roof of the building according to the segmented remote sensing image to be identified; The hybrid framework model encoder of the preset building segmentation model is composed of four Transformer modules connected in series, wherein the Transformer module includes an improved attention mechanism module, a multi-layer perceptron, a convolutional layer, and an activation function; The improved attention mechanism module includes a global branch and a local branch, wherein the global branch includes a window-based multi-head self-attention module and a cross-window context interaction module, and the local branch includes two parallel convolutional layers.

2. The method according to claim 1, characterized in that The training of the preset building segmentation model includes: Generate a model dataset based on the high-resolution image, wherein the model dataset includes at least a training set, a validation set, and a test set; The training set is input into the pre-built preset building segmentation model for training, wherein the preset building segmentation model is built based on a Transformer network.

3. The method according to claim 2, characterized in that Inputting the training set into the pre-built preset building segmentation model for training includes: Determining a cross entropy loss function as a network loss function of the preset building segmentation model; Determining the local optimal solution of the preset building segmentation model through a back propagation algorithm based on an Adam optimizer set with a learning rate; The convolution kernel parameters and the offset size of the preset building segmentation model are adjusted based on the local optimal solution.

4. The method according to claim 1 or 2, characterized in that The decoder of the preset building segmentation model includes a multi-scale feature fusion module and four upsampling blocks.

5. The method according to claim 1, characterized in that: The dependency relationship between the windows of the cross-shaped window context interaction module is at least one of the following: Among them, P1(m,n) represents any point in window 1, P2(m+w,n) represents any point in window 2, w represents the window size, and D represents the attention calculation.

6. A building roof recognition device, characterized in that: The device comprises: An image acquisition module is used to acquire remote sensing images to be identified; A building segmentation module, configured to perform building segmentation on the remote sensing image to be identified based on a preset building segmentation model; A roof recognition module, configured to recognize a building roof based on the segmented remote sensing image to be recognized; The hybrid framework model encoder of the preset building segmentation model is composed of four Transformer modules connected in series, wherein the Transformer module includes an improved attention mechanism module, a multi-layer perceptron, a convolutional layer, and an activation function; The improved attention mechanism module includes a global branch and a local branch, wherein the global branch includes a window-based multi-head self-attention module and a cross-window context interaction module, and the local branch includes two parallel convolutional layers.

7. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the building roof recognition method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the building roof recognition method according to any one of claims 1 to 5 when executed.

Citation Information

Patent Citations

  • Photovoltaic roof resource identification method based on deep learning image segmentation

    CN111191500A

  • Crack identification method and device, medium and equipment

    CN115187539A