Remote sensing image green roof intelligent identification method and system based on improved Mask RCNN
By improving the Mask RCNN algorithm and combining multiple network modules and loss functions, the problem of green roof identification in complex environments was solved, efficient and accurate green roof detection and feature extraction were achieved, and intelligent management at the city level was supported.
Patent Information
- Application Number
- CN202510859977.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing technologies have difficulty accurately identifying green roofs in complex built environments. Traditional field surveys are inefficient and costly. Semantic segmentation methods based on remote sensing images cannot effectively distinguish green roof instances and cannot extract key morphological parameters.
An improved Mask RCNN algorithm is adopted, combined with a residual network, a feature pyramid network, a dual-channel attention enhancement module and a region proposal network. The target edge segmentation accuracy is improved through a joint loss function, and a morphological parameter analysis module is embedded to achieve accurate detection and feature extraction of green roofs.
It realizes large-scale intelligent identification of green roofs, reduces the cost of manual investigation, improves detection efficiency and accuracy, and can extract geometric characteristic parameters such as the area and perimeter of green roofs to support refined analysis.
Smart Images

Figure CN120808139A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of remote sensing image processing, and particularly relates to a remote sensing image green roof intelligent identification method based on an improved Mask RCNN. BACKGROUND
[0002] As the "fifth facade" of the city, green roofs have become an important carrier for breaking through the constraints of two-dimensional space and constructing a three-dimensional ecological network in high-density built-up areas where land resources are scarce. Green roofs can significantly alleviate the urban heat island effect, regulate rainwater runoff, improve air quality, enhance carbon sink function, and provide new ecological corridors for biodiversity protection. Under the combined influence of climate change and rapid urbanization, green roofs play an irreplaceable role in improving the climate resilience of cities. Therefore, accurately identifying the location, quantity and shape of green roofs is crucial for scientifically quantifying their comprehensive ecological and environmental benefits. However, in complex built environments, green roof detection faces many technical challenges, mainly manifested in the complexity of roof structures, the high similarity of green roofs and ground vegetation features, and the interference of factors such as building shadows, obstructions and uneven lighting. These problems together restrict the accurate identification of green roofs, greatly increasing the difficulty of identification.
[0003] Currently, green roof detection mainly relies on traditional field investigation methods, which involve professional personnel conducting on-site surveys, recording and classification, and collecting relevant parameter data. Although field investigation can obtain high-precision data, it has obvious limitations: low data collection efficiency, high labor and material costs, long update cycle, poor timeliness, and difficulty in achieving wide coverage. In contrast, remote sensing image recognition technology based on computer vision has significant advantages such as wide coverage, fast data update, and lower cost. However, only a small number of studies have attempted to apply it to green roof detection. Existing methods for identifying green roofs based on remote sensing images mainly use semantic segmentation, which uses convolutional neural networks to classify remote sensing images at the pixel level, thereby separating green roofs from the background. This method still faces many challenges: first, semantic segmentation cannot distinguish between different green roof instances within the same category, and the recognition accuracy is insufficient for green roofs that are close to each other or partially overlapping; second, the complex structure of building roofs, combined with factors such as shadows and reflections, results in high false detection and missed detection rates for green roofs; finally, this method can only achieve pixel-level classification and cannot effectively extract key shape parameters such as the shape and area of green roofs, limiting the ability of fine-grained analysis. SUMMARY
[0004] In view of the problems in the prior art, the application provides a remote sensing image green roof intelligent identification method based on an improved Mask RCNN algorithm, the feature extraction capability of the algorithm is enhanced by introducing an attention mechanism, the edge segmentation precision of the target is improved by constructing a joint loss function, and the accurate detection of the green roof in a complex built environment is realized. The morphological parameter analysis module is embedded in the mask post-processing process, the geometric feature parameters such as the area, perimeter and shape index of the green roof can be automatically extracted, and the integration of green roof identification and feature parameter extraction is realized. Compared with the traditional method, the application can effectively solve the problems of difficult target instance differentiation and poor environmental adaptability, and provides a reliable technical means for monitoring and ecological environment benefit evaluation of the green roof.
[0005] The technical scheme adopted by the application is as follows: The technical scheme adopted by the application is as follows: Step 1, obtaining a high-resolution remote sensing image data set; Step 2, building an improved Mask RCNN green roof detection model, including integrating a residual network, a feature pyramid network, a double-channel attention enhancement module, a region proposal network and a region of interest alignment module; The high-resolution remote sensing image is first subjected to multi-level feature extraction by the ResNet-101 residual network, and the multi-level features extracted are fused by the feature pyramid network to obtain a multi-scale feature map, then the double-channel attention enhancement module is used to dynamically calibrate the spectral and spatial response related to the green roof in the feature map to obtain an enhanced multi-scale feature map; the region proposal network extracts candidate regions that may contain the green roof from the enhanced multi-scale feature map, then the candidate regions and the feature map generated by the ResNet-101 residual network are input into the region of interest alignment module to realize feature alignment and generate standardized feature representation, and finally, through multi-task collaborative processing, target classification, bounding box regression and instance segmentation mask generation are completed; Step 3, designing a loss function to train and evaluate the improved Mask RCNN green roof detection model; Step 4, using the optimal model trained to realize green roof detection.
[0006] Further, the specific implementation process of step 1 includes: Step 11, obtaining a high-resolution remote sensing image, enhancing the vegetation spectrum and texture features of the green roof in the image through multi-spectral band fusion, vegetation index enhancement, spatial domain enhancement and noise suppression image processing technology; Step 12, the processed high-resolution remote sensing image is divided into blocks using the sliding window method, and the RGB three channels of each block are normalized. The upper left pixel coordinates and corresponding geographic coordinates of each block are recorded while sliding, and a metadata file is generated; Step 13, filter out the effective samples containing clear green roofs from the block images generated in step 12, and use the image labeling tool LabelMe to label the filtered samples. The labeling results are saved in JSON file format, which contains image path, resolution, polygon vertex coordinates and class label; Step 14, convert the generated JSON file into MS COCO2017 dataset format JSON file. The conversion process automatically generates an annotation index file containing image ID, class ID, segmentation mask coordinates and bounding box information for the improved Mask RCNN green roof detection model to read; Step 15, randomly divide all samples into training set, validation set and test set according to a certain proportion. The training set is used for model training, the validation set is used for parameter optimization and performance adjustment, and the test set is used to evaluate the generalization performance of the model.
[0007] Further, the specific process of the dual-channel attention enhancement module is as follows: (A1) The multi-scale feature map output by the feature pyramid network is denoted as Global spatial average pooling and maximum pooling are performed on each channel respectively to generate two channel description vectors and The former represents the global response strength of the channel, and the latter captures the local significant features within the channel; (A2) Input the above two channel description vectors into the multi-layer perceptron (MLP) with shared parameters, and perform channel compression and expansion operations respectively. Through nonlinear transformation, the dependence between channels is established. After adding the results of the two paths and performing Sigmoid normalization, the channel attention weight is generated, which identifies the importance of the channel related to vegetation. The formula is:
[0008] (A3) Multiply the weight with the input feature map channel by channel to suppress the channel response unrelated to the green roof and strengthen the semantic information related to the target, and obtain the channel enhanced feature map . The formula is:
[0009] (A4) Feature map average pooling and max pooling along the channel dimension to generate two single-channel feature maps and represent the average response and the local maximum response at each position, respectively; (A5) The two single-channel feature maps are concatenated along the channel direction, and a 7x7 convolution kernel is used to capture the local spatial context information. Finally, the Sigmoid function is used for normalization to generate the spatial attention weight , which is used to quantify the response strength of each position in the feature map to the green roof target region. The formula is:
[0010] (A6) The spatial attention weight is multiplied element-wise with in the spatial dimension to output the final two-channel enhanced feature map . The formula is:
[0011] In each formula, is the Sigmoid function, represents element-wise multiplication, represents feature extraction using a 7x7 convolution kernel.
[0012] Further, the region proposal network receives the P2, P3, P4, and P5 feature maps of four different scales output by the feature pyramid network and enhanced by the two-channel attention enhancement module as input. Through the sliding window mechanism, it generates a variety of pre-defined anchor boxes of different sizes and aspect ratios on the feature map to cover the various morphological features that the green roof may present. For each anchor box, the region proposal network performs two key tasks: foreground / background classification and bounding box coordinate regression. For foreground / background classification, first, a 3x3 convolution layer is used to extract spatial features, and then a 1x1 convolution layer is used to output the probability value of each anchor box being foreground or background. The bounding box regression task regresses the position and size increment of the anchor box to make it more consistent with the green roof boundary, including the following steps: (B1) The region proposal network uses a 3x3 convolution layer to extract features and generate multiple pre-set anchor boxes at each position , where is the anchor box center coordinate, is the anchor box width and height; (B2) For each anchor box, calculate the geometric difference between its matching real box coordinates , i.e., the real offset . The formula is:
[0013] (B3) The region proposal network predicts the four-dimensional offset of the anchor box through a 1x1 convolutional layer The model-predicted offset is compared with the true offset Through a loss function The prediction error is quantified, and the formula is:
[0014] (B4) The loss function is calculated using the backpropagation algorithm to respond to the model parameters, and the parameter update strategy of the optimizer drives the predicted offset to gradually approach the true offset in a gradient descent manner ; (B5) The predicted offset is used to adjust the anchor box parameters to obtain the refined predicted box coordinates , and the formula is:
[0015] In the above formulas, is the number of anchor boxes participating in regression; is the th anchor box sample; is a smooth L1 function used to segment the error; respectively represent the horizontal coordinate of the center point, the vertical coordinate of the center point, the width of the anchor box, and the height of the anchor box; is the offset of the true anchor box; is the true offset of the th anchor box in dimension , and is the predicted offset of the th anchor box in dimension After classification and regression, the region proposal network sorts the anchor boxes according to the foreground probability value, selects high-score anchor boxes as initial candidate regions, and uses a non-maximum suppression algorithm for post-processing. The candidate regions are sorted from high to low probability, and high-probability regions are retained and regions highly overlapping with them are removed, effectively removing redundant candidate boxes.
[0016] Further, the region of interest alignment module receives the ROI candidate boxes generated by the region proposal network and the feature maps extracted by the backbone network ResNet-101 as input, and according to the downsampling rate N of the backbone network, maps the ROI boundary box coordinates in the original image space to the feature map space, and scales N times to convert to feature map space coordinates ); the region of interest alignment module normalizes each mapped ROI region to generate a fixed scale feature map; The region of interest alignment module is processed differently for different task requirements: for the target detection task, each candidate box region is divided into 7x7 equal size grid units; for the mask prediction task, a 14x14 fine-grained grid division is used to retain detailed features, 4 uniformly distributed sampling points are set in each grid unit, based on the relative distance of the sampling points and the adjacent integer coordinates, the weighted average value is calculated by the bilinear interpolation algorithm as the feature value of the sampling point, and finally the feature values of all sampling points in the grid are aggregated by mean to output 7x7x256 and 14x14x256 normalized feature maps.
[0017] Further, the 7x7x256 feature map output by the region of interest alignment module is compressed into a 1024-dimensional feature vector by a fully connected layer, which encodes the global semantic information of the target region; the feature vector is then split into two parallel processing branches: the classification branch and the bounding box regression branch; the classification branch outputs the confidence score of the target class through the fully connected layer, which is normalized by the Softmax function to filter high probability targets, thereby determining whether the current region belongs to the green roof; the bounding box regression branch outputs the four-dimensional geometric correction amount of the anchor box ), which modifies the candidate box generated by the region proposal network to improve the target positioning accuracy; At the same time, the higher resolution feature map 14x14x256 output by the region of interest alignment module is used for the mask generation branch, which first enhances the local features of the 14x14x256 feature map through four consecutive 3x3 convolution operations, each layer of convolution uses 256 3x3x256 convolution kernels, without changing the spatial size and channel number of the input, thereby gradually strengthening the texture details of the vegetation coverage area and the spectral mutation features of the building boundary while maintaining the original spatial dimensions; then, a single layer 2x2 transpose convolution is used to upsample the feature map, expanding the feature map size from 14x14 to 28x28 to restore the spatial details of the target outline; finally, a 1x1x256 convolution kernel is used to compress the channel number of the feature map from 256 to 1, generating a single-channel 28x28 probability mask for instance-level segmentation prediction.
[0018] Further, in step 3, the loss function is designed as a multi-task weighted loss, including target classification loss, bounding box regression loss and instance segmentation mask loss, wherein the instance segmentation mask loss function is composed of a joint loss of binary cross-entropy loss and gradient constraint loss The specific steps are as follows: (31) Sobel operator is used to calculate horizontal and vertical gradients respectively, which are used to detect the intensity mutation of the mask in horizontal direction and the spectral transition in vertical direction, and the specific definitions are as follows:
[0019] (32) Gradient operator is used to convolve the predicted mask and the real mask respectively, and the gradient responses in horizontal and vertical directions are extracted, and then the responses in two directions are fused to calculate the edge gradient amplitude, and the edge gradient amplitude maps of the predicted mask and the real mask are obtained respectively to improve the edge sharpening degree of the green roof target area, and the formula is as follows:
[0020]
[0021] (33) The traditional BCE loss function is calculated, and the formula is as follows:
[0022] (34) Gradient constraint loss is defined, L1 norm is used to measure the difference between the gradient amplitude of the predicted mask and the real mask, and the model is forced to learn the detailed information of the edge structure, and the formula is as follows:
[0023] (35) Gradient loss and original BCE loss are combined to construct joint loss function, and the formula is as follows:
[0024] In the above formulas, N represents the total number of pixels; N represents the pixel position of the i-th pixel;
[0025] Further, an incremental loss weighting strategy is introduced: in the early stage of training, the weight of gradient loss term is increased to strengthen the learning of edge details; with the gradual convergence of training, the weight is gradually adjusted to a fixed value to balance the stability optimization and convergence effect.
[0026] Further, the outputted sub-block prediction masks are further spliced into a complete remote sensing image mask: firstly, based on a metadata file recording the upper-left corner geographic coordinates of each sub-block image and its pixel position coordinates in the original image, each sub-mask is accurately mapped into the original image coordinate system; then, for the overlapping area between adjacent sub-blocks, a weighted fusion and edge smoothing strategy is adopted to generate a complete mask image with smooth edge characteristics and good spatial continuity; finally, a binary mask image consistent with the spatial range and resolution of the original remote sensing image is generated, 0 represents the background area, 1 represents the green roof area, and is saved in GeoTIFF format, embedding the corresponding geographic coordinate information. After obtaining the final prediction mask, firstly, morphological closing operation is performed on it to fill small holes and remove isolated noise, so as to optimize the integrity and continuity of the mask boundary; then, the processed mask is subjected to contour extraction operation, and the green roof area is vectorized; all vector objects are stored as a vector file in GeoJSON or Shapefile format, and area, centroid coordinates and model prediction confidence information are attached to each object.
[0027] The application also provides a remote sensing image green roof intelligent identification system based on an improved Mask RCNN, which comprises a memory and a processor in communication connection with the memory, the memory stores computer program instructions, and the processor executes the computer program instructions to realize the remote sensing image green roof intelligent identification method based on the improved Mask RCNN.
[0028] Compared with the prior art, the technical scheme has the following advantages: 1. Compared with the traditional field investigation method, the application realizes wide-range intelligent identification and dynamic detection of green roofs, can effectively reduce the time cost and economic investment of artificial investigation, has the technical advantages of rapidness, high efficiency and reusability, and can meet the application requirements of city-level green infrastructure intelligent monitoring and management.
[0029] 2. For the technical problem of green roof detection in complex scenes, an improved Mask RCNN architecture-based green roof high-precision intelligent identification algorithm is proposed, which effectively solves the problem that the traditional semantic segmentation cannot effectively distinguish green roof instances, and the specific technical advantages are as follows: (1) A series module composed of channel attention mechanism and spatial attention mechanism is embedded between FPN and RPN, which enhances the spectral feature response and spatial structure information of the green roof area, suppresses complex background interference, and improves the generation quality of the candidate area.
[0030] (2) To solve the problem of low segmentation accuracy caused by the fuzzy boundary of green roofs, a joint loss function is proposed, which combines gradient constraint loss and binary cross-entropy loss, and introduces a dynamic weight adjustment strategy to adaptively balance the contributions of the two, significantly improving the accuracy and robustness of the segmentation results.
[0031] (3) By fusing geographic registration information and mask post-processing procedures, the prediction results are converted into standardized GeoTIFF mask data and vectorized GeoJSON or Shapefile files, with additional attribute information such as roof area, centroid coordinates, and model confidence, enabling direct loading on mainstream GIS platforms and improving the application value of the results. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 Flowchart of the backbone network ResNet-101 and FPN algorithm; Figure 2 RPN bounding box regression diagram; Figure 3 Flowchart of the improved Mask RCNN algorithm, with the bolded boxes representing the improved DAM and mask modules; Figure 4 Flowchart of the green roof instance segmentation. DETAILED DESCRIPTION
[0033] The technical solutions of the present application will be further described in detail below in conjunction with the drawings and examples.
[0034] The present application aims to provide an intelligent identification method for green roofs in remote sensing images based on an improved target segmentation algorithm, Mask Region Convolutional Neural Network (Mask RCNN). This method can accurately detect and segment green roofs in complex built environments and extract their complete contours, while supporting green roof area measurement and spatial distribution feature analysis. Compared with traditional survey methods, this method has the advantages of low cost, high efficiency, and good accuracy, which can greatly improve the timeliness and accuracy of urban green infrastructure monitoring and provide reliable technical support for intelligent management and ecological benefit evaluation of green roof resources.
[0035] As shown in Figure 4 , the present application provides an intelligent identification method for green roofs in remote sensing images based on an improved Mask RCNN, with the following specific steps: S1: For high spatial resolution remote sensing images, the vegetation spectrum and texture features of green roofs in the images are enhanced through image processing techniques such as multispectral band fusion, vegetation index enhancement, spatial domain enhancement, and noise suppression.
[0036] S2: The pre-processed high-resolution remote sensing image is divided into blocks using the sliding window method, with a window size of 1024x1024 pixels. Starting from the top left corner of the image, the image is cropped row by row and column by column, with an overlap rate of 20% to ensure the integrity of the building edges and avoid cutting the main body of the green roof. The block images are named roof_0001, roof_0002, and so on, with the total number of blocks being NNNN. Then, the RGB three channels of each block are normalized, with the pixel value being linearly mapped from 0-255 to 0-1, and saved as a PNG format. While sliding, the top left pixel coordinates and corresponding geographic coordinates of each block are recorded, and a metadata file is generated to provide positioning basis for subsequent spatial stitching.
[0037] S3: From the block images generated in S2, the effective samples containing clear green roofs are selected (more than 2500 samples), and the invalid data such as background with too large proportion, serious shadow coverage interference, and suspected targets mixed with ground vegetation spectrum are removed. Using the image labeling tool LabelMe, the selected samples are labeled: along the outer contour of the green roof, draw a point by point, form a closed polygon by connecting multiple straight lines, ensure that the contour line strictly adheres to the target edge, and uniformly label the target class as green_roof. The labeling results are saved in JSON file format, which contains image path, resolution, polygon vertex coordinates and class label.
[0038] S4: Using the labelme2coco conversion tool provided by LabelMe, the JSON file generated by labeling is converted into MSCOCO2017 (Microsoft Common Objects in Context) dataset format JSON file, and the conversion process automatically generates annotation index file (annotations.json), which contains image ID, class ID, segmentation mask coordinates and bounding box information, for improved Mask RCNN model reading.
[0039] S5: According to the proportion of 70%:15%:15%, all samples are randomly divided into training set, validation set and test set, where the training set is used for model training, the validation set is used for parameter optimization and performance adjustment, and the test set is used for evaluating the generalization performance of the model.
[0040] S6: Load the Mask RCNN model parameters (i.e. pre-trained weights) pre-trained on the Common Objects in Context (COCO) dataset, and build an improved Mask RCNN green roof detection model by integrating a Residual Network (ResNet-101), a Feature Pyramid Networks (FPN), a dual-channel attention enhancement module, a Region Proposal Network (RPN), and a Region of Interest Alignment module (ROIAlign).
[0041] S6-1: Use a 101-layer dynamic convolution residual network (ResNet-101) as the backbone network architecture of the FPN to serve as the feature extraction part, in order to achieve strong semantic and strong resolution feature extraction of green roofs in complex built environments.
[0042] Specifically, after the block image is input into the Mask RCNN model, it is first subjected to feature extraction by the ResNet-101 residual network. As the backbone network, ResNet-101 extracts multi-level features from shallow to deep through a series of convolutional layers, batch normalization layers, and activation functions. This process forms four feature maps of different dimensions and resolutions: C2 (256x256x256), C3 (128x128x512), C4 (64x64x1024), and C5 (32x32x2048) from shallow to deep. The shallow feature map has a high resolution and contains spatial detail information such as texture, edge, and color, which is used for precise positioning of the green roof boundary; the deep feature map has a low resolution and contains high-level semantic information such as shape, structure, and context relationship, which is used for overall shape recognition of the green roof.
[0043] The multi-level features obtained by the feature extraction network are input into the FPN, and the FPN fully utilizes the features extracted by each layer of the ResNet-101, realizes the fusion of multi-scale features through the top-down information flow and the horizontal connection mechanism. First, in the order from the deep layer to the shallow layer, the deep layer feature map is up-sampled to adjust the spatial size to be consistent with the size of the adjacent shallow layer feature map. For example, the C5 feature map (32x32x2048) is expanded to 64x64x2048 through 2 times nearest neighbor up-sampling to match the spatial dimension of the C4 feature map. At the same time, the channel dimension of each shallow layer feature map is reduced through 1x1 convolution to maintain the semantic information while reducing the calculation complexity. For example, the C4 feature map (64x64x1024) is processed through a 1x1 convolution layer, and the channel number is reduced from 1024 to 256 to generate a reduced feature map (64x64x256). Subsequently, the up-sampled deep layer feature map and the processed shallow layer feature map are added element by element through horizontal connection to pass the deep layer semantic information to the shallow layer. This process finally generates multi-scale feature maps P2 (256x256x256), P3 (128x128x256), P4 (64x64x256) and P5 (32x32x256) with unified dimensions, and the feature maps of different scales correspond to green roof targets of different sizes, so that the network can detect green roofs of various sizes in complex built environments.
[0044] S6-2: To enhance the expression ability of the green roof area and suppress the interference factors such as building shadows and ground vegetation in the complex scene, the present application embeds a dual-channel attention enhancement module (Dual Attention Module, DAM) before the multi-scale feature maps (P2-P5) output by the FPN enter the RPN. The module dynamically calibrates the spectral and spatial response of the feature map related to the green roof through the series mechanism of channel attention and spatial attention, and the specific process is as follows: (1) The multi-scale feature maps output by the FPN are denoted as , and the global spatial average pooling and the maximum pooling are performed on each channel respectively to generate two channel description vectors and , wherein the former represents the global response strength of the channel, and the latter captures the local significant features in the channel.
[0045] (2) The above two channel description vectors are input into a multi-layer perceptron (Multi-Layer Perceptron, MLP) with shared parameters to perform channel compression and expansion operations, and the dependence between channels is established through nonlinear transformation. After adding the two results, the channel attention weight which identifies the importance of the vegetation related channel (for example, the high NDVI value channel weight is close to 1, and the bare roof channel weight tends to 0) is generated through Sigmoid normalization , where is the input feature map, is the weight of the i-th channel, and is the output feature map.
[0046] (3) The weight is multiplied with the input feature map channel by channel to suppress the channel response irrelevant to the green roof and strengthen the semantic information related to the target, to obtain the channel-enhanced feature map , where is the input feature map, is the weight of the i-th channel, and is the output feature map.
[0047] (4) For each spatial position in the feature map , average pooling and maximum pooling are performed along the channel dimension to generate two single-channel feature maps and , representing the average response and local maximum response of each position, respectively.
[0048] (5) The two single-channel feature maps are concatenated along the channel direction, and then a 7×7 convolution kernel is used to capture the local spatial context information (such as edges and textures). Finally, the Sigmoid function is used for normalization to generate the spatial attention weight , which is used to quantify the response strength of each position in the feature map to the green roof target area (high weight area corresponds to the edge or center of the green roof, and low weight area corresponds to the shadow or bare land), where is the input feature map, is the weight of the i-th channel, and is the output feature map.
[0049] (6) The spatial attention weight is multiplied with element by element in the spatial dimension to highlight the geometric structure of the green roof and suppress background noise interference, to output the final dual-channel enhanced feature map , where is the input feature map, is the weight of the i-th channel, and is the output feature map.
[0050] In each formula, is the Sigmoid function, represents element-wise multiplication, represents feature extraction using a 7×7 convolution kernel.
[0051] S6-3: The RPN extracts candidate regions that may contain green roofs from the enhanced multi-scale feature maps (P2~P5) to adapt to the identification needs of green roofs of different sizes and shapes.
[0052] In particular, the RPN receives the feature pyramid network output and takes the DAM-enhanced P2, P3, P4 and P5 four different scale feature maps as input. Through the sliding window mechanism, a variety of sizes and aspect ratios of pre-defined anchors are generated on the feature map to cover a variety of morphological features that green roofs may present. For each anchor, the RPN performs two key tasks: foreground / background binary classification and bounding box coordinate regression. For foreground / background classification, first, spatial features are extracted through a 3x3 convolutional layer, and then a 1x1 convolutional layer is used to output the probability value of each anchor being foreground (containing green roofs) or background (not containing green roofs). The bounding box regression task then regresses the position and size increment of the anchor to make it more consistent with the green roof boundary, including the following steps: (1) The RPN uses a 3x3 convolutional layer to extract features and generate multiple pre-set anchors at each location , where is the anchor center coordinate, is the anchor width and height.
[0053] (2) For each anchor, calculate the geometric difference between its coordinates and the coordinates of the matching real box, i.e., the real offset , the formula is:
[0054] (3) The RPN predicts the four-dimensional offset of the anchor through a 1x1 convolutional layer . Compare the model-predicted offset with the real offset , and use the loss function to quantify the prediction error, the formula is:
[0055] (4) Use the backpropagation algorithm to calculate the response (gradient) of the loss function to the model parameters, and use the parameter update strategy of the optimizer to drive the predicted offset to gradually approach the real offset in a gradient descent manner.
[0056] (5) Use the predicted offset to adjust the anchor parameters to obtain the refined predicted box coordinates , the formula is:
[0057] In the above formulas, is the number of anchors participating in regression; is the th anchor sample; L1 is a smooth function used for piecewise processing of errors; respectively represent the horizontal coordinate of the center point , the vertical coordinate of the center point , the width of the anchor box , the height of the anchor box , and the offset of the anchor box compared to the real anchor box; is the real offset of the th anchor box in dimension , is the predicted offset of the th anchor box in dimension .
[0058] After classification and regression, RPN sorts the anchor boxes according to the foreground probability value and selects high-score anchor boxes as initial candidate regions. To solve the problem of possible large overlap of candidate regions, RPN uses a non-maximum suppression algorithm for post-processing, sorts the candidate regions from high to low according to the probability, retains high-probability regions and eliminates highly overlapping regions, thereby effectively removing redundant candidate boxes. After the above processing, RPN finally outputs high-quality candidate regions containing position information and foreground probability scores.
[0059] S6-4: Input the ROI candidate regions generated by RPN and the feature maps generated by the ResNet-101 backbone network into the ROIAlign module to realize feature alignment and generate standardized feature representations.
[0060] Specifically, the ROIAlign module receives the ROI information generated by RPN and the feature maps extracted by the backbone network ResNet-101 as input. According to the downsampling rate of the backbone network (16 times), the module maps the ROI boundary box coordinates in the original image space to the feature map space, and after scaling by 16 times, converts them to feature map space coordinates . To meet the input size requirements of the detection and segmentation module, ROIAlign performs standardized processing on each mapped ROI candidate region to generate fixed-scale feature maps.
[0061] The ROIAlign is differentially processed according to different task requirements: for the target detection task, each ROI candidate region is divided into 7x7 equal-size grid units; for the mask prediction task, a 14x14 fine-grained grid division is adopted to retain detailed features. Four uniformly distributed sampling points are set in each grid unit. Based on the relative distance between the sampling point and the adjacent integer coordinates, the weighted average value is calculated by the bilinear interpolation algorithm as the feature value of the sampling point. The feature values of the four adjacent integer coordinates around the sampling point are fused by weighting, effectively eliminating the pixel-level misalignment problem caused by twice coordinate quantization in the traditional region of interest pooling module (ROI Pooling). Finally, the feature values of all sampling points in the grid are aggregated by mean, and the standardized feature maps of 7x7x256 (detection) and 14x14x256 (mask) are output respectively. Through this process, the ROIAlign module uniformly converts ROIs of different sizes into standardized feature representations, maintains the spatial correspondence of the features, and eliminates the quantization error in the feature extraction process.
[0062] S6-5: The feature maps output by the ROIAlign are processed in a multi-task cooperative manner to complete target classification, bounding box regression, and instance segmentation mask generation.
[0063] First, the 7x7x256 feature map output by the ROIAlign is compressed into a 1024-dimensional feature vector through a fully connected layer, which encodes the global semantic information of the target region. The feature vector is then split into two parallel processing branches: the classification branch and the bounding box regression branch. The classification branch outputs the confidence score of the target class (green roof / background) through a fully connected layer, which is normalized by the Softmax function to filter high-probability targets, thereby determining whether the current region belongs to the green roof. The bounding box regression branch outputs the four-dimensional geometric correction amount of the anchor box ( ), which is used to modify the candidate box generated by the RPN to improve the target positioning accuracy, as in step S6-3.
[0064] Meanwhile, the higher resolution feature map (14x14x256) output by ROIAlign is used for the mask generation branch. This branch first enhances the local features of the 14x14x256 feature map through four layers of consecutive 3x3 convolution operations. Each layer of convolution uses 256 3x3x256 convolution kernels, without changing the spatial size and the number of channels of the input, so as to gradually strengthen the texture details of the vegetation coverage area and the spectral mutation features of the building boundary while maintaining the original spatial dimensions. Subsequently, a single layer of 2x2 transpose convolution (with a step size of 2) is used to upsample the feature map in space, expanding the feature map size from 14x14 to 28x28, so as to restore the spatial details of the target outline. The main function of the transpose convolution is to increase the spatial size of the feature map, rather than adjusting the channel dimension. Finally, through a 1x1x256 convolution kernel, the number of channels of the feature map is compressed from 256 to 1, generating a single-channel 28x28 probability mask for instance segmentation prediction at the pixel level.
[0065] To improve the edge segmentation accuracy of green roofs, the mask generation task of the present application introduces a gradient constraint loss on the basis of the traditional binary cross entropy (BCE) loss, and constructs a joint loss function. The specific steps are as follows: (1) The Sobel operator is used to calculate the horizontal direction gradient and the vertical direction gradient , which are used to detect the intensity mutation of the mask in the horizontal direction (X axis) and the spectral transition in the vertical direction (Y axis), and are specifically defined as follows:
[0066] (2) The gradient operators and are applied to the predicted mask and the real mask respectively to extract the gradient responses in the horizontal and vertical directions, and then the responses in the two directions are fused to calculate the edge gradient amplitude, and the edge gradient amplitude maps of the predicted mask and the real mask are obtained respectively to improve the edge sharpening degree of the green roof target area, and the formula is:
[0067]
[0068] (3) The traditional BCE loss function is calculated, which is used to supervise the model to learn the global probability distribution of the roof vegetation coverage area and the non-vegetation coverage area, and is helpful for the shape recognition of the main area, and the formula is:
[0069] (4) Define the gradient constraint loss The L1 norm is used to measure the difference between the predicted mask and the real mask gradient amplitude, and the model is forced to learn the detailed information of the edge structure, and the formula is:
[0070] (5) Joint the gradient loss and the original BCE loss to construct a joint loss function The classification accuracy and edge structure retention ability are considered at the same time, and the formula is:
[0071] In the above formulas, N is the total number of pixels; represents the pixel position; and are the weighting coefficients of the classification loss and the edge loss. Experiments show that can effectively balance the classification accuracy and edge sensitivity of the model.
[0072] S7: Set key hyperparameters during training, including weight initialization method, learning rate, batch size, and loss function weight coefficients, and select the optimizer type (Adam optimizer) to improve training stability and convergence speed. Configure data augmentation strategies, including random flipping, rotation, brightness adjustment, etc., to improve the generalization ability of the model.
[0073] The loss function is used to quantify the difference between the model prediction value and the true value, and provides a reference basis for model parameter optimization. The present application designs a multi-task weighted loss strategy, which allocates weights to the losses of classification, regression and segmentation tasks, and the proportion is set to 1:1:1.5. Among them, the loss functions of classification and regression tasks use the existing standard loss functions in Mask RCNN, which are cross-entropy loss and smooth L1 loss, respectively; the loss function of the segmentation task is the joint loss function proposed by the present application, which is composed of binary cross-entropy loss and gradient constraint loss , and the specific form is:
[0074] To improve the expression ability of the model on the edge structure, the present application innovatively introduces a progressive loss weighting strategy: in the early stage of training, the weight of the gradient loss term is increased ( ) to strengthen the learning of edge details; with the gradual convergence of training, it is gradually adjusted to 0.5 to achieve a balance between stable optimization and convergence effect.
[0075] S8: The training set generated in step S5 is input into the green roof detection model constructed in step S6 for training. A phased training strategy is adopted, i.e., first, the backbone network is trained (for 40 rounds), and then the entire network is unfrozen for fine-tuning.
[0076] S9: The verification set generated in step S5 is input into the model under training for evaluation. The mask average precision (mAP) is calculated once after each round of training. An early stopping mechanism is implemented, i.e., when the mAP indicator does not improve for 10 consecutive rounds, the training is terminated, and the performance optimal model weight corresponding to the highest mAP in history is saved.
[0077] S10: The test set is input into the optimal model after training to perform the green roof detection task, so as to evaluate the generalization ability of the model on unseen samples. Specifically, the green roof detection can be reduced to an instance-level binary classification problem. According to the consistency between the model prediction result and the true label (i.e., the class label annotated in S3), the detection result can be divided into the following four categories: True Positive (TP) refers to a positive instance being correctly predicted as positive, i.e., a sample actually being a green roof is correctly identified as a green roof; False Negative (FN) refers to a positive instance being incorrectly predicted as negative, i.e., a sample actually being a green roof is incorrectly identified as background; False Positive (FP) refers to a negative instance being incorrectly predicted as positive, i.e., a sample actually being background is incorrectly identified as a green roof; and True Negative (TN) refers to a negative instance being correctly predicted as negative, i.e., a sample actually being background is correctly identified as background. The following indicators are used for model performance evaluation: The intersection over union (Iou) is the ratio of the intersection and union of the true value and the predicted value two sets, which is used to measure the spatial overlap degree of the predicted mask and the true mask, and the calculation formula is:
[0078] The precision (Precision) represents the proportion of samples actually being a green roof among the samples identified as a green roof, and the formula is:
[0079] The recall (Recall) represents the proportion of samples actually being a green roof that are correctly identified, and the formula is:
[0080] The F1 score (F1 Score) is the harmonic mean of Precision and Recall, which is used to comprehensively evaluate the classification performance of the model, and the formula is:
[0081] S11: Input the high-resolution remote sensing image to be detected into the green roof detection model with the optimal model weight obtained in step S9, complete the target detection and segmentation operations, and obtain the spatial position, category and corresponding mask information of the green roof.
[0082] S12: The present invention further splices the block prediction mask output by step S11 into a complete remote sensing image mask. First, based on the metadata file generated by S2 (recording the geographic coordinates of the upper left corner of each block image and its pixel position coordinates in the original image), each sub-mask is accurately mapped to the original image coordinate system. Then, a weighted fusion and edge smoothing strategy is implemented for the overlapping area (20%) between adjacent blocks to generate a complete mask image with smooth edge features and good spatial continuity to solve the edge discontinuity problem caused by block prediction. Finally, a binary mask map consistent with the spatial range and resolution of the original remote sensing image is generated (0 represents the background area and 1 represents the green roof area), and saved in GeoTIFF format with the corresponding geographic coordinate information embedded to facilitate subsequent spatial analysis and visualization.
[0083] S13: After obtaining the final predicted mask, the present invention first performs a morphological closing operation on it to fill small holes and remove isolated noise, thereby optimizing the integrity and continuity of the mask boundary. A contour extraction operation is then performed on the processed mask to vectorize the green roof area. All vector objects are uniformly stored as vector files in GeoJSON or Shapefile format, and attribute information such as area, centroid coordinates, and model prediction confidence is attached to each object. The resulting vector output can be directly loaded into mainstream GIS platforms for spatial analysis, statistical analysis, and visualization of green roofs.
[0084] The results obtained by the present invention are compared with those of Mask RCNN, and the accuracy ( ), recall rate ( ) and F1 score ( ) are evaluated by three indicators, and the comparison results of the indicators before and after the improvement are shown in Table 1. It can be seen that the improved Mask RCNN model proposed in this invention is significantly better than the original model in all evaluation indicators: Precision is improved by 10.83%, Recall is improved by 11.35%, The results show that the improved Mask RCNN model has higher accuracy, stronger robustness, and better generalization ability in the green roof recognition task in complex built environments, and can more effectively distinguish green roof areas from background areas.
[0085] Table 1 Accuracy comparison of Mask RCNN and improved Mask RCNN models in green roof recognition
[0086] In summary, the improvements of the present application are as follows: 1. Green roof detection and contour extraction process based on improved Mask RCNN algorithm, including remote sensing image preprocessing, model design, post-processing and geographic data generation.
[0087] 2. Green roof feature expression optimization method based on attention mechanism. A series module combining channel attention mechanism and spatial attention mechanism is constructed and embedded between FPN and RPN network to enhance the response ability of target area in feature map.
[0088] 3. Green roof boundary feature extraction method based on joint loss function. A weighted loss function composed of gradient constraint loss and binary cross entropy loss is constructed, combined with dynamic weight adjustment strategy to improve the segmentation accuracy of green roof edge.
[0089] 4. Mask splicing and standardized spatial data generation method based on geographic registration, including edge fusion, vector conversion, attribute information integration and other key technical links.
[0090] On the other hand, the embodiment of the present application also provides a remote sensing image green roof intelligent identification system based on improved Mask RCNN, comprising a memory and a processor in communication connection with the memory, the memory stores computer program instructions, and the processor executes the computer program instructions to realize the remote sensing image green roof intelligent identification method based on improved Mask RCNN.
[0091] The above is only the preferred embodiment of the present application, but the scope of protection of the present application is not limited to this, any modification, equivalent replacement, improvement, etc. made by any person skilled in the art within the technical scope disclosed by the present application shall be included in the protection scope of the present application.
Claims
1. A remote sensing image green roof intelligent recognition method based on improved Mask RCNN, characterized by: The steps include: Step 1: Obtain high-resolution remote sensing image dataset; Step 2: Build an improved Mask RCNN green roof detection model, which includes an integrated residual network, a feature pyramid network, a dual-channel attention enhancement module, a region proposal network, and a region of interest alignment module; The high-resolution remote sensing image is first subjected to multi-level feature extraction using a ResNet-101 residual network. The extracted multi-level features are then fused using a feature pyramid network to generate a multi-scale feature map. A dual-channel attention enhancement module then dynamically calibrates the spectral and spatial responses associated with green roofs in the feature map to generate an enhanced multi-scale feature map. A region proposal network extracts candidate regions that may contain green roofs from the enhanced multi-scale feature map. The candidate regions and the feature map generated by the ResNet-101 residual network are then fed into a region of interest alignment module to achieve feature alignment and generate a standardized feature representation. Finally, through multi-task collaborative processing, object classification, bounding box regression, and instance segmentation mask generation are completed. Step 3: Design a loss function to train and evaluate the improved Mask RCNN green roof detection model; Step 4: Use the trained optimal model to detect green roofs.
2. The method for intelligently identifying green roofs from remote sensing images based on an improved Mask RCNN as claimed in claim 1, wherein: The specific implementation process of step 1 includes: Step 11: Obtain high-resolution remote sensing images and enhance the vegetation spectrum and texture characteristics of the green roof in the images through multispectral band fusion, vegetation index enhancement, spatial domain enhancement, and noise suppression image processing technology; Step 12: Use the sliding window method to divide the processed high-resolution remote sensing image into blocks, perform normalization on the RGB channels of each block, and record the pixel coordinates of the upper left corner of each block and the corresponding geographic coordinates while sliding to generate a metadata file; Step 13: Filter out valid samples containing clear green roofs from the segmented images generated in step 12. Use the image annotation tool LabelMe to annotate the filtered samples. The annotation results are saved in JSON file format, which contains the image path, resolution, polygon vertex coordinates, and category labels. Step 14: Convert the annotated JSON file into a JSON file in the MS COCO2017 dataset format. The conversion process automatically generates an annotation index file containing image ID, category ID, segmentation mask coordinates, and bounding box information for the improved MaskRCNN green roof detection model to read. In step 15, all samples are randomly divided into a training set, a validation set, and a test set according to a certain ratio. The training set is used for model training, the validation set is used for parameter optimization and performance adjustment, and the test set is used to evaluate the generalization performance of the model.
3. The method for intelligently identifying green roofs from remote sensing images based on an improved Mask RCNN according to claim 1, wherein: The specific process of the dual-channel attention enhancement module is as follows: (A1) The multi-scale feature map output by the feature pyramid network is recorded as , perform global spatial average pooling and maximum pooling on each channel, and generate two channel description vectors respectively and , the former characterizes the global response strength of the channel, and the latter captures the local salient features within the channel; (A2) The two channel description vectors are input into the multi-layer perceptron (MLP) with shared parameters, and channel compression and expansion operations are performed respectively. The dependency relationship between channels is established through nonlinear transformation. After adding the two results, the channel attention weights are generated after Sigmoid normalization to identify the importance of vegetation-related channels. , the formula is: (A3) The weight With the input feature map Multiply channel by channel to suppress channel responses irrelevant to the green roof and strengthen the semantic information related to the target to obtain the channel-enhanced feature map , the formula is: (A4) Feature map For each spatial position in , average pooling and maximum pooling are performed along the channel dimension to generate two single-channel feature maps and , respectively representing the average response and local maximum response of each position; (A5) These two single-channel feature maps are spliced in the channel direction, and then a 7×7 convolution kernel is used to capture the local spatial context information. Finally, the sigmoid function is used to normalize the spatial attention weights. , which is used to quantify the response intensity of each location in the feature map to the green roof target area. The formula is: (A6) The spatial attention weight and Multiply element by element in the spatial dimension and output the final dual-channel enhanced feature map , the formula is: Among various is the Sigmoid function, represents element-wise multiplication, Indicates that a 7×7 convolution kernel is used for feature extraction.
4. The method for intelligently identifying green roofs from remote sensing images based on an improved Mask RCNN as claimed in claim 1, wherein: The region proposal network receives as input the feature maps of four different scales (P2, P3, P4, and P5) output by the feature pyramid network and enhanced by the dual-channel attention enhancement module. Through a sliding window mechanism, it generates predefined anchor boxes of various sizes and aspect ratios on the feature maps to cover the various morphological features that may be present on green roofs. For each anchor box, the region proposal network performs two key tasks: foreground / background classification and bounding box coordinate regression. For foreground / background classification, a 3×3 convolutional layer is first used to extract spatial features. Then, a 1×1 convolutional layer is used to output the probability value of each anchor box being foreground or background. The bounding box regression task regresses the position and size increment of the anchor box to make it more closely fit the green roof boundary. The following steps are included: (B1) The region proposal network uses a 3×3 convolutional layer to extract features and generates multiple preset anchor boxes at each location ,in is the center coordinate of the anchor box, is the width and height of the anchor box; (B2) For each anchor box, calculate its coordinates with the matching real box The geometric difference between , the formula is: (B3) The region proposal network predicts the four-dimensional offset of the anchor box through a 1×1 convolutional layer , the offset predicted by the model With the actual offset By contrast, through the loss function Quantify the prediction error, the formula is: (B4) Use the backpropagation algorithm to calculate the response of the loss function to the model parameters, and drive the prediction offset through the optimizer's parameter update strategy Gradually approach the true offset using gradient descent ; (B5) Using the predicted offset Adjust anchor box parameters , get the refined prediction frame coordinates , the formula is: In the above formulas, is the number of anchor boxes involved in regression; For the Anchor box samples; It is a smooth L1 function used to process the error in segments; Represents the horizontal coordinates of the center point , center point vertical coordinate , Anchor box width , Anchor frame height The offset compared to the real anchor box; For the Anchor boxes in dimension The actual offset on For the Anchor boxes in dimension The prediction offset on ; After completing classification and regression, the region proposal network sorts the anchor boxes according to the foreground probability value, selects the high-scoring anchor boxes as the initial candidate regions, and uses the non-maximum suppression algorithm for post-processing to sort the candidate regions from high to low according to the probability, retaining the high-probability regions and eliminating the regions that highly overlap with them, thereby effectively removing redundant candidate boxes.
5. The method for intelligently identifying green roofs from remote sensing images based on an improved Mask RCNN as claimed in claim 1, wherein: The region of interest alignment module receives the ROI candidate box generated by the region proposal network and the feature map extracted by the backbone network ResNet-101 as input, and converts the ROI bounding box coordinates in the original image space ( ) is mapped to the feature map space and converted to feature map space coordinates after scaling N times ( ); The region of interest alignment module normalizes each mapped ROI area to generate a fixed-scale feature map; The region of interest alignment module performs differentiated processing according to different task requirements: for the target detection task, each candidate box area is divided into 7×7 equal-sized grid cells; for the mask prediction task, a 14×14 fine-grained grid division is used to retain detailed features. Four evenly distributed sampling points are set in each grid cell. Based on the relative distance between the sampling point and the adjacent integer coordinates, a weighted average is calculated by the bilinear interpolation algorithm as the eigenvalue of the sampling point. Finally, the eigenvalues of all sampling points in the grid are mean-aggregated, and standardized feature maps of 7×7×256 and 14×14×256 are output, respectively.
6. The method for intelligently identifying green roofs from remote sensing images based on an improved Mask RCNN according to claim 5, wherein: The 7×7×256 feature map output by the region of interest alignment module is compressed into a 1024-dimensional feature vector through a fully connected layer. This feature vector encodes the global semantic information of the target region. This feature vector is then split into two parallel processing branches: the classification branch and the bounding box regression branch; The classification branch outputs the confidence score of the target category through the fully connected layer, and then filters out high-probability targets after normalization by the Softmax function to determine whether the current area belongs to a green roof; The bounding box regression branch outputs the four-dimensional geometric correction of the anchor box ( ), perform secondary correction on the candidate boxes generated by the region proposal network to improve the target positioning accuracy; At the same time, the higher-resolution 14×14×256 feature map output by the region of interest alignment module is used in the mask generation branch. This branch first performs local feature enhancement on the 14×14×256 feature map through four consecutive layers of 3×3 convolution operations. Each convolution layer uses 256 3×3×256 convolution kernels without changing the input spatial size and number of channels. This gradually enhances the texture details of vegetation cover areas and the spectral mutation characteristics of building boundaries while maintaining the original spatial dimensions. Subsequently, a single-layer 2×2 transposed convolution is used to spatially upsample the feature map, expanding the feature map size from 14×14 to 28×28 to restore the spatial details of the target contour; finally, a 1×1×256 convolution kernel is used to compress the number of channels of the feature map from 256 to 1, generating a single-channel 28×28 probability mask for pixel-level instance segmentation prediction.
7. The method for intelligently identifying green roofs from remote sensing images based on an improved Mask RCNN according to claim 1, wherein: In step 3, the designed loss function is a multi-task weighted loss, including target classification loss, bounding box regression loss and instance segmentation mask loss, where the instance segmentation mask loss function is By binary cross entropy loss With gradient constrained loss The specific steps for the joint loss composition are as follows: (31) Sobel operator is used to calculate the horizontal gradient and vertical gradient , used to detect the intensity mutation of the mask in the horizontal direction and the spectral transition in the vertical direction. The specific definition is as follows: (32) respectively predict the mask and the true mask Applying the gradient operator and Perform convolution to extract the gradient response in the horizontal and vertical directions, and then fuse the responses in the two directions to calculate the edge gradient amplitude, and obtain the edge gradient amplitude maps of the predicted mask and the true mask respectively. , in order to improve the edge sharpness of the green roof target area, the formula is: (33) Calculate the traditional BCE loss function , the formula is: (34) Define gradient constrained loss , the L1 norm is used to measure the difference between the predicted mask and the true mask gradient amplitude, forcing the model to learn the details of the edge structure. The formula is: (35) Combine the gradient loss with the original BCE loss to construct a joint loss function , the formula is: In the above formulas, N is the total number of pixels; Indicates the pixel positions; and are the weighting coefficients of classification loss and marginal loss, respectively.
8. The method for intelligently identifying green roofs from remote sensing images based on an improved Mask RCNN as claimed in claim 1, wherein: Introducing a progressive loss weighting strategy: increasing the weight of the gradient loss term in the early stages of training to enhance the learning of edge details; As the training gradually converges, its weights are gradually adjusted to fixed values to achieve a balance between stable optimization and convergence effect.
9. The method for intelligently identifying green roofs from remote sensing images based on an improved Mask RCNN as claimed in claim 1, wherein: The process also involves further stitching the output block prediction masks into a complete remote sensing image mask: first, based on the metadata file, which records the geographic coordinates of the upper left corner of each block image and its pixel position coordinates in the original image, each sub-mask is accurately mapped to the original image coordinate system; then, for the overlapping areas between adjacent blocks, a weighted fusion and edge smoothing strategy is used to generate a complete mask image with smooth edge features and good spatial continuity; finally, a binary mask image is generated that is consistent with the spatial range and resolution of the original remote sensing image, with 0 representing the background area and 1 representing the green roof area, and is saved in GeoTIFF format with the corresponding geographic coordinate information embedded; After obtaining the final predicted mask, a morphological closing operation is first performed on it to fill small holes and remove isolated noise, thereby optimizing the integrity and continuity of the mask boundary; A contour extraction operation is then performed on the processed mask to vectorize the green roof area; All vector objects are uniformly stored as vector files in GeoJSON or Shapefile format, and area, centroid coordinates, and model prediction confidence information are attached to each object.
10. A remote sensing image green roof intelligent recognition system based on improved Mask RCNN, characterized by: The invention comprises a memory and a processor in communication with the memory, wherein the memory stores computer program instructions, and when the processor executes the computer program instructions, the remote sensing image green roof intelligent recognition method based on the improved Mask RCNN according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Method for detecting open-pit mine field in remote sensing image based on deep learning
CN112270280A
Mask RCNN-based image segmentation model training method and particle size detection method
CN113408478A
Mask-RCNN-based multi-target detection method in indoor complex environment
CN115937659A
Efficient high-resolution non-destructive detecting method based on convolutional neural network
GB2610449A
Cited By
Equipment identification and quality grade joint evaluation method for substation inspection image
CN121305453A
Forestry resource protection management method and system and medium
CN121527617A
A method, system, and medium for the protection and management of forestry resources.
CN121527617B
Port facility management and maintenance large model report review intelligent agent construction method and system
CN121581681A
Orthophoto-based roof instance recognition method and system, and related device
CN122368799A