Mine site instance segmentation method

By using the boundary enhancement and optimization module of the mine land occupation instance segmentation model, the problem of segmentation accuracy of complex boundaries of mine land occupation targets is solved, generating high-quality instance masks to support mine monitoring and resource management.

CN120877127BActive Publication Date: 2025-11-28CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511375599.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2025-11-28
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

Existing instance segmentation algorithms struggle to effectively capture the complex boundaries and subtle changes of mining sites, leading to problems such as blurred edges and missing contours in the segmentation results, which affects segmentation accuracy.

Method used

A mine land occupation instance segmentation model is adopted, including a boundary enhancement module and a boundary optimization module. The boundary enhancement module generates an edge attention weight map by extracting edge features and performs context aggregation to generate an initial mask; the boundary optimization module optimizes the boundary blocks by fusing features of different scales through sliding window and multi-resolution branching to generate an accurate instance mask.

Benefits of technology

It improves the accuracy and completeness of mine land occupation target identification, and the generated instance mask more accurately reflects the real shape and boundary of various targets in the mine remote sensing scene, supporting mine monitoring and resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877127B_ABST
    Figure CN120877127B_ABST
Patent Text Reader

Abstract

The application provides a mine land instance segmentation method, and relates to the technical field of image processing. The mine land instance segmentation method comprises the following steps: inputting obtained mine land remote sensing image grid data into a trained mine land instance segmentation model, and outputting a corresponding instance mask, wherein the mine land instance segmentation model comprises a boundary enhancement module and a boundary optimization module; the boundary enhancement module is used for processing a feature map of the mine land remote sensing image grid data to obtain an initial mask of the remote sensing image; the boundary optimization module is used for extracting a boundary block along the boundary of the initial mask according to a preset sliding window, optimizing the boundary block by fusing different scale features through a multi-resolution branch to obtain an optimized boundary block, and splicing the optimized boundary block to generate the instance mask. The application can effectively improve the segmentation precision of the mine land instance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a mine land occupation instance segmentation method. BACKGROUND

[0002] Under the background of global digital transformation and the increasing awareness of ecological protection, with the deep integration of remote sensing technology and artificial intelligence algorithms, quickly and accurately obtaining the spatio-temporal distribution information of mine land occupation has become the core of dynamically monitoring the evolution of mine ecological environment and promoting the sustainable use of mineral resources and the coordinated development of ecological protection. It can quickly lock the mine land occupation target from massive remote sensing images, accurately analyze the target category, spatial position and geometric contour, and significantly improve the fine degree of identification.

[0003] However, compared with the regular geometric targets in natural images (such as rectangular buildings and circular traffic signs), the mine land occupation targets present complex and irregular shapes due to mining activities, topography and other factors. The current mainstream instance segmentation algorithms are mostly optimized for regular geometric targets, which are difficult to effectively capture the subtle changes and complex features of the boundaries of mine land occupation targets, resulting in problems such as blurred edges and missing contours in the segmentation results, which affects the segmentation accuracy of mine land occupation instances. SUMMARY

[0004] The problem solved by the present application is how to improve the segmentation accuracy of mine land occupation instances.

[0005] To solve the above problems, the present application provides a mine land occupation instance segmentation method.

[0006] In a first aspect, the present application provides a mine land occupation instance segmentation method, comprising:

[0007] inputting the obtained mine land occupation remote sensing image grid data into the trained mine land occupation instance segmentation model, and outputting the corresponding instance mask, wherein the mine land occupation instance segmentation model comprises a boundary enhancement module and a boundary optimization module;

[0008] The boundary enhancement module is configured to perform edge feature extraction on a feature map of the mine land occupation remote sensing image grid data to generate an edge attention weight map, perform element point multiplication fusion on the edge attention weight map and the feature map to obtain an attention feature map, perform context aggregation processing on the attention feature map and the edge attention weight map to obtain a context feature, perform point multiplication fusion on the context feature and the edge attention weight map to obtain a spatial context feature, and perform channel splicing on the spatial context feature and the feature map to obtain an initial mask of the remote sensing image.

[0009] The boundary optimization module is configured to extract a boundary block along the boundary of the initial mask according to a preset sliding window, optimize the boundary block by fusing different scale features through a multi-resolution branch to obtain an optimized boundary block, and splice the optimized boundary block to generate the instance mask.

[0010] Optionally, the edge feature extraction on the feature map of the mine site remote sensing image raster data to generate an edge attention weight map comprises:

[0011] The double-path convolution on the feature map comprises a horizontal convolution kernel and a vertical convolution kernel.

[0012] The edge feature map is input into an attention mechanism for processing to obtain the edge attention weight map, wherein the attention mechanism comprises a first layer convolution, a second layer convolution, and a third layer convolution.

[0013] Optionally, the double-path convolution on the feature map to obtain an edge feature map comprises:

[0014] The horizontal gradient value corresponding to each pixel point is obtained according to the horizontal convolution kernel and the pixel point of the feature map, and the vertical gradient value corresponding to each pixel point is obtained according to the vertical convolution kernel and the pixel point of the feature map.

[0015] The horizontal gradient value and the vertical gradient value of the pixel point are combined to obtain the gradient intensity corresponding to the pixel point.

[0016] The remote sensing image is binarized to generate the edge feature map through a comparison result of the gradient intensity of each pixel point and a preset gradient intensity threshold.

[0017] Optionally, the horizontal gradient value satisfies:

[0018] ;

[0019] The vertical gradient value satisfies:

[0020] ;

[0021] The gradient intensity satisfies:

[0022] ;

[0023] wherein G s is the horizontal gradient value, G c is the vertical gradient value, G is the gradient intensity, K s (i, j) is the horizontal convolution kernel, K c(i, j) is the vertical convolution kernel; (x, y) is the position of the current pixel point of the feature map, (i, j) is the offset corresponding to the position of the current pixel point of the feature map, and I(x+i, y+j) is the pixel value of the surrounding pixel point of the current pixel point of the feature map.

[0024] Optionally, the edge attention weight map satisfies:

[0025] ;

[0026] wherein A is the edge attention weight map, Conv 1×1,1 is the first layer convolution, the convolution kernel is 1x1, Conv 1×1,2 is the second layer convolution, the convolution kernel is 1x1, Conv 1×1,3 is the third layer convolution, the convolution kernel is 1x1, E is the edge feature map, and sigma is a Sigmoid activation function, which maps the output to the range of [0, 1].

[0027] Optionally, the spatial context feature satisfies:

[0028] ;

[0029] wherein C(x, y) is the spatial context feature, n is the index of the direction-aware convolution kernel, h is the horizontal direction convolution kernel, v is the vertical direction convolution kernel, d is the diagonal direction convolution kernel, fd is the flipped diagonal direction convolution kernel, is the relative offset coordinate inside the convolution kernel, is the weight parameter of the nth direction-aware convolution kernel at ; is the feature value of the attention feature map at position ; is the weight value of the edge attention weight map at position .

[0030] Optionally, the optimization of the boundary block by fusing different scale features through the multi-resolution branch to obtain an optimized boundary block comprises:

[0031] extracting a low-resolution feature map and a high-resolution feature map from the boundary block respectively;

[0032] performing multi-scale feature fusion on the low-resolution feature map and the high-resolution feature map to obtain a comprehensive feature map corresponding to the boundary block;

[0033] optimizing the boundary block through the comprehensive feature map of the boundary block to obtain an optimized boundary block.

[0034] Optionally, the multi-scale feature fusion on the low-resolution feature map and the high-resolution feature map obtains a comprehensive feature map corresponding to the boundary block, and the comprehensive feature map comprises:

[0035] extracting a low-resolution feature from the low-resolution feature map and a high-resolution feature from the high-resolution feature map;

[0036] fusing the low-resolution feature map with the high-resolution feature through up-sampling to obtain a low-resolution fusion feature map;

[0037] fusing the high-resolution feature map with the low-resolution feature through down-sampling to obtain a high-resolution fusion feature map;

[0038] fusing the low-resolution fusion feature and the high-resolution fusion feature through up-sampling to obtain the comprehensive feature map.

[0039] Optionally, the low-resolution feature map extraction process comprises:

[0040] extracting a sample point feature map corresponding to each sample point center position according to the sample point center position in the boundary block;

[0041] generating the low-resolution feature map of the boundary block according to the sample point feature maps corresponding to all sample point center positions of the boundary block;

[0042] the sample point feature map satisfies:

[0043] ;

[0044] wherein y (P0) is the sample point feature map corresponding to the sample point center position P0, R is a sample point adjacent to the sample point center position, P n is the nth adjacent sample point, △P n is an offset corresponding to the sample point P n , w (P n ) is a weight value on a convolution kernel sampling point during convolution operation, and X is the boundary block.

[0045] Optionally, the training process of the mine land instance segmentation model comprises:

[0046] obtaining a plurality of to-be-processed mine land remote sensing image data, wherein the to-be-processed mine land remote sensing image data comprises to-be-processed remote sensing image raster data and corresponding vector data;

[0047] annotating the corresponding to-be-processed remote sensing image raster data through the vector data to obtain a binary mask of the to-be-processed remote sensing image raster data;

[0048] The boundary extraction of the binary mask through a contour detection algorithm obtains a boundary contour of the mine land target;

[0049] The boundary contour is generated through a convex hull fitting algorithm to generate a candidate box label of the to-be-processed remote sensing image grid data;

[0050] The training set is generated according to the binary mask and the candidate box label corresponding to all the to-be-processed remote sensing image grid data;

[0051] The initial model is trained through the training set to obtain a trained mine land instance segmentation model, wherein the initial model comprises a boundary enhancement module and a boundary optimization module.

[0052] In a second aspect, the present application provides an electronic device comprising a memory and a processor;

[0053] The memory is configured to store a computer program;

[0054] The processor is configured to implement the mine land instance segmentation method according to the first aspect when executing the computer program.

[0055] In a third aspect, the present application provides a computer readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the mine land instance segmentation method according to the first aspect is implemented.

[0056] The beneficial effects of the mine land instance segmentation method of the present application are as follows: the trained mine land instance segmentation model is used to perform instance segmentation on the obtained mine land remote sensing image grid data, so as to obtain an instance mask corresponding to accurate mine remote sensing grid data. The boundary enhancement module of the mine land instance segmentation model generates an edge attention weight map through edge feature extraction, highlights the edge information of the image, and uses context aggregation processing to fuse multi-dimensional information, enhances the semantic association between the edge and the context in the feature map, and generates an initial mask that is more consistent with the target contour. The boundary optimization module of the mine land instance segmentation model is used to solve the problem of inaccurate boundary of the initial mask. The boundary block is extracted by using a sliding window, and different scale features are fused by using a multi-resolution branch. The boundary is further optimized by capturing fine edge details and grasping overall structure trends. The mine land instance segmentation model uses two modules to gradually improve the accuracy and completeness of the mine land target recognition from the whole to the part and from the preliminary sketch to the fine polishing, effectively reduces the segmentation error, and improves the accuracy and completeness of the mine land target recognition. The final output instance mask can more accurately reflect the real shape and boundary of various targets in the mine land remote sensing scene, and provides more reliable data support for mine monitoring and resource management. BRIEF DESCRIPTION OF DRAWINGS

[0057] Figure 1 A flowchart of a mine land example segmentation method according to an embodiment of the present application is shown in FIG. 1.

[0058] Figure 2 A structural diagram of a boundary enhancement module according to an embodiment of the present application is shown in FIG. 2.

[0059] Figure 3 A structural diagram of a boundary optimization module according to an embodiment of the present application is shown in FIG. 3.

[0060] Figure 4 A structural diagram of an electronic device according to an embodiment of the present application is shown in FIG. 4. DETAILED DESCRIPTION

[0061] In order to make the above objectives, features and advantages of the present application more apparent, specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms, and should not be construed as being limited to the embodiments set forth herein, but rather, these embodiments are provided so as to more completely and thoroughly understand the present application. It should be understood that the drawings and embodiments of the present application are for exemplary purposes only, and are not intended to limit the scope of the present application.

[0062] It should be understood that each of the steps described in the method embodiments of the present application can be performed in different orders, and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0063] The term "comprising" and variations thereof as used in the present application are open-ended, that is, "comprising but not limited to"; the term "based on" is "based at least in part on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". Related definitions will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in the present application are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0064] It should be noted that the modification of "one" or "multiple" mentioned in the present application is illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".

[0065] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present application are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0066] In the related art, mine land instance segmentation can quickly and accurately identify mine land targets by using remote sensing images and deep learning technology, not only obtaining the target category, specific location, but also presenting key information such as shape contour, and realizing fine-grained identification of mine land. This technology is crucial for timely grasping the changes of mine ecological environment, and through dynamic monitoring of the spatial and temporal distribution of mine land, it can quickly discover illegal land occupation, illegal mining and other behaviors; at the same time, in promoting the sustainable use of mineral resources, the detailed data provided can assist in optimizing resource exploitation planning and reducing resource waste; in the field of ecological environment protection, accurate segmentation results can help assess the impact of mine development on the surrounding ecology and provide a scientific basis for ecological restoration scheme development.

[0067] However, since conventional segmentation methods are mostly designed for regular geometric shape targets, this approach is difficult to accurately extract target edge details and spatial topological relations when facing mine land targets with complex and variable shapes, and boundary zigzag and broken, lacking adaptive perception ability for irregular target local structure, which may cause edge blur, detail loss, segmentation results and actual boundary misalignment, etc. problems, resulting in the shape features of the mine land target cannot be accurately described, thereby reducing the segmentation accuracy and affecting the accuracy and reliability of subsequent mine land spatial and temporal distribution information extraction.

[0068] To solve the problems in the related art, the present embodiment provides a mine land instance segmentation method.

[0069] As shown in Figure 1 The mine land instance segmentation method provided by the present embodiment comprises:

[0070] S100, input the obtained mine land remote sensing image grid data into the trained mine land instance segmentation model, and output the corresponding instance mask, wherein the mine land instance segmentation model comprises a boundary enhancement module and a boundary optimization module.

[0071] Specifically, the mine land occupation remote sensing image raster data is input into the trained mine land occupation instance segmentation model. The boundary enhancement module in the model first processes the input data. By designing a specific convolution kernel and algorithm, the boundary features of the mine land occupation and the surrounding area are enhanced, the edge details are highlighted, and the model can accurately identify the land occupation contour. Then, the boundary optimization module uses optimization algorithms and post-processing strategies to optimize the preliminary generated boundary features, and finally outputs accurate, continuous and actual instance masks, providing high-quality basic data for subsequent analysis and management of mine land occupation. The mine land occupation remote sensing image raster data is two-dimensional plane data obtained by observing the mine area with sensors carried by remote sensing satellites or aerial vehicles. It divides the mine surface space into regularly arranged grid cells (pixels), each pixel corresponds to a specific geographic location, and records the electromagnetic radiation information (such as the reflection or radiation intensity of visible light, infrared, microwave, etc.) of that location in numerical form. These data have spatial resolution, spectral resolution, and temporal resolution. Spatial resolution determines the actual area size represented by the pixel, which can reflect the clarity of the mine surface details. Spectral resolution reflects the number of spectral bands and wavelength range that the sensor can distinguish, which helps to identify different ground object types (such as ore bodies, vegetation, water, bare land, etc.). Temporal resolution can monitor dynamic processes such as mine exploitation and land occupation range changes through image data from different periods. Instance mask is a two-dimensional matrix or array used for image segmentation tasks. Its function is to accurately label and separate the spatial location of specific target instances in the image. In the context of mine land occupation instance segmentation, the instance mask corresponds to the original mine land occupation remote sensing image raster data one by one, and outlines the contour range of each independent mine land occupation target in the form of a pixel-level mask (usually using different numerical values or colors), so that each mine land occupation instance is uniquely identified in the image. For example, for different types of mine development and land occupation in the remote sensing image, the instance mask can generate independent mask regions for each target, clearly distinguishing adjacent or overlapping land occupation instances. This mask not only visually displays the specific distribution of mine land occupation, but also provides accurate spatial range for subsequent area calculation, boundary analysis, change monitoring, etc. It is a key data achievement to realize fine management and dynamic tracking of mine land occupation.

[0072] S200, the boundary enhancement module, configured to perform edge feature extraction on a feature map of the mine land occupation remote sensing image raster data to generate an edge attention weight map, perform element point multiplication fusion on the edge attention weight map and the feature map to obtain an attention feature map, perform context aggregation processing on the attention feature map and the edge attention weight map to obtain a context feature, perform point multiplication fusion on the context feature and the edge attention weight map to obtain a spatial context feature, and perform channel splicing on the spatial context feature and the feature map to obtain an initial mask of the remote sensing image.

[0073] Specifically, the boundary enhancement module first performs edge feature extraction on a feature map of mine land occupation remote sensing image grid data, accurately captures image edge information, and then generates an edge attention weight map, which highlights the importance of the edge part in the mine remote sensing image. The edge feature extraction method can use the Sobel operator to perform edge feature extraction. Next, the edge attention weight map and the original feature map are element-wise multiplied and fused to enhance the proportion of edge features in the overall feature map, obtaining an attention feature map. Subsequently, the attention feature map and the edge attention weight map are subjected to context aggregation processing, and the context features are obtained by comprehensively considering the correlation between the pixel points and their surrounding environment. The spatial context features obtained by multiplying the context features and the edge attention weight map contain more rich spatial structure information.

[0074] Further, the spatial context features and the original feature map are channel spliced to integrate information of different dimensions together, thereby generating an initial mask of the mine remote sensing image. The spatial context features contain image global structure, pixel semantic correlation and other information obtained by context aggregation processing, which supplements the overall semantics and spatial dependency of the image from a macroscopic level. The original feature map retains the original details, textures and shallow features of the image, and carries rich local information. The channel splicing operation is to stack the two feature maps together along the channel dimension, so that the fused feature has both local details and global semantic information. The integrated multi-dimensional feature data can generate the initial mask of the mine land occupation remote sensing image after subsequent processing (such as convolution, activation, etc.), which provides a basis for subsequent accurate target recognition.

[0075] S300, the boundary optimization module is configured to extract a boundary block along the boundary of the initial mask according to a preset sliding window, optimize the boundary block by fusing different scale features through a multi-resolution branch to obtain an optimized boundary block, and splice the optimized boundary block to generate the instance mask.

[0076] Specifically, the accuracy of the initial mask is further improved by the boundary optimization module to obtain a high-precision instance mask of the mine remote sensing image. First, the sliding window algorithm is used to enhance the boundary of the initial mask output by the boundary enhancement module, and the instance boundary is divided into a plurality of candidate boundary blocks that overlap or connect with each other by traversing with a specific step size and window size. These boundary blocks capture the local features of the boundary and provide initial data for subsequent optimization. However, due to the characteristics of the sliding window, there may be a large amount of redundant, low-quality or repeated parts in the candidate boundary blocks. Therefore, the non-maximum suppression algorithm is used to calculate the score (such as the confidence of containing edge information) of each candidate boundary block, compare the scores of adjacent boundary blocks, suppress the candidate boundary blocks with lower scores, and keep the candidate boundary blocks with higher scores and more representative of the true boundary features as the final boundary blocks. After screening, these high-quality boundary blocks remove redundant information and noise interference, have more accurate boundary features, and are sent as input to the boundary optimization module. The boundary block optimization network in the boundary optimization module can more efficiently perform fine processing on the boundary, and improve the accuracy and integrity of the final instance segmentation.

[0077] Further, since a single scale feature cannot balance details and overall structure, the boundary optimization module adopts a multi-resolution branch approach to extract features from boundary blocks at different scales. Small-scale features focus on the fine details of the edge, and large-scale features grasp the overall shape trend. By fusing these features at different scales, the problem of rough boundary segmentation of the initial mask is improved, ensuring the fit of the edge to the actual target contour, and thus obtaining optimized boundary blocks. Finally, all optimized boundary blocks are spliced in the original position order to generate a complete and accurate instance mask. This mask can more accurately segment each target instance in the mine remote sensing image, and provide a reliable basis for subsequent quantitative analysis and decision support.

[0078] In this embodiment, the obtained mine land remote sensing image raster data is subjected to instance segmentation by using the trained mine land instance segmentation model, so as to obtain an instance mask corresponding to the accurate mine remote sensing raster data. The boundary enhancement module of the mine land instance segmentation model generates an edge attention weight map by edge feature extraction, highlights the edge information of the image, and fuses multi-dimensional information by using context aggregation processing, enhances the semantic association between the edge and the context in the feature map, and thus generates an initial mask more suitable for the target contour. The boundary optimization module of the mine land instance segmentation model is aimed at the possible inaccuracy of the boundary of the initial mask, extracts the boundary block by using a sliding window, fuses different scale features by using a multi-resolution branch, can capture fine edge details and grasp overall structure trends, further optimizes the boundary, and improves the segmentation roughness. The mine land instance segmentation model gradually progresses from the whole to the part, from the preliminary sketch to the fine polishing through the two modules, effectively reduces the segmentation error, improves the accuracy and integrity of the mine land target recognition, and makes the finally output instance mask more accurately reflect the real shape and boundary of various targets in the mine scene, thereby providing more reliable data support for mine monitoring and resource management.

[0079] Optionally, the edge feature extraction on the feature map of the mine land remote sensing image raster data generates an edge attention weight map, and the edge feature extraction includes:

[0080] The feature map is subjected to double-path convolution processing to obtain an edge feature map, wherein the double-path convolution includes a horizontal convolution kernel and a vertical convolution kernel.

[0081] The edge feature map is input into an attention mechanism for processing to obtain the edge attention weight map, wherein the attention mechanism includes a first layer convolution, a second layer convolution, and a third layer convolution.

[0082] Optionally, the edge attention weight map satisfies:

[0083] ;

[0084] wherein A is the edge attention weight map, Conv 1×1,1 is the first layer convolution, the convolution kernel is 1x1, Conv 1×1,2 is the second layer convolution, the convolution kernel is 1x1, Conv 1×1,3 is the third layer convolution, the convolution kernel is 1x1, E is the edge feature map, and sigma is a Sigmoid activation function that maps the output to the range of [0, 1].

[0085] Optionally, the spatial context feature satisfies:

[0086] ;

[0087] wherein C(x, y) is the spatial context feature, n is the index of the direction-aware convolution kernel, h is the horizontal direction convolution kernel, v is the vertical direction convolution kernel, d is the diagonal direction convolution kernel, fd is the flipped diagonal direction convolution kernel, is the relative offset coordinate inside the convolution kernel, is the weight parameter of the nth direction-aware convolution kernel at is the weight parameter of the nth direction-aware convolution kernel at is the feature value of the attention feature map at position is the feature value of the attention feature map at position is the weight value of the edge attention weight map at position is the weight value of the edge attention weight map at position

[0088] Specifically, the feature map is processed by double-path convolution, and convolution operations are performed using horizontal and vertical convolution kernels respectively. The weight distribution of the horizontal convolution kernel usually differs in the horizontal direction (for example, the weights of the middle rows are positive, and the weights of the upper and lower rows are negative, or vice versa), while the weights are consistent (or symmetric) in the vertical direction. When the horizontal convolution kernel is convolved with the vertical edge (i.e., the edge whose pixel value changes in the vertical direction and is continuous in the horizontal direction) in the image, the pixel value at the vertical edge will produce a significant response (such as a high gradient value) due to the weight difference in the horizontal direction, thereby highlighting the vertical edge. Similarly, the weight distribution of the vertical convolution kernel exhibits vertical direction differentiation and horizontal direction consistency (or symmetry), and the vertical convolution kernel is good at capturing horizontal edge information. Through parallel processing of these two direction-sensitive convolution kernels, edge information in different directions of the feature map can be comprehensively extracted, and an edge feature map containing rich edge details can be generated.

[0089] Further, the edge feature map is input into an attention mechanism for optimization. The attention mechanism can adopt a three-layer convolution cascade structure: the first layer of convolution can perform preliminary feature transformation on the edge feature map, compressing the feature dimension while preserving key information; the second layer of convolution can enhance the feature expression ability through a nonlinear activation function (sigmoid), highlighting important edge features; the third layer of convolution maps the features to the 0-1 interval, generating an edge attention weight map. The edge attention weight map assigns a corresponding weight to the edge feature of each position, strengthens the significant edge, and suppresses noise and irrelevant information. Through the combination of double-path convolution and attention mechanism, both multi-directional edge features are extracted, and the attention mechanism is adaptively focused on important edges, providing high-quality input for subsequent edge enhancement and context aggregation.

[0090] In this optional embodiment, dual-path convolution operates on the feature map in parallel with horizontal and vertical convolution kernels, accurately capturing edge information in both vertical and horizontal directions. Compared to single-directional convolution, this multi-directional feature extraction method more comprehensively covers edge details, avoiding information omission. Meanwhile, the attention mechanism composed of three convolutional layers adaptively optimizes the edge feature map generated by dual-path convolution, generating an edge attention weight map. This allows the model to focus on key edges and suppress noise and irrelevant information interference. The two mechanisms work together to not only achieve complete extraction of multi-directional edges but also strengthen important edge features through the attention mechanism, laying the foundation for generating more accurate masks and effectively improving the accuracy and robustness of instance segmentation in mine remote sensing images.

[0091] For example, such as Figure 2 As shown, the feature map is edge-extracted using the Sobel operator to obtain an edge feature map. This edge feature map is then processed through a first, second, and third convolutional layer, followed by a sigmoid activation function, to obtain an edge attention weight map. Each of the three convolutional layers can be a 1×1 convolution. The edge attention weight map is then fused element-wise with the feature map processed by the 3×3 and 1×1 convolutions to obtain an attention feature map. This fused attention feature map is then input into a direction-sensitive context aggregation attention layer. To address the multi-directional characteristics of the target boundary, the context aggregation attention layer uses four different 3×3 convolutions. 3. Convolution operations, namely orientation-aware convolution kernels, including: horizontal convolution kernels to capture vertical features, and convolution kernels as follows: Vertical convolution kernels are used to capture features in the horizontal direction; the convolution kernel is... Diagonal convolutions are used to obtain the oblique features of the target, and the convolution kernel is... Flipping the diagonal convolution kernel enhances its adaptability to mirror-symmetric structures; the convolution kernel is... After context aggregation, context features are obtained. These context features are then multiplied and fused with the edge attention weight map to obtain spatial context features. Finally, the spatial context features processed by 1×1 convolution are concatenated with the initial feature map to obtain the initial mask for the final output. The initial mask obtained through the above processing enhances edge and spatial context information and contains multi-dimensional feature information, such as edge information, spatial context information, and original feature information.

[0092] Optionally, the step of performing dual-path convolution processing on the feature map to obtain the edge feature map includes:

[0093] According to the horizontal convolution kernel and the pixel point of the feature map, a horizontal gradient value corresponding to each pixel point is obtained, and according to the vertical convolution kernel and the pixel point of the feature map, a vertical gradient value corresponding to each pixel point is obtained;

[0094] The horizontal gradient value and the vertical gradient value of the pixel point are combined to obtain a gradient intensity corresponding to the pixel point.

[0095] By comparing the gradient intensity of each pixel point with a preset gradient intensity threshold, the remote sensing image is binarized to generate the edge feature map.

[0096] Optionally, the horizontal gradient value satisfies:

[0097] ;

[0098] The vertical gradient value satisfies:

[0099] ;

[0100] The gradient intensity satisfies:

[0101] ;

[0102] wherein, G s is the horizontal gradient value, G c is the vertical gradient value, G is the gradient intensity, K s (i, j) is the horizontal convolution kernel, K c (i, j) is the vertical convolution kernel; (x, y) is the position of the current pixel point of the feature map, (i, j) is the offset corresponding to the position of the current pixel point of the feature map, and I(x+i, y+j) is the pixel value of the surrounding pixel point of the current pixel point of the feature map.

[0103] Specifically, the horizontal convolution kernel is used to perform convolution operation with the pixel points of the feature map, and the horizontal gradient value of each pixel point is obtained by multiplying the convolution kernel weight and the corresponding pixel value and then summing, which reflects the pixel change intensity of the image in the horizontal direction. Similarly, the vertical convolution kernel is used to perform convolution operation with the pixel points of the feature map to generate the vertical gradient value, which represents the pixel change in the vertical direction. Then, the horizontal gradient value and the vertical gradient value of the same pixel point are combined by vector synthesis (such as square sum and then square root) to obtain the gradient intensity of the pixel point. The greater the gradient intensity, the higher the edge possibility of the position. Finally, a preset gradient intensity threshold is set, and the gradient intensity of each pixel point is compared with the threshold. For example, when the gradient intensity is greater than or equal to the threshold, the pixel point is determined as an edge point and is assigned a value of 1 (or white color), and vice versa, the pixel point is determined as a non-edge point and is assigned a value of 0 (or black color). Through this binary processing, the continuous gradient information is converted into clear edge contour to generate the edge feature map, which provides clear edge position information for the subsequent attention mechanism and context aggregation, and ensures the accurate positioning and highlighting of the edge feature.

[0104] In the optional embodiment, the horizontal and vertical gradient values are calculated by the horizontal and vertical convolution kernels respectively to comprehensively capture the edge information of the image from two key directions, accurately locate the edge position, effectively avoid the edge omission caused by single direction detection, and greatly improve the integrity of the edge detection. The gradient intensity is obtained by combining the horizontal and vertical gradient values, which comprehensively considers the pixel change information in two directions, more accurately measures the edge possibility of each pixel point, enhances the quantitative representation ability of the edge feature, and makes the edge information clearer and more explicit. Through the binary processing by comparing with the preset gradient intensity threshold, the continuous gradient intensity information is converted into a simple and intuitive black and white binary image, the edge region is highlighted, the redundant background and noise interference in the image are effectively removed, the generated edge feature map is simple and the edge feature is prominent, and a high-quality input is provided for the subsequent image processing task, which improves the accuracy and efficiency of the overall processing process.

[0105] Optionally, the optimizing the boundary block by fusing different scale features through the multi-resolution branch to obtain an optimized boundary block comprises:

[0106] extracting a low-resolution feature map and a high-resolution feature map from the boundary block respectively;

[0107] performing multi-scale feature fusion on the low-resolution feature map and the high-resolution feature map to obtain a comprehensive feature map corresponding to the boundary block;

[0108] optimizing the boundary block by using the comprehensive feature map of the boundary block to obtain an optimized boundary block.

[0109] In this optional embodiment, low-resolution feature maps and high-resolution feature maps are extracted for the boundary block respectively. The low-resolution feature maps can capture the overall semantic information and macro-structure of the boundary block, because local details are integrated when the resolution is reduced, highlighting more general features. The high-resolution feature maps retain rich detailed information, such as fine textures of the boundary, subtle shape changes, etc., which are crucial for accurately defining the boundary. Then, through multi-scale feature fusion, the low-resolution feature maps and the high-resolution feature maps are organically combined. During the fusion process, through appropriate algorithms (such as weighted summation, convolution after splicing, etc.), the macro-semantic and micro-detail are complemented to generate a comprehensive feature map corresponding to the boundary block, which has both overall cognition and local details. Finally, the comprehensive feature map is used to optimize the boundary block. Based on the rich and complementary information in the comprehensive feature map, the features of the boundary block are adjusted and improved, such as correcting boundary position deviation, enhancing the recognition of boundary features, etc., so as to obtain an optimized boundary block, providing better basic data for subsequent image segmentation, target recognition, etc.

[0110] Optionally, the multi-scale feature fusion of the low-resolution feature map and the high-resolution feature map obtains a comprehensive feature map corresponding to the boundary block, and the method further comprises:

[0111] extracting low-resolution features from the low-resolution feature map and extracting high-resolution features from the high-resolution feature map;

[0112] up-sampling the low-resolution feature map and fusing the low-resolution feature map with the high-resolution feature to obtain a low-resolution fusion feature map;

[0113] down-sampling the high-resolution feature map and fusing the high-resolution feature map with the low-resolution feature to obtain a high-resolution fusion feature map;

[0114] fusing the low-resolution fusion feature and the high-resolution fusion feature through up-sampling to obtain the comprehensive feature map.

[0115] In this optional embodiment, low-resolution feature maps and high-resolution feature maps are extracted for the boundary block respectively. The low-resolution feature maps can capture the overall semantic information and macro-structure of the boundary block, because local details are integrated when the resolution is reduced, highlighting more general features. The high-resolution feature maps retain rich detailed information, such as fine textures of the boundary, subtle shape changes, etc., which are crucial for accurately defining the boundary. Then, through multi-scale feature fusion, the low-resolution feature maps and the high-resolution feature maps are organically combined. During the fusion process, through appropriate algorithms (such as weighted summation, convolution after splicing, etc.), the macro-semantic and micro-detail are complemented to generate a comprehensive feature map corresponding to the boundary block, which has both overall cognition and local details. Finally, the comprehensive feature map is used to optimize the boundary block. Based on the rich and complementary information in the comprehensive feature map, the features of the boundary block are adjusted and improved, such as correcting boundary position deviation, enhancing the recognition of boundary features, etc., so as to obtain an optimized boundary block, providing better basic data for subsequent image segmentation, target recognition, etc. Figure 1In addition, the low-resolution feature map is fused with the high-resolution feature map to obtain a high-resolution fused feature map, so that semantic information of the low-resolution feature is integrated into details of the high-resolution feature map. The low-resolution fused feature and the high-resolution fused feature are further fused through upsampling to integrate them at the same high-resolution scale, so as to generate a comprehensive feature map which has both macro semantic and micro details, and provides more comprehensive and rich feature information for subsequent optimization of the boundary block and related image processing tasks.

[0116] Optionally, the low-resolution feature map extraction process comprises:

[0117] According to the feature extraction of the sampling point center position in the boundary block, a sampling point feature map corresponding to each sampling point center position is obtained;

[0118] According to the sampling point feature maps corresponding to all sampling point center positions of the boundary block, a low-resolution feature map corresponding to the boundary block is generated;

[0119] The sampling point feature map satisfies:

[0120] ;

[0121] wherein y(P0) is the sampling point feature map corresponding to the sampling point center position P0, R is a sampling point adjacent to the sampling point center position, P n is the nth adjacent sampling point, △P n is an offset corresponding to the sampling point P n , w(P n ) is a weight value on a sampling point of a convolution kernel during convolution operation, and X is the boundary block.

[0122] It should be noted that this part can dynamically adjust the size and position of the convolution kernel through deformable convolution, thereby improving the segmentation capability of the model for irregular geometric deformation targets. The core mechanism is to additionally introduce an offset for each sampling point of the convolution kernel, and the offset is generated by convolution operation on the input boundary block to change the default position of the sampling point. When the offset floating-point coordinate is obtained, the bilinear interpolation algorithm can be used to determine the position of the corresponding sampling pixel point on the boundary block to solve the problem of no actual pixel for the floating-point coordinate. In the training process, the offset vector can be continuously optimized through the back propagation algorithm, so that the sampling position of the convolution kernel can be dynamically adjusted according to the geometric shape of the target to be detected, so as to more accurately capture the irregular boundary and deformation characteristics of the target, effectively enhance the feature extraction capability of the model for complex morphological targets, and improve the accuracy and robustness of the segmentation task.

[0123] In this optional embodiment, feature extraction is performed on the center positions of sampling points within the boundary block. The sampling points can be determined using a fixed regular grid (e.g., a 3×3 grid), with the center position obtained through the grid. The center position of the sampling point is a key location for feature extraction. Using a specific algorithm (e.g., convolution), feature information is extracted from the local region centered on this point. This information encompasses pixel grayscale, texture, and color features surrounding the sampling point, resulting in a feature map corresponding to the center position of each sampling point. Each feature map reflects the local features at its corresponding location. Subsequently, the feature maps corresponding to the center positions of all sampling points within the boundary block are comprehensively processed. Using a specific aggregation method (e.g., average pooling, max pooling), the information from numerous feature maps is integrated and dimensionality reduced, decreasing data dimensionality and resolution while retaining key feature information, thereby generating a low-resolution feature map corresponding to the boundary block. This low-resolution feature map summarizes the comprehensive features of multiple sampling points within the boundary block, presenting the macroscopic features of the boundary block and laying the foundation for subsequent multi-scale feature fusion and boundary block optimization operations.

[0124] For example, such as Figure 3 As shown, the boundary optimization module mainly performs multi-stage feature processing, including stages 1 to 4. Stage 1 is mainly composed of a deformable module, which is based on a deformable convolution structure. Stages 2 to 4 consist of multiple basic blocks. Between stages, interpolation upsampling (solid line) and strided convolution downsampling (dashed line) operations are used to achieve the fusion and transfer of features at different scales, enabling the network to capture multi-scale image features. Within each stage, basic blocks fuse features through an addition operation (⊕) to enhance feature representation.

[0125] Specifically, first, the input is a candidate boundary block, and the important area boundary block is screened from the preliminary candidate boundary block through non-maximum suppression (NMS). Then the boundary block is sequentially subjected to two 3x3 convolution (Conv3) modules for preliminary feature extraction, and then a deformable (Deformneck) module in stage 1 is used to obtain the low-resolution feature map corresponding to the boundary block, and then two branches in stage 2 are used to extract low-resolution features and high-resolution features, respectively. One branch obtains the low-resolution feature map output by stage 1 to extract low-resolution features, and the other branch extracts high-resolution features according to the high-resolution feature map of the boundary block. Then, the low-resolution features and the high-resolution features are fused by upsampling and downsampling, that is, the low-resolution feature map and the high-resolution feature map are fused by upsampling to obtain a low-resolution fusion feature map, and the high-resolution feature map and the low-resolution feature are fused by downsampling to obtain a high-resolution fusion feature map. Similarly, in stage 3, one branch can be added on the basis of stage 2, such as stage 2 for feature extraction and fusion of different resolutions, and stage 4 can be added according to the actual situation, and finally all the fusion feature maps of different resolutions are spliced along the channel by upsampling to obtain a comprehensive feature map.

[0126] Optionally, the training process of the mine land instance segmentation model comprises:

[0127] a plurality of to-be-processed mine land remote sensing image data are acquired, wherein the to-be-processed mine land remote sensing image data comprise to-be-processed remote sensing image raster data and corresponding vector data;

[0128] the to-be-processed remote sensing image raster data corresponding to the vector data are labeled to obtain a binary mask of the to-be-processed remote sensing image raster data;

[0129] the binary mask is subjected to boundary extraction through a contour detection algorithm to obtain a boundary contour of a mine land target;

[0130] all the boundary contours are subjected to a convex hull fitting algorithm to generate a candidate box label of the to-be-processed remote sensing image raster data;

[0131] a training set is generated according to the binary mask and the candidate box label corresponding to all the to-be-processed remote sensing image raster data;

[0132] an initial model is trained through the training set to obtain a trained mine land instance segmentation model, wherein the initial model comprises a boundary enhancement module and a boundary optimization module.

[0133] Specifically, in the process of mine land occupation remote sensing image data processing and model construction, first, data preprocessing is carried out. Using tools such as Geospatial Data Abstraction Library (GDAL), the coordinate system of mine land occupation remote sensing image raster data and mine vector data is registered, and multi-source remote sensing images such as remote sensing image data obtained by different remote sensing satellites and mine vector data are unified to the geocentric coordinate system (WGS84 coordinate system), which can control the registration error within 1 pixel and eliminate geometric distortion; for example, the image is cropped according to the standard size of 800x800 pixels to ensure the integrity of the mine land occupation target; the image containing cloud content is removed to complete data cleaning and ensure data consistency. Then, a mask and a candidate box label are generated through a semi-automatic process. Based on the GDAL library, the mine vector surface data is batch converted into a binary mask, and the topological errors such as holes and overlaps are batch corrected; the contour detection (OpenCV) and convex hull fitting algorithm are used to extract the mask boundary to generate the candidate box. Then, the data is checked for quality, and inconsistent samples are removed. The data can be divided into training set, validation set and test set according to the ratio of 7:2:1. In terms of model construction, an edge-guided attention module is constructed, the edge feature extraction operator is used to extract edge features to generate an attention weight map, and after fusion with the features, a direction-sensitive context aggregation layer is used for secondary attention fusion to strengthen the complex boundary contour feature; a boundary block optimization post-processing module based on deformable convolution is used to further extract fine-grained spatial information on the rough mask to improve the segmentation performance of the model on instances with complex background and severe geometric deformation.

[0134] It should be noted that in the training process of the mine land occupation instance segmentation model, the loss function design is carried out around the multi-task demand. The binary cross-entropy loss function can be used for basic segmentation to solve the pixel-level classification and sample imbalance problem; for complex boundary contours, edge perception loss or structural similarity loss is used to strengthen edge feature learning; if there is a target detection branch, the intersection over union loss is used to optimize the candidate box positioning. Finally, the segmentation, edge, positioning and other multi-task losses are weighted and summed to form the total loss function, which can also be trained with optimizers such as Adaptive Moment Estimation (Adam) or Stochastic Gradient Descent (SGD) and learning rate scheduling strategies, and the generalization ability of the model is improved through data augmentation to realize accurate segmentation and positioning of mine targets.

[0135] For example, Figure 4As shown, the electronic device 400 provided by the embodiment of the present application comprises a memory 410 and a processor 420; the memory 410 is used for storing a computer program; the processor 420 is used for realizing the mine land instance segmentation method as described above when executing the computer program.

[0136] Alternatively, an electronic device 400 comprises a memory 410 and a processor 420 coupled to the memory 410; the memory 410 is configured to store a computer program; the processor 420 is configured to perform the following operations when executing the computer program:

[0137] input the obtained mine land remote sensing image grid data into the trained mine land instance segmentation model, and output a corresponding instance mask, wherein the mine land instance segmentation model comprises a boundary enhancement module and a boundary optimization module;

[0138] The boundary enhancement module is used for performing edge feature extraction on a feature map of the mine land remote sensing image grid data to generate an edge attention weight map, performing element point multiplication fusion on the edge attention weight map and the feature map to obtain an attention feature map, performing context aggregation processing on the attention feature map and the edge attention weight map to obtain a context feature, performing point multiplication fusion on the context feature and the edge attention weight map to obtain a spatial context feature, and performing channel splicing on the spatial context feature and the feature map to obtain an initial mask of the remote sensing image.

[0139] The boundary optimization module is used for extracting a boundary block along a boundary of the initial mask according to a preset sliding window, optimizing the boundary block by fusing different scale features through a multi-resolution branch to obtain an optimized boundary block, and splicing the optimized boundary block to generate the instance mask.

[0140] The computer readable storage medium provided by the embodiment of the present application has a computer program stored thereon, and when the computer program is executed by a processor, the mine land instance segmentation method as described above is realized.

[0141] Alternatively, a non-volatile computer readable storage medium has a computer program stored thereon, and when the computer program is executed by a processor, the processor performs the following operations:

[0142] input the obtained mine land remote sensing image grid data into the trained mine land instance segmentation model, and output a corresponding instance mask, wherein the mine land instance segmentation model comprises a boundary enhancement module and a boundary optimization module;

[0143] The boundary enhancement module is configured to perform edge feature extraction on a feature map of the mine site remote sensing image grid data to generate an edge attention weight map, perform element point multiplication fusion on the edge attention weight map and the feature map to obtain an attention feature map, perform context aggregation processing on the attention feature map and the edge attention weight map to obtain a context feature, perform point multiplication fusion on the context feature and the edge attention weight map to obtain a spatial context feature, and perform channel splicing on the spatial context feature and the feature map to obtain an initial mask of the remote sensing image.

[0144] The boundary optimization module is configured to extract a boundary block along a boundary of the initial mask according to a preset sliding window, optimize the boundary block by fusing different scale features through a multi-resolution branch to obtain an optimized boundary block, and splice the optimized boundary block to generate the instance mask.

[0145] An electronic device 400, which can serve as a server or a client of the present application, will now be described, which is an example of a hardware device that can be applied to aspects of the present application. The electronic device 400 is intended to represent various forms of digital electronic computer devices such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computer devices. The electronic device 400 can also represent various forms of mobile devices such as personal digital assistants, cellular telephones, smart phones, wearable devices, and other like computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0146] The electronic device 400 includes a computing unit that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). Various programs and data required for device operation can also be stored in the RAM. The computing unit, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0147] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware, and the program can be stored in a computer readable storage medium. When the program is executed, the program can include the processes of the above-mentioned embodiment methods. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like. In this application, the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the present application. In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically independently, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0148] Although the present application is disclosed as above, the protection scope of the present application is not limited to this. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application, and these changes and modifications will fall within the protection scope of the present application.

Claims

1. A mine footprint instance segmentation method, characterized by, The method comprises the steps of: inputting the obtained mine remote sensing image raster data into a trained mine land instance segmentation model to output a corresponding instance mask, wherein the mine land instance segmentation model comprises a boundary enhancement module and a boundary optimization module; the boundary enhancement module is used for performing edge feature extraction on a feature map of the mine remote sensing image raster data to generate an edge attention weight map, performing element point multiplication fusion on the edge attention weight map and the feature map to obtain an attention feature map, performing context aggregation processing on the attention feature map and the edge attention weight map to obtain a context feature, performing point multiplication fusion on the context feature and the edge attention weight map to obtain a spatial context feature, and performing channel splicing on the spatial context feature and the feature map to obtain an initial mask of the remote sensing image; the boundary optimization module is used for extracting a boundary block along the boundary of the initial mask according to a preset sliding window, optimizing the boundary block by fusing different scale features through a multi-resolution branch to obtain an optimized boundary block, and splicing the optimized boundary block to generate the instance mask; the spatial context feature satisfies: ; Where C(x, y) is the spatial context feature, n is the index of the orientation-aware convolution kernel, h is the horizontal convolution kernel, v is the vertical convolution kernel, d is the diagonal convolution kernel, and fd is the flipped diagonal convolution kernel. These are the relative offset coordinates inside the convolution kernel. For the nth direction sensing convolution kernel in The weight parameters, For the attention feature map at location eigenvalues ​​at that location The edge attention weight map at position The weight value at that location.

2. The mine footprint instance segmentation method of claim 1, wherein, the edge feature extraction on the feature map of the mine remote sensing image raster data comprises: performing double-path convolution processing on the feature map to obtain an edge feature map, wherein the double-path convolution comprises a horizontal convolution kernel and a vertical convolution kernel; obtaining the edge attention weight map through an attention mechanism, wherein the attention mechanism comprises a first layer convolution, a second layer convolution and a third layer convolution.

3. The mine footprint instance segmentation method of claim 2, wherein, the double-path convolution processing on the feature map to obtain an edge feature map comprises: obtaining a horizontal gradient value corresponding to each pixel point according to the horizontal convolution kernel and the pixel points of the feature map, and obtaining a vertical gradient value corresponding to each pixel point according to the vertical convolution kernel and the pixel points of the feature map; combining the horizontal gradient value and the vertical gradient value of the pixel points to obtain a gradient intensity corresponding to the pixel points; performing binaryzation processing on the remote sensing image through a comparison result of the gradient intensity of each pixel point and a preset gradient intensity threshold to generate the edge feature map.

4. The mine footprint instance segmentation method of claim 3, wherein, the horizontal gradient value satisfies: ; the vertical gradient value satisfies: ; the gradient intensity satisfies: ; wherein G s is the horizontal gradient value, G c is the vertical gradient value, G is the gradient intensity, K s (i, j) is the horizontal convolution kernel, K c (i, j) is the vertical convolution kernel; (x, y) is the position of the current pixel point of the feature map, (i, j) is the offset corresponding to the position of the current pixel point of the feature map, I(x+i, y+j) is the pixel value of the surrounding pixel point of the current pixel point of the feature map.

5. The mine footprint instance segmentation method of claim 2, wherein, the edge attention weight map satisfies: ; where A is the edge attention weight map, Conv 1×1,1 is the first layer convolution, with a 1x1 kernel, 1×1,2 is the second layer convolution, with a 1x1 kernel, 1×1,3 is the third layer convolution, with a 1x1 kernel, E is the edge feature map, and σ is a Sigmoid activation function that maps the output to the range [0, 1].

6. The mine footprint instance segmentation method of claim 1, wherein, the optimization of the boundary block by fusing different scale features through a multi-resolution branch comprises: extracting a low-resolution feature map and a high-resolution feature map from the boundary block respectively; performing multi-scale feature fusion on the low-resolution feature map and the high-resolution feature map to obtain a comprehensive feature map corresponding to the boundary block; optimizing the boundary block through the comprehensive feature map of the boundary block to obtain an optimized boundary block.

7. The mine footprint instance segmentation method of claim 6, wherein, the multi-scale feature fusion on the low-resolution feature map and the high-resolution feature map to obtain a comprehensive feature map corresponding to the boundary block comprises: extracting low-resolution features from the low-resolution feature map and extracting high-resolution features from the high-resolution feature map; The low-resolution feature map and the high-resolution feature are fused through upsampling to obtain a low-resolution fused feature map; The high-resolution feature map and the low-resolution feature are fused through downsampling to obtain a high-resolution fused feature map; The low-resolution fused feature and the high-resolution fused feature are fused through upsampling to obtain the comprehensive feature map.

8. The mine footprint instance segmentation method of claim 6, wherein, The low-resolution feature map extraction process comprises: According to the sampling point center position in the boundary block, a sampling point feature map corresponding to each sampling point center position is obtained through feature extraction; According to the sampling point feature maps corresponding to all sampling point center positions of the boundary block, the low-resolution feature map of the corresponding boundary block is generated; The sampling point feature map satisfies: ; wherein y(P0) is the sampling point feature map corresponding to the sampling point center position P0, R is a sampling point adjacent to the sampling point center position, P n is the nth adjacent sampling point, and n is the sampling point P n corresponding offset, w(P n ) is the weight value on the sampling point of the convolution kernel during convolution operation, and X is the boundary block.

9. The mine footprint instance segmentation method of claim 1, wherein, The training process of the mine land instance segmentation model comprises: Obtaining a plurality of to-be-processed mine remote sensing image data, wherein the to-be-processed mine remote sensing image data comprises to-be-processed remote sensing image raster data and corresponding vector data The to-be-processed remote sensing image raster data is labeled through the vector data to obtain a binary mask of the to-be-processed remote sensing image raster data; The boundary of the mine land target is extracted through the contour detection algorithm to obtain a boundary contour of the mine land target; All the boundary contours are fitted through a convex hull fitting algorithm to generate a candidate box label of the to-be-processed remote sensing image raster data; According to the binary mask and the candidate box label corresponding to all the to-be-processed remote sensing image raster data, a training set is generated; An initial model is trained through the training set to obtain a trained mine land instance segmentation model, wherein the initial model comprises a boundary enhancement module and a boundary optimization module.

Citation Information

Patent Citations

  • Single-stage instance image segmentation method and device and computer equipment

    CN115222946A

  • Single-stage cell nucleus instance segmentation method for medical microscopic image

    CN116309545A