Remote sensing image ground fracture identification method, device, equipment, medium and product
By combining the improved U-Net network with ResNet50 and attention mechanism, the problems of low efficiency and low accuracy of traditional ground fissure identification methods are solved, and efficient and accurate identification and extraction of ground fissures are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-26
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional methods for identifying ground fissures are inefficient and inaccurate, especially in complex environments where they are sensitive to noise and lighting changes, making it difficult to achieve efficient and accurate identification.
An improved U-Net network is used for ground fissure identification. ResNet50 is used as the backbone network to extract deep features, and an attention mechanism is introduced in the decoding stage. By upsampling and convolution of multi-scale deep feature maps and semantic features, combined with the attention module, the ability to identify fissure regions is improved.
It enables accurate identification and continuous extraction of ground fissures of different scales and morphologies, improving identification efficiency and accuracy. It can effectively suppress background interference in complex backgrounds and enhance the ability to express the features of the fissure region.
Smart Images

Figure CN121789079A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and remote sensing image processing, and in particular to a method, apparatus, device, medium, and product for identifying ground fissures in remote sensing images. Background Technology
[0002] Surface fissures, as a direct manifestation of damage caused by coal mining activities, are of significant practical importance for guiding the rational development of coal resources, preventing geological disasters, and protecting the ecological environment. Furthermore, the presence of surface fissures can exacerbate environmental problems such as soil erosion and land degradation, posing a challenge to regional sustainable development. Therefore, effective monitoring and research of mining-induced ground fissures are crucial for guiding the rational development of coal resources, preventing geological disasters, and protecting the ecological environment.
[0003] Traditional methods for identifying ground fissures mainly fall into two categories: manual visual interpretation and semi-automatic identification based on image processing. Manual interpretation relies on expert experience, is inefficient and highly subjective, and is unsuitable for rapid processing of large-scale, high-resolution images. While traditional image processing-based fissure detection methods, such as edge detection, threshold segmentation, and morphological operations, are simple to implement, they are sensitive to noise, lighting variations, and background complexity, resulting in low accuracy and stability. Therefore, a method capable of efficiently and accurately identifying ground fissures is needed. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, equipment, medium, and product for identifying ground fissures in remote sensing images, which can improve the efficiency and accuracy of ground fissure identification.
[0005] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for identifying ground fissures in remote sensing images, including: Acquire actual satellite remote sensing images of the area to be identified; The actual satellite remote sensing images are identified using a ground fissure identification network model to obtain ground fissure identification results. The ground fissure identification network model is a trained improved U-Net network. The improved U-Net network uses ResNet50 as the backbone network to extract deep features of the data in the encoding stage, and introduces an attention mechanism in the decoding stage to improve the extraction effect of surface fissures.
[0006] In one embodiment, the actual satellite remote sensing image is a remote sensing image tile from Google Earth.
[0007] In one embodiment, the actual satellite remote sensing image is identified using a ground fissure identification network model to obtain ground fissure identification results, specifically including: The encoder of the ground fissure identification network model is used to obtain multi-scale deep feature maps and semantic features at different resolutions from actual satellite remote sensing images. The decoder based on the ground fissure recognition network model upsamples and convolves multi-scale deep feature maps and semantic features of different resolutions to obtain ground fissure recognition results.
[0008] In one embodiment, the training process of the ground fissure identification network specifically includes: Acquire raw image data; The original image data is cropped and divided into an initial training set according to the proportions; The initial training set is then characterized, normalized, and augmented to obtain the final training set. Based on the final training set, the improved U-Net network is trained using pixel-level cross-entropy as the loss function to obtain the ground fissure recognition network model.
[0009] In one embodiment, the loss function further includes: Dice Loss and boundary loss.
[0010] In one embodiment, an attention mechanism is introduced after the fused convolution of the improved U-Net network decoder.
[0011] Secondly, this application provides a remote sensing image ground fissure identification device, comprising: The acquisition module is used to acquire actual satellite remote sensing images of the area to be identified; The identification module is used to identify the actual satellite remote sensing image using a ground fissure identification network model to obtain ground fissure identification results. The ground fissure identification network model is a trained improved U-Net network. The improved U-Net network uses ResNet50 as the backbone network to extract deep features of the data in the encoding stage and introduces an attention mechanism in the decoding stage to improve the extraction effect of surface fissures.
[0012] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the remote sensing image ground fissure identification method.
[0013] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the remote sensing image ground fissure identification method.
[0014] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the remote sensing image ground fissure identification method.
[0015] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a method, apparatus, device, medium, and product for identifying ground fissures in remote sensing images. The method involves acquiring actual satellite remote sensing images of the area to be identified; using a ground fissure identification network model to identify the actual satellite remote sensing images to obtain ground fissure identification results; and by introducing an attention mechanism into the improved U-Net network, the network's ability to express fissure region features is effectively enhanced, enabling accurate identification and continuous extraction of ground fissures of different scales and morphologies, thereby improving the efficiency and accuracy of ground fissure identification. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is an application environment diagram of a remote sensing image ground fissure identification method according to an embodiment of this application; Figure 2 A flowchart illustrating a method for identifying ground fissures in remote sensing images, provided as an embodiment of this application; Figure 3 This is a schematic diagram of residual structure 1; Figure 4 This is a schematic diagram of residual structure 2; Figure 5 This is a schematic diagram of the encoding and decoding process results; Figure 6 A schematic diagram of the improved U-Net network structure; Figure 7 A schematic diagram of the functional modules of a remote sensing image ground fissure identification device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] The remote sensing image ground fissure identification method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 101 communicates with server 102 via a network. A data storage system can store the data that server 102 needs to process. The data storage system can be set up independently, integrated into server 102, or placed in the cloud or on another server. Terminal 101 can send actual satellite remote sensing images of the area to be identified to server 102. Server 102 receives the actual satellite remote sensing images of the area to be identified and uses a ground fissure identification network model to identify the ground fissures, obtaining the ground fissure identification result. Server 102 can feed back the obtained ground fissure identification result to terminal 101. Furthermore, in some embodiments, remote sensing image ground fissure identification can also be implemented independently by server 102 or terminal 101. For example, terminal 101 can directly perform remote sensing image ground fissure identification on the actual satellite remote sensing images of the area to be identified, or server 102 can perform remote sensing image ground fissure identification from the data storage system.
[0021] The terminal 101 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 102 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.
[0022] In one exemplary embodiment, such as Figure 2 As shown, a method for identifying ground fissures in remote sensing images is provided. This method is executed by a computer device, specifically a terminal or server, or both. In this embodiment, the method is applied to... Figure 1 Taking server 102 as an example, the following steps are included.
[0023] Step 201: Obtain actual satellite remote sensing images of the area to be identified.
[0024] Step 202: The actual satellite remote sensing image is identified using a ground fissure identification network model to obtain ground fissure identification results; the ground fissure identification network model is a trained improved U-Net network; the improved U-Net network uses ResNet50 as the backbone network to extract deep features of the data in the encoding stage, and introduces an attention mechanism in the decoding stage to improve the extraction effect of surface fissures.
[0025] By implementing the above steps and using the improved U-Net network, the network's ability to express the characteristics of the crack region can be effectively enhanced, enabling accurate identification and continuous extraction of ground cracks of different scales and morphologies, thereby improving the efficiency and accuracy of ground crack identification.
[0026] In one exemplary embodiment, the actual satellite remote sensing image is a Google Earth remote sensing image tile.
[0027] The actual satellite remote sensing image is identified using a ground fissure identification network model to obtain ground fissure identification results. Specifically, this includes: using the encoder of the actual satellite remote sensing image to obtain multi-scale deep feature maps and semantic features at different resolutions; and using the decoder based on the ground fissure identification network model to upsample and convolve the multi-scale deep feature maps and semantic features at different resolutions to obtain ground fissure identification results.
[0028] Specifically, an input image is first scaled to a fixed size and normalized, then fed into an encoder with a ResNet50 backbone. Through a series of convolutions, batch normalization, ReLU activation, max pooling, and residual Bottleneck modules, the original image is progressively transformed into multi-scale deep feature maps, and semantic features at different resolutions are extracted at four scales for subsequent skip connections. These features are then fed into an improved U-Net decoding structure. The decoding part upsamples the deepest features step-by-step through upsampling layers, concatenating them with the encoded features at each level. After concatenation and convolution, an attention module applies attention weights to the fused features, suppressing background redundancy and highlighting elongated texture regions associated with cracks. As the decoder continuously upsamples and convolves, the network eventually recovers a feature map with the same spatial resolution as the input. The number of channels is compressed into the number of categories (cracks and non-cracks) through the last 1×1 convolution, and a pixel-level segmentation mask is output. During training, loss functions such as cross-entropy are used to continuously optimize the parameters on a large number of labeled samples, enabling the model to learn to automatically segment crack regions in complex backgrounds.
[0029] In practical applications, the training process of the ground fissure recognition network specifically includes: acquiring original image data; cropping the original image data and dividing it into an initial training set according to a ratio; performing featureization, normalization, and data augmentation on the initial training set to obtain a final training set; and training an improved U-Net network using pixel-level cross-entropy as the loss function based on the final training set to obtain a ground fissure recognition model. The input of the improved U-Net network is the image data to be recognized, and the output is the ground fissure recognition result.
[0030] As another exemplary embodiment, the loss function further includes: Dice Loss and boundary loss.
[0031] In practical applications, an attention mechanism is introduced after the fusion convolution of the improved U-Net network decoder. This attention mechanism improves the extraction of surface cracks. Specifically, after the stitching and convolution in the decoding stage, the attention module applies attention weights to the fused features, suppressing redundant background information and highlighting the elongated texture regions associated with the cracks.
[0032] To address the shortcomings of existing ground fissure image recognition methods, such as insufficient recognition accuracy, poor robustness to complex backgrounds, and weak ability to extract fissure details, this application introduces an attention mechanism module into a semantic segmentation network. This effectively enhances the network's ability to express fissure region features, enabling accurate identification and continuous extraction of ground fissures of different scales and morphologies. This application is primarily applied to scenarios such as geological disaster monitoring, infrastructure safety inspection, and surface deformation analysis.
[0033] In another exemplary embodiment, this application provides a method for identifying ground fissures in remote sensing images, encompassing the entire process from data acquisition to fissure identification, including the following steps: Data Source: In watershed-scale research on ground fissure identification and extraction, using satellite remote sensing imagery as the primary data source offers significant scientific and engineering advantages. Compared to UAV aerial photography data, satellite remote sensing provides comprehensive advantages in terms of coverage, spatiotemporal continuity, data standardization, and operability. Satellite remote sensing imagery possesses wide-area coverage capabilities, enabling rapid observation of large areas within a single scene or a small number of scenes, significantly improving data acquisition efficiency and spatial continuity. For watershed areas spanning tens to hundreds of square kilometers, using UAV imagery for full coverage not only faces high costs and complex flight organization issues but is also limited by weather conditions, battery life, and airspace control, hindering the systematic identification of large-scale ground fissures. Therefore, Google Earth remote sensing image tiles were selected as the raw image data, with a maximum resolution of 0.15m, encompassing red, green, and blue bands.
[0034] Image preprocessing: Radiometric correction: The purpose of radiometric correction is to eliminate brightness differences caused by sensor characteristics, imaging time, solar elevation angle, and variations in surface reflectivity. Google Earth imagery is a product that has undergone visual enhancement and color balancing, and in crack detection tasks, the model can directly learn pixel texture features, thus eliminating the need for radiometric correction.
[0035] Geometric correction: Geometric correction aims to align the imagery to the true geographic coordinate system, eliminating spatial distortions caused by terrain undulations and sensor orientation. Google Earth's base map imagery has already undergone orthorectification, meaning that geometric distortion has been corrected using topographic elevation (DEM) and sensor parameters. Therefore, no further geometric correction is needed for ground fissure identification tasks.
[0036] Dataset Creation: The original image was padded to a multiple of the crop size in both length and width, and then randomly cropped into 512×512 sub-blocks. These were then divided into training, validation, and test sets in a 7:1.5:1.5 ratio. Finally, other working surfaces were selected for instance extraction. Semantic segmentation, through feature recognition of semantic information in the image, divides the image into regions of different semantic categories. It can recognize and understand the content of each pixel in the image: its semantic region annotation and prediction are pixel-level. This application uses the annotation tool Labelme to perform pixel-level annotation of cracks in the original image. By outlining the crack contours, all crack regions in the image are labeled with a label value of 1. Non-crack regions, i.e., all regions without cracks (the background), have a default label value of 0. The above training and test sets are annotated at the pixel level using the tool.
[0037] Featureization and Normalization: Featureization refers to extracting texture, spectral, morphological, or spatial structural features from raw remote sensing images that are helpful for crack identification. It transforms complex raw image data into feature representations that more easily distinguish between "cracks" and "non-cracks." For grayscale images, the Gray-Level Co-occurrence Matrix (GLCM) is calculated (direction = 0°, 45°, 90°, 135°), and contrast, entropy, and homogeneity are extracted. Since the pixel value ranges differ significantly between image sources and a unified input scale is required, normalization maps image pixel values to a uniform numerical range [0,1] to eliminate brightness and contrast differences between different images. This helps accelerate gradient descent, stabilize training, and allow different images to enter the network at the same scale. Furthermore, the distribution of normalized feature values is more stable, and gradient convergence is faster, avoiding training oscillations caused by excessively large pixel value ranges.
[0038] Data augmentation: Data augmentation involves generating similar but different training samples by applying a series of random changes to the training images, thereby expanding the size of the training set. Furthermore, data augmentation is applied because random alterations to image data not only expand the dataset but also simulate image performance under different external scenarios, reducing the model's dependence on certain attributes. This method employs various augmentation techniques, including flipping and rotation (horizontal / vertical flipping, random angle rotation (45°, 90°, 180°, 270°), translation, affine and perspective transformations, random contrast enhancement, and mesh distortion, allowing the model to experience richer variations in scene, lighting, texture, and scale, thus improving robustness and generalization ability. This makes the model less susceptible to being constrained by specific lighting conditions, sensor characteristics, and background textures during learning, enabling more reliable predictions under new observation conditions in practical applications. Through appropriate augmentation strategies, overfitting caused by scarce labeled data can be mitigated, and the ability to identify minute cracks, complex boundaries, and occlusion can be improved.
[0039] Deep feature extraction: An improved U-Net network is used as the basic structure, which consists of an encoder and a decoder. The encoder consists of multiple convolutions, batch normalization and ReLU activation function, which are used to extract image features layer by layer. Each layer is downsampled through max pooling to gradually increase the receptive field. The decoder gradually restores the spatial resolution of the image through upsampling and uses skip connections to fuse the shallow and deep features of the corresponding layers.
[0040] U-Net is a highly symmetrical U-shaped network structure designed for pixel-level segmentation. It is widely used in tasks such as medical image segmentation and crack segmentation. It has a clear encoding and decoding approach, with skip connections being key. These connections combine high-resolution features from the encoder stage with upsampled features from the decoder stage to preserve positional information and details. In the encoding part, it uses two convolutional layers plus a pooling layer for four downsampling operations, and maintains the same number of sampling operations during decoding. It then uses two ordinary convolutional layers plus a transposed convolutional layer for four upsampling operations. Furthermore, during each upsampling operation, it introduces overlapping skip connections to fuse feature layers of different depths. Spatial domain information is crucial for segmentation tasks. In the encoding part of the network, the feature map resolution is significantly reduced after downsampling. Skip connections introduce high-resolution features from shallower convolutional layers, resulting in better performance in decoding and predicting segmentation results.
[0041] In deep learning research, attention mechanisms adaptively model the importance of features, enabling the network to focus more on task-related information, thereby significantly improving feature representation capabilities. Based on the different objects of attention, attention mechanisms can be mainly divided into three categories: channel attention, spatial attention, and joint channel-space attention.
[0042] AG (Attention Gate) attention mechanism.
[0043] AG (Attention Gating) is a typical spatial attention mechanism that focuses on "which spatial locations in the feature map are more important." Its core objective is to highlight the target region and suppress irrelevant interference in a complex background, thereby improving the model's ability to locate and represent key regions. This application introduces a spatial attention gating module (AttentionGate, AG) in the decoding stage to further filter and enhance the feature map after fusion and convolutional reconstruction at this scale by the decoder. Unlike the common "skip-connection pre-gating," the AG module in this application is located inside the decoding block. After the upsampled features are concatenated and fused with features at the same scale as the encoder, and then refined through two convolutions to obtain intermediate features, spatial attention weighting is applied to these features, thereby achieving saliency enhancement of the target region and suppression of background interference.
[0044] Let the output features of the decoded block after fusion and convolutional reconstruction be... Where H and W are spatial dimensions, and C is the number of channels. The value is a real number. The AG module generates a spatial attention weight map by modeling the importance of spatial locations. The AG module uses additive attention to spatially model the fused features. First, it performs a linear mapping on the input features (typically achieved through 1×1 convolutions for dimensionality reduction): .
[0045] The joint response is obtained by adding the two mapping results and then applying ReLU activation. .
[0046] Then, a spatial weight map is generated using a 1×1 convolutional mapping and the Sigmoid function: .
[0047] This represents a linear transformation operator that maps intermediate features s to a single-channel attention response. In practice, this is typically accomplished by a 1×1 convolutional layer. The above calculation process can also be uniformly represented as: .
[0048] in Let z be the linear mapping matrix of the input feature z. This is used to map intermediate features to a single-channel attention response, where b is a bias term. This represents the Sigmoid function. The final result is... This is a spatial attention map, which characterizes the importance of each spatial location to the current segmentation task. The AG module then applies this weight map to the original features: 。
[0049] in," " indicates a pixel-by-pixel multiplication operation, The weighted output features are shown below. Through this gating mechanism, the network can further highlight the spatial response corresponding to the crack region during the decoding stage, while suppressing the activation intensity of the background region. This makes the subsequent upsampling reconstruction process more focused on the target region, thereby effectively improving the segmentation accuracy and continuity of fine crack structures.
[0050] The AG module, through a spatial-level "gating" mechanism, enables the network to focus on potential target regions in complex scenes. It not only achieves explicit attention modeling in the spatial dimension but also guides low-level features through high-level semantics, realizing a "top-down" region selection mechanism. In the task of segmenting remote sensing images of ground fissures, this mechanism significantly enhances the network's response to fine fissure structures, suppresses interference from complex backgrounds such as soil texture and vegetation shadows, and provides a purer and more discriminative feature base for subsequent fine reconstruction.
[0051] This application introduces Attention Gate (AG) as a spatial attention mechanism in the decoding stage. The main reason for this is that ground fissure targets in remote sensing images are characterized by their small scale, elongated shape, weak contrast, and susceptibility to interference from complex backgrounds. Compared to attention mechanisms that focus on channel recalibration, AG directly gates features in the spatial dimension, which is more suitable for the localization needs of weak targets like fissures. The features in the decoding stage have already incorporated detailed and deep semantic information provided by the encoder, possessing strong discriminative capabilities. At this point, AG performs pixel-level filtering of the fused features, adaptively suppressing invalid responses from background areas such as soil textures, shadows, and gully edges, while significantly enhancing the activation intensity of fissure regions, making the upsampling reconstruction process more focused on the target area. This gating mechanism effectively reduces the false alarm rate and strengthens the continuous representation of small fissure structures without significantly increasing model complexity, thereby improving the accuracy and stability of ground fissure semantic segmentation in complex scenes.
[0052] This application builds upon the standard U-Net encoding / decoding structure by progressively amplifying the deepest features through upsampling layers in the decoding phase. Simultaneously, at each level, these features are concatenated with the corresponding scale's encoded features. After concatenation and convolution, an Attention Gate (AG) module is introduced. This module applies attention weights to the fused features, automatically selecting regions that are helpful for segmentation, suppressing irrelevant information, and improving the ability to identify crack regions. While VGG or ResNet are typically options, this application chooses the deeper ResNet50 as the backbone network for high-level feature extraction. The improved U-Net structure handles progressive downsampling and upsampling, and the AG module is used in the decoding stage to enhance the crack filtering capability.
[0053] Figure 3 and Figure 4 The residual structure units, residual block 1 and residual block 2, are shown. They are the basic units of ResNet. The shortcut convolution in residual block 1 is to ensure that the input and output feature matrices of the residual block maintain the same shape and size when passing through the shortcut. The main branches in the residual structure are all composed of two 1 Convolution with 1 plus 3 in the middle The convolution consists of 3 layers, and after convolution, a batch normalization (BN) layer is connected to normalize the output features.
[0054] The improved U-Net network architecture, including specific network structure and parameters, is as follows: Figure 5 and Figure 6 As shown.
[0055] Overall shape: U-shaped. The left side is the encoder (downsampling path), the right side is the decoder (upsampling path), and there is usually a bottleneck layer in the middle.
[0056] The encoder on the left (downsampling path): multiple convolutions + ReLU + downsampling, the spatial size is halved with each downsampling, and the number of channels increases (e.g., 64 / 128 / 256 / 512).
[0057] In the encoding process, a deeper ResNet50 is used as the backbone network to extract deep features from the data. Regarding the downsampling rate, a secondary downsampling is performed, taking into account the characteristics of the deep network and the dataset features, to increase the receptive field while acquiring deep features. Skip connections are incorporated during the encoding and reconstruction process to extract crack features through sufficiently deep feature layers, combined with high-resolution shallow features for accurate image prediction and reconstruction. The ResNet50 uses Bottleneck residual blocks, with each block containing three convolutional layers. Figure 3 , Figure 4The residual structural units, residual block 1 and residual block 2, are shown. They are the basic units of ResNet. Batch normalization is added after the convolutional layer to make the feature distribution stable and easy to train. ReLU introduces nonlinearity and suppresses useless activation, so that the features are both stable and have nonlinear expressive power.
[0058] The right side is the decoder (upsampling path): layer-by-layer upsampling + convolution. Each layer splices features from the encoder and the upsampled output of the next layer, aligns spatial dimensions, fuses boundary details, and adds an attention mechanism module. This module automatically selects regional features that are helpful for the segmentation task, suppresses background noise, and makes the network more focused on narrow, irregularly shaped crack targets, thus improving the ability to identify ground cracks.
[0059] Output layer: 1×1 convolutions are used to map features to the required number of classes, thus obtaining the class probability distribution for each pixel. Finally, loss functions such as pixel-level cross-entropy are used for training.
[0060] Input and output: The input is an original image, and the output is a pixel-level segmentation map of the same size as the input (each pixel belongs to a certain category).
[0061] The encoder (shrinking path) progressively reduces the spatial dimensionality of the feature map through multiple layers of convolution, non-linear activation, and downsampling, while extracting increasingly abstract features. This step aims to gain an understanding of the global context of the input while preserving layer-by-layer local details.
[0062] Central bottleneck layer: The connection point between encoding and decoding, usually a stack of several convolutional layers, providing the deepest feature representation.
[0063] Decoder (Expansion Path): Gradually restores spatial resolution through upsampling. Each level makes skip connections with the corresponding encoder layer, concatenating or adding the high-resolution features from the encoding stage with the upsampled features of the current layer, and then fusing them through convolution. This skip connection ensures that the decoding stage can utilize the detailed information from lower layers, improving boundary localization and detail recovery capabilities.
[0064] AG module structure and principle.
[0065] AG input: Features from the left-side encoding and the right-side upper layer decoding.
[0066] Main steps: Feature fusion: Channel mapping and spatial alignment are performed on encoder and decoder features, and then the number of channels is reduced by linear transformation (usually 1x1 convolution) to facilitate subsequent fusion.
[0067] Activation & Weighting: The two features are summed and then activated by ReLU, and then a weight map with the same space is generated by a small convolution and sigmoid gate.
[0068] Feature selection: Element-wise multiplication of the weight map and encoder features is performed to retain regional features that are useful for the task and suppress redundancy and background.
[0069] Output: The decoded features after convolution and attention modules are passed to the next upsampling module for further upsampling and convolution.
[0070] Inputs and outputs of encoders and decoders.
[0071] Encoder: Input: Original image, size 512×512×3 (height×width×channels, RGB three channels). After one 7×7 convolution, the output is 256×256×64, and after one max pooling, the output is 128×128×64. Output of each layer: As the number of layers increases, the spatial size is halved layer by layer, and the number of channels gradually increases.
[0072] Encoder-1 Input: 128×128×64 Output: 128×128×256.
[0073] Encoder-2 Input: 128×128×256 Output: 128×128×256.
[0074] Encoder-3 Input: 128×128×256 Output: 64×64×512.
[0075] Encoder-4 Input: 64×64×512 Output: 32×32×1024.
[0076] Bottleneck Input: 32×32×1024 Output: 16×16×2048 (last layer of the main trunk).
[0077] Decoder: Input: Central bottleneck feature, size 16×16×2048, and decoded features from the previous layer after convolution and attention modules. Output of each layer: Spatial size increases progressively, while the number of channels decreases.
[0078] Decoder-4 input: Bottleneck output 32×32×2048, Encoder-4 feature 32×32×1024, output: 32×32×512.
[0079] Decoder-3 input: Upsampled output 64×64×512, Encoder-3 feature 64×64×512; Output: 64×64×256.
[0080] Decoder-2 input: Upsampled output 128×128×256, Encoder-2 feature 128×128×256 Output: 128×128×128.
[0081] Decoder-1 input: upsampled output 256×256×128, Encoder-1 feature 256×256×64; Output: 256×256×64.
[0082] Output layer: The Decoder-1 output is upsampled once to 512×512×64, and finally convolved by 1×1 to output 512×512×2 (pixel-level category distribution, where 2 is the number of segmentation categories, including cracks and background).
[0083] Training process: 1. Data preparation.
[0084] The input data consisted of raw images (remote sensing images of ground fissures) with dimensions of 512×512×3. Labels were assigned to each pixel as a category (total of 2 categories: fissures and background).
[0085] The dataset is split into training and validation sets, and preprocessed during loading, including normalization and enhancement (such as random flipping and random contrast enhancement). Data loading is performed using PyTorch's DataLoader, with num_workers=4 threads.
[0086] 2. Model initialization and loading.
[0087] An improved U-Net architecture is adopted, with ResNet50 as the backbone, a network input size of 512×512, and 2 output classes. The main intervention training weights are loaded first.
[0088] 3. Optimizer and loss function settings.
[0089] The optimizer chosen is Adam, with an initial learning rate of 0.0001 and a minimum learning rate of 0.01. The momentum parameter is 0.9. Adaptive Moment Estimation (Adam) is another method for calculating the adaptive learning rate for each parameter. This algorithm calculates the exponential moving average of the gradient and the squared gradient. and The decay rate of these moving averages was controlled: .
[0090] .
[0091] and These are the estimates of the first moment (mean) and second moment (non-central variance) of the gradient at the t-th iteration, respectively. Set the current gradient g t and past gradient average By weight and By weighting, a smooth gradient estimate is formed to guide parameter updates. This is more stable and less susceptible to disturbances by single noisy gradients than using instantaneous gradients directly. This can be understood as the "smooth mean of the gradient magnitude," used to characterize the fluctuations in the gradient. The optimizer uses this value when updating parameters. Using it as the denominator is equivalent to "adaptively scaling" the gradient. The step size for parameters with large gradient fluctuations becomes smaller, while the step size for parameters with small fluctuations becomes relatively larger, thus obtaining an adaptive learning rate. It is the gradient. The optimizer adjusts the parameters along the "opposite direction of the gradient" to gradually reduce the loss function, thereby allowing the model to fit the training set better. This is an operator for calculating the gradient (partial derivative vector) of a parameter W. Let be the loss function, representing the loss function under the current parameters. The magnitude of the error in the model. This represents the complete set of parameters of the model at the t-th iteration. and These are parameters used to control the first-order moment decay rate and the second-order moment decay rate. and Vectors initialized to 0, especially in the initial few steps of initialization, and with a small decay rate (i.e. and Approaching point 1) presents a significant bias. Therefore, these biases are offset by calculating the bias-corrected first and second moment estimates: .
[0092] and This is the result after deviation correction. and Hyperparameters and The power of t, that is, and The number obtained by multiplying itself by t times.
[0093] Adam has new update rules based on the updated parameters: .
[0094] The learning rate; It is a very small constant used to prevent division by zero. This represents the model parameters at the (t+1)th iteration, and is generally recommended. The default value is 0.9. The default value is 0.999. Generally, Adam is highly robust to hyperparameters and not very sensitive to the learning rate, so it is widely used in deep learning models.
[0095] The learning rate is scheduled using cosine annealing, gradually reducing the learning rate as training progresses.
[0096] The loss function is mainly pixel-level cross-entropy, and auxiliary terms such as Dice Loss, Focal Loss, and class weights can be added as needed.
[0097] Cross-entropy loss is one of the most commonly used loss functions in deep learning. Its specific calculation form in semantic segmentation and multi-class pixel-level classification tasks is as follows: .
[0098] in, Let N be the cross-entropy, N be the number of pixels involved in the calculation, and C be the number of categories. Predict whether the classification of the i-th pixel is correct. The i-th pixel is classified into categories. The probability of.
[0099] 4. Staged training of freezing / thawing the main trunk.
[0100] It supports freezing parameters in the backbone network, training only the decoder to accelerate convergence initially. After training to a specified epoch, the backbone is unfrozen, and the entire network is jointly optimized to improve global segmentation accuracy. An epoch is the basic unit for representing the number of training rounds in deep learning.
[0101] The initial training epoch is set to 0, the trunk is frozen for training epoch=50, and the joint training is unfrozen until epoch=100.
[0102] 5. Main loop training process.
[0103] The training process for each epoch includes the following steps: Batch data is read and input into the improved U-Net network to complete forward inference and obtain the class probability for each pixel.
[0104] The loss is calculated based on the loss function, and backpropagation is performed. The optimizer updates parameters and learning rate according to the set strategy (Adam, cosine annealing, etc.). The loss and evaluation metrics (such as mIoU, Dice) are recorded and periodically written to a log file.
[0105] The weights are saved every 5 epochs, and the model performance is evaluated on the validation set. The best model weights are saved.
[0106] 6. Results and Model Output.
[0107] All stage weights and logs are automatically saved, facilitating subsequent model reproduction and inference applications.
[0108] After training, the optimized and improved U-Net parameter model is output, which can then be used for actual binary classification image segmentation tasks.
[0109] By introducing a residual network structure and an attention mechanism module, this application enables the model to adaptively highlight crack region features, significantly improve the model's ability to identify small cracks, and preserve the continuous morphology of long and thin cracks. Moreover, the model is lightweight and easy to deploy, making it suitable for rapid crack extraction from large-scale remote sensing images or UAV images in practical engineering applications.
[0110] In addition, this application provides the following detailed questions.
[0111] 1. Data annotation precision and quality.
[0112] Ground fissure targets are usually narrow and irregularly distributed, so pixel-level ground truth annotations need to be extremely accurate to ensure clear boundaries; otherwise, the model is prone to misclassification or omission.
[0113] The ratio of label categories (cracks / background) is extremely unbalanced. It is recommended to sample and augment "small sample" crack images to avoid the model focusing only on the background.
[0114] 2. Data augmentation strategies.
[0115] Employing highly targeted enhancement methods, such as random rotation, affine transformation, lighting perturbation, and contrast enhancement, helps improve the model's adaptability to cracks in various surface environments.
[0116] Add appropriate amounts of noise, blur, and other special effects to simulate complex scene data such as those in the wild and remote sensing, thereby improving robustness.
[0117] 3. Input preprocessing and size selection.
[0118] If the ground fissure features are subtle, a higher resolution input (e.g., 512×512~1024×1024) is recommended, but attention should be paid to balancing GPU memory and training speed. Image normalization (0~1 or mean-variance) has a positive effect on model stability.
[0119] 4. Attention module parameters and design.
[0120] An overly weak attention module may result in excessive attenuation of jump features, leading to the loss of details; an overly strong module, on the other hand, introduces redundancy. It is recommended to fine-tune the attention channel compression ratio and activation strength on the validation set.
[0121] You could try further optimizing the position of the attention module, such as using the attention module only for the lowest level jump or using a tiered setup.
[0122] 5. Selection of loss function / evaluation metric.
[0123] Since ground fissures are small targets with extremely fine boundaries, it is recommended to add Dice Loss and boundary loss in addition to the main loss to improve sensitivity to small targets.
[0124] The evaluation uses comprehensive metrics such as mIoU, F1-score, and Precision-Recall, and does not only consider the total pixel accuracy.
[0125] 6. Overfitting and generalization.
[0126] Due to the complexity of the crack scene and background, the backbone can be frozen / thawed in stages for training to improve the generalization effect.
[0127] Overfitting in complex scenarios can be prevented by using techniques such as Dropout and L2 regularization.
[0128] 7. Post-processing methods.
[0129] The prediction results need to be combined with morphological post-processing, such as erosion, dilation, edge refinement, area filtering, etc., to help eliminate false cracks and correct details.
[0130] 8. The complexity of real-world scenarios.
[0131] Ground fissures are often mixed with interference from vegetation, shadows, and water bodies. The model needs to be combined with multimodal features or multi-temporal remote sensing sequences to further improve the accuracy of identification.
[0132] Based on the same inventive concept, this application also provides a remote sensing image ground fissure identification device for implementing the remote sensing image ground fissure identification method described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more remote sensing image ground fissure identification device embodiments provided below can be found in the limitations of the remote sensing image ground fissure identification method above, and will not be repeated here.
[0133] In one exemplary embodiment, such as Figure 7 As shown, a remote sensing image ground fissure identification device is provided, comprising: The acquisition module is used to acquire actual satellite remote sensing images of the area to be identified.
[0134] The identification module is used to identify the actual satellite remote sensing image using a ground fissure identification network model to obtain ground fissure identification results. The ground fissure identification network model is a trained improved U-Net network. The improved U-Net network uses ResNet50 as the backbone network to extract deep features of the data in the encoding stage and introduces an attention mechanism in the decoding stage to improve the extraction effect of surface fissures.
[0135] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 8 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores remote sensing image data for identifying ground fissures. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When the computer program is executed by the processor, it implements a method for identifying ground fissures from remote sensing images.
[0136] Those skilled in the art will understand that Figure 8 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method embodiments.
[0137] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the above-described method embodiments.
[0138] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method embodiments.
[0139] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0140] In this application, all actions to acquire signals, information, or data are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.
[0141] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0142] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0143] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0144] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for identifying ground fissures in remote sensing images, characterized in that, The method for identifying ground fissures in remote sensing images includes: Acquire actual satellite remote sensing images of the area to be identified; The actual satellite remote sensing images are identified using a ground fissure identification network model to obtain ground fissure identification results. The ground fissure identification network model is a trained improved U-Net network. The improved U-Net network uses ResNet50 as the backbone network to extract deep features of the data in the encoding stage, and introduces an attention mechanism in the decoding stage to improve the extraction effect of surface fissures.
2. The method for identifying ground fissures in remote sensing images according to claim 1, characterized in that, The actual satellite remote sensing images mentioned are remote sensing image tiles from Google Earth.
3. The method for identifying ground fissures in remote sensing images according to claim 1, characterized in that, The actual satellite remote sensing images are used to identify ground fissures using a ground fissure identification network model to obtain ground fissure identification results, specifically including: The encoder of the ground fissure identification network model is used to obtain multi-scale deep feature maps and semantic features at different resolutions from actual satellite remote sensing images. The decoder based on the ground fissure recognition network model upsamples and convolves multi-scale deep feature maps and semantic features of different resolutions to obtain ground fissure recognition results.
4. The method for identifying ground fissures in remote sensing images according to claim 1, characterized in that, The training process of the ground fissure identification network specifically includes: Acquire raw image data; The original image data is cropped and divided into an initial training set according to the proportions; The initial training set is then characterized, normalized, and augmented to obtain the final training set. Based on the final training set, the improved U-Net network is trained using pixel-level cross-entropy as the loss function to obtain the ground fissure recognition network model.
5. The method for identifying ground fissures in remote sensing images according to claim 4, characterized in that, The loss function also includes: Dice Loss and boundary loss.
6. The method for identifying ground fissures in remote sensing images according to claim 1, characterized in that, An attention mechanism is introduced after the fused convolution of the improved U-Net network decoder.
7. A remote sensing image ground fissure identification device, characterized in that, The remote sensing image ground fissure identification device includes: The acquisition module is used to acquire actual satellite remote sensing images of the area to be identified; The identification module is used to identify the actual satellite remote sensing image using a ground fissure identification network model to obtain ground fissure identification results. The ground fissure identification network model is a trained improved U-Net network. The improved U-Net network uses ResNet50 as the backbone network to extract deep features of the data in the encoding stage and introduces an attention mechanism in the decoding stage to improve the extraction effect of surface fissures.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the remote sensing image ground fissure identification method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the remote sensing image ground fissure identification method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the remote sensing image ground fissure identification method according to any one of claims 1-6.