A slope crack detection method and device based on super-resolution processing
By employing multimodal image acquisition, super-resolution processing, and multi-scale feature enhancement techniques, combined with a target detection model, the problem of insufficient accuracy and efficiency in slope crack detection has been solved, achieving high-precision identification and localization of fine cracks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YUNNAN HUADIAN LUDILA HYDROPOWER CO LTD
- Filing Date
- 2024-12-16
- Publication Date
- 2026-05-15
AI Technical Summary
Existing slope crack detection methods are difficult to accurately capture and identify small cracks in complex environments, and suffer from insufficient detection accuracy and efficiency.
Multimodal slope images are acquired simultaneously using multiple image acquisition devices, and super-resolution processing and image detail enhancement are performed. By utilizing feature extraction models and multi-scale feature enhancement techniques, combined with target detection models, crack identification and localization are achieved.
It significantly improves the accuracy and efficiency of slope crack detection, can accurately identify and locate minute cracks, reduce background noise interference, and ensure high reliability of detection results.
Smart Images

Figure CN120031788B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image analysis technology, and more specifically, to a method and apparatus for detecting slope cracks based on super-resolution processing. Background Technology
[0002] Slope crack detection refers to the use of certain technical means to monitor and analyze whether there are cracks in natural or artificial slopes in engineering projects such as reservoirs, mountains, construction sites, and highways, in order to ensure slope safety and prevent disasters.
[0003] Existing crack detection methods mainly rely on traditional visual inspection, manual inspection, or simple image processing techniques. While these methods can identify cracks to some extent, they often face several challenges. First, traditional image acquisition techniques struggle to accurately capture and identify crack features in low-resolution images when dealing with crack detection in complex environments. This is due to various interference factors on slope surfaces, such as changes in lighting and complex surface textures, which affect detection accuracy and reliability. Furthermore, existing crack feature extraction methods are weak in identifying small cracks and, in complex scenarios, struggle to provide sufficiently accurate crack localization and quantitative analysis. Therefore, there is still room for improvement in the accuracy and efficiency of slope crack detection technologies.
[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this disclosure is to provide a slope crack detection method, a slope crack detection device, an electronic device, and a computer-readable storage medium based on super-resolution processing, thereby overcoming, to at least a certain extent, the problems of insufficient accuracy and efficiency in related slope crack detection methods.
[0006] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0007] According to a first aspect of the present disclosure, a slope crack detection method based on super-resolution processing is provided, comprising: simultaneously acquiring data of a target slope area using multiple image acquisition devices to obtain multimodal slope images; performing super-resolution processing and image detail enhancement on the slope images to obtain high-resolution images; inputting the high-resolution images into a feature extraction model to extract image features, and optimizing the image features through multi-scale feature enhancement; and inputting the optimized image features into a target detection model to identify and locate cracks.
[0008] According to a second aspect of the present disclosure, a slope crack detection device based on super-resolution processing is provided. The device includes: an image acquisition module for simultaneously acquiring data from a target slope area using multiple image acquisition devices to obtain multimodal slope images; an image processing module for performing super-resolution processing and image detail enhancement on the slope images to obtain high-resolution images; a feature extraction module for inputting the high-resolution images into a feature extraction model to extract image features and optimizing the image features through multi-scale feature enhancement; and a crack detection module for inputting the optimized image features into a target detection model to identify and locate cracks.
[0009] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory storing computer-readable instructions, which, when executed by the processor, implement the slope crack detection method based on super-resolution processing described above.
[0010] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the slope crack detection method based on super-resolution processing as described above.
[0011] The technical solutions provided in this disclosure may have the following beneficial effects:
[0012] The slope crack detection method based on super-resolution processing in the exemplary embodiments of this disclosure has several advantages. Firstly, by simultaneously acquiring data from the target slope area using multiple image acquisition devices, multimodal slope images can be obtained simultaneously, thereby capturing multi-dimensional information about the slope. Simultaneous acquisition of multimodal images reduces spatiotemporal inconsistencies between data, providing a high-quality input foundation for subsequent processing and solving the problem of insufficient information from a single modality. Secondly, super-resolution processing of the multimodal slope images significantly improves image clarity by enhancing spatial resolution. Detail enhancement techniques further highlight the edges and textures of cracks, allowing even subtle cracks that are difficult to identify in the image to be well represented, thus laying a solid foundation for accurate crack detection and avoiding feature loss due to insufficient resolution. Thirdly, multi-scale feature enhancement techniques further optimize feature representation at different scales, ensuring that the extracted image features contain both global information and capture subtle crack features, effectively suppressing background noise interference and improving the accuracy and usability of the features. Furthermore, by inputting the optimized image features into the target detection model, the detection algorithm can automatically identify and locate cracks, avoiding the inefficiency of manual identification and ensuring high reliability of the detection results. This, in turn, improves the accuracy and efficiency of slope crack detection to a certain extent.
[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0015] Figure 1 The illustration schematically shows a flowchart of a slope crack detection method based on super-resolution processing according to some embodiments of the present disclosure.
[0016] Figure 2 The schematic diagram illustrates a process diagram for multi-scale feature enhancement according to some embodiments of the present disclosure.
[0017] Figure 3 The diagram illustrates a block diagram of a slope crack detection device based on super-resolution processing according to some embodiments of the present disclosure.
[0018] Figure 4The schematic diagram illustrates the structural schematic of a computer system of an electronic device according to some embodiments of the present disclosure.
[0019] Figure 5 A schematic diagram of a computer-readable storage medium according to some embodiments of the present disclosure is shown.
[0020] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0021] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.
[0022] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0023] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0024] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.
[0025] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0026] Furthermore, the accompanying drawings are for illustrative purposes only and are not necessarily drawn to scale. The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0027] Figure 1 The illustration schematically shows a flowchart of a slope crack detection method based on super-resolution processing according to some embodiments of the present disclosure. Reference Figure 1 As shown, the slope crack detection method based on super-resolution processing may include the following steps:
[0028] Step S110: Simultaneously acquire data from the target slope area using multiple image acquisition devices to obtain multimodal slope images;
[0029] Step S120: Perform super-resolution processing and image detail enhancement on the slope image to obtain a high-resolution image;
[0030] Step S130: Input the high-resolution image into the feature extraction model to extract image features, and optimize the image features through multi-scale feature enhancement;
[0031] Step S140: Input the optimized image features into the target detection model to identify and locate the cracks.
[0032] The following section provides a further explanation of the slope crack detection method based on super-resolution processing.
[0033] Step S110: Simultaneously acquire data of the target slope area using multiple image acquisition devices to obtain multimodal slope images.
[0034] Image acquisition equipment can refer to sensor devices or imaging equipment capable of acquiring data of a target slope area and generating image outputs, such as visible light cameras, infrared imaging devices, point cloud sensors, depth cameras, UAV-mounted devices, and other suitable image acquisition equipment. The target slope area can refer to a naturally or artificially formed sloping surface portion that requires crack detection, monitoring, or assessment in engineering applications. Multimodal slope images can represent various forms of image data acquired through different types of image acquisition equipment, such as RGB images, infrared images, depth images, point cloud images, and multispectral images. These image modalities reflect different physical properties of the slope, such as color, depth, temperature, texture, and three-dimensional morphology.
[0035] Step S120: Perform super-resolution processing and image detail enhancement on the slope image to obtain a high-resolution image.
[0036] Super-resolution processing can be defined as the process of transforming a low-resolution image into a high-resolution image using algorithms. Image detail enhancement can be defined as the process of highlighting key features in an image, such as edges, textures, and contours, using algorithms to make the details of the target area clearer. For example, super-resolution processing can be implemented using generative adversarial networks (GANs) or convolutional neural networks (CNNs) to transform acquired multimodal slope images from low resolution to high resolution and supplement detailed information in crack areas. Through image detail enhancement operations, edge detection algorithms, texture enhancement techniques, and high-pass filtering are used to highlight the edges, width, and depth information of cracks, significantly improving the visibility of crack areas in the image and generating slope images with high spatial resolution and enhanced details.
[0037] Step S130: Input the high-resolution image into the feature extraction model to extract image features, and optimize the image features through multi-scale feature enhancement.
[0038] Feature extraction models can represent computational methods or deep learning structures used to extract meaningful features from input images, such as convolutional neural network models, residual network models, and Transformer models. Multi-scale feature enhancement can represent a method that analyzes and optimizes image features at different scales, aiming to highlight key information about cracks. Cracks typically have different widths, depths, and orientations. Multi-scale feature enhancement techniques capture the representation of these features at different resolutions, enabling a more accurate description of the complex morphological information of cracks during feature extraction.
[0039] For example, the feature extraction model can employ a convolutional neural network to extract key features from high-resolution images through multi-layer convolutional operations, including the edges, textures, orientations, and morphological information of cracks. The extracted image features are optimized using multi-scale feature enhancement techniques, which utilize convolutional kernels of different sizes to capture the behavior of cracks across multiple scales. This can be combined with a feature pyramid network to fuse local and global information, enhancing the ability to identify cracks of different widths, depths, and orientations, while suppressing background noise interference with feature extraction, ensuring the accuracy and comprehensiveness of image features.
[0040] Step S140: Input the optimized image features into the target detection model to identify and locate the cracks.
[0041] Among them, the target detection model can be represented as a computational model for locating and identifying specific targets in an input image. It completes the automatic identification and localization of targets by extracting features from the image, generating candidate regions, classifying target types, and regressing target locations.
[0042] For example, the optimized image features are input into the target detection model, which identifies and locates cracks through a comprehensive process of candidate region generation, classification, and location regression. The target detection model can use a region proposal network to generate candidate regions containing cracks, a classification network to determine the crack category within these candidate regions, and a regression network to optimize the crack's coordinates, outputting the precise location and category label of the crack. This crack identification and localization process can adapt to slope scenarios of varying sizes and complex backgrounds, providing reliable detection results for subsequent slope monitoring and engineering decisions.
[0043] Next, the above-described slope crack detection method based on super-resolution processing will be further described in other embodiments of this disclosure.
[0044] In some embodiments, the target slope area is simultaneously acquired using multiple image acquisition devices to obtain multimodal slope images. Specifically, the following steps are included: simultaneously acquiring image data of the target slope area using infrared imaging devices, lidar devices, and depth imaging devices to obtain multimodal raw image data containing temperature, depth, and spatial information; and performing denoising and standardization processing on the multimodal raw image data to obtain multimodal slope images.
[0045] Specifically, infrared imaging equipment can acquire thermal distribution information of slope areas, reflecting temperature gradient changes and indirectly reflecting thermal characteristics such as cracks and crack propagation. LiDAR equipment can collect three-dimensional spatial information of the target slope area, providing the geometry, surface irregularities, and spatial distribution of potential cracks on the slope surface. Depth imaging equipment provides depth information of objects, i.e., the distance from each pixel to the camera, thus providing more spatial dimensions to the image. Especially when detecting complex or minute cracks, depth information can significantly improve the structural resolution of the image. In practical implementation, infrared imaging equipment can use thermal infrared sensors, effectively capturing temperature distribution images even in no-light or low-light conditions. The acquired infrared images contain differences in radiated heat from different object surfaces; cracks, faults, and other locations often exhibit different temperature gradients. LiDAR equipment can utilize laser scanning principles, emitting laser beams and receiving reflected signals to accurately acquire the three-dimensional coordinates of the target slope area. Depth imaging equipment can use sensors to emit beams and receive their reflected signals, measuring the depth information of each point on the object's surface based on the structured light or time-of-flight principles of depth sensors.
[0046] In addition, when denoising and standardizing multimodal raw image data, denoising algorithms based on autoencoders and standardization algorithms based on adaptive modal fusion can be used.
[0047] Specifically, the process of denoising multimodal raw image data using an autoencoder-based denoising algorithm includes:
[0048] An autoencoder (AE) is an unsupervised learning model that can be applied to tasks such as feature learning and image denoising. An autoencoder maps a noisy image to a denoised, clear image. The model structure consists of an encoder and a decoder. Assume the input multimodal original image data is... It is an image with noise, in which and Representing pixel coordinates, the denoising model adds random noise. To the original image What was obtained:
[0049]
[0050] in, This represents the raw image data with noise. Indicates the original, clear image. The noise term can be composed of Gaussian noise, salt-and-pepper noise, etc., representing random interference in the image.
[0051] In the encoder stage, the encoder portion of the autoencoder is used to process the multimodal raw image data. Mapping to latent space This latent space represents the feature information contained in the image. The encoder consists of a multi-layer neural network, using a weight matrix... The input image is subjected to convolution or fully connected transformation to obtain the latent feature vector. The specific process can be represented as follows:
[0052]
[0053] in, The characteristic representation of the latent space, This represents the activation function of the encoder. This represents the weight matrix in the encoder, responsible for transforming image features. The encoder's role is to capture important feature information by reducing the dimensionality of the input image, providing a meaningful latent representation for the decoding process.
[0054] In the decoder stage, the decoder receives the output of the encoder. The image is then reconstructed into a denoised image using a decoder network. The decoder uses the weight matrix. The process of inversely mapping the representation of the latent space back to the image space can be represented as:
[0055]
[0056] in, This represents the image after denoising. This represents the activation function of the decoder. This represents the weight matrix in the decoder, used to map latent features back to the image space. The goal of the decoder is to transform the latent representation... To restore a clear image, which means eliminating noise in the input image as much as possible.
[0057] Furthermore, in a denoising autoencoder, the training objective is to minimize the loss function. To optimize model parameters. Loss function. The mean squared error can be constructed using the following method:
[0058]
[0059] in, This represents the number of pixels in the image. The optimization objective of this loss function is to minimize the difference between the original image and the denoised image, and to update the model parameters through the backpropagation algorithm.
[0060] Furthermore, the process of standardizing the denoised multimodal original image data using the standardization algorithm based on adaptive modality fusion includes:
[0061] First, the denoised multimodal image data is calibrated. Let the denoised infrared image, point cloud image, and depth image be respectively... , and .
[0062] Then, each modal image is normalized, mapping its pixel values to a standard range. To standardize the scale, the normalization formula is as follows:
[0063]
[0064]
[0065]
[0066] in, , These represent the minimum and maximum pixel values of the infrared image, respectively. , These represent the minimum and maximum pixel values of the point cloud image, respectively. , These represent the minimum and maximum pixel values of the depth image, respectively.
[0067] Next, to adjust the contribution of each modal image in the final fusion based on its quality, such as sharpness and contrast, an adaptive modal fusion (AMF) technique is employed. The quality of each modal image is determined by a quality metric function. Given, among which These represent infrared image, point cloud image, and depth image, respectively. Assume a quality metric. It is an index calculated based on factors such as image sharpness, contrast, and noise. Its specific calculation method can be based on features such as image gradient and information entropy. Adaptive weights for each modality of image. According to their respective quality metrics The calculation is performed using the following formula:
[0068]
[0069] in: Representing modes Adaptive weights, Representing modes Quality measurement, This represents the sum of all modal quality metrics, used to normalize the weights. Through this calculation, the higher the quality metric of each modal image, the greater its corresponding weight, ensuring that high-quality modal images contribute more in the subsequent fusion process.
[0070] Finally, in adaptive weights Once determined, the weights are used to perform weighted fusion of the normalized modal images to obtain a multimodal slope image. The fusion formula can be expressed as:
[0071]
[0072] in: Represents the normalized mode The pixel values of the image, Representing modes Adaptive weights, This represents the fused, standardized image, i.e., the multimodal slope image. This weighted fusion process fully considers the quality and features of each modal image, ensuring that each modal image can exert maximum efficiency during the fusion process.
[0073] In some embodiments, reference Figure 2 As shown, super-resolution processing and image detail enhancement are performed on the slope image to obtain a high-resolution image. The specific steps include the following:
[0074] Step S210: Input the slope image into the generative adversarial network for super-resolution processing to generate a resolution-optimized image.
[0075] Step S220: Perform a convolution operation on the resolution-optimized image to obtain a feature-enhanced image.
[0076] Step S230: Perform image fusion on the multimodal feature-enhanced images to obtain a high-resolution image.
[0077] Specifically, firstly, the slope image is input into a generative adversarial network (GAN) for super-resolution processing to generate a resolution-optimized image. This GAN learns the mapping relationship between low-resolution and high-resolution images, achieving detail restoration and sharpness enhancement of the slope image. Next, convolution operations are performed on the resolution-optimized image to extract high-order features and enhance details. This operation uses convolution kernels to filter the image, enhancing local features through layer-by-layer convolution operations and improving detail representation. Finally, the multimodal feature-enhanced images are fused to obtain a high-resolution image. Image fusion integrates information from different modalities, such as infrared, depth, and point cloud data, using a weighted approach to further improve image quality and accuracy, ensuring that the final high-resolution image fully displays the characteristic information of slope cracks.
[0078] In some embodiments, the slope image is input into a generative adversarial network (GAN) for super-resolution processing to generate a resolution-optimized image. Specifically, this includes the following technical steps: inputting the slope image into the generator of the GAN for feature extraction and nonlinear mapping to obtain a first optimized image; inputting the first optimized image into a discriminator for similarity determination to obtain an evaluation value corresponding to the first optimized image; and optimizing the parameters of the generator and discriminator based on the generator loss function and the discriminator loss function to generate the resolution-optimized image. The generator loss function is expressed as: , Represents the generator loss function. This represents the first optimized image. This represents the data distribution of the slope image. The evaluation value is represented by: The discriminator loss function is expressed as: , This represents the discriminator loss function. Represents a true high-resolution image. This represents the discriminator's evaluation value for a true high-resolution image. Data distribution representing a true high-resolution image, Represents a slope image. Represents a generator. This represents the discriminator.
[0079] Specifically, firstly, the slope image is input into the generator of a generative adversarial network (GAN) for feature extraction and nonlinear mapping. During this process, the generator processes the input slope image through a series of convolutional layers, activation layers, and other neural network layers, extracting high-dimensional features and transforming them into an approximately high-resolution image—the first optimized image—through nonlinear mapping. The generator's task is to learn how to transform a low-resolution input image into a clearer, more detailed high-resolution image. Next, the first optimized image is input into a discriminator for similarity assessment to obtain an evaluation value. The discriminator determines whether the input image closely approximates the real high-resolution image. In this step, the discriminator analyzes the similarity between the input image and the real high-resolution image based on various features and outputs an evaluation value. This evaluation value guides the generator to further optimize, aiming to improve the realism and clarity of the generated image. If the discriminator deems the generated image significantly different from the real image, the generator needs to adjust its parameters to improve image quality. Furthermore, the generator loss function and the discriminator loss function are fed back based on the quality of the current generated image, thereby optimizing the parameters of the generator and discriminator. This optimization process, based on the loss functions of the generator and discriminator, gradually adjusts the network weights using gradient descent or other optimization algorithms to reduce the difference between the generated image and the real image. Through these technical steps, the generator can produce resolution-optimized images, completing the super-resolution image processing task and effectively improving the image resolution, allowing potential cracks, textures, and other details in slope images to be presented more clearly.
[0080] In some embodiments, a convolution operation is performed on the resolution-optimized image to obtain a feature-enhanced image, specifically including the following technical steps:
[0081] The first step, according to:
[0082]
[0083] Edge detection is performed on the resolution-optimized image to obtain an edge image, where, Represents the edge image. Indicates the resolution-optimized image at position Pixel value at that location, This represents the gradient of the resolution-optimized image in the horizontal direction. This represents the gradient of the resolution-optimized image in the vertical direction. Specifically, the image gradient is calculated using difference operations in both the horizontal and vertical directions. These gradients reflect the degree of change in pixel values within the image, thus highlighting edge features. The gradient values are derived using a synthesis formula and further squared to obtain the gradient magnitude for each pixel. The final edge image is the result of this gradient magnitude image, representing regions of significant change in the image, thereby effectively identifying edge components.
[0084] The second step is to binarize the edge image to obtain a mask image. Each pixel value in the edge image is processed according to a set threshold. If the gradient magnitude is greater than the threshold, the corresponding pixel is set to 1; if the gradient magnitude is less than the threshold, the corresponding pixel is set to 0, thus transforming the edge image into a mask image containing only edge and non-edge regions.
[0085] The third step, according to:
[0086]
[0087] A weighted convolution operation is performed on the resolution-optimized image to obtain the convolution result, where, This indicates the position of the output image of the convolution operation. pixel values, and These represent the convolution kernels at... direction and Half width of the direction, Represents the convolution kernel In position The value at the specified location. The convolution result is obtained by locally convolving the convolution kernel with the resolution-optimized image. The convolution kernel is a matrix that performs a weighted product operation with each pixel of the input image and its surrounding neighborhood to generate a new pixel value. The size and shape of the convolution kernel can be adjusted according to actual needs. In this process, specific features in the image, such as edges and corners, are highlighted by weighting the convolution kernel.
[0088] The fourth step involves multiplying the convolution result pixel-by-pixel with the mask image to obtain the feature-enhanced image. In this process, the mask image plays a selective enhancement role, preserving and enhancing only the regions with a pixel value of 1 in the mask image, while leaving the regions with a pixel value of 0 unprocessed. This operation effectively focuses on edge regions in the image, enhancing the detail in these areas, thereby improving image quality and resolution.
[0089] In some embodiments, the multimodal feature-enhanced images are fused to obtain a high-resolution image, specifically including the following technical steps:
[0090] First, according to:
[0091]
[0092] Spatial alignment is performed on the feature-enhanced infrared image, the feature-enhanced point cloud image, and the feature-enhanced depth image. Indicates image alignment. , and These represent the spatial transformation matrices of the feature-enhanced infrared image, the feature-enhanced point cloud image, and the feature-enhanced depth image, respectively. , and These represent the pixel values of the feature-enhanced infrared image, the feature-enhanced point cloud image, and the feature-enhanced depth image, respectively. and These represent the width and height of the image, respectively. Specifically, during spatial alignment, the images of different modalities are first adjusted using a preset spatial transformation matrix so that pixels correspond to the same physical location. Pixel values from the infrared image, point cloud image, and depth image undergo coordinate mapping and interpolation operations using their respective spatial transformation matrices, ensuring that pixel values at the same location in the aligned image reflect multiple information from different modalities. Ultimately, the generated aligned image contains multimodal information from the infrared image, point cloud image, and depth image, and this information is accurately located at the same spatial position, ensuring the effectiveness of image fusion.
[0093] Then, the aligned multimodal feature-enhanced images are weighted and fused to obtain a high-resolution image. Specifically, in the weighted fusion process, each modality's image is processed using different weighting coefficients, which are determined based on the quality, reliability, and contribution of each modality's image to the final target. Using the aligned images, the pixel values of each modality are summed according to their corresponding weighting coefficients to obtain the fused pixel values. This weighted fusion method effectively reflects the contribution of each modality during the fusion process, preserving the feature advantages of each modality while eliminating potential noise or missing information from a single modality.
[0094] In some embodiments, a high-resolution image is input into a feature extraction model to extract image features, and the image features are optimized through multi-scale feature enhancement. Specifically, the technical steps include: inputting a high-resolution image into a feature extraction model and extracting image features through convolution operations; performing multi-scale convolution operations on the image features to obtain multi-scale enhanced features; and optimizing the multi-scale enhanced features based on an attention mechanism.
[0095] Specifically, firstly, a high-resolution image is input into a feature extraction model, where convolution operations are performed to extract image features. The convolution operation utilizes convolution kernels to extract local textures, edges, and other key features from the image, particularly shape and texture information related to cracks. Next, multi-scale convolution operations are performed on the extracted image features. By using convolution kernels of different sizes, the model can simultaneously capture both small crack features and large crack regions in the image. Multi-scale convolution operations enable the network to effectively learn image features at multiple scales, thereby improving the detection capability for small cracks. Finally, the multi-scale enhanced features are optimized based on an attention mechanism. The attention mechanism learns the importance of features and dynamically adjusts the weights of specific regions in the image. By focusing on key crack regions, the attention mechanism effectively reduces interfering information, improves crack detection accuracy, and enhances detection robustness.
[0096] In some embodiments, multi-scale convolution operations are performed on image features to obtain multi-scale enhanced features, specifically including the following technical steps:
[0097] according to:
[0098]
[0099] Multi-scale convolution operations are performed on image features to obtain multi-scale enhanced features, where, Indicates multi-scale enhancement features, Indicates the weighting coefficient. The scale is represented as convolution kernel, and Indicates the position of the convolution kernel. Representing scale The bias value, Indicates the scale quantity. This represents the activation function. Indicates image features at location The pixel value.
[0100] In some embodiments, feature optimization of multi-scale enhanced features based on an attention mechanism may specifically include the following technical steps:
[0101] First, channel attention weights are calculated based on the given multi-scale enhancement features. Global average pooling is used to compress the feature information of each channel into a scalar. ,in The channel index of the feature map is calculated using the following formula:
[0102]
[0103] in, and These represent the height and width of the image, respectively. Indicates the location The first The pixel value of the channel. A scalar obtained through global average pooling. It reflects the overall information of the channel.
[0104] Next, channel attention weights are generated. A fully connected network layer is used to perform a non-linear transformation on the features of each channel to generate channel attention weights. This process can be represented by the following formula:
[0105]
[0106] in, and These represent the weights and biases of the fully connected layer, respectively. This represents the activation function of the fully connected layer, which can be ReLU or other non-linear activation functions. Attention weights. Used to weight and adjust the features of each channel.
[0107] Next, based on the calculated channel attention weights The feature map of each channel is weighted to optimize the importance of each channel. The weighting process can be represented as:
[0108]
[0109] in, Indicates the location The feature value is optimized through the channel attention mechanism.
[0110] In addition, for each location Spatial attention weights are generated by calculating the comprehensive information of all channels in the feature map. The global information of all channels in the feature map is calculated using the following formula:
[0111]
[0112] in, This represents the number of channels in the feature map. This represents the feature value after channel attention optimization. This represents a non-linear activation function. The spatial attention weights are obtained by applying a non-linear activation function to the sum of all channels. .
[0113] Finally, the spatial attention weights were calculated. The image features at each location are weighted, and the weighting process can be represented as follows:
[0114]
[0115] in, This represents the optimized image features. This represents the feature values optimized using channel attention. Spatial attention is used to weight important regions in the image, enhancing the saliency of key features.
[0116] In some embodiments, the optimized image features are input into a target detection model to identify and locate cracks. Specifically, this includes the following technical steps: inputting the optimized image features into a target detection model, extracting crack candidate regions based on a region proposal network; classifying and regressing the crack candidate regions to determine the crack candidate regions containing cracks and the crack locations. Specifically:
[0117] The first step is to optimize the image features. The input is fed into the Region Proposal Network (RPN) of the object detection model. The purpose of the RPN is to generate multiple potential crack candidate regions from image features, specifically including the following steps:
[0118] For each sliding window position A set of possible candidate boxes is generated through convolution operations, and the size of the candidate boxes is determined by... The width and height are represented by these four parameters. The position of the candidate box is determined by these four parameters. express, and The coordinates of the center of the candidate box. and Here, represents the width and height of the candidate box. The process of generating the candidate box can be represented by the following formula:
[0119]
[0120] in, Indicates the first The coordinates and dimensions of each candidate box. These represent the coordinates and size of the candidate box, respectively.
[0121] RPN uses fully connected layers to integrate convolutional features. The classification and regression outputs are mapped to each candidate region. The generated candidate regions are then used for subsequent classification and regression operations. The generated candidate regions include not only regions that may contain cracks, but also non-crack regions.
[0122] The second step involves classifying and regressing the crack candidate regions generated by the RPN to further determine which candidate regions contain cracks and their specific locations. The classification process includes binary classification of each candidate region as either cracked or non-cracked, while the regression process adjusts the position of the candidate boxes to more accurately pinpoint the crack location.
[0123] The classification process, based on the differences between the features of the candidate regions and the background, obtains the class probability through the Softmax activation function. Let... Indicates the first The classification results of the candidate boxes, ,in This indicates that the area does not contain cracks. This indicates that the area contains cracks. The classification process is as follows:
[0124]
[0125] in, This represents the weight matrix of the classification layer. For bias terms, This represents the candidate region features input to the classification layer. This represents the activation function of the classification layer.
[0126] The regression process can optimize the accuracy of candidate regions by predicting the positional changes of each candidate region, i.e., the offset relative to the initial bounding box. Let... For the first The regression prediction values for each candidate region represent the changes in horizontal and vertical offsets, as well as width and height, relative to the initial bounding box. The regression operation can be calculated using the following formula:
[0127]
[0128] in, Indicating a return to the internet, The features representing the candidate regions. The regressed candidate boxes can be represented by the following formula:
[0129]
[0130] in, Indicates the first after regression A precise candidate bounding box coordinate and size, This represents the center coordinates and dimensions after regression.
[0131] The third step involves determining candidate regions containing cracks and their precise locations based on non-maximum suppression (NMS). Since multiple overlapping boxes may exist within the candidate boxes generated by the region proposal network, NMS is used to remove redundant boxes, retaining only the most representative ones. Specifically, for each candidate box, its overlap with other boxes is calculated, i.e., the Intersection over Union (IoU), and the box with the highest confidence is retained. The calculation of the IoU can be expressed as follows:
[0132]
[0133] in, Indicates the first and the The intersection-union ratio of candidate boxes, Indicates the first after regression The positions and sizes of the candidate boxes are determined. Non-maximum suppression is used to remove candidate regions with high overlap and low confidence.
[0134] Through the above steps, the region containing the crack and its precise location are finally obtained. These precisely located areas are the final detection results of the cracks.
[0135] Furthermore, in other embodiments of this disclosure, after crack identification and localization of multimodal slope images, more detailed and dynamic visualization and interactive display can be performed based on the crack information of the multimodal data to enhance the intelligence and accuracy of crack monitoring.
[0136] Specifically, by fusing multimodal data from various types, such as infrared images, point cloud images, and depth images, a unified 3D scene model can be constructed. This model can then be used to display the spatial distribution, morphological characteristics, and evolutionary trends of cracks in real time. This 3D scene not only presents the surface morphology of cracks but also displays their depth characteristics based on depth information, providing more accurate spatial location and depth profile analysis of cracks. Furthermore, combining crack detection results with relevant risk assessment data, the visualization system can dynamically generate a risk heatmap of cracks and display different colors or textures on the 3D visualization interface according to different risk levels, allowing users to quickly identify the severity of cracks and potential safety hazards. In addition, users can adjust the viewing angle, zoom level, and transparency in real time through the interactive interface to view the location and morphology of cracks from all angles. Simultaneously, by combining time-series data, users can dynamically play back the crack change process at different time points, accurately tracking the crack's expansion, development trend, and historical changes, thereby predicting the crack's possible future trajectory.
[0137] During the visualization process, each detection result and real-time data of the crack can be correlated with relevant engineering information and environmental data, such as soil moisture, rainfall, and temperature changes, to generate comprehensive monitoring reports or dynamic charts. This helps engineering monitoring personnel analyze the causes, patterns of change, and impacts on the engineering structure of the cracks from multiple perspectives. Furthermore, the visualization system also features intelligent alarm functionality. When the risk level of a crack reaches a set threshold, the system will automatically generate an alarm and push it to the monitoring personnel's mobile devices or computer terminals in real time, ensuring timely response.
[0138] The slope crack detection method based on super-resolution processing in the above embodiments firstly acquires data from the target slope area simultaneously using multiple image acquisition devices, obtaining multimodal image data containing different information. Specifically, infrared imaging, lidar, and depth imaging devices provide temperature, depth, and spatial information, respectively. These different modal image data are spatially aligned and fused, comprehensively utilizing the advantages of various sensors to improve the accuracy and adaptability of crack detection. Furthermore, the acquired multimodal image data undergoes super-resolution processing. This step uses a generative adversarial network (GAN) to reconstruct the low-resolution image at a high resolution, significantly improving image details, especially the fine structure of the crack area. The generator in the GAN is responsible for generating optimized images through nonlinear mapping, while the discriminator evaluates the realism of the generated images. Through adversarial training between the generator and the discriminator, not only is the image resolution enhanced, but the high quality of the generated images is also ensured. Super-resolution processing not only improves image clarity but also provides more refined image information for subsequent crack feature extraction.
[0139] In terms of image detail enhancement, convolution operations are used to process the optimized image to further highlight key features such as cracks. By performing edge detection on the image and combining it with weighted convolution operations, edge information in the slope image can be extracted and fused with other image features to enhance crack identifiability. Edge detection utilizes the gradient information of the image to extract edge contours, while convolution operations further enhance the image's detail representation, making cracks more prominent in the image. By processing image features with convolution kernels of different scales, crack features at different scales can be captured, effectively representing the features of both small and large cracks. This technique not only improves the extraction effect of crack features but also enhances the image's adaptability to different crack sizes. Combining an attention mechanism to further optimize the features after multi-scale convolution ensures the prominence of important features and suppresses the interference of irrelevant features. This step significantly improves the accuracy and recall of crack detection. Finally, by inputting the optimized image features into the target detection model for crack identification and localization, accurate detection of slope cracks is ultimately achieved. The target detection model uses a region proposal network to generate crack candidate regions and accurately locates these regions through classification and regression operations. The classification step ensures that each candidate region contains a crack, while the regression step fine-tunes the position of the candidate boxes to ensure the accuracy of crack localization.
[0140] It should be noted that although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0141] Next, in this embodiment of the disclosure, a slope crack detection device based on super-resolution processing is also provided, with reference to... Figure 3 As shown, the slope crack detection device 300 based on super-resolution processing can be composed of an image acquisition module 301, an image processing module 302, a feature extraction module 303, and a crack detection module 304. Specifically: the image acquisition module 301 can be used to simultaneously acquire data from the target slope area using multiple image acquisition devices to obtain multimodal slope images; the image processing module 302 can be used to perform super-resolution processing and image detail enhancement on the slope images to obtain high-resolution images; the feature extraction module 303 can be used to input the high-resolution image into a feature extraction model to extract image features and optimize the image features through multi-scale feature enhancement; and the crack detection module 304 can be used to input the optimized image features into a target detection model to identify and locate cracks.
[0142] It should be noted that the specific details of each part of the slope crack detection device based on super-resolution processing have been described in detail in the implementation method section of the slope crack detection method based on super-resolution processing. For any undisclosed details, please refer to the implementation method section, and therefore will not be repeated here.
[0143] Furthermore, in an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described slope crack detection method based on super-resolution processing is also provided.
[0144] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be embodied in the following forms: a completely hardware embodiment, a completely software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0145] The following reference Figure 4 To describe an electronic device 400 according to such an embodiment of the present disclosure. Figure 4 The electronic device 400 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0146] like Figure 4 As shown, the electronic device 400 is manifested in the form of a general-purpose computing device. The components of the electronic device 400 may include, but are not limited to: at least one processing unit 410, at least one storage unit 420, a bus 430 connecting different system components (including storage unit 420 and processing unit 410), and a display unit 440.
[0147] The storage unit stores program code, which can be executed by the processing unit 410, causing the processing unit 410 to perform the steps described in the "Exemplary Methods" section above according to various exemplary embodiments of this disclosure.
[0148] Storage unit 420 may include readable media in the form of volatile storage units, such as random access memory (RAM) 421 and / or cache memory 422, and may further include read-only memory (ROM) 423.
[0149] Storage unit 420 may also include a program / utility 424 having a set (at least one) of program modules 425, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0150] Bus 430 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0151] Electronic device 400 can also communicate with one or more external devices 470 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 400, and / or with any device that enables electronic device 400 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 450. Furthermore, electronic device 400 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 460. As shown, network adapter 460 communicates with other modules of electronic device 400 via bus 430. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0152] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0153] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure.
[0154] refer to Figure 5As shown, a program product 500 for implementing the above-described slope crack detection method based on super-resolution processing, according to an embodiment of the present disclosure, is described. It may employ a portable compact disk read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0155] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory, read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0156] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0157] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, electromagnetic waves, etc., or any suitable combination thereof.
[0158] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0159] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0160] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0161] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0162] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A slope crack detection method based on super-resolution processing, characterized in that, include: Multimodal slope images are obtained by simultaneously acquiring data from the target slope area using multiple image acquisition devices. The slope image is input into the generator of the generative adversarial network for feature extraction and nonlinear mapping to obtain the first optimized image; The first optimized image is input into the discriminator for similarity determination to obtain an evaluation value corresponding to the first optimized image. The evaluation value is used to guide the generator optimization. Based on the generator loss function and the discriminator loss function, the parameters of the generator and the discriminator are optimized to generate a resolution-optimized image. Edge detection is performed on the resolution-optimized image to obtain an edge image; The edge image is binarized to obtain a mask image; A weighted convolution operation is performed on the resolution-optimized image to obtain the convolution result; the convolution result is then multiplied pixel-by-pixel with the mask image to obtain the feature-enhanced image. Spatial alignment is performed on the feature-enhanced infrared image, feature-enhanced point cloud image, and feature-enhanced depth image based on the spatial transformation matrix; The aligned multimodal feature-enhanced images are then weighted and fused to obtain a high-resolution image; The high-resolution image is input into the feature extraction model, and image features are extracted through convolution operations. Based on the activation function and convolution kernel, multi-scale convolution operations are performed on the image features to obtain multi-scale enhanced features. Based on the multi-scale enhancement features, the feature information of each channel is compressed into a scalar through global average pooling; the scalar reflects the global information of the channel; and a fully connected network is used to perform a nonlinear transformation on the features of each channel to generate channel attention weights. Based on the channel attention weights, the feature maps of each channel are weighted to optimize the importance of each channel; Spatial attention weights are generated by calculating the comprehensive information of all channels of the feature map; The spatial attention weights are used to weight the image features at each location to enhance the saliency of key features in the image. The optimized image features are input into the target detection model to identify and locate the cracks; The process of inputting the optimized image features into the target detection model to identify and locate cracks includes: Optimized image features The region proposal network, input to the object detection model, is used for each sliding window position. Candidate boxes are generated through convolution operations. The process of generating candidate boxes is represented by the following formula: in, Indicates the first The coordinates and dimensions of each candidate box. These represent the coordinates and size of the candidate boxes, respectively; the region proposal network uses fully connected layers to integrate convolutional features. Classification and regression outputs mapped to each candidate region; Crack candidate regions generated by the region proposal network are classified and regressed; the classification process is based on the differences between the features of the candidate regions and the background, and the class probability is obtained through the Softmax activation function; let... Indicates the first The classification results for each candidate box are as follows: in, This represents the weight matrix of the classification layer. For bias terms, This represents the candidate region features input to the classification layer. This represents the activation function of the classification layer; set up For the first The regression prediction values for each candidate region represent the changes in horizontal and vertical offsets, as well as width and height, relative to the initial bounding box; the regression operation is calculated using the following formula: in, Indicating a return to the internet, The features representing the candidate regions; the candidate bounding boxes after regression are represented by the following formula: in, Indicates the first after regression A precise candidate bounding box coordinate and size, Indicates the center coordinates and dimensions after regression; Candidate regions containing cracks and their precise locations are determined based on the nonmaximum suppression method: For each candidate box, its intersection-over-union (IoU) ratio with other boxes is calculated. The calculation process for the IoU ratio is as follows: in, Indicates the first and the The intersection-union ratio of candidate boxes, Indicates the first after regression The position and size of each candidate box.
2. The slope crack detection method based on super-resolution processing according to claim 1, characterized in that, The process of simultaneously acquiring data from the target slope area using multiple image acquisition devices to obtain multimodal slope images includes: Image data of the target slope area is acquired simultaneously by infrared imaging equipment, lidar equipment and depth imaging equipment to obtain multimodal raw image data containing temperature, depth and spatial information; The original multimodal image data is denoised and standardized to obtain the multimodal slope image.
3. A slope crack detection device based on super-resolution processing, used to implement the slope crack detection method based on super-resolution processing as described in any one of claims 1 to 2, characterized in that, The device includes: The image acquisition module is used to simultaneously acquire data from the target slope area using multiple image acquisition devices to obtain multimodal slope images; The image processing module is used to perform super-resolution processing and image detail enhancement on the slope image to obtain a high-resolution image; The feature extraction module is used to input the high-resolution image into the feature extraction model, extract image features, and optimize the image features through multi-scale feature enhancement; The crack detection module is used to input the optimized image features into the target detection model to identify and locate cracks.