Crack extraction method and system and storage medium
By using the HRNet deep learning network and wavelet downsampling enhancement module, the problems of severe noise and incomplete multi-scale extraction in crack detection are solved, enabling refined crack extraction in complex backgrounds, improving crack recognition accuracy and system ease of use.
Patent Information
- Application Number
- CN202511052181.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies for crack detection suffer from problems such as severe noise, incomplete extraction of multi-scale cracks, and poor extraction results for cracks with varied and tortuous shapes.
We employ the HRNet deep learning network, combined with wavelet downsampling enhancement, morphological adaptive convolution, and an efficient multi-scale attention module, to construct a crack refinement extraction model. Through iterative training, we generate refined crack extraction results.
It improves the accuracy of crack extraction, effectively identifies cracks of different scales and complex shapes, reduces background interference, enhances the ability to perceive irregular crack boundaries, and provides a complete system solution from data acquisition to practical application.
Smart Images

Figure CN120931940A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of crack detection technology, and particularly relates to a crack extraction method, system and storage medium. Background Technology
[0002] With the continuous advancement of urbanization and infrastructure construction, infrastructure such as roads, bridges, and residential buildings are generally threatened by structural defects such as cracks during their long-term service due to environmental factors, material aging, or load changes. Cracks are not only one of the most common surface defects, but also an important indicator reflecting declining material performance, weakened durability, and even safety hazards. Therefore, efficient and accurate identification and segmentation of cracks are of significant practical importance for facility operation and maintenance management, life assessment, and disaster early warning.
[0003] Currently, crack detection mainly relies on manual inspection and traditional image processing methods. Manual methods are inefficient and highly subjective, making them unsuitable for large-scale detection needs. To address this, traditional image processing methods are applied in crack segmentation, such as crack detection methods based on thresholding, edge analysis, region analysis, and matching algorithms. However, while these methods can meet the needs of large-scale detection, they typically rely on manually defined crack features for extraction. This leads to poor recognition performance when dealing with complex background interference, lighting variations, or irregular crack shapes, especially with serious misjudgments at crack edges, narrow slits, and noisy regions.
[0004] In recent years, with the continuous development of deep learning methods, crack segmentation technology based on convolutional neural networks (CNNs) has made significant progress, resulting in numerous models for crack segmentation, such as U-Net, DeepCrack, SegNet, and DeepLabV3+. Based on this, a series of modules for crack feature enhancement, boundary awareness, and feature fusion have emerged to improve the model's crack segmentation accuracy. While the aforementioned deep learning-based networks and modules have achieved good results in crack segmentation, they still have the following shortcomings in practical applications: First, existing networks have certain difficulties in extracting cracks of different scales, and small-scale cracks are prone to missed detections or incomplete extraction. Second, in special environments with severe background noise interference and diverse crack morphologies, the network's insufficient anti-interference ability leads to severe noise, making it difficult to achieve refined crack extraction. Summary of the Invention
[0005] In view of this, the present invention provides a crack extraction method, system and storage medium to solve the problems of severe noise in crack extraction results, incomplete extraction of multi-scale cracks and poor extraction effect of cracks with tortuous and diverse morphologies.
[0006] To achieve the above objectives, in a first aspect, the technical solution of the present invention to solve the technical problem is to provide a crack extraction method, comprising: acquiring a publicly available crack dataset and collecting crack images, and constructing a sample dataset based on the publicly available crack dataset and crack images; creating an initial network model based on the HRNet deep learning network; inputting the crack sample dataset into the initial network model for iterative training to obtain a crack refinement extraction model based on optimal weights; inputting real-time crack data into the crack refinement extraction model and outputting crack extraction results.
[0007] Furthermore, the creation of the initial network model includes: the encoder is based on the HRNet multi-branch structure and includes multiple feature extraction modules and a wavelet downsampling enhancement module; the decoder includes multiple dual-stream feature fusion modules and a prediction module; wherein, the feature extraction module is used to perform upsampling, multi-scale feature extraction, multi-scale feature interaction, and feature identity mapping operations on the crack image data to obtain first feature information at different scales; the wavelet downsampling enhancement module is used to extract frequency domain information in the crack, downsample the first feature information, and supplement the first feature information with frequency domain information to form second feature information; the dual-stream feature fusion module is used to fuse the second feature information, perform upsampling and feature alignment, and output feature output results at a unified spatial resolution; the prediction module is used to generate classification results of crack pixels and background pixels based on the feature output results.
[0008] Furthermore, the feature extraction module includes: a basic feature extraction module, a morphological adaptive convolution module, and an efficient multi-scale attention module; wherein, the basic feature extraction module is constructed based on a residual structure, performs convolution, batch normalization, and nonlinear activation on the input image, and extracts the first feature information; the morphological adaptive convolution module, using the first feature information as input, perceives irregular crack boundaries and curve structures, and captures crack morphological changes and topological features in the spatial dimension; the efficient multi-scale attention module is used to weight and adjust the first feature information, enhance the attention to key regions, integrate spatial perception and channel perception, strengthen the expression of salient regions, and suppress background interference.
[0009] Furthermore, the operation of the morphological adaptive convolution module includes: predicting the offset parameter of each pixel on the first feature information using a lightweight convolution structure to guide the convolution kernel to adaptively sample; resampling the neighboring pixels at a specified position in the first feature information according to the offset to construct an adaptive sampling region; performing a weighted convolution operation on the adaptive sampling region to output the enhanced first feature information.
[0010] Furthermore, the operation of the efficient multi-scale attention module includes: using different pooling and convolution methods to achieve spatial alignment and feature integration at different scales for the first feature information, and obtaining attention weights of a unified dimension; and reconstructing the first feature information by channel or spatial dimension weighting according to the attention weights, thereby enhancing the expression of salient features and suppressing redundant information.
[0011] Furthermore, the operation of the wavelet downsampling enhancement module includes: performing a two-dimensional discrete wavelet transform on the first feature information to decompose it into a low-frequency component and multiple high-frequency components; extracting basic features with compression characteristics from the low-frequency component to achieve spatial downsampling of the first feature information; and weighted fusing the low-frequency features with some high-frequency features to form the second feature information.
[0012] Furthermore, the operation of the dual-stream feature fusion module includes: performing channel concatenation on the second feature information at different scales, generating channel attention weights through pooling, convolution, and activation functions, and performing channel-weighted enhancement on the concatenated feature map; performing 1x1 convolution on the second feature information of different network layers respectively, concatenating and activating them to generate spatial attention weights, and performing spatial weighting adjustment on the channel-enhanced second feature information to obtain the fused feature map.
[0013] Furthermore, the operation of the prediction module includes: upsampling and aligning the outputs of different branches and fusing them into a feature map of uniform size; using 1x1 convolution to classify and predict cracks and background pixels, and outputting a refined crack extraction result in the form of a binarized mask.
[0014] Secondly, the present invention also provides a crack extraction system, including a database construction module for acquiring publicly available crack datasets and collecting crack images, and constructing a sample dataset based on the publicly available crack datasets and crack images; an initial network model construction module for creating an initial network model based on the HRNet deep learning network; a network iterative training module for inputting the crack sample dataset into the initial network model for iterative training to obtain a crack refinement extraction model based on optimal weights; and an extraction module for inputting real-time crack data into the crack refinement extraction model and outputting crack extraction results.
[0015] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a crack extraction method.
[0016] Compared with the prior art, the crack extraction method, system, and storage medium provided by the present invention have the following beneficial effects: Based on the HRNet deep learning model, a refined crack extraction network comprising encoder and decoder modules was constructed. The encoder module utilizes the HRNet multi-branch structure for feature extraction and combines wavelet downsampling enhancement to obtain feature information at different scales. The decoder module, through dual-stream feature fusion and prediction, fuses features from different scales to generate refined crack extraction results. This achieves a complete workflow from data acquisition, preprocessing, network construction and training to crack segmentation, solving the problems of severe noise in crack extraction results under complex backgrounds, incomplete extraction of small cracks, and poor crack extraction performance with diverse and tortuous shapes. Simultaneously, it can effectively extract crack feature information at different scales and has good recognition performance for small and complex cracks, improving crack extraction accuracy. Furthermore, the introduction of a morphological adaptive convolution module and an efficient multi-scale attention module enhances the network's ability to perceive irregular crack boundaries and curved structures, strengthens the representation of salient regions, and reduces background interference. Finally, a complete system from data acquisition, annotation, network construction and training to practical application is provided, facilitating deployment and application in real-world engineering projects. Attached Figure Description
[0017] Figure 1 This is a flowchart of the crack extraction method provided in the first embodiment of the present invention; Figure 2 A schematic diagram of the modules of the crack extraction method provided in the first embodiment of the present invention; Figure 3 This is a schematic diagram of the binary segmentation results for crack extraction from a portion of the experimental data. Figure 4 This is a schematic diagram of the crack extraction system provided in the second embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0019] like Figures 1-2 As shown, the first embodiment of this application provides a crack extraction method, which includes: S100: Obtain publicly available crack datasets and collect crack images, and construct a sample dataset based on the publicly available crack datasets and crack images. S200, based on the HRNet deep learning network, creates an initial network model; S300: Input the sample dataset of cracks into the initial network model for iterative training to obtain a crack refinement extraction model based on optimal weights; S400 inputs real-time crack data into the crack refinement extraction model and outputs crack extraction results; In this embodiment, firstly, a publicly available crack dataset is acquired, and high-resolution crack images are captured through photography. These high-resolution crack images are then combined with the publicly available crack dataset to create a multi-scene crack sample dataset. Next, based on the HRNet deep learning model, wavelet enhancement, feature fusion, attention mechanisms, and morphological adaptive convolution operations are integrated into the network to create an initial network model. Then, the initial model is iteratively trained using the multi-scene crack sample dataset to obtain a refined crack extraction model based on optimal weights. Finally, real-time crack data is input into the refined crack extraction model, and the crack extraction results are output.
[0020] As a preferred embodiment, constructing the sample dataset includes: The collected crack image data is preprocessed; By combining publicly available crack datasets and preprocessed crack images, and dividing them into training, validation, and test sets, a sample dataset for training is formed.
[0021] In a specific example, high-resolution images of cracks in the background of roads, bridges, and building surfaces can be captured using photographic devices such as mobile phones, cameras, and drones. These high-resolution crack images can then be preprocessed to obtain an image dataset.
[0022] The preprocessing includes data cleaning, data annotation, and data augmentation. Data cleaning is used to initially screen the acquired crack images, removing invalid image data that is blurry, incomplete, lacks cracks, or contains severe noise, which would affect subsequent training results. Data annotation uses the Labelme image annotation tool to perform pixel-level annotation processing on the screened crack images to generate images that are consistent with the original. Figure 1 A corresponding binary mask image is used for supervised learning of the crack region. Data augmentation expands the number of training samples and increases the diversity of image data by performing various operations on the image, including random rotation, image enhancement, brightness adjustment, and affine transformation, thereby enhancing the robustness and generalization ability of the constructed network.
[0023] In a specific example, the publicly available crack dataset includes CFD, DeepCrack, etc., totaling 4258 images, and high-resolution crack images taken by the camera device are named MCD (Multi-scenario Crack Data), totaling 729 images, which together constitute the sample dataset.
[0024] After data preprocessing, the constructed sample dataset is divided into training set, validation set and test set in a ratio of 7:2:1, which are used for network training, network validation and network testing respectively.
[0025] As a preferred embodiment, creating the initial network model includes: A network architecture consisting of an encoder and a decoder is constructed. The encoder is based on the HRNet multi-branch structure and includes multiple feature extraction modules and wavelet downsampling enhancement modules. The decoder includes multiple dual-stream feature fusion modules and prediction modules. The feature extraction module performs upsampling, multi-scale feature extraction, multi-scale feature interaction, and feature identity mapping operations on the crack image data to obtain first feature information at different scales. The wavelet downsampling enhancement module extracts frequency domain information from the crack, downsamples the first feature information, and supplements the first feature information with frequency domain information to form second feature information.
[0026] The dual-stream feature fusion module fuses second-level feature information, performs upsampling and feature alignment, and outputs feature results at a uniform spatial resolution. The prediction module generates classification results for crack pixels and background pixels based on the feature output results.
[0027] In a specific example, the initial model is based on the HRNet deep learning model, combined with wavelet enhancement, feature fusion, attention mechanism and morphological adaptive convolution module to build the initial network model.
[0028] The initial network model first receives two-dimensional crack image data in three channels (R, G, B) and inputs it into the HRNet base network. The base network includes four branches, each with the same resolution and composed of multiple stacked feature extraction modules. The number of stacked modules in the highest resolution branch, the second highest resolution branch, the second lowest resolution branch, and the lowest resolution branch are 4, 3, 2, and 1, respectively. The resolution of each branch is one-quarter of that of the previous branch. An attention-guided wavelet downsampling enhancement module is used as a feature supplement to form the initial feature map of each branch.
[0029] In a preferred embodiment, the feature extraction module includes: a basic feature extraction module, a morphological adaptive convolution module, and an efficient multi-scale attention module; The feature extraction module is built upon a residual structure, comprising multiple convolutional layers, batch normalization layers, and nonlinear activation functions. This enhances feature representation and prevents gradient vanishing, extracting multi-level, multi-scale feature information (the first feature information) from the input image. The morphologically adaptive convolutional module, taking the feature output from the basic module as input (i.e., the first feature information), enhances the network's perception of irregular crack boundaries and curved structures, flexibly capturing crack morphological changes and topological features in spatial dimensions. The efficient multi-scale attention module weights and adjusts the multi-scale feature information to enhance the network's focus on key regions, fusing spatial and channel perception capabilities to strengthen the representation of salient regions and reduce background interference.
[0030] In a specific example, the basic feature extraction module consists of a 3×3 basic convolution, batch normalization, and ReLU activation function, followed by a morphological adaptive convolution module, batch normalization, an efficient multi-scale attention module, a residual connection layer, and another ReLU activation function.
[0031] In a preferred embodiment, the operation of the morphological adaptive convolution module includes: The offset prediction module uses a lightweight convolutional structure to predict the offset parameter of each pixel in the input feature map, i.e., the first feature information, and guides the convolutional kernel to adaptively sample. The dynamic sampling module resamples the neighboring pixels of the feature map, i.e., the position specified by the first feature information, according to the offset to construct an adaptive sampling region. The convolution calculation module performs weighted convolution operations on the adaptive sampling region and outputs the enhanced feature map, that is, the enhanced first feature information.
[0032] To illustrate with a specific example, a standard 3×3 convolution kernel K is represented as follows: To make the convolution kernel more flexible and focus on the complex topological features of the crack structure, inspired by deformable convolution, a deformation offset Δ is introduced. Taking the x-axis direction as an example, the center position of the convolution kernel is used as the position reference, and each position away from this grid depends on the position of the previous grid. K i+1 Relative to K i Increased the offset in the y direction. Furthermore, the current position offset is an accumulation process based on the offset of the previous position, thus making the convolution kernel conform to the linear shape characteristics.
[0033] The change in the x-axis direction is as follows: Where c represents the horizontal distance of each grid cell from the center grid cell; The change along the y-axis is as follows: Where c represents the vertical distance of each grid cell from the center grid cell; Since the offset is usually not an integer, it is bilinearly interpolated to meet the coordinate requirements, as shown below: Where K represents the decimal position of the sampling point in the x and y directions. All integer spatial locations are listed, and B is the bilinear interpolation kernel, i.e.: The offset prediction and adaptive sampling process is completed by transforming the position of the above sampling points, and then the feature extraction within the convolution kernel is completed by convolution operation.
[0034] As a preferred embodiment, the operation of the efficient multi-scale attention module includes: By using a multi-scale feature aggregation unit, different pooling and convolution methods are applied to the first feature information to achieve spatial alignment and feature integration at different scales, thereby obtaining attention weights of a unified dimension. The feature recalibration unit reconstructs the input feature map, i.e. the first feature information, based on attention weights, by channel or spatial dimension weighting, thereby enhancing the expression of salient regions and suppressing redundant information.
[0035] In a specific example, the input feature map first follows the rules. The input feature map is divided into g subgroups, each containing C / g channels. The scale of the feature map after grouping is expressed as follows: Each sub-feature map after segmentation is input into the attention mechanism structure for modeling.
[0036] To aggregate multi-scale spatial structure information, attention weights for grouped features are extracted using three parallel path branches, including a single 3×3 convolution branch and two parallel 1D global average pooling branches. The design of parallel branches also improves the running efficiency to a certain extent.
[0037] For the 1×1 branch, the features are encoded along the horizontal and vertical directions using two 1D global average pooling operations, respectively, resulting in two 1D feature encoding vectors. Specifically, the 1D global average pooling operation along the horizontal direction is defined as follows: in, This represents the feature value of the c-th channel at position i. Similarly, the 1D global average pooling operation along the vertical direction is defined as follows: Two 1D feature encoding vectors are concatenated and processed through a 1×1 convolution, followed by vector decomposition and nonlinear fitting using the Sigmoid function. Then, the two channel-level attention maps within each group are aggregated using simple multiplication. Group normalization is then performed, and the data is fed into a two-dimensional global average pooling process. This operation aims to encode global information and model long-range dependencies, where the two-dimensional global average pooling is defined as follows: To enhance the modeling ability of local structural features, a 3×3 convolutional branch is introduced into the parallel path to perceive local contextual information and compensate for the loss of spatial location information in the global attention path, thereby improving the network's response to fine-grained crack regions.
[0038] 3×3 convolution is used to capture local features and complete cross-channel feature interaction. Then, global average pooling and Softmax activation function are also performed. Thus, the global spatial information encoding at two different scales is completed.
[0039] Finally, the global spatial encoding information obtained from each branch is multiplied by the initial features of the other branch to obtain two cross-scale interactive spatial attention maps. The attention maps of the two branches are then aggregated and multiplied by the original features. After passing through the Sigmoid function, the final feature optimization result is output.
[0040] In a preferred embodiment, the operation of the wavelet downsampling enhancement module includes: The input feature map, i.e. the first feature information, is decomposed into a low-frequency component and multiple high-frequency components by using a wavelet decomposition unit. By using the low-frequency feature extraction unit, basic features with compression characteristics are extracted from the low-frequency components, thereby achieving spatial downsampling of the first feature information and reducing information loss. The feature fusion unit weighted and fused low-frequency features with some high-frequency features to form second feature information, thereby improving the feature representation capability after downsampling.
[0041] In a specific example, the input feature map is first... The frequency band was divided into four sub-bands using wavelet transform: low-frequency sub-band. and three high-frequency sub-bands , , The obtained sub-band feature map .
[0042] Next, the three high-frequency sub-bands were analyzed separately. , , By concatenating the features, a feature map is obtained that conforms to the specifications. The feature map is fed into a 1×1 convolutional layer for channel dimension compression, and then a sigmoid activation function is applied to generate an attention weight map. Attention weight map High-frequency features after splicing Element-wise multiplication is performed to enhance high-frequency detail information.
[0043] The enhanced feature map is then subjected to spatial average pooling to extract its global feature response, and an enhanced high-frequency representation is generated through a 1×1 convolution operation. This enhanced representation is then compared with the low-frequency subband. The fusion is achieved by adding elements one by one, so as to effectively integrate high and low frequency features.
[0044] Finally, the fused feature map is refined through a 1×1 convolutional layer, and the output size conforms to... The downsampling feature enhancement results are used as the final output of this module.
[0045] In a preferred embodiment, the operation of the dual-stream feature fusion module includes: The channel attention generation module performs channel concatenation on feature maps of different scales, i.e., the second feature information. Through pooling, convolution and activation functions, channel attention weights are generated to enhance the concatenated feature maps. The spatial attention generation module performs 1×1 convolution on feature maps at different levels, i.e., the second feature information at different levels, and then concatenates and activates them to generate spatial attention weights. The second feature information after channel enhancement is spatially weighted and adjusted to obtain the fused feature map.
[0046] In a specific example, the feature map is first input from the low-resolution branch, and its size conforms to... It is upsampled and aligned with the feature map size of the previous branch.
[0047] As the first feature optimization path, the two branch feature maps are fused across channels using 1×1×1 convolution and then dimensionality reduced. After that, the features are summed and then processed by the Sigmoid activation function to generate a spatial attention weight map.
[0048] As the second feature optimization flow, the two-branch aligned feature maps are first concatenated in the channel dimension. After global average pooling and 1×1 convolution feature refinement, the channel attention weight map is output through the Sigmoid activation function. After dot product operation with the original concatenated result, the channel dimension reduction is performed by 1×1×1 convolution to restore the original feature map size.
[0049] Finally, the spatial attention weights and the feature maps after channel attention are multiplied by a dot product to obtain the final feature fusion result.
[0050] In a preferred embodiment, the operation of the prediction module includes: The feature fusion unit upsamples and aligns the outputs of different branches, fusing them into a feature map of uniform size. The classification output unit uses 1×1 convolution to classify and predict crack and background pixels, and outputs a refined crack extraction result in the form of a binary mask.
[0051] In a specific example, the prediction module is the result of integrating the two-stream feature fusion module. The output feature map is restored to the input size through two upsampling operations, and 1×1 convolution is used to refine the features and reduce the dimensionality to form the final output result.
[0052] In a specific example, when inputting the crack sample dataset into the initial network model for iterative training, the experimental environment and parameter settings for the training process include: The experimental hardware environment consisted of an Intel Core i9-7900X CPU and an NVIDIA Quadro P6000 GPU, while the software environment consisted of a Windows 10 operating system, Python 3.8, PyTorch 2.4.1, and related neural network Python libraries. The network uses the Adam optimizer, with an initial learning rate of 0.0005. The first momentum parameter of the Adam optimizer is specified as 0.99, the second momentum parameter as 0.999, the weight decay coefficient as 1e-4, the batch size as 4, and 200 epochs. The optimal weights are saved after each training iteration.
[0053] In a specific example, real-time crack data is input into the crack refinement extraction model, and the output is the crack extraction result. That is, the network, after iterative training, performs fine segmentation on the real crack image, forming a result such as... Figure 3 The extraction results are shown.
[0054] like Figure 4 As shown, the second embodiment of this application provides a crack extraction system 500, which includes: Database construction module 501 is used to obtain publicly available crack datasets and collect crack images, and to construct a sample dataset based on the publicly available crack datasets and crack images. Initial network model building module 502 is used to create an initial network model based on the HRNet deep learning network; The network iterative training module 503 is used to input the sample dataset of cracks into the initial network model for iterative training, so as to obtain a crack refinement extraction model based on the optimal weights. The extraction module 504 is used to input real-time crack data into the crack refinement extraction model and output crack extraction results.
[0055] The third embodiment of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.
[0056] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one of relational and non-relational databases. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these. The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described; however, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0057] Compared with existing technologies, the crack extraction method, system, and storage medium provided by this invention construct a refined crack extraction network based on the HRNet deep learning model, comprising an encoder module and a decoder module. The encoder module utilizes the HRNet multi-branch structure for feature extraction and combines wavelet downsampling enhancement to obtain feature information at different scales. The decoder module, through dual-stream feature fusion and prediction modules, fuses features at different scales and generates refined crack extraction results. This achieves a complete process from data acquisition, preprocessing, network construction and training to crack segmentation, solving the problems of severe noise in crack extraction results under complex backgrounds, incomplete extraction of small cracks, and poor crack extraction performance for cracks with diverse and tortuous shapes. Simultaneously, it can effectively extract crack feature information at different scales and has good recognition performance for small and complex cracks, improving the accuracy of crack extraction. Furthermore, the introduction of a morphological adaptive convolution module and an efficient multi-scale attention module enhances the network's ability to perceive irregular crack boundaries and curved structures, strengthens the expression of salient regions, and reduces background interference. Finally, it provides a complete system from data acquisition, annotation, network construction and training to practical applications, facilitating deployment and application in practical engineering.
[0058] The specific embodiments of the present invention described above do not constitute a limitation on the scope of protection of the present invention. Any other corresponding changes and modifications made in accordance with the technical concept of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for extracting cracks, characterized in that, include: Obtain publicly available crack datasets and collect crack images, and construct a sample dataset based on the publicly available crack datasets and crack images; An initial network model was created based on the HRNet deep learning network. The sample dataset of cracks is input into the initial network model for iterative training to obtain a crack refinement extraction model based on optimal weights; Input real-time crack data into the crack refinement extraction model and output the crack extraction results.
2. The crack extraction method as described in claim 1, characterized in that, The creation of the initial network model includes: The encoder is based on the HRNet multi-branch structure and includes multiple feature extraction modules and wavelet downsampling enhancement modules; the decoder includes multiple dual-stream feature fusion modules and prediction modules. The feature extraction module is used to perform upsampling, multi-scale feature extraction, multi-scale feature interaction, and feature identity mapping operations on the crack image data to obtain first feature information at different scales; the wavelet downsampling enhancement module is used to extract frequency domain information in the crack, downsample the first feature information, and supplement the first feature information with frequency domain information to form second feature information. The dual-stream feature fusion module is used to fuse the second feature information, perform upsampling and feature alignment, and output feature output results under a unified spatial resolution. The prediction module is used to generate classification results for crack pixels and background pixels based on the feature output results.
3. The crack extraction method as described in claim 2, characterized in that, The feature extraction module includes: The module consists of a basic feature extraction module, a morphological adaptive convolution module, and an efficient multi-scale attention module. The feature extraction module is built based on a residual structure and performs convolution, batch normalization and nonlinear activation on the input image to extract the first feature information. The morphological adaptive convolution module, taking the first feature information as input, perceives the irregular crack boundary and curved structure, and captures crack morphological changes and topological features in the spatial dimension. The high-efficiency multi-size attention module is used to weight and adjust the first feature information, enhance the attention to key regions, integrate spatial perception and channel perception, strengthen the expression of salient regions and suppress background interference.
4. The crack extraction method as described in claim 3, characterized in that, The operations of the shape-adaptive convolution module include: A lightweight convolutional structure is used to predict the offset parameter of each pixel on the first feature information, guiding the convolutional kernel to adaptively sample. Based on the offset, the neighboring pixels at the specified position in the first feature information are resampled to construct an adaptive sampling region; A weighted convolution operation is performed on the adaptive sampling region to output the enhanced first feature information.
5. The crack extraction method as described in claim 3, characterized in that, The operation of the efficient multi-scale attention module includes: Different pooling and convolution methods are used on the first feature information to achieve spatial alignment and feature integration at different scales, thereby obtaining attention weights of a unified dimension. Based on the attention weights, the first feature information is reconstructed by channel or spatial dimension weighting to enhance the expression of salient features and suppress redundant information.
6. The crack extraction method as described in claim 2, characterized in that, The operation of the wavelet downsampling enhancement module includes: Perform a two-dimensional discrete wavelet transform on the first feature information to decompose it into a low-frequency component and multiple high-frequency components. Extract basic features with compression characteristics from low-frequency components to achieve spatial downsampling of the first feature information; The low-frequency features are weighted and fused with some high-frequency features to form the second feature information.
7. The crack extraction method as described in claim 2, characterized in that, The operation of the dual-stream feature fusion module includes: Channel concatenation is performed on the second feature information at different scales. Channel attention weights are generated through pooling, convolution, and activation functions to enhance the concatenated feature map. The second feature information of different network layers is concatenated and activated after being subjected to 1x1 convolution to generate spatial attention weights. The second feature information after channel enhancement is spatially weighted and adjusted to obtain a fused feature map.
8. The crack extraction method as described in claim 2, characterized in that, The operation of the prediction module includes: The outputs of different branches are upsampled and aligned, and then fused into a feature map of uniform size. A 1x1 convolution is used to classify and predict cracks and background pixels, and the cracks are extracted in the form of a binarized mask.
9. A crack extraction system, characterized in that, include: The database construction module is used to obtain publicly available crack datasets and collect crack images, and to construct a sample dataset based on the publicly available crack datasets and crack images. The initial network model building module is used to create an initial network model based on the HRNet deep learning network. The network iterative training module is used to input the sample dataset of cracks into the initial network model for iterative training, so as to obtain a crack refinement extraction model based on the optimal weights; The extraction module is used to input real-time crack data into the crack refinement extraction model and output crack extraction results.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 8.
Citation Information
Cited By
Image feature enhancement model target detection method based on RTDETR network
CN121582559A
Construction road crack detection method and system, readable storage medium and computer
CN121811262A
Construction road crack detection method, system, readable storage medium and computer
CN121811262B