A crack detection method, device, equipment and medium based on feature enhancement

CN122820549APending Publication Date: 2026-09-25SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610765437.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0002]在现有的裂缝检测技术中,传统方法过度依赖人工设计特征,检测精度与效率易受人为因素影响而难以保证,当前主流深度学习方法仍存在明显不足,如FlexiCrackNet上采样设计简单,未显式建模裂缝边缘,易导致定位模糊;CrackFormer虽能提升裂缝连续性,但计算量大、对细微裂缝捕捉能力弱

Benefits of technology

[0024]本发明利用预训练边缘分割模型构建多尺度特征金字塔,同时保留细小组节、中尺度结构、全局语义信息,对复杂背景、污渍、光照变化、低对比度裂缝具备提升抗干扰能力。再通过渐进式门控聚合模型逐尺度、自适应加权融合特征,避免无效信息干扰,突出裂缝关键特征,提升对宽窄不一及渐变裂缝的统一检测能力。随后,针对裂缝细长性、连续性、拓扑性进行定向增强,强化裂缝连通性与结构完整性,改善裂缝检测中易出现的断裂、断续、断点问题,输出贴近真实物理形态的裂缝结构。在此基础上,通过多尺度边缘检测分支并行处理不同宽度裂缝,配合几何约束细化修正,实现边缘定位更准、轮廓更光滑、宽度更真实。从整体上有效解决了裂缝检测结果易出现断续现象的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820549A_ABST
    Figure CN122820549A_ABST
Patent Text Reader

Abstract

The application provides a feature enhancement-based crack detection method and device, equipment and medium, and relates to the technical field of crack detection, which comprises the following steps: performing bilinear interpolation scaling on a crack image to construct a standard crack image tensor; inputting the standard crack image tensor into a pre-trained edge segmentation model to extract multi-level features and obtain a multi-scale feature pyramid; fusing the multi-scale feature pyramid according to a progressive gated aggregation model to obtain a progressive fusion feature map; performing shape perception enhancement on the progressive fusion feature map based on a cross-modal adaptive attention mechanism to obtain a shape perception feature map; and performing edge perception positioning decoding on the shape perception feature map, extracting crack edges of different widths in parallel through a multi-scale edge detection branch, and geometrically constraining, refining and correcting the crack edges to obtain a crack detection segmentation map. The application solves the problem that the crack detection result is prone to discontinuity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of crack detection technology, and more specifically, to a crack detection method, apparatus, equipment, and medium based on feature enhancement. Background Technology

[0002] In existing crack detection technologies, traditional methods rely excessively on manually designed features, making it difficult to guarantee detection accuracy and efficiency due to human factors. Current mainstream deep learning methods still have significant shortcomings. For example, FlexiCrackNet's upsampling design is simple and does not explicitly model crack edges, easily leading to blurred localization. While CrackFormer can improve crack continuity, it has high computational cost and weak ability to capture fine cracks. Overall, existing technologies generally suffer from the common defects of losing fine crack feature information during downsampling, fixed convolution kernels being unable to adapt to crack morphologies with arbitrary orientations, and uniform upsampling strategies failing to distinguish the feature distribution differences between smooth and edge regions, thus causing edge blurring. Ultimately, this leads to the problem of discontinuous crack detection results.

[0003] Therefore, there is an urgent need for a crack detection method, device, equipment, and medium based on feature enhancement, which solves the problem of discontinuous crack detection results. Summary of the Invention

[0004] The purpose of this invention is to provide a crack detection method, apparatus, device, and medium based on feature enhancement to improve the aforementioned problems. To achieve the above objective, the technical solution adopted by this invention is as follows:

[0005] Firstly, this application provides a crack detection method based on feature enhancement, comprising:

[0006] Acquire images of cracks on the surface of a concrete structure;

[0007] The crack image is scaled using bilinear interpolation to construct a standard crack image tensor;

[0008] The standard crack image tensor is input into a pre-trained edge segmentation model for multi-level feature extraction to obtain a multi-scale feature pyramid.

[0009] The multi-scale feature pyramid is subjected to multi-scale feature progressive gating fusion according to the progressive gating aggregation model to obtain a progressive fused feature map;

[0010] Based on a cross-modal adaptive attention mechanism, the progressively fused feature map is enhanced with elongation, continuity, and topology to obtain a morphology-aware feature map.

[0011] Edge-aware localization and decoding are performed on the morphological feature map. Crack edges of different widths are extracted in parallel through multi-scale edge detection branches and geometrically constrained for refinement and correction to obtain a crack detection segmentation map.

[0012] Secondly, this application also provides a crack detection device based on feature enhancement, comprising:

[0013] The acquisition module is used to acquire images of cracks on the surface of concrete structures.

[0014] A standard module is used to perform bilinear interpolation scaling on the crack image to construct a standard crack image tensor;

[0015] The extraction module is used to input the standard crack image tensor into a pre-trained edge segmentation model for multi-level feature extraction, thereby obtaining a multi-scale feature pyramid.

[0016] The fusion module is used to perform multi-scale feature progressive gating fusion on the multi-scale feature pyramid according to the progressive gating aggregation model to obtain a progressive fused feature map.

[0017] The enhancement module is used to perform morphological perception enhancement on the progressively fused feature map based on a cross-modal adaptive attention mechanism to improve its elongation, continuity, and topology, thereby obtaining a morphological perception feature map.

[0018] The correction module is used to perform edge-aware localization and decoding on the morphological perception feature map. It extracts crack edges of different widths in parallel through multi-scale edge detection branches and refines and corrects them with geometric constraints to obtain a crack detection segmentation map.

[0019] Thirdly, this application also provides a crack detection device based on feature enhancement, comprising:

[0020] Memory, used to store computer programs;

[0021] A processor is configured to implement the steps of the feature-enhanced crack detection method when executing the computer program.

[0022] Fourthly, this application also provides a medium on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described feature-enhanced crack detection method.

[0023] The beneficial effects of this invention are as follows:

[0024] This invention utilizes a pre-trained edge segmentation model to construct a multi-scale feature pyramid, preserving small segments, mesoscale structures, and global semantic information, thus enhancing its anti-interference capabilities against complex backgrounds, stains, lighting variations, and low-contrast cracks. A progressively gated aggregation model then adaptively weights and fuses features scale-by-scale to avoid interference from invalid information, highlighting key crack features and improving the unified detection capability for cracks of varying widths and gradients. Subsequently, targeted enhancements are applied to the crack's elongation, continuity, and topology, strengthening crack connectivity and structural integrity, and improving issues such as breaks, discontinuities, and breakpoints that easily occur in crack detection, outputting crack structures that closely resemble real physical morphology. Building upon this, multi-scale edge detection branches process cracks of different widths in parallel, combined with geometric constraint refinement, achieving more accurate edge localization, smoother contours, and more realistic widths. Overall, this effectively solves the problem of discontinuous crack detection results.

[0025] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a schematic diagram of the feature-enhanced crack detection method described in an embodiment of the present invention;

[0028] Figure 2 This is a schematic diagram of the processing flow of the cross-modal adaptive attention mechanism described in this embodiment of the invention;

[0029] Figure 3 This is a schematic diagram of the edge-aware positioning and decoding process described in an embodiment of the present invention;

[0030] Figure 4 This is a schematic diagram of the feature-enhanced crack detection device described in an embodiment of the present invention.

[0031] The diagram is labeled as follows: 800, feature-enhanced crack detection device; 801, processor; 802, memory; 803, multimedia component; 804, I / O interface; 805, communication component. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0033] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0034] In engineering environments involving crack detection on concrete structures, when using deep neural networks for crack detection, downsampling operations tend to filter out minute crack features with widths less than 2 pixels as noise, causing the deep network to fail to perceive them, resulting in the loss of small cracks and low detection recall. Secondly, the upsampling process, due to the use of uniform bilinear interpolation or deconvolution, fails to distinguish between smooth and edge regions, resulting in jagged and blurry recovered crack edges, making it difficult to achieve sub-pixel-level width measurement and insufficient positioning accuracy. Finally, due to interference from stains and shadows, as well as the arbitrary direction and complex topological structure of cracks, coupled with the fixed receptive field of the convolution kernel, which makes it difficult to match slender geometric features, and the model's lack of long-distance dependency modeling ability, the detected cracks may appear as discontinuous segments, ultimately leading to technical problems related to complex morphology and topological fractures.

[0035] Example 1:

[0036] This embodiment provides a crack detection method based on feature enhancement.

[0037] See Figure 1 The figure shows that the method includes steps S1 to S6, including:

[0038] S1: Obtain images of cracks on the surface of the concrete structure;

[0039] This step involves using a high-definition camera to capture on-site images of cracked areas on the surface of concrete structures such as bridges, roads, and buildings, obtaining original RGB color images with complete crack information as crack images.

[0040] S2: Perform bilinear interpolation scaling on the crack image to construct a standard crack image tensor;

[0041] In this step, the crack images are subjected to bilinear interpolation scaling to uniformly adjust all crack images to a standard size of 512×512 pixels. At the same time, channel-level normalization is performed to eliminate the magnitude difference of image pixel values ​​between different channels, and finally output a standard crack image tensor with a dimension of 512×512×3.

[0042] S3: Input the standard crack image tensor into a pre-trained edge segmentation model to perform multi-level feature extraction and obtain a multi-scale feature pyramid.

[0043] To clarify the specific method for obtaining the multi-scale feature pyramid, step S3 includes S31 to S33, specifically:

[0044] S31: Input the standard crack image tensor into the multi-cascade stage module of the pre-trained edge segmentation model to perform preliminary feature mapping and size regularization to obtain the initial feature map;

[0045] In this step, the standard crack image tensor is input into a multi-cascaded stage module consisting of four stages in the pre-trained EdgeSAM encoder. Each stage performs a 2x downsampling operation through a convolution with Stage=2, sequentially performing preliminary convolutional feature mapping and layer-by-layer size regularization on the standard crack image tensor, completing the preliminary transformation from the original image to the feature map, and obtaining the initial feature map.

[0046] S32: Based on the reparameterizable VGG block of the multi-cascaded stage module, the initial feature map is convolved and feature enhanced to extract multi-level features from details to semantics to obtain shallow feature map and deep feature map.

[0047] In this step, the initial feature map is sequentially subjected to convolution operations and feature enhancement processing based on the two reparameterizable VGG blocks (RepVGG) contained in each stage of the multi-cascaded stage module. Image features are extracted step by step from the bottom layer to the top layer. The first two stages output a first shallow feature map S1 (128×128×64) and a second shallow feature map S2 (64×64×128) that retains the texture and detail information of the crack edges. The last two stages output a third deep feature map S3 (32×32×256) and a fourth deep feature map S4 (16×16×512) that contains the semantic abstract information of the crack.

[0048] S33: Perform multi-scale fusion and channel dimension splicing on the shallow feature map and the deep feature map to obtain a multi-scale feature pyramid.

[0049] In this step, the shallow feature map and the deep feature map are integrated at multiple scales, retaining the unique information and related features of each scale feature. At the same time, channel-dimensional splicing is performed to form a multi-scale feature pyramid of full-dimensional feature information from detailed texture to semantic abstraction. The multi-scale feature pyramid is based on feature maps of different scales under the first shallow feature map S1, the second shallow feature map S2, the third deep feature map S3, and the fourth deep feature map S4.

[0050] S4: Perform multi-scale feature progressive gating fusion on the multi-scale feature pyramid according to the progressive gating aggregation model to obtain a progressive fused feature map;

[0051] In this step, the multi-scale feature pyramid includes a first, second, third, and fourth feature map, where the first shallow feature map S1 is the first feature map, the second shallow feature map S2 is the second feature map, the third deep feature map S3 is the third feature map, and the fourth deep feature map S4 is the fourth feature map.

[0052] To clarify the specific method for obtaining the progressively fused feature map, step S4 includes S41 to S45, specifically:

[0053] S41: The spatial resolution of the first feature map is adjusted by upsampling according to the progressive gated aggregation model, and then fused with the second feature map to obtain the first fused feature map;

[0054] In this step, the fourth deep feature map S4 (16×16×512) in the multi-scale feature pyramid is upsampled according to the progressive gating aggregation model. The spatial resolution of the fourth deep feature map S4 is adjusted to be consistent with that of the third deep feature map S3 (32×32×256). The adjusted fourth deep feature map S4 and the third deep feature map S3 are fused through the first gating aggregation unit IGAM-1 of the progressive gating aggregation model to obtain the first fused feature map F1 with a dimension of 32×32×256.

[0055] S42: Input the first fused feature map and the second feature map into the second gated aggregation unit for feature reuse enhancement fusion to obtain the second fused feature map;

[0056] In this step, the first fused feature map F1 and the third deep feature map S3 are input again into the second gated aggregation unit IGAM-2 of the progressive gated aggregation model. The third deep feature map S3 is reused and enhanced by fusion with the first fused feature map F1, thereby mining the association information of deep semantic features. The output dimension is still the second fused feature map F2 with a dimension of 32×32×256.

[0057] S43: After upsampling the spatial resolution of the second fused feature map, it is combined with the second feature map and input into the third gated aggregation unit for feature fusion to obtain the third fused feature map;

[0058] In this step, an upsampling operation is performed on the second fused feature map F2, increasing its spatial resolution from 32×32 to 64×64, which is then matched with the resolution of the second shallow feature map S2 (64×64×128). Subsequently, the second shallow feature map S2 and the second fused feature map F2 are input into the third gated aggregation unit IGAM-3 for cross-scale feature fusion, resulting in a third fused feature map F3 with a dimension of 64×64×128, thus achieving the initial combination of deep semantic features and shallow detail features.

[0059] S44: Input the third fused feature map and the third feature map into the fourth gated aggregation unit to enhance shallow details and deeply combine semantic information to obtain the fourth fused feature map;

[0060] In this step, the third fused feature map F3 and the second shallow feature map S2 are input into the fourth gated aggregation unit IGAM-4. The shallow edge texture details of the second shallow feature map S2 are reused and enhanced. At the same time, the second shallow feature map S2 is deeply combined with deep semantic information to improve the detail expression and semantic association ability of the feature map. The output is a fourth fused feature map F4 with a dimension of 64×64×128.

[0061] S45: Input the fourth fused feature map and the fourth feature map into the fifth gated aggregation unit to perform fine edge texture, mid-level structure and high-level semantic fusion to obtain a progressive fused feature map.

[0062] In this step, the fourth fused feature map F4 is upsampled to increase its spatial resolution from 64×64 to 128×128, maintaining consistency with the resolution of the first shallow feature map S1 (128×128×64). Then, the first shallow feature map S1 and the fourth fused feature map F4 are input into the fifth gated aggregation unit IGAM-5 for comprehensive progressive fusion of the fine edge texture features of cracks in the first shallow feature map S1, the mid-level structural features of the second shallow feature map S2, and the high-level semantic features of the third deep feature map S3 and the fourth deep feature map S4, ultimately resulting in a progressive fused feature map F5 with dimensions of 128×128×64.

[0063] S5: Based on the cross-modal adaptive attention mechanism, the progressively fused feature map is enhanced with elongation, continuity and topology to obtain a morphologically perceptual feature map;

[0064] To clarify the specific method for acquiring the morphologically perceived feature map, step S5 includes S51 to S54, specifically:

[0065] S51: Based on the cross-modal adaptive attention mechanism, the progressively fused feature map is input into the multi-directional directional convolution and deformable convolution in the slender structure detection model to perform dual slender feature enhancement, thereby obtaining slender structure enhancement features;

[0066] To clarify the specific method for obtaining the enhancement features of the slender structure, step S51 includes S511 to S514, specifically:

[0067] S511: The progressively fused feature map is processed based on the long strip convolution kernel in the multi-directional directional convolution to generate response maps of crack directions with different orientations;

[0068] like Figure 2 As shown, the progressively fused feature map is convolved based on the elongated convolution kernel in the multi-directional directional convolution. Each convolution kernel extracts the crack features of the corresponding direction, and finally generates 8 response maps that can represent cracks of different directions, which are used to capture the morphological information of cracks in various directions.

[0069] The multi-directional directional convolution branch uses eight pre-defined elongated convolution kernels with angles ranging from 0° to 157.5° and intervals of 22.5°.

[0070] S512: Channel splicing and convolution fusion are performed on the response maps of the cracks with different orientations to obtain the features of cracks with different orientations;

[0071] In this step, the response maps of cracks with different orientations are stitched together along the channel dimension, and the crack feature information of all directions is integrated. Then, a 1×1 convolutional layer is used to compress the dimensions of the stitched features and fuse the features, thereby eliminating feature redundancy and strengthening the correlation of crack features with different orientations, and outputting crack features with different orientations, thus achieving effective extraction of crack directional features.

[0072] S513: Based on the deformable convolution, predict the spatial offset of each sampling point of the progressively fused feature map to obtain the crack geometric adaptive feature;

[0073] In this step, the progressively fused feature map is input into the deformable convolution branch of the slender structure detection model (slender structure detector). The spatial offset of each pixel sampling point on the progressively fused feature map is predicted by the auxiliary convolutional layer, so that the sampling points of the convolution kernel can complete adaptive displacement according to the actual bending and twisting shape of the crack, thereby fitting the geometric contour of the crack and outputting the crack geometric adaptive feature, which adapts to the complex geometric shape of the crack.

[0074] S514: The crack features with different orientations and the crack geometry adaptive features are fused to obtain the slender structure enhancement features.

[0075] In this step, the crack features with different orientations and the crack geometric adaptive features are fused element-wise by adding them together. Combining the advantages of directional feature extraction from multi-directional directional convolution and the geometric adaptive advantages of deformable convolution, slender structure enhancement features are generated, which solves the problem of insufficient identification of slender cracks.

[0076] S52: The elongated structure enhancement features are spatially scanned and continuously enhanced according to the bidirectional gated loop unit to obtain the crack spatial continuity features;

[0077] To clarify the specific method for obtaining the spatial continuity characteristics of the crack, step S25 includes S521 to S523, specifically:

[0078] S521: The elongated structure enhancement features are horizontally scanned according to the bidirectional gated loop unit. By splicing the hidden states of the horizontal scan results into bidirectional hidden states, the horizontally dependent features are obtained.

[0079] In this step, the elongated structure enhancement features are input into a set of bidirectional gated recurrent units (GRUs). The two-dimensional feature map of the elongated structure enhancement features is expanded into a one-dimensional sequence by rows and a forward scan from left to right and a backward scan from right to left are performed respectively to capture the long-distance spatial dependency of the crack in the horizontal direction. Subsequently, the bidirectional hidden states obtained from the forward and backward scans are spliced ​​and integrated and restored to the form of a two-dimensional feature map to obtain the horizontal dependency features. The horizontal dependency features are used to characterize the global correlation information of the crack in the horizontal direction.

[0080] S522: The elongated structure enhancement feature is vertically scanned according to the bidirectional gated loop unit. By splicing the hidden states of the vertical scan result into bidirectional hidden states, the vertical direction dependent feature is obtained.

[0081] In this step, the elongated structure enhancement features are input into another set of bidirectional gated recurrent units (GRUs). The two-dimensional feature map of the elongated structure enhancement features is expanded into a one-dimensional sequence by columns and a forward scan from top to bottom and a backward scan from bottom to top are performed to obtain the spatial continuity association information of the crack in the vertical direction. Then, the forward and backward bidirectional hidden states of the scanning process are stitched together and restored to a two-dimensional feature map to obtain the vertical direction dependent features. The vertical direction dependent features are used to characterize the global association information of the crack in the vertical direction.

[0082] S523: Channel splicing and convolution fusion are performed on the horizontal and vertical dependent features to obtain the spatial continuity features of the crack.

[0083] In this step, the horizontal and vertical dependent features are concatenated along the channel dimension to integrate the global spatial dependency information of the crack in both the horizontal and vertical directions. Then, a 1×1 convolutional layer is used to compress the dimensions and fuse the features after concatenation, while strengthening the feature correlation between different directions. Finally, the fused features are added to the original elongated structure enhancement features by residual addition to compensate for the information loss during the feature fusion process. This completes the crack spatial continuity modeling and yields a crack spatial continuity feature with dimensions of 128×128×64 that can effectively fill the visual breakpoints of the crack.

[0084] S53: Input the continuous features of the crack space into the topology preservation model, and perform parallel prediction through the skeleton extraction branch and the connectivity modeling branch to obtain the skeleton probability map and the connected component probability map.

[0085] To clarify the specific methods for obtaining the skeleton probability graph and the connected component probability graph, step S53 includes S531 to S534, specifically:

[0086] S531: Input the spatial continuity features of the crack into the skeleton extraction branch of the topology-preserving model, and predict the crack centerline through three convolutional layers and logistic functions to obtain the initial skeleton probability map.

[0087] In this step, the spatially continuous features of the crack are input into the skeleton extraction branch in the topology-preserving model. The core structural features of the crack are extracted layer by layer through three convolutional layers to capture the centerline information of the crack. Then, the output of the three convolutional layers is compressed to the probability interval of [0, 1] by the logistic function (Sigmoid) to generate an initial skeleton probability map. The initial skeleton probability map is used to characterize the confidence distribution of each pixel as the centerline of the crack.

[0088] S532: Based on the skeletonization algorithm, extract the skeleton as the skeleton supervision signal to perform topology-guided enhancement on the initial skeleton probability map to obtain the skeleton probability map;

[0089] In this step, a precise crack centerline with a width of 1 pixel is extracted from the real crack annotation data based on the skeletonization algorithm as a skeleton supervision signal. The skeleton supervision signal and the initial skeleton probability map are subjected to supervised learning and feature correction. The confidence of the crack centerline in the initial skeleton probability map is enhanced and the noise confidence of the background area is suppressed to obtain a precise skeleton probability map.

[0090] S533: Input the continuous features of the crack space into the connectivity modeling branch of the topology preservation model, and predict the pixel connectivity attribution through three convolutional layers and logistic functions to obtain the initial connected domain probability map.

[0091] In this step, the continuous features of the crack space are input into the connectivity modeling branch of the topology-preserving model. The same three-layer convolutional layer structure as the skeleton extraction branch is used to extract the connectivity features between pixels. Then, the output of the three-layer convolutional layer is mapped to the interval [0, 1] by the logistic function (Sigmoid) to generate an initial connected component probability map. The initial connected component probability map is used to characterize the confidence that each pixel belongs to the same connected crack region, reflecting the connectivity relationship of the crack pixels.

[0092] S534: Based on the connected component labeling algorithm, generate labels as connectivity monitoring signals to perform topology-guided enhancement on the initial connected component probability graph, and obtain the connected component probability graph.

[0093] In this step, the actual labeled data of the cracks is divided into connected components based on the connected component labeling algorithm, and connectivity labels are generated. The connectivity labels are used as connectivity supervision signals to supervise and optimize the initial connected component probability map, thereby strengthening the pixel confidence of the connected crack region and reducing the confidence of isolated noise and background regions, completing the topology-guided enhancement process, and obtaining the connected component probability map.

[0094] S54: Perform element-wise processing on the skeleton probability map, the connected component probability map, and the continuous feature of the crack space to obtain a morphological perception feature map.

[0095] In this step, the skeleton probability map and the connected component probability map are used as dual attention masks and multiplied element-wise with the continuous features of the crack space. The continuous features of the crack space are weighted and filtered through the attention mask, which enhances the feature response of the crack skeleton and connected regions and effectively suppresses the background noise of non-crack regions. This completes the enhancement of morphological perception in three dimensions: elongation, continuity, and topology, resulting in a morphological perception feature map with dimensions of 128×128×64.

[0096] S6: Perform edge-aware localization decoding on the morphological perception feature map, extract crack edges of different widths in parallel through multi-scale edge detection branches and refine and correct them with geometric constraints to obtain a crack detection segmentation map.

[0097] To clarify the specific method for obtaining the crack detection segmentation map, step S6 includes S61 to S66, specifically:

[0098] S61: The shape-aware feature map is used to predict the edge probability map through lightweight convolution. By performing edge-aware sampling on the smooth and edge regions of the predicted edge probability map, full-resolution decoding features are obtained.

[0099] like Figure 3 As shown, the morphologically aware feature map is input into the edge-aware upsampling stage of the edge-aware localization decoder (EPLD decoder). First, a lightweight convolutional layer is used to calculate and generate an edge probability map representing the confidence level of a pixel as an edge region. (Between 0 and 1), the predicted edge probability map is divided into smooth regions and edge regions. PixelShuffle subpixel convolution upsampling is applied to the smooth regions to preserve texture continuity. Transposed convolution is applied to the edge regions to adjust the number of channels, followed by nearest neighbor interpolation upsampling to maintain edge sharpness. Then, the upsampling results of the two regions are adaptively fused by weighted edge response values. Two levels of upsampling are completed sequentially from 128×128 to 256×256 and then to 512×512, restoring the feature resolution to the input corresponding size of 512×512×16, and finally obtaining the full-resolution decoded features.

[0100] S62: Extract crack edges of different widths in parallel from the full-resolution decoded features based on the convolution in the multi-scale edge detection branch to obtain comprehensive edge features;

[0101] In this step, the full-resolution decoding features are processed by convolution in the multi-scale edge detection branch. Edge response features corresponding to 1-2 pixels fine edges, 3-5 pixels medium edges, and more than 5 pixels coarse edges are extracted in parallel using 3×3 convolution, 5×5 convolution, and 7×7 convolution. Then, the edge response features of the three scales are weighted and merged by adaptive fusion weights to generate comprehensive edge features that are adapted to the needs of crack detection of different widths.

[0102] The adaptive fusion weight expression for the edge-aware upsampling step is as follows:

[0103] ;

[0104] In the above formula, This is an edge-aware upsampled feature map. To smooth out the region confidence level, This is a marginal probability map. For the upsampled feature map of the smooth region, For the upsampled feature map of the edge region, This is for element-wise multiplication.

[0105] S63: Based on distance field prediction as an auxiliary task, predict the distance from the pixel in the full-resolution decoded features to the nearest crack boundary to obtain a distance field map;

[0106] In this step, distance field prediction is used as an auxiliary task. The full-resolution decoded features are input into a lightweight convolutional network. The lightweight convolutional network learns to predict the Euclidean distance from each pixel to the nearest crack boundary, generating a distance field map. The distance field map is used to represent the spatial positional relationship between pixels and crack boundaries.

[0107] S64: Calculate the gradient field pointing to the crack boundary based on the distance field map, and integrate the geometric prior information of the gradient field into the decoding features to obtain the geometrically constrained decoding features;

[0108] In this step, gradient calculation is performed based on the distance field map to obtain a gradient field containing information about the direction of the pixel pointing to the nearest crack boundary. The dual geometric prior information of the distance field map and the gradient field is integrated into the full-resolution decoding feature and geometrically constrained to obtain the geometrically constrained decoding feature.

[0109] S65: Perform range field geometric constraints on the geometrically constrained decoding features and coarsely segment the head to output the initial segmentation mask image;

[0110] In this step, the geometrically constrained decoded features are fed into a coarse segmentation head constructed by 1×1 convolution under the geometric prior constraints of the distance field for feature mapping and binarization, generating an initial segmentation mask map with dimensions of 512×512×1.

[0111] S66: Based on the initial segmentation mask and the integrated edge features, a lightweight residual network is used to iteratively predict the boundary displacement correction to obtain the crack detection segmentation map.

[0112] In this step, the initial segmentation mask and the integrated edge features are input into a lightweight residual network (RefineNet). The network is refined in three iterations using residual learning. In each iteration, the boundary offset of the current mask is predicted and the mask edge is precisely corrected. The initial coarse mask is gradually optimized into a fine mask, and finally a 512×512 crack detection segmentation map with clear edge contours is obtained.

[0113] Example 2:

[0114] This embodiment provides a crack detection device based on feature enhancement, the device comprising:

[0115] The acquisition module is used to acquire images of cracks on the surface of concrete structures.

[0116] A standard module is used to perform bilinear interpolation scaling on the crack image to construct a standard crack image tensor;

[0117] The extraction module is used to input the standard crack image tensor into a pre-trained edge segmentation model for multi-level feature extraction, thereby obtaining a multi-scale feature pyramid.

[0118] The fusion module is used to perform multi-scale feature progressive gating fusion on the multi-scale feature pyramid according to the progressive gating aggregation model to obtain a progressive fused feature map.

[0119] The enhancement module is used to perform morphological perception enhancement on the progressively fused feature map based on a cross-modal adaptive attention mechanism to improve its elongation, continuity, and topology, thereby obtaining a morphological perception feature map.

[0120] To clarify the specific methods for obtaining the enhancement modules, the following are provided:

[0121] The elongated unit is used to input the progressively fused feature map into the multi-directional directional convolution and deformable convolution in the elongated structure detection model based on the cross-modal adaptive attention mechanism to perform dual elongated feature enhancement and obtain elongated structure enhanced features;

[0122] A continuous unit is used to perform spatial scanning and continuous enhancement of the elongated structure enhancement features according to the bidirectional gated loop unit, so as to obtain the crack spatial continuity feature;

[0123] Topological units are used to input the continuous features of the crack space into the topology-preserving model, and perform parallel prediction through skeleton extraction branches and connectivity modeling branches to obtain skeleton probability maps and connected component probability maps.

[0124] The processing unit is used to process the skeleton probability map, the connected component probability map and the continuous features of the crack space element by element to obtain a morphological perception feature map.

[0125] The correction module is used to perform edge-aware localization and decoding on the morphological perception feature map. It extracts crack edges of different widths in parallel through multi-scale edge detection branches and refines and corrects them with geometric constraints to obtain a crack detection segmentation map.

[0126] To clarify the specific methods for obtaining the correction module, the following are provided:

[0127] The first prediction unit is used to predict the edge probability map by performing lightweight convolution on the morphologically aware feature map, and to obtain full-resolution decoded features by performing edge-aware sampling on the smooth regions and edge regions of the predicted edge probability map.

[0128] The extraction unit is used to extract crack edges of different widths in parallel from the full-resolution decoded features based on the convolution in the multi-scale edge detection branch, so as to obtain comprehensive edge features;

[0129] The second prediction unit is used to predict the distance from the pixel to the nearest crack boundary in the full-resolution decoded features based on distance field prediction as an auxiliary task, so as to obtain a distance field map.

[0130] The calculation unit is used to calculate the gradient field pointing to the crack boundary based on the distance field map, and to integrate the geometric prior information of the gradient field into the decoding features to obtain the geometrically constrained decoding features.

[0131] The segmentation unit is used to perform range field geometric constraints on the geometrically constrained decoding features and coarsely process the segmentation head, and output an initial segmentation mask image;

[0132] The correction unit is used to perform lightweight residual network iteration to predict the boundary displacement correction amount based on the initial segmentation mask map and the integrated edge features, so as to obtain the crack detection segmentation map.

[0133] It should be noted that the specific manner in which each module performs its operation in the apparatus described in the above embodiments has been described in detail in the embodiments of the method, and will not be elaborated here.

[0134] Example 3:

[0135] Corresponding to the above method embodiments, this embodiment also provides a feature-enhanced crack detection device. The feature-enhanced crack detection device described below and the feature-enhanced crack detection method described above can be referred to and correspond to each other.

[0136] Figure 4 This is a block diagram illustrating a feature-enhanced crack detection device 800 according to an exemplary embodiment. Figure 4 As shown, the feature-enhanced crack detection device 800 may include a processor 801 and a memory 802. The feature-enhanced crack detection device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.

[0137] The processor 801 controls the overall operation of the feature-enhanced crack detection device 800 to complete all or part of the steps in the aforementioned feature-enhanced crack detection method. The memory 802 stores various types of data to support the operation of the feature-enhanced crack detection device 800. This data may include, for example, instructions for any application or method operating on the feature-enhanced crack detection device 800, as well as application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 805 is used for wired or wireless communication between the feature-enhanced crack detection device 800 and other devices. Wireless communication includes Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, or an NFC module.

[0138] In an exemplary embodiment, the feature-enhanced crack detection device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the feature-enhanced crack detection method described above.

[0139] Example 4:

[0140] Corresponding to the above method embodiments, this embodiment also provides a medium. The medium described below can be referred to in relation to the feature-enhanced crack detection method described above.

[0141] A medium storing a computer program, which, when executed by a processor, implements the steps of the feature-enhanced crack detection method described in the above method embodiments.

[0142] The medium can specifically be any medium capable of storing program code, such as a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0143] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0144] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A crack detection method based on feature enhancement, characterized in that, include: Acquire images of cracks on the surface of a concrete structure; The crack image is scaled using bilinear interpolation to construct a standard crack image tensor; The standard crack image tensor is input into a pre-trained edge segmentation model for multi-level feature extraction to obtain a multi-scale feature pyramid. The multi-scale feature pyramid is subjected to multi-scale feature progressive gating fusion according to the progressive gating aggregation model to obtain a progressive fused feature map; Based on a cross-modal adaptive attention mechanism, the progressively fused feature map is enhanced with elongation, continuity, and topology to obtain a morphology-aware feature map. Edge-aware localization and decoding are performed on the morphological feature map. Crack edges of different widths are extracted in parallel through multi-scale edge detection branches and geometrically constrained for refinement and correction to obtain a crack detection segmentation map.

2. The crack detection method based on feature enhancement according to claim 1, characterized in that, Based on a cross-modal adaptive attention mechanism, the progressively fused feature map is enhanced with elongation, continuity, and topology to obtain a morphologically-aware feature map, including: Based on the cross-modal adaptive attention mechanism, the progressively fused feature map is input into the multi-directional directional convolution and deformable convolution in the slender structure detection model to perform dual slender feature enhancement, thereby obtaining slender structure enhanced features; The elongated structure enhancement features are spatially scanned and continuously enhanced using a bidirectional gated loop unit to obtain the crack spatial continuity features. The continuous features of the crack space are input into the topology-preserving model, and parallel prediction is performed through the skeleton extraction branch and the connectivity modeling branch to obtain the skeleton probability map and the connected component probability map. The skeleton probability map, the connected component probability map, and the continuous spatial features of the crack are processed element-wise to obtain a morphological perception feature map.

3. The crack detection method based on feature enhancement according to claim 2, characterized in that, Based on the aforementioned cross-modal adaptive attention mechanism, the progressively fused feature map is input into the multi-directional directional convolution and deformable convolution in the slender structure detection model for dual slender feature enhancement, resulting in slender structure enhanced features, including: The progressively fused feature map is processed based on the long strip convolution kernel in the multi-directional directional convolution to generate response maps of crack directions with different orientations; The response maps of cracks with different orientations are spliced ​​together and convolved to obtain the characteristics of cracks with different orientations. Based on the deformable convolution, the spatial offset of each sampling point in the progressively fused feature map is predicted to obtain the crack geometry adaptive feature; The features of cracks with different orientations and the adaptive crack geometry are fused together to obtain the slender structure enhancement features.

4. The crack detection method based on feature enhancement according to claim 2, characterized in that, The continuous features of the crack space are input into the topology-preserving model, and parallel prediction is performed through a skeleton extraction branch and a connectivity modeling branch to obtain a skeleton probability map and a connected component probability map, including: The spatially continuous features of the crack are input into the skeleton extraction branch of the topology-preserving model. The crack centerline is predicted by three convolutional layers and a logistic function to obtain the initial skeleton probability map. The skeleton extracted based on the skeletonization algorithm is used as a skeleton supervision signal to perform topology-guided enhancement on the initial skeleton probability map, thereby obtaining the skeleton probability map; The continuous features of the crack space are input into the connectivity modeling branch of the topology preservation model. The pixel connectivity attribution is predicted by three convolutional layers and the logistic function to obtain the initial connected domain probability map. The initial connected component probability graph is enhanced by generating labels based on the connected component labeling algorithm as connectivity monitoring signals, thereby obtaining the connected component probability graph.

5. The crack detection method based on feature enhancement according to claim 1, characterized in that, The morphologically aware feature map is subjected to edge-aware localization and decoding. Crack edges of different widths are extracted in parallel through multi-scale edge detection branches and refined using geometric constraints to obtain a crack detection segmentation map, including: The morphologically aware feature map is used to predict edge probability maps through lightweight convolution. By performing edge-aware sampling on the smooth and edge regions of the predicted edge probability map, full-resolution decoding features are obtained. Based on the convolution in the multi-scale edge detection branch, crack edges of different widths are extracted in parallel from the full-resolution decoded features to obtain comprehensive edge features; Based on distance field prediction as an auxiliary task, the distance from the pixel in the full-resolution decoded feature to the nearest crack boundary is predicted to obtain the distance field map; The gradient field pointing to the crack boundary is calculated based on the distance field map. The geometric prior information of the gradient field is incorporated into the decoding features to obtain the geometrically constrained decoding features. The geometrically constrained decoding features are subjected to range field geometric constraints and coarse segmentation head processing to output an initial segmentation mask image; Based on the initial segmentation mask and the integrated edge features, a lightweight residual network is used to iteratively predict the boundary displacement correction, resulting in a crack detection segmentation map.

6. A crack detection device based on feature enhancement, characterized in that, include: The acquisition module is used to acquire images of cracks on the surface of concrete structures. A standard module is used to perform bilinear interpolation scaling on the crack image to construct a standard crack image tensor; The extraction module is used to input the standard crack image tensor into a pre-trained edge segmentation model for multi-level feature extraction, thereby obtaining a multi-scale feature pyramid. The fusion module is used to perform multi-scale feature progressive gating fusion on the multi-scale feature pyramid according to the progressive gating aggregation model to obtain a progressive fused feature map. The enhancement module is used to perform morphological perception enhancement on the progressively fused feature map based on a cross-modal adaptive attention mechanism to improve its elongation, continuity, and topology, thereby obtaining a morphological perception feature map. The correction module is used to perform edge-aware localization and decoding on the morphological perception feature map. It extracts crack edges of different widths in parallel through multi-scale edge detection branches and refines and corrects them with geometric constraints to obtain a crack detection segmentation map.

7. The crack detection device based on feature enhancement according to claim 6, characterized in that, The enhancement module includes: The elongated unit is used to input the progressively fused feature map into the multi-directional directional convolution and deformable convolution in the elongated structure detection model based on the cross-modal adaptive attention mechanism to perform dual elongated feature enhancement and obtain elongated structure enhanced features; A continuous unit is used to perform spatial scanning and continuous enhancement of the elongated structure enhancement features according to the bidirectional gated loop unit, so as to obtain the crack spatial continuity feature; Topological units are used to input the continuous features of the crack space into the topology-preserving model, and perform parallel prediction through skeleton extraction branches and connectivity modeling branches to obtain skeleton probability maps and connected component probability maps. The processing unit is used to process the skeleton probability map, the connected component probability map and the continuous features of the crack space element by element to obtain a morphological perception feature map.

8. The crack detection device based on feature enhancement according to claim 6, characterized in that, The correction module includes: The first prediction unit is used to predict the edge probability map by performing lightweight convolution on the morphologically aware feature map, and to obtain full-resolution decoded features by performing edge-aware sampling on the smooth regions and edge regions of the predicted edge probability map. The extraction unit is used to extract crack edges of different widths in parallel from the full-resolution decoded features based on the convolution in the multi-scale edge detection branch, so as to obtain comprehensive edge features; The second prediction unit is used to predict the distance from the pixel to the nearest crack boundary in the full-resolution decoded features based on distance field prediction as an auxiliary task, so as to obtain a distance field map. The calculation unit is used to calculate the gradient field pointing to the crack boundary based on the distance field map, and to integrate the geometric prior information of the gradient field into the decoding features to obtain the geometrically constrained decoding features. The segmentation unit is used to perform range field geometric constraints on the geometrically constrained decoding features and coarsely segment the head, and output an initial segmentation mask image; The correction unit is used to perform lightweight residual network iteration to predict the boundary displacement correction amount based on the initial segmentation mask map and the integrated edge features, so as to obtain the crack detection segmentation map.

9. A crack detection device based on feature enhancement, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the feature-enhanced crack detection method as described in any one of claims 1 to 5 when executing the computer program.

10. A medium, characterized in that: The medium stores a computer program that, when executed by a processor, implements the steps of the feature-enhanced crack detection method as described in any one of claims 1 to 5.