External facade crack detection method and device based on multi-scale feature fusion, equipment and medium

Through the crack detection method of multi-scale feature fusion, combined with the MEEMO optimization algorithm and the HMFF feature fusion framework, adaptive crack detection in complex environments is achieved, which improves detection accuracy and efficiency, solves the problems of insufficient detection accuracy and resource waste in existing technologies, and provides reliable technical support for building structure health monitoring.

CN120635015AInactive Publication Date: 2025-09-12刘滨睿
View PDF 0 Cites 18 Cited by

Patent Information

Application Number
CN202510731201.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing crack detection technology lacks detection accuracy and robustness in actual engineering environments with complex lighting, texture interference or high background noise, making it difficult to meet high-standard industry requirements. It also lacks an adaptive multi-scale feature fusion mechanism, resulting in waste of computing resources and low detection efficiency.

Method used

A crack detection method based on multi-scale feature fusion is adopted. Through multi-dimensional complexity evaluation and dynamic classification mechanism, combined with the MEEMO optimization algorithm and the HMFF feature fusion framework, parametric enhancement and multi-threshold fusion processing are performed on low-complexity images, and multi-modal signal enhancement and asymmetric convolution processing are performed on high-complexity images. The crack probability map is generated and subjected to probability normalization and connected domain analysis to produce the final crack detection results.

Benefits of technology

It significantly improves the robustness and accuracy of crack detection, can adaptively process in complex scenarios, effectively suppress noise interference, enhance edge response in small crack areas, extract multi-scale texture features, solve the problem of false detection in dense texture or shadow coverage scenes, ensure the geometric integrity and spatial consistency of the output results, and optimize computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635015A_ABST
    Figure CN120635015A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-scale feature fusion-based facade crack detection method, apparatus and device, and a medium. The method comprises the steps of performing multi-dimensional complexity evaluation on a to-be-detected image, and generating a classification label by combining information entropy, edge density, texture and gradient variance; dividing the image into high / low-complexity data based on a dynamic threshold, and processing the high / low-complexity data by adopting parameterized enhancement and multi-modal feature fusion strategies: for the low-complexity image, suppressing noise through multi-threshold segmentation and morphological optimization, and extracting fine crack features; for a high-complexity image, in combination with multi-scale feature extraction, asymmetric convolution solution and attention weight fusion, crack response under a complex background is enhanced; and finally, carrying out normalization and geometric verification on the two types of probability graphs, and outputting accurate crack positions and forms. According to the invention, through a complexity-driven differential processing mechanism, the detection robustness in a complex illumination and texture interference scene is significantly improved, and the consumption of computing resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of crack detection technology, and in particular to a method, device, equipment and medium for detecting facade cracks by fusion of multi-scale features. Background Art

[0002] Facade crack detection is one of the key technologies in building structural health monitoring and maintenance. By analyzing building surface images, the location, shape, and expansion trend of cracks can be identified, providing an important basis for subsequent maintenance decisions. With the development of computer vision technology, crack detection methods based on image processing have gradually replaced traditional manual inspections, showing significant advantages in efficiency and consistency. Existing methods mostly rely on single-scale feature analysis or fixed-threshold segmentation strategies. Although they can achieve basic detection in simple scenarios, in actual engineering environments with complex lighting, texture interference, or high background noise, the detection accuracy and robustness are greatly reduced, making it difficult to meet the high-standard industry requirements.

[0003] Currently, mainstream crack detection technologies face two major limitations: First, for low-complexity images, such as those with uniform backgrounds and high-contrast cracks, traditional methods are susceptible to noise interference, resulting in missed detection of small cracks or broken edges. Second, for high-complexity images, such as those with dense textures and shadow coverage, existing algorithms struggle to effectively distinguish real cracks from background artifacts, significantly increasing the false detection rate. Furthermore, most methods lack adaptive multi-scale feature fusion mechanisms, making it impossible to dynamically adjust processing strategies based on image complexity, resulting in wasted computing resources and low detection efficiency. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to provide a facade crack detection method, device, equipment and medium that can take into account the complexity of different scenes and realize multi-scale feature fusion of accurate feature extraction and adaptive fusion.

[0005] The purpose of the present invention is achieved by the following scheme:

[0006] In a first aspect, the present invention provides a facade crack detection method using multi-scale feature fusion, comprising the following steps:

[0007] S1: Obtain the image set to be detected input by the user, and calculate the information entropy, edge density, fusion texture complexity and gradient variance of the image set to be detected and perform multi-dimensional complexity evaluation to generate an image dataset with complexity score;

[0008] S2: performing image classification on the image dataset based on a preset dynamic threshold to generate high-complexity image data and low-complexity image data;

[0009] S3: Based on the MEEMO optimization algorithm, parametric enhancement and multi-threshold fusion processing are performed on low-complexity image data to highlight small area features and capture edge changes, optimize the crack area separation process, and generate the first crack probability map. The first crack probability map is used to indicate the crack distribution in low-complexity scenes;

[0010] S4: Based on the HMFF feature fusion framework, multimodal signal enhancement and asymmetric convolution processing are performed on high-complexity image data. A second crack probability map is generated through multi-scale feature extraction, attention weight fusion, and dual-branch decision making. The second crack probability map is used to indicate the crack distribution in high-complexity scenarios.

[0011] S5: Process the first crack probability map and the second crack probability map, and generate crack detection results through Sigmoid normalization and connected domain analysis. The crack detection results are used to indicate the final crack position and geometric shape.

[0012] In one embodiment, S1 of a facade crack detection method provided by the present invention that integrates multi-scale features specifically includes the following steps:

[0013] S11: Based on the information entropy evaluation technology, the information volume of the image set to be detected is analyzed and processed, a smoothing factor is introduced to filter the zero probability items, and an information entropy score is generated;

[0014] S12: Quantify the edge distribution of the image set to be detected based on edge density detection technology, extract the edge area through the gradient operator and calculate the pixel ratio to generate an edge density score;

[0015] S13: Based on texture fusion analysis technology, local and global texture features are extracted from the image set to be detected, and the texture complexity score is generated by combining LBP variance and GLCM contrast;

[0016] S14: performing local structural change analysis on the image set to be detected, and generating a gradient variance score by calculating the gradient amplitude variance;

[0017] S15: Based on the weighted fusion algorithm, the information entropy score, edge density score, texture complexity score and gradient variance score are comprehensively evaluated and processed, and an image dataset containing a complexity score is generated in combination with the image set to be detected.

[0018] In one embodiment, S3 of a facade crack detection method fused with multi-scale features provided by the present invention specifically includes the following steps:

[0019] S31: performing parametric enhancement processing on the low-complexity image data, generating a parametric enhanced image through contrast gain and brightness offset;

[0020] S32: extracting candidate crack regions from the parameterized enhanced image, performing region segmentation using a combination of a global threshold, a fixed ratio threshold, and a local adaptive threshold to generate a candidate mask set;

[0021] S33: Perform noise suppression on the candidate mask set based on the multi-directional morphological optimization technique, and perform opening operations using 9×1, 1×9, and 5×5 structure elements to generate denoising masks;

[0022] S34: Perform multi-feature fusion processing on the denoising mask, combine dark area enhancement, gradient intensity calculation and skeleton structure extraction to generate a first crack probability map.

[0023] In one embodiment, S32 of the facade crack detection method provided by the present invention incorporating multi-scale features specifically includes the following steps:

[0024] S321: Perform full-image threshold segmentation on the parameterized enhanced image based on the global threshold segmentation technique, maximize the inter-class variance of the image foreground and background regions, and generate a global threshold mask;

[0025] S322: performing grayscale truncation processing on the parametric enhanced image, dividing high and low grayscale areas according to a preset ratio, and generating a fixed threshold mask;

[0026] S323: Perform dynamic threshold calculation processing on the parameterized enhanced image and generate an adaptive threshold mask according to the local neighborhood grayscale distribution;

[0027] S324: Based on the multi-mask fusion technology, the global threshold mask, the fixed threshold mask and the adaptive threshold mask are jointly processed to generate a candidate mask set.

[0028] In one embodiment, S4 of the facade crack detection method provided by the present invention incorporating multi-scale features specifically includes the following steps:

[0029] S41: Perform color space conversion on high-complexity image data, separate brightness information to extract LAB brightness channel and multi-level CLAHE processing to generate multi-scale brightness feature map;

[0030] S42: Based on the direction-sensitive feature extraction technology, the multi-scale brightness feature map is edge-enhanced. By using Gaussian difference filtering and introducing spatial anisotropic sharpening, a direction-sensitive feature map is generated.

[0031] S43: Perform multi-directional feature extraction on the direction-sensitive feature map based on asymmetric convolution decomposition technology, capturing horizontal, vertical and omnidirectional features to generate a multi-directional feature map;

[0032] S44: Based on context-aware technology, spatial relationship modeling is performed on multi-directional feature maps, and multi-angle dilated convolution is introduced to generate spatial context feature maps;

[0033] S45: processing the spatial context feature map based on a preset adaptive feature fusion mechanism to generate a multi-scale fusion feature map;

[0034] S46: Predict the crack probability of the multi-scale fusion feature map, generate normalized weights through 1×1 convolution and Softmax function, dynamically adjust the contribution of features at different scales according to the feature responses at different positions, and generate a second crack probability map.

[0035] In one embodiment, S45 of the facade crack detection method provided by the present invention incorporating multi-scale features specifically includes the following steps:

[0036] S451: performing channel compression processing on the spatial context feature map, extracting channel mean features and channel extreme value features respectively through global average pooling and global maximum pooling, and generating a channel statistical feature map;

[0037] S452: Perform nonlinear mapping on the channel statistical feature map, realize channel dimension compression and feature interaction through multi-layer perceptron, and generate channel attention weight. The expression of channel attention weight is:

[0038] w c =σ*(MLP(GAP(F))+MLP(GMP(F)))

[0039] Among them, w c is the channel attention weight, σ is the Sigmoid activation function, MLP is the multi-layer perceptron operation, which means nonlinear transformation of data, GAP(F) represents the global average pooling result of the input spatial context feature map F in the spatial dimension, and GMP(F) represents the global maximum pooling result of the input spatial context feature map F in the spatial dimension;

[0040] S453: Perform spatial domain conversion on the spatial context feature map, restore the channel dimension through convolution operation, and generate spatial attention weights. The expression of spatial attention weights is:

[0041] w s =Conv(Concat(AvgPool(F),MaxPool(F)))

[0042] Among them, w sis the spatial attention weight, Conv is the convolution operation, Concat is the channel dimension splicing operation, AvgPool(F) is the local spatial average pooling result of F, and MaxPool(F) is the local spatial maximum pooling result of F;

[0043] S454: Dynamically weight the channel attention weights and spatial attention weights, multiply the elements together, and weight them element-by-element with the spatial context feature map to generate a multi-scale fusion feature map.

[0044] In one embodiment, S5 of the facade crack detection method provided by the present invention incorporating multi-scale features specifically includes the following steps:

[0045] S51: performing normalization processing on the first crack probability map and the second crack probability map based on a probability normalization technology, mapping the probability values ​​to a preset interval through a nonlinear function, and generating a standard probability map;

[0046] S52: Performing connected domain extraction processing on the standard probability map, generating a set of candidate connected domains through region growing and boundary tracking;

[0047] S53: performing noise suppression processing on the candidate connected domain set, and generating a denoised connected domain set by using corrosion and expansion operations of structural elements in different directions;

[0048] S54: Perform post-processing verification on the denoised connected domain set, and generate crack detection results through area threshold and aspect ratio screening.

[0049] In a second aspect, the present invention provides a facade crack detection device with multi-scale feature fusion, which is configured with the following modules:

[0050] The data acquisition and scoring module is used to obtain the image set to be detected input by the user, and calculate the information entropy, edge density, fusion texture complexity and gradient variance of the image set to be detected, and perform multi-dimensional complexity evaluation to generate an image dataset containing complexity scores;

[0051] A complexity classification module is used to classify the image data set based on a preset dynamic threshold to generate high-complexity image data and low-complexity image data;

[0052] The MEEMOModule module is used to perform parametric enhancement and multi-threshold fusion processing on low-complexity image data based on the MEEMO optimization algorithm, highlighting small area features and capturing edge changes, optimizing the crack area separation process, and generating the first crack probability map. The first crack probability map is used to indicate the crack distribution in low-complexity scenes;

[0053] The HMFFFramework module is used to perform multimodal signal enhancement and asymmetric convolution processing on high-complexity image data based on the HMFF feature fusion framework. It generates a second crack probability map through multi-scale feature extraction, attention weight fusion, and dual-branch decision-making. The second crack probability map is used to indicate the crack distribution in high-complexity scenarios;

[0054] The normalization and analysis module is used to process the first crack probability map and the second crack probability map, and generate crack detection results through Sigmoid normalization and connected domain analysis. The crack detection results are used to indicate the final crack position and geometric shape.

[0055] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements any one of the above-mentioned multi-scale feature fusion facade crack detection methods.

[0056] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned facade crack detection methods using multi-scale feature fusion.

[0057] In summary, the facade crack detection method with multi-scale feature fusion provided by the present invention can realize adaptive processing in complex scenes through multi-dimensional complexity evaluation and dynamic classification mechanism, and significantly improve the robustness and accuracy of facade crack detection. For low-complexity images, the processing flow based on parameterized enhancement and multi-threshold fusion can effectively suppress noise interference, strengthen the edge response and continuity of small crack areas, and accurately capture tiny cracks in high-contrast environments, overcoming the edge breakage or missed detection problems caused by fixed threshold segmentation in traditional methods; for high-complexity images, combined with multimodal signal enhancement and asymmetric convolution decomposition technology, it can effectively extract multi-scale texture features and suppress background artifacts, and achieve a dynamic balance of deep and shallow features through attention weight fusion and dual-branch decision-making mechanism to solve the problem of false detection of cracks and background confusion in scenes with dense textures or shadow coverage. Further, through probability normalization and geometric continuity verification, it can achieve fine reconstruction of crack morphology to ensure the geometric integrity and spatial consistency of the output results. This method uses a complexity-driven differentiated processing strategy to achieve a reasonable allocation of computing resources, optimize computing efficiency while ensuring high-precision detection, overcome the shortcomings of traditional algorithms in adapting to complex engineering environments, and provide reliable technical support for building structure health monitoring.

[0058] For better understanding and implementation, the present invention is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1A schematic flow chart of a facade crack detection method using multi-scale feature fusion provided in an embodiment of the present application;

[0060] Figure 2 A module structure diagram of the ARSO unit provided in an embodiment of the present application;

[0061] Figure 3 A schematic diagram of a process for generating a first crack probability map provided in an embodiment of the present application;

[0062] Figure 4 A black hat transformation slice diagram provided in an embodiment of the present application;

[0063] Figure 5 A schematic diagram of a process for generating a second crack probability map provided in an embodiment of the present application;

[0064] Figure 6 A schematic structural diagram of a facade crack detection device with multi-scale feature fusion provided in another embodiment of the present application. DETAILED DESCRIPTION

[0065] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate preferred embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure.

[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0067] In one embodiment, Figure 1 As shown, a method for detecting facade cracks using multi-scale feature fusion is provided. This embodiment uses the method applied to a terminal as an example. It is understandable that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0068] S1: Obtain the image set to be detected input by the user, and calculate and evaluate the information entropy, edge density, fusion texture complexity and gradient variance of the image set to be detected, and generate an image dataset with complexity score.

[0069] Specifically, the system first obtains the image set to be detected input by the user, and then performs a multidimensional complexity assessment on each image in the image set. During this process, the system calculates the information entropy, edge density, fused texture complexity, and gradient variance of the image respectively. Information entropy is used to measure the complexity of the pixel grayscale distribution in the image. Its calculation is based on the Shannon entropy formula and is obtained through statistical analysis of the image grayscale values. Edge density reflects the richness of edge features in the image. The system determines the edge pixels in the image through the edge detection algorithm and calculates their proportion in the entire image. Fusion texture complexity combines two texture analysis methods, local binary pattern (LBP) and gray-level co-occurrence matrix (GLCM), to comprehensively evaluate the texture characteristics of the image. Among them, LBP is mainly used to capture the non-uniformity of local texture, while GLCM focuses more on analyzing the spatial relationship of texture.

[0070] Gradient variance reflects the severity of grayscale changes between pixels in an image. The system determines this by calculating the image's gradient amplitude. After calculating the four aforementioned metrics, the system then weights and fuses them based on preset weighting coefficients to generate a comprehensive complexity score. This score comprehensively reflects the overall complexity of the image, thereby constructing an image dataset containing complexity scores, providing a quantitative basis for subsequent classification based on varying levels of complexity.

[0071] S2: Perform image classification on the image dataset based on a preset dynamic threshold to generate high-complexity image data and low-complexity image data.

[0072] Specifically, low-complexity image data typically features a relatively simple background and high crack-background contrast, making crack features more distinct and easier to detect. High-complexity image data, on the other hand, often contains multiple interfering factors, such as complex background textures, dense patterns, shadow coverage, and uneven lighting. This makes it difficult to distinguish crack features from the background, significantly increasing the difficulty of detection. By classifying image datasets, the system can adopt more targeted processing strategies for image data of varying complexity, thereby ensuring detection accuracy while improving overall detection efficiency and providing more accurate data support for subsequent crack detection processes.

[0073] Preferably, the dynamic threshold setting can be flexibly adjusted based on the overall statistical characteristics of the image dataset and the needs of the actual application scenario to ensure that the classification results accurately reflect the differences in image complexity. The system compares each image in the image dataset with the dynamic threshold and classifies the images into high-complexity and low-complexity image data based on the complexity score.

[0074] S3: Based on the MEEMO optimization algorithm, parametric enhancement and multi-threshold fusion processing are performed on low-complexity image data to highlight small area features and capture edge changes, optimize the crack area separation process, and generate the first crack probability map. The first crack probability map is used to indicate the crack distribution in low-complexity scenarios.

[0075] Specifically, the system performs parametric image enhancement, adjusting parameters such as contrast and brightness to highlight detailed features and enhance the visibility of small areas, thereby better capturing edge variations. Furthermore, the system performs multi-threshold fusion processing, segmenting the image using multiple thresholds. Combined with edge detection and other techniques, the system extracts crack features at different levels, achieving precise separation of crack regions. Specifically, the system uses a black-hat transform to highlight small areas darker than the background, enhancing the characteristics of faint cracks. It then uses Canny edge detection and the Sobel operator to extract edge information and gradient features, respectively, and then fuses these features. Canny edge detection uses thresholds of 60 and 140 to detect edges in the preprocessed image, and then fuses these features with the morphological results using cv2.bitwise_or. This combination fully utilizes edge information to compensate for details lost in morphological processing, improving the perception of shallow cracks. The Sobel operator uses a 3x3 kernel to calculate the x- and y-direction gradients of the original grayscale image. The combined gradient magnitudes are then normalized to [0, 1], providing additional texture information for probability map generation. Finally, the extraction results of the crack area are further optimized through morphological operations and connected domain analysis to generate the first crack probability map. This probability map can clearly indicate the distribution of cracks in low-complexity scenarios. The probability value of each pixel reflects the possibility that the location belongs to a crack.

[0076] S4: Based on the HMFF feature fusion framework, multimodal signal enhancement and asymmetric convolution processing are performed on high-complexity image data. A second crack probability map is generated through multi-scale feature extraction, attention weight fusion and dual-branch decision making. The second crack probability map is used to indicate the crack distribution in high-complexity scenarios.

[0077] Specifically, the system converts the image from BGR to LAB color space through color space transformation and implements multi-level CLAHE processing with parameterized differences on the luminance channel. This parameter gradient processing effectively addresses the scale uncertainty and contrast variability of crack features. Furthermore, to address edge frequency characteristics, the system introduces a multi-scale Difference of Gaussians (DoG) filter bank to construct a bandpass filter response set and optimize the frequency domain representation of crack edges. Furthermore, the introduction of directional convolution kernels enhances the directional selectivity of linear structures and improves edge location accuracy.

[0078] In the feature extraction stage, such as Figure 2 As shown, for the original U 2 -NET symmetric convolution design is not good at extracting directional features of cracks. This embodiment proposes an asymmetric recursive neural unit (ARSO). The system uses ARSO to combine heterogeneous convolution decomposition with multi-angle dilated convolution to enhance the model's perception of crack directional characteristics and enhance the sensitivity to directional features, thereby achieving multi-scale feature extraction of crack features. The attention weight fusion mechanism combines channel attention and spatial attention to dynamically weight the extracted multi-scale features, highlight crack-related features, and suppress background noise. Finally, by coordinating the advantages of traditional image processing and deep learning, the traditional processing branch and the deep learning branch independently generate crack detection results, and then perform regional adaptive weight fusion to generate a second crack probability map. This probability map is used to indicate the distribution of cracks in highly complex scenes, effectively improving the detection accuracy of cracks in complex scenes.

[0079] S5: Process the first crack probability map and the second crack probability map, and generate crack detection results through Sigmoid normalization and connected domain analysis. The crack detection results are used to indicate the final crack position and geometric shape.

[0080] Specifically, the system normalizes the probability map using the Sigmoid function, mapping the probability values ​​to the interval [0,1], making the probability values ​​more comparable and interpretable. Next, the system performs a connected domain analysis, screening out connected domains that meet the crack characteristics based on preset conditions such as area, aspect ratio, and length. During the connected domain analysis process, the system analyzes the shape, size, and other features of each connected domain to determine the location of the crack, and obtains the geometric morphology of the crack by analyzing features such as the shape change of the connected domain. Finally, the system obtains the crack detection results based on the above processing. The results indicate the final location of the cracks in the image and the geometric morphology of the cracks in an intuitive way, providing an accurate basis for subsequent building structure health assessments and maintenance decisions. In this way, the system can effectively integrate crack detection information in scenarios of different complexity and generate comprehensive and accurate crack detection results.

[0081] In summary, the facade crack detection method with multi-scale feature fusion provided by the present invention can realize adaptive processing in complex scenes through multi-dimensional complexity evaluation and dynamic classification mechanism, and significantly improve the robustness and accuracy of facade crack detection. For low-complexity images, the processing flow based on parameterized enhancement and multi-threshold fusion can effectively suppress noise interference, strengthen the edge response and continuity of small crack areas, and accurately capture tiny cracks in high-contrast environments, overcoming the edge breakage or missed detection problems caused by fixed threshold segmentation in traditional methods; for high-complexity images, combined with multimodal signal enhancement and asymmetric convolution decomposition technology, it can effectively extract multi-scale texture features and suppress background artifacts, and achieve a dynamic balance of deep and shallow features through attention weight fusion and dual-branch decision-making mechanism to solve the problem of false detection of cracks and background confusion in scenes with dense textures or shadow coverage. Further, through probability normalization and geometric continuity verification, it can achieve fine reconstruction of crack morphology to ensure the geometric integrity and spatial consistency of the output results. This method uses a complexity-driven differentiated processing strategy to achieve a reasonable allocation of computing resources, optimize computing efficiency while ensuring high-precision detection, overcome the shortcomings of traditional algorithms in adapting to complex engineering environments, and provide reliable technical support for building structure health monitoring.

[0082] In one embodiment, S1 of a facade crack detection method provided by the present invention that integrates multi-scale features specifically includes the following steps:

[0083] S11: Based on the information entropy evaluation technology, the information amount of the image set to be detected is analyzed and processed, a smoothing factor is introduced to filter the zero probability items, and an information entropy score is generated.

[0084] Specifically, during this process, the system uniformly converts the input color images in RGB or BGR format into grayscale images. This conversion simplifies the computational complexity by reducing the dimensionality from three channels to a single channel, while retaining key information about the distribution of grayscale values, laying the foundation for subsequent steps such as information entropy, edge detection, and texture analysis. Compared with complex methods that directly process color images, this grayscale strategy avoids the interference of color noise on complexity indicators while maintaining evaluation efficiency, and exhibits higher stability and consistency, especially when processing images with diverse lighting conditions. Image preprocessing and basic transformation provide a reliable input basis for subsequent complexity evaluation through efficient color space conversion and data preparation, reflecting the importance of data standardization in computer vision.

[0085] After completing image preprocessing, the system analyzes the information content of the image set to be detected based on information entropy evaluation technology. The information entropy calculation is based on the Shannon entropy formula. The system generates a 256-level grayscale histogram through an efficient method, then uses vectorized operations to calculate the probability distribution and introduces a smoothing factor to filter zero probability items to avoid the mathematical error of log(0). Finally, the entropy value is calculated. The calculation formula for the information entropy score is:

[0086]

[0087] Where H is the information entropy score, p(i) represents the probability of a pixel with grayscale value i appearing in the image, h(i) is the number of pixels with grayscale value i, and N is the total number of pixels in the image.

[0088] Compared to traditional pixel-by-pixel statistics, this vectorized implementation significantly improves computational efficiency. It also reduces memory usage through a histogram approach, making the system suitable for large-scale image processing. Entropy reflects the uniformity of grayscale distribution, with high entropy values ​​corresponding to complex images and low entropy values ​​indicating simple structures. The system normalizes entropy values ​​to the range [0, 10] (based on the maximum entropy value of 8 for 8-bit images) to generate an information entropy score, which provides a quantitative basis for subsequent comprehensive evaluation.

[0089] S12: Based on the edge density detection technology, the edge distribution of the image set to be detected is quantified. The edge area is extracted through the gradient operator and the pixel ratio is calculated to generate an edge density score.

[0090] Specifically, the system first performs image preprocessing to convert the input image into a grayscale image. It then implements Canny edge detection to extract edge and gradient features. Canny edge detection processes grayscale images with thresholds of 50 and 150, and combines Gaussian smoothing, gradient calculation, and dual-threshold filtering to generate an edge map. The edge pixel ratio is calculated and normalized to the [0, 10] range. The edge density score is calculated as follows:

[0091]

[0092] Among them, D edge represents the edge density score, ∑ x,y E(x,y) represents the sum of the edge pixel values ​​E(x,y) at all coordinates (x,y) in the image. Here, E(x,y) takes the value 1 at edges and 0 at non-edges. The sum is the total number of edge pixels in the image. Compared to simple Sobel edge detection methods, Canny's non-maximum suppression and connectivity analysis offer advantages in detecting subtle structures and provide a reliable measure of edge density.

[0093] S13: Based on texture fusion analysis technology, local and global texture features are extracted from the image set to be detected, and the texture complexity score is generated by combining LBP variance and GLCM contrast.

[0094] Specifically, when the system performs local and global texture feature extraction, it first generates a texture description map through the local binary pattern (LBP). Specifically, the system uses a certain radius and number of sampling points as parameters, uses a specific function to generate a texture description map, and calculates the LBP histogram. After that, the texture complexity is quantified by the uniformity formula. The expression of texture complexity is 10(1-uniformity), where uniformity is an indicator of texture uniformity. This method can capture the non-uniformity of local textures in more detail and is suitable for complexity assessment of complex surfaces.

[0095] At the same time, the system also calculates the gray-level co-occurrence matrix (GLCM), setting the distance to 1, the angles to 0°, 45°, 90°, and 135°, and the grayscale to 256. Contrast and correlation features are extracted from the co-occurrence matrix, the contrast is normalized to the interval [0,5] (divided by 100), and the correlation is converted to a complexity index using 5×(1-correlation). Compared with a single statistical feature, GLCM has more advantages in spatial relationship analysis and reflects the contrast and consistency of textures. Among them, the calculation formula for the texture complexity score is:

[0096]

[0097] Among them, C texture is the texture complexity index, U LBP is the uniformity index calculated based on the local binary pattern, C GLCM is the contrast eigenvalue calculated by GLCM, R GLCM is the correlation eigenvalue calculated by GLCM.

[0098] Finally, the system averages and fuses the texture features obtained from LBP and GLCM, restricting them to the interval [0, 10], to generate a texture complexity score. This fusion method combines local and global texture information to more comprehensively and accurately reflect the texture complexity of the image, providing rich and accurate texture quantification results for subsequent comprehensive evaluation.

[0099] S14: Perform local structural change analysis on the image set to be detected, and generate a gradient variance score by calculating the gradient amplitude variance.

[0100] Specifically, the Sobel operator extracts the x- and y-direction gradients using a 3x3 kernel in the gradient variance calculation. The gradient magnitudes are then combined to calculate the variance, which is then normalized to the range [0, 10]. Compared to direct grayscale differentiation, this method more accurately captures dynamic changes between pixels and reflects the contrast complexity of the image. The combination of these two methods assesses complexity from two dimensions: structural density and gradient distribution. This ensures comprehensive and robust feature extraction and provides multidimensional data support for subsequent comprehensive scoring.

[0101] S15: Based on the weighted fusion algorithm, the information entropy score, edge density score, texture complexity score and gradient variance score are comprehensively evaluated and processed, and an image dataset containing a complexity score is generated in combination with the image set to be detected.

[0102] Specifically, the system integrates the information entropy score, edge density score, texture complexity score and gradient variance score through a weighted formula (the default entropy weight is 0.4, edge 0.25, texture 0.25, and gradient 0.1). The entropy value is normalized to [0,10] based on the maximum theoretical value of 8, and other indicators have been standardized to the same range. The final score is compared with the standardized threshold and the "high" or "low" complexity classification is output. This weighted strategy is relatively simple averaging and can flexibly reflect the importance of each indicator and image characteristics, providing a scientific basis for the dynamic selection of crack detection modules.

[0103] During the comprehensive evaluation process, the system first normalizes each scoring item to ensure they are within the same dimensional range. Then, based on preset weight coefficients, the normalized scores are weighted and summed to calculate a comprehensive complexity score for each image. Finally, the system combines these scores with the set of images to be inspected to generate an image dataset with complexity scores. This dataset not only contains the original image data but also attaches a complexity score to each image, providing a critical reference for subsequent image classification and crack detection. In this way, the system can rationally allocate computing resources and select appropriate processing strategies based on the image complexity score, thereby improving the efficiency and accuracy of the entire crack detection process.

[0104] In one embodiment, Figure 3 As shown, S3 of the facade crack detection method provided by the present invention for multi-scale feature fusion specifically includes the following steps:

[0105] S31: Performing parametric enhancement processing on the low-complexity image data to generate a parametric enhanced image through contrast gain and brightness offset.

[0106] Specifically, in the image preprocessing and enhancement stage, the input color image is first converted into a grayscale image through a color space conversion method, which simplifies subsequent calculations while retaining key brightness information, providing a stable input basis for subsequent processing steps. Subsequently, a parameterized linear transformation contrast enhancement method (alpha = 1.9, beta = 14) is used to stretch the grayscale value range of the grayscale image. Compared with the traditional global equalization technology, the visibility of shallow cracks is significantly improved, and noise interference is effectively controlled through parameter optimization. The image is further subjected to local contrast equalization processing through a block adaptive enhancement strategy. Compared with the global method, this strategy accurately enhances the crack details under complex backgrounds while avoiding excessive enhancement of background areas. Then, a smoothing method combining color and spatial distance weighting is used to effectively suppress noise while maintaining the sharpness of the crack edges, providing a high-quality image foundation for subsequent processing.

[0107] In the feature extraction stage, the system uses black hat transformation as the core feature extraction method, such as Figure 4 As shown in the figure, the system uses a 9x9 elliptical structuring element to highlight small areas in the image that are darker than the background. It then uses contrast enhancement technology to further enhance the characteristics of faint cracks. Compared with traditional methods, this strategy is more targeted at detecting dark cracks in complex backgrounds. The coordinated design of these preprocessing and enhancement techniques significantly improves the module's robustness and detection sensitivity in processing low-contrast and noisy images, fully covering the full process optimization from basic brightness preservation and noise suppression to feature enhancement.

[0108] S32: Extract candidate crack regions from the parameterized enhanced image, perform region segmentation using a combination of a global threshold, a fixed ratio threshold, and a local adaptive threshold to generate a candidate mask set.

[0109] Specifically, the candidate mask set is obtained through the following steps:

[0110] S321: Perform full-image threshold segmentation on the parametric enhanced image based on the global threshold segmentation technique to maximize the inter-class variance of the image foreground and background regions and generate a global threshold mask.

[0111] Specifically, the system divides the pixels in the image into two categories, foreground and background, by calculating the grayscale histogram of the image. Specifically, the system achieves the best separation effect between the foreground cracks and background non-cracks based on a preset optimal separation threshold. During the calculation process, the system analyzes the pixel distribution of each grayscale level in the image and looks for the grayscale value that can maximize the inter-class variance as the threshold. The larger the inter-class variance, the more obvious the distinction between the foreground and the background, and the better the segmentation effect. The global threshold mask generated in this way can provide preliminary crack area division results for subsequent processing steps, helping the system to preliminarily lock in the possible locations of cracks. This technology is widely used in the field of image segmentation because it can effectively separate the target object and the background on a global scale, laying the foundation for subsequent feature extraction and analysis.

[0112] S322: performing grayscale truncation processing on the parameterized enhanced image, dividing high and low grayscale areas according to a preset ratio, and generating a fixed threshold mask.

[0113] Specifically, the system pre-sets a fixed ratio based on experience or specific scene requirements, and divides the grayscale value range of the image into high grayscale areas and low grayscale areas. Generally, high grayscale areas are considered to be potential crack areas, while low grayscale areas are considered to be background or other non-crack feature areas. During the processing process, the system will traverse each pixel in the image and compare its grayscale value with the preset high and low grayscale boundaries to divide the pixel into high grayscale or low grayscale areas. The advantage of this processing method is that it can quickly determine the areas in the image that may contain cracks, providing a preliminary screening result for subsequent detailed analysis. The setting of a fixed ratio is of great significance for the extraction of crack features under stable conditions, especially when the image quality and lighting conditions are relatively consistent, which can effectively improve the efficiency and accuracy of crack detection.

[0114] S323: Perform dynamic threshold calculation processing on the parameterized enhanced image and generate an adaptive threshold mask according to the local neighborhood grayscale distribution.

[0115] Specifically, in this embodiment, different areas of the image may be affected by different degrees of illumination, resulting in large differences in the grayscale value distribution of local areas. Traditional global threshold methods often have difficulty accurately segmenting crack features in this situation because they cannot adapt to these local changes. However, through dynamic threshold calculation, the system can calculate the threshold that is most suitable for each local area based on the characteristics of the area, thereby ensuring that candidate crack areas can be accurately extracted under different lighting conditions. During the calculation process, the system analyzes the grayscale distribution of pixels in each local neighborhood, taking into account factors such as its average grayscale and standard deviation, and thus determines an optimal threshold for each local area. This adaptive threshold generation mechanism makes the system more flexible and accurate when processing images under complex lighting conditions, providing a more reliable mask basis for subsequent crack detection.

[0116] S324: Based on the multi-mask fusion technology, the global threshold mask, the fixed threshold mask and the adaptive threshold mask are jointly processed to generate a candidate mask set.

[0117] Specifically, a logical OR operation is performed on the global threshold mask, the fixed threshold mask, and the adaptive threshold mask to comprehensively judge the status of each pixel in the three masks. If a pixel is marked as a foreground area in any mask, it will also be marked as a foreground area in the final candidate mask set. This method ensures that no possible crack features are missed, while minimizing false detections and missed detections. The formula is as follows:

[0118] B(x,y)=B OTSU-low (x,y)∨B OTSU (x,y)∨B adaptive (x,y)

[0119] Among them, B(x,y) represents the pixel state in the final generated candidate mask set, B OTSU-low (x,y) and B OTSU (x, y) represent the low threshold and high threshold masks calculated based on the OTSU method, respectively. adaptive (x,y) represents the adaptive threshold mask. This allows the system to more comprehensively and accurately identify potential crack areas in the image, providing more options for areas likely to contain cracks in subsequent processing steps, ensuring the accuracy and reliability of detection results. The application of multi-mask fusion technology not only improves the robustness of crack detection but also enhances the system's adaptability to complex image conditions, making crack detection results more stable and reliable.

[0120] S33: Perform noise suppression on the candidate mask set based on multi-directional morphological optimization technology, and perform opening operations using 9×1, 1×9, and 5×5 structure elements to generate denoising masks.

[0121] Specifically, the opening operation, through erosion followed by dilation, effectively removes small noise points and discontinuous areas from the image while preserving the continuity and integrity of linear features such as cracks. 9×1 and 1×9 linear structuring elements perform morphological operations in the horizontal and vertical directions, respectively, effectively suppressing noise in these directions. The 5×5 square structuring element is suitable for processing noise in diagonal or other complex directions. By performing multi-directional morphological optimization using these structuring elements of different shapes and sizes, the system can more comprehensively suppress noise, generate cleaner and clearer denoising masks, and reduce false detections and missed detections in subsequent processing steps.

[0122] S34: Perform multi-feature fusion processing on the denoising mask, combine dark area enhancement, gradient intensity calculation and skeleton structure extraction to generate a first crack probability map.

[0123] Specifically, dark area enhancement technology makes dark crack features that were originally difficult to distinguish more obvious by increasing the brightness and contrast of dark areas in the image, thereby improving the detectability of cracks. Gradient intensity calculation highlights areas in the image with drastic grayscale changes by calculating the gradient amplitude of each pixel in the image. These areas are usually the edges of the cracks or the cracks themselves. Skeleton structure extraction technology is used to extract the center line of the crack and form the skeleton structure of the crack, which helps in the subsequent analysis of the shape and direction of the crack. By fusing these different types of features, the system can generate a first crack probability map, which intuitively represents the possibility that each pixel in the image belongs to a crack in the form of probability, providing an important basis for subsequent crack detection and analysis, helping the system to more accurately identify and locate crack features in low-complexity images, and ensure the reliability and accuracy of the detection results.

[0124] The above-mentioned facade crack detection method with multi-scale feature fusion can achieve refined detection and complete reconstruction of geometric features of small cracks in low-complexity images through the innovative design of MEEMOModule, so as to solve the problems of missed detection and edge fracture caused by noise interference in traditional methods under uniform background or high-contrast scenes. Based on the parameterized preprocessing process, through the linear transformation optimization of contrast gain and brightness offset, it can effectively enhance the visibility of shallow crack areas, and combine adaptive histogram equalization and nonlinear filtering to achieve uneven illumination correction and noise suppression, providing enhanced images with high signal-to-noise ratio for subsequent feature extraction; through the multi-level threshold fusion strategy, combined with global, fixed-ratio and local adaptive segmentation technology, it can achieve full coverage extraction of crack candidate areas, overcoming the lack of sensitivity of single threshold segmentation to weak edges or low-contrast cracks;

[0125] Furthermore, the system introduces multi-directional morphological optimization technology, utilizing directional erosion and dilation operations on linear structuring elements to effectively filter out isolated noise points and maintain the topological continuity of crack edges, avoiding the geometric distortion caused by traditional square structuring elements. Finally, through a multi-feature fusion strategy, combined with dark area enhancement, gradient intensity calculation, and skeletonization processing, it is possible to achieve dynamic weighted generation of crack probability maps, retaining the original probability information in the central region while weighting the expanded region, thereby improving detection sensitivity while ensuring the geometric integrity of the crack morphology. Through the coordinated optimization of multi-level enhancement, noise suppression, and feature fusion, this process can significantly improve the robustness of crack detection in low-complexity scenarios, providing high-precision basic data support for building apparent health assessment.

[0126] In one embodiment, Figure 5 As shown, S4 of the facade crack detection method provided by the present invention with multi-scale feature fusion specifically includes the following steps:

[0127] S41: Perform color space conversion on high-complexity image data, separate brightness information to extract LAB brightness channel and multi-level CLAHE processing, and generate multi-scale brightness feature map.

[0128] Specifically, the system performs color space conversion on highly complex image data, converting the image from its original color space to the LAB color space to separate luminance information and extract the LAB luminance channel. The LAB color space is device-independent and better represents human visual perception of color. Through this conversion, the system can effectively separate the image's luminance and chrominance information, facilitating subsequent luminance processing. The system then performs multi-level adaptive histogram equalization (CLAHE) on the luminance channel.

[0129] CLAHE enhances image contrast by performing histogram equalization on different regions of the image, making details more visible. During this process, the system uses a multi-stage CLAHE process with parameterized differences, setting different parameters to account for scale uncertainty and contrast variability in crack features, generating a multi-scale brightness feature map.

[0130] S42: Based on the direction-sensitive feature extraction technology, the multi-scale brightness feature map is edge-enhanced, and the direction-sensitive feature map is generated by Gaussian difference filtering and the introduction of spatial anisotropic sharpening.

[0131] Specifically, the system processes the feature map using a Gaussian difference filter to highlight edge and contour information in the image. The Gaussian difference filter is an edge detection filter that effectively enhances edge features in an image while suppressing noise. To further enhance the directional characteristics of edges, the system introduces spatial anisotropic sharpening technology. This technology enhances the gradient information of the image in a specific direction, making the edges clearer and sharper, thereby generating a direction-sensitive feature map. The direction-sensitive feature map can highlight edge features in different directions in the image, providing support for subsequent directional analysis and crack detection.

[0132] Specifically, the system introduces a multi-scale Difference of Gaussian (DoG) filter bank to construct a bandpass filter response set to optimize the frequency domain representation of crack edges. The directional convolution kernel introduces spatial anisotropic sharpening, enhancing the directional selectivity of linear structures and improving edge location accuracy. The specific DoG function is defined as follows:

[0133]

[0134] Here, σ1 and σ2 represent the standard deviations of two Gaussian filters of different scales, respectively. Finally, the multi-channel enhancement map generated by the module serves as an extended input to the deep network, forming an end-to-end optimization path from feature enhancement to feature extraction. The multi-channel enhanced feature map is fused with the original image and then input into the deep network, forming an end-to-end optimization path from preprocessing to feature learning. By jointly optimizing illumination robustness and directional sensitivity, the module improves the signal-to-noise ratio at crack edges, providing high-purity feature input for subsequent models to accurately locate microcracks in complex backgrounds.

[0135] S43: Based on the asymmetric convolution decomposition technology, multi-directional feature extraction is performed on the direction-sensitive feature map to capture horizontal, vertical and omnidirectional features and generate a multi-directional feature map.

[0136] Specifically, if Figure 2 As shown in the figure, the system performs multi-directional feature extraction on the direction-sensitive feature map based on the asymmetric convolution decomposition technology. In this process, the system uses ARSO to significantly improve the perception of crack directional characteristics through the combination of special-shaped convolution decomposition and multi-angle dilated convolution. The traditional standard 3×3 convolution is decomposed into a parallel combination of three convolution kernels: 1×3, 3×1 and 3×3. This decomposition method not only reduces the number of parameters, but also enhances the ability to extract directional features, which are used to capture horizontal, vertical and omnidirectional features respectively. In the bottleneck layer of the ARSO module, the system introduces rotation invariance enhancement technology, and constructs a full-range direction perception network through multi-angle (0°, 45°, 90°, 135°) dilated convolution, which effectively overcomes the problem of insufficient accuracy in diagonal crack detection. The formula for dilated convolution is expressed as:

[0137] (F* d k)(p)=∑ s+dt=p F(s)*k(t)

[0138] Among them, F represents the input feature map, k represents the convolution kernel, and d is the expansion rate, which controls the spacing between the convolution kernel elements, thereby expanding the receptive field of the convolution operation and capturing a wider range of contextual information.

[0139] S44: Based on context-aware technology, spatial relationship modeling is performed on multi-directional feature maps, and multi-angle dilated convolution is introduced to generate spatial context feature maps.

[0140] Specifically, context-aware technology aims to capture the spatial relationships and contextual information between pixels in an image, which is crucial for understanding the overall structure and spatial distribution of crack features. Multi-angle dilated convolution further expands the receptive field of the convolution kernel by using different dilation rates and angles in the convolution operation, enabling the network to capture the relationship between pixels at a greater distance. Compared with traditional convolution operations, this method can effectively capture long-range dependencies in the image without increasing too much computational complexity, thereby generating a feature map containing rich spatial contextual information. This feature map not only highlights the crack features, but also incorporates information about its surrounding environment, providing a more comprehensive feature representation for subsequent feature fusion and crack detection. In this way, the system can more accurately identify and locate crack features in highly complex images, ensuring the accuracy and reliability of the detection results.

[0141] S45: Processing the spatial context feature map based on a preset adaptive feature fusion mechanism to generate a multi-scale fusion feature map.

[0142] Specifically, the adaptive feature fusion mechanism aims to dynamically adjust the weights between different feature maps to achieve the optimal feature fusion effect. The system dynamically evaluates the importance of features by constructing a feature reweighting module enhanced by channel-space dual attention. The channel attention branch captures channel statistical information through global average pooling and maximum pooling, and generates a channel weight vector through the MLP network to realize the importance evaluation of the feature channel. The spatial attention branch generates a spatial weight map through channel compression and convolution operations to highlight the spatial distribution characteristics of the crack area. The two attention mechanisms work together to form a channel-space joint attention module, which adaptively reweights the input features, effectively suppresses background noise, and highlights crack-related features. When fusing cross-scale features, the system introduces a scale-adaptive weighting strategy to dynamically adjust the contribution of features of different scales according to the feature responses at different positions to generate a multi-scale fusion feature map. This fusion strategy can adaptively integrate multi-scale features based on content differences, effectively improving the detection accuracy of cracks of various widths and directions. Preferably, the multi-scale fusion feature map is obtained through the following steps:

[0143] S451: Perform channel compression processing on the spatial context feature map, extract channel mean features and channel extreme value features through global average pooling and global maximum pooling respectively, and generate a channel statistical feature map.

[0144] Specifically, the system uses global average pooling (GAP) and global maximum pooling (GMP) to extract the mean and extreme features of the channel, respectively. GAP can capture the overall trend of the channel, while GMP can highlight the significant feature points in the channel. The system combines the results of these two pooling operations to generate a channel statistical feature map, thereby obtaining more comprehensive statistical information in the channel dimension. This channel compression processing not only reduces the amount of data, but also retains key feature information, providing a basis for subsequent nonlinear mapping processing.

[0145] S452: Perform nonlinear mapping processing on the channel statistical feature map, realize channel dimension compression and feature interaction through multi-layer perceptron, and generate channel attention weight.

[0146] Specifically, through the operation of a multi-layer perceptron (MLP), the system performs nonlinear transformations on the data, achieving compression of the channel dimension and feature interaction. The MLP consists of multiple fully connected layers and can learn complex feature relationships. The system inputs the results of GAP and GMP into the MLP separately and then combines the outputs of the two. The combined expression is:

[0147] w c =σ*(MLP(GAP(F))+MLP(GMP(F)))

[0148] Among them, w cis the channel attention weight, σ is the Sigmoid activation function, MLP is the multi-layer perceptron operation, and represents a nonlinear transformation of the data. GAP(F) represents the global average pooling result of the input spatial context feature map F in the spatial dimension, and GMP(F) represents the global maximum pooling result of the input spatial context feature map F in the spatial dimension. This nonlinear mapping process enables the system to generate channel attention weights, thereby dynamically adjusting the importance of each channel and highlighting feature channels related to crack detection.

[0149] S453: Perform spatial domain conversion on the spatial context feature map, restore the channel dimension through convolution operation, and generate spatial attention weights.

[0150] Specifically, the system first performs local spatial average pooling (AvgPool) and local spatial maximum pooling (MaxPool) on the spatial context feature map F to obtain the mean feature and extreme feature of the local area respectively. Then, the two pooling results are concatenated (Concat) in the channel dimension, and finally a convolution operation (Conv) is performed to generate the spatial attention weight. The expression of the spatial attention weight is:

[0151] w s =Conv(Concat(AvgPool(F),MaxPool(F)))

[0152] Among them, w s is the spatial attention weight, Conv is the convolution operation, Concat is the channel dimension splicing operation, AvgPool(F) is the local spatial average pooling result of F, and MaxPool(F) is the local spatial maximum pooling result of F.

[0153] S454: Dynamically weight the channel attention weights and spatial attention weights, multiply the elements together, and weight them element-by-element with the spatial context feature map to generate a multi-scale fusion feature map.

[0154] Specifically, the system dynamically weights the channel attention weights and spatial attention weights. First, the two weights are element-wise multiplied to comprehensively consider the importance of both the channel and spatial dimensions. The resulting combined weights are then element-wise weighted with the spatial context feature map to generate a multi-scale fused feature map. This process enables the system to dynamically adjust the contribution of different features and adaptively fuse multi-scale features based on image content. In this way, the system can more accurately capture crack features and improve the accuracy and robustness of crack detection. The multi-scale fused feature map not only contains rich feature information but also highlights crack-related regions and channels, providing high-quality feature input for subsequent crack probability prediction.

[0155] S46: Predict the crack probability of the multi-scale fusion feature map, generate normalized weights through 1×1 convolution and Softmax function, dynamically adjust the contribution of features at different scales according to the feature responses at different positions, and generate a second crack probability map.

[0156] Specifically, the system first introduces a scale-adaptive weighting strategy. To address the characteristics of deep features, which contain strong semantic information but lack spatial details, and shallow features, which retain details but have weak semantics, a feature importance self-learning module is designed. A normalized weight map is generated through 1×1 convolution and Softmax operations, and the contribution of features at different scales is dynamically adjusted based on the feature response at different locations. The formula is as follows:

[0157]

[0158] Among them, α and β represent the weights of features of different scales, and the Softmax function is used to reduce the number of channels of the feature map to the same as the original feature. Figure 1 The size of the same, Conv represents the convolution operation, which is used to learn the relationship between features. deep represents the deep features, F shallow In this way, the system can dynamically adjust the weight according to the importance of the feature to ensure the reasonable fusion of deep features and shallow features. The fusion result is:

[0159] F fused =α*F deep +β*F shallow

[0160] Among them, F fused This refined feature fusion strategy enables the model to adaptively integrate multi-scale features based on content differences, effectively improving the detection accuracy of cracks of various widths and directions.

[0161] To enhance position sensitivity, the system also introduces residual connections enhanced with position encoding. By incorporating positional prior knowledge, the system improves the ability to distinguish crack features in areas with similar background textures. By adding positional information to the feature map, the system enables the model to better understand the spatial distribution of crack features in the image, thereby improving crack detection accuracy.

[0162] The above-mentioned multi-scale feature fusion facade crack detection method can achieve accurate detection and interference suppression of fine cracks in complex backgrounds through a multi-level processing flow of high-complexity images, so as to solve the technical difficulties of high false detection rate and high missed detection rate of traditional methods in scenes with dense texture, shadow coverage or noise interference. Through color space conversion and multi-level contrast adaptive enhancement, the proposed method effectively separates brightness information and enhances the local contrast of crack regions. Combined with Gaussian difference filtering and spatial anisotropic sharpening, it strengthens the gradient response of crack edges and suppresses background texture interference, providing a high signal-to-noise ratio brightness feature map for subsequent feature extraction. Asymmetric convolution decomposition technology is used to fine-tune the modeling of multi-directional edge features, enabling directional enhancement of horizontal, vertical, and omnidirectional crack features. Multi-angle dilated convolution is used to model spatial contextual relationships, effectively capturing the geometric continuity of cracks and alleviating localized fractures. A channel-spatial dual attention mechanism and a scale-adaptive weighting strategy are used to dynamically fuse the advantages of multi-scale features, enhancing the collaborative expression of deep semantic features and shallow detail features, thereby addressing the lack of adaptability of single-scale features to complex morphological cracks. Furthermore, a dual-branch decision-making mechanism dynamically weights the fusion of traditional morphological processing results and deep learning prediction probabilities, combining the advantages of prior knowledge and data-driven models. The weight distribution is dynamically adjusted according to regional feature responses, thereby retaining weak edge cracks while suppressing background artifacts, achieving a balance between detection sensitivity and specificity. Through the collaborative optimization of multimodal signal enhancement, direction-sensitive feature extraction, context-aware fusion, and adaptive decision-making, this process can significantly improve the geometric integrity and spatial consistency of crack detection in highly complex scenarios, providing highly reliable technical support for building structural health monitoring.

[0163] In one embodiment, S5 of the facade crack detection method provided by the present invention incorporating multi-scale features specifically includes the following steps:

[0164] S51: performing normalization processing on the first crack probability map and the second crack probability map based on a probability normalization technology, mapping the probability values ​​to a preset interval through a nonlinear function, and generating a standard probability map.

[0165] Specifically, the system uses nonlinear functions, such as Sigmoid or Tanh, to transform probability values. These functions can effectively compress the range of probability values ​​to fit within a preset interval, such as [0, 1] or [0, 255]. This operation not only ensures the comparability and consistency of probability values ​​across different probability maps, but also enhances the contrast of the probability maps, making the boundaries between crack and non-crack areas clearer. The standard probability map generated in this way provides a unified benchmark for subsequent connected domain extraction. Probability normalization technology is of great significance in the field of image processing. It can eliminate the differences in numerical ranges caused by different algorithms or processing steps, ensuring the accuracy of subsequent analysis.

[0166] S52: Performing connected domain extraction processing on the standard probability map, and generating a set of candidate connected domains through region growing and boundary tracking.

[0167] Specifically, the region growing algorithm starts from a seed point and gradually merges adjacent pixels with similar attributes to form connected domains. This algorithm can effectively identify potential crack areas in the image. The boundary tracing algorithm searches along the boundaries of the connected domain, recording the location and connectivity of boundary pixels. The combination of these two algorithms can not only accurately identify candidate crack regions in the image but also represent them as a set of connected domains. During the connected domain extraction process, the system records the geometric features of each connected domain in detail, such as area, perimeter, and center coordinates. This information provides critical data support for subsequent noise suppression and post-processing verification. Through connected domain extraction, the system can preliminarily determine the location and shape of the crack, laying the foundation for subsequent detailed analysis.

[0168] S53: Perform noise suppression processing on the candidate connected domain set, and generate a denoised connected domain set by using corrosion and expansion operations of structural elements in different directions.

[0169] Specifically, the erosion operation removes noisy pixels by eliminating small holes and fractures in the connected domain, while the dilation operation restores the continuity of the cracks by expanding the scope of the connected domain. Structural elements in different directions, such as horizontal, vertical, and diagonal directions, can effectively handle noise and fractures in different directions, ensuring the geometric integrity of the cracks. After the erosion and dilation operations, the system generates a denoised connected domain set, reducing the possibility of false detection and missed detection. Noise suppression processing is crucial in image processing. It can effectively improve image quality, remove interference information, and make subsequent crack detection more accurate and reliable. Through this processing step, the system can ensure that the connected domain set is purer and the crack features are more prominent, thereby improving the robustness and accuracy of the entire detection process.

[0170] S54: Perform post-processing verification on the denoised connected domain set, and generate crack detection results through area threshold and aspect ratio screening.

[0171] Specifically, the system sets an area threshold to remove connected domains with too small an area, which are usually considered to be noise or non-crack features. Aspect ratio screening is used to remove connected domains whose shapes do not conform to crack characteristics, such as areas that are too circular or square. Through these screening conditions, the system can accurately identify true crack areas and separate them from the background and other non-crack features. The final crack detection results can be represented in a variety of ways, such as binary images, annotated images, or text reports, which intuitively indicate key information such as the location, length, and width of the cracks. Post-processing verification is an important link in the entire crack detection process. It ensures the accuracy and reliability of the detection results and provides an important reference for building structure health monitoring and maintenance decisions.

[0172] Preferably, if Figure 6 As shown, the present invention provides a facade crack detection device 600 with multi-scale feature fusion, which is configured with the following modules:

[0173] The data acquisition and scoring module 610 is used to obtain the image set to be detected input by the user, and calculate the information entropy, edge density, fusion texture complexity and gradient variance of the image set to be detected and perform multi-dimensional complexity evaluation to generate an image data set containing a complexity score;

[0174] A complexity classification module 620 is configured to perform image classification on the image data set based on a preset dynamic threshold value to generate high-complexity image data and low-complexity image data;

[0175] MEEMOModule 630 is used to perform parameter enhancement and multi-threshold fusion processing on low-complexity image data based on the MEEMO optimization algorithm, highlight small area features and capture edge changes, optimize the crack area separation process, and generate a first crack probability map. The first crack probability map is used to indicate the crack distribution in the low-complexity scene;

[0176] HMFFFramework module 640 is used to perform multimodal signal enhancement and asymmetric convolution processing on high-complexity image data based on the HMFF feature fusion framework, and generate a second crack probability map through multi-scale feature extraction, attention weight fusion and dual-branch decision making. The second crack probability map is used to indicate the crack distribution in high-complexity scenes;

[0177] The normalization and analysis module 650 is used to process the first crack probability map and the second crack probability map, and generate crack detection results through Sigmoid normalization and connected domain analysis. The crack detection results are used to indicate the final crack position and geometric shape.

[0178] In summary, the multi-scale feature-fused facade crack detection device provided by the present invention, through a multi-dimensional complexity assessment and dynamic classification mechanism, can achieve adaptive processing in complex scenarios, significantly improving the robustness and accuracy of facade crack detection. The multi-dimensional image complexity assessment framework based on the complexity classification module, through the weighted fusion of information entropy, edge density, texture complexity, and gradient variance and dynamic threshold classification, can achieve accurate complexity classification of image data, thereby resolving the classification bias problem caused by single-metric evaluation in traditional methods and providing a reliable basis for subsequent differentiated processing. For low-complexity images, the MEEMOModule module can effectively enhance the visibility of shallow cracks and suppress background noise through a parameterized enhancement process and a multi-threshold fusion strategy. Combined with multi-directional morphological optimization and multi-feature fusion, it can achieve geometrically complete reconstruction of small cracks, overcoming the defects of traditional algorithms that are prone to missed detection or edge breakage under uniform backgrounds. For high-complexity images, the HMFFFramework framework can achieve refined extraction of multi-scale directional features through multimodal signal enhancement and asymmetric convolution decomposition. Combined with the channel-space dual attention mechanism and dual-branch decision strategy, it can effectively distinguish between real cracks and complex background artifacts to solve the problem of false detection caused by dense texture or shadow interference. Finally, through the probabilistic fusion and geometric verification of the normalization and analysis modules, the accurate output of crack location and morphology can be achieved. Through the technical architecture of complexity-driven, low / high scene division, and deep and shallow feature collaboration, this invention can significantly improve the detection sensitivity and adaptability in complex engineering environments, providing highly reliable technical support for building structure health monitoring.

[0179] Preferably, the data acquisition and scoring module 610 provided in this application is configured with the following units:

[0180] Information entropy score generation unit: Based on the information entropy evaluation technology, the information amount of the image set to be detected is analyzed and processed, a smoothing factor is introduced to filter the zero probability items, and the information entropy score is generated.

[0181] Edge density score generation unit: Based on edge density detection technology, the edge distribution of the image set to be detected is quantified, the edge area is extracted through the gradient operator and the pixel ratio is calculated to generate an edge density score.

[0182] Texture complexity score generation unit: Based on texture fusion analysis technology, local and global texture features are extracted from the image set to be detected, and the texture complexity score is generated by combining LBP variance and GLCM contrast.

[0183] Gradient variance score generation unit: performs local structural change analysis on the image set to be detected, and generates a gradient variance score by calculating the gradient amplitude variance.

[0184] Image dataset generation unit: Based on the weighted fusion algorithm, the information entropy score, edge density score, texture complexity score and gradient variance score are comprehensively evaluated and processed, and an image dataset containing a complexity score is generated in combination with the image set to be detected.

[0185] Preferably, the MEEMOModule 630 provided in this application is configured with the following units:

[0186] A parameterized enhanced image generation unit is used to perform parameterized enhancement processing on low-complexity image data and generate a parameterized enhanced image through contrast gain and brightness offset;

[0187] A candidate mask set generation unit is used to extract candidate crack regions from the parameterized enhanced image, perform region segmentation by combining the global threshold, fixed ratio threshold, and local adaptive threshold to generate a candidate mask set;

[0188] The denoising mask generation unit is used to perform noise suppression on the candidate mask set based on the multi-directional morphological optimization technology, and generate the denoising mask by performing opening operations using 9×1, 1×9, and 5×5 structure elements;

[0189] The first crack probability map generating unit is used to perform multi-feature fusion processing on the denoising mask, and jointly perform dark area enhancement, gradient intensity calculation and skeleton structure extraction to generate the first crack probability map.

[0190] Preferably, the candidate mask set generation unit includes a global threshold mask generation subunit, a fixed threshold generation subunit, an adaptive generation subunit, and a joint mask generation subunit. The global threshold mask generation subunit is used to perform full-image threshold segmentation processing on the parametric enhanced image based on the global threshold segmentation technology, maximize the inter-class variance of the image foreground and background areas, and generate a global threshold mask; the fixed threshold generation subunit is used to perform grayscale truncation processing on the parametric enhanced image, divide the high and low grayscale areas according to a preset ratio, and generate a fixed threshold mask; the adaptive generation subunit is used to perform dynamic threshold calculation processing on the parametric enhanced image, and generate an adaptive threshold mask based on the local neighborhood grayscale distribution; the joint mask generation subunit is used to perform joint processing on the global threshold mask, the fixed threshold mask, and the adaptive threshold mask based on the multi-mask fusion technology to generate a candidate mask set.

[0191] Preferably, the HMFFFramework module 640 provided in this application is configured with the following units:

[0192] Multi-scale brightness feature map generation unit: performs color space conversion processing on high-complexity image data, separates brightness information to extract LAB brightness channel and multi-level CLAHE processing, and generates multi-scale brightness feature map.

[0193] Direction-sensitive feature map generation unit: Based on the direction-sensitive feature extraction technology, the multi-scale brightness feature map is edge-enhanced, and a direction-sensitive feature map is generated by Gaussian difference filtering and the introduction of spatial anisotropic sharpening.

[0194] Multi-directional feature map generation unit: Based on the asymmetric convolution decomposition technology, multi-directional feature extraction processing is performed on the direction-sensitive feature map to capture horizontal, vertical and omnidirectional features and generate a multi-directional feature map.

[0195] Spatial context feature map generation unit: Based on context-aware technology, spatial relationship modeling is performed on multi-directional feature maps, and multi-angle dilated convolution is introduced to generate spatial context feature maps.

[0196] Multi-scale fusion feature map generation unit: processes the spatial context feature map based on the preset adaptive feature fusion mechanism to generate a multi-scale fusion feature map.

[0197] Second crack probability map generation unit: Crack probability prediction is performed on the multi-scale fusion feature map, normalized weights are generated through 1×1 convolution and Softmax function, the contribution of features at different scales is dynamically adjusted according to the feature responses at different positions, and the second crack probability map is generated.

[0198] Preferably, the multi-scale fusion feature map generation unit includes a statistical feature map generation subunit, a channel weight generation subunit, a spatial weight generation subunit and a fusion feature map generation subunit. Among them, the statistical feature map generation subunit is used to perform channel compression processing on the spatial context feature map, extract channel mean features and channel extreme features respectively through global average pooling and global maximum pooling, and generate a channel statistical feature map; the channel weight generation subunit is used to perform nonlinear mapping processing on the channel statistical feature map, realize channel dimension compression and feature interaction through a multi-layer perceptron, and generate channel attention weights; the spatial weight generation subunit is used to perform spatial domain conversion processing on the spatial context feature map, restore the channel dimension through a convolution operation, and generate spatial attention weights; the fusion feature map generation subunit is used to perform dynamic weighting processing on the channel attention weights and the spatial attention weights, and perform element-by-element weighting on the spatial context feature map after element-by-element multiplication to generate a multi-scale fusion feature map.

[0199] Preferably, the normalization and analysis module 650 provided in this application is configured with the following units:

[0200] Standard probability map generating unit: performs standardization processing on the first crack probability map and the second crack probability map based on probability normalization technology, maps the probability value to a preset interval through a nonlinear function, and generates a standard probability map.

[0201] Candidate connected domain set generation unit: performs connected domain extraction processing on the standard probability map, and generates a candidate connected domain set through region growing and boundary tracking.

[0202] Denoising connected domain set generation unit: performs noise suppression on the candidate connected domain set, and generates the denoising connected domain set by using corrosion and expansion operations of structural elements in different directions.

[0203] Crack detection result generation unit: Post-process and verify the denoised connected domain set, and generate crack detection results through area threshold and aspect ratio screening.

[0204] In one embodiment, the present application further provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned facade crack detection method using multi-scale feature fusion when executing the computer program.

[0205] In one embodiment, the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the above-mentioned facade crack detection method using multi-scale feature fusion is implemented.

[0206] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and integrate different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0207] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0208] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A facade crack detection method based on multi-scale feature fusion, characterized in that: The following steps are involved: S1: Obtain the image set to be detected input by the user, and calculate the information entropy, edge density, fusion texture complexity and gradient variance of the image set to be detected and perform multi-dimensional complexity evaluation to generate an image dataset containing complexity scores; S2: performing image classification on the image dataset based on a preset dynamic threshold to generate high-complexity image data and low-complexity image data; S3: performing parameter enhancement and multi-threshold fusion processing on the low-complexity image data based on the MEEMO optimization algorithm to highlight small area features and capture edge changes, optimize the crack area separation process, and generate a first crack probability map, which is used to indicate crack distribution in the low-complexity scene; S4: performing multimodal signal enhancement and asymmetric convolution processing on the high-complexity image data based on the HMFF feature fusion framework, generating a second crack probability map through multi-scale feature extraction, attention weight fusion and dual-branch decision making, wherein the second crack probability map is used to indicate the crack distribution in the high-complexity scene; S5: Processing the first crack probability map and the second crack probability map, and generating a crack detection result through Sigmoid normalization and connected domain analysis. The crack detection result is used to indicate the final crack position and geometric shape.

2. The method according to claim 1, characterized in that Said S1 comprises: S11: performing information analysis on the image set to be detected based on information entropy evaluation technology, introducing a smoothing factor to filter zero-probability items, and generating an information entropy score; S12: performing edge distribution quantification processing on the image set to be detected based on edge density detection technology, extracting edge areas through a gradient operator and calculating pixel ratios to generate edge density scores; S13: performing local and global texture feature extraction processing on the image set to be detected based on texture fusion analysis technology, and generating a texture complexity score by combining LBP variance and GLCM contrast; S14: performing local structural change analysis on the image set to be detected, and generating a gradient variance score by calculating the gradient amplitude variance; S15: performing a comprehensive evaluation process on the information entropy score, the edge density score, the texture complexity score, and the gradient variance score based on a weighted fusion algorithm, and generating an image data set containing a complexity score in combination with the image set to be detected.

3. The method according to claim 1, characterized in that The S3 includes: S31: performing parametric enhancement processing on the low-complexity image data to generate a parametric enhanced image through contrast gain and brightness offset; S32: extracting candidate crack regions from the parameterized enhanced image, performing region segmentation using a global threshold, a fixed ratio threshold, and a local adaptive threshold to generate a candidate mask set; S33: performing noise suppression processing on the candidate mask set based on a multi-directional morphological optimization technique, performing opening operations using 9×1, 1×9, and 5×5 structure elements to generate a denoising mask; S34: performing multi-feature fusion processing on the denoising mask, combining dark area enhancement, gradient intensity calculation and skeleton structure extraction to generate a first crack probability map.

4. The method according to claim 3, characterized in that The S32 includes: S321: performing full-image threshold segmentation processing on the parameterized enhanced image based on a global threshold segmentation technique, maximizing the inter-class variance of the image foreground and background regions, and generating a global threshold mask; S322: performing grayscale truncation processing on the parametric enhanced image, dividing high and low grayscale areas according to a preset ratio, and generating a fixed threshold mask; S323: Perform dynamic threshold calculation processing on the parameterized enhanced image, and generate an adaptive threshold mask according to the local neighborhood grayscale distribution; S324: Jointly process the global threshold mask, the fixed threshold mask, and the adaptive threshold mask based on a multi-mask fusion technology to generate a candidate mask set.

5. The method according to claim 1, wherein The S4 includes: S41: performing color space conversion processing on the high-complexity image data, separating brightness information to extract LAB brightness channels and performing multi-level CLAHE processing to generate a multi-scale brightness feature map; S42: performing edge enhancement processing on the multi-scale brightness feature map based on a direction-sensitive feature extraction technology, performing Gaussian difference filtering and introducing spatial anisotropic sharpening to generate a direction-sensitive feature map; S43: performing multi-directional feature extraction processing on the direction-sensitive feature map based on an asymmetric convolution decomposition technology, capturing horizontal, vertical, and omnidirectional features, and generating a multi-directional feature map; S44: performing spatial relationship modeling processing on the multi-directional feature map based on context-aware technology, introducing multi-angle dilated convolution, and generating a spatial context feature map; S45: processing the spatial context feature map based on a preset adaptive feature fusion mechanism to generate a multi-scale fusion feature map; S46: Perform crack probability prediction on the multi-scale fusion feature map, generate normalized weights through 1×1 convolution and Softmax function, dynamically adjust the contribution of features at different scales according to feature responses at different positions, and generate a second crack probability map.

6. The method according to claim 5, characterized in that The S45 includes: S451: performing channel compression processing on the spatial context feature map, extracting channel mean features and channel extreme value features respectively through global average pooling and global maximum pooling, and generating a channel statistical feature map; S452: Perform nonlinear mapping processing on the channel statistical feature map, implement channel dimension compression and feature interaction through a multi-layer perceptron, and generate a channel attention weight. The expression of the channel attention weight is: w c Nσ*(MLP(GAP(F))+MLP(GMP(F))) Among them, w c is the channel attention weight, σ is the Sigmoid activation function, MLP is the multi-layer perceptron operation, which means nonlinear transformation of data, GAP(F) represents the global average pooling result of the input spatial context feature map F in the spatial dimension, and GMP(F) represents the global maximum pooling result of the input spatial context feature map F in the spatial dimension; S453: Perform spatial domain conversion processing on the spatial context feature map, restore the channel dimension through convolution operation, and generate spatial attention weights. The expression of the spatial attention weights is: w s =Conv(Concat(AvgPool(F),MaxPool(F))) Among them, w s is the spatial attention weight, Conv is the convolution operation, Concat is the channel dimension splicing operation, AvgPool(F) is the local spatial average pooling result of F, and MaxPool(F) is the local spatial maximum pooling result of F; S454: Dynamically weight the channel attention weight and the spatial attention weight, multiply the elements together and weight them element-by-element with the spatial context feature map to generate a multi-scale fusion feature map.

7. The method according to any one of claims 1 to 6, characterized in that The S5 includes: S51: normalizing the first crack probability map and the second crack probability map based on a probability normalization technology, mapping the probability values ​​to a preset interval using a nonlinear function, and generating a standard probability map; S52: Performing connected domain extraction processing on the standard probability map, and generating a set of candidate connected domains through region growing and boundary tracking; S53: performing noise suppression processing on the candidate connected component set, generating a denoised connected component set by using corrosion and expansion operations of structural elements in different directions; S54: performing post-processing verification on the denoised connected domain set, and generating crack detection results through area threshold and aspect ratio screening.

8. A facade crack detection device with multi-scale feature fusion, characterized in that: The device comprises: The data acquisition and scoring module is used to obtain the image set to be detected input by the user, and calculate the information entropy, edge density, fusion texture complexity and gradient variance of the image set to be detected and perform multi-dimensional complexity evaluation to generate an image data set containing a complexity score; a complexity classification module, configured to perform image classification on the image data set based on a preset dynamic threshold value to generate high-complexity image data and low-complexity image data; A MEEMOModule module is used to perform parameter enhancement and multi-threshold fusion processing on the low-complexity image data based on the MEEMO optimization algorithm, highlight the features of small areas and capture edge changes, optimize the crack area separation process, and generate a first crack probability map, which is used to indicate the crack distribution in the low-complexity scene; The HMFFFramework module is used to perform multimodal signal enhancement and asymmetric convolution processing on the high-complexity image data based on the HMFF feature fusion framework, and generate a second crack probability map through multi-scale feature extraction, attention weight fusion and dual-branch decision making. The second crack probability map is used to indicate the crack distribution in the high-complexity scene; The normalization and analysis module is used to process the first crack probability map and the second crack probability map, and generate a crack detection result through Sigmoid normalization and connected domain analysis. The crack detection result is used to indicate the final crack position and geometric shape.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Visual segmentation counting method and system suitable for disorderly stacked parts

    CN120931936A

  • Multi-level concrete crack detection method based on unmanned aerial vehicle

    CN121147227A

  • GLCM texture feature extraction method and system based on dynamic multi-scale weighting

    CN121147547A

  • GLCM texture feature extraction method and system based on dynamic multi-scale weighting

    CN121147547B

  • Visual identification method and system for grinding quality of R corner of edge of cover plate glass

    CN121258944A