Concrete surface multi-defect intelligent identification system based on multi-scale feature fusion

By employing a multi-scale feature fusion method, which utilizes wavelet frequency separation and attention mechanisms to fuse feature information, the problem of insufficient accuracy in identifying concrete surface defects is solved. This enables precise identification and combination of various defects, improving the completeness and efficiency of detection.

CN121660978APending Publication Date: 2026-03-13GUANGDONG GUANGMEI KEZHU CONSTR TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

The accuracy of concrete surface defect identification is insufficient, and the interference of complex textures makes it difficult to distinguish between real defects and natural textures, affecting the integrity and reliability of the detection.

Method used

A multi-scale feature fusion method is adopted, which combines multi-scale wavelet frequency separation and texture-driven random high-frequency denoising with an attention mechanism to fuse feature information of different sizes, and constructs a defect recognition model to identify various defects.

Benefits of technology

It improves the accuracy and reliability of concrete surface defect identification, enables precise differentiation and combination identification of different types of defects, reduces reliance on manual inspection, and lowers the risk of human error.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121660978A_ABST
    Figure CN121660978A_ABST
Patent Text Reader

Abstract

The invention discloses a concrete surface multi-defect intelligent identification system based on multi-scale feature fusion, and relates to the technical field of surface defect identification. The system comprises a concrete surface image acquisition module which is used for acquiring a concrete surface image and an angle, a distance and a space coordinate during imaging through imaging equipment; the texture and defect separation module is used for performing multi-scale wavelet frequency separation on the concrete surface image; the image multi-scale feature fusion module is used for extracting feature information of different sizes from the concrete surface image and fusing the feature information through an attention mechanism; and the multi-defect intelligent identification module is used for obtaining a judgment combination of defect identification by analyzing the comprehensive feature map, and identifying various different defects according to the judgment combination of the defect identification. According to the method, the concrete surface image is subjected to multi-scale wavelet frequency separation, so that the concrete surface defect recognition accuracy is improved, and the problem of insufficient concrete surface defect recognition accuracy in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of surface defect recognition technology, and in particular to an intelligent recognition system for multiple defects on concrete surfaces based on multi-scale feature fusion. Background Technology

[0002] Concrete components may develop various defects during use, such as cracks, spalling, honeycombing, and voids. These defects not only individually affect the structural safety but also accumulate, accelerating material degradation and leading to an overall decline in performance. Therefore, single-defect detection is prone to overlooking other potential problems and cannot comprehensively reflect the structural condition. Multi-defect identification can simultaneously capture the characteristic information of different types of defects, improving the completeness and accuracy of detection and providing a reliable basis for structural safety assessment. Furthermore, different defect types have varying degrees of impact and require different treatment methods; multi-defect identification can provide a reference for scientifically formulating maintenance and reinforcement plans, optimizing maintenance decisions. At the same time, automated multi-defect identification can reduce reliance on manual inspections, improve detection efficiency, reduce labor costs, and reduce the risk of human oversight or misjudgment.

[0003] Multi-scale feature fusion is a core technology widely used in image understanding tasks, especially in complex visual scenes such as object detection, image segmentation, and defect recognition. Specifically, multi-scale feature fusion refers to the effective integration of feature information extracted from different scales (or different semantic levels) in an image, so as to take into account both local details and global semantics, enabling the model to have both small target recognition capabilities and large target semantic understanding capabilities.

[0004] The rough surface, uneven particle distribution, and varying mortar-aggregate ratios of concrete result in complex and highly irregular natural textures. These complex textures manifest as high-frequency details and edge variations in images, easily confused with the edge and texture features of defects such as cracks, voids, and spalling. This makes it difficult for the system to distinguish between genuine defects and the material's natural texture during feature extraction. This interference can lead to missed detection of minute cracks or shallow spalling, or misclassification of natural particles or surface irregularities as defects, significantly reducing the accuracy and reliability of defect identification. Furthermore, complex textures weaken the distinguishability of different defect types in the feature space, making multi-defect classification more challenging.

[0005] In view of this, the present invention proposes an intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion to solve the above problems. Summary of the Invention

[0006] This application provides an intelligent recognition system for multiple defects on concrete surfaces based on multi-scale feature fusion, which solves the problem of insufficient accuracy in recognizing concrete surface defects in the prior art and achieves the effect of improving the accuracy of concrete surface defect recognition.

[0007] This application provides an intelligent recognition system for multiple defects on concrete surfaces based on multi-scale feature fusion, comprising: a concrete surface image acquisition module: acquiring concrete surface images through an imaging device, simultaneously collecting the angle, distance, and spatial coordinates during imaging, and preprocessing the concrete surface images, which include visible light images and infrared images of the concrete surface; a texture and defect separation module: performing multi-scale wavelet frequency separation on the concrete surface images, separating the frequencies of the concrete surface images into texture-dominated random high frequencies and defect-dominated structured high frequencies, applying soft thresholding to the texture-dominated random high frequencies for noise reduction, and performing frequency enhancement on the defect-dominated structured high frequencies to obtain a concrete surface image with natural texture removed; an image multi-scale feature fusion module: extracting feature information of different sizes from the concrete surface image with natural texture removed, fusing the feature information of different sizes through an attention mechanism to obtain a comprehensive feature map of the concrete surface image with natural texture removed, where the feature information of different sizes includes low-level features (i.e., small-scale features) and high-level features (i.e., large-scale features); and a multi-defect intelligent recognition module: analyzing the edge information and texture features in the comprehensive feature map of the concrete surface image with natural texture removed to obtain a judgment combination for defect recognition, and recognizing multiple different defects based on the judgment combination for defect recognition.

[0008] Furthermore, multi-scale wavelet frequency separation is performed on the concrete surface image, including: performing discrete wavelet transform on the concrete surface image to obtain a first-level decomposition result, which includes a low-frequency sub-band and a high-frequency sub-band; performing multiple decompositions on the low-frequency sub-band to obtain a multi-scale wavelet frequency decomposition result; and using the multi-scale wavelet frequency decomposition result as the processed low-frequency sub-band.

[0009] Furthermore, soft thresholding is applied to the texture-dominated random high frequencies for denoising, and frequency enhancement is performed on the defect-dominated structured high frequencies. This includes: analyzing the high-frequency subbands to obtain texture-dominated random high frequencies and defect-dominated structured high frequencies; using a soft thresholding function to attenuate the texture-dominated random high frequencies until they are removed from the high-frequency subbands; performing gradient direction statistics on the defect-dominated structured high frequencies; selecting the high-frequency components in the main direction for filter weighting, while keeping other directions unchanged, to obtain the processed high-frequency subbands; and fusing the processed low-frequency subbands and the processed high-frequency subbands to obtain a concrete surface image with natural texture removed.

[0010] Furthermore, feature information of different sizes is extracted from the concrete surface image after removing natural textures. This includes: layering the concrete surface image after removing natural textures into low-level feature images, mid-level feature images, and high-level feature images according to their size; extracting features from each layer to obtain low-level feature maps, mid-level feature maps, and high-level feature maps; and aligning the feature maps of different levels through upsampling and downsampling. The specific process for aligning the feature maps of different levels through upsampling and downsampling is as follows: obtaining a preset size for the feature maps; upsampling and aligning all feature maps smaller than the preset size; downsampling and aligning all feature maps larger than the preset size; and outputting all the aligned feature maps as a feature map set. The low-level feature images are pixel-level microscopic images, the mid-level feature images are local pixel-level small region images, and the high-level feature images are global-level entire images.

[0011] Furthermore, feature extraction is performed on each layer of the image, including: extracting microscopic feature information such as edges, texture and brightness from low-level feature images; extracting local structural features such as local shape, texture distribution and local contrast from mid-level feature images; and extracting global features such as overall morphology, spatial distribution of defects and large-scale texture trends from high-level feature images.

[0012] Furthermore, feature information of different sizes is fused through an attention mechanism, including: inputting a set of feature maps after feature alignment, calculating attention weight coefficients for each feature map, fusing the feature maps with the attention weight coefficients to obtain a fused comprehensive feature map of the concrete surface image with natural texture removed; calculating attention weight coefficients for each feature map, specifically: compressing the feature map information, performing global average pooling on the compressed feature map to obtain a compressed feature vector, constructing an attention mapping function, inputting the compressed feature vector into the attention mapping function to obtain the attention weight coefficients for each feature map.

[0013] Furthermore, the compressed feature map is subjected to global average pooling, including: extracting the height, width and channels from the compressed feature map; inputting the height, width and number of channels of the compressed feature map into the global average pooling function; taking the global average value of all pixels in each channel; and combining the global average values ​​of all channels in order to form a feature vector as the feature compression vector.

[0014] Furthermore, the feature maps are fused with attention weight coefficients, including multiplying the feature map of each channel with the corresponding channel weight coefficient channel by channel to obtain a weighted feature map, and stitching all the weighted feature maps together along the channel dimension to obtain a fused comprehensive feature map of the concrete surface image with natural texture removed.

[0015] Furthermore, the edge information and texture features in the comprehensive feature map of the concrete surface image after removing natural textures are analyzed, including: performing edge detection operators and texture analysis on the comprehensive feature map, extracting edge information and texture from the comprehensive feature map of the concrete surface image after removing natural textures, stitching the edge information and texture features in the channel dimension, and constructing a judgment combination for defect recognition based on the stitching result.

[0016] Furthermore, based on the judgment combination of defect identification, various different defects can be identified, including: constructing a defect identification model based on the judgment combination of defect identification, training the defect identification model, inputting the fused feature map into the defect identification model for region-by-region analysis, and outputting the spatial location and type label of the defect.

[0017] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0018] 1. By performing multi-scale wavelet frequency separation on concrete surface images, applying soft thresholding to the texture-dominated random high frequencies for noise reduction, and enhancing the defect-dominated structured high frequencies, a concrete surface image with natural textures removed is obtained. This improves the accuracy of concrete surface defect identification and effectively solves the problem of insufficient accuracy in concrete surface defect identification in existing technologies.

[0019] 2. By acquiring images of the concrete surface through imaging equipment and simultaneously collecting the angle, distance, and spatial coordinates during imaging, the concrete surface images are preprocessed, thereby achieving the effect of restoring the geometric proportions of the concrete surface in real space, effectively solving the problem of insufficient realism of concrete surface images in existing technologies.

[0020] 3. By analyzing the edge information and texture features in the comprehensive feature map of the concrete surface image after removing natural texture, a judgment combination for defect identification is obtained. Based on the judgment combination for defect identification, a variety of different defects can be identified, thereby achieving accurate differentiation and combination identification of different types of defects. Attached Figure Description

[0021] Figure 1 A schematic diagram of the structure of the intelligent recognition system for multiple defects on concrete surfaces based on multi-scale feature fusion provided in this application embodiment;

[0022] like Figure 2 The image shown is an image of four sub-bands of discrete wavelet transform in the intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion provided in this application embodiment.

[0023] like Figure 3The diagram shown is a flowchart illustrating the processing of the comprehensive feature map in the intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion provided in this application embodiment. Detailed Implementation

[0024] This application provides an intelligent recognition system for multiple defects on concrete surfaces based on multi-scale feature fusion, which solves the problem of insufficient accuracy in recognizing concrete surface defects in the prior art. By performing multi-scale wavelet frequency separation on the concrete surface image, applying soft threshold denoising to the texture-dominated random high frequencies, and enhancing the frequency of the defect-dominated structured high frequencies, a concrete surface image with natural texture removed is obtained, thereby improving the accuracy of concrete surface defect recognition.

[0025] like Figure 1 The diagram shown is a structural schematic of the intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion provided in this application embodiment. The intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion provided in this application embodiment is implemented in detail through the following steps:

[0026] First, the dual-spectral imaging device is fixed on the drone to acquire images of the concrete surface to be identified. While acquiring images, the shooting angle, the distance between the camera and the concrete surface, and the three-dimensional spatial coordinates of the shooting point are collected. The image is then subjected to perspective transformation based on the shooting angle to eliminate imaging distortion caused by the imaging angle. The actual physical size corresponding to the unit area of ​​pixels in the image is calculated based on the distance between the camera and the concrete surface. Finally, the image is spatially aligned based on the three-dimensional spatial coordinates of the shooting point.

[0027] The image is transformed by perspective based on the shooting angle, specifically: the shooting angle includes pitch angle, yaw angle and roll angle. A perspective transformation matrix is ​​established from the current imaging plane to the ideal frontal plane. The perspective transformation matrix is ​​used to perform pixel-level affine remapping on the image. In the transformed image, the photographed concrete surface will show a geometrically corrected shape, such as curved lines becoming straight, deformed patterns returning to normal proportions, and crack length and direction becoming more realistic and quantifiable.

[0028] Pitch, yaw, and roll angles are collectively known as Euler angles, representing the rotation of an imaging device in three-dimensional space relative to its ideal attitude. Pitch angle is the yaw angle on the X-axis relative to the ideal attitude, roll angle is the yaw angle on the Y-axis relative to the ideal attitude, and yaw angle is the yaw angle on the Z-axis relative to the ideal attitude. The perspective transformation matrix is ​​a 3×3 matrix representing the projection relationship of an image from one plane to another (caused by changes in viewpoint). It maps a point (x, y) in one image to a point (x', y') in another image and is typically used for image correction. Affine remapping is a process that recalculates the geometric position of each pixel in an image based on the transformation matrix, ultimately generating a new image.

[0029] The specific process of establishing the perspective transformation matrix is ​​as follows: Based on the pitch angle, yaw angle and roll angle, construct rotation matrices around the X-axis, Y-axis and Z-axis respectively, combine the three rotation matrices into an overall rotation matrix, and combine the overall rotation matrix with the intrinsic parameter matrix of the imaging device to construct the perspective transformation matrix.

[0030] The actual physical size per unit area of ​​a pixel in the image is calculated based on the distance between the camera and the concrete surface. Specifically, the actual physical length of a single pixel on the image sensor is calculated using the following formula:

[0031]

[0032] In the formula, C x Represented as the horizontal size in pixels, C y T is represented as the vertical size in pixels. x I represents the actual physical length of the sensor's photosensitive area in the horizontal direction. x T is the horizontal length of the image itself. y I represents the actual physical length of the sensor's photosensitive area in the vertical direction. y This represents the vertical length of the image itself.

[0033] The perspective projection relationship is established based on the focal length of the imaging device and the object distance during imaging, using the following formula:

[0034]

[0035] In the formula, F x F represents the actual horizontal size of an image pixel within the sensor. y C represents the actual vertical size of an image pixel within the sensor. x Represented as the horizontal size in pixels, C y The vertical size of the pixel is represented by D, which is the object distance, that is, the straight-line distance between the principal point (optical center) of the camera lens and the object being photographed (such as a concrete surface), and F is the focal length, which is the distance from the principal point of the lens to the photosensitive surface of the image sensor.

[0036] Calculate the actual physical area corresponding to a pixel region. For example, if a region has A pixels, the actual physical area = A * (C) x *C y By obtaining the distance between the camera and the concrete surface, and combining the camera's focal length, sensor size, and image resolution, a mapping relationship between pixel size and actual physical size is established. This enables accurate conversion of the physical length and area of ​​any region in the image in the real world, effectively solving the problems of inability to quantify defects and inaccurate dimensions in traditional visual recognition, and providing a reliable basis for subsequent engineering evaluation and maintenance decisions.

[0037] Discrete wavelet transform (DWT) is performed on the concrete surface image using a wavelet basis, yielding four sub-bands: low-frequency (LL), horizontal high-frequency (LH), vertical high-frequency (HL), and diagonal high-frequency (HH). Each sub-band represents different image attributes: LL preserves overall structure and brightness distribution; LH represents edges and horizontal texture; HL represents vertical texture and cracks; and HH represents diagonal texture and detail interference. The LL sub-band obtained in the first-level decomposition is used as input for further DWT, repeated 2-3 times to obtain multi-level LL and corresponding high-frequency sub-bands. These multi-level LL represent the multi-scale wavelet frequency decomposition result. Wavelet decomposition separates the image into different frequency regions. Structural defects (such as crack edges) are concentrated in the mid-to-high frequencies, while natural textures are more widely distributed. Layered processing can clearly separate these two types of defects, thus enabling better identification of image defects.

[0038] Discrete wavelet transform is performed on a concrete surface image using wavelet basis functions. A specific example is as follows: The Haar wavelet is selected as the wavelet basis function. A low-pass filter is used to perform directional convolution on the concrete surface image to obtain a low-pass column filter. A high-pass filter is then used to perform directional convolution on the concrete surface image to obtain a high-pass column filter. The low-pass and high-pass column filters are then convolved in column directions to obtain the low-frequency subband LL and the high-frequency subbands LH, HL, and HH, respectively. Figure 2 The image shown is an image of four sub-bands of the discrete wavelet transform in the intelligent recognition system for multiple defects on concrete surfaces based on multi-scale feature fusion provided in this application embodiment. The four sub-band images are obtained by discrete wavelet transform from a concrete surface image with an image size of 256*256, where LL: 128×128 (representing overall texture), LH: 128×128 (representing horizontal crack details), HL: 128×128 (representing vertical crack details), and HH: 128×128 (representing diagonal details).

[0039] Energy distribution analysis is used to classify high-frequency subbands into texture-dominated random high frequencies and defect-dominated structured high frequencies. A soft thresholding function is applied to the pixel coefficients of texture-dominated random high frequencies to significantly attenuate random texture noise while preserving edge features. Gradient direction statistical analysis is performed on structured high frequencies to identify local dominant directions (such as crack directions). Only high-frequency components along the dominant directions are weighted and amplified, while non-dominant directions remain unchanged. The low-frequency subbands after multi-level processing are then subjected to inverse wavelet transform with the processed high-frequency subbands to obtain an image that removes natural texture interference and enhances defect structure edges. Traditional filtering methods process high-frequency components uniformly, which can easily lead to the loss of defect information. This method separates and processes random and structured high frequencies in the frequency domain, selectively denoising and enhancing them to ensure that key identification features are preserved, significantly improving recognition accuracy.

[0040] A soft thresholding function is applied to pixel coefficients in texture-dominated random high-frequency areas. Specifically, for example, with an image of a concrete surface, a wavelet transform is performed, and the high-frequency subbands are summarized into a high-frequency coefficient matrix. A soft thresholding function is retrieved from a database, and applied to each high-frequency coefficient in the image to obtain the processed high-frequency coefficients, which are then summarized into a processed high-frequency coefficient matrix. Soft thresholding, by "weakening rather than completely truncating" high-frequency coefficients, can suppress small-amplitude random noise while preserving large-amplitude texture information. It also smooths the amplitude, reduces artifacts, and makes image edges and texture transitions natural, solving the problem that traditional mean or hard thresholding denoising easily erases real texture or retains too much noise.

[0041] Statistical analysis of gradient directions in structured high frequencies is performed. Specifically, the gradient direction and magnitude of each pixel in the image are calculated. The gradient direction refers to the direction of gray-level change, and the magnitude refers to the intensity of the gray-level change. The gradient direction of local regions is quantized to a certain angle range, for example, 0°–180° divided into 9 intervals, each 20°. The frequency of occurrence in each interval is counted, and the direction with the most occurrences is identified as the dominant direction. By performing a sliding window statistical analysis on the entire image, a structured high-frequency direction map of the entire image can be obtained, which can be used for texture enhancement, denoising, or feature extraction. By statistically analyzing high-frequency gradient directions and identifying local dominant directions, structural information can be preferentially preserved in the dominant texture direction while suppressing random noise in non-dominant directions, ensuring clear and natural textures. This solves the problem that traditional smoothing or denoising methods easily blur local textures and edges, leading to the loss of image details.

[0042] Based on size, the concrete surface image after removing natural textures is layered into low-level feature images, mid-level feature images, and high-level feature images. The features extracted at different levels have differences in semantic depth and spatial granularity. A unified feature map size is preset, for example, 256*256 pixels. All low-level feature maps (e.g., 512×512) are downsampled to the target size, and all high-level feature maps (e.g., 128×128) are upsampled to the target size. The feature map after all features are aligned is output as a feature map set. Defects of different scales cannot be fully expressed in a single layer. After alignment and fusion, a unified perception from details to the whole can be achieved, improving the defect discrimination capability.

[0043] Local texture filters are used to extract microscopic features, including edges, texture, and brightness, from low-level feature images. Medium-receptive-field convolutional layers are used to extract local structural features, including local shape, texture distribution, and local contrast, from mid-level feature images. Deep convolutional layers are used to extract global features, including overall morphology, spatial distribution of defects, and large-scale texture trends, from high-level feature images. These features are then combined to form a semantically clear and scale-complementary feature map set. This multi-layer feature extraction strategy effectively solves the recognition difficulties caused by large scale differences and complex structures in the identification of various concrete defects, improving the comprehensiveness of defect perception.

[0044] Edges refer to areas with significant changes in brightness, such as object outlines, crack edges, and particle interfaces, for example, extracting crack edges from a concrete surface or the outline of scratches on a metal surface. Texture refers to the repetitive patterns or directional changes of pixels in a local area, such as the unevenness of a wall surface. Brightness refers to the distribution of local pixel intensity, reflecting local lighting or material characteristics; smooth areas have more uniform brightness than rough areas. Local shape refers to small geometric forms formed by the combination of low-level edges and textures, such as the curved shape of a local crack or the outline of a particle cluster. Texture distribution refers to the arrangement or density of textures in a local area, such as dense cracks in one area and sparse cracks in another. Local contrast refers to the amplitude of local brightness variations, used to distinguish the strength or unevenness of textures; for example, pitted areas have higher contrast than flat areas. Overall morphology refers to the global geometric structure or object shape of an image, such as the overall shape of a crack network or the overall shape of an object outline. Defect spatial distribution refers to the distribution pattern of defects throughout the image or region, such as the path of a crack along a surface or the distribution density of corrosion spots. Large-scale texture trends refer to the directional, density variations, or structural tendencies of textures over a large area, such as the direction of the overall texture of a material surface or the roughness gradient.

[0045] For each feature map in the feature map set, channel-oriented compression is performed to preserve channel structure and reduce spatial dimension interference. Height, width, and number of channels are extracted from the compressed feature map; for example, height G = 256, width W = 256, and number of channels C = 64. The height, width, and number of channels of the compressed feature map are then input into a global average pooling function, as detailed below:

[0046]

[0047] In the formula, V c F represents the global average value of the c-th channel. c (i, j) represents the feature value of the c-th channel at position (i, j). The global average values ​​of all channels are sorted in order, V = [V1, V2, V3, ..., V...]. c For the low-level feature map, the feature compression vector V is obtained. low For the middle layer feature map, the feature compression vector V is obtained. mid For the high-level feature map, the feature compression vector V is obtained. high By compressing two-dimensional spatial features into one-dimensional channel vectors, computational costs are reduced, making it more suitable for attention weight learning. The extracted compressed feature vectors can be directly used in the attention mapping function. The system automatically learns the importance of low-level, mid-level, and high-level features in different defect scenarios. Furthermore, through global average pooling, local noise is smoothed, enhancing the expressive power of defect signals at the channel level.

[0048] A lightweight fully connected network is used to construct an attention mapping function. The compressed feature vector is input into the attention mapping model to obtain the attention weight coefficients corresponding to the feature map. Specifically, the attention mapping function is constructed by using a ReLU function for non-linear transformation between two fully connected layers, and then using a Sigmoid function to map the result to the range of 0-1, ensuring the correct range of weight values. The attention mechanism dynamically evaluates the importance of each feature map layer, automatically assigns importance to scale features, weakens the influence of low-quality feature maps, reduces redundant feature interference, and improves representation quality. For unimportant or highly interfering feature layers, the attention weights are low to reduce misidentification and avoid information dilution and noise amplification problems caused by feature concatenation.

[0049] The ReLU and Sigmoid functions are activation functions that can be called directly in Python.

[0050] Each channel's feature map is multiplied by its corresponding channel weight coefficient channel by channel to obtain a weighted feature map. All weighted feature maps are then concatenated along the channel dimension to obtain the fused, composite feature map of the concrete surface image after removing natural textures. For example, a low-level feature map:

[0051] Flow ∈R 256*256*64 Weight: α low =0.75, mid-layer feature map:

[0052] F mid ∈R 256*256*64 Weight: α mid =0.85, High-level feature map:

[0053] F high ∈R 256*256*64 Weight: α high =0.65. Weighting is applied to each channel of each feature map. ,

[0054] Once the weighted feature maps are obtained, they are concatenated along the channel dimension.

[0055] Concat represents stacking feature maps along the channel dimension, outputting a fused feature map with a spatial size of H*W and the number of channels C1+C2+C3, thus obtaining the fused comprehensive feature map F of the concrete surface image after removing natural textures. fusion , will F fusion The input is processed through the multi-defect intelligent identification module for defect identification. Through this process, the system can adaptively determine the importance of features at different scales, strengthen useful information, weaken interfering information, and unify microscopic details, local structure, and global semantics into a fusion feature map, thereby improving the accuracy of multi-defect identification.

[0056] Concat represents stacking feature maps along the channel dimension. For example, given feature map A[H, W, C1] and feature map B[H, W, C2], where H represents image brightness, W represents image width, and C1 and C2 represent the number of image channels, Concat(A, B) = C[H, W, C1+C2], we get the fused feature map C.

[0057] like Figure 3The diagram illustrates the process of processing the comprehensive feature map in the intelligent recognition system for multiple defects on concrete surfaces based on multi-scale feature fusion provided in this application embodiment. Edge detection operators are applied to the comprehensive feature map, and gradients and edge strengths are calculated for each map. An edge map is output based on the gradients and edge strengths. The edge map captures the contour information of defects such as cracks and spalling, highlighting high-frequency structures. Texture analysis is performed on the comprehensive feature map, using convolutional filtering to extract the texture response of each channel. A texture map is output based on the texture response, reflecting the local material surface texture distribution and distinguishing between defect-dominated textures and residual natural textures. The edge map and texture map are concatenated in the channel dimension to form a fused feature map containing both shape and texture information. Edge information and texture features are concatenated in the channel dimension, and a judgment combination for defect recognition is constructed based on the concatenation result. This system does not rely on a single feature but combines edges and textures, improving the recognition rate of micro-cracks and shallow spalling. Low-level micro-details, mid-level local structures, and high-level global information simultaneously participate in feature construction, enhancing the defect discrimination capability.

[0058] A convolutional neural network was chosen to construct the defect recognition model. Spatial locations and type labels of defects corresponding to historical images were obtained from a database. The fused feature maps were divided into training and testing sets in a 7:3 ratio. The defect recognition model was trained using the training set, and the accuracy was calculated on the testing set after each training round. The optimal model was selected as the final defect recognition model. For test images, the fused feature maps were input into the defect recognition model by region. The model output the defect probability for each region, classifying each region to obtain the defect type and its spatial coordinates. By combining multi-scale fusion, natural texture removal, and edge-texture judgment, the recognition ability of micro-cracks, shallow peeling, and honeycomb voids was improved. The model outputs the spatial coordinates of the defects, achieving accurate localization of defects in images or actual tunnel locations.

[0059] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0060] This invention is described with reference to flowchart illustrations and / or block diagrams of systems, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0061] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0062] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0063] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0064] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A multi-defect intelligent identification system for concrete surfaces based on multi-scale feature fusion, characterized in that, include: Concrete surface image acquisition module: Acquires concrete surface images through imaging equipment, synchronously collects the angle, distance and spatial coordinates during imaging, and preprocesses the concrete surface images, which include visible light images and infrared images of the concrete surface. Texture and Defect Separation Module: Performs multi-scale wavelet frequency separation on the concrete surface image, separating the frequency of the concrete surface image into texture-dominated random high frequencies and defect-dominated structured high frequencies. Apply soft thresholding to the texture-dominated random high frequencies for noise reduction, and perform frequency enhancement on the defect-dominated structured high frequencies to obtain a concrete surface image with natural texture removed. Image multi-scale feature fusion module: Extracts feature information of different sizes from concrete surface images with natural texture removed, and fuses the feature information of different sizes through an attention mechanism to obtain a comprehensive feature map of the concrete surface image with natural texture removed. The feature information of different sizes includes low-level features, i.e. small-scale features, and high-level features, i.e. large-scale features. Multi-defect intelligent identification module: By analyzing the edge information and texture features in the comprehensive feature map of the concrete surface image after removing natural texture, the module obtains the judgment combination for defect identification and identifies a variety of different defects based on the judgment combination.

2. The intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion as described in claim 1, characterized in that, The multi-scale wavelet frequency separation of the concrete surface image includes: Discrete wavelet transform is performed on the concrete surface image to obtain a first-level decomposition result, which includes a low-frequency sub-band and a high-frequency sub-band. The low-frequency sub-band is decomposed multiple times to obtain a multi-scale wavelet frequency decomposition result, which is then used as the processed low-frequency sub-band.

3. The intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion as described in claim 1, characterized in that, The process of applying soft thresholding to texture-dominated random high frequencies and frequency enhancement to defect-dominated structured high frequencies includes: The high-frequency subband is analyzed to obtain texture-dominated random high frequencies and defect-dominated structured high frequencies. The texture-dominated random high frequencies are attenuated using a soft thresholding function until they are removed from the high-frequency subband. Gradient direction statistics are performed on the defect-dominated structured high frequencies, and the high-frequency components in the main direction are selected for filter weighting while other directions remain unchanged to obtain the processed high-frequency subband. The processed low-frequency subband and the processed high-frequency subband are fused to obtain an image of the concrete surface with natural texture removed.

4. The intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion as described in claim 1, characterized in that, The extraction of feature information of different sizes from concrete surface images with natural textures removed includes: The concrete surface image after removing natural texture is layered according to size, including low-level feature image, middle-level feature image and high-level feature image. Feature extraction is performed on each layer to obtain low-level feature map, middle-level feature map and high-level feature map. The feature maps of different levels are aligned by upsampling and downsampling. The feature maps at different levels are aligned by upsampling and downsampling. The specific process is as follows: obtain the preset size of the feature map, upsample and align all feature maps smaller than the preset size, downsample and align all feature maps larger than the preset size, and output the feature map set after all feature alignment. The low-level feature image is a pixel-level microscopic image, the mid-level feature image is a local pixel-level small region image, and the high-level feature image is a global-level entire image.

5. The intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion as described in claim 4, characterized in that, The feature extraction for each image layer includes: Microscopic features, including edges, texture, and brightness, are extracted from low-level feature images. Local structural features, including local shape, texture distribution, and local contrast, are extracted from mid-level feature images. Global features, including overall morphology, spatial distribution of defects, and large-scale texture trends, are extracted from high-level feature images.

6. The intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion as described in claim 1, characterized in that, The process of fusing feature information of different sizes through an attention mechanism includes: The input feature map set after feature alignment is used to calculate the attention weight coefficient for each feature map. The feature map and the attention weight coefficient are then fused to obtain the comprehensive feature map of the fused concrete surface image with natural texture removed. The specific process for calculating the attention weight coefficient for each feature map is as follows: compress the information of the feature map, perform global average pooling on the compressed feature map to obtain the feature compression vector, construct the attention mapping function, input the feature compression vector into the attention mapping function, and obtain the attention weight coefficient for each feature map.

7. The intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion as described in claim 6, characterized in that, The step of performing global average pooling on the compressed feature map includes: The height, width, and number of channels are extracted from the compressed feature map. The height, width, and number of channels of the compressed feature map are input into the global average pooling function. The global average value of all pixels in each channel is taken, and the global average values ​​of all channels are combined in order to form a feature vector as the feature compression vector.

8. The intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion as described in claim 6, characterized in that, The process of fusing the feature map with the attention weight coefficients includes: The feature map of each channel is multiplied by the corresponding channel weight coefficient channel by channel to obtain a weighted feature map. All weighted feature maps are then concatenated along the channel dimension to obtain the fused comprehensive feature map of the concrete surface image after removing natural textures.

9. The intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion as described in claim 1, characterized in that, The analysis of edge information and texture features in the comprehensive feature map of the concrete surface image after removing natural texture includes: Edge detection operators and texture analysis are performed on the comprehensive feature map to extract edge information and texture from the comprehensive feature map of the concrete surface image after removing natural texture. The edge information and texture features are then stitched together in the channel dimension, and a judgment combination for defect recognition is constructed based on the stitching result.

10. The intelligent identification system for multiple defects on concrete surfaces based on multi-scale feature fusion as described in claim 1, characterized in that, The method of identifying multiple different defects based on defect identification judgments includes: A defect recognition model based on defect identification judgment combination is constructed. The defect recognition model is trained, and the fused feature map is input into the defect recognition model for region-by-region analysis, outputting the spatial location and type label of the defect.