Concrete crack analysis method and system based on multi-modal image fusion
By using multimodal image fusion technology and combining visible light and structured light depth images, a triangular reference network is constructed to correct surface distortion. This solves the problem of crack width error caused by the angle between the optical axis and the normal in traditional methods, and realizes accurate quantification and type and grade assessment of concrete cracks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN JINLU FANGYUAN ENG EXPERIMENT CHECKING & MEASURATION CO LTD
- Filing Date
- 2026-06-03
- Publication Date
- 2026-07-03
AI Technical Summary
Traditional image analysis methods, when dealing with surface cracks in concrete structures, suffer from systematic projection compression errors due to the non-linear mapping between pixel span and actual physical width caused by the change in the angle between the optical axis and the surface normal. This makes it impossible to accurately identify construction joints and cracks, leading to misjudgments and delayed maintenance.
Multimodal image fusion technology is used to combine visible light and structured light depth images to construct a triangular reference network to correct surface distortion. Through adaptive fusion of high-frequency texture and surface normal gradient, a compensation correction matrix is generated, the crack feature tensor is remapped, the continuous crack skeleton is extracted, and the physical width is measured.
It eliminates surface geometric distortion, achieves precise quantification of crack width and automatic determination of crack type level, overcomes skeleton fracture and artifact interference, and improves the accuracy and reliability of crack analysis.
Smart Images

Figure CN122335633A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method and system for analyzing concrete cracks based on multimodal image fusion. Background Technology
[0002] During the inspection of cracks at the bottom of the crossbeams on the concrete main tower of the cross-river cable-stayed bridge, due to the 6.2-meter radius of curvature on the curved surface at the bottom of the crossbeam, the angle between the optical axis of the visible light camera and the local normal at various points on the concrete surface dynamically changes with the position of the curved surface, with a deviation range of 10 to 18 degrees. Traditional image analysis methods use a planar perspective projection model to process crack width inversion, without establishing a geometric constraint benchmark for the true curved surface of the concrete structure. This results in a non-linear mapping relationship between the pixel span of the crack in the image and the actual physical width when there is an angle between the optical axis and the surface normal. The larger the angle, the more severely the pixel span is compressed. In a typical inspection, a microcrack approximately 0.8 meters long was found on the arc transition section at the bottom of a crossbeam. The visible light image showed the crack's mid-segment pixel span as 3 pixels, which translates to 0.36 mm at a resolution of 0.12 mm / pixel. However, due to the 15-degree angle between the optical axis and the surface normal at this location, and the lack of geometric compensation for the arc curvature in the planar projection model, the calculated width value contained a systematic projection compression error of approximately 0.18 mm. The actual width of the crack was 0.54 mm, exceeding the 0.5 mm threshold for structural crack identification in highway bridge maintenance specifications. Because of projection distortion, the width quantification value was underestimated by approximately 33%, leading to the crack being misjudged as a shallow microcrack that did not require repair, and grouting treatment was not performed in a timely manner.
[0003] The construction joints and pre-embedded anchor ends at the bottom of the beam exhibit grayscale gradient features highly similar to crack textures in visible light images. The local image contrast is below 0.12. Traditional feature extraction methods cannot effectively distinguish between crack edges and joint artifacts in low-contrast areas. When the crack skeleton passes through the joint intersection area, its edge signal is submerged by the strong gradient interference of the joint. The skeleton sequence is discrete and segmented at the joint, and the topological continuity is destroyed. In addition, the abrupt change in local curvature in the joint intersection area causes a mismatch in coordinate mapping between the two images, further exacerbating the interruption of edge fitting. At the intersection of the construction joint and the crack at the bottom of the beam, the crack was truncated by the joint artifact, and the skeleton sequence was broken into two independent segments, north and south. The upper half of the crack was correctly identified, but the lower half of the crack was misjudged as surface attachments due to the partial lack of depth data at the joint groove. As a result, the grouting construction only covered the upper half of the crack, and the lower half of the crack continued to expand over the next three months, eventually extending to the side wall of the beam. This caused the area to be upgraded from a routine inspection to a key monitoring section, delaying the best maintenance window and increasing the complexity of subsequent treatment. Summary of the Invention
[0004] This invention provides a method and system for analyzing concrete cracks based on multimodal image fusion. By fusing depth features of visible light and structured light, a triangular reference network is constructed to correct surface distortion, thereby achieving crack geometric correction, continuous skeleton extraction, and accurate quantification of opening width, solving the problems of projection distortion error and skeleton fracture.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a concrete crack analysis method based on multimodal image fusion, the method comprising: Visible light images and structured light depth images of the same concrete structure surface are acquired. Spatial alignment is performed on the visible light images and structured light depth images to obtain a bimodal aligned image set. The texture high-frequency tensor and surface normal gradient tensor of the bimodal aligned image set are extracted. The texture high-frequency tensor and surface normal gradient tensor are adaptively fused to obtain a fused feature tensor. Based on the structured light depth image, the three-dimensional coordinates of three structural reference points are located: the center point of the bottom casting joint of the pier column, the reference point of the side edge of the transverse prestressed anchorage zone, and the intersection point of the bridge deck drainage channel facade. A triangular reference network is constructed with the three-dimensional coordinates as vertices and stretched into a triangular prism along the surface normal direction. The local projection deviation of each pixel in the fused feature tensor onto the triangular reference network within the triangular prism is calculated, and a compensation correction matrix is generated. Based on the compensation and correction matrix, spatial remapping is performed on the fused feature tensor to obtain the corrected fused feature tensor. The corrected fused feature tensor is then fitted to the sub-pixel edge along the crack direction and connected to the fracture gap to obtain a continuous crack skeleton sequence. Based on the local neighborhood point cloud fitting tangent plane at each point in the continuous crack skeleton sequence, the physical width from each point to the crack edge line on both sides is measured along the normal of the tangent plane to obtain the opening width quantization vector set; The continuous crack skeleton sequence and the opening width quantization vector set are concatenated into a crack feature vector. The crack feature vector is then used to determine the mode, output the concrete crack type and grade assessment, and complete the concrete crack analysis.
[0006] Secondly, a concrete crack analysis system based on multimodal image fusion includes: The acquisition module is used to acquire visible light images and structured light depth images of the same concrete structure surface, perform spatial alignment on the visible light images and structured light depth images to obtain a bimodal aligned image set, extract the texture high-frequency tensor and surface normal gradient tensor of the bimodal aligned image set, and adaptively fuse the texture high-frequency tensor and surface normal gradient tensor to obtain a fused feature tensor. The correction module is used to locate the three-dimensional coordinates of three structural reference points based on the structured light depth image: the center point of the bottom casting joint of the pier, the side edge reference point of the transverse prestressed anchorage zone, and the intersection point of the bridge deck drainage channel facade. It constructs a triangular reference network with the three-dimensional coordinates as vertices and stretches it into a triangular prism along the surface normal direction. It calculates the local projection deviation of each pixel in the fused feature tensor onto the triangular reference network within the triangular prism and generates a compensation correction matrix. The fusion module is used to perform spatial remapping on the fusion feature tensor according to the compensation correction matrix to obtain the corrected fusion feature tensor. The corrected fusion feature tensor is then fitted to the sub-pixel edge along the crack direction and connected to the fracture gap to obtain a continuous crack skeleton sequence. The calculation module is used to fit a tangent plane to the local neighborhood point cloud at each point in the continuous crack skeleton sequence, measure the physical width from each point to the crack edge line on both sides along the normal of the tangent plane, and obtain the opening width quantization vector set. The analysis module is used to concatenate the continuous crack skeleton sequence with the opening width quantization vector set into a crack feature vector, perform mode determination on the crack feature vector, output the concrete crack type and grade assessment, and complete the concrete crack analysis.
[0007] The above-described solution of the present invention has at least the following beneficial effects: By fusing structured light depth images and visible light images, and constructing a triangular reference network using three structural reference points for spatial remapping, the systematic compression error of crack width caused by the angle between the optical axis and the normal on curved surfaces in traditional planar projection models is overcome. This achieves the technical effect of eliminating surface geometric distortion and restoring the true physical width of cracks. By extracting high-frequency textures and performing adaptive fusion with surface normal gradients, the technical problems of similar artifact features between cracks in low-contrast areas and construction joints, and misjudgment caused by skeleton fracture are overcome. This achieves high-precision positioning of crack edges and stable extraction of continuous skeleton sequences. By fitting a tangent plane based on the local neighborhood point cloud of the skeleton and measuring the physical width along the normal vector, the technical problems of missing depth information in the two-dimensional pixel width conversion and inability to distinguish between structural cracks and surface attachments are overcome. This achieves a multi-dimensional and accurate evaluation effect of continuous quantification of crack opening width and automatic determination of type and level. Attached Figure Description
[0008] Figure 1 This is a schematic flowchart of a concrete crack analysis method based on multimodal image fusion provided by an embodiment of the present invention.
[0009] Figure 2 This is a schematic diagram of a concrete crack analysis system based on multimodal image fusion provided in an embodiment of the present invention.
[0010] Figure 3This is a curve comparing the crack width error before and after triangular reference network correction in the concrete crack analysis method based on multimodal image fusion provided by an embodiment of the present invention.
[0011] Figure 4 This is a graph showing the trend of crack width variation along the extension direction in a concrete crack analysis method based on multimodal image fusion provided by an embodiment of the present invention. Detailed Implementation
[0012] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0013] like Figure 1 As shown, embodiments of the present invention propose a concrete crack analysis method based on multimodal image fusion, the method comprising the following steps: Visible light images and structured light depth images of the same concrete structure surface are acquired. Spatial alignment is performed on the visible light images and structured light depth images to obtain a bimodal aligned image set. The texture high-frequency tensor and surface normal gradient tensor of the bimodal aligned image set are extracted. The texture high-frequency tensor and surface normal gradient tensor are adaptively fused to obtain a fused feature tensor. Based on the structured light depth image, the three-dimensional coordinates of three structural reference points are located: the center point of the bottom casting joint of the pier column, the reference point of the side edge of the transverse prestressed anchorage zone, and the intersection point of the bridge deck drainage channel facade. A triangular reference network is constructed with the three-dimensional coordinates as vertices and stretched into a triangular prism along the surface normal direction. The local projection deviation of each pixel in the fused feature tensor onto the triangular reference network within the triangular prism is calculated, and a compensation correction matrix is generated. Based on the compensation and correction matrix, spatial remapping is performed on the fused feature tensor to obtain the corrected fused feature tensor. The corrected fused feature tensor is then fitted to the sub-pixel edge along the crack direction and connected to the fracture gap to obtain a continuous crack skeleton sequence. Based on the local neighborhood point cloud fitting tangent plane at each point in the continuous crack skeleton sequence, the physical width from each point to the crack edge line on both sides is measured along the normal of the tangent plane to obtain the opening width quantization vector set; The continuous crack skeleton sequence and the opening width quantization vector set are concatenated into a crack feature vector. The crack feature vector is then used to determine the mode, output the concrete crack type and grade assessment, and complete the concrete crack analysis.
[0014] In this embodiment of the invention, by using visible light and structured light depth image fusion and constructing a triangular reference network using three structural reference points for spatial remapping, the technical problems of projection distortion and crack skeleton fracture on the arc surface are overcome, achieving the effects of eliminating geometric deformation, extracting continuous crack skeleton and accurately quantifying physical width; by fitting the tangent plane of the skeleton local point cloud to measure the opening width and splicing feature vectors for pattern determination, the problems of missing depth information and misjudgment of type are overcome, realizing multi-dimensional evaluation of crack type level automatic determination.
[0015] In a preferred embodiment of the present invention, step 1 above includes: By simultaneously acquiring visible light images and structured light depth images of the concrete structure surface under the same field of view using a visible light camera and a structured light depth module mounted on the same acquisition platform, a raw dual-modal image pair is obtained. Specifically, this involves using a visible light imaging device and a structured light depth detection module fixed on the same mobile inspection platform, ensuring that the optical axes of the two devices are parallel and their field of view coverages completely overlap, and controlling the movement speed of the inspection platform to [missing information]. The imaging equipment maintains a distance from the concrete structure surface. Simultaneous imaging was performed on the same target area of the concrete bridge piers and deck structure at the same time. The imaging trigger signal adopted a hardware synchronous triggering method, and the synchronization error was controlled within a certain range. Within this range, raw visible light images under natural lighting were acquired, with a resolution set at 1920×1080 pixels and a pixel size of 3.45. ×3.45 The original structured light depth image, containing spatial distance information, has a depth measurement range of 0.5. ~3 Depth measurement accuracy ±0.1 These are combined to form the original bimodal image pair.
[0016] Lens distortion correction and histogram equalization are performed on the visible light image of the original bimodal image pair to obtain a preprocessed visible light image. Median depth filtering and invalid depth point neighborhood interpolation filling are performed on the structured light depth image of the original bimodal image pair to obtain a depth preprocessed image. Specifically, this includes: acquiring the original pixel spatial distribution data of the visible light image; obtaining the lens radial and tangential distortion parameters from multi-angle images acquired through standard calibration; establishing a spatial mapping relationship from the original pixel position to the target pixel position based on the optical center position of the image sensor; performing a point-by-point traversal operation on all pixel regions of the visible light image; completing the geometric remapping of the original pixel positions according to the spatial mapping relationship; eliminating radial dilation and tangential deflection deformation introduced by the lens; performing histogram equalization after distortion correction; and statistically analyzing the grayscale value range of all pixels in the visible light image. The total number of grayscale levels is set to 256, corresponding to grayscale values from 0 to 255, and the number of pixels corresponding to each grayscale level is... , =0,1,...,255 The image uses a grayscale index and has a total of [number] pixels. The gray level probability is calculated as the ratio of the number of pixels at a single gray level to the total number of pixels in the image. The cumulative value of grayscale distribution is obtained by accumulating the values at each level. The accumulated value and the maximum grayscale value are multiplied and rounded. In this embodiment, the maximum grayscale value is 255, resulting in the equalized grayscale value. ,in, The grayscale value after equalization This represents the cumulative value of the grayscale distribution. This represents the grayscale probability. This is the standard rounding function, based on The grayscale values of each pixel are redistributed to increase the grayscale difference of pixels in low-contrast areas, resulting in a visible light preprocessed image. The original structured light depth image within the original bimodal image is then subjected to depth median filtering and invalid depth point neighborhood interpolation to suppress noise and repair missing data. The specific implementation process is as follows: Depth median filtering uses a 3×3 square neighborhood window. A square neighborhood with a side length of 3 is defined with each depth pixel to be processed as the center. All valid depth values within the window are traversed and arranged in ascending order. The value at the middle position of the sorted sequence is selected as the filtered output value of the current pixel. This output value is used to replace the depth values at the corresponding positions in the original image one by one. This can effectively suppress salt-and-pepper noise and random measurement outliers in the depth image. The noise suppression rate of the filtered depth image is not less than 85%. For invalid depth points without data caused by surface reflection, occlusion, etc. during structured light acquisition, they are uniformly marked as -1. The eight nearest valid depth points around each invalid depth point are selected as neighborhood reference points. The arithmetic mean of the depth values of these eight reference points is calculated. This mean is used as the depth value after filling the corresponding invalid depth point. The calculated filling value is used to fill all invalid depth positions point by point to repair the data missing areas of the depth image. It is ensured that the data integrity rate after filling is not less than 99.5%, and finally the depth preprocessed image is obtained.
[0017] Local feature descriptors are extracted between the visible light preprocessed image and the depth preprocessed image. Feature point matching is performed based on the local feature descriptors. Rigid body transformation parameters between the visible light preprocessed image and the depth preprocessed image are calculated. Based on the rigid body transformation parameters, the visible light preprocessed image and the depth preprocessed image are mapped to the same pixel coordinate system to obtain a dual-modal aligned image set. Specifically, this includes: multi-scale local feature extraction of the visible light preprocessed image and the depth preprocessed image, and constructing multi-scale Gaussian difference pyramids for each image. The Gaussian difference image calculation formula is as follows: ; in, Coordinates are The pixels at scale The Gaussian difference image values below, For the original image at scale Gaussian smoothing results As a scale factor, The scaling factor is used; corner points and edge inflection points are detected in the difference of Gaussian pyramid as feature points, requiring that the gray value of the feature point is greater than the gray value of its 8 neighboring pixels and greater than the gray value of the corresponding pixels at the adjacent scales above and below; the gray-level gradient distribution of pixels in a fixed 16×16 pixel neighborhood around each feature point is statistically encoded, and the gradient magnitude of each pixel in the neighborhood is calculated. With gradient direction ,in, pixel coordinates The gradient magnitude at a given point represents the degree of drastic change in the grayscale of that pixel. The image after Gaussian smoothing at the pixel level The partial derivative along the lateral direction; The image after Gaussian smoothing at the pixel level The partial derivative along the longitudinal direction; pixel coordinates The gradient direction, representing the direction of grayscale change, is divided into 8 intervals, each corresponding to 4 pixels. The sum of the gradient magnitudes in each interval is calculated to form a 128-dimensional local feature descriptor vector with rotation and illumination invariance properties. Feature point matching is performed based on the extracted local feature descriptor, using Euclidean distance as the similarity metric. A certain feature descriptor vector in the visible light preprocessed image is set as... The feature description vector in the depth preprocessed image is The sum of the absolute values of the differences in the corresponding dimensions of the two feature description vectors is: Set the matching threshold to , The value was determined through multiple experiments. In this embodiment... =150, where This is the threshold for determining feature point matching, used to distinguish between matching and non-matching feature point pairs; when the sum is accumulated... Less than the preset threshold At that time, two sets of features are determined to be matching feature point pairs; simultaneously, the ratio method is used to eliminate mismatched points, and the best matching point and the second best matching point for each feature point are calculated. The ratio of the feature points is used to retain the matching pair when the ratio is less than 0.8. Multiple reliable sets of identically named feature points are selected, with a matching accuracy ≥ 90%. The selected sets of matching feature point pairs are then used to calculate the rigid body transformation parameters between the two images. The rigid body transformation only includes translation and rotation, without changing the image scaling or shape. The lateral translation amount of the image plane is set to... The vertical translation of the image is The image plane is rotated by an angle of 100°. The coordinate transformation relationship of a rigid body is: ; in, The pixel coordinates of the depth preprocessed image pixels are mapped to the visible light preprocessed image coordinate system; The original pixel coordinates of the depth preprocessed image; This is the rotation angle of the image plane, used to achieve rotation transformation of the depth image; Rotate the image plane The cosine value; Let H be the sine value of the rotation H of the image plane; This represents the horizontal translation of the image plane. To determine the vertical translation of the image plane, select at least four pairs of non-collinear matching feature points, and then... and Substituting the above relationships, we construct an overdetermined system of equations and use the least squares method to solve for the rigid body transformation parameters E, F, and H. The least squares objective function is as follows: ; in, To match the number of feature point pairs, For the first In the group matching feature point pairs, the horizontal coordinates of the pixels in the depth preprocessed image are mapped to the coordinate system of the visible light preprocessed image. For the first In the group matching feature point pairs, the horizontal coordinates of the original pixels in the depth preprocessed image are used. For the first In the group matching feature point pairs, the vertical coordinates of the original pixels in the depth preprocessed image are used; For the first In the group matching feature point pairs, the vertical coordinates of the pixels in the depth preprocessed image are mapped to the coordinate system of the visible light preprocessed image. is the rotation angle of the image plane; E is the lateral translation of the image plane. The image plane is represented by the longitudinal translation amount; the objective function obtains its optimal parameters by minimizing the square of the total error, and the optimal transformation parameters are obtained by solving for them. Based on the rigid body transformation parameters obtained from the solution, the coordinates of any original pixel in the depth preprocessed image are... Substitute the coordinate transformation equations sequentially to complete the rotation transformation calculation, and then superimpose the horizontal translation E and the vertical translation E. After completing the coordinate offset correction, all pixels of the depth preprocessed image are uniformly mapped to the same pixel coordinate system of the visible light preprocessed image, so that the pixels of the two images correspond one-to-one and the registration error is ≤1 pixel. After completing the mapping and matching, a complete bimodal aligned image set is obtained.
[0018] A multi-layer Gaussian image sequence is constructed based on visible light images from a dual-modal aligned image set. Inter-layer differencing is performed on each layer of the Gaussian image sequence to obtain the texture high-frequency tensor. Surface normal vectors are calculated point-by-point from the structured light depth images in the dual-modal aligned image set, and the spatial gradient derivative of the surface normal vectors is calculated to obtain the surface normal gradient tensor. Specifically, this includes: constructing a multi-layer Gaussian image sequence based on preprocessed visible light images from the dual-modal aligned image set, and setting the Gaussian smoothing kernel size to [value missing]. In this embodiment That is, a 5×5 Gaussian kernel, where Let be the size of the Gaussian smoothing kernel. The Gaussian smoothing kernel function is as follows: ; in, Gaussian smoothing kernel in coordinates The core value at the location; is the base of the natural logarithm. The standard deviation is Gaussian; the standard deviation is taken layer by layer. Gaussian convolution operations are applied layer by layer to smooth and reduce noise in the visible light preprocessed image. The convolution calculation formula is: ,in, For the first After Gaussian smoothing, pixel coordinates The grayscale value at that location; Visible light preprocessed image pixel coordinates The original grayscale value at that location; This represents the convolution operation; after each smoothing layer, a Gaussian smoothed image with progressively larger scale is generated. These images are arranged in ascending order of scale to form a five-layer Gaussian image sequence. Inter-layer difference operations are performed on each of the constructed Gaussian image sequences to extract the high-frequency texture tensor and suppress low-frequency background noise. The difference is solved by traversing all adjacent Gaussian images layer by layer. The difference results of all layers are combined and arranged according to the dimension to form the high-frequency texture tensor.
[0019] Based on the structured light depth preprocessed image from the dual-modal aligned image set, the surface normal vector is calculated point-by-point to characterize the geometry of the concrete surface. A 3×3 neighborhood 3D point cloud of the current depth pixel is selected, comprising 9 points. The 3D coordinates of each point are set as follows: in, Let be the index of a point in the neighborhood. Fit a local spatial plane, and the equation of the spatial plane is: ,in, is the lateral coefficient of the plane equation, representing the degree of inclination of the plane in the X direction; The longitudinal coefficient of the plane equation represents the degree of inclination of the plane in the Y direction; is the vertical coefficient of the plane equation, representing the degree of inclination of the plane in the Z direction; For the three-dimensional coordinates of a point in space; The constant term in the plane equation is used to adjust the spatial position of the plane. The plane coefficients are solved using the least squares method, and the objective function is as follows: ; in, The index of the point within the neighborhood; is the lateral coefficient of the plane equation, representing the degree of inclination of the plane in the X direction; The longitudinal coefficient of the plane equation represents the degree of inclination of the plane in the Y direction; is the vertical coefficient of the plane equation, representing the degree of inclination of the plane in the Z direction. These are constant terms in the plane equation, used to adjust the spatial position of the plane. For the first The three-dimensional coordinates of each point; the objective function obtains the optimal plane coefficients by minimizing the square of the total distance, and the plane coefficients are obtained by solving the problem. Then, the plane coefficients are normalized to obtain the unitized surface normal vectors. The surface normal vectors for each pixel are calculated by traversing all pixels in the depth image, forming a three-dimensional normal vector field. Spatial gradient differentiation is performed on all the obtained surface normal vectors to construct the surface normal gradient tensor. The rate of change of the normal vectors in the horizontal and vertical directions of the image are calculated separately. The difference between the normal vectors of adjacent pixels in the horizontal direction is set to... Use the difference Divide by step size Obtain the horizontal gradient value ,in, pixel coordinates The gradient value of the normal vector in the horizontal direction represents the rate of change of the normal vector along the X direction; It represents the difference in normal vectors between horizontally adjacent pixels; For a unit spatial step size, similarly, calculate the difference in normal vectors between vertically adjacent pixels, and the vertical gradient value. ,in, pixel coordinates The gradient value of the normal vector in the vertical direction represents the rate of change of the normal vector along the Y direction; It represents the difference in the normal vectors of adjacent pixels in the vertical direction; For a unit spatial step size, the gradient values of the two directions are combined into gradient feature quantities, which are then arranged regularly according to the pixel spatial position to form the surface normal gradient tensor.
[0020] Based on the gray-level difference amplitude of each pixel's neighborhood in the texture high-frequency tensor and the depth change amplitude at the corresponding position in the surface normal gradient tensor, weighting coefficients for the texture high-frequency tensor and the surface normal gradient tensor are calculated pixel-by-pixel. These weighting coefficients are then used to perform pixel-by-pixel weighted fusion of the texture high-frequency tensor and the surface normal gradient tensor to obtain a fused feature tensor. Specifically, this includes: calculating the gray-level difference amplitude within the neighborhood of the texture high-frequency tensor pixel-by-pixel to characterize the clarity of texture details; defining a 3×3 neighborhood around the current pixel; calculating the gray-level values of all pixels within the neighborhood; and setting the maximum gray-level value in the neighborhood to be [value missing]. The minimum grayscale value is The grayscale difference amplitude of the current pixel is obtained by subtracting the minimum grayscale value from the maximum grayscale value. Simultaneously extract the depth change amplitude of the pixel corresponding to the surface normal gradient tensor to characterize the degree of surface geometric abrupt change. Set the maximum value of the normal gradient within a 3×3 neighborhood of the current position as... The minimum value is Subtract the minimum value from the maximum value to obtain the current pixel depth change magnitude. Based on the grayscale difference amplitude and depth change amplitude obtained above, the weighting coefficients of the texture high-frequency tensor and the surface normal gradient tensor are calculated pixel by pixel, and the texture weight benchmark coefficient is set as follows. The depth weighting baseline coefficient is Calibration through testing , It balances the weighting of texture details and geometric features, among which, This is the baseline coefficient for texture weights, used to set the basic weights of the high-frequency texture tensor; The depth weight baseline coefficient is used to set the basic weights of the surface normal gradient tensor. It is obtained by summing the gray-level difference amplitude of a single pixel with the depth change amplitude, and then calculating the pixel weighting coefficients corresponding to the high-frequency tensor of the texture. ,in, These are the pixel weighting coefficients of the high-frequency texture tensor, representing the weight ratio of texture features in the fusion process; This represents the magnitude of the grayscale difference of the current pixel. This is the baseline coefficient for texture weights; The formula for calculating the pixel weighting coefficients of the surface normal gradient tensor, which is the sum of magnitudes, is as follows: ,in, , where is the pixel weighting coefficient of the surface normal gradient tensor, representing the weight ratio of depth geometric features in the fusion; This represents the magnitude of the depth change for the current pixel. This is the baseline coefficient for depth weighting; The sum of amplitudes, where To determine the rationality of the weighted fusion, the weights of texture and depth features are summed to 1. Using the weighting coefficients calculated above, a pixel-by-pixel weighted fusion is performed on the texture high-frequency tensor and the surface normal gradient tensor. Let the pixel value corresponding to the texture high-frequency tensor be... The pixel value corresponding to the surface normal gradient tensor is The high-frequency tensor value of the texture is multiplied pixel by pixel with the corresponding texture weighting coefficient to obtain... ,in, This represents the feature value of the high-frequency texture tensor at the current pixel. The pixel weighting coefficients of the high-frequency texture tensor; This is the weighted contribution value of the texture features, obtained by multiplying the surface normal gradient tensor value with the corresponding depth weighting coefficient. ,in, This represents the feature value of the surface normal gradient tensor at the current pixel. These are the pixel weighting coefficients of the surface normal gradient tensor; The weighted contribution value of depth geometric features is calculated by summing the results of multiplication of two terms at the same location, and the fusion formula is as follows: ,in, The resulting pixel feature values combine texture details and depth geometry features. Weighted contribution value for texture features; The weighted contribution value of the depth geometric features is used to perform adaptive weighted fusion calculation for each pixel. After traversing all pixels of the entire tensor, the fused feature tensor of the fused texture details and depth geometric features is obtained after all calculations are completed.
[0021] In this embodiment of the invention, the visible light image is distorted and histogram equalized by checkerboard calibration, which overcomes the problems of original image distortion, missing depth data and modal mismatch, and achieves high-quality alignment and adaptive enhancement of dual-modal features.
[0022] In a preferred embodiment of the present invention, step 2 above includes: Based on the two joint edge lines at the bottom of the pier in the structured light depth image, the spatial intersection of the two joint edge lines is solved to obtain the three-dimensional coordinates of the center point of the bottom of the pier's casting joint. Specifically, this involves three steps: edge extraction, line fitting, and intersection point solving, based on the two joint edge lines at the bottom of the pier in the structured light depth image, to obtain the three-dimensional coordinates of the center point of the bottom of the pier's casting joint. The specific implementation process is as follows: Edge detection is performed on the preprocessed structured light depth image, with a high threshold of 80 and a low threshold of 40. By traversing the bottom region of the pier in the image, the two continuous edge lines at the casting joint are accurately identified and denoted as edge lines. and edge line Among them, the edge line All are spatial polylines, composed of the three-dimensional coordinates of several consecutive pixels, with each pixel's coordinates denoted as... ;in, The pixel index is used for the edge line. Based on the edge line extraction, spatial line fitting is performed on both edge lines. The least squares method is used to eliminate pixel discrepancy errors, fitting the discrete edge line pixels to a continuous spatial line, thus obtaining the spatial line equations for the two edge lines. For all pixels on the edge lines, an optimization objective is constructed based on the least squares principle: minimizing the sum of squared perpendicular distances from each pixel to the fitted spatial line. In this process, only the geometric projection relationship from the spatial point to the line needs to be established to define the error function. By setting the partial derivatives of the error function with respect to the direction and position parameters to zero, a system of linear equations is constructed and solved to obtain the coefficients of the spatial line equations. The solved coefficients are then converted into point-to-point equations of the spatial line and further expressed as parameter equations. Based on the parametric equations corresponding to the two edge lines, a system of simultaneous equations with independent variables is constructed. Since the two edge lines represent the physical boundaries of the cast-in-place joint at the bottom of the pier, they have a unique intersection point in spatial geometry. The Gaussian elimination method is directly used to solve the system of simultaneous equations to obtain a unique solution for the parameters. The obtained parameter values are substituted into the parametric equation of any edge line to calculate the three-dimensional coordinates of the intersection point, which is the center point of the cast-in-place joint at the bottom of the pier. To ensure the robustness of the coordinate calculation, if the system of equations does not have a unique solution during the simultaneous solution process, the closest pair of points between the two spatial lines is extracted in space, and the coordinates of the midpoint of the closest pair are calculated. The coordinates of the midpoint are used as the final three-dimensional coordinates of the center point of the cast-in-place joint at the bottom of the pier, thus obtaining the three-dimensional coordinates of the intersection point C. The three-dimensional coordinates of the center point of the bottom casting joint of the pier column were solved.
[0023] Based on the depth abrupt change bands on the side edge of the transverse prestressed anchorage zone in the structured light depth image, the extreme points of the depth gradient are extracted along the direction of the depth abrupt change bands to obtain the three-dimensional coordinates of the reference points on the side edge of the transverse prestressed anchorage zone. Specifically, after solving for the coordinates of the center point of the casting joint at the bottom of the pier, in order to construct a complete triangular reference network, it is necessary to further extract the feature reference points on the side edge of the transverse prestressed anchorage zone. Based on the depth abrupt change bands on the side edge of the transverse prestressed anchorage zone in the structured light depth image, the three-dimensional coordinates of the reference points on the side edge of the transverse prestressed anchorage zone are obtained through depth abrupt change band identification, depth gradient calculation, and gradient extreme point extraction. The specific implementation process is as follows: Gradient calculation is performed on the preprocessed structured light depth image. The Sobel gradient operator is used as a first-order edge detection operator for discrete difference to calculate the transverse gradient of the image. and longitudinal gradient gradient magnitude Set gradient magnitude threshold Through multiple tests and calibrations, it was determined that it can effectively distinguish between abrupt changes and noise when the pixel gradient magnitude... When a pixel is identified as a depth abrupt change point, all depth abrupt change points are arranged continuously according to their spatial location, forming a depth abrupt change zone on the side edge of the transverse prestressed anchorage zone, denoted as... Based on the identification of the depth abrupt change zone M, the depth gradient change is calculated point-by-point along the extension direction of the abrupt change zone using interpolation, thereby characterizing the degree of depth change at each point within the abrupt change zone; specifically, three consecutive adjacent pixels within the abrupt change zone are grouped together, and the pixel coordinates are set as follows: , , Calculate the depth difference between adjacent pixels Combining the unit spatial step size, with 1 pixel corresponding to a physical distance of 3.45μm, the depth gradient value of each point is obtained. Traversing deep mutation zones For all pixels, the point with the largest gradient value is selected as the characteristic reference point of the lateral edge of the transverse prestressed anchorage zone, denoted as . The point has three-dimensional coordinates as During the extraction process, the 5×5 neighborhood around the extreme point is verified with a verification accuracy of ≤1 pixel, thus completing the extraction of the three-dimensional coordinates of the reference point on the side edge of the transverse prestressed anchorage zone.
[0024] Based on the intersection line between the bridge deck drainage channel elevation and the pier surface in the structured light depth image, the lowest point of the intersection line is extracted to obtain the three-dimensional coordinates of the intersection point of the bridge deck drainage channel elevation. Specifically, after extracting the coordinates of the reference points on the side edge of the transverse prestressed anchorage zone, the triangulation reference network still lacks a key reference point, requiring further extraction of the three-dimensional coordinates of the intersection point of the bridge deck drainage channel elevation. Based on the intersection line between the bridge deck drainage channel elevation and the pier surface in the structured light depth image, the specific implementation process is as follows: The structured light depth preprocessed image is segmented into regions using a threshold segmentation method, with a depth threshold set. Based on the depth difference between the drainage ditch and the pier, this embodiment is calibrated. Distinguish between the bridge deck drainage channel elevation area and the pier surface area; extract the boundary line between the two areas, which is the intersection line between the bridge deck drainage channel elevation and the pier surface, and denoted as the intersection line. Intersection It is a spatial polyline, composed of the three-dimensional coordinates of several consecutive pixels, with the coordinates of each pixel denoted as . ,in, The pixel number of the intersection line; after the intersection line is completed Based on the extraction, traverse the intersection lines. The three-dimensional coordinates of all pixels are analyzed, with a focus on extracting the vertical depth value of each pixel. Because the intersection line between the drainage ditch facade and the pier surface is sloping, its lowest point corresponds to the pixel with the smallest vertical depth value; for the intersection line... Smoothing was performed using a 3×3 Gaussian filter with a Gaussian standard deviation. Consistent with the Gaussian smoothing parameters mentioned earlier, discrete noise points are removed; the intersection lines are based on the smoothed lines described earlier. Pixel data, filtering out intersection lines based on the smoothed data described above. Pixel data, filter out vertical depth values The smallest pixel, which is the intersection of the bridge deck drainage channel facade, is denoted as point D, and its three-dimensional coordinates are: The point was verified to confirm that the depth values of the surrounding pixels were all greater than that point, and that the point was located at the actual intersection of the drainage channel and the pier. The coordinate extraction error was determined to be ≤0.3mm, and the three-dimensional coordinates of the intersection point of the bridge deck drainage channel facade were extracted.
[0025] Connect the three-dimensional coordinates of the center point of the cast-in-place joint at the bottom of the pier, the reference point of the side edge of the transverse prestressed anchorage zone, and the intersection point of the bridge deck drainage channel elevation to form a spatial triangle, thus obtaining a triangular reference network. Extract the unit normal vector of the plane containing the triangular reference network, and stretch the triangular reference network to a preset height along the direction of the unit normal vector to obtain a triangular prism. Specifically, this includes: connecting the three reference points obtained above, point C... Point P Point D The three-dimensional coordinates are connected sequentially to form a closed spatial triangle, which is the triangular reference network, denoted as . ; For the triangular reference network Spatial fitting is performed on the plane to solve for the unit normal vector of the plane; two vectors in the plane are calculated, which together constitute the support vector of the plane containing the triangulation reference network; the plane normal vector is calculated by vector cross product. ,in For the triangular reference network The normal vector of the plane is used to characterize the spatial orientation of the plane, and its cross product is calculated as follows: ; in, , , Normal vector The three components correspond to the components in the X, Y, and Z directions, respectively, and are used to quantize the spatial orientation of the normal vector; The three-dimensional coordinate components of the center point C of the joint at the bottom of the pier; The three-dimensional coordinate components of the reference point P on the side edge of the transverse prestressed anchorage zone; The three-dimensional coordinate components of the intersection point D of the bridge deck drainage channel elevation; , , These are the unit vectors in the X, Y, and Z directions, forming the first row of the determinant; the normal vector... Normalization is performed to obtain the unit normal vector. ,in Normal vector The modulus is used for normalization to make the unit normal vector... The mold length is 1; the preset stretching height is set to... In this embodiment According to the size of the detection area, the calibration is performed. The translation distance of the triangular reference network along the unit normal vector direction, i.e., the height of the triangular prism, is used to determine the spatial height range of the triangular prism; along the unit normal vector... The direction of the triangular reference network The entire plane is translated equidistantly, with the translation distance equal to the preset stretching height. The coordinates of the three reference points after translation are as follows: , , , , , These are the corresponding points after the translation of points E, F, and G. Their spatial positions correspond one-to-one with the original reference points, and the translation direction is the same as the unit normal vector. Consistent; Original triangulation reference network With the translated triangular reference network Connect the corresponding vertices in sequence ( , , ( ), which encloses and forms a closed spatial polyhedron, namely a triangular prism.
[0026] The two-dimensional coordinates of each pixel in the fused feature tensor are mapped onto a triangular prism. The vertical projection coordinates of each pixel along the height of the prism onto the triangular reference network are calculated. Based on the vertical projection coordinates and the original spatial position of the corresponding pixel in the fused feature tensor, the local projection deviation of each pixel is calculated point by point. The local projection deviations of all pixels are combined to obtain the compensation and correction matrix. Specifically, this includes mapping the two-dimensional coordinates of each pixel in the fused feature tensor onto a triangular reference network. Mapped into the body of the triangular prism, and combined with the coordinate mapping relationship obtained from the previous dual-modal registration, the two-dimensional pixel coordinates are converted into corresponding three-dimensional spatial coordinates. ; Calculate the position of each pixel along the height direction of the triangular prism on the triangular reference mesh. Vertical projection point coordinates The projection calculation uses the formula for the perpendicular projection from a spatial point to a plane, based on the coordinates of the perpendicular projection point. Original spatial coordinates of the corresponding pixel Calculate the local projection deviation point by point. The deviation calculation formula is: The local projection deviation value of each pixel is obtained; the local projection deviations of all pixels are arranged in a regular manner according to the pixel spatial position to form a compensation and correction matrix with the same dimension as the fused feature tensor.
[0027] In this embodiment of the invention, the three-dimensional coordinates of three structural reference points—the center point of the bottom casting joint of the pier column, the side edge reference point of the transverse prestressed anchorage zone, and the intersection point of the bridge deck drainage channel elevation—are extracted by structural light depth images. This overcomes the technical problem of systematic projection compression error of crack width caused by the change of the angle between the optical axis and the surface normal on the curved surface of the concrete in the traditional planar projection model, and achieves the technical effect of eliminating surface distortion based on the real structural geometry.
[0028] In a preferred embodiment of the present invention, step 3 above includes: The local projection deviation value of each pixel in the compensation correction matrix is pixel-by-pixel superimposed onto the original spatial coordinates of the corresponding pixel in the fused feature tensor to obtain the corrected fused feature tensor. Specifically, this includes: defining each key parameter, where... To compensate for the coordinates in the correction matrix The local projection deviation value of the pixel. To fuse coordinates in feature tensors The original three-dimensional spatial coordinates of the pixel. This refers to the three-dimensional spatial coordinates of the corresponding pixels in the fused feature tensor after correction. Let be the unit normal vector of the plane containing the triangulation reference network, used to determine the direction of deviation superposition. Since local projection deviations are generated along the direction of the unit normal vector, deviation superposition must be performed in conjunction with the direction of the normal vector. The specific superposition formula is as follows: ; ; ; The above superposition calculation is performed pixel by pixel, traversing all pixel points of the fused feature tensor. After the calculation is completed, all the corrected pixel 3D coordinates are arranged in a regular manner according to their original spatial positions, and combined with the original feature components to form the corrected fused feature tensor.
[0029] The magnitude and direction of the gray-level gradient in the neighborhood of each pixel in the fused feature tensor after correction are calculated. Pixels with the largest gray-level gradient magnitude exceeding a preset gradient threshold are selected to obtain the crack edge pixel set. This specifically includes setting key parameters, where... Coordinates in the fused feature tensor after correction The grayscale gradient magnitude of the pixel at that location. This represents the grayscale gradient direction of the pixel. To preset the grayscale gradient threshold, this embodiment calibrates it through experiments. To effectively distinguish crack edges from background pixels, the Sobel gradient operator is used to calculate the gray-level gradient of the corrected fused feature tensor. After calculating the gray-level gradient magnitude and direction pixel by pixel, pixels that meet the requirements are selected. Furthermore, the pixel with the maximum gradient magnitude in its 3×3 neighborhood is considered a candidate point for the crack edge. All candidate points are then arranged in a regular spatial coordinate pattern to form a set of crack edge pixels.
[0030] A cubic spline function is fitted piecewise along the direction of the crack edge pixel set to determine the precise sub-pixel position of each edge pixel between adjacent pixels, thus obtaining the sub-pixel edge point set. Specifically, this includes: [the process of fitting the crack edge pixel set...] Preprocessing is performed by sorting the edge points along their direction to remove discrete noise points. A 3×3 neighborhood mean filter is then used to eliminate isolated points more than 0.5 mm away from their neighbors, resulting in a sorted edge point sequence. This sequence is then divided into fitting segments, grouped into sets of three adjacent points. Let the coordinates of three consecutive edge points within a fitting segment be... , , Construct a cubic spline fitting function, where the fitting function in the X direction is... The fitting function in the Y direction is The fitting function in the Z direction is ,in For fitting parameters ∈[0,1], The fitting coefficients are the boundary conditions, taking into account the coordinates of the three edge points. =0 corresponds to the first point. When =1, it corresponds to the third point; the middle point corresponds to... =0.5, solve for the coefficients of each fitting segment to obtain a continuous cubic spline fitting curve. Along the fitting curve, sample at a step size of 0.1 pixels corresponding to a physical distance of 0.345μm to obtain the sub-pixel coordinates of each sampling point. These sampling points are the sub-pixel edge points. Integrate the sub-pixel edge points of all fitting segments, remove duplicate points, and obtain the sub-pixel edge point set.
[0031] Based on the sub-pixel edge point set, the spatial distance between two adjacent sub-pixel edge points is calculated, and adjacent edge point pairs whose spatial distance exceeds a preset distance threshold are identified to obtain a continuous crack skeleton sequence. This specifically includes setting key parameters, where... sub-pixel edge point set The spatial distance between the q-th point and the (q+1)-th point in the array. In this embodiment, a preset distance threshold is used. Based on sub-pixel precision and crack width calibration, it is used to determine whether edge points are broken. To find the number of interpolation points, iterate through the sub-pixel edge point set. Calculate the spatial distance between two adjacent sub-pixel edge points pair by pair, for those satisfying Adjacent edge point pairs are identified as crack edge fractures and require interpolation for completion. During interpolation, the spatial distance between the two points and a preset step size of 0.2 are used. / point, calculate the number of interpolation points required. By using linear interpolation, the three-dimensional coordinates of the interpolation points are calculated point by point along the line connecting the two edge points. The interpolation formula is as follows: ; ; ; in, , , These are the X, Y, and Z three-dimensional spatial coordinates of the interpolation point, respectively. , , These are the first sub-pixel edge points in the set. The X, Y, and Z three-dimensional spatial coordinates of each sub-pixel edge point; , , These are the first sub-pixel edge points in the set. The X, Y, and Z three-dimensional spatial coordinates of each sub-pixel edge point; The interpolation point number has a value range of 1, 2, ... Each interpolation point is assigned a number to distinguish it from other interpolation points. To determine the number of interpolation points, all interpolation points are added to the sub-pixel edge point set and reordered according to their spatial location to obtain a complete continuous edge point sequence. Then, the central skeleton point of this sequence is extracted and integrated to form a continuous crack skeleton sequence.
[0032] In this embodiment of the invention, by superimposing the local projection deviation value of each pixel in the compensation correction matrix pixel by pixel along the normal vector direction to the original spatial coordinates of the fused feature tensor, the technical problem of distortion of crack spatial position caused by surface geometric distortion is overcome, and the technical effect of accurate matching between the spatial position of the fused feature tensor after correction and the geometric topology of the real concrete surface is achieved. By calculating the magnitude and direction of the gray-level gradient of each pixel neighborhood in the fused feature tensor after correction, the technical problem of missed detection caused by the crack edge being submerged by background noise in the low contrast area is overcome, and the reliable extraction of high confidence candidate points of crack edge is realized, achieving the technical effect of subpixel-level smooth and continuous characterization of crack edge.
[0033] In a preferred embodiment of the present invention, step 4 above includes: Based on the spatial coordinates of each skeleton point in the continuous crack skeleton sequence, depth points within a sphere of a preset radius centered at each skeleton point are extracted from the structured light depth image to obtain the local neighborhood point cloud of each skeleton point. Specifically, this includes: based on the continuous crack skeleton sequence, the surrounding depth point cloud data needs to be extracted for each skeleton point, assuming... For the skeleton point number, In this embodiment, the spherical cutoff radius is preset. Based on the crack width detection range and sub-pixel accuracy calibration, the crack area that can completely cover the periphery of the skeleton point is determined. The local neighborhood point cloud of a skeleton point consists of the three-dimensional coordinates of several depth points. This process iterates through each skeleton point in a continuous crack skeleton sequence. With the skeletal point as the center of the sphere, a preset radius is defined. Let be the radius, and construct a spherical region in space. The mathematical expression for the spherical region is: ; Traverse all depth points in the structured light depth image, filter out depth points whose coordinates satisfy the above spherical region expression, and arrange these depth points in a regular spatial coordinate pattern to form the first... The local neighborhood point cloud of a skeleton point is composed of the three-dimensional coordinates of several depth points.
[0034] For each skeleton point, a plane fitting is performed on the local neighborhood point cloud. The normal vector of the fitted plane is extracted as the surface normal vector of the corresponding skeleton point, thus obtaining the surface normal vector for each skeleton point. Specifically, this involves: performing plane fitting using the least squares method on the local neighborhood point cloud of each skeleton point to eliminate the influence of discrete noise at depth points, obtaining the fitted plane containing the local neighborhood point cloud, and extracting the normal vector of the fitted plane as the surface normal vector of the skeleton point. For each depth point in the local neighborhood point cloud, an optimization objective is constructed using the least squares principle, namely, minimizing the sum of squared vertical distances from each depth point to the fitted plane. By setting the partial derivative of the error function with respect to the coefficients of each plane equation and setting it equal to zero, a system of linear equations is constructed and solved to obtain the coefficients of the plane equations. The normal vector of the fitted plane is directly determined by the coefficients of the plane equations. The normal vector of the fitted plane is normalized to obtain the unit surface normal vector corresponding to each skeleton point. The above operations are repeated to perform plane fitting and normal vector extraction on the local neighborhood point cloud of each skeleton point one by one, thus obtaining the unit surface normal vector corresponding to each skeleton point, completing the extraction of surface normal vectors for all skeleton points.
[0035] Starting from each skeleton point, the scanning and acquisition proceeds pixel by pixel along the corresponding surface normal vector direction to both sides. After acquisition and correction, the gray values at each advancing position in the fused feature tensor are used to locate the pixel positions where the gray values change abruptly as crack edge points, thus obtaining the left and right crack edge points corresponding to each skeleton point. Specifically, this includes: obtaining the unit surface normal vector of each skeleton point. Then, using the normal vector direction as the scanning reference, the left and right edges of the crack corresponding to each skeleton point are located, providing a boundary basis for calculating the opening width. The specific implementation process is as follows: Determine the key parameters of the advancing scan, where the advancing step size is 0.1 pixels, consistent with the sub-pixel step size mentioned earlier, corresponding to a physical distance of 0.345μm. The grayscale abrupt change threshold is used in this embodiment. Through experimental calibration, it was ensured that the grayscale difference between cracks and the background could be effectively distinguished. The grayscale abrupt change judgment criterion was: the absolute value of the difference in grayscale values between two consecutive advancement positions was greater than 1. ,Right now ,in , The first Step, First The grayscale value at the +1 step advancement position, where, To advance the step size sequence, for each skeleton point Starting from that point, along its unit surface normal vector The positive direction is denoted as the right side, and the negative direction as the left side. A pixel-by-pixel scanning process is performed, with the scanning direction determined by the surface normal vector. During the scanning process, the grayscale value of the corresponding scanning position in the fused feature tensor after correction is gradually acquired. The difference in grayscale values between two adjacent scanning positions is calculated in real time to determine whether the grayscale change criterion is met. When a grayscale change is detected, the scanning in that direction is stopped, and the pixel position at the grayscale change is taken as the crack edge point in that direction. The edge point obtained by scanning along the positive direction of the surface normal vector is the right crack edge point, denoted as... The edge points obtained by scanning along the negative direction of the surface normal vector are the left crack edge points, denoted as... The located edge points are verified by scanning again within their 3×3 neighborhood. After verification, the left crack edge points corresponding to each skeleton point are obtained. and the edge of the crack on the right Complete the positioning of all skeleton points corresponding to the crack edge points.
[0036] Calculate the pixel distance from each skeleton point to the corresponding left crack edge point and the pixel distance to the corresponding right crack edge point. Summate the pixel distances on both sides to obtain the pixel span. Convert the pixel span to physical width to obtain the opening width value corresponding to each skeleton point. Specifically, based on the located left and right crack edge points, calculate the crack opening width at each skeleton point through distance calculation and unit conversion, achieving point-by-point quantization of the width. The specific implementation process is as follows: Determine the distance calculation and conversion standard, where pixel distance is calculated using Euclidean distance, and the physical width conversion is based on the previously defined unit pixel physical distance, where 1 pixel corresponds to 3.45. For each skeleton point Calculate the distance from the left crack edge point to each point. and the edge of the crack on the right The pixel distance is calculated using the Euclidean distance method described above. The distance calculated here is the pixel-level Euclidean distance, in pixels. The pixel distances on both sides are summed to obtain the crack pixel span corresponding to the skeleton point. ,in, For the first The crack pixel span corresponding to each skeleton point, that is, the pixel width of the crack at that skeleton point. For the skeleton point number, These are the pixel distances from the skeleton points to the left and right crack edge points, respectively. The pixel span is converted to physical width using the following formula: ; For the first The physical value of the crack opening width corresponding to each skeleton point, in units of ; Let be the crack pixel span of the skeleton point, where For the first The crack opening width value corresponding to each skeleton point, in units of It can be converted to according to needs. 1 =1000 The converted physical width is corrected, and the error caused by background noise is deducted. The correction value is 0.05. After multiple tests and calibrations, the final opening width value was obtained. , For the first The final crack opening width value corresponding to each skeleton point, in units of ; This is the converted physical width, 0.05. For noise correction, ensure the error in the opening width value is ≤0.1. To meet the sub-pixel quantization requirements, repeat the above operation to obtain the opening width value corresponding to each skeleton point. .
[0037] The opening width value corresponding to each skeleton point is organized along the extension order of the continuous crack skeleton sequence to obtain an opening width quantization vector set. Specifically, this involves: determining the correspondence between the extension order of the continuous crack skeleton sequence and the skeleton point numbers; arranging the skeleton points sequentially from 1 to the total number according to the natural extension direction of the crack; ensuring the arrangement order of the opening width values is consistent with the crack extension direction to fully characterize the variation law of the crack opening width along the path; and arranging the opening width value corresponding to each skeleton point sequentially according to the number to construct an opening width quantization vector set. The dimension of this vector set is consistent with the total number of skeleton points, and each element corresponds to the opening width value of a skeleton point, with the order completely consistent with the extension order of the continuous crack skeleton sequence. The opening width quantization vector set is then normalized, using the three-standard-deviation criterion to remove outliers, eliminating opening width values exceeding the mean ± three standard deviations; and filling missing values with the mean of the opening width values of two adjacent skeleton points to ensure the completeness and accuracy of the quantization vector set. After normalization, the final opening width quantization vector set is obtained.
[0038] In this embodiment of the invention, local neighborhood point clouds are extracted centered on each point of the continuous crack skeleton sequence, and surface normal vectors are extracted by fitting the tangent plane using the least squares method. This overcomes the technical problem of deviation in normal vector calculation caused by depth noise interference and achieves high-precision local adaptive solution of surface normal vectors at each skeleton point.
[0039] In a preferred embodiment of the present invention, step 5 above includes: The extension length of each skeleton point from the starting point is calculated based on the continuous crack skeleton sequence to obtain the skeleton extension length sequence. This sequence is then paired point-by-point with the opening width quantization vector set according to the skeleton point positions to obtain the crack feature vector. Specifically, this involves: determining the starting point of the continuous crack skeleton sequence, using the first skeleton point as the starting point, which is the initial position of the crack extension; calculating the extension length of all skeleton points based on this starting point; and for each skeleton point in the continuous crack skeleton sequence, sequentially calculating the spatial straight-line distance between the skeleton point and the starting point along the crack extension direction, which is the extension length of that skeleton point from the starting point. The extension length values of all skeleton points are arranged sequentially according to the order of the corresponding skeleton points in the continuous crack skeleton sequence to form a skeleton extension length sequence. This sequence has the same number of skeleton points as the continuous crack skeleton sequence and the opening width quantization vector set, and each extension length value corresponds one-to-one with a corresponding skeleton point. The skeleton extension length sequence and the opening width quantization vector set are paired point by point, that is, the extension length value corresponding to each skeleton point corresponds one-to-one with the opening width value corresponding to that skeleton point. Each pair of paired extension length values and opening width values is used as a feature element. All feature elements are integrated according to the order of skeleton points to form a crack feature vector.
[0040] Based on the opening width and skeleton extension length values in the crack feature vector, the ratio of the opening width to the skeleton extension length is calculated. This ratio is then compared with a preset structural crack determination value to obtain the concrete crack type. Specifically, this involves extracting each pair of data from the crack feature vector, i.e., the opening width value and the extension length value from that point to the starting point for each skeleton point. For each pair of data, the ratio of the opening width to the extension length is calculated. During the calculation, the units of measurement are kept consistent, using [missing information - likely a specific unit]. or The system calls a preset structural crack judgment value, which is calibrated through numerous concrete crack tests and set based on the characteristics of different types of concrete cracks. This value can be flexibly adjusted according to the type of concrete component, such as piers or beams. The preset structural crack judgment value is 0.001, a fixed threshold specifically used to distinguish between structural and non-structural cracks. Each calculated ratio is compared with the preset structural crack judgment value of 0.001. If a ratio is greater than 0.001, it indicates that the crack width at that location increases rapidly with the extension length, consistent with the characteristics of a structural crack. If the ratio is less than or equal to 0.001, it indicates that the crack width at that location increases gradually with the extension length, belonging to a non-structural crack. Based on the comparison results of all ratios, if more than 80% of the ratios are greater than 0.001, the concrete crack is determined to be a structural crack; if more than 80% of the ratios are less than or equal to 0.001, the concrete crack is determined to be a non-structural crack, thus completing the determination of the concrete crack type.
[0041] Based on the maximum value of the opening width in the quantitative vector set and the total extension length of the continuous crack skeleton sequence, the concrete crack grade assessment result is obtained by comparing it with the concrete crack grade judgment standard. The concrete crack type and concrete crack grade assessment result are output to complete the concrete crack analysis. Specifically, this includes: extracting all opening width values from the quantitative vector set and selecting the maximum value, which is the maximum opening width of the crack along the extension direction and is one of the core indicators for assessing the severity of the crack; calculating the total extension length of the continuous crack skeleton sequence, which is the total length of the crack. The specific calculation process is to use the spatial straight-line distance calculation method mentioned above, taking the starting point of the continuous crack skeleton sequence as the reference, and calculating the spatial straight-line distance between the starting point and the last skeleton point in the sequence. This distance is the total extension length of the continuous crack skeleton sequence. This total length value is used as an auxiliary indicator for crack grade assessment. The preset concrete crack grade judgment standard is called. This standard is formulated with reference to relevant industry specifications and determines the crack grades corresponding to different maximum crack widths and total extension lengths, such as four grades: minor, moderate, severe, and dangerous. Different grades correspond to different crack hazards and treatment requirements. The maximum opening width of the selected cracks and the calculated total extension length of the cracks are compared with the concrete crack grade assessment standard to determine the corresponding grade of the concrete crack. If both the maximum opening width and the total extension length fall within the assessment range of a certain grade, the crack is directly classified as that grade. If they fall within different grade ranges, the higher grade range is used as the final assessment result, thus obtaining the concrete crack grade evaluation result. Finally, the concrete crack type determined above and the concrete crack grade evaluation result obtained this time are output through a preset output interface. The output content clearly marks the crack type, maximum opening width, total length, and corresponding grade, and can also simultaneously output the core information of the crack feature vector, completing the entire concrete crack analysis.
[0042] In this embodiment of the invention, the extension length of each skeleton point from the starting point is calculated by a continuous crack skeleton sequence, overcoming the technical problem that analyzing width or length alone cannot reflect the morphological changes along the crack, and realizing the spatial correlation representation of crack opening width and extension length; by comparing the ratio of width value to extension length value in the crack feature vector with the preset structural crack judgment value, the technical problem of misjudging crack type caused by a single width threshold or experience judgment is overcome, realizing automatic and quantitative judgment of crack type based on actual morphological characteristics; by extracting the maximum opening width and the total extension length of the crack, the technical problem of insufficient maintenance decision-making basis due to the lack of graded quantitative indicators is overcome, and the scientific grading of crack hazard degree and the visualization output of diagnosis results are completed.
[0043] like Figure 2As shown, embodiments of the present invention also provide a concrete crack analysis system based on multimodal image fusion, comprising: The acquisition module is used to acquire visible light images and structured light depth images of the same concrete structure surface, perform spatial alignment on the visible light images and structured light depth images to obtain a bimodal aligned image set, extract the texture high-frequency tensor and surface normal gradient tensor of the bimodal aligned image set, and adaptively fuse the texture high-frequency tensor and surface normal gradient tensor to obtain a fused feature tensor. The correction module is used to locate the three-dimensional coordinates of three structural reference points based on the structured light depth image: the center point of the bottom casting joint of the pier, the side edge reference point of the transverse prestressed anchorage zone, and the intersection point of the bridge deck drainage channel facade. It constructs a triangular reference network with the three-dimensional coordinates as vertices and stretches it into a triangular prism along the surface normal direction. It calculates the local projection deviation of each pixel in the fused feature tensor onto the triangular reference network within the triangular prism and generates a compensation correction matrix. The fusion module is used to perform spatial remapping on the fusion feature tensor according to the compensation correction matrix to obtain the corrected fusion feature tensor. The corrected fusion feature tensor is then fitted to the sub-pixel edge along the crack direction and connected to the fracture gap to obtain a continuous crack skeleton sequence. The calculation module is used to fit a tangent plane to the local neighborhood point cloud at each point in the continuous crack skeleton sequence, measure the physical width from each point to the crack edge line on both sides along the normal of the tangent plane, and obtain the opening width quantization vector set. The analysis module is used to concatenate the continuous crack skeleton sequence with the opening width quantization vector set into a crack feature vector, perform mode determination on the crack feature vector, output the concrete crack type and grade assessment, and complete the concrete crack analysis.
[0044] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0045] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0046] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0047] Experimental example: To further verify the technical effectiveness of the crack detection method based on multimodal data fusion proposed in this application under complex scenarios with curved surfaces, a field experiment was conducted at the bottom of the crossbeam of the main concrete tower of a cross-river cable-stayed bridge with a radius of curvature of 6.2m. The experimental scenario selected was a representative curved concrete surface affected by solar radiation temperature difference and structural stress, and a high-resolution visible light camera and a structured light depth module were deployed.
[0048] Step 1: Dual-modal acquisition and feature fusion experiment: A 20-megapixel visible light camera with a 12mm focal length and a structured light depth module with an accuracy of ±0.1mm are integrated on the same inspection platform. Hardware synchronization of dual-modal data is achieved at a 1.2m acquisition distance, with a synchronization error of <8ms. Distortion correction is performed on the visible light image, and the radial distortion coefficient is... Histogram equalization was applied to improve the contrast ratio in low-contrast areas from 0.12 to 0.35. Simultaneously, the depth image underwent 3×3 median filtering to achieve 89% noise suppression and 99.6% invalid point filling. Based on this, the high-frequency texture tensor in the visible light domain and the surface normal gradient tensor in the depth domain were extracted, and texture weights were determined using adaptive weighting coefficients. Depth weights Pixel-by-pixel fusion is performed to generate a fusion feature tensor with dimensions of 1920×1080×10, which effectively solves the problem of blurred crack edges caused by seam artifact interference in a single visible light mode.
[0049] Step 2: Triangulation reference network construction and distortion correction experiment: Based on the precise location of three structural reference points using structured light depth images—the center point C (1250.3, 860.5, 320.2) of the cast-in-place joint at the bottom of the pier, the reference point P (1350.8, 920.7, 325.6) of the lateral edge of the transverse prestressed anchorage zone, and the intersection point D (1180.5, 780.2, 315.8) of the bridge deck drainage channel elevation—a triangular reference network is constructed using these three-dimensional coordinates as vertices and along the unit normal vector. =(0.12,0.08,0.99) stretched by 50mm to form a triangular prism spatial constraint; by calculating the local projection deviation of each pixel in the fused feature tensor within the triangular prism, a compensation correction matrix is generated. Experimentally, the average projection deviation is controlled at 0.08mm, and the maximum deviation does not exceed 0.15mm. Compared with the traditional planar projection model, which produces an error of 0.12~0.22mm on curved surfaces with an average of 0.18mm, this method compresses the measurement error to the range of -0.03~0.05mm with an average of 0.01mm, successfully eliminating the systematic projection compression distortion caused by the change in the angle between the optical axis and the surface normal. Figure 3The curve showing the comparison of crack width error before and after correction using the triangular reference network clearly demonstrates the excellent performance of this method, which significantly narrows the error fluctuation to within ±0.05mm within a crack extension length of 0~0.8m.
[0050] Step 3: Continuous crack skeleton extraction experiment: Using the corrected fused feature tensor, the Sobel operator and gradient thresholding are applied. =3.2, calculate the gray-level gradient of each pixel's neighborhood, and extract a set of crack edge pixels containing 1280 candidate points; for the edge fracture problem of curved surfaces and joints, a cubic spline function is used to fit piecewise along the crack direction to obtain 2560 sub-pixel level accurate edge points, and linear interpolation is performed to complete the fracture gap with a spacing >0.8mm, inserting 12 points with a spacing of 0.2mm; the generated continuous crack skeleton sequence has a total length of 0.82m, which is completely consistent with the measured length, and completely solves the problem of skeleton fracture and topological discontinuity caused by artifact interference when crossing construction joints in traditional methods.
[0051] Step 4: Quantitative experiment on opening width: Centered on each point in the continuous crack skeleton sequence, a local neighborhood point cloud within a spherical area with a radius of 3 mm is extracted. The tangent plane is fitted using the least squares method to obtain the surface normal vector. Subsequently, starting from the skeleton points, a scan is performed along the positive and negative directions of the normal vector with a step size of 0.1 pixels, based on the gray-level abrupt change threshold. The method accurately locates the edges of the cracks on both sides, with a single-point measurement time of only 0.8ms. The experimentally obtained quantized vector set of the crack width shows that the crack width smoothly transitions from 0.32mm to 0.57mm along the extension direction, with an average value of 0.45mm. Furthermore, no sudden drop in width, common in traditional methods, is observed at the construction joint, verifying the continuity and anti-interference capability of the width quantization. Figure 4 The diagram showing the trend of crack width variation along the extension direction quantitatively depicts a smooth, monotonous increasing trend in width, gradually increasing from 0.32 mm at the beginning to 0.57 mm at the end.
[0052] Step 5: Crack Type and Grade Assessment Experiment Based on the continuous crack skeleton sequence, the extension length of each point from the starting point was calculated, and a crack feature vector containing spatial location and width attributes was constructed. The average ratio of width to extension length was calculated to be 0.00069, which is far below the structural crack judgment threshold of 0.001. Based on this, the crack was judged to be a non-structural crack with a judgment accuracy of 100%. In conjunction with the highway bridge maintenance specifications, based on the maximum crack width of 0.57mm, which exceeds the 0.5mm threshold and the total length of 0.82m, it was rated as a medium crack and a re-inspection was recommended within 3 months. The consistency between this assessment result and the human expert assessment result reached 100%, effectively avoiding maintenance decision errors that might be caused by relying solely on a single width threshold.
[0053] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A concrete crack analysis method based on multimodal image fusion, characterized in that, The method includes: Visible light images and structured light depth images of the same concrete structure surface are acquired. Spatial alignment is performed on the visible light images and structured light depth images to obtain a bimodal aligned image set. The texture high-frequency tensor and surface normal gradient tensor of the bimodal aligned image set are extracted. The texture high-frequency tensor and surface normal gradient tensor are adaptively fused to obtain a fused feature tensor. Based on the structured light depth image, the three-dimensional coordinates of three structural reference points are located: the center point of the bottom casting joint of the pier column, the reference point of the side edge of the transverse prestressed anchorage zone, and the intersection point of the bridge deck drainage channel facade. A triangular reference network is constructed with the three-dimensional coordinates as vertices and stretched into a triangular prism along the surface normal direction. The local projection deviation of each pixel in the fused feature tensor onto the triangular reference network within the triangular prism is calculated, and a compensation correction matrix is generated. Based on the compensation and correction matrix, spatial remapping is performed on the fused feature tensor to obtain the corrected fused feature tensor. The corrected fused feature tensor is then fitted to the sub-pixel edge along the crack direction and connected to the fracture gap to obtain a continuous crack skeleton sequence. Based on the local neighborhood point cloud fitting tangent plane at each point in the continuous crack skeleton sequence, the physical width from each point to the crack edge line on both sides is measured along the normal of the tangent plane to obtain the opening width quantization vector set; The continuous crack skeleton sequence and the opening width quantization vector set are concatenated into a crack feature vector. The crack feature vector is then used to determine the mode, output the concrete crack type and grade assessment, and complete the concrete crack analysis.
2. The concrete crack analysis method based on multimodal image fusion according to claim 1, characterized in that, Visible light images and structured light depth images of the same concrete structure surface are acquired. Spatial alignment is performed on the visible light images and structured light depth images to obtain a bimodal aligned image set, including: By using a visible light camera and a structured light depth module mounted on the same acquisition platform, visible light images and structured light depth images of the concrete structure surface under the same field of view are acquired simultaneously to obtain the original dual-modal image pair. Lens distortion correction and histogram equalization are performed on the visible light image in the original bimodal image pair to obtain a visible light preprocessed image. Median depth filtering and invalid depth point neighborhood interpolation filling are performed on the structured light depth image in the original bimodal image pair to obtain a depth preprocessed image. Local feature descriptors are extracted between the visible light preprocessed image and the depth preprocessed image. Feature point matching is performed based on the local feature descriptors. Rigid body transformation parameters between the visible light preprocessed image and the depth preprocessed image are calculated. Based on the rigid body transformation parameters, the visible light preprocessed image and the depth preprocessed image are mapped to the same pixel coordinate system to obtain a dual-modal aligned image set.
3. The concrete crack analysis method based on multimodal image fusion according to claim 1, characterized in that, Extract the texture high-frequency tensor and surface normal gradient tensor from the bimodal aligned image set, and adaptively fuse the texture high-frequency tensor and surface normal gradient tensor to obtain the fused feature tensor, including: A multi-layer Gaussian image sequence is constructed based on the visible light images in the dual-modal aligned image set. Inter-layer difference operations are performed on each layer of the Gaussian image sequence to obtain the texture high-frequency tensor. The surface normal vector is calculated point by point based on the structured light depth image in the dual-modal aligned image set. The spatial gradient derivative of the surface normal vector is obtained to obtain the surface normal gradient tensor. Based on the grayscale difference amplitude of each pixel neighborhood in the texture high-frequency tensor and the depth change amplitude at the corresponding position in the surface normal gradient tensor, the weighting coefficients of the texture high-frequency tensor and the surface normal gradient tensor are calculated pixel by pixel. The texture high-frequency tensor and the surface normal gradient tensor are then weighted and fused pixel by pixel using the weighting coefficients to obtain the fused feature tensor.
4. The concrete crack analysis method based on multimodal image fusion according to claim 3, characterized in that, Based on the structured light depth image, the three-dimensional coordinates of three structural reference points are determined: the center point of the cast-in-place joint at the bottom of the pier, the reference point on the side edge of the transverse prestressed anchorage zone, and the intersection point of the bridge deck drainage channel elevation. A triangular reference network is constructed using these three-dimensional coordinates as vertices and stretched into a triangular prism along the surface normal direction. The local projection deviation of each pixel in the fused feature tensor onto the triangular reference network within the triangular prism is calculated, generating a compensation and correction matrix, including: Based on the two joint edge lines at the bottom of the pier in the structured light depth image, the spatial intersection of the two joint edge lines is solved to obtain the three-dimensional coordinates of the center point of the bottom of the pier. Based on the depth abrupt change zone at the side edge of the transverse prestressed anchorage zone in the structured light depth image, the extreme points of the depth gradient are extracted along the direction of the depth abrupt change zone to obtain the three-dimensional coordinates of the reference point at the side edge of the transverse prestressed anchorage zone. Based on the intersection line between the bridge deck drainage channel elevation and the pier surface in the structured light depth image, the lowest point of the intersection line is extracted to obtain the three-dimensional coordinates of the intersection point of the bridge deck drainage channel elevation. Connect the three-dimensional coordinates of the center point of the bottom casting joint of the pier column, the reference point of the side edge of the transverse prestressed anchorage zone, and the intersection point of the bridge deck drainage channel elevation to form a spatial triangle to obtain a triangular reference network. Extract the unit normal vector of the plane where the triangular reference network is located, and stretch the triangular reference network to a preset height along the direction of the unit normal vector to obtain a triangular prism. The two-dimensional coordinates of each pixel in the fused feature tensor are mapped onto the body of a triangular prism. The coordinates of the vertical projection point of each pixel along the height direction of the triangular prism on the triangular reference network are calculated. Based on the coordinates of the vertical projection point and the original spatial position of the corresponding pixel in the fused feature tensor, the local projection deviation of each pixel is calculated point by point. The local projection deviations of all pixels are combined to obtain the compensation correction matrix.
5. The concrete crack analysis method based on multimodal image fusion according to claim 4, characterized in that, Spatial remapping is performed on the fused feature tensor based on the compensation and correction matrix to obtain the corrected fused feature tensor. The corrected fused feature tensor is then fitted to the sub-pixel edges along the crack direction and connected to the fracture gaps to obtain a continuous crack skeleton sequence, including: The local projection deviation value of each pixel in the compensation correction matrix is superimposed pixel by pixel onto the original spatial coordinates of the corresponding pixel in the fused feature tensor to obtain the corrected fused feature tensor. Calculate the magnitude and direction of the gray-level gradient in the neighborhood of each pixel in the fused feature tensor after correction, and select the pixel point with the largest gray-level gradient magnitude that exceeds the preset gradient threshold to obtain the crack edge pixel point set. By performing cubic spline function fitting segment by segment along the direction of the crack edge pixel set, the precise sub-pixel position of each edge pixel between adjacent pixels is determined, thus obtaining the sub-pixel edge point set; Based on the sub-pixel edge point set, the spatial distance between two adjacent sub-pixel edge points is calculated, and adjacent edge point pairs whose spatial distance exceeds a preset distance threshold are identified to obtain a continuous crack skeleton sequence.
6. The concrete crack analysis method based on multimodal image fusion according to claim 5, characterized in that, Based on the local neighborhood point cloud of each point in the continuous crack skeleton sequence, a tangent plane is fitted, and the physical width from each point to the crack edge lines on both sides is measured along the normal of the tangent plane to obtain the quantized vector set of the opening width, including: Based on the spatial coordinates of each skeleton point in the continuous crack skeleton sequence, depth points within a sphere with a preset radius centered on each skeleton point are extracted from the structured light depth image to obtain the local neighborhood point cloud of each skeleton point. For each skeleton point, perform plane fitting on the local neighborhood point cloud, extract the normal vector of the fitted plane as the surface normal vector of the corresponding skeleton point, and obtain the surface normal vector corresponding to each skeleton point. Starting from the spatial coordinates of each skeleton point, the scanning and acquisition are performed pixel by pixel along the direction of the corresponding surface normal vector. After acquisition and correction, the gray values at each advancing position in the feature tensor are fused. The pixel positions where the gray values change abruptly are located as crack edge points, thus obtaining the left and right crack edge points corresponding to each skeleton point. Calculate the pixel distance from each skeleton point to the corresponding left crack edge point and the pixel distance to the corresponding right crack edge point. Add the pixel distances on both sides to get the pixel span. Convert the pixel span to the physical width to get the opening width value corresponding to each skeleton point. The opening width value corresponding to each skeleton point is organized along the extension order of the continuous crack skeleton sequence to obtain the opening width quantization vector set.
7. The concrete crack analysis method based on multimodal image fusion according to claim 6, characterized in that, The continuous crack skeleton sequence and the quantized vector set of opening width are concatenated to form a crack feature vector. A pattern determination is performed on the crack feature vector, and the concrete crack type and grade assessment are output to complete the concrete crack analysis, including: The extension length of each skeleton point from the starting point is calculated based on the continuous crack skeleton sequence to obtain the skeleton extension length sequence. The skeleton extension length sequence is then paired with the opening width quantization vector set point by point according to the skeleton point position to obtain the crack feature vector. Based on the opening width value and skeleton extension length value in the crack feature vector, calculate the ratio of the opening width value to the skeleton extension length value, compare the ratio with the preset structural crack judgment value, and obtain the concrete crack type. Based on the maximum value of the opening width in the quantitative vector set and the total extension length of the continuous crack skeleton sequence, the concrete crack grade assessment result is obtained by comparing with the concrete crack grade judgment standard; the concrete crack type and concrete crack grade assessment result are output to complete the concrete crack analysis.
8. A concrete crack analysis system based on multimodal image fusion, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The acquisition module is used to acquire visible light images and structured light depth images of the same concrete structure surface, perform spatial alignment on the visible light images and structured light depth images to obtain a bimodal aligned image set, extract the texture high-frequency tensor and surface normal gradient tensor of the bimodal aligned image set, and adaptively fuse the texture high-frequency tensor and surface normal gradient tensor to obtain a fused feature tensor. The correction module is used to locate the three-dimensional coordinates of three structural reference points based on the structured light depth image: the center point of the bottom casting joint of the pier, the side edge reference point of the transverse prestressed anchorage zone, and the intersection point of the bridge deck drainage channel facade. It constructs a triangular reference network with the three-dimensional coordinates as vertices and stretches it into a triangular prism along the surface normal direction. It calculates the local projection deviation of each pixel in the fused feature tensor onto the triangular reference network within the triangular prism and generates a compensation correction matrix. The fusion module is used to perform spatial remapping on the fusion feature tensor according to the compensation correction matrix to obtain the corrected fusion feature tensor. The corrected fusion feature tensor is then fitted to the sub-pixel edge along the crack direction and connected to the fracture gap to obtain a continuous crack skeleton sequence. The calculation module is used to fit a tangent plane to the local neighborhood point cloud at each point in the continuous crack skeleton sequence, measure the physical width from each point to the crack edge line on both sides along the normal of the tangent plane, and obtain the opening width quantization vector set. The analysis module is used to concatenate the continuous crack skeleton sequence with the opening width quantization vector set into a crack feature vector, perform mode determination on the crack feature vector, output the concrete crack type and grade assessment, and complete the concrete crack analysis.
9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.