Adaptive sensing-based lightweight monitoring method for fine crack in complex background region

Through the PTZ camera and lightweight crack segmentation network, combined with multi-scale template matching and European-style distance similarity classification, the efficient monitoring of fine cracks in complex background areas inside the bridge is solved, and automated, precise and dynamic tracking of cracks inside the bridge is achieved, and detection efficiency and accuracy are improved.

WO2025161130A1PCT designated stage Publication Date: 2025-08-07SOUTHEAST UNIV

Patent Information

Application Number
PCT/CN2024/087869
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-31
Filing Date
2024-04-16
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently monitor fine cracks in complex background areas inside bridges, especially large-area cracks inside box girders, and traditional methods cannot capture the evolution of cracks, resulting in low detection efficiency and insufficient accuracy.

Method used

The PTZ camera sensor is used for area division and acquisition, combined with multi-scale template matching and lightweight crack segmentation network, crack tracking is carried out through European-style distance similarity classification, and a lightweight crack monitoring method is built to realize automated, precise monitoring and dynamic tracking of cracks inside bridges.

Benefits of technology

Without reducing the recognition accuracy, the scope of the identification area is expanded, the scale distortion problem in complex backgrounds is overcome, efficient and accurate monitoring and dynamic analysis of internal cracks of the bridge are achieved, and the computing resource requirements are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024087869_07082025_PF_FP_ABST
    Figure CN2024087869_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to an adaptive sensing-based lightweight monitoring method for a fine crack in a complex background region. The method comprises the following steps: step S1, on the basis of region division, performing automatic acquisition of crack information, wherein PTZ camera sensors are used to automatically perform block-wise acquisition on crack regions; step S2, performing an adaptive complex scale calibration process, using a multi-scale template matching algorithm to adaptively correct distortion information of all regions, and performing real-scale conversion from pixel precision; step S3, constructing a lightweight crack segmentation network to process data processed in step S2; and step S4, by means of a quantitative crack-tracking algorithm based on Euclidean distance similarity classification, performing real-time monitoring on each piece of crack dynamic information. Compared with the prior art, the present invention has advantages such as achieving efficient, accurate, and online monitoring and analysis of cracks.
Need to check novelty before this filing date? Find Prior Art

Description

Lightweight monitoring method for fine cracks in complex background areas based on adaptive perception Technical Field

[0001] The present invention relates to the field of structural health detection in civil engineering, and in particular to a lightweight monitoring method for fine cracks in complex background areas based on adaptive perception. Background Art

[0002] Bridge structures play a critical role in transportation systems, making their long-term safe operation crucial. These structures are subject to a complex mix of factors during their daily operation, including direct effects such as the structure's own weight and vehicle loads, as well as indirect effects such as temperature and humidity fluctuations, concrete shrinkage, and creep. Bridge structures are typically large and complex, and as a result, long-term damage can be widespread. This can sometimes lead to difficult-to-reach or overlooked locations, such as the base of the box girder and piers on the exterior of the bridge, and the web and top plates of the box girder on the interior. Cracks are one of the most common types of damage in these locations, and these cracks often develop continuously, eroding the structure and compromising the safety and durability of the bridge. Therefore, crack detection is a key component of long-term bridge maintenance. Utilizing inspection data, engineers can promptly formulate appropriate maintenance measures to maximize the lifespan of the bridge and ensure safe and smooth transportation.

[0003] Traditional manual crack detection methods suffer from low efficiency and a difficulty covering many difficult-to-detect or missed areas. To address these issues, researchers have begun introducing intelligent equipment into bridge engineering to enable rapid and comprehensive inspection of difficult-to-detect or under-covered areas of bridge structures. These equipment, such as rope-climbing robots, unmanned aerial vehicles (UAVs) with flexible flight-climbing mode switching, and mobile inspection robots, have significantly benefited field engineering. However, when dealing with large and complex crack data, engineers still require significant time and human resources for in-house analysis, which can easily lead to oversights. Manual analysis can also be subject to subjective factors, resulting in varying results between operators. With the advancement of computer vision technology, researchers have proposed a series of crack detection methods based on digital image processing to address these issues. However, these methods face significant challenges, including sensitivity to noise, high parameter dependence, and difficulty handling complex backgrounds. These limitations significantly impact their practicality and robustness in engineering applications. In recent years, deep learning technology has spurred cutting-edge technological innovations across a wide range of disciplines. Vision-based deep learning techniques primarily encompass image-level classification networks, object-level targeting networks, and pixel-level segmentation networks. However, due to the small proportion of cracks in images, classification and targeting networks are not suitable for accurate crack analysis. Typically, further operations are performed on the results of these networks. Therefore, pixel-level segmentation networks are considered the most suitable tool for crack analysis. These networks can identify and locate cracks in images with high accuracy and resolution, providing a suitable solution for crack detection and analysis.

[0004] Intelligent detection equipment has made significant progress in data collection efficiency. However, these devices are generally limited to collecting data on external bridge defects and can only provide defect data at a specific moment in time. Cracks are a type of defect with a growing property, and conventional detection methods cannot capture their evolutionary process, making it impossible to intervene at the optimal maintenance time. At the same time, current crack detection algorithms often rely on complex model design and attention mechanisms to improve detection accuracy, but this limits their feasibility in practical engineering applications with limited computing resources. In addition, existing crack quantitative analysis algorithms focus primarily on extracting precise crack parameters, with less attention paid to accurately tracking the crack evolution process.

[0005] Therefore, the task of monitoring crack areas inside bridge box girders has become a key problem that engineers are currently concerned about. The engineering community urgently needs new long-term intelligent monitoring equipment suitable for the interior of bridges to monitor these areas that are difficult to monitor manually or with other intelligent equipment.

[0006] After searching Chinese patent publication number CN114913162A, a bridge concrete crack detection method and device based on a lightweight Transformer is disclosed. Specifically, the method and device use a drone to collect bridge concrete crack images and transmit the collected data back; the images are preprocessed and converted into a fixed input format to establish a data set; a lightweight Transformer network is constructed, and the images are trained and tested to obtain a recognition model. The lightweight Transformer adopts a two-stage patch transformation method, which discards part of the Transformer input sequence during training to reduce overfitting problems and enhance the diversity between patches; the recognition model is deployed to carry out recognition of the bridge surface image to be identified.

[0007] However, the existing patent is limited to the collection of external bridge defects, and is implemented by using human-machine image collection. This method still has problems such as low recognition accuracy.

[0008] Summary of the Invention

[0009] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a lightweight monitoring method for fine cracks in complex background areas based on adaptive perception, aiming to fully combine intelligent detection equipment with efficient analysis algorithms to improve the monitoring efficiency of bridge structure cracks.

[0010] The purpose of the present invention can be achieved by the following technical solutions:

[0011] According to one aspect of the present invention, a method for lightweight monitoring of fine cracks in complex background areas based on adaptive perception is provided, the method comprising the following steps:

[0012] Step S1, automatic collection of crack information based on regional division, wherein a PTZ camera sensor is used to automatically collect crack areas in blocks, overcoming the inherent contradiction between the field of view and recognition accuracy of crack collection under the fixed focal length of traditional cameras;

[0013] Step S2: Perform an adaptive complex scale calibration process. Since the PTZ sensor's acquisition mechanism causes each area of ​​the image to have varying degrees of tilt distortion, a multi-scale template matching algorithm is used to adaptively correct the distortion information of all areas and perform true scale conversion with pixel accuracy.

[0014] Step S3: constructing a lightweight crack segmentation network to process the data processed in step S2. The lightweight crack segmentation network adopts a compact structure and lightweight module construction method to significantly reduce the complexity of the model while maintaining excellent crack identification performance;

[0015] In step S4, a quantitative crack tracking algorithm based on Euclidean distance similarity classification is used to monitor the dynamic information of each crack with a large number and complex development trend in real time.

[0016] As a preferred technical solution, the automatic collection of crack information based on regional division in step S1 is specifically as follows:

[0017] Traditional vision sensors improve crack recognition accuracy by reducing the acquisition field of view. However, if cracks are widely distributed, a single sensor's field of view may not cover the entire crack. To overcome the inherent contradiction between the acquisition field of view and crack recognition accuracy, a PTZ acquisition sensor with integrated control, acquisition, and transmission functions was introduced.

[0018] Through area division and path pre-planning, and selecting identifiable auxiliary markers as target areas, PTZ acquisition sensors are used to conduct targeted acquisition of crack areas, expanding the original recognition area without reducing recognition accuracy.

[0019] As a preferred technical solution, the adaptive complex scale calibration process in step S2 specifically includes:

[0020] Step S21, a multi-scale template matching process based on color features, thereby achieving automatic and stable extraction of the target area;

[0021] Step S22: Adaptive image calibration process, thereby achieving orthorectification of complex scales.

[0022] As a preferred technical solution, the multi-scale template matching process based on color features in step S21 specifically includes:

[0023] The Hue-Saturation-Value (HSV) color space, which is consistent with the human eye's color perception, is used as the color representation method;

[0024] First, the template image is discretized and histogram estimated to create a descriptive representation of the features; then the continuous HSV space is mapped to the discrete histogram space; finally, at each position, the template model is compared with the corresponding part of the target image, and the HSV probability value is calculated. The area with high matching degree is output as the matching result;

[0025] Template images with sizes 0.25 and 0.75 times the original template are added to extract features from the target image. A weighted fusion strategy is used to combine matching results from different scales into a comprehensive result. To maintain fairness, the fusion weight coefficient of templates at each scale is set to 1 / 3 to clearly indicate that the matching results at each scale have equal importance.

[0026] The adaptive image calibration process in step S22 specifically includes:

[0027] First, the inner circle area is extracted from the target area. The center point is determined by the connection line between the four equally divided points on the inner circle and used as the matching feature point. The other three matching feature points in the complete image are adaptively extracted according to the same process. Then, a homography mapping relationship is established by combining the coordinate information of each feature point in real space, which is expressed as:

[0028] Where (u t ,v t ) represents the image coordinates of the four matching feature points in the source image, (u r ,v r ) represents the plane coordinates of the four matching feature points in the target image in real space, t is the scale factor, represents the homography matrix;

[0029] Then, the least squares method is used to estimate the homography matrix in the above equation. The core goal of this method is to achieve this by minimizing the reprojection error, that is, mapping the points on the source image to the target image through the estimated homography matrix, and then comparing it with the actual corresponding pixel points to find the optimal homography matrix parameters. Once this optimal homography matrix is ​​obtained, the tilted and distorted source images under various acquisition angles can be quickly corrected to obtain a corrected image with real size. The specific calculation is as follows:

[0030] As a preferred technical solution, the lightweight crack segmentation network in step S3 includes:

[0031] Bilateral downsampling and upsampling blocks are used for dimensionality reduction extraction and dimensionality increase learning of crack features.

[0032] Compressed sensing and attention fusion module, used to improve the model's training and inference speed and recognition accuracy;

[0033] The output module is used to output the prediction results of crack segmentation detection.

[0034] As a preferred technical solution, the bilateral downsampling and upsampling block specifically includes a downsampling block and an upsampling block:

[0035] The downsampling block uses a downsampling ratio of less than 32 times (3 bilateral downsampling blocks) to reduce the scale of the feature map to better preserve and extract key feature details. The downsampling block adopts a dual-path parallel structure, namely a convolutional downsampling branch and a pooling downsampling branch with encoding;

[0036] The front end and back end of the convolutional downsampling branch are respectively equipped with a 1×1 mapping convolution (M Conv). The middle part contains batch normalization (BN) for standardizing the distribution of channel features, a 2×2 depth convolution (D Conv), and a ReLU activation function that introduces nonlinear properties. D Conv is used instead of the standard convolution structure. This choice helps to reduce the number of model parameters and computational complexity while maintaining the performance of the model.

[0037] The convolution downsampling branch CDB is expressed as the following formula:

[0038] in represents the input feature map, For BN processing, is the ReLU activation function, where DConv is depthwise convolution, MConv is mapping convolution, N, C, H, and W represent the batch, number of channels, height, and width of the input feature map, respectively;

[0039] The pooling downsampling branch PDB uses a pooling layer with a pooling kernel size of 2×2 to downsample the feature map, and records the position of the maximum value for feature restoration in the subsequent upsampling process;

[0040] We adopt a strategy of fusing CDB and PDB and downsampling the input features to effectively protect important details while maintaining lightweight. The specific technical process is as follows:

[0041] Among them BDB i is a bilateral downsampling module, represents the input features of the CDB branch of the i-th BDB module, represents the input features of the PDB branch of the i-th BDB module, represents the connection along the channel dimension, K i represents the index encoding output of the i-th BDB module, C cu with C pu It is a parameter obtained by processing the feature map channel input to the j-th BUB module through the allocation coefficient λ. The specific value of λ will be described in the subsequent network structure design table.

[0042] As a preferred technical solution, the compressed sensing and attention fusion module includes a compressed sensing unit and an attention fusion unit;

[0043] The compressed sensing unit includes two key stages, namely the asymmetric convolution stage and the asymmetric dilated convolution stage;

[0044] In the asymmetric convolution stage, the input features are first Compression processing is performed by introducing the compression coefficient α to divide X into the features to be convolved along the direction of the feature channel With compression features Then a set of 3×1 and 1×3 asymmetric convolutions are used to extract X c The characteristic information of X s The shortcut path designed will be used at the output stage with X' c Merge in the channel direction to obtain the features of stage one

[0045] In the asymmetric dilated convolution stage, the same feature compression operation as the asymmetric convolution stage is first performed, and the compression coefficient α is used to compress the input features. Divide and obtain the features to be convolved With compression features Then use a set of (3+d i,z )×1 and 1×(3+d i,z ) combined with asymmetric dilated convolution to X' c Perform feature extraction, where d i,z represents the void rate of the zth compressed sensing unit in the i-th compressed sensing module; at the same time, X' s The shortcut path of the design will be in the output stage with X" c Merge in the channel direction to obtain the features of stage 2

[0046] The overall design of the perception unit draws on the structure of the residual network. In the final output stage, the goals of perception and feature enhancement are achieved through pixel-level feature fusion.

[0047] The compressed sensing unit is CPM i The basic components of CPM1 are as follows, where i = 1, 2; CPM1 consists of 4 CS units, and the dilation rate of each CS unit is set to 0, which is used to capture the spatial characteristics of cracks; CPM2 consists of 8 CS units, and the dilation rate of each CS unit follows the following value rule:

[0048] where d i,z represents the void ratio of the compressed sensing unit, and z is the adjustment coefficient for controlling the size of the void ratio;

[0049] According to the above value selection rules, CPM2 can be regarded as a complex consisting of two multi-scale pyramid structures connected in series.

[0050] The attention fusion unit introduces a lightweight channel attention mechanism based on the compressed sensing unit structure. The specific process is as follows:

[0051] First, along the input features Perform average pooling operation on the length and width direction of

[0052] The input features is global average pooling;

[0053] Then take two band matrices and Continuous Learning Channel attention in

[0054] where k s1 ,k s2 Indicates the channel range covered by each row in the band matrix, d s1 ,d s2 represents the cross-channel interaction rate within the coverage range of each row of channels in the band matrix, and k in the text s1 ,k s2 Both are set to 3, d s1 ,d s2 Set to 1, 3, w respectively c,c Represents the band matrix coefficient of the c-th row and c-th column, where c is the number of channels of the input feature;

[0055] Use a one-dimensional convolution kernel with a dilation rate right Perform channel attention weight The specific calculation process of the extraction and attention module is as follows:

[0056] in Represents element-by-element multiplication, σ represents the Sigmoid function, For BN processing, is the output feature, is the input feature, AFM i It is constructed by an AFU and two CPUs and is located in the decoding stage of the network structure, where i = 1, 2. These two modules each fuse the different scale feature outputs from CPM1 and BDB1.

[0057] As a preferred technical solution, the output module includes three components: a main output and two auxiliary outputs;

[0058] The main output part consists of a transposed convolution with a kernel size of 2×2; the two auxiliary output modules consist of a 1×1 mapping convolution and a transposed convolution with a kernel size of 2×2. The network output Supervised by true labels It should be emphasized that the auxiliary output module is only enabled during the training phase and does not add additional computational burden during the inference phase.

[0059] As a preferred technical solution, step S4 specifically includes:

[0060] Step S41: Establishing a crack classification criterion to clarify the crack change conditions involved in the crack evolution process, specifically:

[0061] Cracks are divided into three categories: new cracks, extended cracks, and merged cracks. New cracks refer to cracks that did not exist in a specific area at time point t and appeared at time point t+1, indicating a new damage event. Extended cracks refer to cracks that existed at time point t and further expanded along the original development path at time point t+1, indicating an escalation of the damage event. Merged cracks refer to two or more cracks that existed at time point t and merged into one crack at time point t+1, indicating regional expansion of damage.

[0062] Step S42, crack skeletonization, reduces the complexity of crack feature processing, specifically:

[0063] Simplify the fracture object into a more manageable central skeleton structure; the skeletonization process not only significantly reduces the computational burden but also helps to more accurately identify the topological features of the fracture, thereby highlighting the morphological characteristics of the fracture development;

[0064] Step S43: Establishing a dynamic analysis criterion of Euclidean distance similarity, further reducing the complex crack morphology information into a simpler similarity metric, and simplifying the crack analysis process, specifically:

[0065] Euclidean distance similarity The specific calculation formula for crack analysis is as follows:

[0066] in represents the i-th coordinate point on the center line of the m-th crack at time t, represents the i-th coordinate point on the centerline of the n-th crack at time t+1, i∈[0,N], N represents the length of the crack centerline;

[0067] This formula mainly uses the inverse distance transformation to convert distance into similarity. The characteristic of this transformation is that when the distance is smaller, the similarity value is closer to 1, indicating high similarity, and when the distance is larger, the similarity value is closer to 0, indicating low similarity. m =N n When N m ≠N n , we need to further consider how to comprehensively calculate the similarity between cracks of different lengths. Therefore, this method is based on the idea of ​​convolution operation, and here we assume that N m >N n , the shorter crack centerline coordinate set is regarded as 1×N n The convolution kernel of dimension 1×N is used, and the coordinate set of the crack centerline with a longer length is regarded as 1×N m The convolution target of dimension is set. Set the moving step of the pseudo convolution kernel to 1, and replace the convolution calculation with the similarity calculation, and finally get N m -N n +1 similarity index, and select the maximum value as the final similarity result. This calculation method is also applicable to N m <N n The crack similarity solution process.

[0068] The similarity metric effectively establishes correlations between crack images at different times, thereby forming a sequence of corresponding crack matching combinations. This establishes a one-to-one correspondence between newly emerged cracks and previously observed cracks. It's important to note that newly emerged cracks constitute a special case in this matching sequence because they lack similarity to any cracks from the previous time point. In other words, the crack matching combination contains only one crack. Therefore, the similarity matching results enable rapid extraction of newly emerged cracks. The crack matching combination also includes extended cracks and merged cracks. Based on their evolutionary characteristics, these two types of cracks can be easily separated and extracted from the crack matching combination. Merged cracks can only be extracted if they meet the following criteria: At least two cracks with high similarity matching exist at the previous time point, compared to the merged crack observed at the current time point. This means that the number of identical new crack matching combinations is greater than two. This principle ensures reliable detection of merged cracks. On the other hand, the classification criteria for extended cracks differ significantly from those for merged cracks. For extended cracks, only one crack with high similarity matching exists at the current time point. Therefore, extended cracks can be extracted based on the principle that the number of matching combinations for the new crack is equal to one.

[0069] The similarity-based crack classification method assigns a unique category attribute to each crack in an image, allowing accurate tracking of all cracks in the image. To comprehensively describe the development characteristics of cracks, six characteristic attributes, namely length, area, width, length change, area change, and width change, are introduced on the basis of the category attribute. The length of a crack usually refers to the length of the crack centerline, which can be obtained by measuring the length of the crack skeleton. The area of ​​a crack represents the pixel area occupied by the crack in the image, which can be obtained by performing regional statistics on the crack segmentation image. Based on the obtained crack length and area, the average width of the crack can be further calculated. Finally, the changing trends of length, area, and width over time are analyzed to gain a deeper understanding of the dynamic evolution of cracks.

[0070] As a preferred technical solution, the method further includes constructing a crack monitoring sensor network architecture, as follows:

[0071] Based on a dynamic perception framework, a comprehensive crack monitoring sensor network was constructed. This network consists of data acquisition nodes, data processing nodes, a data transmission layer, and an information aggregation layer. Its key features include modularity, high efficiency, security, and reliability. In terms of modularity, each data acquisition node consists of two crack monitoring sensors, forming a modular collection process. The data acquisition process is isolated between nodes, ensuring the security of the data collection process. To improve the efficiency of the data processing process, the data collected by each acquisition node will be processed by the corresponding data processing node (edge ​​computer). The data transmission layer, namely the server located at the construction site, will be responsible for establishing a reliable communication medium between the enclosed space inside the bridge and the outside world, transmitting data from each data processing node to a cloud service platform. As the information aggregation layer, the cloud service platform will comprehensively analyze and evaluate all data, providing engineers with intuitive and reliable support to support the implementation of operation and maintenance strategies.

[0072] Compared with the prior art, the present invention has the following advantages:

[0073] 1) Current intelligent detection equipment is usually limited to the collection of external bridge defects and can only provide defect data at a certain moment. Cracks are a type of defect with growth properties, and conventional detection methods cannot capture its evolution process, so it is impossible to intervene at the best maintenance time. At the same time, traditional visual sensors improve crack recognition accuracy by reducing the size of the acquisition field of view. However, if the crack object is widely distributed, the field of view of a single sensor cannot cover the entire crack object. In order to overcome the inherent contradiction between the crack acquisition field of view and recognition accuracy, the present invention introduces a PTZ acquisition sensor with integrated control, acquisition and transmission functions, which can overcome the challenges of collecting large-area crack areas in the internal space of the bridge, such as box girders. The PTZ acquisition sensor expands the original recognition area range without reducing the recognition accuracy through area division and path pre-planning methods.

[0074] 2) The complex multi-angle and multi-focal length combination acquisition mechanism introduces different degrees of scale distortion. When the central axis of the lens is pitched around x in the zy plane, different pitch angles will cause the image to have different degrees of directional tilt distortion. Similarly, when the central axis of the lens is yawed around y in the zx plane, different yaw angles will cause the image to have different degrees of directional tilt distortion. In addition, when the central axis of the lens is rotated around the origin of the coordinate system in the xyz space for shooting, the image will have different degrees of uv direction tilt distortion as the rotation angle changes. In order to solve the above-mentioned complex tilt distortion problem, the present invention proposes a simple complex scale calibration method based on targeted perception, which consists of two steps: automatic extraction of target area and adaptive scale calibration. It can extract the features of the target area online and adaptively correct the image based on the extracted features.

[0075] 3) Current crack detection algorithms usually rely on complex model design and attention mechanisms to improve detection accuracy, but this limits their feasibility in practical engineering applications with limited computing resources. Therefore, unlike previous complex network designs, the present invention focuses on building a lightweight network model suitable for large-scale images and proposes a simple, lightweight, high-performance segmentation network, aiming to get rid of the dependence on high-performance hardware to achieve fast and accurate segmentation of crack images. By cleverly conceiving the network architecture and carefully designing lightweight modules, the redundant information in the feature map is successfully reduced, which has a significant impact on the efficiency of the network. At the same time, the multi-level scale feature information generated in the downsampling and encoding convolution process is more effectively utilized, which not only helps to capture the shallow detail information of the image, but also effectively integrates this information with deep semantic information, thereby improving the accuracy and robustness of the segmentation task and enabling the model to understand the image content more comprehensively.

[0076] 4) The primary goal of the present invention is to carry out crack monitoring, focusing on the evolution tracking of dynamic cracks. The evolution process of cracks involves changeable and complex situations. Therefore, the present invention divides cracks into three categories, namely new cracks, extended cracks and merged cracks, in order to more accurately describe their characteristics. According to a clearly defined classification system, the present invention proposes a crack classification algorithm based on similarity analysis, which aims to achieve high-precision control of all cracks in a complex evolutionary environment. The similarity-based crack classification method gives each crack in an image a unique category attribute, so that all cracks in the image can be accurately tracked. In order to comprehensively describe the development characteristics of cracks, six characteristic attributes of length, area, width, length change, area change, and width change are introduced on the basis of the category attributes. This method is no longer limited to the simple statistics of the overall crack index, but can capture the development details of each crack in a highly accurate manner, thereby providing a more specific and intuitive evolutionary representation of the crack area. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] FIG1 is a schematic diagram of the overall framework of the present invention;

[0078] FIG2 is a schematic diagram of input, matching, fusion and extraction of the adaptive complex scale calibration proposed in the present invention;

[0079] FIG3 is a schematic diagram of a conversion matrix for adaptive complex scale calibration according to the present invention;

[0080] FIG4 is a diagram of a lightweight crack network architecture proposed in the present invention;

[0081] FIG5 is a schematic diagram of a dynamic crack tracking algorithm based on similarity classification proposed in the present invention;

[0082] FIG6 is a diagram of a network scheme of a crack monitoring sensor network architecture proposed in the present invention;

[0083] FIG7 is a detailed diagram of the bilateral downsampling and upsampling module proposed in the present invention;

[0084] FIG8 is a schematic diagram of the compressed sensing and attention fusion module proposed in the present invention;

[0085] Figure 9 shows the monitoring bridge structure and equipment layout for the example;

[0086] FIG10 is a schematic diagram of the implementation process of the crack sensor proposed in the example;

[0087] FIG11 is a schematic diagram of the second implementation process of the crack sensor proposed in the example;

[0088] FIG12 is a schematic diagram of a calibration image of all sub-regions mentioned in the example;

[0089] Figure 13 is a schematic diagram of the second calibration image of all sub-regions mentioned in the example;

[0090] Figure 14 is a schematic diagram of the crack monitoring results in the area mentioned in the example;

[0091] FIG15 is a comparison table of the results of the algorithm framework proposed in the present invention and other advanced segmentation algorithms. DETAILED DESCRIPTION

[0092] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0093] The present invention is based on a lightweight monitoring method for fine cracks in complex background areas with adaptive perception. First, in order to overcome the challenge of collecting data in large crack areas in the interior space of a bridge, such as in a box girder, a set of PTZ acquisition sensors with integrated control, acquisition and transmission functions is introduced. A strategy based on regional division is adopted to carry out targeted collection of crack areas according to predefined routes. Subsequently, an efficient and compact crack segmentation network structure is designed and constructed, and a variety of independently designed lightweight feature extraction and fusion modules are used. This design not only significantly reduces the complexity of the model, but also greatly improves the crack detection performance. This study pays special attention to the dynamic tracking and analysis of cracks, and adopts the Euclidean distance similarity classification method to ensure robust tracking and accurate quantitative analysis of each crack in the monitoring area.

[0094] The present invention is based on a lightweight monitoring method for fine cracks in complex background areas based on adaptive perception. The overall flow chart of this solution is shown in Figure 1 and includes the following components:

[0095] Component 1: Automatic Crack Information Collection Based on Regional Division: This method uses a PTZ camera sensor to automatically collect crack areas in blocks, overcoming the inherent contradiction between the field of view and recognition accuracy of crack collection under the fixed focal length of traditional cameras.

[0096] Component 2: Adaptive Complex Scale Calibration Method. The PTZ sensor's acquisition mechanism causes varying degrees of tilt distortion in each image region. A multi-scale template matching algorithm is used to adaptively correct this distortion across all regions and achieve pixel-accurate true scale conversion. This is shown in Figures 2 and 3.

[0097] Component 3: Lightweight Crack Segmentation Network (PCS-Net). It utilizes a compact structure and lightweight modular construction method to significantly reduce model complexity while maintaining excellent crack identification performance. The specific modules are shown in Figure 4.

[0098] Component 4: Dynamic Crack Tracking Algorithm Based on Similarity Classification. This quantitative crack tracking algorithm, based on Euclidean distance similarity classification, enables real-time control of dynamic information for each crack, which is numerous and has complex development trends. The specific process is shown in Figure 5.

[0099] Component 5: Crack Monitoring Sensing Network Architecture. Combining the aforementioned methods with the hardware platform, we propose an online perception and monitoring network system integrating sensing, computing, transmission, and analysis. This system aims to overcome communication barriers within the box girder and achieve comprehensive control of critical crack areas within the bridge. The specific architecture is shown in Figure 6.

[0100] Furthermore, the specific details of the automatic crack information collection method based on regional division described in Component 1 are as follows:

[0101] Traditional vision sensors improve crack recognition accuracy by reducing their field of view. However, if cracks are widely distributed, a single sensor's field of view may not cover the entire area. To overcome the inherent conflict between crack acquisition field of view and recognition accuracy, a PTZ acquisition sensor with integrated control, acquisition, and transmission capabilities was introduced. By using area division and path pre-planning, the original recognition area was expanded without compromising recognition accuracy.

[0102] Furthermore, the adaptive complex scale calibration method described in Component 2 includes the following steps: Step 1: a multi-scale template matching method based on color features: achieving automatic and stable extraction of the target area; Step 2: adaptive image calibration: achieving complex scale orthorectification.

[0103] Furthermore, the specific method of step 1 of the adaptive complex scale calibration method described in component 2: multi-scale template matching based on color features is as follows:

[0104] The Hue-Saturation-Value (HSV) color space, which aligns with human color perception, is used as a color representation. Template matching based on HSV features begins with discretizing the template image's histogram to create a descriptive representation of the features. The continuous HSV space is then mapped to a discrete histogram space. At each location, the template model is compared with the corresponding portion of the target image, and an HSV probability value is calculated. Regions with high matching scores are output as matching results. Template images with 0.25 and 0.75 times the original template are added to extract features from the target image. A weighted fusion strategy is employed to combine matching results from different scales into a comprehensive result. To ensure fairness, the fusion weight coefficient for templates at each scale is set to 1 / 3 to clearly indicate that matching results at each scale are equally important.

[0105] Furthermore, the specific method of step 2 of the adaptive complex scale calibration method described in component 2: adaptive image calibration is as follows:

[0106] First, the inner circle area is extracted from the target area. Then, the center point is determined based on the lines connecting the four equally divided points on the inner circle and used as the matching feature point. The other three matching feature points in the complete image are adaptively extracted using the same method. Then, a homography mapping relationship is established by combining the coordinate information of each feature point in real space. It can be expressed as:

[0107] Here (u t ,v t ) represents the image coordinates of the four matching feature points in the source image, (u r ,v r ) represents the plane coordinates of the four matching feature points in the target image in real space, t is the scale factor, represents the homography matrix.

[0108] The least squares method is used to estimate the homography matrix in the above equation. The core goal of this method is to achieve this by minimizing the reprojection error, that is, mapping the points on the source image to the target image through the estimated homography matrix, and then comparing it with the actual corresponding pixel points to find the optimal homography matrix parameters. Once this optimal homography matrix is ​​obtained, the tilted and distorted source images under various acquisition angles can be quickly corrected to obtain a corrected image with real size. The specific calculation formula is as follows:

[0109] Furthermore, the lightweight crack segmentation network (PCS-Net) described in component three includes the following steps:

[0110] Step 1: Bilateral downsampling and upsampling blocks are responsible for dimensionality reduction extraction and dimensionality increase learning of crack features; Step 2: Compressed sensing and attention fusion modules improve the model's training and inference speed and recognition accuracy; Step 3: Output module realizes the output of prediction results for crack segmentation detection.

[0111] Furthermore, the specific design details of the bilateral downsampling and upsampling blocks in step 1 of the lightweight crack segmentation network (PCS-Net) described in component 3 are shown in Figure 7. The specific method is as follows:

[0112] Bilateral downsampling uses a downsampling ratio of less than 32 times (3 bilateral downsampling blocks) to reduce the scale of the feature map to better preserve and extract key feature detail information. The downsampling block adopts a dual-path parallel structure, which is divided into a convolutional downsampling path and a pooling downsampling path with encoding. A 1×1 mapping convolution (MConv) is set at the front and back ends of the convolutional downsampling path to play the role of feature integration and channel adjustment. The middle part contains batch normalization (BN) for standardizing the distribution of channel features, a 2×2 depth convolution (D Conv), and the introduction of a nonlinear ReLU activation function. And using D Conv instead of the standard convolution structure helps to reduce the number of model parameters and computational complexity while maintaining the performance of the model. The convolutional downsampling branch (CDB) can be expressed as the following formula:

[0113] here represents the input feature map, For BN processing, is the ReLU activation function; the pooling-downsampling branch (PDB) uses a pooling layer with a pooling kernel size of 2×2 to downsample the feature map, while recording the location of the maximum value for feature restoration in the subsequent upsampling process. This method adopts a strategy that combines CDB and PDB, and simultaneously downsamples the input features to effectively preserve important details while maintaining lightweightness. BDB i The calculation process can be expressed as:

[0114] here represents the input features of the CDB branch of the i-th BDB module, represents the input features of the PDB branch of the i-th BDB module, represents the connection along the channel dimension, K i Represents the index encoding output of the i-th BDB module.

[0115] The dual-branch upsampling path follows the same design principles as the downsampling path and is divided into a deconvolution upsampling path (CUB) and a pooling upsampling path (PUB), aiming to better balance model accuracy and efficiency. Specifically, the deconvolution upsampling path consists of a 3×3 transposed convolution (T Conv). Meanwhile, the pooling upsampling path consists of a depooling layer (UP) with a pooling kernel of 2×2. TBDB i The calculation process can be expressed as:

[0116] here represents the input features of the CUB branch of the j-th BUB module, represents the input features of the PUB branch of the j-th BUB module, K j represents the index encoding output of the j+1th BDB module. Here C cu with C pu It is obtained by processing the feature map channel input to the j-th BUB module through the allocation coefficient λ. The specific value of λ will be described in the subsequent network structure design table.

[0117] Furthermore, the specific scheme of the compressed sensing and attention fusion module in step 2 of the lightweight crack segmentation network (PCS-Net) described in component 3 is shown in Figure 8. The specific method is as follows:

[0118] The compressed sensing unit is divided into two key stages, namely the asymmetric convolution stage and the asymmetric dilated convolution stage. In the asymmetric convolution stage, the input features are first By introducing the compression coefficient α, X is divided into the features to be convolved along the direction of the feature channel. With compression features The specific value of the compression coefficient α will be described in the subsequent network structure design table. Next, a set of 3 1 and 1 3 asymmetric convolutions are used to extract X c At the same time, X s The shortcut path designed will be used at the output stage with X' c Merge in the channel direction to obtain the features of stage one In the asymmetric dilated convolution stage, the same feature compression operation as in stage 1 is required, using the compression coefficient α to compress the input features. Divide and obtain the features to be convolved With compression features Next, use a set of (3+d i,z )×1 and 1×(3+d i,z ) combined with asymmetric dilated convolution to X' c Perform feature extraction, where d i,z represents the hole rate of the zth compressed sensing unit in the i-th compressed sensing module. At the same time, X' s Will be designed through the shortcut path in the output stage with X" c Merge in the channel direction to obtain the features of stage 2 The overall design of the perception unit draws on the structure of the residual network. In the final output stage, the goals of perception and feature enhancement are achieved through pixel-level feature fusion.

[0119] The compressed sensing unit is CPM i(i=1,2) basic components. CPM1 consists of 4 CS units, each with a void ratio of 0. Its main task is to capture the spatial characteristics of cracks. CPM2 consists of 8 CS units, and the void ratio of each CS unit follows the following value rules:

[0120] According to the above value selection rules, CPM2 can be regarded as a complex consisting of two multi-scale pyramid structures connected in series.

[0121] Based on the compressed sensing unit structure, this method introduces a very lightweight channel attention mechanism to construct an attention fusion unit. First, along the input feature Perform average pooling operation on the length and width direction of

[0122] here Global Average Pooling (GAP), input features Next, the method uses two special band matrices and Continuous Learning Channel attention in .

[0123] Here k s1 ,k s2 Indicates the channel range covered by each row in the band matrix, d s1 ,d s2 represents the cross-channel interaction rate within the coverage range of each row of channels in the band matrix, and k in the text s1 ,k s2 Both are set to 3, d s1 ,d s2 Set to 1, 3 respectively. In the implementation stage, a one-dimensional convolution kernel with a void rate can be used right Perform channel attention weight The specific calculation process of the extraction and attention module is as follows:

[0124] here Represents element-by-element multiplication, σ represents the Sigmoid function. AFM i (i=1,2) consists of an AFU and two CPUs, located in the decoding stage of the network structure. These two modules each fuse the different scale feature outputs from CPM1 and BDB1 to more effectively recover the spatial details of crack features.

[0125] Furthermore, the specific method of the output module in step 3 of the lightweight crack segmentation network (PCS-Net) described in component 3 is as follows:

[0126] The output of PCS-Net consists of three components: a main output and two auxiliary outputs. The main output consists of a transposed convolution with a kernel size of 2×2. The two auxiliary output modules consist of a 1×1 mapped convolution and a transposed convolution with a kernel size of 2×2. Network output Supervised by true labels It should be emphasized that the auxiliary output module is only enabled during the training phase and does not add additional computational burden during the inference phase.

[0127] Furthermore, the dynamic crack tracking algorithm based on similarity classification described in component four includes the following steps: Step 1: Establishing crack classification criteria to clarify the crack change situations involved in the crack evolution process; Step 2: Crack skeletonization to reduce the complexity of crack feature processing; Step 3: Euclidean distance similarity dynamic analysis criterion to further reduce the complex crack morphology information into a simpler similarity measurement and simplify the crack analysis process.

[0128] Furthermore, in step 1 of the dynamic crack tracking algorithm based on similarity classification described in component 4, the specific method for establishing the crack classification criteria is as follows:

[0129] This study categorizes cracks into three types: new cracks, extended cracks, and merged cracks. New cracks are cracks that did not exist in a specific area at time point t but appear at time point t+1, representing a new damage event. Extended cracks are cracks that existed at time point t and further expanded along their original development path at time point t+1, indicating an escalation of the damage event. Merged cracks are two or more cracks that existed at time point t and merged into a single crack at time point t+1, indicating regional expansion of damage.

[0130] Furthermore, the specific method of crack skeletonization in step 2 of the dynamic crack tracking algorithm based on similarity classification described in component 4 is as follows:

[0131] This method uses a fracture skeletonization approach to simplify the fracture object into a more manageable central skeleton structure. This skeletonization process not only significantly reduces the computational burden but also helps to more accurately identify the topological features of the fracture, thereby highlighting the morphological characteristics of the fracture development.

[0132] Furthermore, in step 3 of the dynamic crack tracking algorithm based on similarity classification described in component 4, the specific method of the Euclidean distance similarity dynamic analysis criterion is as follows:

[0133] This method uses Euclidean distance similarity The specific calculation formula for crack analysis is as follows:

[0134] here represents the i-th coordinate point on the center line of the m-th crack at time t, Represents the i-th coordinate point on the center line of the n-th crack at time t+1, i∈[0,N], N represents the length of the crack center line. This formula mainly uses the inverse distance transformation to convert distance to similarity. The characteristic of this transformation is that when the distance is smaller, the similarity value is closer to 1, indicating high similarity, and when the distance is larger, the similarity value is closer to 0, indicating low similarity. When N m =N n When N m ≠N n , we need to further consider how to comprehensively calculate the similarity between cracks of different lengths. Therefore, this method is based on the idea of ​​convolution operation, and here we assume that N m >N n , the shorter crack centerline coordinate set is regarded as 1×N n The convolution kernel of dimension 1×N is used, and the coordinate set of the crack centerline with a longer length is regarded as 1×N m The convolution target of dimension is set. Set the moving step of the pseudo convolution kernel to 1, and replace the convolution calculation with the similarity calculation, and finally get N m -N n +1 similarity index, and select the maximum value as the final similarity result. This calculation method is also applicable to N m <N n The crack similarity solution process.

[0135] The similarity metric effectively establishes correlations between crack images at different times, thereby forming a sequence of corresponding crack matching combinations. This establishes a one-to-one correspondence between newly emerged cracks and previously observed cracks. It's important to note that newly emerged cracks constitute a special case in this matching sequence because they lack similarity to any cracks from the previous time point. In other words, the crack matching combination contains only one crack. Therefore, the similarity matching results enable rapid extraction of newly emerged cracks. The crack matching combination also includes extended cracks and merged cracks. Based on their evolutionary characteristics, these two types of cracks can be easily separated and extracted from the crack matching combination. Merged cracks can only be extracted if they meet the following criteria: At least two cracks with high similarity matching exist at the previous time point, compared to the merged crack observed at the current time point. This means that the number of identical new crack matching combinations is greater than two. This principle ensures reliable detection of merged cracks. On the other hand, the classification criteria for extended cracks differ significantly from those for merged cracks. For extended cracks, only one crack with high similarity matching exists at the current time point. Therefore, extended cracks can be extracted based on the principle that the number of matching combinations for the new crack is equal to one.

[0136] The similarity-based crack classification method assigns a unique category attribute to each crack in an image, allowing accurate tracking of all cracks in the image. To comprehensively describe the development characteristics of cracks, six characteristic attributes, namely length, area, width, length change, area change, and width change, are introduced on the basis of the category attribute. The length of a crack usually refers to the length of the crack centerline, which can be obtained by measuring the length of the crack skeleton. The area of ​​a crack represents the pixel area occupied by the crack in the image, which can be obtained by performing regional statistics on the crack segmentation image. Based on the obtained crack length and area, the average width of the crack can be further calculated. Finally, the changing trends of length, area, and width over time are analyzed to gain a deeper understanding of the dynamic evolution of cracks.

[0137] Furthermore, the crack monitoring sensor network architecture described in Component 5 includes the following details: Based on the dynamic sensing method framework, a comprehensive crack monitoring sensor network is proposed, consisting of data acquisition nodes, data processing nodes, a data transmission layer, and an information aggregation layer. Its key features include modularity, high efficiency, security, and reliability. In terms of modularity, each data acquisition node consists of two crack monitoring sensors, forming a modular collection process. The data acquisition process is isolated between nodes, ensuring data security. To improve data processing efficiency, the data collected by each acquisition node will be processed by the corresponding data processing node (edge ​​computer). The data transmission layer, namely the server located at the construction site, will be responsible for establishing a reliable communication medium between the enclosed space inside the bridge and the outside world, transmitting data from each data processing node to a cloud service platform. As the information aggregation layer, the cloud service platform will conduct comprehensive analysis and evaluation of all data, providing engineers with intuitive and reliable support to support the implementation of operation and maintenance strategies.

[0138] Example 1

[0139] The proposed scheme was tested on the Humen Bridge, as shown in Figure 9. This bridge section utilizes a three-span prestressed concrete box-beam continuous rigid frame with a substructure of double-column hollow thin-walled piers, resulting in a total span of 570 meters. The first crack monitoring area is located at section 6 of beam section 19, near the bridge support and therefore identified as a key monitoring area, as shown in Figure 10(a). The second crack monitoring area is located at section 17 of beam section 20, three-quarters of the way along the beam. This section, with numerous cracks, is considered a key monitoring area by engineers, as shown in Figure 11(a). The two monitoring areas are divided into nine sub-areas of equal size, as shown in Figures 10(c) and 11(c). The captured images are automatically calibrated to assign each pixel a true physical size of ρ = 0.16 mm. The calibrated images are shown in Figures 12 and 13. This paper selects sub-area 19-6 as a case study for dynamic monitoring results. Figure 14(a) shows the original image of the 19-6 area collected on August 3, 2023. This time point is set as the starting time of the monitoring and analysis, and the images at subsequent times are compared with this. Subsequently, the PCSNet algorithm proposed in Component 3 of the invention is used to process the image of the 19-6 area, and the crack extraction results are presented in Figure 14(b). In order to clearly display the topological structure of the cracks in the area, the crack extraction results are marked one by one, as shown in Figure 14(c). In order to present the results of the dynamic crack analysis more clearly, this case only shows the crack monitoring data for two weeks. The dynamic crack analysis algorithm introduced in Component 4 of the invention is used to calculate the dynamic evolution of all cracks in the area at subsequent times, as shown in Figure 14(d).

[0140] We also compared the proposed method with a number of popular crack segmentation detection models, including BiSeNet, CGNet, and Fast-SCNN, on the same dataset. The precision evaluation metrics chosen were the mean intersection over union (MIoU) and Dice coefficient, commonly used in the field of deep learning, to measure the precision and accuracy of image segmentation. MIoU evaluates performance by calculating the ratio of the intersection over union (IoU) between the predicted segmentation result and the true segmentation mask, and is therefore better able to handle cases with imbalanced classes or uneven object sizes. The specific calculation formula is as follows:

[0141] K represents the number of categories, represents the number of pixels where the model predicts m as n, that is, the number of pixels that are not correctly classified as cracks (FN), represents the number of pixels where the model predicts n as m, that is, the number of pixels misclassified as cracks (FP), It represents the number of pixels for which the model predicts m as m and the number of pixels correctly classified as background (TP).

[0142] Dice is an indicator that evaluates performance by calculating the ratio of the overlap between two sets to their average size. The specific calculation formula is as follows:

[0143] Dice also considers the degree of overlap between the segmentation result and the true label, and performs well in imbalanced segmentation problems. MIoU emphasizes global coverage, while the Dicee coefficient focuses on local similarity. By combining them, a more comprehensive performance evaluation can be obtained, taking into account both the global and local characteristics of the segmentation task.

[0144] A comparison of prediction results from different methods is shown in Figure 15 below. The present invention uses parameter count and FLOPs as metrics for evaluating model efficiency. FLOPs represents the computational complexity of the model. By comprehensively considering the two efficiency metrics and two accuracy metrics in the table, PCS-Net achieves optimal object recognition performance while minimizing parameter count and computational complexity. Compared to ENet, which has the fewest parameters, PCS-Net's MIoU and Dice metrics are 2.48% and 3.64% higher, respectively. Compared to ERFNet, which has the highest accuracy in the table, PCS-Net has only 11.17% of the parameters and 24.45% of the computational complexity, but its MIoU and Dice are 1.29% and 1.87% higher, respectively.

[0145] In summary, the specific embodiments have verified the effectiveness of the solution proposed in the present invention and its applicability to complex projects.

[0146] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A lightweight monitoring method for fine cracks in complex background areas based on adaptive perception, characterized by: The method comprises the following steps: Step S1, automatically collecting crack information based on regional division, wherein a PTZ camera sensor is used to automatically collect crack area in blocks; Step S2: performing an adaptive complex scale calibration process, using a multi-scale template matching algorithm to adaptively correct distortion information in all regions and perform pixel-accurate true scale conversion; Step S3, constructing a lightweight crack segmentation network to process the data processed in step S2; Step S4: Real-time monitoring of the dynamic information of each crack is performed using a crack tracking quantitative algorithm based on Euclidean distance similarity classification.

2. The method for lightweight monitoring of fine cracks in complex background areas based on adaptive perception according to claim 1 is characterized in that: The automatic collection of crack information based on regional division in step S1 is specifically as follows: Through area division and path pre-planning, and selecting identifiable auxiliary markers as target areas, PTZ acquisition sensors are used to conduct targeted acquisition of crack areas.

3. The method for lightweight monitoring of fine cracks in complex background areas based on adaptive perception according to claim 1 is characterized in that: The adaptive complex scale calibration process in step S2 specifically includes: Step S21, a multi-scale template matching process based on color features, thereby achieving automatic and stable extraction of the target area; Step S22: Adaptive image calibration process, thereby achieving orthorectification of complex scales.

4. The method for lightweight monitoring of fine cracks in complex background areas based on adaptive perception according to claim 3 is characterized in that: The multi-scale template matching process based on color features in step S21 specifically includes: First, the template image is discretized and histogram estimated to create a descriptive representation of the features; then the continuous HSV space is mapped to the discrete histogram space; finally, at each position, the template model is compared with the corresponding part of the target image, and the HSV probability value is calculated. The area with high matching degree is output as the matching result; Based on the original template, template images with sizes of 0.25 and 0.75 times are added to extract features from the target image. In the matching result output stage, a weighted fusion strategy is used to merge matching results from different scales into a comprehensive result. The adaptive image calibration process in step S22 specifically includes: First, the inner circle area is extracted from the target area. The center point is determined by the connection line between the four equally divided points on the inner circle and used as the matching feature point. The other three matching feature points in the complete image are adaptively extracted according to the same process. Then, a homography mapping relationship is established by combining the coordinate information of each feature point in real space, which is expressed as: Where (u t ,v t ) represents the image coordinates of the four matching feature points in the source image, (u r ,v r ) represents the plane coordinates of the four matching feature points in the target image in real space, t is the scale factor, represents the homography matrix; Then, the least squares method is used to estimate the homography matrix in the above equation. The specific calculation is as follows:

5. The method for lightweight monitoring of fine cracks in complex background areas based on adaptive perception according to claim 1 is characterized in that: The lightweight crack segmentation network in step S3 includes: Bilateral downsampling and upsampling blocks are used for dimensionality reduction extraction and dimensionality increase learning of crack features. Compressed sensing and attention fusion module, used to improve the model's training and inference speed and recognition accuracy; The output module is used to output the prediction results of crack segmentation detection.

6. The method for lightweight monitoring of fine cracks in complex background areas based on adaptive perception according to claim 5 is characterized in that: The bilateral downsampling and upsampling block specifically includes a downsampling block and an upsampling block: The downsampling block uses a downsampling ratio of less than 32 times to reduce the scale of the feature map. The downsampling block adopts a dual-path parallel structure, namely a convolutional downsampling branch and a coding-containing pooling downsampling branch; The front end and back end of the convolutional downsampling branch are respectively equipped with a 1×1 mapping convolution. The middle part includes batch normalization for normalizing the distribution of channel features, a 2×2 depth convolution, and a ReLU activation function that introduces nonlinear properties. The convolution downsampling branch CDB is expressed as the following formula: in represents the input feature map, For BN processing, is the ReLU activation function, where D Conv is the depthwise convolution, M Conv is the mapping convolution, N, C, H, and W represent the batch, number of channels, height, and width of the input feature map, respectively; The pooling downsampling branch PDB uses a pooling layer with a pooling kernel size of 2×2 to downsample the feature map, and records the position of the maximum value for feature restoration in the subsequent upsampling process; The strategy of fusing CDB and PDB is adopted, and the input features are downsampled at the same time. The specific technical process is as follows: Among them BDB i is a bilateral downsampling module, represents the input features of the CDB branch of the i-th BDB module, represents the input features of the PDB branch of the i-th BDB module, represents the connection along the channel dimension, K i represents the index encoding output of the i-th BDB module, C cu with C pu It is the parameter obtained by processing the feature map channel input to the j-th BUB module through the allocation coefficient λ.

7. The method for lightweight monitoring of fine cracks in complex background areas based on adaptive perception according to claim 5 is characterized in that: The compressed sensing and attention fusion module includes a compressed sensing unit and an attention fusion unit; The compressed sensing unit includes two key stages, namely the asymmetric convolution stage and the asymmetric dilated convolution stage; In the asymmetric convolution stage, the input features are first Compression processing is performed by introducing the compression coefficient α to divide X into the features to be convolved along the direction of the feature channel With compression features Then a set of 3×1 and 1×3 asymmetric convolutions are used to extract X c The characteristic information of X s The shortcut path designed will be used at the output stage with X' c Merge in the channel direction to obtain the features of stage one In the asymmetric dilated convolution stage, the same feature compression operation as the asymmetric convolution stage is first performed, and the compression coefficient α is used to compress the input features. Divide and obtain the features to be convolved With compression features Then use a set of (3+d i,z )×1 and 1×(3+d i,z ) combined with asymmetric dilated convolution to X' c Perform feature extraction, where d i,z represents the void rate of the zth compressed sensing unit in the i-th compressed sensing module; at the same time, X' s The shortcut path designed will be used at the output stage with X' c , merge in the channel direction to obtain the features of stage 2 The compressed sensing unit is CPM i The basic components of CPM1 are as follows, where i = 1, 2; CPM1 consists of 4 compressed sensing units, and the void ratio of each compression unit is set to 0, which is used to capture the spatial characteristics of cracks; CPM2 consists of 8 compressed sensing units, and the hole rate of the compressed sensing unit follows the following value rules: where d i,z represents the void ratio of the compressed sensing unit, and z is the adjustment coefficient for controlling the size of the void ratio; The attention fusion unit introduces a lightweight channel attention mechanism based on the compressed sensing unit structure. The specific process is as follows: First, along the input features Perform average pooling operation on the length and width direction of The input features is global average pooling; Then take two band matrices and Continuous Learning Channel attention in where k s1 ,k s2 Indicates the channel range covered by each row in the band matrix, d s1 ,d s2 represents the cross-channel interaction rate within the coverage range of each row of channels in the band matrix, w c,c Represents the band matrix coefficient of the c-th row and c-th column, where c is the number of channels of the input feature; Use a one-dimensional convolution kernel with a dilation rate right Perform channel attention weight The specific calculation process of the extraction and attention module is as follows: in Represents element-by-element multiplication, σ represents the Sigmoid function, For BN processing, is the output feature, is the input feature, AFM i It is constructed by an AFU and two CPUs and is located in the decoding stage of the network structure, where i = 1, 2. These two modules each fuse the different scale feature outputs from CPM1 and BDB1.

8. The method for lightweight monitoring of fine cracks in complex background areas based on adaptive perception according to claim 5 is characterized in that: The output module includes three components: a main output and two auxiliary outputs; The main output part consists of a transposed convolution with a kernel size of 2×2; the two auxiliary output modules consist of a 1×1 mapping convolution and a transposed convolution with a kernel size of 2×2. The network output Y m , Y ao1 , Supervised by true labels 9. The method for lightweight monitoring of fine cracks in complex background areas based on adaptive perception according to claim 1 is characterized in that: The step S4 specifically includes: Step S41: Establishing a crack classification criterion to clarify the crack change conditions involved in the crack evolution process, specifically: Cracks are divided into three categories: new cracks, extended cracks, and merged cracks. New cracks refer to cracks that did not exist in a specific area at time point t and appeared at time point t+1, indicating a new damage event. Extended cracks refer to cracks that existed at time point t and further expanded along the original development path at time point t+1, indicating an escalation of the damage event. Merged cracks refer to two or more cracks that existed at time point t and merged into one crack at time point t+1, indicating regional expansion of damage. Step S42, crack skeletonization, reduces the complexity of crack feature processing, specifically: Simplify the crack object into a more manageable central skeleton structure; Step S43: Establishing a dynamic analysis criterion of Euclidean distance similarity, further reducing the complex crack morphology information into a simpler similarity metric, and simplifying the crack analysis process, specifically: Euclidean distance similarity The specific calculation formula for crack analysis is as follows: in represents the i-th coordinate point on the center line of the m-th crack at time t, represents the i-th coordinate point on the centerline of the n-th crack at time t+1, i∈[0,N], N represents the length of the crack centerline; Based on the idea of convolution operation, assuming N m >n n , the shorter crack centerline coordinate set is regarded as 1×n n The convolution kernel of dimension 1×n is used, and the coordinate set of the crack centerline with a longer length is regarded as 1×n m The convolution target of dimension is set to 1, and the moving step of the pseudo convolution kernel is replaced by similarity calculation. Finally, N m -N n +1 similarity index, and select the maximum value as the final similarity result; The correlation between crack images at different times is established through the similarity index, thereby forming a corresponding crack matching combination sequence, that is, forming a one-to-one matching relationship between the newly appeared cracks and the previous cracks.

10. The method for lightweight monitoring of fine cracks in complex background areas based on adaptive perception according to claim 1, characterized in that: The method further includes constructing a crack monitoring sensor network architecture as follows: Based on the dynamic perception method framework, a comprehensive crack monitoring sensing network is constructed, which consists of data acquisition nodes, data processing nodes, data transmission layer and information aggregation layer.

Citation Information

Patent Citations

  • PTZ camera moving target detection and recognition method based on dynamic background compensation and deep learning

    CN111738211A

  • Crack measurement method, device and equipment and storage medium

    CN116416233A

  • Crack estimation device and crack estimation method

    US20230251170A1

  • KR20230140621A

Cited By

  • Underwater crack segmentation-oriented color correction and texture sharpening double-branch enhancement system

    CN120766127A

  • Wood structure ancient building crack identification positioning and repairing guidance system and method

    CN120781441A

  • Table area positioning correction method based on image edge detection

    CN120807564A

  • Feature image automatic detection and comparison method

    CN120931964A

  • A feature image automatic detection comparison method

    CN120931964B