Intelligent recording and analyzing system and method for whole process of structural test

Through computer vision technology, combined with YOLOv8, DeepLabV3+ and DeepSORT models, efficient identification and real-time dynamic analysis of cracks in reinforced concrete structural components are achieved, solving the problems of low efficiency, poor accuracy and safety risks in existing technologies, and realizing the synchronous display and full process recording of component property changes.

CN119624940BActive Publication Date: 2025-10-17BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411844088.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-15
Publication Date
2025-10-17
Estimated Expiration
2044-12-15

AI Technical Summary

Technical Problem

In the existing technology, crack inspection of reinforced concrete structural components relies on manual marking, which is inefficient and has poor accuracy. It is also unable to achieve synchronous display and real-time recording of actual changes in component properties during loading, posing a safety risk.

Method used

A computer vision-based intelligent recording and analysis system for the entire structural test process is adopted, including a camera, a video acquisition module, a crack location module, a crack segmentation module and a parameter measurement module. The improved YOLOv8 and DeepLabV3+ models are used for crack identification and segmentation, combined with the DeepSORT method for crack tracking, to achieve real-time dynamic analysis of cracks.

Benefits of technology

It realizes the synchronous display and real-time recording of the actual property changes of the components during the loading process, improves the accuracy and safety of crack identification, and enables intuitive grasp of the entire process of overall damage development of the test components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624940B_ABST
    Figure CN119624940B_ABST
Patent Text Reader

Abstract

The present application relates to structural test whole process intelligent record analysis system and method, belong to structural test monitoring technical field.It is based on computer vision, including camera and video acquisition module, crack positioning module, crack segmentation module, parameter measurement module and crack tracking module which are communicated with the camera;In addition, it also has corresponding structural test whole process intelligent record analysis method, which includes crack positioning method, crack segmentation method and crack tracking method in turn.The present application can intuitively master the whole process of the test component overall damage development, including same frequency display and real-time record, and has broad application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a structural test recording method, in particular to a computer vision-based intelligent recording and analysis system and method for the whole process of structural test, and belongs to the technical field of structural test monitoring. BACKGROUND

[0002] Crack measurement and analysis of structural members in civil engineering structural tests are important for structural damage identification and evaluation. During the test process, recording the cracking and damage of the members can serve as an important basis for studying the performance and behavior of the structure. For example, in reinforced concrete structures, by tracking the appearance and development of cracks throughout the test process, the mode and rate of crack propagation can be explored, and the durability and remaining life of the structure can be predicted.

[0003] Although structural engineering tests are widely conducted, there are still some shortcomings in the test process and test result recording:

[0004] (1) The crack inspection of typical reinforced concrete members still uses manual marking with a hand-drawing pen. As the size and number of cracks continue to develop and evolve during the test process, the manual marking method consumes a large amount of work time for the test personnel and has low work efficiency.

[0005] (2) The traditional manual crack inspection results are not objective, the inspection accuracy for fine cracks is low, the crack recording error is large, and the size, direction, and other information of the cracks cannot be accurately represented.

[0006] (3) During the crack inspection process, the member is still under load, especially when it is close to the failure state, the member may experience rapid changes in deflection and deformation, and the risk of manual inspection and recording is high.

[0007] In addition, the test output of traditional reinforced concrete structural members is usually the load-displacement curve or load-strain curve, etc. Although this method can reflect the mechanical properties of the member, it cannot achieve the same frequency display and real-time recording of the actual behavior changes of the member during the loading process, and lacks intuitive understanding of the overall damage development process of the test member. SUMMARY

[0008] The purpose of the present application is to provide a computer vision-based intelligent recording and analysis system and method for the whole process of structural test, which solves or overcomes the above technical problems existing in the prior art, i.e., intuitive understanding of the overall damage development process of the test member, including same frequency display and real-time recording.

[0009] To achieve the above purpose, the technical solution adopted by the present application is as follows:

[0010] The structural test whole-process intelligent recording and analyzing system is based on computer vision, and comprises a camera and a video acquisition module, a crack positioning module, a crack segmentation module, a parameter measurement module and a crack tracking module which are in communication connection with the camera.

[0011] The camera records the test process based on the following processing mode: in order to meet the crack identification ability of the component, a plurality of cameras are connected in parallel, the number of cameras is configured as needed, the shooting parameters are adjusted, and the supporting light supplementing conditions are set for the whole test component and the key parts;

[0012] The video acquisition module is used to acquire the video images of the whole test process based on the following processing: time code synchronization is performed on the plurality of cameras, which is used to accurately trace back the test process from multiple angles; after the unified clock source, different cameras are triggered to image using the same instruction, and data is collected at the same time; finally, video fusion is performed to seamlessly splice the video streams from different cameras, so as to provide more comprehensive test scene coverage and richer structural state change information;

[0013] The crack positioning module is used to identify the cracking of the component based on the following processing: the test monitoring video acquired by the video acquisition module is processed frame by frame, and an improved YOLOv8 crack detection model is used to analyze the monitoring image to quickly locate the cracking position of the test component;

[0014] The crack segmentation module is used to finely extract crack information based on the following processing: for the cracks at the key positions of the test component located, an improved DeeplabV3+ crack segmentation model is used for pixel-level crack segmentation to exclude the interference of background noise and improve the accuracy and precision of crack identification;

[0015] The parameter measurement module is used to quantitatively analyze crack parameters based on the following processing: first, a scale factor is defined to convert the pixel scale of the crack into the real physical scale; the physical length, width and area are calculated based on the length, width, area and scale factor of the crack on the pixel level, and the crack distribution density is calculated based on the total surface area of the crack and the total area of the apparent image of the test component;

[0016] The crack tracking module is used to dynamically analyze the crack evolution process based on the following processing: after the crack is identified and the features are extracted by the model, a motion model is established based on the DeepSORT method to synchronize it with the test loading process, a matching matrix between cracks is constructed in the tracking process and the Hungarian algorithm is used to ensure the accuracy of cross-frame crack tracking, finally, the crack propagation path is presented in a visual way, and the time series data of the crack features changing with time is generated, which provides support for the analysis and prediction of the crack propagation law; wherein, the identification of the crack refers to the detection by the crack positioning module and the segmentation by the crack segmentation module, and the feature extraction refers to the length and width features of the crack obtained by the parameter measurement module.

[0017] Further, the camera parallel mode and the data transmission method use WiFi technology, Zigbee technology or Bluetooth technology; for data storage, the video data of multiple cameras can be stored in the same storage device, and the storage method is in chronological order, or a distributed storage method is used, and the data of part of the cameras is stored in the local device, and the rest is stored in the cloud server, to improve the security and availability of data.

[0018] Further, the number of cameras is determined according to the stress characteristics of different test components, so as to comprehensively obtain the whole process change of the component. The installation method can use a tripod or be fixed on the surrounding wall. The shooting parameters are determined according to actual needs, such as resolution, frame rate, focal length, viewing angle, etc. to accurately capture the cracks of the component; the matching light supplementing condition can be a white light supplementing lamp. The angle of the light supplementing lamp needs to match the shooting angle of the camera, to ensure that the light supplementing area is the area shot by the camera, and to avoid uneven light or no light in some areas.

[0019] Further, in the video acquisition module, the time code synchronization method uses NTP network time protocol. After connecting each camera to the network, accurate time is obtained by configuring the NTP server, then the camera sends a request to the server at regular intervals to obtain the latest time and calibrate it; on this basis, a trigger script is written using Python language, which iterates through all the IP addresses or device identifiers of the cameras, uses a network communication library and the API of the camera, and sends a start recording instruction to multiple cameras at the same time.

[0020] Based on the above time synchronization method for video acquisition, after the acquisition is completed, the videos of each camera are spliced and fused to represent the change of the component at the same time. First, the video is decomposed into frames, then each frame is used to extract features using the scale-invariant feature transform algorithm or the speeded-up robust features algorithm, find the matching points between different frames, splice these frames together by calculating the transformation matrix, and recombine the spliced frames into a video.

[0021] Further, in the parameter measurement module, the definition method of the scale factor a is as follows:

[0022] a = d d real

[0023] d pixel

[0024] where d real is the real physical size of the test component, d pixel is the corresponding component pixel size in the image, then the physical size calculation method of the crack is:

[0025] D real =αD pixel

[0026] wherein D can represent length, width and area respectively;

[0027] The model-identified crack image is converted into a binary image, white is set as a crack and black is set as a background, the area of the crack is calculated by counting the total number of white pixels in the crack region of the binary image; the crack length is calculated by using a skeletonization algorithm of the crack, which is an iterative thinning algorithm or a distance transform method, the crack region is extracted as a single-pixel-wide line and the total pixel length of the skeleton line is calculated; the crack width is calculated by the nearest distance between the crack region and the skeleton; and the crack distribution density is obtained by dividing the total area of the crack by the total area of the apparent image of the test member.

[0028] A structure test whole-process intelligent recording and analyzing method adopts the structure test whole-process intelligent recording and analyzing system, and sequentially comprises a crack positioning method, a crack segmentation method and a crack tracking method.

[0029] Further, the crack positioning method comprises the following steps:

[0030] S11. Constructing a crack data set according to the crack open source image collected in advance and the test monitoring image; performing geometric transformation, color transformation and adding noise on the crack image to expand the number of the data set and increase the diversity of the data;

[0031] S12. Marking a boundary box on the data set to clearly indicate the position of the crack and mark the crack category and boundary box coordinate content, and establishing a crack detection data set;

[0032] S13. Improving the YOLOv8 crack detection model algorithm and constructing a fused CFE module. Wherein, the FasterNet light and efficient convolutional neural network is added to reduce the resource consumption of the model and reduce the running time required for detection; at the same time, the efficient multi-scale attention module of EMA cross-space learning is introduced, and the batch reconstruction is performed on part of the channels, so that the YOLOv8 model is more targeted when extracting crack features;

[0033] S14. Dividing the data set and training the model, using the distributed focal loss and the complete intersection over union loss as the loss function, and using the precision, recall and average precision to evaluate the performance of the model;

[0034] S15. According to the actual demand, deploying the trained improved YOLOv8 crack detection model to the camera, considering the computing power and storage resources of the device, appropriately compressing and optimizing the improved YOLOv8 crack detection model to ensure that the model can run efficiently.

[0035] Further, the crack segmentation method comprises the following steps:

[0036] S21. On the basis of the above crack data set, a data enhancement method is used to ensure that the image has different crack shapes, sizes, directions and different background environments;

[0037] S22. The crack image is pixel-level labeled, wherein the crack part is the foreground and the remaining part is the background, and a crack segmentation data set is established;

[0038] S23. The DeepLabV3+ model algorithm is improved, and MobilenetV2 is used to replace the original backbone feature extraction network to reduce the number of model parameters and improve the operation speed; in addition, an ECANet cross-channel interaction attention module is added to make the model more focused on crack information and enhance the performance of crack segmentation;

[0039] S24. The data set is divided and the model is trained, cross-entropy loss and Dice loss are used as loss functions, and class average pixel accuracy and average intersection over union are used to evaluate the performance of the model;

[0040] S25. The improved DeepLabV3+ crack segmentation model trained in S24 is deployed to the camera to perform real-time crack segmentation; at the same time, data feedback in the actual test process is collected, and the model is continuously optimized according to the feedback to adapt to different crack segmentation environments and requirements.

[0041] Further, the crack tracking method comprises the following steps:

[0042] S31. Combining the motion information and appearance features of the crack, the DeepSORT method is used to realize accurate association across time frames; after identifying and extracting the features of the crack, a Kalman filter is used to establish a motion model of the crack, and the crack position and size in the current frame are predicted according to the state of the previous frame to provide a preliminary association reference; at the same time, the crack appearance features extracted by the above module are embedded; during tracking, the motion model distance predicted by the Kalman filter and the similarity of the appearance features are combined to construct a matching matrix between cracks; the Hungarian algorithm is used to detect the optimal allocation of the bounding box and the tracking trajectory, ensuring that the same crack remains consistent in consecutive frames;

[0043] S32. After matching, the state of the crack is dynamically updated, and the unmatched crack is initialized as a new track, and finally the tracking number of the crack and its evolution track are output, the crack propagation path is presented through visualization, and time series data of crack features over time are generated.

[0044] Furthermore, in said S11, the geometric transformation method is rotation, flipping, cropping, and scaling; the color transformation method is color adjustment, brightness adjustment, and contrast adjustment; and the noise addition method is Gaussian noise and salt and pepper noise;

[0045] In S12, the Make Sense annotation tool is used to perform bounding box annotation and category annotation on the crack information in the dataset, clarifying the location of the cracks and indicating the crack categories and bounding box coordinates, so as to directly generate the label format required by the YOLO series algorithm and establish a crack detection dataset;

[0046] In S13, the FasterNet network is integrated with the EMA attention mechanism to construct a CFE module and apply it to the backbone network and neck network of the model. FasterNet, as a lightweight and efficient convolutional neural network, is used to reduce the resource consumption of the model and shorten the running time required for detection. EMA, as an efficient multi-scale attention module for cross-space learning, can batch reconstruct some channels, making the model more targeted when extracting crack features.

[0047] In S14, the crack detection dataset is divided into a training set and a test set according to a certain ratio, wherein the training set is used to train the model, and the test set is used to evaluate the performance of the model;

[0048] Use distribution focus loss and complete intersection-over-union loss as loss functions;

[0049] The distribution focus loss is described as follows: for a continuous value y∈R, R is a set of real numbers; it is discretized into the integer interval [l,u], where l is the maximum integer not greater than y rounded down; u is the minimum integer not less than y rounded up; assuming that the predicted distribution probability p = {p0, p1, ..., p n}, corresponding to the discrete position {0,1,…,n}, the calculation formula of the distribution focus loss is:

[0050] L DFL =-((uy)log(p l )+(yl)log(p u ))

[0051] Among them, y is the target value, p l and p u is the predicted probability corresponding to the interval [l,u]; uy and yl are weights, reflecting the degree of deviation of the target value within the interval;

[0052] Complete Intersection-over-Union loss assumes the predicted box is B p , the target box is B g, the overlapping area is Area, the Euclidean distance between the center points of the prediction box and the target box is p (l p , l g ), the diagonal length of the minimum closed box enclosing the prediction box and the target box is d, the height and width of the prediction box and the target box are h p , w p and h g , w g , the weight factor is a, used to balance the influence of the aspect ratio term and other terms, then the calculation formula is:

[0053]

[0054] In the S14, in the evaluation of the model performance by using the precision, recall and average precision,

[0055] TP represents the number of samples that are actually cracks and are correctly detected as cracks, FN represents the number of samples that are actually cracks but are incorrectly detected as backgrounds, FP represents the number of samples that are actually backgrounds but are incorrectly detected as cracks, and TN represents the number of samples that are actually backgrounds and are correctly detected as backgrounds.

[0056] The precision is the proportion of the number of samples that are actually cracks and are detected by the model TP to the total number of samples that are detected as cracks TP+FP, and is represented as:

[0057]

[0058] The recall is the proportion of the number of samples that are actually cracks and are detected by the model TP to the total number of samples that are actually cracks TP+FN, and is represented as:

[0059]

[0060] The average precision is the area below the precision-recall curve, and the higher the value represents that the model is easier to maintain high precision at high recall, and the calculation formula is:

[0061]

[0062] In the S15, after the model is iteratively trained and the performance meets the actual requirements, it is encapsulated and deployed in the camera by using TorchScript or TensorFlow Lite to execute;

[0063] In the S23, firstly, a lightweight neural network MobileNetV2 is used instead of the original DeepLabV3+ semantic segmentation model backbone feature extraction network Xception, which is used to effectively reduce the number of parameters, balance the calculation accuracy and speed, and make the model more suitable for embedded devices such as cameras; secondly, an attention module ECANet is added after the ASPP hollow space pyramid pooling module extracts features and before the low-order features are spliced, so that the model can better focus on the crack information of the test component and suppress invalid and complex background features, and improve the accuracy of the model;

[0064] In the S24, the cross-entropy loss L CE measures the difference between the predicted probability distribution and the true distribution, and the calculation formula is:

[0065]

[0066] where N is the total number of pixels in the image, C is the number of categories, y i,c is the true label of the i-th pixel belonging to category c, is the predicted probability of the i-th pixel belonging to category c;

[0067] The Dice loss is based on the Dice similarity coefficient L Dice , which is used to measure the overlap between the predicted and true segmentation, and the formula is:

[0068]

[0069] where p i is the predicted probability of the i-th pixel belonging to the crack, and g i is the true label of the i-th pixel belonging to the crack;

[0070] In the S24, the class average pixel accuracy MPA is used to measure the proportion of pixels classified correctly in each class, and then the average of all classes is calculated:

[0071]

[0072] where TP c represents the number of pixels whose true label is category c and whose prediction is also category c, and FN c represents the number of pixels whose true label is category c but is predicted to be other categories;

[0073] The mean intersection over union MIoU measures the ratio of the intersection to the union of the predicted result and the true segmentation, and the average of all categories is calculated:

[0074]

[0075] where FP crepresents the number of pixels for which the predicted label is class c but the real label is not c;

[0076] The performance of the model reaches the expected performance on the crack segmentation dataset, and then the model size is reduced and the inference speed is improved by pruning and quantization methods, and then converted into ONNX format or TensorRT format suitable for deployment, encapsulated and deployed in the camera to perform;

[0077] In the S31, the DeepSORT crack tracking method mainly refers to combining a simple and efficient Kalman filter motion model and deep feature measurement to perform real-time crack expansion state tracking;

[0078] The Kalman filter assumes that the crack state is a vector x k , and the observation is z k , then the change of the crack state is described by the state transition equation:

[0079] x k = F k x k-1 + Bu k + w k k

[0080] Where x k represents the state vector of the crack at time k, F k is the state transition matrix, u k is the control input vector, B k is the control input matrix, and w k is the process noise.

[0081] The motion model of the crack uses Kalman filtering to model the geometric features and position state of the crack: the state vector is defined as the center point position (x, y) of the crack, the size of the bounding box and its change speed; Kalman filtering predicts the possible position and size of the crack in the next frame according to the current frame state, providing a reference for association;

[0082] The crack matching and association method is to match the new cracks detected in each frame of image with the cracks tracked in the last frame, and to realize it by combining distance measurement;

[0083] The distance measurement adopts motion model distance and appearance feature distance; the motion model distance calculates the Euclidean distance between the state predicted by Kalman filtering and the actual detection result, and the appearance feature distance calculates the cosine distance or Euclidean distance of the features of the current frame and the last frame crack by embedding the extracted features;

[0084] ​The two distances are weighted and fused to construct a matching matrix. The Hungarian algorithm is then used to achieve the optimal assignment of crack identification results to tracking trajectories. This involves finding a minimum set of matching pairs on the matching matrix, such that each trajectory matches at most one detection box, and each detection box matches at most one trajectory. Unmatched detection boxes are initialized as new crack trajectories.

[0085] In S32, the dynamic update method updates the position, size, and appearance characteristics of successfully matched crack trajectories to keep them consistent with the actual state of the cracks. For unmatched crack trajectories, the tracking continues based on the prediction of the Kalman filter. If no new detection results are matched within a period of time, the crack is considered to have stopped expanding.

[0086] After tracking is completed, the dynamic trajectory of the crack is visualized and output in the following ways: the crack number and trajectory line are displayed on the original image, the curves of the crack width and length changing with time are drawn, and an animation of the crack expansion is generated to demonstrate its dynamic evolution process.

[0087] After adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art:

[0088] The present invention can effectively realize the synchronous display and real-time recording of the actual property changes of the component during the loading process, and can intuitively grasp the entire process of the overall damage development of the test component. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 This is a schematic structural diagram of an embodiment of the present invention;

[0090] Figure 2 Schematic diagram of the arrangement of the test component and the camera of the present invention;

[0091] Figure 3 This is a schematic diagram of the improved YOLOv8 crack detection network of the present invention;

[0092] Figure 4 Schematic diagram of the improved DeepLabV3+ crack segmentation network of the present invention;

[0093] Figure 5 This is the core flow chart of the crack dynamic tracking method of the present invention. DETAILED DESCRIPTION

[0094] The following is combined with Figures 1-5 The present invention will be further described in detail with specific implementations to facilitate a clear understanding of the present invention, but they do not constitute a limitation to the present invention.

[0095] In the description of the present application, it should be noted that the terms "upper", "lower", "front", "back", "left", "right", "vertical", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.

[0096] In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0097] As shown in the accompanying drawings Figures 1-2 The structural test whole-process intelligent recording and analyzing system of the embodiment is based on computer vision, and includes a camera 1 and a video acquisition module, a crack positioning module, a crack segmentation module, a parameter measurement module and a crack tracking module in communication connection with the camera 1;

[0098] The camera 2 records the test process based on the following processing mode: in order to meet the crack identification ability of the component, a plurality of cameras 1 are connected in parallel, the number of cameras is configured as needed, the shooting parameters are adjusted, and the supporting light supplementing conditions are set for the whole and key parts of the test component 2; as shown in the accompanying drawings Figure 2 The camera 1 parallel connection mode and the data transmission method use WiFi technology, Zigbee technology or Bluetooth technology; for data storage, the video data of the plurality of cameras 1 can be stored in the same storage device in time sequence, or a distributed storage mode is adopted, the data of part of the cameras are stored in a local device, and the rest are stored in a cloud server, so as to improve the security and availability of the data. In the embodiment, the number of cameras 1 is determined according to the stress characteristics of different test components 2, and the installation mode can use a tripod or be fixed on the surrounding wall. The shooting parameters are determined according to actual needs, and the resolution, frame rate, focal length, viewing angle and the like are determined according to the ability to accurately capture the cracks of the component; the supporting light supplementing conditions can be a white light supplementing lamp, the angle of the light supplementing lamp needs to match the shooting angle of the camera, to ensure that the area of light supplementing is the area shot by the camera, avoiding the situation that the light is not uniform or part of the area is not illuminated.

[0099] The video acquisition module is used to acquire video images of the whole test process based on the following processing: time code synchronization is performed on the laid multiple cameras 1, which is used to accurately trace the test process from multiple angles; after the unified clock source, the same instruction is used to trigger the imaging of different cameras 1, and data is collected at the same time line; finally, video fusion is performed to seamlessly splice the video streams from different cameras 1 to provide more comprehensive test scene coverage and richer structural state change information; wherein, in the video acquisition module, the time code synchronization method adopts NTP network time protocol, after connecting each camera 1 to the network, the accurate time is obtained by configuring the NTP server, then the camera 1 sends a request to the server at regular intervals to obtain the latest time and calibrate; on this basis, a trigger script is written using Python language, the IP addresses or device identifiers of all cameras 1 are iterated, and the network communication library and the API of the camera 1 are used to send the start recording instruction to multiple cameras 1 at the same time.

[0100] Based on the above time synchronization method for video acquisition, after the acquisition is completed, the videos of each camera 1 are spliced and fused to represent the change of the member at the same time; first, the video is decomposed into frames, then each frame is used to extract features using the scale-invariant feature transform algorithm or the speeded-up robust features algorithm, find the matching points between different frames, splice these frames together by calculating the transformation matrix, and recombine the spliced frames into a video.

[0101] The crack positioning module is used to identify the cracking of the member based on the following processing: the test monitoring video acquired by the video acquisition module is processed frame by frame, and the improved YOLOv8 crack detection model is used to analyze the monitoring image to quickly locate the cracking position of the test member 2;

[0102] The crack segmentation module is used to finely extract crack information based on the following processing: for the cracks at the key positions of the test member 2 located, the improved DeeplabV3+ crack segmentation model is used for pixel-level crack segmentation to exclude the interference of background noise and improve the accuracy and accuracy of crack identification;

[0103] The parameter measurement module is used to quantitatively analyze the crack parameters based on the following processing: first, define the scale factor for converting the pixel scale of the crack into the real physical scale; calculate the physical length, width and area based on the length, width, area and scale factor of the crack on the pixel level, and calculate the crack distribution density based on the total surface area of the crack and the total area of the test member apparent image;

[0104] Wherein, the definition method of the scale factor a is as follows:

[0105]

[0106] wherein d real is the real physical size of the test component, d pixel is the corresponding component pixel size in the image, then the physical size calculation method of the crack is:

[0107] D real = aD pixel

[0108] wherein D can represent length, width and area respectively;

[0109] The crack image identified by the model is converted into a binary image, white is set as the crack and black is set as the background, the area of the crack is calculated by counting the total number of white pixels in the crack region of the binary image; the length of the crack is calculated by using the skeletonization algorithm of the crack, which is an iterative thinning algorithm or a distance transform method, the crack region is extracted as a single-pixel-wide line and the total pixel length of the skeleton line is calculated; the width of the crack is calculated by the nearest distance between the crack region and the skeleton; the distribution density of the crack is obtained by dividing the total area of the crack by the total area of the apparent image of the test component.

[0110] The crack tracking module is used to analyze the crack evolution process dynamically based on the following processing: after the crack is identified and the features are extracted by the model, a motion model is established based on the DeepSORT method to synchronize it with the test loading process, a matching matrix between cracks is constructed during the tracking process and the Hungarian algorithm is used to ensure the accuracy of the cross-frame crack tracking, finally the crack propagation path is presented in a visual way and the time series data of the crack features changing with time is generated, which provides support for the analysis and prediction of the crack propagation law. Wherein, the identification of the crack refers to the detection by the crack positioning module and the segmentation by the crack segmentation module, and the feature extraction refers to the length and width features of the crack obtained by the parameter measurement module.

[0111] Embodiment 2

[0112] In this embodiment, it is a whole-process intelligent recording and analysis method for structural test, which adopts each module of the whole-process intelligent recording and analysis system for structural test in embodiment 1, and the whole-process intelligent recording and analysis method for structural test comprises a crack positioning method, a crack segmentation method and a crack tracking method in sequence.

[0113] The crack positioning method comprises the following steps:

[0114] S11. Construct a crack dataset according to the pre-collected crack open source images, such as CFD dataset or SDNET dataset, and test monitoring images; perform geometric transformation, color transformation and noise addition on the crack images to expand the number of the dataset and increase the diversity of the data; the geometric transformation methods are rotation, flipping, cropping and scaling, the color transformation methods are color adjustment, brightness adjustment and contrast adjustment, and the noise addition methods are Gaussian noise and salt and pepper noise;

[0115] S12. Perform boundary box labeling on the dataset to clearly indicate the position of the crack and mark the crack category and boundary box coordinate content, and establish a crack detection dataset; specifically, use the Make Sense labeling tool to perform boundary box labeling and category labeling on the crack information in the dataset to clearly indicate the position of the crack and mark the crack category and boundary box coordinate content, to directly generate the label format required by the YOLO series algorithm, and establish a crack detection dataset;

[0116] S13. Improve the YOLOv8 crack detection model algorithm and construct a fused CFE module; wherein, by increasing the FasterNet lightweight and efficient convolutional neural network, the resource consumption of the model is reduced, and the running time required for detection is reduced; at the same time, the efficient multi-scale attention module of EMA cross-space learning is introduced, and part of the channel is reconstructed in batches, so that the YOLOv8 model is more targeted when extracting crack features; in the S13, as shown in the figure, Figure 3 the FasterNet network is fused with the EMA attention mechanism to construct a CFE module and apply it to the backbone network part and the neck network part of the model; wherein, the FasterNet is a lightweight and efficient convolutional neural network, which is used to reduce the resource consumption of the model and reduce the running time required for detection; the EMA is an efficient multi-scale attention module of cross-space learning, which can reconstruct part of the channel in batches, so that the model is more targeted when extracting crack features;

[0117] S14. Divide the dataset and train the model, use the distribution focal loss and the complete intersection over union loss as the loss function, and use the precision, recall and average precision to evaluate the performance of the model; wherein

[0118] in the S14, the crack detection dataset is divided into a training set and a test set according to a certain proportion (9:1 / 8:2 / 7:3), wherein the training set is used to train the model, and the test set is used to evaluate the performance of the model;

[0119] the distribution focal loss (Distribution Focal Loss) and the complete intersection over union loss (Complete Intersection over Union Loss) are used as the loss function;

[0120] Distribution Focal Loss (DFL) is described as, for a continuous value y∈R, R is the set of real numbers; it is discretized into an integer interval [l, u], where Assuming the predicted distribution probability p = {p0, p1, …, pn} corresponds to the discrete positions {0, 1, …, n}, the calculation formula of the distribution focal loss is: n

[0121] L DFL = -((u-y)log(p l )+(y-l)log(p u ))

[0122] Where y is the target value, p l and p u are the predicted probabilities corresponding to the interval [l, u]; u-y and y-l are weights, representing the degree of deviation of the target value in the interval;

[0123] The Complete Intersection Over Union Loss (CIoUL) assumes that the predicted box is B p , the target box is B g , the overlapping area is Area, the Euclidean distance between the centers of the predicted box and the target box is ρ(λ p ,λ g ), the diagonal length of the smallest closed frame enclosing the predicted box and the target box is d, and the height and width of the predicted box and the target box are h p , w p and h g , w g , the weight factor is α, used to balance the influence of the height-width ratio term and other terms, then the calculation formula is:

[0124]

[0125] In the S14, the Precision, Recall and Average Precision (AP) are used to evaluate the performance of the model,

[0126] TP represents the number of samples that are actually cracks and are correctly detected as cracks, FN represents the number of samples that are actually cracks but are incorrectly detected as background, FP represents the number of samples that are actually background but are incorrectly detected as cracks, and TN represents the number of samples that are actually background and are correctly detected as background.

[0127] The precision is the proportion of the number of samples that are actually cracks TP to the total number of samples that are detected as cracks TP+FP, and is expressed as:

[0128]

[0129] ​The recall rate refers to the proportion of the number of samples TP actually being cracks detected by the model to the number of samples TP actually being cracks + FN, and is expressed as:

[0130]

[0131] The average precision is the area below the precision-recall curve, and the higher the value, the easier the model is to maintain high precision at high recall. The calculation formula is:

[0132]

[0133] S15. According to the actual demand, the trained YOLOv8 crack detection model is deployed to the camera, considering the computing power and storage resources of the device, the model is appropriately compressed and optimized to ensure that the model can run efficiently. That is, after the model is iteratively trained and its performance meets the actual demand, it is encapsulated and deployed in the camera (1) using TorchScript or TensorFlow Lite to execute.

[0134] In this embodiment, the crack segmentation method comprises the following steps:

[0135] S21. On the basis of the above crack data set, use data enhancement method to ensure that the image has different crack shapes, sizes, directions and different background environment;

[0136] S22. Pixel-level labeling is performed on the crack image, where the crack part is the foreground and the rest is the background, and a crack segmentation data set is established;

[0137] S23. The DeepLabV3+ model algorithm is improved, using MobilenetV2 to replace the original backbone feature extraction network, to reduce the number of model parameters and improve the operation speed; In addition, the ECANet cross-channel interaction attention module is added, so that the model pays more attention to crack information and enhances the performance of crack segmentation; as Figure 4 As shown in the figure, in S23, first, the lightweight neural network MobileNetV2 is used to replace the original backbone feature extraction network Xception in the DeepLabV3+ semantic segmentation model, to effectively reduce the number of parameters, balance the calculation accuracy and speed, and make the model more suitable for embedded devices such as cameras; Second, the attention module ECANet is added after the ASPP hollow space pyramid pooling module extracts the features and before the low-order features are spliced, so that the model can better focus on the crack information of the test component (2) and suppress invalid and complex background features, improving the accuracy of the model;

[0138] S24. The data set is divided and the model is trained, cross-entropy loss and Dice loss are used as loss functions, and the model performance is evaluated by using class average pixel accuracy (MPA) and mean intersection over union (MIoU); specifically, preferably, cross-entropy loss (Cross-Entropy Loss) and Dice loss (Dice Loss) are used as loss functions.

[0139] In the S24, the cross-entropy loss L CE The difference between the predicted probability distribution and the true distribution is measured, and the calculation formula is:

[0140]

[0141] Where N is the total number of pixels in the image, C is the number of classes, y i,c is the true label of the i-th pixel belonging to class c, is the predicted probability of the i-th pixel belonging to class c;

[0142] The Dice loss is based on the Dice similarity coefficient L Dice , which is used to measure the overlap between the prediction and the true segmentation, and the formula is:

[0143]

[0144] Where p i is the predicted probability of the i-th pixel belonging to the crack, g i is the true label of the i-th pixel belonging to the crack;

[0145] In the S24, the class average pixel accuracy MPA is used to measure the proportion of pixels classified correctly in each class, and then the average of all classes is calculated:

[0146]

[0147] Where TP c represents the number of pixels whose true label is class c and whose prediction is also class c, FN c represents the number of pixels whose true label is class c but is predicted to be other classes;

[0148] The mean intersection over union MIoU measures the ratio of the intersection to the union of the predicted result and the true segmentation, and the average of all classes is calculated:

[0149]

[0150] Where FP c represents the number of pixels whose predicted label is class c but whose true label is not c;

[0151] The performance of the model reaches the expected performance on the crack segmentation dataset, and then the model size is reduced and the inference speed is improved through pruning and quantization methods, and then converted into ONNX format or TensorRT format suitable for deployment, encapsulated and deployed in the camera (1) to perform;

[0152] S25. The improved DeepLabV3+ crack segmentation model trained in S24 is deployed in the camera for real-time crack segmentation; at the same time, data feedback during the actual test process is collected, and the model is continuously optimized according to these feedbacks to adapt to different crack segmentation environments and requirements.

[0153] Figure 5 A core flowchart of the crack dynamic tracking method is given, and the crack tracking method comprises the following steps:

[0154] S31. Combined with the motion information and appearance features of the cracks, the DeepSORT method is used to realize accurate association across time frames; after the cracks are identified and features are extracted, the Kalman filter is used to establish a motion model of the cracks, the crack position and size in the current frame are predicted according to the state of the previous frame, and a preliminary association reference is provided; at the same time, the crack appearance features extracted by the above module are embedded; during the tracking process, the motion model distance predicted by the Kalman filter and the similarity of the appearance features are combined to construct a matching matrix between the cracks; the Hungarian algorithm is used to detect the optimal allocation of the bounding box and the tracking trajectory, so that the same crack remains consistent in consecutive frames;

[0155] In S31, the DeepSORT crack tracking method mainly refers to combining a simple and efficient Kalman filter motion model with deep feature measurement to realize real-time crack expansion state tracking;

[0156] The Kalman filter assumes that the crack state is a vector x k , and the observation value is z k , then the change of the crack state is described by the state transition equation:

[0157] x k =F k x k-1 +B k u k +w k

[0158] Where x k represents the state vector of the crack at time k, F k is the state transition matrix, u k is the control input vector, B k is the control input matrix, and w k is the process noise.

[0159] The motion model of the crack utilizes Kalman filtering to model the geometric features and position state of the crack: a state vector is defined as the position (x, y) of the center point of the crack, the size of the bounding box and the change speed thereof; the Kalman filtering predicts the possible position and size of the crack in the next frame according to the state of the current frame, to provide a reference for association;

[0160] The crack matching and association method is to match the newly detected crack in each frame of image with the crack tracked in the last frame, and to realize the matching in combination with distance measurement;

[0161] The distance measurement adopts motion model distance and appearance feature distance; the motion model distance is calculated by the Euclidean distance between the state predicted by the Kalman filtering and the actual detection result, and the appearance feature distance is calculated by the cosine distance or Euclidean distance of the features embedded in the current frame and the last frame of the crack;

[0162] The two kinds of distances are weighted and fused to construct a matching matrix, and the optimal allocation of the crack recognition result and the tracking trajectory is realized by the Hungarian algorithm, that is, a set of minimum matching pairs is found on the matching matrix, so that each trajectory matches at most one detection box, and each detection box matches at most one trajectory; the unmatched detection box is initialized as a new crack trajectory;

[0163] S32. After the matching is completed, the state of the crack is dynamically updated, the unmatched crack is initialized as a new trajectory, and the tracking number of the crack and the evolution trajectory thereof are finally output, the crack propagation path is presented in a visual manner, and the time sequence data of the crack features changing over time are generated. Specifically, in the S32, for the successfully matched crack trajectory, the position, size and other states and appearance features thereof are updated to keep consistent with the actual state of the crack. For the unmatched crack trajectory, the prediction by the Kalman filtering is continued, and if no new detection result is matched for a period of time, it is considered that the crack stops expanding;

[0164] After the tracking is completed, the dynamic trajectory of the crack is visually output, and the output mode is that the number and trajectory line of the crack are displayed on the original image, the curves of the width and length of the crack changing over time are drawn, and the animation of the crack propagation is generated to demonstrate the dynamic evolution process thereof.

[0165] The above is only a preferred embodiment of the present application, and does not limit the structure of the present application in any form. The arrangement type and the number of uses of the present application are not limited to the example, and can be optimized and selected according to the actual engineering. Any modification, equivalent change and decoration of the above embodiment according to the technical principle of the present application, which does not deviate from the technical solution of the present application, is still within the scope of the technical solution of the present application.

Claims

1. An intelligent recording and analysis system for the entire structural test process, based on computer vision, featuring: It includes a camera (1) and a video acquisition module, a crack positioning module, a crack segmentation module, a parameter measurement module and a crack tracking module that are connected to the camera; The camera (1) records the test process based on the following processing method: in order to meet the crack recognition capability of the component, multiple cameras (1) are connected in parallel, and the number of cameras is configured as needed for the entire test component (2) and key parts, and the shooting parameters are adjusted and the supporting fill light conditions are set; The video acquisition module is used to obtain video images of the entire test process based on the following processing: time code synchronization of multiple cameras (1) arranged to accurately trace the test process from multiple angles; After unifying the clock source, the same command is used to trigger different cameras (1) to image and collect data simultaneously on the same timeline; finally, video fusion is performed to seamlessly splice the video streams from different cameras (1) to provide more comprehensive test scene coverage and richer structural state change information; The crack locating module is used to identify component cracks based on the following processing: processing the test monitoring video acquired by the video acquisition module frame by frame, analyzing the monitoring image using an improved YOLOv8 crack detection model to quickly locate the crack position of the test component (2); The crack segmentation module is used to extract crack information in a refined manner based on the following processing: for cracks at key positions of the located test component (2), pixel-level crack segmentation is performed using an improved DeeplabV3+ crack segmentation model to eliminate interference from background noise and improve the precision and accuracy of crack identification; The parameter measurement module is used to quantitatively analyze crack parameters based on the following processing: first, a scaling factor is defined to convert the pixel scale of the crack into a real physical scale; the physical length, width, and area of ​​the crack are calculated based on the length, width, and area of ​​the crack at the pixel level and the scaling factor; and the crack distribution density is calculated based on the total surface area of ​​the crack and the total area of ​​the surface image of the test component; The crack tracking module is used to dynamically analyze the crack evolution process based on the following processing: after the crack is identified and its features are extracted by the model, a motion model is established based on the DeepSORT method to synchronize it with the test loading process. During the tracking process, a matching matrix between cracks is constructed and the Hungarian algorithm is used to ensure the accuracy of cross-frame crack tracking. Finally, the crack propagation path is presented in a visual manner, and time series data of crack characteristics changing over time is generated to support the analysis and prediction of crack propagation patterns. Among them, crack identification refers to detection by the crack positioning module and segmentation by the crack segmentation module, and feature extraction refers to the length and width characteristics of the crack obtained by the parameter measurement module.

2. The intelligent recording and analysis system for the entire structural test process according to claim 1 is characterized by: The parallel connection mode of the cameras (1) and the data transmission method use WiFi technology, Zigbee technology or Bluetooth technology; for data storage, the video data of multiple cameras (1) can be stored in the same storage device in a chronological order, or a distributed storage method is used to store the data of some cameras in a local device and the rest in a cloud server, so as to improve the security and availability of the data.

3. The intelligent recording and analysis system for the entire structural test process according to claim 2 is characterized by: The number of cameras (1) is determined according to the stress characteristics of different test components (2), so as to be able to fully capture the changes of the components throughout the entire process. The cameras can be installed using a tripod or fixed on the surrounding walls. The shooting parameters are determined according to actual needs, and the resolution, frame rate, focal length, and viewing angle are determined so as to be able to accurately capture the cracks of the components. The supporting fill light condition can be a white light fill light, and the angle of the fill light needs to match the shooting angle of the camera to ensure that the fill light area is the area shot by the camera, so as to avoid uneven lighting or a situation where some areas have no light.

4. The intelligent recording and analysis system for the entire structural test process according to claim 3 is characterized by: In the video acquisition module, the time code synchronization method adopts the NTP network time protocol. After each camera (1) is connected to the network, the accurate time is obtained by configuring the NTP server. Then, the camera (1) sends a request to the server at regular intervals to obtain the latest time and calibrate it. On this basis, a trigger script is written in Python language to loop through the IP addresses or device identifiers of all cameras (1). Using the network communication library and the API of the camera (1), a command to start recording is sent to multiple cameras (1) at the same time. Based on the above-mentioned time synchronization method, video acquisition is performed, and after acquisition is completed, the videos of each camera (1) are spliced ​​and fused to represent the changes of components in the same time; first, the video is decomposed into frames, and then a scale-invariant feature transformation algorithm or an accelerated robust feature algorithm is used to extract features for each frame, and matching points between different frames are found. These frames are spliced ​​together by calculating the transformation matrix, and the spliced ​​frames are recombined into a video.

5. The intelligent recording and analysis system for the entire structural test process according to claim 4 is characterized by: In the parameter measurement module, the scaling factor α is defined as follows: Among them, d real is the actual physical size of the test component, d pixel is the pixel size of the corresponding component in the image, then the physical size of the crack is calculated as: D real =αD pixel Among them, D can represent length, width and area respectively; The crack image after model identification is converted into a binary image, with white set as the crack and black set as the background. The area of ​​the crack is calculated by counting the total number of white pixels in the crack area in the binary image; the crack length is calculated using the crack skeletonization algorithm, which is an iterative refinement algorithm or distance transformation method. The crack area is extracted as a single-pixel wide line and the total pixel length of the skeleton line is calculated; the crack width is calculated by the closest distance between the crack area and the skeleton; the crack distribution density is obtained by dividing the total area of ​​the crack by the total area of ​​the apparent image of the test component.

6. A method for intelligently recording and analyzing the entire process of a structural test, which uses the intelligent recording and analyzing system for the entire process of a structural test according to claim 5, characterized in that: The intelligent recording and analysis method for the entire structural test process includes a crack positioning method, a crack segmentation method, and a crack tracking method in sequence.

7. The method for intelligent recording and analysis of the entire structural test process according to claim 6, characterized in that: The crack locating method comprises the following steps: S11. Construct a crack dataset based on previously collected open-source crack images and experimental monitoring images. Perform geometric transformations, color transformations, and add noise to the crack images to expand the dataset and increase data diversity. S12. Annotate the dataset with bounding boxes, clearly identify the locations of cracks, indicate the crack categories and bounding box coordinates, and establish a crack detection dataset; S13. Improve the YOLOv8 crack detection model algorithm and construct a fused CFE module. By adding the lightweight and efficient FasterNet convolutional neural network, the model's resource consumption is reduced, shortening the detection runtime. Furthermore, an efficient multi-scale attention module based on EMA cross-space learning is introduced to batch reconstruct some channels, making the YOLOv8 model more targeted when extracting crack features. S14. Divide the dataset and train the model, using distributional focal loss and complete intersection-over-union loss as loss functions, and evaluate model performance using precision, recall, and average precision. S15. Based on actual needs, deploy the trained improved YOLOv8 crack detection model to the camera. Considering the computing power and storage resources of the device, appropriately compress and optimize the improved YOLOv8 crack detection model to ensure that the model can run efficiently.

8. The method for intelligent recording and analysis of the entire structural test process according to claim 7 is characterized in that: The crack segmentation method comprises the following steps: S21. Based on the above crack dataset, use data augmentation methods to ensure that the images have different crack shapes, sizes, directions, and different background environments; S22. Annotate the crack image at the pixel level, where the crack portion is the foreground and the rest is the background, to establish a crack segmentation dataset; S23. Improve the DeepLabV3+ model algorithm by replacing the original backbone feature extraction network with MobilenetV2 to reduce the number of model parameters and improve computing speed. In addition, add the ECANet cross-channel interactive attention module to make the model more focused on crack information and enhance crack segmentation performance. S24. Divide the dataset and train the model, using cross-entropy loss and Dice loss as loss functions. Evaluate model performance using category-average pixel accuracy and average intersection-over-union (IoU) accuracy. S25. Deploy the improved DeepLabV3+ crack segmentation model trained in S24 into the camera for real-time crack segmentation. At the same time, collect data feedback from the actual test process and continuously optimize the model based on this feedback to adapt to different crack segmentation environments and requirements.

9. The method for intelligent recording and analysis of the entire structural test process according to claim 8, characterized in that: The crack tracking method comprises the following steps: S31. Combining the motion information and appearance features of cracks, the DeepSORT method is used to achieve accurate association across time frames. After crack identification and feature extraction, a Kalman filter is used to establish a crack motion model. The crack position and size in the current frame are predicted based on the state of the previous frame, providing a preliminary association reference. Simultaneously, the crack appearance features extracted by the above module are embedded. During tracking, a matching matrix between cracks is constructed by combining the motion model distance predicted by the Kalman filter and the similarity of appearance features. The Hungarian algorithm is used to optimally allocate detection boxes and tracking trajectories to ensure that the same crack remains consistent in consecutive frames. S32. After the matching is completed, the status of the crack is dynamically updated, and the unmatched cracks are initialized as new trajectories. Finally, the tracking number of the crack and its evolution trajectory are output, and the crack extension path is presented in a visual manner. At the same time, time series data of the crack characteristics changing over time is generated.

10. The method for intelligent recording and analyzing the entire structural test process according to claim 9, characterized in that: In S11, the geometric transformation method is rotation, flipping, cropping, and scaling; the color transformation method is color adjustment, brightness adjustment, and contrast adjustment; and the noise addition method is Gaussian noise and salt and pepper noise; In S12, the Make Sense annotation tool is used to perform bounding box annotation and category annotation on the crack information in the dataset, clarifying the location of the cracks and indicating the crack categories and bounding box coordinates, so as to directly generate the label format required by the YOLO series algorithm and establish a crack detection dataset; In S13, the FasterNet network is integrated with the EMA attention mechanism to construct a CFE module and apply it to the backbone network and neck network of the model. FasterNet, as a lightweight and efficient convolutional neural network, is used to reduce the resource consumption of the model and shorten the running time required for detection. EMA, as an efficient multi-scale attention module for cross-space learning, can batch reconstruct some channels, making the model more targeted when extracting crack features. In S14, the crack detection dataset is divided into a training set and a test set according to a certain ratio, wherein the training set is used to train the model, and the test set is used to evaluate the performance of the model; Use distribution focus loss and complete intersection-over-union loss as loss functions; The distribution focus loss is described as follows: for a continuous value y∈R, R is a set of real numbers; discretize it into the integer interval [l,u], where l is the maximum integer not greater than y rounded down; u is the minimum integer not less than y rounded up; assuming that the predicted distribution probability p = {p0, p1, ..., p n }, corresponding to the discrete position {0,1,…,n}, the calculation formula of the distribution focus loss is: L DFL =-((u-y)log(p l )+(y-l)log(p u )) Among them, y is the target value, p l and p u is the predicted probability corresponding to the interval [l,u]; uy and yl are weights, reflecting the degree of deviation of the target value within the interval; Complete Intersection-over-Union loss assumes the predicted box is B p , the target box is B g , the overlapping area is Area, and the Euclidean distance between the center point of the prediction box and the target box is ρ(λ p ,λ g ), the diagonal length of the minimum closed box surrounding the prediction box and the target box is d, and the height and width of the prediction box and the target box are h respectively p , w p and h g , w g , the weight factor is α, which is used to balance the influence of the aspect ratio term and other terms. The calculation formula is: In S14, the model performance is evaluated using precision, recall and average precision. Precision refers to the ratio of the number of samples that the model detects as cracks (TP) to the total number of samples detected as cracks (TP+FP), expressed as: The recall rate refers to the ratio of the number of samples that the model detects as actually cracks, TP, to the number of samples that are actually cracks, TP+FN, and is expressed as: Average precision is the area under the precision-recall curve. The higher the value, the easier it is for the model to maintain high precision at high recall. The calculation formula is: TP represents the number of samples that are actually cracks and are correctly detected as cracks, FN represents the number of samples that are actually cracks but are incorrectly detected as background, and FP represents the number of samples that are actually background but are incorrectly detected as cracks. In said S15, after the model is iteratively trained and its performance meets the actual requirements, it is encapsulated using TorchScript or TensorFlowLite and deployed in the camera (1) for execution; In said S23, firstly, a lightweight neural network MobileNetV2 is used to replace the backbone feature extraction network Xception in the original DeepLabV3+ semantic segmentation model, so as to effectively reduce the number of parameters and balance the calculation accuracy and speed, so as to make the model more suitable for camera embedded devices; secondly, an attention module ECANet is added after the ASPP void space pyramid pooling module extracts features and before the low-order features are spliced, so that the model can better focus on the crack information of the test component (2) and suppress invalid and complex background features, thereby improving the accuracy of the model; In S24, the cross entropy loss L CE Measures the difference between the predicted probability distribution and the true distribution, calculated as: Where N is the total number of pixels in the image, C is the number of categories, and y i,c is the true label of the i-th pixel belonging to category c, is the predicted probability that the i-th pixel belongs to category c; Dice loss is based on the Dice similarity coefficient L Dice , which is used to measure the overlap between the prediction and the true segmentation. The formula is: Among them, p i is the predicted probability that the i-th pixel belongs to a crack, g i is the true label of the i-th pixel belonging to the crack; In S24, the category average pixel accuracy MPA is used to measure the proportion of pixels in each category that are correctly classified, and then averaged over all categories: Among them, TP c Indicates the number of pixels whose true label is category c and whose prediction is also category c, FN c Indicates the number of pixels whose true label is category c but is predicted to be other categories; The mean intersection over union (MIoU) measures the ratio of the intersection and union of the predicted result to the true segmentation, averaged over all categories: Among them, FP c Indicates the number of pixels whose predicted label is category c but the true label is not c; After the model achieves the desired performance on the crack segmentation dataset, the model size is reduced and the inference speed is improved through pruning and quantization methods, and then converted into ONNX format or TensorRT format suitable for deployment, which is packaged and deployed in the camera (1) for execution; In S31, the DeepSORT crack tracking method mainly refers to combining a simple and efficient Kalman filter motion model and deep feature measurement to perform real-time crack extension state tracking; Kalman filtering assumes that the crack state is a vector x k , the observed value is z k , then the change of crack state is described by the state transition equation: x k =F k x k-1 +B k u k +w k Among them, x k represents the state vector of the crack at time k, F k is the state transfer matrix, u k Control input vector, B k Control input matrix, w k is the process noise; The crack motion model uses a Kalman filter to model the geometric characteristics and position state of the crack. The state vector is defined as the center point position (x, y) of the crack, the size of the bounding box, and its changing speed. The Kalman filter predicts the possible position and size of the crack in the next frame based on the state of the current frame, providing a reference for association. The crack matching and association method is that in each frame of the image, the new crack detected needs to be matched with the crack tracked in the previous frame, which is achieved by combining distance measurement; The distance measurement uses motion model distance and appearance feature distance. The motion model distance calculates the Euclidean distance between the state predicted by Kalman filter and the actual detection result. The appearance feature distance uses the extracted feature embedding to calculate the cosine distance or Euclidean distance between the crack features of the current frame and the previous frame. The two distances are weighted and fused to construct a matching matrix. The Hungarian algorithm is then used to achieve the optimal assignment of crack identification results to tracking trajectories. This involves finding a minimum set of matching pairs on the matching matrix, such that each trajectory matches at most one detection box, and each detection box matches at most one trajectory. Unmatched detection boxes are initialized as new crack trajectories. In S32, the dynamic update method updates the position, size, and appearance characteristics of successfully matched crack trajectories to keep them consistent with the actual state of the cracks; for unmatched crack trajectories, the tracking is continued based on the prediction of the Kalman filter. If no new detection results are matched within a period of time, the crack is considered to have stopped expanding. After tracking is completed, the dynamic trajectory of the crack is visualized and output in the following ways: the crack number and trajectory line are displayed on the original image, the curves of the crack width and length changing with time are drawn, and an animation of the crack expansion is generated to demonstrate its dynamic evolution process.

Citation Information

Patent Citations

  • Method for identifying concrete cracks based on yolov3 deep learning model

    AU2020101011A4

  • Pavement crack detection system based on YOLOv8 semantic segmentation

    CN117237639A