Tower crane risk source detection transferable system based on multi-task learning
By applying a multi-task learning-based detection system on the tower crane, using drones and deep learning technology to simultaneously detect bolt loosening, cracks and rust in the tower crane, the problems of low efficiency and inability to multi-task detection in the existing technology are solved, and efficient and low-cost tower crane disease detection and early warning are achieved.
Patent Information
- Application Number
- CN202510301516.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is inefficient in safety detection of tower cranes and cannot detect multiple diseases at the same time, resulting in high risk and cost of manual inspection.
The tower crane risk source detection migration system based on multi-task learning is adopted, and image data is obtained through the drone equipped with a camera, combined with the multi-task deep learning visual model to detect bolt loosening, cracks and rust, and the detection results are visualized and early warning in the BIM model.
It realizes efficient and multi-task tower crane disease detection, reduces the risks and costs of manual inspection, and the system engineering is available, migratory, miniaturized, low-cost, and easy to commercially apply.
Smart Images

Figure CN119990777A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of construction machinery risk source detection and management at smart construction sites, and in particular to a transferable tower crane risk source detection system based on multi-task learning. Background Art
[0002] With the rapid implementation of smart construction site projects, the corresponding technologies are developing rapidly. However, the safety inspection of tower cranes is often done manually, which is inefficient and not conducive to personal safety. Therefore, the demand for machine inspection has emerged. Existing technologies often focus on algorithm implementation and accuracy improvement, and existing technologies often only focus on the detection of one disease of the tower crane, and cannot detect multiple diseases at one time. Summary of the invention
[0003] The present invention aims to address the deficiencies of the prior art and provides a transferable system for tower crane risk source detection based on multi-task learning. It performs multi-task detection from the perspective of multi-task visual learning and attaches importance to technology implementation, forming a set of engineering-usable, transferable, miniaturized, and low-cost tower crane disease detection platform. It also solves the problem of high cost of non-standard automated migration of traditional engineering intelligent detection solutions, is easier to commercialize, and has broad application potential.
[0004] In order to achieve the above object, the present invention adopts the following technical solutions:
[0005] A tower crane risk source detection transferable system based on multi-task learning, including a risk source perception module, a data transmission and calculation module, a BIM visualization module and a risk warning module; the risk source perception module is connected with the data transmission and calculation module, and the data transmission and calculation module is connected with the BIM visualization module and the risk warning module;
[0006] The risk source perception module includes a tower crane perception unit and an unmanned drone system; the data transmission and calculation module includes a data transmission unit and a data calculation unit; the BIM visualization module includes an image information transmission unit and a BIM model;
[0007] The risk source perception module uses a drone equipped with a movable camera to capture images along a preset trajectory, and obtains the coordinate information of the drone during shooting. The image is the input data set of the algorithm, and the coordinate information is used for synchronous positioning of the BIM model; the data transmission and calculation module is used to transmit image information over long distances. The image information is processed by the server of the multi-task deep learning visual model carried by the data calculation unit to obtain risk source information; the BIM visualization module is used to display the risk source information in the BIM model, realize the calibration of the risk source position in the BIM model and the visualization of the risk; the risk warning module is used to evaluate and classify the severity of risk points, automatically generate warnings, and notify relevant personnel.
[0008] The risk source perception module can sense loose bolts, cracks, and rust.
[0009] The tower crane sensing unit includes a positive angle sensing subunit and a negative angle sensing subunit. The positive angle sensing subunit senses the two side columns on the far side of the tower crane from the building, while the negative angle sensing subunit senses the two side columns on the near side of the tower crane from the building. The optical network camera carried by the drone is used for data collection.
[0010] The unmanned drone system includes a meteorological monitoring module, a network card module, an RTK module, a data preprocessing and calculation module, and a cabin cooling module.
[0011] The risk source perception module captures image data sets and stores the location information and angle information of the drone. The location information of the drone is represented by a three-dimensional Cartesian coordinate system, and the three-dimensional vector (x, y, z) is used to represent the position of the drone on the longitude x-axis, latitude y-axis, and altitude z-axis respectively. The angle information of the drone is described by Euler angles, and the three-dimensional vector (φ, θ, ψ) is used to represent the camera posture, where φ represents the roll angle, θ represents the pitch angle, and ψ represents the yaw angle.
[0012] The risk source perception module uses drones to obtain image information. The drone is an unmanned drone, and the drone battery is a wireless charging battery. The drone platform includes an airtight and automatically rechargeable cabin. The operation process of the drone platform is as follows: engineers use the built-in function of the drone to manually fly and set the route and observation point. At this time, the drone records six variables including position and rotation angle. When unmanned, the drone connects to the Internet to check the weather, battery and satellite status. After takeoff, it flies to the set point to take pictures, completes the work in sequence, and then returns to the return point to automatically charge and close the cabin; the drone obtains the image interval, and obtains the first image collected by the risk source perception module at each first distance L1 interval. H is the standard section height of the tower crane.
[0013] The data transmission unit is used to remotely transmit images with the cabin platform through the Wi-Fi data transmission system installed in the drone. The cabin platform uses a network card module with a built-in SIM card to transmit the pre-processed data to the server. The data computing unit uses a server equipped with two A6000 parallel computing GPUs for the deployment and data calculation of multi-task deep learning vision models.
[0014] The image information transmission unit is used to transmit the risk source image information and the risk source location information to the BIM model; the BIM model is used for engineers to view the risk source data information in real time, including the location of loose bolts, cracks, and corrosion areas.
[0015] The specific process of the risk source perception module taking image data sets, storing the location information of the drone, and the image information transmission unit transmitting information to the BIM model is as follows:
[0016] Get the camera position parameters. When shooting the data set through the tower crane sensing unit, get the position information of the three parameters of the camera. According to the longitude, latitude and altitude of the drone, the coordinates are generated by the GPS built into the drone. Store this information in a database or file for subsequent use. Add a marker. In the BIM model, according to the camera position parameter information, add a marker at the corresponding coordinate to indicate that there is a camera or camera-related information at that location. Interactive design. When the engineer clicks on the marker, a pop-up window displays the image taken by the camera.
[0017] Adding markers requires coordinate transformation. The GPS data collected by the drone is based on the WGS84 coordinate system, and the BIM model relies on the local coordinate system. The GPS data of a feature point on the tower crane or around the tower crane and the coordinate information in the BIM model are obtained respectively. The coordinate transformation matrix is obtained by calculation, and then the camera coordinate information is converted into BIM model information according to the matrix.
[0018] The risk warning module uses a deep learning warning model to evaluate and classify the severity of risk points, generate warning levels, and then push warning information via SMS, email or application to ensure that relevant personnel are notified in time. The specific process is as follows:
[0019] For the target detection corresponding to the crack, the accuracy is used as an indicator, and the first threshold M1 is manually set. The image is scanned by the target detection algorithm to obtain the crack prediction box Bounding Box. The accuracy is used as the first probability P1. When P1≥M1, the first warning prompt is sent; for the rusted part of the tower crane, the second threshold M2 is manually set. The Deeplabv3+ algorithm is first used to process the image to obtain the preliminary recognition result of the rusted area. On this basis, the connected area analysis is performed on the identified rusted area, and the ratio of the number of pixels in the largest connected rusted area to the total number of pixels in the image is calculated as the second probability P2. When P2≥M2, the second warning prompt is sent; for the loose part of the tower crane bolt, the third threshold M3 is manually set. The algorithm identifies a point on the upper edge of the screw in the image as the first key point, a point on the lower edge of the tower crane socket as the second key point, and a point on the upper edge of the tower crane socket as the third key point. The number of straight line pixels between the first key point and the second key point is the first spacing L1, and the number of pixels between the second key point and the third key point is the second spacing L2. When When the alarm is triggered, a third warning prompt is sent.
[0020] Multi-task deep learning includes the following algorithms: CNN algorithm, ResNet algorithm, YOLOv8 algorithm, Faster-RCNN algorithm, Deeplabv3+ algorithm, YOLOv8-Pose algorithm, image processing algorithm, which can simultaneously perform target detection, semantic segmentation, and key point detection;
[0021] The structure of the multi-task deep learning vision model includes: region of interest extraction layer, feature extraction layer, object detection branch, semantic segmentation branch, and key point detection branch;
[0022] The region of interest extraction layer uses the Faster-RCNN algorithm to determine the regions of interest for corrosion, cracks, and bolt loosening. Other algorithms are performed on the regions of interest to narrow the detection range;
[0023] Feature extraction layer, using pre-trained CNN and ResNet algorithms to extract features from images. Different branches are added on top to handle different tasks. Each branch includes a specific head to output task-specific results.
[0024] The target detection branch uses the YOLOv8 algorithm to perform crack detection;
[0025] The semantic segmentation branch uses the Deeplabv3+ algorithm for rust area detection;
[0026] The key point detection branch uses the YOLOv8-Pose algorithm for bolt loosening detection;
[0027] The loss functions are designed according to each branch. The target detection branch adopts the combined loss function used in Faster-RCNN or YOLOv8, including classification loss and positioning loss. The semantic segmentation branch uses the pixel-level cross entropy loss function to measure the difference between the category distribution of each pixel predicted by the model and the true label. The key point detection branch uses the key point positioning error as the loss function.
[0028] The loss function formula is:
[0029] Total Loss=λ1·Detection Loss+λ2·Segmentation Loss+λ3·KeypointLoss;
[0030] Among them, λ1, λ2, and λ3 are the weights of the loss function of each task, which are adjusted according to the importance of the task;
[0031] Specifically,
[0032] Detection Loss = classification loss + positioning loss;
[0033] The classification loss is expressed as
[0034] N obj is the number of bounding boxes marked as objects; p i is the probability that the predicted bounding box i belongs to the correct category; It is the indicator function of the true category, 1 represents the correct category, and 0 represents other categories;
[0035] The localization loss is expressed as
[0036] λ coord is the weight of the positioning loss; (x i ,y i ,w i ,h i ) are the center coordinates and width and height of the predicted bounding box; are the center coordinates and width and height of the real bounding box;
[0037]
[0038] N is the number of pixels, C is the number of categories, y ij is the indicator function that pixel i belongs to category j in the true label, is the probability predicted by the model that pixel i belongs to category j;
[0039]
[0040] N kp is the number of key points, is the position of key point i predicted by the model, p i is the position of the true key point i.
[0041] The training process of the multi-task deep learning vision model is as follows: task-specific training only updates the weights of specific branches, using the parameter freezing method to fix parameters that do not need to be updated, not participating in back propagation and gradient updates, and assigning different optimizers to different tasks for weight updates;
[0042] The method for deploying a multi-task deep learning vision model in a server includes: converting a trained deep learning model into a deployment format, setting up a model deployment service on the server side, receiving and preprocessing input data from a client, performing model inference through an inference service, and returning a prediction result; optimizing server hardware resources and network connections, and monitoring model performance and server load in real time;
[0043] The deployment process uses ONNX for model conversion to convert model parameters and weights into a format supported by the C++ language, which is recorded as the middleware model. The structured pruning channel pruning method that can retain the complete deep learning model structure is used to count the absolute value of the weight of each convolutional layer followed by the BN layer as the scaling factor of the BN layer. The pruning weight threshold is determined according to the array composed of the scaling factors of all BN layers, which is recorded as the first threshold thre_0. During the pruning process, all channels with weight values lower than the threshold are pruned by setting the weights in these channels to zero to achieve sparseness of the model.
[0044] The beneficial effects of the present invention are as follows: the present invention performs multi-task detection from the perspective of multi-task visual learning, forming a set of engineering-usable, transferable, miniaturized, low-cost tower crane disease detection system, which solves the problem of high cost of non-standard automation migration of traditional engineering intelligent detection solutions, is easier to commercialize, and has broad application potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic diagram of the overall structure of the present invention;
[0046] Figure 2 This is a schematic diagram of the architecture of the risk source perception module in the present invention;
[0047] Figure 3 A schematic diagram of the architecture of the data transmission and calculation module in the present invention;
[0048] Figure 4 This is a schematic diagram of the architecture of the BIM visualization module in the present invention;
[0049] Figure 5 This is a schematic diagram of an unattended drone system in the present invention;
[0050] Figure 6 This is a schematic diagram of the overall process steps for detecting risk sources in the present invention;
[0051] Figure 7 This is a schematic diagram of the steps of the server deployment method when detecting risk sources in the present invention;
[0052] Figure 8 This is a schematic diagram of the migration steps when detecting risk sources in the present invention;
[0053] Fig. 9 This is a schematic diagram of the image processing flow when detecting risk sources in the present invention;
[0054] Fig.10 This is a schematic diagram of the backbone of the multi-task learning algorithm for detecting risk sources in the present invention;
[0055] In the figure: 1-risk source perception module; 2-data transmission and calculation module; 3-BIM visualization module; 4-risk warning module;
[0056] 11-Tower crane sensing unit; 12-Unmanned aerial vehicle system;
[0057] 21-data transmission unit; 22-data calculation unit;
[0058] 31-image information transmission unit; 32-BIM model;
[0059] The following is a detailed description of the embodiments of the present invention with reference to the accompanying drawings. DETAILED DESCRIPTION
[0060] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The embodiments are only used to explain the present invention and are not used to limit the scope of the present invention. The present invention is described in more detail by way of example with reference to the accompanying drawings in the following paragraphs. The advantages and features of the present invention will become clearer according to the following description. It should be noted that the accompanying drawings are all in a very simplified form and are not in precise proportions, and are only used to facilitate and clearly assist in explaining the purpose of the embodiments of the present invention.
[0061] It should be noted that when a component is referred to as being "fixed to" another component, it may be directly on the other component or there may also be a component centered. When a component is considered to be "connected to" another component, it may be directly connected to the other component or there may also be a component centered. When a component is considered to be "set on" another component, it may be directly set on the other component or there may also be a component centered. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0063] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:
[0064] A transferable system for tower crane risk source detection based on multi-task learning, such as Figures 1 to 10As shown, it includes a risk source perception module 1, a data transmission and calculation module 2, a BIM visualization module 3 and a risk warning module 4; the risk source perception module 1 is connected to the data transmission and calculation module 2, and the data transmission and calculation module 2 is connected to the BIM visualization module 3 and the risk warning module 4;
[0065] The risk source perception module 1 includes a tower crane perception unit 11 and an unmanned drone system 12; the data transmission and calculation module 2 includes a data transmission unit 21 and a data calculation unit 22; the BIM visualization module 3 includes an image information transmission unit 31 and a BIM model 32;
[0066] The risk source perception module 1 uses a drone equipped with a movable camera to capture images along a preset trajectory, and simultaneously obtains the coordinate information of the drone during shooting. The image is the input data set of the algorithm, and the coordinate information is used for synchronous positioning of the BIM model 32; the data transmission and calculation module 2 is used to transmit image information over long distances, and the image information is processed by a server of a multi-task deep learning visual model carried by the data calculation unit 22 to obtain risk source information; the BIM visualization module 3 is used to display the risk source information in the BIM model 32, to achieve risk source position calibration and risk visualization in the BIM model 32; the risk warning module 4 is used to evaluate and classify the severity of risk points, automatically generate warnings, and notify relevant personnel.
[0067] The scope of risk source perception module 1 includes loose bolts, cracks, and rust;
[0068] The tower crane sensing unit 11 includes a positive angle sensing subunit and a negative angle sensing subunit. The positive angle sensing subunit senses two side columns on the far side of the tower crane from the building, and the negative angle sensing subunit senses two side columns on the near side of the tower crane from the building. The optical network camera carried by the drone is used for data collection.
[0069] The unmanned aerial vehicle system 12 includes a meteorological monitoring module, a network card module, an RTK module, a data preprocessing and calculation module, and a cabin cooling module.
[0070] The risk source perception module 1 captures an image data set and stores the position information and angle information of the drone. The position information of the drone is represented by a three-dimensional Cartesian coordinate system, and the three-dimensional vector (x, y, z) is used to represent the position of the drone on the longitude x-axis, latitude y-axis, and altitude z-axis respectively. The angle information of the drone is described by Euler angles, which include roll angle, pitch angle, and yaw angle, and are used to represent the rotation angle of the drone around the three axes in the fixed coordinate system; the roll angle represents the rotation angle around the longitudinal axis of the aircraft; the pitch angle represents the rotation angle around the transverse axis of the aircraft; the yaw angle represents the rotation angle around the vertical axis of the aircraft; the camera posture is represented by a three-dimensional vector (φ, θ, ψ), φ represents the roll angle, θ represents the pitch angle, and ψ represents the yaw angle;
[0071] The risk source perception module 1 uses a drone to obtain image information. The drone is an unmanned drone. The drone battery is a wireless charging battery. The drone platform includes an airtight and automatically rechargeable cabin. The operation process of the drone platform is as follows: engineers use the built-in function of the drone to manually fly and set the route and observation point. At this time, the drone records six variables including position and rotation angle. When unmanned, the drone connects to the Internet to check the weather, battery and satellite status. After taking off, it flies to the set point to take pictures, completes the work in sequence, and then returns to the return point to automatically charge and close the cabin; the drone obtains the image interval, and obtains the first image collected by the risk source perception module 1 at each first distance L1 interval. H is the standard section height of the tower crane.
[0072] The data transmission unit 21 is used to perform remote image transmission with the cabin platform through the Wi-Fi data transmission system installed in the drone. The cabin platform uses a network card module with a built-in SIM card to transmit the pre-processed data to the server. The drone can either return to the cabin and transmit data to the cabin, or use the Wi-Fi system for remote image transmission; the data computing unit 22 uses a server equipped with two A6000 parallel computing GPUs for the deployment and data calculation of multi-task deep learning vision models.
[0073] The data transmission unit 21 adopts a Wi-Fi data transmission system for long-distance data transmission, provides 2.4GHz and 5.8GHz dual-band access, and achieves a maximum transmission distance of 4 kilometers. It has an embedded SIM card for accessing the 5G network and a high-performance processor for processing data.
[0074] The data computing unit 22 includes a high-performance server for deploying visual models and real-time computing.
[0075] The image information transmission unit 31 is used to transmit the risk source image information and the risk source location information corresponding to the BIM model 32; the BIM model 32 is used for engineers to view the risk source data information in real time, including the bolt loosening location, crack location, and corrosion area range.
[0076] It also provides historical data comparison of risk sources, allowing users to track the development trend of risk sources and take timely preventive and maintenance measures.
[0077] The specific process of the risk source perception module 1 taking an image data set, storing the location information of the drone, and the image information transmission unit 31 transmitting the information to the BIM model 32 is as follows:
[0078] Obtain the camera position parameters. When shooting the data set through the tower crane sensing unit 11, obtain the position information of the three parameters of the camera. According to the position of the drone in longitude, latitude and altitude (x, y, z), the coordinates are generated by the GPS built into the drone. Store this information in a database or file for subsequent use; add a marker. In the BIM model 32, according to the camera position parameter information, add a marker at the corresponding coordinate to indicate that there is a camera or information related to the camera at the location; interactive design. When the engineer clicks on the marker, a pop-up window displays the image taken by the camera;
[0079] The addition of markers requires coordinate transformation. The GPS data collected by the drone is based on the WGS84 coordinate system, and the BIM model 32 relies on the local coordinate system. The GPS data of a feature point on the tower crane or around the tower crane and the coordinate information in the BIM model 32 are obtained respectively. The coordinate transformation matrix is obtained by calculation, and then the camera coordinate information is converted into the information of the BIM model 32 according to the matrix.
[0080] Risk warning module 4 uses a deep learning warning model to evaluate and classify the severity of risk points, generate warning levels, and then push warning information via SMS, email or application to ensure that relevant personnel are notified in time. The specific process is as follows:
[0081] For the target detection corresponding to the crack, the accuracy is used as an indicator, and the first threshold M1 is manually set. The image is scanned by the target detection algorithm to obtain the crack prediction box Bounding Box. The accuracy is used as the first probability P1. When P1≥M1, the first warning prompt is sent; for the rusted part of the tower crane, the second threshold M2 is manually set. The Deeplabv3+ algorithm is first used to process the image to obtain the preliminary recognition result of the rusted area. On this basis, the connected area analysis is performed on the identified rusted area, and the ratio of the number of pixels in the largest connected rusted area to the total number of pixels in the image is calculated as the second probability P2. When P2≥M2, the second warning prompt is sent; for the loose part of the tower crane bolt, the third threshold M3 is manually set. The algorithm identifies a point on the upper edge of the screw in the image as the first key point, a point on the lower edge of the tower crane socket as the second key point, and a point on the upper edge of the tower crane socket as the third key point. The number of straight line pixels between the first key point and the second key point is the first spacing L1, and the number of pixels between the second key point and the third key point is the second spacing L2. When When the alarm is triggered, a third warning prompt is sent.
[0082] Multi-task deep learning includes the following algorithms: CNN algorithm, ResNet algorithm, YOLOv8 algorithm, Faster-RCNN algorithm, Deeplabv3+ algorithm, YOLOv8-Pose algorithm, image processing algorithm, which can simultaneously perform target detection, semantic segmentation, and key point detection;
[0083] The structure of the multi-task deep learning vision model includes: region of interest extraction layer, feature extraction layer, object detection branch, semantic segmentation branch, and key point detection branch;
[0084] The region of interest extraction layer uses the Faster-RCNN algorithm to determine the regions of interest for corrosion, cracks, and bolt loosening. Other algorithms are performed on the regions of interest to narrow the detection range;
[0085] Feature extraction layer, using pre-trained CNN and ResNet algorithms to extract features from images. Different branches are added on top to handle different tasks. Each branch includes a specific head to output task-specific results.
[0086] The target detection branch uses the YOLOv8 algorithm to perform crack detection;
[0087] The semantic segmentation branch uses the Deeplabv3+ algorithm for rust area detection;
[0088] The key point detection branch uses the YOLOv8-Pose algorithm for bolt loosening detection;
[0089] The loss functions are designed according to each branch. The target detection branch adopts the combined loss function used in Faster-RCNN or YOLOv8, including classification loss and positioning loss. The semantic segmentation branch uses the pixel-level cross entropy loss function to measure the difference between the category distribution of each pixel predicted by the model and the true label. The key point detection branch uses the key point positioning error as the loss function.
[0090] The loss function formula is:
[0091] Total Loss=λ1·Detection Loss+λ2·Segmentation Loss+λ3·KeypointLoss;
[0092] Among them, λ1, λ2, and λ3 are the weights of the loss function of each task, which are adjusted according to the importance of the task;
[0093] Specifically,
[0094] Detection Loss = classification loss + positioning loss;
[0095]
[0096] N obj is the number of bounding boxes marked as objects; p i is the probability that the predicted bounding box i belongs to the correct category; p i * is the indicator function of the true category, 1 represents the correct category, and 0 represents other categories;
[0097]
[0098] λ coord is the weight of the positioning loss; (x i ,y i ,w i ,h i ) are the center coordinates and width and height of the predicted bounding box; are the center coordinates and width and height of the real bounding box;
[0099]
[0100] N is the number of pixels, C is the number of categories, y ij is the indicator function that pixel i belongs to category j in the true label, is the probability predicted by the model that pixel i belongs to category j;
[0101]
[0102] N kp is the number of key points, is the position of key point i predicted by the model, p i is the position of the true key point i.
[0103] The training process of the multi-task deep learning vision model is as follows: task-specific training only updates the weights of specific branches, using the parameter freezing method to fix parameters that do not need to be updated, not participating in back propagation and gradient updates, and assigning different optimizers to different tasks for weight updates;
[0104] The method for deploying a multi-task deep learning vision model in a server includes: converting a trained deep learning model into a deployment format, setting up a model deployment service on the server side, receiving and preprocessing input data from a client, performing model inference through an inference service, and returning a prediction result; optimizing server hardware resources and network connections, and monitoring model performance and server load in real time;
[0105] The deployment process uses ONNX for model conversion to convert model parameters and weights into a format supported by the C++ language, which is recorded as the middleware model. The structured pruning channel pruning method that can retain the complete deep learning model structure is used to count the absolute value of the weight of each convolutional layer followed by the BN layer as the scaling factor of the BN layer. The pruning weight threshold is determined according to the array composed of the scaling factors of all BN layers, which is recorded as the first threshold thre_0. During the pruning process, all channels with weight values lower than the threshold are pruned by setting the weights in these channels to zero to achieve sparseness of the model. Specific embodiment 1:
[0107] Figure 1 The present invention is a structural diagram of a tower crane risk source detection transferable system based on multi-task learning, including: a risk source perception module 1, a data transmission and calculation module 2, a BIM visualization module 3 and a risk warning module 4.
[0108] The risk source perception module 1 is connected to the data transmission and calculation module 2, and is used for the movable camera carried by the drone to locate the six key coordinates captured according to the preset trajectory and obtain the image and the coordinate information of the camera during shooting, and transmit and process the key information through the data transmission and calculation module 2;
[0109] The data transmission and calculation module 2 is connected to the BIM visualization module 3, and is used to transmit the data and information processed by the server and display them in the BIM model 32;
[0110] The risk warning module 4 is connected to the data transmission and calculation module 2 , and the data processed and generated by the data transmission and calculation module 2 is used to issue a warning through the risk warning module 4 .
[0111] In this embodiment, the possible risks of the construction tower crane can be automatically determined through image data, and then an early warning can be issued and displayed in the BIM model 32 for visualization, thereby effectively checking for construction hazards, reducing manual labor, and improving monitoring efficiency.
[0112] Figure 2 Schematic diagram of the structure of the risk source perception module 1, which includes:
[0113] The tower crane sensing unit 11 includes a positive angle sensing subunit and a negative angle sensing subunit. The positive angle sensing subunit senses two side columns on the far side of the tower crane from the building, and the negative angle sensing subunit senses two side columns on the near side of the tower crane from the building. The optical network camera carried by the drone is used for data collection.
[0114] The unmanned aerial vehicle system 12 includes: a weather monitoring module, a network card module, an RTK module, a data preprocessing calculation module, and a cabin cooling module. Figure 5 shown.
[0115] Figure 3 2 is a schematic diagram of the structure of the data transmission and calculation module 2, which includes:
[0116] A data transmission unit 21, used for transmitting data via a Wi-Fi data transmission system installed in the drone;
[0117] The data computing unit 22 uses a server equipped with two A6000 parallel computing GPUs for the deployment and data computing of multi-task deep learning vision models.
[0118] Figure 4 Schematic diagram of the structure of BIM visualization module 3, BIM visualization module 3 includes:
[0119] The image information transmission unit 31 is used to transmit the risk source image information and the risk source location information to the BIM model 32.
[0120] BIM model 32 is used for engineers to view risk source data information in real time, including bolt loosening location, crack location, corrosion area range, etc. In addition, the system also provides historical data comparison of risk sources, allowing users to track the development trend of risk sources and take timely preventive and maintenance measures.
[0121] The risk warning module uses a deep learning warning model to evaluate and classify the severity of risk points, automatically generate warning levels, and then push warning information via SMS, email or application to ensure that relevant personnel are notified in time.
[0122] In order to make the purpose, technical solutions and advantages of the present invention more clear, Figure 6 The present invention is described in further detail.
[0123] like Figure 6 As shown in the figure, a transferable system for tower crane risk source detection based on multi-task learning specifically detects risk sources in the following steps:
[0124] Step S1: Acquire a first image collected by the risk source perception module at a first distance L1 interval. H is the standard section height of the tower crane; a multi-task deep learning visual model is trained and established, which is the first model, and each static image is identified using the first model to obtain the target risk source in each static image;
[0125] Specifically, the first image is obtained by the UAV system, and the first image is an image obtained from different angles, different distances, different lighting, different weather conditions and different construction progresses. In order to improve the robustness of the training results, the original data is expanded using random data enhancement, and the expansion method includes but is not limited to the use of "random flip", "random horizontal movement", "random rotation", "random contrast enhancement", and "random brightness enhancement"; then the static image is labeled as "Crack", "Corrosion", and "Loose"; the first model preferably uses pre-trained CNN and ResNet as feature extraction layers, and adds different branches on top to process different tasks. Each branch includes a specific head for outputting task-specific results, and adds a classification head, a detection head, a segmentation head, a posture estimation head, and an object key point detection head respectively; the processed data is trained through pre-training weights, and the set experimental hyperparameters include learning rate, Batch size, Epochs, Momentum, and image size. The learning rate in this embodiment is 0.01, and Batch The value of size is 32, the value of Epochs is 300, the value of Momentum is 0.937, the value of image size is 640×640, the confidence threshold is 0.5, the IOU non-maximum suppression threshold is 0.25, data enhancement in the inference stage is enabled, FP16 half-precision inference is disabled, and the detection time for each image is about 30ms.
[0126] Further, such as Fig. 9 , Fig.10As shown, the first model structure includes: a region of interest extraction layer, which uses the Faster-RCNN algorithm to determine the region of interest ROI of corrosion, cracks and bolt loosening as the second image, and other algorithms are performed on the second image; a feature extraction layer, which uses CNN and ResNet algorithms for image feature extraction; a target detection branch, which uses the YOLOv8 algorithm for crack detection; a semantic segmentation branch, which uses the Deeplabv3+ algorithm for rust area detection; a key point detection branch, which uses the YOLOv8-Pose algorithm for bolt loosening detection; using the shared feature extraction layer in the multi-task deep learning visual model, multiple tasks share information on the underlying feature representation; using the trained model to predict unlabeled data, and using the prediction results as pseudo labels in the training process; during the training process, the model simultaneously learns multiple tasks and self-supervised learning tasks to improve the model's ability to represent and generalize data; during training, specific task training only updates specific branch weights, and uses the parameter freezing method to fix the parameters that do not need to be updated, and does not participate in back propagation and gradient update, and assigns different optimizers to different tasks for weight update;
[0127] Specifically, the loss function is designed according to each branch. The target detection branch adopts a combined loss function, including positioning loss and classification loss. The classification loss uses cross entropy loss to measure the accuracy of the predicted category, and the mean square error loss is used to measure the positioning accuracy of the position and size of the bounding box. The semantic segmentation branch uses a pixel-level cross entropy loss function to measure the difference between the category distribution of each pixel predicted by the model and the true label. The key point detection branch uses the key point positioning error as the loss function.
[0128] The loss function formula is:
[0129] Total Loss=λ1·Detection Loss+λ2·Segmentation Loss+λ3·KeypointLoss;
[0130] Among them, λ1, λ2, and λ3 are the weights of the loss function of each task, which are adjusted according to the importance of the task;
[0131] Specifically,
[0132] Detection Loss = classification loss + positioning loss;
[0133] The classification loss is expressed as
[0134] N obj is the number of bounding boxes marked as objects; p i is the probability that the predicted bounding box i belongs to the correct category; It is the indicator function of the true category, 1 represents the correct category, and 0 represents other categories;
[0135] The localization loss is expressed as
[0136] λ coord is the weight of the positioning loss; (x i ,y i ,w i ,h i ) are the center coordinates and width and height of the predicted bounding box; are the center coordinates and width and height of the real bounding box;
[0137]
[0138] N is the number of pixels, C is the number of categories, y ij is the indicator function that pixel i belongs to category j in the true label, is the probability predicted by the model that pixel i belongs to category j;
[0139]
[0140] N kp is the number of key points, is the position of key point i predicted by the model, p i is the position of the true key point i.
[0141] Step S2, using the drone to enter the trajectory and shooting points;
[0142] Specifically, the UAV is a UAV using an unmanned cabin, and the cabin includes: a weather monitoring module, a network card module, an RTK module, a data preprocessing calculation module, and a cabin cooling module, such as Figure 5 As shown. The drone can remotely formulate flight plans, automatically execute tasks, free up complicated labor, and support rapid deployment. The drone can cover a radius of four kilometers and can work at -35℃-50℃, covering most civil engineering use scenarios. The meteorological monitoring module is used to monitor the current weather and prepare for flight. The network card module cooperates with the SIM card to realize data transmission. The RTK model is used to provide the drone with accurate geographic information data and initial calibration. The data preprocessing calculation module is used to preprocess media files and reduce the size of data, such as media data compression. The cabin cooling module uses two industrial air conditioners to quickly reduce the battery temperature.
[0143] The operation process of the drone platform is as follows: engineers use the built-in function of the drone to manually fly to set the route and observation points, and record six variables of the drone's position and rotation angle; use a three-dimensional vector (x, y, z) to represent the drone's position on the x, y, and z axes, namely longitude, latitude, and altitude; the camera posture can be represented by a three-dimensional vector (φ, θ, ψ), where φ represents the roll angle, θ represents the pitch angle, and ψ represents the yaw angle.
[0144] Furthermore, when unattended, the drone connects to the Internet to check weather, battery and satellite status, flies and records video according to preset tracks and waypoints after takeoff, and completes the work in sequence;
[0145] Step S3, the drone takes off, cruises, senses and shoots image data, and obtains location information during shooting;
[0146] Specifically, this embodiment takes cracks, corrosion, and loosening of connecting bolts of the tower crane as monitoring targets, and the distance between monitoring points is the first distance.
[0147] Step S4, performing data transmission of up to 4 km via the data transmission system Wi-Fi;
[0148] Specifically, the Wi-Fi data transmission system installed in the drone is used to transmit images remotely with the cabin platform. The cabin platform uses a network card module with a built-in SIM card to transmit pre-processed data to the server. The drone can either return to the cabin and transmit data to the cabin, or use the Wi-Fi system for remote image transmission.
[0149] Step S5, using a server equipped with two A6000 parallel computing GPUs to deploy, infer, and back up data of the first model, such as Figure 7 As shown;
[0150] Specifically, the first model reasoning step is:
[0151] Step S501, using ONNX to perform model conversion to convert model parameters and weights into a format supported by the C++ language, recorded as a middleware model;
[0152] Step S502, a model lightweight pruning operation is performed, using a structured pruning channel pruning method that can retain the complete deep learning model structure, and the absolute value of the weight of each convolutional layer followed by the BN layer, that is, the gamma parameter of the BN layer, is counted as a scaling factor of the BN layer;
[0153] Step S503, determine the pruning weight threshold according to the array composed of the scaling coefficients of all BN layers, denoted as the first threshold thre_0. During the pruning process, all channels with weight values lower than the threshold will be pruned by setting the weights in these channels to zero, thereby achieving sparseness of the model.
[0154] Step S6, BIM model 32 display and risk warning prompt, including bolt loosening location, crack location, corrosion area range, etc.; Step S6 includes two sub-steps of BIM model 32 display and risk warning prompt, which are performed in parallel;
[0155] Step S61, BIM model 32 display, includes the following steps:
[0156] Step S611, obtain the camera position parameters. When shooting the data set through the risk source perception unit 1, obtain the position information of the three parameters of the camera. According to the position representation (x, y, z) of the drone in longitude, latitude and altitude, the coordinates are generated by the GPS built into the drone. Store this information in a database or file for subsequent use. In this embodiment, the GPS data collected by the drone or the end of the robotic arm is based on the WGS84 coordinate system, and BIM relies on the local coordinate system to obtain the GPS data of a feature point on the tower crane or around the tower crane and the coordinate information in the BIM model, and obtain the coordinate conversion matrix by calculation. Then, the camera coordinate information can be converted into the information of the BIM model point according to the matrix.
[0157] Step S612, adding a marker. In the BIM model, according to the camera position parameter information, a marker is added at the corresponding coordinates to indicate that there is a camera or camera-related information at the location;
[0158] Step S613, interactive design, when the engineer clicks on the marker, a window will pop up to display the picture taken by the camera.
[0159] Step S62, risk warning prompt, includes the following steps:
[0160] Step S621, for target detection corresponding to the crack, the accuracy is used as an indicator, and the first threshold M1 is manually set. The image is scanned by the target detection algorithm to obtain the prediction box Bounding Box, and its accuracy is used as the first probability P1. When P1≥M1, the first warning prompt is sent;
[0161] Step S622: For the corroded part of the tower crane, manually set the second threshold M2, first use the Deeplabv3+ algorithm to process the image, obtain the preliminary recognition result of the corroded area, and then perform connected area analysis on the identified corroded area, and calculate the ratio of the number of pixels in the largest connected corroded area to the number of all pixels in the image as the second probability P2. When P2≥M2, send the second warning prompt;
[0162] Step S623, for the loose part of the tower crane bolt, manually set the third threshold M3, the algorithm identifies a point on the upper edge of the screw in the image as the first key point, a point on the lower edge of the tower crane socket as the second key point, and a point on the upper edge of the tower crane socket as the third key point. The number of straight line pixels between the first key point and the second key point is the first spacing L1, and the number of pixels between the second key point and the third key point is the second spacing L2; when When the third warning prompt is sent;
[0163] Step S7, the drone automatically returns according to the built-in return point, and the following steps are performed synchronously at the end of the return: the cabin is closed, the drone is quickly charged, the cooling module is turned on, the data preprocessing calculation module preprocesses the compressed media files and transmits the data to the server through the network card module.
[0164] Step S8, migration to other projects, other risk sources, recorded as the second and third tasks, etc. Figure 8 As shown;
[0165] Specifically, using this system to expand usage scenarios includes the following steps:
[0166] Step S801, collect new risk source data, annotate and directly train it in the first model. If the first model cannot meet the requirements, redeploy the second model, and set the second model according to actual needs;
[0167] Step S802, the route and waypoints are manually re-set and entered as the second route and waypoints;
[0168] Step S803, cabin adaptation, including RTK debugging, signal debugging, etc.;
[0169] Step S804, second task test;
[0170] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as various improvements are made using the method concept and technical solution of the present invention, or are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A transferable system for tower crane risk source detection based on multi-task learning, characterized in that: It comprises a risk source perception module (1), a data transmission and calculation module (2), a BIM visualization module (3) and a risk warning module (4); the risk source perception module (1) is connected to the data transmission and calculation module (2), and the data transmission and calculation module (2) is connected to the BIM visualization module (3) and the risk warning module (4); The risk source perception module (1) includes a tower crane perception unit (11) and an unmanned aerial vehicle system (12); the data transmission and calculation module (2) includes a data transmission unit (21) and a data calculation unit (22); the BIM visualization module (3) includes an image information transmission unit (31) and a BIM model (32); The risk source perception module (1) uses a drone equipped with a movable camera to capture images according to a preset trajectory and obtains coordinate information of the drone during the capture. The image is the input data set of the algorithm and the coordinate information is used for synchronous positioning of the BIM model (32); The data transmission and calculation module (2) is used to transmit image information over long distances. The image information is processed by a server of a multi-task deep learning visual model installed in a data calculation unit (22) to obtain risk source information. The BIM visualization module (3) is used to display the risk source information in the BIM model (32), thereby realizing the location calibration of the risk source and the visualization of the risk in the BIM model (32); The risk warning module (4) is used to evaluate and classify the severity of risk points, automatically generate warnings, and notify relevant personnel.
2. According to the multi-task learning-based tower crane risk source detection transferable system according to claim 1, it is characterized in that: The risk source perception module (1) senses loose bolts, cracks, and rust; The tower crane sensing unit (11) comprises a positive angle sensing subunit and a negative angle sensing subunit. The positive angle sensing subunit senses two side columns on the far side of the tower crane from the building, and the negative angle sensing subunit senses two side columns on the near side of the tower crane from the building. The optical network camera mounted on the drone is used to collect data. The unmanned aerial vehicle system (12) comprises a meteorological monitoring module, a network card module, an RTK module, a data preprocessing calculation module, and a cabin cooling module.
3. According to the multi-task learning-based tower crane risk source detection transferable system according to claim 2, it is characterized in that: The risk source perception module (1) captures an image data set and stores the position information and angle information of the drone. The position information of the drone is represented by a three-dimensional Cartesian coordinate system, and the three-dimensional vector (x, y, z) is used to represent the position of the drone on the longitude x-axis, latitude y-axis, and altitude z-axis respectively. The angle information of the drone is described by Euler angles, and the three-dimensional vector (φ, θ, ψ) is used to represent the camera posture, where φ represents the roll angle, θ represents the pitch angle, and ψ represents the yaw angle; The risk source perception module (1) uses a drone to obtain image information. The drone is an unmanned drone. The drone battery is a wireless charging battery. The drone platform includes an airtight cabin that can be automatically charged. The operation process of the drone platform is as follows: an engineer uses the built-in function of the drone to manually fly and set the route and observation point. At this time, the drone records six variables, including position and rotation angle. When unmanned, the drone is connected to the Internet to check the weather, battery and satellite status. After taking off, it flies to the set point to take pictures, completes the work in sequence, and then returns to the return point to automatically charge and close the cabin; the drone obtains the image interval, and obtains the first image collected by the risk source perception module (1) at each first distance L1 interval. The first distance H is the standard section height of the tower crane.
4. According to the multi-task learning-based tower crane risk source detection transferable system of claim 3, it is characterized in that: The data transmission unit (21) is used to perform remote image transmission with the cabin platform through a Wi-Fi data transmission system installed in the drone, and a network card module with a built-in SIM card is used inside the cabin platform to transmit the pre-processed data to the server; the data computing unit (22) uses a server equipped with two A6000 parallel computing GPUs for the deployment and data computing of multi-task deep learning visual models.
5. According to the multi-task learning-based tower crane risk source detection transferable system according to claim 4, it is characterized in that: The image information transmission unit (31) is used to transmit the risk source image information and the risk source location information to the BIM model (32); the BIM model (32) is used for engineers to view the risk source data information in real time, including the bolt loosening location, crack location, and corrosion area range.
6. According to the multi-task learning-based tower crane risk source detection transferable system of claim 5, it is characterized in that: The specific process of the risk source perception module (1) capturing the image data set, storing the location information of the drone, and the image information transmission unit (31) transmitting the information to the BIM model (32) is as follows: Obtaining the camera position parameters, when shooting the data set through the tower crane sensing unit (11), obtaining the position information of the three parameters of the camera, according to the longitude, latitude and altitude position of the drone (x, y, z), the coordinates are generated by the GPS built into the drone, and the information is stored in a database or file for subsequent use; adding a marker, in the BIM model (32), according to the camera position parameter information, adding a marker at the corresponding coordinate to indicate that there is a camera or information related to the camera at the location; interactive design, when the engineer clicks on the marker, a pop-up window displays the image taken by the camera; The addition of markers requires coordinate transformation. The GPS data collected by the drone is based on the WGS84 coordinate system, and the BIM model (32) relies on the local coordinate system. The GPS data of a feature point on the tower crane or around the tower crane and the coordinate information in the BIM model (32) are obtained respectively, and the coordinate transformation matrix is obtained by calculation. Then, the camera coordinate information is converted into the information of the BIM model (32) according to the matrix.
7. The multi-task learning-based tower crane risk source detection transferable system according to claim 6 is characterized in that: The risk warning module (4) uses a deep learning warning model to evaluate and classify the severity of risk points, generate warning levels, and then push warning information via SMS, email or application to ensure that relevant personnel are notified in a timely manner. The specific process is as follows: For target detection corresponding to cracks, the first threshold M1 is manually set with accuracy as an indicator. The image is scanned by the target detection algorithm to obtain the crack prediction box Bounding Box. The accuracy is used as the first probability P1. When P1≥M1, the first warning prompt is sent; for the rusted part of the tower crane, the second threshold M2 is manually set. The Deeplabv3+ algorithm is first used to process the image to obtain the preliminary recognition result of the rusted area. On this basis, the connected area analysis is performed on the identified rusted area, and the ratio of the number of pixels in the largest connected rusted area to the total number of pixels in the image is calculated as the second probability P2. When P2≥M2, the second warning prompt is sent; for the loose part of the tower crane bolt, the third threshold M3 is manually set. The algorithm identifies a point on the upper edge of the screw in the image as the first key point, a point on the lower edge of the tower crane socket as the second key point, and a point on the upper edge of the tower crane socket as the third key point. The number of straight line pixels between the first key point and the second key point is the first spacing L1, and the number of pixels between the second key point and the third key point is the second spacing L2. When When the alarm is triggered, a third warning prompt is sent.
8. The multi-task learning-based tower crane risk source detection transferable system according to claim 7 is characterized in that: Multi-task deep learning includes the following algorithms: CNN algorithm, ResNet algorithm, YOLOv8 algorithm, Faster-RCNN algorithm, Deeplabv3+ algorithm, YOLOv8-Pose algorithm, image processing algorithm, which can simultaneously perform target detection, semantic segmentation, and key point detection; The structure of the multi-task deep learning vision model includes: region of interest extraction layer, feature extraction layer, object detection branch, semantic segmentation branch, and key point detection branch; The region of interest extraction layer uses the Faster-RCNN algorithm to determine the regions of interest for corrosion, cracks, and bolt loosening. Other algorithms are performed on the regions of interest to narrow the detection range; Feature extraction layer, using pre-trained CNN and ResNet algorithms to extract features from images. Different branches are added on top to handle different tasks. Each branch includes a specific head to output task-specific results. The target detection branch uses the YOLOv8 algorithm to perform crack detection; The semantic segmentation branch uses the Deeplabv3+ algorithm for rust area detection; The key point detection branch uses the YOLOv8-Pose algorithm for bolt loosening detection; The loss functions are designed according to each branch. The target detection branch adopts the combined loss function used in Faster-RCNN or YOLOv8, including classification loss and positioning loss. The semantic segmentation branch uses the pixel-level cross entropy loss function to measure the difference between the category distribution of each pixel predicted by the model and the true label. The key point detection branch uses the key point positioning error as the loss function.
9. The multi-task learning-based tower crane risk source detection transferable system according to claim 8 is characterized in that: The loss function formula is: Total Loss=λ1·Detection Loss+λ2·Segmentation Loss+λ3·Keypoint Loss; Among them, λ1, λ2, and λ3 are the weights of the loss function of each task, which are adjusted according to the importance of the task; Specifically, Detection Loss = classification loss + positioning loss; The classification loss is expressed as N obj is the number of bounding boxes marked as objects; p i is the probability that the predicted bounding box i belongs to the correct category; It is the indicator function of the true category, 1 represents the correct category, and 0 represents other categories; The localization loss is expressed as λ coord is the weight of the positioning loss; (x i ,y i ,w i ,h i ) are the center coordinates and width and height of the predicted bounding box; are the center coordinates and width and height of the real bounding box; N is the number of pixels, C is the number of categories, y ij is the indicator function that pixel i belongs to category j in the true label, is the probability predicted by the model that pixel i belongs to category j; N kp is the number of key points, is the position of key point i predicted by the model, p i is the position of the true key point i.
10. The multi-task learning-based tower crane risk source detection transferable system according to claim 9, characterized in that: The training process of the multi-task deep learning vision model is as follows: task-specific training only updates the weights of specific branches, using the parameter freezing method to fix parameters that do not need to be updated, not participating in back propagation and gradient updates, and assigning different optimizers to different tasks for weight updates; The method for deploying a multi-task deep learning vision model in a server includes: converting a trained deep learning model into a deployment format, setting up a model deployment service on the server side, receiving and preprocessing input data from a client, performing model inference through an inference service, and returning a prediction result; optimizing server hardware resources and network connections, and monitoring model performance and server load in real time; The deployment process uses ONNX for model conversion to convert model parameters and weights into a format supported by the C++ language, which is recorded as the middleware model. The structured pruning channel pruning method that can retain the complete deep learning model structure is used to count the absolute value of the weight of each convolutional layer followed by the BN layer as the scaling factor of the BN layer. The pruning weight threshold is determined according to the array composed of the scaling factors of all BN layers, which is recorded as the first threshold thre_0. During the pruning process, all channels with weight values lower than the threshold are pruned by setting the weights in these channels to zero to achieve sparseness of the model.