Tower crane safety warning method based on small target visual recognition and virtual entity mapping

By improving the YOLOv5 feature fusion network and risk coefficient calculation, the SEC4-YOLOv5 model was built, which solved the problem of small target recognition accuracy at the tower crane construction site, real-time monitoring of construction site safety was achieved, and the risk of safety accidents was reduced.

CN116129135BActive Publication Date: 2025-08-26HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211343504.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-08-26
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

There are great safety hazards on the construction site of tower cranes, the existing target detection algorithm model has low recognition accuracy and poor generalization capabilities, especially for small target recognition accuracy, which cannot meet the real-time requirements.

Method used

Improve the YOLOv5 feature fusion network, add a scale detection layer, adopt the SENet attention mechanism and CIoU_LOSS function to build the SEC4-YOLOv5 model, combine the elimination of error detection module and interpolation function to calculate the hazard coefficient for early warning.

Benefits of technology

It improves the accuracy of small target recognition and model robustness, realizes real-time monitoring of safety at construction sites, and reduces the risk of safety accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116129135B_ABST
    Figure CN116129135B_ABST
Patent Text Reader

Abstract

The present invention discloses a tower crane safety early warning method based on small target visual recognition and virtual entity mapping, comprising the following steps: step 1, constructing a SEC4-YOLOv5 model suitable for small target visual recognition; step 2, fitting an interpolation function of hook height and hook size; step 3, using the weight file of the SEC4-YOLOv5 model to identify monitoring video data, and obtain construction personnel position information and hook data; step 4, demarcating a dangerous area according to the hook position and height; step 5, calculating a danger coefficient in combination with the construction personnel position information; step 6, determining whether an alarm is needed by comparing the calculated danger coefficient with a preset safety alarm threshold. It can be seen that the present invention adopts an optimized SEC4-YOLOv5 model to realize feature channel self-calibration. The model recognition detection accuracy is improved, and a balance between the speed and accuracy of the target detection model is achieved, with strong robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a tower crane safety early warning method, which is implemented based on small target visual recognition and virtual entity mapping technology. Background Art

[0002] Tower cranes are widely used for material handling on construction sites. Crane operators often make subjective judgments about hook position through observation, and the reliability of hook positioning depends on their own experience. Complex construction site environments, densely populated areas, and overlapping work processes create significant safety hazards during tower crane lifting operations. Therefore, to eliminate these hazards, it is crucial to quickly identify tower crane safety zones and establish an early warning mechanism.

[0003] In recent years, deep learning models using neural networks have been widely used as a new technical tool to assist in construction site safety management. The large amount of image data captured by image sensors at construction sites provides a rich data set for training object detection algorithms. Advances in computer vision technology have made the use of digital imaging techniques to rapidly identify construction site safety a viable solution. However, there is currently no research on applying deep learning technology to detect the safety of tower crane construction sites.

[0004] Traditional target detection algorithms require manually designed feature operators to extract image features, but this approach has limitations, resulting in low model recognition accuracy and poor generalization. Deep learning-based target detection algorithms use convolutional neural networks to self-learn image features, significantly improving model speed and accuracy while also demonstrating strong robustness. These algorithms include two-stage target detection algorithms such as R-CNN and single-stage target detection algorithms such as YOLO. While the R-CNN family of algorithms offers high detection accuracy but slow speed, they cannot meet real-time requirements. The YOLO family of algorithms directly regresses the category and position of targets in the input image, offering the fastest detection speed and ease of engineering deployment. However, their recognition accuracy for small targets is slightly lower. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this paper provides a tower crane safety warning method based on small-target visual recognition and virtual entity mapping. This method improves the feature fusion network of YOLOv5, adds a scale detection layer, and utilizes the SENet attention mechanism network to improve the original model and achieve self-calibration of feature channels. Furthermore, it provides an improved risk factor calculation model that takes monitoring time into account, further improving the accuracy of safety warning results and reducing the risk of safety accidents.

[0006] In order to achieve the above technical objectives, the present invention will adopt the following technical solutions:

[0007] A tower crane safety warning method based on small target visual recognition and virtual entity mapping includes the following steps:

[0008] Step 1: Build a SEC4-YOLOv5 model for small target visual recognition

[0009] The SENet attention mechanism network is used to improve the existing YOLOv5 feature fusion network, and the CIoU_LOSS function is used to replace the GIoU_LOSS function in the existing YOLOv5 feature fusion network to obtain the SEC4-YOLOv5 model. The SEC4-YOLOv5 model is trained with the training set and verified with the validation set, thereby obtaining the weight file for detecting hooks and construction workers.

[0010] Add a false detection removal module to the SEC4-YOLOv5 model to remove falsely detected hook targets;

[0011] Step 2: Fitting the interpolation function of hook height and hook size;

[0012] Step 3: Identify surveillance video data

[0013] The SEC4-YOLOv5 model's weight file is used to identify surveillance video data captured by the camera on the tower crane's luffing trolley. The false detection rejection module is used to remove falsely detected hook targets. This allows the construction worker's location information, the unique hook's location information, and its dimensions to be uploaded to a cloud database.

[0014] Step 4: Delineate the danger zone based on the hook position and height

[0015] Based on the position information and size data of the hook identified in step 3, the height of the falling object is calculated using an interpolation function to delineate the danger zone;

[0016] Step 5: Calculate the risk factor based on the construction workers' location information

[0017] Calculate the risk factor based on the height of the falling object, the location of the construction workers, and the monitoring time;

[0018] Step 6: Early Warning Judgment

[0019] By comparing the risk factor calculated in step 5 with the preset safety alarm threshold, it is determined whether an alarm is needed.

[0020] As a further improvement to the tower crane safety early warning method, in step 1, the training set and validation set used are constructed in the following way:

[0021] Step 1.1.1. Acquire images of a tower crane luffing trolley hook and images of construction workers at the construction site where the tower crane luffing trolley is located to construct an image dataset; divide the image dataset into two categories: training set images and validation set images; the training set images and validation set images are stored in an image training set folder and an image validation set folder, respectively;

[0022] Step 1.1.2: Label the image dataset constructed in step 1.1 to obtain the corresponding label file; the label file of the training set images is stored in the training set label folder, and the label file of the validation set images is stored in the validation set label folder;

[0023] The training set images in the image training set folder in step 1.1.1 and the label files in the training set label folder in step 1.1.2 constitute the training set;

[0024] The validation set consists of the validation set images in the image validation set folder in step 1.1.1 and the label files in the validation set label folder in step 1.1.2.

[0025] As a further improvement to the tower crane safety warning method described above, in step 1.1.1, by fixing the shooting angle of the camera on the tower crane's luffing trolley so that the camera is directly above the hook, images of the hook and construction workers captured under different lighting and weather conditions are collected to construct an image dataset.

[0026] In step 1.1.2, use the LabelImg annotation tool to annotate the hook and construction worker images to obtain the corresponding label files.

[0027] As a further improvement to the tower crane safety warning method described above, in step 1, the training process of the SEC4-YOLOv5 model specifically includes the following steps:

[0028] Step 1.2.1, preprocess the input image;

[0029] Step 1.2.2: Input the preprocessed image into the backbone feature extraction network Backbone for feature extraction; the backbone feature extraction network is finally embedded into the channel attention SENet module;

[0030] Step 1.2.3: The extracted features are transferred and fused through the Neck layer through the FPN+PAN structure. A small target detection scale is added to the Neck layer to enhance the ability to express small target features, and four multi-scale feature maps are output.

[0031] In step 1.2.4, in the Prediction layer, each pixel of the feature map at each scale has three corresponding prior boxes. The prior boxes are mapped to the input image to regress the category and position of the target object to obtain the predicted box of the target object.

[0032] Step 1.2.5: Construct a loss function LOSS based on the predicted box and the true box of the target object, and train until the loss function LOSS converges; the bounding box regression loss function in the loss function LOSS uses the CIoU_LOSS function.

[0033] As a further improvement to the tower crane safety warning method described above, in step 1.2.2, a channel attention SELayer module is added after the fourth BottleneckCSP structure in Backbone. At the same time, the number of channels in the global average pooling is specified as the number of channels in the BottleneckCSP output feature map. The SELayer module is used to strengthen the network's learning of important channel feature information.

[0034] In step 1.2.3, a small target detection scale is added. The specific method is as follows: upsampling is performed three times in the FPN of the Neck layer and downsampling is performed three times in the PAN. By further deepening the network, the detection scale is expanded to four to enhance the ability to express small target features.

[0035] As a further improvement to the above tower crane safety warning method, in step 1, the false detection elimination module first uses the SEC4-YOLOv5 model to identify the hooks in the image to be detected, and determines whether the number of detected hooks is greater than 1. If it is not greater than 1, the identified hook result is output. If it is greater than 1, the detection frame closest to the image center is retained based on the position information of the hook detection frame, thereby eliminating false detections to the greatest extent. The calculation formula is as follows:

[0036]

[0037]

[0038]

[0039]

[0040] p=min(d i )

[0041] Where w te 、h te is the width and height of the image to be detected, w in 、h in is the width and height of the image after inputting the model, x l 、y lis the coordinate of the upper left point of the i-th hook detection frame in each detection image, x r 、y r is the coordinate of the lower right point of the i-th hook detection frame in each detection image, x i 、y i is the coordinate of the center point of the i-th hook detection frame in each detection image, d i is the Euclidean distance between the center of the i-th hook detection frame and the center of the detection image in each detection image, and p is the only detection frame finally retained.

[0042] As a further improvement to the above tower crane safety warning method, in step 2, determining the interpolation function specifically includes the following steps:

[0043] Step 2.1, collect video data of a hook rising from the lowest point to the highest point at a constant speed;

[0044] Step 2.2: Capture images at the same time interval and record discrete data points with hook height information and hook size information. The hook size information can be obtained by using the label file obtained by the LabelImg annotation tool. The calculation formula is as follows:

[0045]

[0046] Where w is the width of the circumscribed rectangle of the hook, h is the height of the circumscribed rectangle of the hook, and size is the diagonal size of the circumscribed rectangle of the hook;

[0047] Step 2.3: Based on the discrete data points, interpolate the continuous function. By comparing the fitting degree of the linear interpolation function, the second-order spline curve interpolation function and the three-bound spline curve interpolation function, the second-order spline curve interpolation function is finally selected as the interpolation function of the hook height and the hook size.

[0048] As a further improvement to the above-mentioned tower crane safety warning method, in step 4, the range of the danger zone is determined based on the hook height information height. Specifically, when the height is less than 40 meters, the radius of the danger zone is a fixed value; when the height is greater than 40 meters, the radius of the danger zone is a linear function that increases with height; the center of the danger zone is the center of the hook detection frame.

[0049] As a further improvement to the above tower crane safety early warning method, in step 5, the risk factor D is calculated by the following formula:

[0050] D=L×E×C×R1×R2

[0051] L is the probability of an accident;

[0052] E is the frequency of being in a dangerous environment, calculated by establishing a correspondence between the distance between the construction worker and the center of the dangerous area and the probability density function; the construction worker's position is identified by the SEC4-YOLOv5 model, and the center of the dangerous area is the center of the hook detection box identified by the SEC4-YOLOv5 model;

[0053] C is the accident consequence, which is approximately determined based on the height of the falling object;

[0054] R1 is the correction coefficient related to the month in which the video surveillance time occurs; where: when the video surveillance time occurs in a month with a high incidence of accidents, R1 = 1.2; when the video surveillance time occurs in a month with frequent accidents, R1 = 1.1; when the video surveillance time occurs in other months, R1 = 1;

[0055] R2 is a correction coefficient related to the time of video surveillance; among which: when the video surveillance time is at a high-incidence time of accidents, R2 = 1.2; when the video surveillance time is at a frequent accident time, R2 = 1.1; when the video surveillance time is at other times, R1 = 1.

[0056] As a further improvement of the above tower crane safety warning method, in step 6, the safety alarm threshold D max =50; when the real-time calculated risk level D exceeds the safety alarm threshold D max When the system sends an alarm to the personnel in the danger zone, asking them to leave the danger zone.

[0057] Based on the above technical objectives, the present invention has the following advantages over the prior art:

[0058] 1. Improved YOLOv5's feature fusion network, adding a scale detection layer. The SENet attention mechanism network was used to improve the original model, enabling self-calibration of feature channels. To address the slow convergence of GIoU_LOSS, which degenerates to IoU_LOSS when the predicted and target boxes overlap, CIoU_LOSS was used to replace GIoU_LOSS, accelerating model convergence and improving bounding box positioning accuracy. By using the SEC4-YOLOv5 small object recognition optimization algorithm, the model's recognition and detection accuracy was improved at a minimal cost, achieving a balance between speed and accuracy for the object detection model, resulting in strong robustness.

[0059] 2. A tower crane safety warning method based on virtual entity mapping is proposed. It effectively utilizes surveillance video information and uses computer vision technology for calculation and processing, so that the danger level of construction workers during operation can be presented in real time in the surveillance video, assisting safety management personnel at the construction site to scientifically and efficiently supervise the safety of the tower crane construction site and reduce the risk of safety accidents. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a schematic flow diagram of the present invention.

[0061] Figure 2 This is a structural diagram of the SEC4-YOLOv5 small target recognition optimization algorithm of the present invention.

[0062] Figure 3 This is the hook height detection result of an embodiment of the present invention.

[0063] Figure 4 This is the tower crane safety early warning detection result of the embodiment of the present invention. DETAILED DESCRIPTION

[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work are within the scope of protection of the present invention. Unless otherwise specified, the relative arrangement of components and steps, expressions and numerical values ​​described in these embodiments do not limit the scope of the present invention. Technologies, methods and equipment known to ordinary technicians in the relevant fields may not be discussed in detail, but where appropriate, the technologies, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed here, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of the exemplary embodiments may have different values.

[0065] Reference Figure 1 As shown, the present invention constructs a tower crane safety warning method based on small target visual recognition and virtual entity mapping, which includes the following steps:

[0066] Step 1: Build a SEC4-YOLOv5 model for small target visual recognition

[0067] The SENet attention mechanism network is used to improve the existing YOLOv5 feature fusion network, and the CIoU_LOSS function is used to replace the GIoU_LOSS function in the existing YOLOv5 feature fusion network to obtain the SEC4-YOLOv5 model. A false detection removal module that can remove falsely detected hook targets is added to the SEC4-YOLOv5 model.

[0068] The SEC4-YOLOv5 model was trained using the training set and verified using the validation set to obtain the weight files for detecting hooks and construction workers.

[0069] Step 2: Fitting the interpolation function of hook height and hook size;

[0070] Step 3: Identify surveillance video data

[0071] The SEC4-YOLOv5 model's weight file is used to identify surveillance video data captured by the camera on the tower crane's luffing trolley. The false detection rejection module is used to remove falsely detected hook targets. This allows the construction worker's location information, the unique hook's location information, and its dimensions to be uploaded to a cloud database.

[0072] Step 4: Delineate the danger zone based on the hook position and height

[0073] Based on the position information and size data of the hook identified in step 3, the height of the falling object is calculated using an interpolation function to delineate the danger zone;

[0074] Step 5: Calculate the risk factor based on the construction workers' location information

[0075] Calculate the risk factor based on the height of the falling object, the location of the construction workers, and the monitoring time;

[0076] Step 6: Early Warning Judgment

[0077] The system compares the risk factor calculated in step 5 with the preset safety alarm threshold to determine whether an alarm is necessary. If the risk factor exceeds the safety alarm threshold, the system sends an alert to personnel within the danger zone, requesting them to exit. The safety alarm threshold is determined based on a balance between operational efficiency and safety.

[0078] In step 1, the training set and validation set used are constructed in the following manner: the shooting angle of the camera on the tower crane's luffing trolley is fixed so that the camera is located directly above the hook, and 1,000 images of the hook and construction workers at different heights taken under different lighting conditions and weather conditions are collected to construct an image dataset. The images are manually annotated using the LabelImg annotation tool. When annotating, the minimum circumscribed rectangular box of the hook in the image is drawn and its category is named "hook", and the minimum circumscribed rectangular box of the construction worker in the image is drawn and its category is named "person". The corresponding label file in txt format is obtained. The content of the label file is, in order, the category ID number, the normalized target frame center x coordinate, the normalized target frame center y coordinate, the normalized target frame width w, and the normalized target frame height h, that is, the real frame information of the target. The image dataset folder (images) is divided into a training set folder (train) and a validation set folder (val). The training set folder (train) stores the training set images (800 images), and the validation set folder (val) stores the validation set images (200 images). The ratio of the number of training sets to the validation set is 8:2. Similarly, the label folder (labels) is divided into the corresponding training set label folder (train) and validation set label folder (val).

[0079] In the step 1, the present invention refers to Figure 2 The SEC4-YOLOv5s small target recognition optimization algorithm (SEC4-YOLOv5 model) is improved by using the YOLOv5s algorithm for training, which specifically includes the following steps:

[0080] Step 1.2.1: Preprocess the input image, including mosaic data augmentation, random image flipping, and HSV color space enhancement, to enhance the robustness of the model.

[0081] Step 1.2.2: Input the preprocessed image into the backbone feature extraction network Backbone for feature extraction. The backbone feature extraction network consists of four modules: Focus module, CONV module, BottleneckCSP module, and SPP module. Furthermore, a channel attention SENet module is embedded at the end of the backbone feature extraction network Backbone to strengthen the network's learning of important channel feature information.

[0082] Furthermore, in step 1.2.2, the five modules in the backbone feature extraction network Backbone are as follows:

[0083] Step 1.2.2.1. The Focus module performs a slicing operation on the input, dividing the default input 640*640*3 3-channel image into 4 slices by taking alternate pixel values. Each slice is 320*320*4 in size. The 4 slices are concatenated using the Concat operation to obtain a 320*320*12 feature map. After convolution with 32 convolution kernels, the final feature map is 320*320*32, which reduces the amount of model calculation and speeds up the calculation speed.

[0084] Step 1.2.2.2, CONV module performs convolution, normalization, and activation operations on the input in order to deepen the network;

[0085] Step 1.2.2.3, the BottleneckCSP module is mainly composed of a Bottleneck structure. The BottleneckCSP module is a residual structure. After convolution with a 1*1 convolution kernel and a 3*3 convolution kernel, the final output is added to the initial input. The BottleneckCSP module divides the input into two branches. Branch one is convolved with a 1*1 convolution kernel to reduce the number of feature map channels by half. Branch two is convolved with a 1*1 convolution kernel, a BottleneckCSP structure, and a 1*1 convolution kernel to reduce the number of feature map channels by half. The outputs of branches one and two are connected using the Concat function. After normalization, activation, and convolution with a 1*1 convolution kernel, the final output is the output. The size of the output feature map is the same as the input feature map of the BottleneckCSP module. This module aims to better extract deep features of the image.

[0086] Step 1.2.2.4: The input feature map of the SPP module is 20*20*512 in size. First, it is convolved with a 1*1 convolution kernel to obtain a 20*20*256 feature map. Then, it is sampled and concatenated with three parallel maximum pooling layers (5*5, 9*9, and 13*13). Finally, it is convolved with 512 convolution kernels to obtain a 20*20*512 feature map. This module realizes the fusion of local features and global features, enriching the expressive power of the feature map.

[0087] Step 1.2.2.5. Add the channel attention SELayer module after the fourth BottleneckCSP structure in Backbone, and specify the number of channels of the global average pooling to be 512, the number of channels of the BottleneckCSP output feature map. The SELayer module includes squeeze, excitation, and scale. The squeeze operation compresses the original input feature into a one-dimensional vector along the channel direction through global average pooling, that is, compresses a 20×20×512 original input feature into a 1×1×512 feature to obtain the global information of each channel; the excitation operation compresses the original input feature into a one-dimensional vector along the channel direction through global average pooling, that is, compresses a 20×20×512 original input feature into a 1×1×512 feature to obtain the global information of each channel; The feature channel is reduced and increased in dimension by using the scaling factor r through two fully connected layers. In this embodiment, the scaling factor r is 16. The first fully connected layer reduces the feature channel from 1×1×512 to 1×1×32 to reduce the amount of calculation. The second fully connected layer restores the feature channel from 1×1×32 to 1×1×512. The scalar obtained is the weight of each channel; the scaling operation performs a weighted fusion of the weights of each channel obtained by the excitation operation (1×1×512) and the original input features (20×20×512), outputs the weighted feature map, and uses the SELayer module to strengthen the network's learning of important channel feature information. The calculation formula of the SELayer module is as follows:

[0088]

[0089]

[0090]

[0091] u′ c =n c u c

[0092] Where c is the number of channels, H and W are the height and width of the original input features, and u c (i, j) is the original input feature, z c is the result after the squeezing operation, W1 and W2 are the weights of the two fully connected layers used for dimensionality reduction and dimensionality increase, σ is the ReLU activation function, is the result after the first fully connected layer is activated by the ReLU activation function, δ is the Sigmoid activation function, r is the scaling factor used to control the complexity of the model, and n c is the weight obtained after the second fully connected layer is nonlinearly activated by the Sigmoid activation function, u′ c It is the weighted feature, i and j are the height and width coordinates of the feature pixel respectively.

[0093] Step 1.2.3, the extracted features are transferred and fused through the FPN+PAN structure in the Neck layer. After two upsamplings in the Neck layer to obtain an 80*80 feature map, convolution and upsampling are continued to obtain a 160*160 feature map, which is then spliced ​​and fused with the feature map of the same scale of the backbone feature extraction network to obtain a 160*160 detection scale. At the same time, the original 80*80 detection scale obtained by the fusion feature map after two upsamplings is converted into a downsampling fusion feature map after three upsamplings. The other two detection scales (40*40, 20*20) are deduced in this way. By further deepening the network to enhance the expression ability of small target features, four scale feature maps (20*20, 40*40, 80*80, 160*160) are output.

[0094] In step 1.2.4, the Prediction layer outputs feature maps of four scales and generates prior frames of three corresponding initial anchor frame sizes. The prior frames of different sizes are mapped to the input image to regress the category and position of the target object to obtain the predicted frame of the target object. The result represents whether the three prior frames at each grid point contain an object, the type of object, and the adjustment parameters of the prior frames. The 12 groups of initial anchor frame sizes are shown in the following table:

[0095] Table 1

[0096]

[0097] Step 1.2.5. Construct a loss function LOSS based on the predicted box and the true box of the target object. The loss function LOSS includes the category loss function, the confidence loss function, and the bounding box regression loss function. Furthermore, the bounding box regression loss function is changed from GIoU_LOSS to CIoU_LOSS. Train until LOSS converges. The calculation formula of the bounding box loss function CIoU_LOSS is as follows:

[0098]

[0099]

[0100]

[0101]

[0102] Where B p is the prediction box, B gt is the real frame, w p and h p Represents the width and height of the prediction box, w gt and h gtThey represent the width and height of the real box respectively, ρ represents the Euclidean distance between the center point of the predicted box and the real box, c represents the diagonal distance of the minimum circumscribed rectangular box, and v is a parameter to measure the consistency of the aspect ratio.

[0103] In the second step, the interpolation function is determined, specifically including the following steps:

[0104] Step 2.1, collect video data of a hook rising from the lowest point to the highest point at a constant speed;

[0105] Step 2.2: Capture images at the same time interval (3 seconds) and record discrete data points with hook height and hook size information. The hook size information is obtained by using the label file generated by the LabelImg annotation tool. When calculating, consider that the hook size in the label file is a normalized value, so multiply it by the expansion factor 1000. The calculation formula is as follows:

[0106]

[0107] Where w is the width of the circumscribed rectangle of the hook, h is the height of the circumscribed rectangle of the hook, and size is the diagonal size of the circumscribed rectangle of the hook;

[0108] Step 2.3: Based on the discrete data points, a continuous function is interpolated using a second-order spline curve as the interpolation function of the hook height (height) and the hook size (size).

[0109] In step 1, the false detection elimination module first uses the trained YOLOv5s model to identify hooks in the image to be detected, and determines whether the number of detected hooks is greater than 1. If not, the module outputs the identified hook result. If greater than 1, the module retains the detection frame closest to the center of the image based on the position information of the hook detection frame, thereby eliminating false detections to the greatest extent. The calculation formula is as follows:

[0110]

[0111]

[0112]

[0113]

[0114] p=min(d i )

[0115] Where w te 、h te is the width and height of the image to be detected, w in 、h in is the width and height of the image after inputting the model, x l 、yl is the coordinate of the upper left point of the i-th hook detection frame in each detection image, x r 、y r is the coordinate of the lower right point of the i-th hook detection frame in each detection image, x i 、y i is the coordinate of the center point of the i-th hook detection frame in each detection image, d i is the Euclidean distance between the center of the i-th hook detection frame and the center of the detection image in each detection image, and p is the only detection frame finally retained.

[0116] In step three, during the detection phase, the surveillance video is fed into the detection model. The weight file of the trained SEC4-YOLOv5s model is used to identify the hook and construction workers in each frame of RGB image data. The false detection elimination module removes falsely detected hook targets, obtaining the unique hook location information and size data (size). By calling the cloud data and substituting it into the second-order spline curve interpolation function, the weight drop height (height) is calculated, completing the hook detection task. The range of the danger zone is determined based on the hook height information (height). When the height is less than 40 meters, the danger zone radius is a fixed value. When the height is greater than 40 meters, the danger zone radius is a linear function that increases with height. The danger zone is delineated with the center of the hook detection frame as the center of the circle, and displayed in real time in the surveillance video, completing the danger zone delineation task.

[0117] In step 5, the risk factor is obtained by the modified operating condition risk assessment method, and the calculation formula is as follows:

[0118] D=L×E×C×R1×R2

[0119] L is the possibility of an accident. Because it is applied to the hazard level evaluation of a specific tower crane object strike accident scenario, the value 1 according to the original definition of the formula indicates that the possibility is very small and unexpected. E is the frequency of being in a dangerous environment, which can be calculated using a probability density function that obeys the normal distribution. That is, a corresponding function between the distance between the personnel and the center of the dangerous area and the probability density function is established. The SEC4-YOLOv5 small target recognition optimization algorithm can be used to identify the dangerous area and the personnel position, and then the distance between the personnel and the center of the dangerous area can be calculated and substituted into the function to determine the E value. (The E value range is within the original definition of 0-10); C is the consequence of the accident, which can be approximately determined by the height of the falling object, with a value of 5 below 20m, 15 between 20 and 50m, and 40 above 50m; R1 is the correction coefficient. By analyzing the investigation reports of 117 object-strike accidents, it was found that the accident peaks in May, July, September, and October. Therefore, based on the video surveillance time, the correction coefficient R1 corresponding to the high-incidence months is 1.2, the correction coefficient R1 corresponding to the accident-frequent months in June and December is 1.1, and the correction coefficient R1 = 1 for the remaining months; R2 is the correction coefficient. By analyzing the investigation reports of 117 object-strike accidents, it was found that the accident peaks in 8:00, 9:00, and 11:00. Therefore, based on the video surveillance time, the correction coefficient R2 corresponding to the high-incidence times is 1.2, the correction coefficient R2 corresponding to the accident-frequent times is 10:00, 14:00, 15:00, 16:00, 17:00, and 18:00, and the correction coefficient R2 = 1 for the remaining times.

[0120] Determine the highest risk level D that the project object is willing to bear based on its risk tolerance for strike accidents. max =50, when the real-time calculated risk level D exceeds the safety alarm threshold D max When the system sends an alarm to the personnel in the danger zone, asking them to leave the danger zone.

[0121] In order to verify the performance of the SEC4-YOLOv5s small target recognition optimization algorithm in this embodiment, a comparative test was carried out under the premise of using the same hardware configuration and experimental parameters. Part of the training environment: the graphics card is NVIDIA GeForce GTX1650, the deep learning framework is Pytorch1.5, the language environment is Python3.7, CUDA10.2 and Cudnn7.6.5 are used for GPU acceleration, the input image size is 640*640, the optimization function is SGD, the training momentum is 0.937, the initial learning rate is 0.01, the weight decay is 0.0005, the batch size is 8, and the number of training rounds is 100. The precision P (Precision), recall R (Recall), average precision AP (Average precision), weight file size, and single image detection time are used as model performance evaluation indicators. The compared network models include YOLOv5s, YOLOv5m, and the SEC4-YOLOv5s algorithm used in this embodiment. The model performance evaluation indicators obtained by training are shown in Table 2.

[0122] Table 2 Comparison of experimental data of different models

[0123]

[0124] As shown in Table 2, YOLOv5m has improved to a certain extent in all indicators compared to YOLOv5s, but the size of the YOLOv5m model is three times that of YOLOv5s, and the detection speed is slightly slower. The AP_0.5 value of YOLOv5m is improved by 1.15% compared to the YOLOv5s model, and the AP_0.5-0.95 value is improved by 2.66%. The SEC4-YOLOv5s model described in this embodiment has an AP_0.5 value improvement of 3.18% and an AP_0.5-0.95 value improvement of 2.99% compared to the YOLOv5s model, and the model weight file only increases by 2MB. The single image detection speed is comparable to that of the YOLOv5s model. The results show that the SEC4-YOLOv5s model has significantly improved model performance while only increasing the computational overhead to ensure that the model detection speed meets the requirements of engineering deployment.

[0125] When the YOLOv5s model and SEC4-YOLOv5s model were tested respectively using a small target test set (300 test images, which do not overlap with the training set and validation set), the results showed that the YOLOv5s model missed detections when detecting and identifying small target hooks, and its overall accuracy was poor. However, the SEC4-YOLOv5s model had enhanced feature extraction capabilities, effectively improved the problem of missed detection of small targets, and significantly improved prediction accuracy.

[0126] The SEC4-YOLOv5s model was tested using a surveillance video. The model was able to accurately identify the hook and the construction worker. Figure 3 As shown in the figure, when the hook height is low, the test results have slight fluctuations, and the overall accuracy is high. Figure 4 As shown in the figure, the model can calculate and demarcate the danger zone (the red circle area in the figure) based on the hook target, and calculate the danger level in real time. When the model determines that the danger level of the construction worker's position is high, the construction worker prediction rectangle is displayed in red, and the overall accuracy of the model is high.

Claims

1. A tower crane safety warning method based on small target visual recognition and virtual entity mapping, characterized in that: The steps include: Step 1: Build a SEC4-YOLOv5 model for small target visual recognition The SENet attention mechanism network is used to improve the existing YOLOv5 feature fusion network, and the CIoU_LOSS function is used to replace the GIoU_LOSS function in the existing YOLOv5 feature fusion network to obtain the SEC4-YOLOv5 model. The SEC4-YOLOv5 model is trained with the training set and verified with the validation set, thereby obtaining the weight file for detecting hooks and construction workers. In the SEC4-YOLOv5 model, a channel attention SELayer module is added after the fourth BottleneckCSP structure in Backbone. The number of channels in the global average pooling is specified as the number of channels in the BottleneckCSP output feature map. The SELayer module is used to strengthen the network's learning of important channel feature information. A small target detection scale is added to the Neck layer by performing three upsamplings in the FPN of the Neck layer and three downsamplings in the PAN. By further deepening the network, the scale is expanded to four to enhance the representation of small target features. A false detection removal module is added to the SEC4-YOLOv5 model to remove falsely detected hook targets. Step 2: Fitting the interpolation function of the hook height and the hook size; the interpolation function uses a second-order spline curve interpolation as the interpolation function of the hook height height and the hook size size; Step 3: Identify surveillance video data The SEC4-YOLOv5 model's weight file is used to identify surveillance video data captured by the camera on the tower crane's luffing trolley. The false detection rejection module is used to remove falsely detected hook targets. This allows the construction worker's location information, the unique hook's location information, and its dimensions to be uploaded to a cloud database. Step 4: Delineate the danger zone based on the hook position and height Based on the position information and size data of the hook identified in step 3, the height of the falling object is calculated using an interpolation function to delineate the danger zone; Step 5: Calculate the risk factor based on the construction workers' location information The risk factor D is calculated based on the height of the falling object, the location information of the construction workers, and the monitoring time. It is calculated using the following formula: D=L×E×C×R1×R2 Where: L is the probability of an accident; E is the frequency of being in a dangerous environment, calculated by establishing a correspondence between the distance between the construction worker and the center of the dangerous area and the probability density function; the construction worker's position is identified by the SEC4-YOLOv5 model, and the center of the dangerous area is the center of the hook detection box identified by the SEC4-YOLOv5 model; C is the accident consequence, which is approximately determined based on the height of the falling object; R1 is the correction coefficient related to the month of video surveillance; when the video surveillance time is in the month with a high incidence of accidents, R1 = 1.2; when the video surveillance time is in the month with frequent accidents, R1 = 1.1; When the video surveillance time is in the rest of the month, R1=1; R2 is the correction coefficient related to the time of video monitoring; when the video monitoring time is at the time of high accident incidence, R2 = 1.2; when the video monitoring time is at the time of frequent accidents, R2 = 1.1; When the video monitoring time is at other moments, R1=1; Step 6: Early Warning Judgment By comparing the risk factor calculated in step 5 with the preset safety alarm threshold, it is determined whether an alarm is needed.

2. The tower crane safety early warning method based on small target visual recognition and virtual entity mapping according to claim 1 is characterized in that: In step 1, the training set and validation set used are constructed as follows: Step 1.1.1, obtain images of the tower crane luffing trolley hook and images of construction workers at the construction site where the tower crane luffing trolley is located to construct an image dataset; and divide the image dataset into two categories, namely training set images and validation set images; the training set images and validation set images are stored in the image training set folder and image validation set folder respectively; Step 1.1.2: Label the image dataset constructed in step 1.1 to obtain the corresponding label file; the label file of the training set images is stored in the training set label folder, and the label file of the validation set images is stored in the validation set label folder; The training set images in the image training set folder in step 1.1.1 and the label files in the training set label folder in step 1.1.2 constitute the training set; The validation set consists of the validation set images in the image validation set folder in step 1.1.1 and the label files in the validation set label folder in step 1.1.

2.

3. The tower crane safety early warning method based on small target visual recognition and virtual entity mapping according to claim 1 is characterized in that: In step 1.1.1, the camera on the tower crane's luffing trolley is positioned at a fixed angle, positioned directly above the hook. Images of the hook and construction workers are collected under different lighting and weather conditions to construct an image dataset. In step 1.1.2, use the LabelImg annotation tool to annotate the hook and construction worker images to obtain the corresponding label files.

4. The tower crane safety warning method based on small target visual recognition and virtual entity mapping according to claim 1 is characterized in that: In step 1, the training process of the SEC4-YOLOv5 model includes the following steps: Step 1.2.1, preprocess the input image; Step 1.2.2: Input the preprocessed image into the backbone feature extraction network Backbone for feature extraction; the backbone feature extraction network is finally embedded into the channel attention SENet module; Step 1.2.3: The extracted features are transferred and fused through the Neck layer through the FPN+PAN structure. A small target detection scale is added to the Neck layer to enhance the ability to express small target features, and four multi-scale feature maps are output. In step 1.2.4, in the Prediction layer, each pixel of the feature map at each scale has three corresponding prior boxes; the prior boxes are mapped to the input image to regress the category and position of the target object. Get the predicted box of the target object; Step 1.2.5: Construct a loss function LOSS based on the predicted box and the true box of the target object, and train until the loss function LOSS converges; the bounding box regression loss function in the loss function LOSS uses the CIoU_LOSS function.

5. The tower crane safety warning method based on small target visual recognition and virtual entity mapping according to claim 1 is characterized in that: In step 1, the false detection elimination module first uses the SEC4-YOLOv5 model to identify hooks in the image to be detected, and determines whether the number of detected hooks is greater than 1. If not, the module outputs the identified hook result. If greater than 1, the module retains the detection frame closest to the center of the image based on the position information of the hook detection frame, thereby eliminating false detections to the greatest extent possible. The calculation formula is as follows: p=min(d i ) Where w te 、h te is the width and height of the image to be detected, w in 、h in is the width and height of the image after inputting the model, x l 、y l is the coordinate of the upper left point of the i-th hook detection frame in each detection image, x r 、y r is the coordinate of the lower right point of the i-th hook detection frame in each detection image, x i 、y i is the coordinate of the center point of the i-th hook detection frame in each detection image, d i is the Euclidean distance between the center of the i-th hook detection frame and the center of the detection image in each detection image, and p is the only detection frame finally retained.

6. The tower crane safety warning method based on small target visual recognition and virtual entity mapping according to claim 1 is characterized in that: In step 2, the interpolation function is determined, which specifically includes the following steps: Step 2.1, collect video data of a hook rising from the lowest point to the highest point at a constant speed; Step 2.2: Capture images at the same time interval and record discrete data points with hook height information and hook size information. The hook size information can be obtained by using the label file obtained by the LabelImg annotation tool. The calculation formula is as follows: Where w is the width of the circumscribed rectangle of the hook, h is the height of the circumscribed rectangle of the hook, and size is the diagonal size of the circumscribed rectangle of the hook; Step 2.3: interpolate the continuous function based on the discrete data points by comparing the fitting degree of the linear interpolation function, the second-order spline curve interpolation function and the three-bound spline curve interpolation function.

7. The tower crane safety warning method based on small target visual recognition and virtual entity mapping according to claim 1 is characterized in that: In step 4, the size of the danger zone is determined based on the hook height information height. Specifically, when the height is less than 40 meters, the radius of the danger zone is a fixed value; when the height is greater than 40 meters, the radius of the danger zone is a linear function that increases with height; the center of the danger zone is the center of the hook detection frame.

8. The tower crane safety warning method based on small target visual recognition and virtual entity mapping according to claim 1 is characterized in that: In step 6, the security alarm threshold D max =50; when the real-time calculated risk level D exceeds the safety alarm threshold D max When the system sends an alarm to the personnel in the danger zone, asking them to leave the danger zone.