Elevator speed detection method based on improved YOLOv5 and Deepsort
Through the improved YOLOv5 and DeepSORT technology, real-time monitoring and calculation of the movement speed of drilling cranes has been solved, and the problem of insufficient real-time and accuracy of lifting crane speed monitoring in the existing technology has been solved, and effective protection of the physical safety of derrick workers has been achieved.
Patent Information
- Application Number
- CN202510060052.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-03
AI Technical Summary
The existing oilfield drilling monitoring system has poor real-time and insufficient accuracy in monitoring the lifting motion speed, and has failed to effectively protect the personal safety of the second-story derrick staff.
The improved YOLOv5 and DeepSORT technology are used to detect and identify the object of the lifting card through real-time monitoring video, calculate the speed of the lifting card in real time, and determine whether an acoustic and light alarm is issued based on the set speed safety threshold, and start the dynamic braking system or emergency stop device.
Real-time and accurate monitoring of the movement status of the lifting card and timely warning, effectively protecting the safety of the derrick workers, and avoiding equipment damage and personal dangers caused by the fast lifting card.
Smart Images

Figure CN120088289A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the safety detection of the operation of drilling rig equipment, especially the field of detecting the excessive speed of elevator links. Specifically, it is a method for identifying, measuring the speed, and warning of the elevator link equipment on the second floor of the drilling rig based on the improved YOLOv5 and DeepSORT technologies. In particular, it is applied to the field of monitoring the movement of oilfield drilling elevator links and warning of excessive speed, belonging to the technical field of intelligent monitoring systems and their applications. Background Art
[0002] During the oilfield drilling process, the elevator link undertakes important operation tasks. The excessive speed of the elevator link may cause equipment damage and even endanger the safety of the staff. Therefore, it is of great significance to monitor the movement state of the elevator link in real time, especially to monitor the movement speed of the elevator link to protect the safety of the staff on the second floor and the normal operation of the equipment. The existing oilfield drilling monitoring systems usually rely on manual observation or traditional sensor monitoring, which have poor real-time performance, insufficient accuracy, and are only used for equipment safety protection, without considering the threat to the personal safety of the staff working on the second floor of the derrick due to the excessive speed of the elevator link.
[0003] With the development of computer vision and deep learning technologies, the intelligent detection and warning methods based on video monitoring have gradually become an effective way to solve such problems. Summary of the Invention
[0004] The present invention provides a method for identifying, measuring the speed, and warning of elevator links based on the improved YOLOv5 and DeepSORT. This method can use the improved YOLOv5 model to detect and identify the elevator links in the video through the real-time monitoring video of the second floor of the oilfield drilling rig, and transmit the identified elevator link information to the DeepSORT model for target tracking, so as to calculate the speed of the elevator link in real time. According to the set safety threshold of the elevator link speed, it judges the operating position of the derrickman. When it detects that the speed of the elevator link is too fast, it will issue an audible and visual alarm to remind and protect the personal safety of the derrickman. At the same time, it starts the dynamic braking system or the emergency stop (E-Stop) device.
[0005] The technical solution of the present invention includes the following steps:
[0006] S1: Obtain the elevator link data set: Obtain the picture data set of the elevator link equipment, label the target detection objects in the sample data set, and divide them into a training set, a test set, and a validation set;
[0007] S2: Target Detection Module: Build a lifting clamp target recognition network model based on the improved YOLOv5 model. The lifting clamp target recognition network model specifically includes Input, Backbone, Neck, and Output. The Backbone is an improved CSPDarknet network. The Neck includes FPN and PAN modules, and a shallow feature extraction layer is added. The lightweight operator CARAFE is introduced in the upsampling process of FPN to optimize the full-image semantic information in the upsampling process.
[0008] S3: Target Tracking Module: Establish a DeepSort lifting clamp tracking model. The DeepSort model is built to perform target tracking on the results of target detection based on the improved YOLOv5. When the position of the same target changes at different times, target association is carried out through the Hungarian algorithm and the Kalman filter algorithm.
[0009] S4: Speed Calculation Module: Calculate the speed of the detected lifting clamp. Through the target tracking trajectory of DeepSORT and combined with the timestamp information of each frame, calculate the real-time movement speed of the lifting clamp.
[0010] S5: Warning Module: Setting the speed safety threshold and alarming for over-speed of the lifting clamp. According to the on-site operation safety specifications, set the speed safety threshold of the lifting clamp. When the speed of the target exceeds the safety threshold, the system will trigger an audible and visual alarm, and at the same time start the dynamic braking system or the emergency stop device, and send an alarm signal to the monitoring center.
[0011] Furthermore, the steps for building the lifting clamp target recognition network model based on the improved YOLOv5 model in S2 are as follows:
[0012] S201: Introduce the GC (Global Context) module into the C3 module in the Backbone to build a new C3GC module. The GC module consists of three parts: global attention pooling, bottleneck feature transformation, and feature aggregation.
[0013] S202: Introduce the lightweight operator CARAFE in the upsampling process of FPN to replace the original Upsample operator to enhance the upsampling operation of the convolutional neural network feature map.
[0014] S203: Further process the feature map in the detection head of the Output part to generate the final detection result, and use non-maximum suppression to remove duplicate detection results.
[0015] S204: Train the lifting clamp target recognition network model: Use the training set to train the lifting clamp device detection network model, obtain the various parameters of the network model, and get the trained safety device detection network model. Use the test data set to test the trained lifting clamp device detection network model and evaluate the test results.
[0016] Further, in S3, establish the DeepSort lifting clamp tracking model: Specifically, it includes the following steps:
[0017] S301: First, initialize a Kalman filter for each detected target to predict the motion state of the target. The state vector of this Kalman filter includes position and velocity;
[0018] S302: Use the ReID deep learning model to extract feature vectors from the target image region for target association;
[0019] S303: Use the Hungarian algorithm to calculate the association between the detection results of the current frame and the predicted results of the tracker, and construct an association matrix based on the Mahalanobis distance and the deep feature distance;
[0020] S304: Use the Hungarian algorithm to solve the minimum value in the association matrix, that is, the optimal detection result. Find the optimal detection result to match the tracker;
[0021] S305: Update the state and covariance matrix of the Kalman filter according to the association results to improve the accuracy and robustness of multi-target tracking;
[0022] Further, in S4, calculate the use of the detected lifting clamp speed module, specifically including the following steps:
[0023] S401: Convert the upper left corner coordinates (x1, y1) and the lower right corner coordinates (x2, y2) of the target box into the center coordinates and width and height (x, y, w, h). The conversion formula is as shown in Equation (1):
[0024]
[0025] S402: Calculate the scale factor ratio according to the weighted ratio of the current target box width and the previous target box width, as shown in Equations (2) and (3):
[0026]
[0027] Among them, target_actual_width is the actual width of the target, target_distance is the target distance, focal_length is the camera focal length, w cur and w last are the widths of the target boxes in the current frame and the previous frame respectively;
[0028] S403: Calculate the time interval for each frame in seconds, and the formula is as shown in Equation (4):
[0029] t = 1.0 / fps (4)
[0030] S404: Calculate the speed of the elevating work platform according to the speed formula as shown in Equation (5):
[0031]
[0032] where is the pixel displacement between the target center points of the current frame and the previous frame. Multiplying by the scale factor ratio can convert the pixel displacement into the actual distance.
[0033] Furthermore, the use of the warning module in S5 specifically includes the following steps:
[0034] S501: Set a speed safety threshold for the elevating work platform according to the on-site operation safety specifications. When the speed of the target exceeds the safety threshold, the system will trigger an alarm mechanism;
[0035] S502: Set different alarm notification methods, such as local alarms, that is, sound and light alarms are carried out within the detection range, and at the same time, the dynamic braking system or the emergency stop device is activated. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a schematic diagram of the system layering provided by the embodiment of the present invention;
[0037] Figure 2 is a schematic diagram of the improved YOLOv5 model structure provided by the embodiment of the present invention;
[0038] Figure 3 is a schematic diagram of the Deepsort elevating work platform target tracking algorithm process provided by the embodiment of the present invention;
[0039] Figure 4 is an output result diagram of the elevating work platform target recognition based on the improved YOLOv5 model provided by the embodiment of the present invention;
[0040] Figure 5 is a speed measurement result diagram of the elevating work platform provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] The present invention provides a method for identifying, speed measuring and warning of elevator links based on improved YOLOv5 and DeepSORT, as Figure 1 shown, which specifically includes:
[0043] S1: Obtain an elevator link data set, label the target detection objects in the sample data set, and divide them into a training set, a test set and a validation set according to a ratio of 8:1:1;
[0044] S2: Target detection module: Use the YOLOv5 model for target detection, output the bounding box, confidence level and class probability, decode the predicted bounding box coordinates from the feature map space to the original image space, and perform non-maximum suppression on the predicted bounding boxes to remove duplicate detection results;
[0045] S3: Target tracking module: Use the DeepSORT algorithm to track the elevator link targets detected by YOLOv5;
[0046] S4: Speed calculation module: Calculate the moving speed of the elevator link through the target tracking trajectory of DeepSORT, combined with the timestamp information of each frame; Set a speed threshold for the elevator link according to the on-site operation safety specifications. When the speed of the target exceeds this threshold, the system will trigger an alarm mechanism and send an alarm signal to the monitoring center.
[0047] S5: Warning module: Setting of the speed safety threshold and alarm for over-speed of the elevator link. Set the speed safety threshold for the elevator link according to the on-site operation safety specifications. When the speed of the target exceeds the safety threshold, the system will trigger an audible and visual alarm, start the dynamic braking system or the emergency stop device at the same time, and send an alarm signal to the monitoring center.
[0048] In this solution: The YOLOv5 target detection algorithm is used to improve the detection accuracy, so that the system has a faster speed and higher detection accuracy. The DeepSORT is used as the target tracking algorithm, so that the system has strong tracking performance and realizes the ability of elevator link target tracking recognition and speed measurement.
[0049] The specific steps for obtaining the elevator link device picture data set in S1 are as follows:
[0050] S101: Use a high-resolution camera to collect video data of the drilling platform area in real time;
[0051] S102: Adjust the camera to ensure that all the areas required for monitoring the drilling platform are covered, avoid blind spots and overlaps, and use a bracket to maintain the stability of the camera;
[0052] S103: Use the OpenCV library to read the obtained video and extract frames (i.e., pictures);
[0053] S104: Use the Labelimg software to label the target positions in the pictures, add category labels, export the labeled pictures as xml format files, and finally divide them into a training set, a test set, and a validation set according to a ratio;
[0054] The use of the target detection module in S2 specifically includes the following steps:
[0055] S201: Use the backbone network CSPDarknet to extract the basic features of the video frames obtained by the data processing module, that is, use the backbone network to perform operations such as multi-layer convolution, normalization, and activation on the input image, and finally generate low-level, intermediate-level, and high-level feature maps of the input image;
[0056] Specifically in the CBS module, first input the picture x and perform the first-layer convolution operation. The specific operation is as shown in Equation (6):
[0057]
[0058] Among them, x is the input feature map with a size of H*W*C 1 ; C 1 is the number of input channels; C 2 is the number of output channels; k is the convolution kernel size; W is the convolution kernel weight with a size of
[0059] (g is the number of groups); b k is the bias term of the convolution; y is the output feature map with a size of H’*W’*C 2 .
[0060] Then, perform the batch normalization operation. The formula is as shown in Equation (7):
[0061]
[0062] Among them, x i is each element of the input feature map; μ i , are the mean and variance of the current batch of small samples respectively; γ and β are the learnable scaling and offset parameters respectively; ∈ is a stable constant to prevent division-by-zero errors.
[0063] Finally, perform the SiLU activation function operation. The formula is as shown in Equation (8):
[0064]
[0065] In the C3GC module, there are five steps of operations:
[0066] The first step: Given the input feature map X, it is divided into two parts:
[0067] X 1 ,X 2 = Split(X) (9)
[0068] The second step: The main branch starts to pass through the first GC module. First, global pooling is performed, and the formula is as follows:
[0069]
[0070] Context feature generation is performed, and the formula is as follows:
[0071] c 1 = Conv 1×1 (z 1 ) (11) Feature fusion is performed, and the formula is as follows:
[0072] X’ 1 = X 1 ·c 1 + X 1 (12) Then, after operating through the Bottleneck module, the main branch outputs, and the formula is as follows:
[0073] Y 1 = Bottleneck(X’ 1 ) (13)
[0074] The third step: The short branch passes through the second GC module, and global pooling is performed. The formula is as follows:
[0075]
[0076] Context feature generation is performed, and the formula is as follows:
[0077] c 2 = Conv 1×1 (z 2 ) (15)
[0078] Feature fusion is performed, and the formula is as follows:
[0079] X’ 2 = X 2 ·c 2 + X 2 (16)
[0080] Finally, the short branch outputs, and the formula is as follows:
[0081] Y 2 = X' 2 (17)
[0082] Step 4: Feature fusion of two branches, the formula is as follows:
[0083] Y = Concat(Y 1 , Y 2 ) (18)
[0084] Step 5: Channel compression, the formula is as follows:
[0085] Y out = Conv 1×1 (Y) (19)
[0086] After repeating the operations of multiple CBS modules and C3GC modules, high-level features of the input image are generated for subsequent detection tasks.
[0087] S202: Use a feature pyramid network to fuse features of different scales and improve the detection performance for small targets, that is, extract feature maps of multiple levels from the backbone network and fuse these feature maps through a top-down path and lateral connections;
[0088] S203: Use a detection head to further process the features to generate the final detection results, that is, bounding boxes, classes, and confidence levels, generate a series of bounding boxes and corresponding confidence levels and class probabilities, remove duplicate detection results, and use the non-maximum suppression method. The specific formula is as follows:
[0089]
[0090] For each pair of bounding boxes A and B, calculate their intersection over union (IoU). If IoU > threshold, keep the bounding box with the higher confidence level and suppress the bounding box with the lower confidence level;
[0091] S204: Use a test dataset to evaluate the trained detection network model for hanging card devices. The specific evaluation formula is as follows:
[0092]
[0093] Among them, P is the precision rate, R is the recall rate, mAP is the mean average precision, TP is the number of positive samples correctly classified, FP is the number of negative samples misjudged as positive samples, FN is the number of positive samples misjudged as negative samples, n is the number of samples in the test set, P(i) is the size of the precision rate P when simultaneously identifying i samples, ΔR(i) is the change in the recall rate when detecting i samples, and C is the total number of types;
[0094] The specific steps of S202 are as follows:
[0095] S2021: The backbone network outputs feature maps at three scales: X 1 , X 2 , X 3 , where X 1 is the feature map of the topmost layer, with the lowest resolution and the most channels; X 3 is the feature map of the bottommost layer, with the highest resolution and the fewest channels. The operation of the top-down path is to gradually upsample the upper-layer feature map to a higher resolution and fuse it with the lower-layer feature map. The specific upsampling formula is as follows:
[0096]
[0097] S2022: To ensure that features at different scales can be fused with the same number of channels, for each feature map X i , lateral connections are used to adjust the number of channels through 1×1 convolution. The specific formula is:
[0098] L i = Conv 1×1 (X i ) (25)
[0099] S2023: Then, the feature map X i with adjusted number of channels is fused with the lowermost feature map after upsampling. The specific formula is as follows:
[0100]
[0101] S2024: Then, lateral connections are used to fuse the bottom-up feature map with the top-down feature map and finally output fused feature maps at multiple scales. The bottom-up path will perform an addition operation on the feature map transmitted from the low layer to the high layer and the upsampled high-layer feature map. The specific formula is as follows:
[0102]
[0103] The use of the target tracking module in S3 specifically includes the following steps:
[0104] S301: First, initialize a Kalman filter for each detected target to predict the motion state of the target. The state vector of this Kalman filter includes position and velocity,
[0105]
[0106] where u k , v k are the center coordinates of the detection box respectively, r k is the aspect ratio of the detection box, and h k is the height of the detection box. They are the speeds at the center positions of the detection boxes, They are the aspect ratio and the rate of change of the height respectively. The Kalman filter predicts the position and speed of the target at the next moment through the state transition model, and its formula is as follows:
[0107] X k = F·X k-1 + w k (29)
[0108] Among them, X k is the state vector at the current moment, F is the state transition matrix, X k-1 is the state vector at the previous moment, and w k is the process noise. The measurement value z k of the target can be directly obtained from the detection box provided by YOLOv5, and z k includes u k , v k , r k , h k . The measurement model formula is as follows:
[0109] z k = H·X k + γ k (30)
[0110] Among them, H is the measurement matrix, and γ k is the measurement noise;
[0111] S302: Use the ReID deep learning model to extract feature vectors from the target image region for target association;
[0112] S303: Use the Hungarian algorithm to perform target association between the detection results of the current frame and the prediction results of the tracker. The construction of the association matrix is based on the Mahalanobis distance and the deep feature distance. The specific formula for the feature distance is as follows;
[0113]
[0114] Among them, z i is the measurement value of the i-th detection box, is the state vector predicted by the j-th tracker, and S is the measurement residual covariance matrix;
[0115] Then, calculate the feature distance through the cosine distance, and the formula is as follows:
[0116]
[0117] Construct the overall association matrix, and the formula is as follows:
[0118] d(i,j) = α·dmahalanobis (i,j)+(1-α)d feature (i,j) (33)
[0119] Among them, α is the weight coefficient, which balances the influence of the position and appearance features.
[0120] S304: Use the Hungarian algorithm to find the best match between the target detection box and the tracking target on the association matrix.
[0121] S305: Correct the state estimation of the target through the update step of the Kalman filter. The specific update formula is as follows:
[0122]
[0123] Among them, K k is the Kalman gain, and P k is the updated error covariance matrix. Finally, the accurate tracking state of the target is obtained;
[0124] The use of the speed calculation module in S4 is as follows:
[0125] S401: Convert the upper left corner coordinates (x1, y1) and the lower right corner coordinates (x2, y2) of the target box into the center coordinates and width and height (x, y, w, h). Among them, the conversion formula is:
[0126]
[0127] S402: By updating the actual width of the target in real time, calculate the scale factor according to the weighted ratio of the current target box width and the previous target box width:
[0128]
[0129] S403: Calculate the time interval of each frame in seconds. The formula is as follows:
[0130] t = 1.0 / fps (40)
[0131] S404: Calculate the speed of the elevator clamp according to the speed formula. The formula is as follows:
[0132]
[0133] The use of the warning module in S5 specifically includes the following steps:
[0134] S501: Set a speed safety threshold for the elevator clamp according to the on-site operation safety specifications. When the speed of the target exceeds the safety threshold, the system will trigger the alarm mechanism;
[0135] S502: Set different alarm notification methods, such as local alarm, that is, perform acoustic and optical alarms within the detection range, simultaneously activate the dynamic braking system or emergency stop device, and conduct remote notification to relevant staff via text messages, emails, or a dedicated application.
[0136] Integrating the improved YOLOv5 and DeepSORT algorithms into the monitoring system can realize the safety early warning and supervision system.
[0137] Trigger the early warning mechanism by identifying the over-speed of the elevator spider, and feedback the situation to the user to achieve the purpose of maintaining the safety of the well site.
Claims
1. A method for detecting elevator speed based on improved YOLOv5 and Deepsort, characterized in that: The following steps are involved: S1: Prepare the data set: obtain the elevator equipment image data set, annotate the target detection objects in the sample data set, and divide it into training set, test set and validation set; S2: Construct an elevator target recognition network model based on the improved YOLOv5 model: the elevator target recognition network model specifically includes input Input, backbone network Backbone, neck Neck and output Output; the backbone network Backbone is an improved CSPDarknet network; the neck Neck includes FPN and PAN modules, and a shallow feature extraction layer is added, and a lightweight operator CARAFE is introduced in the FPN upsampling process to optimize the full-image semantic information of the upsampling process; S3: Establish DeepSort Elevator Tracking Model: Construct DeepSort model to track the target based on the improved YOLOv5 target detection results. When the position of the same target changes at different times, the target is associated through Hungarian algorithm and Kalman filter algorithm; S4: Calculate the speed of the detected elevator: Calculate the real-time movement speed of the elevator through the target tracking trajectory of DeepSORT and the timestamp information of each frame; S5: Speed safety threshold setting and elevator overspeed alarm: Set the speed safety threshold of the elevator according to the on-site operation safety regulations. When the speed of the target exceeds the safety threshold, the system will trigger the sound and light alarm, start the dynamic braking system or emergency stop device, and send an alarm signal to the monitoring center.
2. The elevator speed detection method based on improved YOLOv5 and Deepsort according to claim 1, characterized in that: The specific method of step S1 is as follows: S101: Use high-resolution cameras to collect video data of the drilling platform area in real time; S102: Adjust the camera to ensure that it covers all the areas that need to be monitored on the drilling platform, avoid blind spots and overlaps, and use a bracket to maintain the stability of the camera; S103: Use the OpenCV library to read the obtained video and extract frames (i.e., pictures); S104: Use Labelimg software to mark the target position in the image and add category labels, export the marked image into an XML format file, and finally divide it into a training set, a test set, and a validation set according to the proportion.
3. The elevator speed detection method based on improved YOLOv5 and Deepsort according to claim 2, characterized in that: The specific method of step S2 is as follows: S201: Use the backbone network CSPDarknet to extract the basic features of the video frame obtained by the data processing module, that is, use the backbone network to perform multi-layer convolution, normalization and activation operations on the input image, and finally generate low-level, medium-level and high-level feature maps of the input image; S202: Use a feature pyramid network to fuse features of different scales to improve the detection performance of small targets, that is, extract feature maps of multiple levels from the backbone network and fuse these feature maps through top-down paths and lateral links; S203: using the detection head to further process the features and generate the final detection results, i.e., bounding boxes, categories, and confidences, to generate a series of bounding boxes and corresponding confidences and category probabilities, remove duplicate detection results, and use a non-maximum suppression method; S204: Using the test data set to evaluate the trained elevator equipment detection network model.
4. The elevator speed detection method based on improved YOLOv5 and Deepsort according to claim 3, characterized in that: The use of the target tracking module in S3 includes the following steps: S301: First, a Kalman filter is initialized for each detected target to predict the motion state of the target, and the state vector of the Kalman filter includes position and velocity; S302: extracting feature vectors from the target image region using a ReID deep learning model for target association; S303: Use the Hungarian algorithm to perform target association between the detection result of the current frame and the prediction result of the tracker; S304: Use the Hungarian algorithm to find the best target detection box and tracking target match on the correlation matrix; S305: Correct the target state estimation through the update step of the Kalman filter, and finally obtain the accurate tracking state of the target.
5. The elevator speed detection method based on improved YOLOv5 and Deepsort according to claim 4, characterized in that: The specific steps for using the S4 elevator speed module are as follows: S401: Convert the upper left corner coordinates (x1, y1) and the lower right corner coordinates (x2, y2) of the target box to the center coordinates and width and height (x, y, w, h); S402: Calculate a proportional factor according to a weighted ratio of a current target frame width and a previous target frame width by updating the target actual width in real time; S403: Calculate the time interval of each frame in seconds; S404: Calculate the speed of the elevator according to the speed formula.
6. The elevator speed detection method based on improved YOLOv5 and Deepsort according to claim 5, characterized in that: The use of the early warning mechanism in S5 specifically includes the following steps: S501: According to the on-site operation safety regulations, set a speed safety threshold of the elevator. When the speed of the target exceeds the safety threshold, the system will trigger the alarm mechanism; S502: Setting different alarm notification modes, such as local alarm, that is, sound and light alarm within the detection range, and starting the dynamic braking system or emergency stop device at the same time.