Method for measuring distance between people around excavator based on improved YOLOv3 and binocular vision

By improving the YOLOv3 deep learning convolutional network and binocular vision technology, the problem that remote-controlled excavators have difficulty identifying the distance to surrounding personnel has been solved. This has achieved high recognition rate and accurate distance measurement of people around the excavator, improving safety.

CN114140609BActive Publication Date: 2025-09-05SHANDONG CHANGLIN MACHINERY GRP +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111363241.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-17
Publication Date
2025-09-05
Estimated Expiration
2041-11-17

AI Technical Summary

Technical Problem

It is difficult for a remote-controlled excavator to accurately identify and calculate the distance to people around the excavator, resulting in safety hazards.

Method used

The improved YOLOv3 deep learning convolutional network model and binocular vision technology are used to calculate the distance between the excavator and the personnel through image enhancement, annotation, model training and feature point matching.

Benefits of technology

It achieves high recognition rate and accurate distance measurement of people around the excavator, improving the safety of remote-controlled excavators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140609B_ABST
    Figure CN114140609B_ABST
Patent Text Reader

Abstract

The present invention provides a method for measuring the distance between people in the vicinity of an excavator based on an improved YOLOv3 and binocular vision. The method comprises the following steps: 1. collecting images containing people in the excavator's working environment and establishing an image database using an image enhancement method; 2. labeling the people in the images; 3. improving a YOLOv3 deep learning convolutional network model based on the actual operating conditions of the excavator; 4. training the model and selecting the optimal model; 5. detecting and matching feature points in the left and right views; and 6. calculating the distance between the excavator and the people. The method, based on the improved YOLOv3 and binocular vision, allows a remotely controlled excavator to identify people in the vicinity and calculate the distance between the people and the excavator during operation, thereby preventing safety accidents during the operation of the excavator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of excavator remote control, and in particular to a method for measuring the distance between people around an excavator based on improved YOLOv3 and binocular vision. Background Art

[0002] When a remote-controlled excavator is operating, the presence of people in its surroundings is unavoidable. If the remote operator fails to monitor those nearby, accidents can occur. While most remote-controlled excavators are equipped with cameras around the excavator, allowing the operator to monitor the excavator's surroundings in real time, it can be difficult to accurately determine the distance between the excavator and the operator. Therefore, identifying nearby personnel and calculating their distance from the excavator can provide a reference for remote operation.

[0003] US Patent Application No. 716526 relates to a construction machinery alarm system. The excavator includes an upper structure swingably supported by a lower structure. The working position sensor consists of multiple radio frequency transceivers installed in the excavator and one carried by each excavator worker working within the excavator's working range. A signal processing unit determines whether the relative distance between each worker and the excavator is short or long, and identifies each worker's position relative to each predetermined identification zone. The determination signal from the signal processing unit is provided to a control unit. The control unit is connected to a drive unit that includes electro-hydraulic proportional valves for actuators used to position, swing, and travel the excavator. The control unit is also connected to machine sensors including a swing angle sensor, a travel level sensor, and a swing lever sensor. The control unit determines whether the machine is moving close to a worker. The control unit controls the machine so that the lower structure stops or moves slowly, or, when the machine approaches a worker, the upper structure swings slowly, so that the machine's motion remains constant when moving away from the worker.

[0004] Chinese patent application number CN201811518087.6 describes a facial recognition system and method for intelligent excavators, belonging to the field of intelligent excavator application technology. The system comprises a video acquisition module, a face detection module, a face recognition module, a distance measurement module, an execution module, and a master control module. The video acquisition module first acquires real-time image information, then performs face detection using a trained face detection model. The aligned face image is then input into the trained face recognition model for face recognition and distance calculation. Finally, the execution module determines whether the person is a worker and whether the person is within a safe distance, issuing an alert. This invention enables real-time monitoring of personnel at the construction site, promptly detecting any personnel approaching the excavator and implementing avoidance measures to ensure construction safety. The detection algorithm achieved an 84.23% detection accuracy on the FDDB dataset and a 99% recognition accuracy on a self-built face data recognition dataset.

[0005] Chinese patent application number CN201310176951.X describes an onboard video surveillance system for construction machinery with face detection. This system monitors the working environment, identifies the presence of people within the environment, and proactively issues warning signals, improving on-site safety. The system specifically includes a main controller that analyzes onboard data to determine the machine's operating status; a data acquisition module, including a video monitoring device and a status monitoring device. The former uses a camera placed in the driver's blind spot to collect working environment information, while the latter uses sensors to collect operating data from the powertrain, hydraulic system, and electronic control system. A face detection module processes the collected video data, identifies the presence of people, marks their faces, and issues warning signals. A human-machine interaction module displays the onboard and video data on an LCD screen, providing signals indicating personnel warnings, fault information, and operating mode.

[0006] The above existing technologies are significantly different from the present invention and fail to solve the technical problem we want to solve. To this end, we invented a new method for measuring the distance between people around an excavator based on improved YOLOv3 and binocular vision. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for measuring the distance between people around an excavator based on improved YOLOv3 and binocular vision, which can identify people around the excavator and calculate the distance between the people and the excavator during the operation of a remote-controlled excavator, thereby avoiding safety accidents during the operation of the excavator.

[0008] The purpose of the present invention can be achieved by the following technical measures: a method for measuring the distance between people around an excavator based on improved YOLOv3 and binocular vision, the method for measuring the distance between people around an excavator based on improved YOLOv3 and binocular vision includes:

[0009] Step 1: Collect images containing people in the excavator working environment and build an image database using image enhancement methods;

[0010] Step 2: Label the person targets in the image;

[0011] Step 3: Improve the YOLOv3 deep learning convolutional network model based on the actual operation conditions of the excavator;

[0012] Step 4: Train the model and select the optimal model;

[0013] Step 5: Detect and match feature points of the left and right views;

[0014] Step 6: Calculate the distance between the excavator and the human target.

[0015] The purpose of the present invention can also be achieved by the following technical measures:

[0016] In step 1, images containing people are captured in the excavator working environment. According to the actual working environment of the excavator, the subjects to be captured should include children, young people, middle-aged people, and elderly people; the postures of the subjects to be captured should include squatting, standing, front, and back; and the shooting environment should include sunny, cloudy, and rainy days.

[0017] In step 1, the collected images are enhanced, including adjusting the brightness and saturation of the images, and performing different degrees of scaling and rotation operations on the images. Through the above operations, an image database of people around the excavator is established.

[0018] In step 2, use the image annotation tool labelImg to annotate the people in the image. During the annotation process, use a rectangular box to frame a person in the image and mark it as the person class. The number of people in the image should be equal to the number of annotated boxes, and the annotated rectangular boxes should be of the appropriate size to the people in the image.

[0019] In step 2, after an image is annotated, a corresponding XML file will be generated. This file includes the image name, storage location, image width, image height, and annotation information. The annotation information includes the category of each person in the image, the xy coordinates of the upper left corner of the annotation box, and the xy coordinates of the lower right corner of the annotation box.

[0020] In step 3, the YOLOv3 deep learning convolutional network model resizes the input images of different resolutions to a uniform size of 416×416, and then evenly divides the image into S×S squares. The center point of each square is used to predict the target within the square; when it is considered that a target exists, the model will further calculate the position coordinates, height and width of the predicted box, and the confidence level; when the confidence level is greater than a certain threshold, the model will consider the target to be a real target.

[0021] In step 3, the model defaults to S of 13, 26, and 52. 13×13 is used to predict large targets, 26×26 is used to predict medium targets, and 52×52 is used to predict small targets. Since this ranging method is actually used to detect people close to the excavator, people farther away from the excavator, that is, smaller targets in the image, can be ignored. To address this situation, the 52×52 structure for detecting small targets in the YOLOv3 model is removed, and large and medium targets in the image are directly detected.

[0022] In step 4, during the initial training process, the model compares the predicted information with the actual labeled information and calculates the loss value. During the continuous iteration process, the parameters are automatically adjusted to reduce the loss value until the loss value converges. Finally, the model with the smallest loss value is selected to predict the target. After that, before each excavator work, the system will automatically train the images with human targets saved during the previous work process and automatically select the model with the smallest loss value.

[0023] In step 5, the SUFT algorithm is used to detect and match the feature points in the prediction frames of the left and right camera images. There will be many incorrect matching points in the feature point matching process. To this end, the following matching conditions need to be met:

[0024]

[0025] Among them, v l 、v r are the ordinates of the corresponding feature points in the left and right views, N and M are the number of prediction boxes in the left and right views, and I lj Used to determine whether the feature point in the left view is in the jth prediction frame, I rk It is used to determine whether the feature point in the right view is in the kth prediction frame, and T is the matching threshold.

[0026] In step 6, the parallax is calculated based on multiple pairs of matched feature points of the same target and the principle of binocular vision ranging, and the average value of the parallax is taken as the distance between the excavator and the human target. The calculation formula is as follows:

[0027]

[0028]

[0029]

[0030] Among them, (u1, v1) and (u2, v2) are the coordinates corresponding to the pixel coordinate system in the left and right views respectively, (X w , Y w , Z w ) is the point coordinate in the world coordinate system, R l 、R r are the rotation matrices from the world coordinate system to the left and right camera coordinate systems, T l 、T r are the translation matrices from the world coordinate system to the left and right camera coordinate systems, K l , K r are the left and right camera intrinsic parameter matrices, respectively, f x 、f y are the focal lengths of the left and right cameras, respectively; B is the distance between the left and right camera baselines; L is the distance between the excavator and the human target; and n is the number of matching feature points of the same target in the left and right views.

[0031] This method, based on an improved YOLOv3 and binocular vision, is designed to identify people in the vicinity of a remotely controlled excavator and calculate their distance to the excavator during operation, thus preventing accidents. The method involves collecting images containing people in the excavator's working environment and establishing an image database using image enhancement methods; labeling the people in the images using labelImg; improving the YOLOv3 deep learning convolutional network model; training the model and selecting the optimal one; detecting and matching feature points in the left and right views; and calculating the distance between the excavator and the people. This method achieves a high recognition rate for people and provides accurate distance measurement, ensuring the safety of people around the excavator. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a flowchart of a specific embodiment of the method for measuring distance between people around an excavator based on improved YOLOv3 and binocular vision of the present invention;

[0033] Figure 2 This is a structural diagram of an improved YOLOv3 deep learning convolutional network model in a specific embodiment of the present invention;

[0034] Figure 3 Schematic diagram of binocular vision ranging in a specific embodiment of the present invention;

[0035] Figure 4 This is a curve diagram of the change of model training loss value in a specific embodiment of the present invention. DETAILED DESCRIPTION

[0036] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0037] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations and / or combinations thereof.

[0038] like Figure 1 As shown, Figure 1 This is a flow chart of the method for measuring the distance between people around an excavator based on improved YOLOv3 and binocular vision. The method for measuring the distance between people around an excavator based on improved YOLOv3 and binocular vision includes:

[0039] (1) Collect images containing people in the excavator working environment and establish an image database through image enhancement methods;

[0040] (2) Use labelImg to label the person targets in the image;

[0041] (3) Improve the YOLOv3 deep learning convolutional network model;

[0042] (4) Train the model and select the optimal model;

[0043] (5) Left and right view feature point detection and matching;

[0044] (6) Calculate the distance between the excavator and the human target.

[0045] The following are several specific embodiments of the present invention.

[0046] Example 1

[0047] In a specific embodiment 1 of the present invention, the method for measuring the distance between people around an excavator based on improved YOLOv3 and binocular vision includes:

[0048] Step 1: Collect images containing people in the excavator's working environment and establish an image database using image enhancement methods. Use a mobile phone or industrial camera to capture images containing people in the excavator's working environment. Based on the actual working environment of the excavator, the subjects to be photographed should include children, young people, middle-aged people, and the elderly; the postures of the subjects to be photographed should include squatting, standing, front view, and back view; and the shooting environment should include sunny, cloudy, and rainy days. To further improve the adaptability of the model, the collected images are enhanced, including adjusting the image brightness and saturation, and performing various degrees of scaling and rotation operations. Through these operations, a database of images of people around the excavator is established.

[0049] Step 2: Use labelImg to label the people in the image. Use the image labeling tool labelImg to label the people in the image. During the labeling process, a person in the image is framed with a rectangular box and labeled as "person." The number of people in the image should be equal to the number of labeled boxes, and the labeled rectangular boxes should be the appropriate size for the people in the image. After an image is labeled, an XML file is generated. This file includes the image name, storage location, image width, image height, and labeling information. The labeling information includes the "person" category corresponding to each person in the image, the xy coordinates of the upper left corner of the labeled box, and the xy coordinates of the lower right corner of the labeled box.

[0050] Step 3: Improve the YOLOv3 deep learning convolutional network model; this model is modified for the actual operation of an excavator. The YOLOv3 deep learning convolutional network model resizes input images of varying resolutions to a uniform 416×416 size. It then evenly divides the image into S×S squares, with the center point of each square used to predict the object within. When an object is identified, the model further calculates the location coordinates, height and width of the predicted box, and the confidence score. When the confidence score exceeds a certain threshold, the model considers the object to be a true target. The model defaults to S of 13, 26, and 52: 13×13 is used to predict large objects, 26×26 is used to predict medium objects, and 52×52 is used to predict small objects. Since this ranging method is actually used to detect people in the vicinity of an excavator, people farther away from the excavator—those smaller objects in the image—can be ignored. To address this situation, the 52×52 small object detection structure in the YOLOv3 model is removed, allowing direct detection of large and medium objects in the image. from Figure 2As can be seen, the model resizes the image to 416×416 pixels. Since the image is scaled to 416 pixels based on the longest side of the input image, the shorter side of the image is smaller than 416 pixels. The model automatically pads the image to 416×416 pixels with grayscale bars. The image then passes through the Darknet-53 feature extraction network. The features output from res4 undergo six layers of DBL and one convolution, outputting 13×13 features. These are used to predict large objects closer to the excavator. Separately, the image passes through the Darknet-53 feature extraction network, six layers of DBL, and upsampling. The features are then concatenated with the features output from res8 in the feature extraction network. Six layers of DBL and one convolution are then passed through to output 26×26 features. These are used to predict medium-sized objects in the image. DBL includes convolution, batch normalization, and the Leaky ReLU activation function. res unit is the residual block, which is the sum of the features obtained after two layers of DBL. resn includes zero padding, DBL, and n layers of res unit.

[0051] Step 4: Train the models and select the optimal one. Improve the YOLOv3 deep learning convolutional network model. Train the models and select the optimal one. During the initial training process, the model compares the predicted information with the actual annotated information and calculates the loss value. During the iterative process, the model automatically adjusts the parameters to reduce the loss value until the loss value converges. The model with the lowest loss value is ultimately selected to predict the target. Thereafter, before each excavator operation, the system automatically trains on images of people saved from the previous operation and automatically selects the model with the lowest loss value.

[0052] Step 5: Left and right view feature point detection and matching. Use the SUFT algorithm to detect and match the feature points in the prediction frames of the left and right camera images. There will be many incorrect matching points during the feature point matching process. To this end, the following matching conditions need to be met:

[0053]

[0054] Among them, v l 、v r are the ordinates of the corresponding feature points in the left and right views, N and M are the number of prediction boxes in the left and right views, and I lj Used to determine whether the feature point in the left view is in the jth prediction frame, I rk It is used to determine whether the feature point in the right view is in the kth prediction frame, and T is the matching threshold.

[0055] Step 6: Calculate the distance between the excavator and the human target.

[0056] The parallax is calculated based on multiple pairs of matched feature points of the same target and the principle of binocular vision ranging, and the average value of the parallax is taken as the distance between the excavator and the human target. The calculation formula is as follows:

[0057]

[0058]

[0059]

[0060] like Figure 3 As shown, (u1, v1) and (u2, v2) are the coordinates corresponding to the pixel coordinate system in the left and right views respectively, (X w , Y w , Z w ) is the point coordinate in the world coordinate system, R l 、R r are the rotation matrices from the world coordinate system to the left and right camera coordinate systems, T l 、T r are the translation matrices from the world coordinate system to the left and right camera coordinate systems, K l , K r are the left and right camera intrinsic parameter matrices, respectively, f x 、f y are the focal lengths of the left and right cameras, respectively; B is the distance between the left and right camera baselines; L is the distance between the excavator and the human target; and n is the number of matching feature points of the same target in the left and right views.

[0061] Example 2:

[0062] In the specific embodiment 2 of the present invention, the number of model training iterations is too low, and the obtained model detection effect is poor. Figure 4 As can be seen from the figure, when the number of iterations is less than 100, the model's loss value continues to decrease but has not converged. To verify the model's detection performance, models with 50 and 100 iterations were selected for target detection and comparison. The results show that the detection accuracy rates of the two models are 61.3% and 79.5%, respectively.

[0063] Example 3:

[0064] In the specific embodiment 3 of the present invention, during the model training process, when the loss value converges, the obtained model detection effect is better. Figure 4 As can be seen from the figure, the loss converges after 150 iterations. The model with 200 iterations was selected for target detection, and the results showed that the detection accuracy was 98.3%.

[0065] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art may modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features therein. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

[0066] Except for the technical features described in the specification, all other technical features are known technologies to those skilled in the art.

Claims

1. The method of measuring distance between people around an excavator based on improved YOLOv3 and binocular vision is characterized by: The method for measuring the distance between people around an excavator based on improved YOLOv3 and binocular vision includes: Step 1: Collect images containing people in the excavator working environment and build an image database using image enhancement methods; Step 2: Label the person targets in the image; Step 3: Improve the YOLOv3 deep learning convolutional network model based on the actual operation conditions of the excavator; Step 4: Train the model and select the optimal model; Step 5: Detect and match feature points of the left and right views; Step 6, calculate the distance between the excavator and the human target; In step 6, the parallax is calculated based on multiple pairs of matched feature points of the same target and the principle of binocular vision ranging, and the average value of the parallax is taken as the distance between the excavator and the human target. The calculation formula is as follows: Among them, (u1, v1) and (u2, v2) are the coordinates corresponding to the pixel coordinate system in the left and right views respectively, (X w , Y w , Z w ) is the point coordinate in the world coordinate system, R l 、R r are the rotation matrices from the world coordinate system to the left and right camera coordinate systems, T l 、T r are the translation matrices from the world coordinate system to the left and right camera coordinate systems, K l , K r are the left and right camera intrinsic parameter matrices, respectively, f x 、f y are the focal lengths of the left and right cameras, respectively; B is the distance between the left and right camera baselines; L is the distance between the excavator and the human target; and n is the number of matching feature points of the same target in the left and right views.

2. The method for measuring distance between people around an excavator based on improved YOLOv3 and binocular vision according to claim 1, characterized in that: In step 1, images containing people are captured in the excavator working environment. According to the actual working environment of the excavator, the subjects to be captured should include children, young people, middle-aged people, and elderly people; the postures of the subjects to be captured should include squatting, standing, front, and back; and the shooting environment should include sunny, cloudy, and rainy days.

3. The method for measuring distance between people around an excavator based on improved YOLOv3 and binocular vision according to claim 2, characterized in that: In step 1, the collected images are enhanced, including adjusting the brightness and saturation of the images, and performing different degrees of scaling and rotation operations on the images. Through the above operations, an image database of people around the excavator is established.

4. The method for measuring distance between people around an excavator based on improved YOLOv3 and binocular vision according to claim 1, characterized in that: In step 2, use the image annotation tool labelImg to annotate the people in the image. During the annotation process, use a rectangular box to frame a person in the image and mark it as the person class. The number of people in the image should be equal to the number of annotated boxes, and the annotated rectangular boxes should be of the appropriate size to the people in the image.

5. The method for measuring distance between people around an excavator based on improved YOLOv3 and binocular vision according to claim 4 is characterized in that: In step 2, after an image is annotated, a corresponding XML file will be generated. This file includes the image name, storage location, image width, image height, and annotation information. The annotation information includes the category of each person in the image, the xy coordinates of the upper left corner of the annotation box, and the xy coordinates of the lower right corner of the annotation box.

6. The method for measuring distance between people around an excavator based on improved YOLOv3 and binocular vision according to claim 1, characterized in that: In step 3, the YOLOv3 deep learning convolutional network model resizes the input images of different resolutions to a uniform size of 416×416, and then evenly divides the image into S×S squares. The center point of each square is used to predict the target within the square; when it is considered that a target exists, the model will further calculate the position coordinates, height and width of the predicted box, and the confidence level; when the confidence level is greater than a certain threshold, the model will consider the target to be a real target.

7. The method for measuring distance between people around an excavator based on improved YOLOv3 and binocular vision according to claim 6, characterized in that: In step 3, the model defaults to S of 13, 26, and 52. 13×13 is used to predict large targets, 26×26 is used to predict medium targets, and 52×52 is used to predict small targets. Since this ranging method is actually used to detect people close to the excavator, people farther away from the excavator, that is, smaller targets in the image, can be ignored. To address this situation, the 52×52 structure for detecting small targets in the YOLOv3 model is removed, and large and medium targets in the image are directly detected.

8. The method for measuring distance between people around an excavator based on improved YOLOv3 and binocular vision according to claim 1, characterized in that: In step 4, during the initial training process, the model compares the predicted information with the actual labeled information and calculates the loss value. During the continuous iteration process, the parameters are automatically adjusted to reduce the loss value until the loss value converges. Finally, the model with the smallest loss value is selected to predict the target. After that, before each excavator work, the system will automatically train the images with human targets saved during the previous work process and automatically select the model with the smallest loss value.

9. The method for measuring distance between people around an excavator based on improved YOLOv3 and binocular vision according to claim 1, characterized in that: In step 5, the SUFT algorithm is used to detect and match the feature points in the prediction frames of the left and right camera images. There will be many incorrect matching points in the feature point matching process. To this end, the following matching conditions need to be met: Among them, v l 、v r are the ordinates of the corresponding feature points in the left and right views, N and M are the number of prediction boxes in the left and right views, and I lj Used to determine whether the feature point in the left view is in the jth prediction frame, I rk Used to determine whether the feature point in the right view is in the kth prediction frame, and T is the matching threshold.

Citation Information

Patent Citations

  • Engineering machinery onboard video monitoring system with face detection function

    CN103324158A

  • Syringe-bulb.

    US716526A

  • Face recognition system and method for an intelligent excavator

    CN109657592A

  • Left and right view feature point matching processing method, terminal and system in binocular ranging

    CN111882618A