A visual-based monitoring method for a working pose of an excavator

By using a monocular camera and vision technology to monitor the excavator's working posture, the problems of easy sensor damage and low algorithm accuracy in traditional methods are solved, enabling low-cost, real-time excavator status monitoring and management.

CN115588043BActive Publication Date: 2026-03-27HUNAN PROVINCE LAND & RESOURCES PLANNING INST +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing methods for monitoring the working posture of excavators suffer from problems such as easily damaged sensors, high maintenance costs, low algorithm accuracy, and limited applicability. They are difficult to achieve accurate monitoring in harsh environments, especially in small or temporary work areas where it is difficult to deploy equipment.

Method used

The system uses a monocular camera to capture real-time video of the excavator. The positions of the bucket and boom are defined by a trained model. Combined with color image back projection or Unet image segmentation technology, the operating pose of the excavator is calculated. No calibration or target matching is required, thus improving monitoring accuracy.

Benefits of technology

It enables low-cost, real-time monitoring of excavator operating posture, accurately identifies excavator status in harsh environments, reduces equipment deployment difficulty and maintenance costs, and supports unified online management for excavator owners.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115588043B_ABST
    Figure CN115588043B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on vision's excavator operation pose monitoring method.The steps are: using monocular camera real-time acquisition excavator, small arm and the video of excavating surface, frame the position of excavator and small arm in monitoring video, obtain the image coordinate information of target detection frame;The image in small arm target detection frame is segmented;Identify the position and size of small arm in image;Then according to the real size of small arm and excavator, and the relative position and size of small arm and excavator in image, the distance of small arm, excavator and camera is fitted to obtain, the position of excavator and cab is determined, and action state;The position of excavating point and unloading point of excavator is calculated by positioning device, and the dynamic monitoring of excavator operation pose is realized.The application can realize the online monitoring of administrative department to the development and utilization of natural resources by excavator and other engineering equipment, and the online unified management of excavator owner to its equipment, real-time master the working state of its equipment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of excavators, and particularly relates to a method for monitoring the working pose of an excavator based on vision. BACKGROUND

[0002] An excavator is one of the most widely used engineering machines, and as a mainstream product of engineering machines, it plays an extremely important role in industrial and civil construction, transportation, water conservancy and power engineering, mining and military engineering construction. In the development and utilization of land resources, excavators are also the most important natural resource development and utilization engineering equipment. Strengthening the management and dynamic monitoring of engineering equipment operation, realizing the dynamic identification of the position and attitude of engineering equipment, can effectively identify the development and utilization connotation of engineering equipment, and further realize the strengthening of natural resource supervision and protection, the improvement of the intensive utilization degree of natural resources, and the improvement of the modernization level of natural resource governance ability and means.

[0003] In recent years, a large number of enterprises and individuals at home and abroad have carried out research work on the pose identification and positioning of excavators. For example, Caterpillar, Komatsu, Volvo, Hitachi, Doosan, and domestic engineering machinery manufacturers such as Xugong, Sany, Zoomlion, Shanhe Intelligence, and LiuGong have carried out a lot of research work on excavator pose identification and positioning to realize automatic operation, and some products have realized automatic operation or remote control operation. At present, the identification of the pose of the excavator is mainly completed by the GPS navigation system to position the machine body, and the action pose is identified by the wire pull sensor installed on the boom, dipper and bucket or the inclination sensor or optical encoder installed at the hinge. Common inclination measurement or distance measurement methods include ultrasonic ranging, laser pulse ranging, infrared ranging, optical ranging and stereoscopic ranging. In the process of obtaining the pose information of the excavator by this method, in the actual working process of the excavator, due to the relatively harsh working environment, the boom, dipper and bucket will inevitably contact and collide with the soil, rock or water body, and the sensors installed thereon are easy to be damaged and corroded; at the same time, the sensors installed thereon are also easy to be damaged in the long-term vibration environment; in addition, these sensors are relatively high in cost and high in maintenance cost.

[0004] Wang Haibo et al. proposed an invention patent of "non-contact excavator working device posture measurement method based on visual measurement" (201410079270.6) in 2014. A circular sheet with obvious image saddle point features is pasted on the working device as a marker feature point, then an industrial camera on the cab frame is used to take an image of the excavator working device, and a saddle point detection method is applied to detect all image saddle points on the working device image including the feature sheet center point in real time. The image saddle points of non-feature sheet center points are filtered out through the distance between the feature sheets, and the inclination angle of the connecting line between the feature sheet center points is used to obtain the inclination angle of each component of the working device, so as to measure the posture of the working device. Although this method can effectively monitor the posture of the excavator, the working environment of the excavator is relatively harsh, and the marker feature points on the working device are easy to be contaminated, resulting in monitoring failure.

[0005] A similar method is also proposed in an engineering machinery system with intelligent monitoring (202011084597.4), which proposes a positioning scanning device placed on the outer periphery of the construction area to monitor and scan the terrain model of the construction area, sense the distance from the excavating device, and indicate the moving direction of the excavating device. The excavating device is configured to excavate the construction area according to the terrain model scanned by the positioning scanning device. These methods all need to arrange multiple cameras, laser radars, sound wave ranging and other sensing devices in appropriate positions in the work area in advance, which has certain practical significance for the monitoring and management of large project work areas, but for the vast rural and town areas, there are many small work areas and temporary work areas, and it is difficult to arrange cameras for each work area. In addition, the advance arrangement of multiple sensing devices requires a certain amount of surveying and mapping and equipment management work, increasing the additional workload.

[0006] In summary, the traditional excavator working posture monitoring method is mainly based on the traditional image processing method. Some feature information is extracted through image grayscale processing, threshold segmentation and other methods, and then the feature information is input into a support vector machine for classification by means of machine learning. This kind of traditional method often has small amount of calculation, but the algorithm accuracy is low, and the applicability of the classifier is not strong, which can easily cause misidentification when the environmental factors change in actual operation. In addition, the traditional excavator working posture monitoring needs to arrange monitoring equipment in the work area or needs to arrange marker feature points on the excavator, resulting in high monitoring cost and monitoring failure. SUMMARY

[0007] In order to solve the above technical problems existing in the prior art, to realize the online monitoring of the development and utilization of natural resources by the administrative department on the excavator and other engineering equipment, and the real-time online unified management of the equipment by the owner of the excavator. The application provides a kind of monitoring method of excavator working pose based on vision. When testing the visual ranging accuracy, the application does not need joint calibration, does not need target matching, and can guarantee the accuracy of data, improves the accuracy of monitoring.

[0008] The technical scheme of the application for solving the above technical problems is: a kind of monitoring method of excavator working pose based on vision, comprising the following steps:

[0009] Step S1, a monocular camera is used to collect the video of the bucket, the arm and the digging surface in real time, and the video is sliced and divided into pictures at different time points, a trained model is used to frame the position of the bucket and the arm in the monitoring video, the state of the bucket is determined and the image coordinate information of the bucket and the arm target detection frame is obtained;

[0010] Step S2, the image in the arm target detection frame is segmented;

[0011] Step S3, the position and width of the arm in the image are calculated;

[0012] Step S4, the actual position of the bucket from the cab is estimated;

[0013] Step S5, the position of the excavator working point is calculated;

[0014] Further, the trained model in step S1 is obtained by the following steps:

[0015] A large number of picture target positions are collected by the camera in the cab of the excavator and are labeled, and the collected bucket and arm image data are divided into independent and non-repeated verification set and test set according to a certain proportion by random sampling;

[0016] The input image is processed, mainly including Mosaic data enhancement, adaptive anchor frame calculation and adaptive image scaling;

[0017] Training and feature extraction are carried out;

[0018] Target detection.

[0019] Further, the image segmentation in step S2 is carried out by color image back projection or Unet image segmentation method.

[0020] Further, the image segmentation step in step S2 using the color image back projection segmentation method is: image magnification of the detection range; preprocessing of image center region pixel extraction; template mean and standard deviation calculation; Gaussian probability density calculation; back projection image and segmentation.

[0021] Further, the image segmentation step in step S2 using the Unet image segmentation method is: preprocessing image extraction and magnification to obtain a preprocessed image; constructing a unet training set and a validation set in the preprocessed image; and segmenting the preprocessed image using a unet network weight file to directly obtain the width of the small arm in the image.

[0022] Further, the specific steps of estimating the actual position of the bucket relative to the cab in step S4 are:

[0023] Estimation of the distance between the small arm end and the cab:

[0024]

[0025] where W represents the actual width of the small arm, F is the focal length of the camera, w t is the width of the small arm in the image;

[0026] Estimation of the distance between the small arm end and the ground:

[0027] According to similar triangles, that is, where w t is the width of the small arm in the image, W is the actual width of the small arm, y t is the vertical y coordinate of the small arm end in the image, y m represents the midpoint coordinate in the vertical direction of the image, and Δy represents the height of the small arm end relative to the camera, that is, the height H of the small arm end is: H = h - Δy, where h is the height of the camera relative to the ground;

[0028] Correction of the distance and height of the small arm end:

[0029] The horizontal distance Dz of the small arm end is: Dz = D x cos(θ) x cos(Φ), where θ is the inclination of the excavator walking and the horizontal plane measured by the level, and D represents the distance from the camera to the small arm.

[0030] The actual height Hz of the small arm end is: Hz = H x Dz / D, where Hz represents the actual height of the small arm end, H represents the height of the small arm end in the image, D represents the distance from the camera to the small arm, and Dz represents the horizontal distance of the small arm end.

[0031] The beneficial effects of the present application: the present application adopts monocular camera to collect the video of the bucket, the arm and the digging surface in real time, frames the position of the bucket and the arm in the monitoring video, obtains the image coordinate information of the target detection frame, and realizes the global real-time dynamic monitoring. It has the advantages of low monitoring cost and does not affect the construction of the excavator. The owner of the excavator can realize real-time online unified management of the equipment, and real-time grasp of the position and working state of the equipment. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 is a flowchart of a visual-based excavator working pose monitoring method of the present application;

[0033] Figure 2 is a training effect evaluation diagram of the bucket and arm target detection model in the excavator in the present application;

[0034] Figure 3 is a bucket and arm target detection effect diagram in the working of the excavator in the present application;

[0035] Figure 4 is a small arm target segmentation effect diagram in the working of the excavator in the present application;

[0036] Figure 5 is a loading point and unloading point position distribution diagram of the bucket in the excavation of the foundation of the excavator at a certain place in the present application. DETAILED DESCRIPTION

[0037] The present application will be further described in detail below in combination with the drawings and specific embodiments.

[0038] As shown in Figure 1 , the present application provides a visual-based excavator working pose monitoring method, which comprises the following steps:

[0039] Step 1, video image bucket state and arm labeling and training.

[0040] (1) Image labeling and processing

[0041] The selected labeled pictures or image data are mainly from internet and on-site video collection data. The video stream data is extracted and converted into photos to form a picture set for training and labeling. Labelme software is used to label the data set, and the arm in the image set is framed and labeled one by one. The empty state and loaded state of the bucket are labeled respectively, wherein the empty state is that the bucket exposes the bucket tooth and there is no soil in the bucket, and the loaded state is that the bucket does not expose the bucket tooth and there is soil in the bucket. The labeling of the photos is completed. The labeled photos are divided into verification set and training set in the ratio of 1:9 according to random sampling.

[0042] (2) Image preprocessing

[0043] Before the model starts training, the image is uniformly processed, the picture size is unified to 640x640, Mosaic data enhancement processing is uniformly performed, clustering analysis is performed to obtain the classification of the small arm prior box, and the classification of the small arm prior box is obtained by clustering analysis of the length and width information of the small arm annotation box formed in the annotation process. The average length and width data of each type of box are calculated respectively, and the adaptive anchor box is obtained as [8, 11; 15, 32; 27, 41; 30, 61; 62, 53; 62, 124; 104, 108; 144, 206; 363, 322].

[0044] (3) Training and feature extraction

[0045] In order to obtain faster convergence speed and reduce training time cost, the embodiment is based on YOLOv5s pre-training weight under YOLOv5 framework for transfer learning. Among them, the hardware platform parameter central processing unit is Intel i7-7700k, the graphics computing card is Nvidia Ge Force GTX 1060Ti, the batch number is 16, the training is 120 rounds, and the total iteration is 92160 times.

[0046] The target detection of Yolov5 is mainly realized through the Head output layer, and the width, height and center point coordinates of the small arm and the bucket boundary box are obtained. The loss function GIOU_Loss or CIOU_Loss of the output layer is calculated, and finally the weight file obtained by training is obtained.

[0047] Step two, bucket state and small arm target detection in video image.

[0048] The weight file yolov5s.pt obtained by training on the computer is deployed on the Nvidia JETSON NANO development board 4GB core module for target detection. Specifically, the USB interface of the JETSON NANO development board 4GB core module is connected with the camera for video signal acquisition, the collected data is sliced into pictures and loaded into the yolov5 detection module, the detection module combines the weight file yolov5s.pt to predict the position of the small arm on the picture; when the bucket is loaded with earthwork, it is determined as the loaded state, and when the bucket is not loaded with earthwork and the bucket teeth are exposed, it is determined as the loaded state.

[0049] Step three, image segmentation of small arm.

[0050] (1) Extraction and enlargement of preprocessed image

[0051] According to the range of the target box, the small arm image detected in the detection range in step one is intercepted, the length and width information is read, and the intercepted image length and width are enlarged by 3 times synchronously, and the preprocessed image is obtained.

[0052] (2) Template pixel extraction

[0053] Extract the template pixel sample in the direction of the long and short sides of the center point of the preprocessed image, where the center point of the preprocessed image is the center point coordinate of the boundary box of the arm and the bucket obtained by target detection in step one.

[0054] (3) Calculation of template mean and standard deviation

[0055] Convert the preprocessed image to HSV format data, and calculate the corresponding mean and standard deviation for the H and S channels respectively, and remove the influence of the V channel (brightness) to reduce the influence of light brightness on segmentation.

[0056] (4) Calculation of Gaussian probability density

[0057] Calculate the product of P(H) and P(S) according to the Gaussian probability density formula for each pixel point of the preprocessed image.

[0058] (5) Back projection image and segmentation

[0059] After normalizing the H and S channels in the preprocessed image, the output result is the final back projection image based on Gaussian PDF, and the final color model object segmentation is obtained according to the mask to obtain the small arm boundary range in the image.

[0060] Step four, calculation of the position and width of the small arm in the image.

[0061] Read the width w, height h, and lower right corner coordinates (Xmax, Ymax) of the preprocessed image in step three; compare the total number of pixels in the boundary range in the segmented small arm image obtained in step three with the total number of pixels in the preprocessed image in step three to obtain k, which is the width w of the small arm in the image t = k x w.

[0062] Step five: estimation of the actual position of the bucket relative to the cab

[0063] (1) Estimation of the distance between the small arm end (bucket) and the cab

[0064] Since the bucket and the small arm are moving in a vertical plane relative to the cab, a monocular distance measurement algorithm can be used, i.e. the distance from the camera to the small arm where W represents the actual width of the small arm, F is the focal length of the camera, and w t is the width of the small arm in the image.

[0065] (2) Estimation of the distance between the small arm end (bucket) and the ground

[0066] According to similar triangles, that is, where w t is the width of the arm in the image, W is the actual width of the arm, y t is the vertical y coordinate of the end of the arm in the image, y m represents the midpoint coordinate in the image vertically, and Δy represents the height of the end of the arm relative to the camera, that is, the height H of the end of the arm is: H = h - Δy, where h is the height of the camera from the ground.

[0067] (3) Distance and height correction of the end of the arm (excavator bucket)

[0068] The distance and height correction of the end of the arm (excavator bucket) is mainly through the inclination θ of the level in the excavator and the angle Φ between the inclination of the level and the orientation of the excavator. The horizontal distance Dz of the end of the arm (excavator bucket) is: Dz = D x cos(θ) x cos(Φ), where θ is the inclination of the excavator measured by the level, and D represents the distance from the camera to the arm.

[0069] The actual height Hz of the end of the arm (excavator bucket) is: Hz = H x Dz / D, where Hz represents the actual height of the end of the arm (excavator bucket), H represents the height of the end of the arm (excavator bucket) in the image, D represents the distance from the camera to the arm, and Dz represents the horizontal distance of the end of the arm (excavator bucket).

[0070] Step six: calculation of the working point position of the excavator

[0071] The positioning information (X1, Y1, Z1) of the cab is obtained by combining the GPS / Beidou chip installed in the cab, and the azimuth angle information (azimuth angle is α) of the cab is obtained by the electronic compass. The state information of the actual working point of the bucket is calculated, that is, the actual coordinates of the bucket are X = X1 + D x sin(α), X = X1 + D x cos(α), and Z = Z1 + Hz x cos(α).

[0072] Through the identification of the empty state and the loaded state of the bucket, when the bucket changes from the empty state to the loaded state, it is determined as a loading point, and the coordinates, time point, and state (loading) of the bucket are recorded. When the bucket changes from the loaded state to the empty state, it is determined as an unloading point, and the coordinates, time point, and state (unloading) of the bucket are recorded. The recorded data is transmitted in real time to the cloud platform through 4G / 5G Internet of Things signals for analysis.

[0073] Example 2:

[0074] In step three, for image segmentation of the arm, a unet model is used for semantic segmentation, and the specific steps can be divided into the following steps:

[0075] (1) Extraction and enlargement of the preprocessed image

[0076] According to the range of the target frame, the small arm image detected in the range is detected, the length-width information thereof is read, and the length and the width of the intercepted image are amplified by 3 times respectively to obtain a pretreatment image.

[0077] (2) Constructing unet training set and verification set in the pretreatment image

[0078] According to the obtained pretreatment image set of 116 images, 100 images are selected as the training set, and the remaining 16 images are used as the verification set. Similarly, the Labelme software is used for labeling to form a Voc2007 format data set, which is sent into the vgg16-unet model for training, and the frozen learning mode is used to freeze the parameters of the first several generations of training, and the pre-training weight optimized for the unet network is used through the migration learning mode.

[0079] (3) Image segmentation

[0080] The pretreatment image is segmented according to the weight file of the unet network, the segmentation is actually a binary classification problem of the image pixels, a black and white binary image is output as the segmentation result, finally, all the segmented ROI region images and the original image are sent into an image fusion module for fusion to obtain the final segmentation result, and the width w of the small arm in the image is directly obtained. t .

[0081] It should be emphasized that the examples described in the present application are illustrative rather than limiting, and therefore the present application is not limited to the examples described in the specific embodiments, and any other embodiments derived by those skilled in the art according to the technical solutions of the present application, without departing from the purpose and scope of the present application, whether modified or replaced, also belong to the protection scope of the present application.

Claims

1. A vision-based method for monitoring the working posture of an excavator, characterized in that, Includes the following steps: Step S1: Use a monocular camera to capture real-time video of the bucket, boom, and excavation face, and slice the video into images at different time points. Use a trained model to frame the position of the bucket and boom in the monitoring video, determine the state of the bucket, and obtain the image coordinate information of the target detection boxes of the bucket and boom. Step S2: Segment the image within the forearm target detection box; Step S3: Calculate the position and width of the forearm in the image; Step S4: Estimate the actual position of the bucket from the cab; The specific steps are as follows: Estimated distance between the end of the forearm and the cab: D= ; Where W represents the actual width of the forearm, F is the camera focal length, and w t The width of the forearm in the image; Estimation of the distance between the end of the forearm and the ground: Calculate based on similar triangles, that is ,in W represents the width of the forearm in the image, and W represents the actual width of the forearm. Let y be the vertical y-coordinate of the forearm tip in the image. This represents the coordinates of the midpoint of the image in the vertical direction. The height H of the forearm end relative to the camera is expressed as: H = h - , where h is the height of the camera above the ground; Forearm distal distance and height correction: The horizontal distance Dz at the end of the forearm is: ,in The angle between the excavator's travel and the horizontal plane, as measured by the level, is where D represents the distance from the camera to the boom, and Φ represents the angle between the level's inclination and the excavator's orientation. Actual height of the forearm end for: ; Step S5, calculate the location of the excavator's working point; the specific steps are as follows: The positioning information (X1, Y1, Z1) of the cab is obtained by combining the GPS / Beidou chip installed in the cab and the azimuth information of the cab obtained by the electronic compass. The status information of the actual working point of the excavator bucket is calculated, where the azimuth is α. The actual coordinates of the bucket are: X = X1 + D × sin(α), Y = Y1 + D × cos(α), Z = Z1 + Hz × cos(α); Determining the bucket's status includes: identifying the bucket's empty and loaded states; when the bucket changes from empty to loaded, it is determined as a loading point, and the bucket's coordinates, time point, and loading status are recorded; when the bucket changes from loaded to empty, it is determined as an unloading point, and the bucket's coordinates, time point, and unloading status are recorded. When the bucket is loaded with soil, it is considered to be in a loading state; when the bucket is unloaded and the bucket teeth are exposed, it is considered to be in an unloading state.

2. The vision-based excavator working posture monitoring method according to claim 1, characterized in that, The training of the model in step S1 includes: A large number of images are collected from the excavator's cab using cameras, and the target locations are then labeled. The collected bucket and boom image data were divided into independent and non-repeating validation and test sets according to a certain proportion using random sampling. The input image is processed, including Mosaic data augmentation, adaptive anchor box calculation, and adaptive image scaling.

3. The vision-based excavator working posture monitoring method according to claim 1, characterized in that, In step S2, image segmentation is performed using color image back projection or the Unet image segmentation method.

4. The vision-based excavator working posture monitoring method according to claim 3, characterized in that, The steps for image segmentation using the color image back projection segmentation method are as follows: Image magnification within the detection range; pixel extraction from the center region of the preprocessed image; determination of template mean and standard deviation; calculation of Gaussian probability density; back-projection and segmentation of the image; The image magnification of the detection range is achieved by cropping the forearm image detected within the detection range of step S1 according to the range of the target box, reading its length and width information, and simultaneously magnifying the length and width of the cropped image by 3 times to obtain the preprocessed image. Pixel extraction of the center region of the preprocessed image involves simultaneously expanding outward by 1 / 10 of the length and width along the length and width of the center point of the preprocessed image to obtain template pixel samples. The template mean and standard deviation are obtained by converting the preprocessed image into HSV format data and calculating the corresponding mean and standard deviation only for the H and S channels. The Gaussian probability density is calculated by multiplying P(H) and P(S) for each pixel of the preprocessed image according to the Gaussian probability density formula. Back-projection image segmentation is the output result after normalizing the H and S channels of the preprocessed image, which is the final back-projection image based on Gaussian PDF. The final color model object segmentation is obtained based on the Mask, and the forearm boundary range in the image is obtained.

5. The vision-based excavator working posture monitoring method according to claim 3, characterized in that, The Unet image segmentation method performs the following image segmentation steps: The preprocessed image is extracted and enlarged to obtain the preprocessed image; a UET training set and a validation set are constructed in the preprocessed image; the preprocessed image is segmented according to the weight file of the UET network to directly obtain the width of the forearm in the image.

Citation Information

Patent Citations

  • Attitude measurement method of non-contact excavator working device based on vision measurement

    CN103900497B

  • Engineering machinery system with intelligent monitoring function

    CN112195989A

  • Monocular distance measurement method in intelligent driving environment

    CN111046843A

  • KR20210068174A