Excavator bucket pose estimation method and equipment based on time sequence neural network
By integrating visual data with inertial measurement unit data and using the timing neural network model, the position estimation problem when the bucket is blocked is solved, and high-precision and low-cost bucket posture detection is achieved. It is suitable for different models of excavators, with strong occlusion resistance and high robustness, and is suitable for embedded equipment storage.
Patent Information
- Application Number
- CN202510513828.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-11
AI Technical Summary
The existing excavator position estimation method is difficult to accurately obtain the position information of the bucket when it is blocked, which affects the stability of the automatic control system, and the sensor is vulnerable to damage or is affected by environmental factors.
The visual data is fused with inertial measurement unit data and combined with the timing neural network model, and the improved object detection network and timing neural network model are used to predict the position information of the bucket, avoid direct dependence on the bucket IMU sensor, and use the fixed geometric relationship between the bucket and the forearm for high-precision estimation.
When the bucket is blocked, it can still estimate the position information with high accuracy, which reduces the risk of sensor damage and reduces detection costs. It is suitable for different models of excavators, with strong occlusion resistance and high robustness. The model is lightweight and suitable for embedded equipment storage, meeting real-time detection requirements.
Smart Images

Figure CN120296677A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent control of construction machinery, and in particular to a method and device for estimating the pose of an excavator bucket based on a temporal neural network. Background Art
[0002] As an important earthwork machinery, excavators are widely used in fields such as mines, infrastructure construction, and engineering construction, and usually need to work in harsh environments, such as dangerous scenes like open-pit mines and chemical mines. However, in these complex working conditions, the operation ability and work safety of the driver become important factors affecting construction efficiency. Therefore, in recent years, the research on automated excavators has gradually received attention and is considered an effective means to improve work safety and production efficiency.
[0003] In the research of automated excavators, accurately controlling the trajectory of the working device of the excavator is the basis for achieving efficient operation. The working device of the excavator consists of multiple independently moving components, including the boom, arm, and bucket, and its spatial pose changes dynamically over time. To achieve precise automatic control, it is necessary to accurately obtain the pose information of these components, especially the pose of the bucket. However, during operation, the bucket comes into contact with media such as soil and rock, resulting in frequent occlusion, making pose measurement extremely difficult.
[0004] Currently, the pose estimation methods for the working device of excavators are mainly divided into two categories: contact sensor-based methods and vision-based methods. Contact sensor-based methods mainly rely on sensors such as encoders and inertial measurement units (IMUs) installed at the joints of the excavator to obtain the motion states of each joint. However, in the construction environment, these sensors are easily affected by factors such as dust and vibration, resulting in the accumulation of measurement errors. Moreover, since the bucket is in direct contact with the media, the sensors installed on the bucket are easily damaged, leading to high maintenance costs. In addition, the IMU sensor has the problem of cumulative drift, and long-term use will affect the pose estimation accuracy.
[0005] In contrast, computer vision-based methods do not need to rely on additional physical sensors. They can capture the motion state of the excavator through the camera and calculate its posture information through image processing technology. This type of method can be divided into fixed external camera solutions and airborne camera solutions according to the different camera placement methods. Chinese patent application CN117058619A discloses a deep learning-based excavator posture detection method, which uses an improved YOLOV5 algorithm to detect the posture of the entire body of the excavator, but the method in this patent does not care about the occlusion in actual work. The fixed external camera solution usually requires setting up multiple cameras in the working environment to capture the working state of the excavator from different angles, but this method is greatly affected by factors such as ambient lighting, occlusion, and viewing angle range, and it is difficult to operate stably for a long time. Although the airborne camera solution can adjust the viewing angle as the excavator moves, it still has problems such as the influence of fuselage vibration, camera calibration error, and high computational complexity. In addition, whether it is a fixed camera or an airborne camera, when the bucket is blocked, it is difficult to guarantee the accuracy of posture estimation by relying solely on visual data, which affects the stability of the automatic control system.
[0006] At present, the existing excavator posture estimation methods still face great challenges when the bucket is obscured, and lack a stable and reliable estimation solution. Therefore, how to accurately obtain its posture information when the bucket is obscured and use it as feedback data for automatic control is still a key problem that needs to be solved urgently. Summary of the invention
[0007] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide an excavator bucket posture estimation method and equipment based on a time series neural network. The method adopts the fusion of visual data and inertial measurement unit data and combines it with the prediction method of the time series neural network model. It can still estimate the bucket posture information with high precision without the need for a bucket IMU (inertial measurement unit) sensor.
[0008] The purpose of the present invention can be achieved by the following technical solutions:
[0009] A method for estimating the position and posture of an excavator bucket based on a temporal neural network, the method comprising:
[0010] Obtain visual data of the excavator under test and inertial measurement unit data of the arm and boom;
[0011] Using the visual data, based on the improved target detection network model, the position information of the connection between the bucket and the forearm, the forearm and the boom of the excavator to be tested is obtained, and the preliminary bucket angle in the two-dimensional plane is calculated;
[0012] fusing the preliminary bucket angle in the two-dimensional plane with the inertial measurement unit data to obtain an input vector;
[0013] Using the input vector, based on the trained time series neural network model, predict the actual bucket angle;
[0014] Convert the actual bucket angle into the bucket pose.
[0015] Further, the inertial measurement unit data includes the angular velocity, acceleration, and attitude angle of the forearm and the boom.
[0016] Further, the position information at the connection between the bucket and the forearm and at the connection between the forearm and the boom includes the key point coordinates at the connection between the bucket and the forearm and the connection point coordinates between the boom and the forearm.
[0017] Even further, the calculation process of the preliminary bucket angle in the two-dimensional plane includes:
[0018] Obtain the self-lengths, link lengths, joint spacings, and position information of the bucket, forearm, and boom of the excavator to be measured; according to the position information, obtain the angle between the connection between the bucket and the forearm and the horizontal plane of the machine body and the angle between the forearm and the horizontal plane of the machine body; according to the fixed geometric relationship between the key points at the connection between the bucket and the forearm, calculate the angle between the bucket tip and the horizontal plane, that is, the angle between the bucket surface and the horizontal plane of the machine body, and output it as the preliminary bucket angle in the two-dimensional plane.
[0019] Even further, the calculation expression for the angle between the bucket tip and the horizontal plane is:
[0020] ∠QNM = π - ∠2 - ∠FNO + ∠1,
[0021]
[0022]
[0023]
[0024] ∠FQV = 2π - ∠NQF - ∠MQN - ∠MQK - ∠KQV,
[0025] ∠3 = ∠2 + ∠FQV - π,
[0026] where M, N, Q, K are 4 key points at the connection between the bucket and the forearm, V is the bucket tip point, F is the connection point between the boom and the forearm, ∠1 is the angle between the line segment MN and the horizontal plane of the machine body, that is, the angle between the connection between the bucket and the forearm and the horizontal plane of the machine body, ∠2 is the angle between the line segment FN and the horizontal plane of the machine body, that is, the angle between the forearm and the horizontal plane of the machine body, and ∠3 is the angle between the line segment QV and the horizontal plane of the machine body, that is, the angle between the bucket tip and the horizontal plane.
[0027] Further, the training process of the time series neural network model includes:
[0028] Obtain the visual data of the excavator for training and the inertial measurement unit data of the bucket, forearm, and boom.
[0029] Utilize the visual data to obtain the position information of the connection between the bucket and the forearm and the positions of the forearm and the boom of the excavator for training based on an improved object detection network model, and calculate the preliminary bucket angle in the two-dimensional plane.
[0030] Fuse the preliminary bucket angle in the two-dimensional plane with the inertial measurement unit data of the forearm and the boom, and construct a mapping with the inertial measurement unit data of the bucket to obtain a time-series training dataset.
[0031] Construct a time-series neural network model and train the time-series neural network model using the time-series training dataset.
[0032] Furthermore, the object detection network model includes a head part and a C2F module. The improvement of the object detection network model includes:
[0033] Add at least 2 shared convolutional layers in the head part, and replace the batch normalization in the head part with group normalization.
[0034] In the C2F module, use a RepConv structure and remove the activation function in the convolutional layer. The parameters between the convolutional kernels of the C2F module are shared.
[0035] After training the object detection network model, perform pruning on the model.
[0036] Furthermore, the time-series neural network model includes a fully connected neural network and a recurrent neural network.
[0037] Furthermore, the bucket pose includes the attitude angle of the bucket tip and the position coordinates of the bucket tip.
[0038] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-mentioned method for estimating the bucket pose of an excavator based on a time-series neural network.
[0039] Compared with the prior art, the beneficial effects of the present invention include:
[0040] 1. The present invention adopts a prediction method that fuses visual data with inertial measurement unit data and combines a time-series neural network model. Without the need for a bucket IMU (inertial measurement unit) sensor, it can still accurately estimate the pose information of the bucket. In actual use, only sensors need to be installed on the forearm and the boom, which greatly reduces the risk of sensor damage and the detection cost.
[0041] 2. In the detection of the bucket pose, the present invention mainly uses the fixed geometric relationship at the connection between the bucket and the forearm to calculate the initial bucket angle in a two-dimensional plane. Even if the bucket is occluded, the pose information of the bucket can still be estimated with high precision by deriving the angle through the unoccluded connection between the bucket and the forearm. The present invention has strong anti-occlusion ability for bucket pose detection.
[0042] 3. The present invention is applicable to excavators of different models. By adjusting the set geometric parameters and model training, it can quickly adapt to new models and has strong generalization ability.
[0043] 4. Considering actual deployment and application, the present invention has made lightweight improvements to the target detection network model. It introduces shared convolutions in the head part, enabling feature maps of different scales to share information in the head part, improving the consistency and coherence of the model in multi-scale feature processing. It uses group normalization to replace batch normalization, enhancing the accuracy and robustness of the model in key point detection and pose estimation tasks. In the C2F module, the RepConv structure is used to enhance feature diversity, ensuring the accuracy of excavator key point detection. In some convolutional layers, the activation function is removed to reduce the computational complexity and improve the inference speed. Shared convolutions are introduced to enhance the robustness and generalization ability of the model. The convolutional kernel design with parameter sharing is adopted to ensure that the network can maintain consistency during the multi-scale feature extraction process and achieve higher accuracy in complex pose estimation tasks. After training the target detection network model, the LAMP pruning method is used to prune the model. Without sacrificing accuracy, the number of model parameters is reduced from 3.08M to 0.25M, a reduction of 91.88%, and the FLOPs are reduced from 8.3GFLOP to 2.5GFLOP, a reduction of 69.88%. The model volume is greatly reduced, making it suitable for storage in embedded devices for excavator bucket pose estimation. At the same time, the model calculation efficiency is improved, meeting the low-latency requirements for real-time detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flow chart of the method of the present invention;
[0045] Figure 2 It is the key node of the bucket and the bucket connecting rod of the present invention;
[0046] Figure 3 It is a schematic diagram of the physical object of the geometric relationship between the bucket and the bucket connecting rod of the present invention;
[0047] Figure 4 It is a schematic diagram of the geometric relationship line between the bucket and the bucket connecting rod of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] Embodiment 1
[0050] This embodiment aims to disclose a method for estimating the pose of an excavator bucket based on a temporal neural network. The flow of this method is as Figure 1 shown, and the specific steps of the method include:
[0051] Step S1, obtain the visual data of the excavator to be measured and the inertial measurement unit data of the forearm and the boom;
[0052] Step S2, use the visual data to obtain the position information of the connection between the bucket and the forearm, and the forearm and the boom based on an improved object detection network model, and calculate the preliminary bucket angle in the two-dimensional plane;
[0053] Step S3, fuse the preliminary bucket angle in the two-dimensional plane with the inertial measurement unit data to obtain an input vector;
[0054] Step S4, use the input vector to predict the actual bucket angle based on a trained temporal neural network model;
[0055] Step S5, convert the actual bucket angle into the bucket pose.
[0056] The bucket pose includes the attitude angle of the bucket tip and the position coordinates of the bucket tip.
[0057] In step S1, the inertial measurement unit data is obtained from an IMU (inertial measurement unit) sensor installed on the excavator to be measured. The inertial measurement unit data includes attitude data such as the angular velocity, acceleration, and attitude angle of the forearm and the boom. The visual data of the excavator to be measured is obtained through a visual sensor installed on the side of the excavator.
[0058] In step S2, the position information of the connection between the bucket and the forearm, and the forearm and the boom obtained based on the improved object detection network model includes the key point coordinates of the connection between the bucket and the forearm and the connection point coordinates of the boom and the forearm. What the object detection network model can also detect are the bucket surface and forearm surface of the excavator and the bucket tip coordinate points, and these data are used for the training of the temporal neural network model.
[0059] In the research of the existing technology, the key points detected by vision algorithms are usually the joint connections of the excavator bucket, forearm, and boom, as well as the key points on the rest of the body. However, due to the occlusion of media such as soil, the feasibility of these methods in the actual working environment of the excavator is relatively low.
[0060] In this embodiment, we observed that during the operation of the excavator, the bucket will be occluded by the excavation materials, but the quadrilateral at the connection between the bucket and the forearm ( Figure 3 the MNQK quadrilateral in Figure 3 ) is generally not occluded, and there is a definite geometric relationship between the angle at this connection and the bucket tooth tip in the two-dimensional plane, as Figure 4 shown. This geometric relationship must exist in the two-dimensional plane, where the lengths of each connection are simply measurable, and the angles can also be obtained through angle calculations after obtaining the coordinates of each point. The geometric relationship diagram after its mathematical line simplification is as
[0061] shown. Therefore, after detecting the coordinates and plane angles of points M, N, Q, and K through the improved object detection network model and combining with the physical parameters of the excavator, the initial bucket angle in the two-dimensional plane can be obtained through calculation.
[0062] In the improvement of the object detection network model, shared convolutions are introduced in the head part to reduce the number of parameters, thereby realizing the lightweight of the model. Two 3x3 shared convolution layers are added to the head part, achieving shared convolutions between feature maps of three different scales: large, medium, and small. This design enables different-scale feature maps to share information in the detection head, thereby improving the consistency and coherence of the model in multi-scale feature processing. Whether it is a large target, a medium target, or a small target, the same image features (feature) are learned for classification. To address the possible differences in processing target scales by different detection heads, in this embodiment, the Scale layer is used to scale the features;
[0063] During the lightweight process, the feature extraction ability of the model may weaken. If no remedial measures are taken, the performance may decline significantly. Therefore, in order to maintain the detection accuracy after shared convolutions, in this embodiment, group normalization (GN) is used to replace batch normalization (BN) in the head part. GN can provide more stable performance between small-batch training or different feature map scales, thereby improving the accuracy and robustness of the model in key point detection and pose estimation tasks;
[0064] Through the improvement of the head part, the fusion of multi-scale features can be effectively enhanced, and at the same time, the stability under different feature scales can be ensured, making the model more suitable for the detection of key points on the excavator bucket surface;
[0065] In addition, in the C2F module of the object detection network model, the RepConv (repeated convolution) structure, that is, a multi-branch convolution path, is adopted. In the C2F module, one branch uses a 3x3 convolution, and the other branch uses a 1x1 convolution. This design can not only effectively capture local detail features but also fuse global features in a larger range, thus ensuring the accuracy of excavator key point detection;
[0066] To reduce the computational complexity and improve the inference speed, the activation function is removed from some convolutional layers in the C2F module, reducing the non-linear interference, thereby effectively reducing the computational overhead and making the model more efficient in task execution, especially suitable for tasks that require real-time response such as excavator key point detection;
[0067] Introducing shared convolution in the C2F module can obtain consistent feature representations at different scales, enhancing the robustness of the model. A convolutional kernel design with parameter sharing is also adopted, enabling feature maps at different scales to share the learned feature representations;
[0068] After training the object detection network model, the LAMP pruning method is used to prune the model.
[0069] In the improved model, while maintaining the accuracy, the number of model parameters is reduced from 3.08M to 0.25M, a reduction of 91.88%. The model volume is significantly reduced, suitable for storage on embedded devices. The model FLOPs are reduced from 8.3 GFLOP to 2.5 GFLOP, a reduction of 69.88%, and the computational efficiency is improved, meeting the low-latency requirements for real-time detection.
[0070] After detecting the coordinates and plane angles of points M, N, Q, and K, the preliminary bucket angle in the two-dimensional plane is calculated.
[0071] The calculation process of the preliminary bucket angle in the two-dimensional plane includes:
[0072] Obtain the self-lengths, link lengths, joint spacings, and position information of the bucket, forearm, and boom of the excavator to be measured; according to the position information, obtain the angles between the connection of the bucket and the forearm and the horizontal plane of the body and the angle between the forearm and the horizontal plane of the body; according to the fixed geometric relationship between the key points at the connection of the bucket and the forearm, calculate the angle between the bucket tip and the horizontal plane, that is, the angle between the bucket surface and the horizontal plane of the body, and output it as the preliminary bucket angle in the two-dimensional plane.
[0073] In this embodiment, the key nodes of the bucket and the bucket links (boom and forearm) are as Figure 2 shown. Point F represents the connection point of the boom and the forearm. Points M, N, Q, and K are the four vertices of the quadrilateral at the connection of the bucket and the forearm, and point V is the bucket tip. Figure 3Among them, the ray FY is drawn from point F and is parallel to the XY plane of the body coordinate system; similarly, the ray NM is drawn from point N and is parallel to the body's horizontal plane, and the ray QZ is drawn from point Q and is parallel to the body's horizontal plane. ∠1 is the angle between the line segment MN and the body's horizontal plane, which is a fixed value in a single detection; ∠2 is the angle between the line segment FN and the body's horizontal plane, which is also a fixed value in a single detection; ∠3 is the angle between the bucket surface QV and the body's horizontal plane, that is, the bucket angle we need to solve.
[0074] The calculation expression for the angle between the bucket tooth tip and the horizontal plane is:
[0075] ∠QNM = π - ∠2 - ∠FNO + ∠1,
[0076]
[0077]
[0078]
[0079] ∠FQV = 2π - ∠NQF - ∠MQN - ∠MQK - ∠KQV,
[0080] ∠3 = ∠2 + ∠FQV - π,
[0081] Among them, M, N, Q, K are 4 key points at the connection between the bucket and the forearm, V is the bucket tooth tip point, F is the connection point between the boom and the forearm, ∠1 is the angle between the line segment MN and the body's horizontal plane, that is, the angle between the connection between the bucket and the forearm and the body's horizontal plane, ∠2 is the angle between the line segment FN and the body's horizontal plane, that is, the angle between the forearm and the body's horizontal plane, and ∠3 is the angle between the line segment QV and the body's horizontal plane, that is, the angle between the bucket tooth tip and the horizontal plane.
[0082] In S2, the preliminary bucket angle in the two-dimensional plane is obtained by analyzing and calculating the visual data. However, if the angle of the bucket tooth tip is directly derived from the detection results of the visual algorithm, the error will be relatively large. This is because the geometric relationship between the components of the excavator only holds completely when reduced to a two-dimensional plane. In actual experiments, since the camera is installed on the excavator body and cannot be completely parallel to the two-dimensional plane, a non-linear error is introduced, and the calculation of the plane cannot be directly applied to the three-dimensional reality.
[0083] To solve this problem, this application introduces a neural network method to eliminate the non-linear error. During the actual operation of an excavator, operations such as digging, lifting, and dumping all involve short-term time-dependence. Considering the actual deployment requirements, the network structure needs to be simple and efficient. Therefore, in this embodiment, a fully connected neural network (FCN) is used to solve the non-linear error problem, and a recurrent neural network (RNN) is combined to process sequence data with short-term dependence. To enhance the efficient transmission of information in the network, this embodiment further introduces a residual connection to alleviate the problem of gradient disappearance.
[0084] In this embodiment, the time-series neural network model for predicting the actual bucket angle includes a fully connected neural network and a recurrent neural network, and its training process includes:
[0085] Obtain the visual data of the training excavator and the inertial measurement unit data of the bucket, forearm, and boom;
[0086] Using the visual data, based on the improved object detection network model, obtain the position information of the connection between the bucket and the forearm, and the forearm and the boom of the training excavator, and calculate the preliminary bucket angle in the two-dimensional plane;
[0087] Fuse the preliminary bucket angle in the two-dimensional plane with the inertial measurement unit data of the forearm and the boom, and construct a mapping with the inertial measurement unit data of the bucket to obtain a time-series training dataset;
[0088] Construct a time-series neural network model, and use the time-series training dataset to train the time-series neural network model.
[0089] During the actual process of using the time-series neural network model for prediction in step S4, the input vector first extracts spatial features through the FCN layer, and the activation function introduces non-linear transformation; the feature sequence of the input vector is then input into the RNN layer to update the hidden state to capture time-dependence; finally, the residual connection adds the output of the FCN layer and the RNN output to obtain the final predicted value of the bucket angle.
[0090] Embodiment 2
[0091] This embodiment aims to give a practical application example of the method based on the excavator bucket pose estimation method using a time-series neural network in the above Embodiment 1, and conduct tests on the laboratory excavator prototype to prove that this method is feasible.
[0092] The excavator experimental prototype used in this embodiment has a compact structural size. Among them, the boom length is 407 mm, the forearm length is 198 mm, and the bucket length is 158 mm. The horizontal offset from the base to the boom hinge point is 47.6 mm, and the fixed height offset at the bucket end is 184 mm. The above structural parameters form the basis for the bucket geometric modeling and pose estimation.
[0093] In this embodiment, the non-working state (stationary) and three common operations of an excavator during work, namely fixed-point excavation (without slewing condition), slewing excavation (with slewing condition), and leveling, are considered to evaluate and discuss the accuracy and robustness of the proposed method in estimating the attitude of the excavator bucket. The combination of these three operations covers the common work processes of an excavator, thus being able to comprehensively reflect the practicality and performance of this solution. In all three experiments, the IMU output value of the bucket position is used as the ground truth, and the result output by the temporal neural network model is used as the comparison value to prove the feasibility of the solution.
[0094] Under the stationary condition, the angle changes of the model output and the ground truth within 0 - 100 seconds are basically coincident, and the overall error range fluctuates between approximately -0.5° and 0.4°, indicating that their following situation is good and the difference is small. Under the non-slewing excavation condition, the angle change trends of the model output and the ground truth during the entire operation cycle are highly consistent, and the values are basically coincident. The error mainly fluctuates within ±5°, indicating that the system has good following ability and robustness and can accurately reflect the attitude change of the bucket under this condition. Under the excavation condition with slewing, the model output and the ground truth also maintain a good following trend, and the angle curves are basically synchronized during the two excavation operations. Most of the errors are controlled within ±5°. From the statistical results, the average value of the model output angle is approximately -135.25°, the average value of the ground truth is -131.88°, and the overall error is approximately -3.38°, indicating that the model prediction is slightly lower than the true value, but the overall difference is small. Under the leveling condition, the angle changes of the model output and the ground truth are basically coincident and the fluctuations are synchronized, indicating that the model also has good dynamic tracking ability in small-amplitude and continuous operations. From the data, the average value of the model angle is -98.48°, while the average value of the ground truth is -102.00°, indicating that the model output is slightly higher than the true value, with an error of approximately 3.52°, and the overall shows good consistency.
[0095] Embodiment 3
[0096] Based on Embodiment 1, this embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the excavator bucket pose estimation method based on the temporal neural network as described above.
[0097] At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned method for estimating the pose of the excavator bucket based on the temporal neural network. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.
[0098] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0099] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0100] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An excavator bucket pose estimation method based on a temporal neural network, characterized in that, The method includes: Obtaining the visual data of the excavator to be measured and the inertial measurement unit data of the forearm and the boom; Using the visual data, based on an improved object detection network model, obtaining the position information of the connection between the bucket and the forearm and the forearm and the boom of the excavator to be measured, and calculating the preliminary bucket angle in the two-dimensional plane; Fusing the preliminary bucket angle in the two-dimensional plane with the inertial measurement unit data to obtain an input vector; Using the input vector, based on a trained time series neural network model, predicting the actual bucket angle; Converting the actual bucket angle into a bucket pose.
2. The method for estimating the pose of an excavator bucket based on a temporal neural network according to claim 1, wherein The inertial measurement unit data includes the angular velocity, acceleration, and attitude angle of the forearm and the boom.
3. A method for estimating the pose of an excavator bucket based on a temporal neural network according to claim 1, wherein The position information of the connection between the bucket and the forearm and the forearm and the boom includes the key point coordinates of the connection between the bucket and the forearm and the connection point coordinates of the boom and the forearm.
4. The method for estimating the pose of the excavator bucket based on the temporal neural network according to claim 3, characterized in that, The calculation process of the preliminary bucket angle in the two-dimensional plane includes: Obtaining the self-length, connecting rod length, joint spacing, and position information of the bucket, forearm, and boom of the excavator to be measured; according to the position information, obtaining the angle between the connection between the bucket and the forearm and the horizontal plane of the machine body and the angle between the forearm and the horizontal plane of the machine body; according to the fixed geometric relationship between the key points at the connection between the bucket and the forearm, calculating the angle between the bucket tip and the horizontal plane, that is, the angle between the bucket surface and the horizontal plane of the machine body, and outputting it as the preliminary bucket angle in the two-dimensional plane.
5. A method for estimating the pose of an excavator bucket based on a temporal neural network according to claim 4, characterized in that, The calculation expression of the angle between the bucket tip and the horizontal plane is: ∠FQV = 2π - ∠NQF - ∠MQN - ∠MQK - ∠KQV, ∠3 = ∠2 + ∠FQV - π, where M, N, Q, K are 4 key points at the connection between the bucket and the forearm, V is the bucket tip point, F is the connection point between the boom and the forearm, ∠1 is the angle between the line segment MN and the horizontal plane of the machine body, that is, the angle between the connection between the bucket and the forearm and the horizontal plane of the machine body, ∠2 is the angle between the line segment FN and the horizontal plane of the machine body, that is, the angle between the forearm and the horizontal plane of the machine body, and ∠3 is the angle between the line segment QV and the horizontal plane of the machine body, that is, the angle between the bucket tip and the horizontal plane.
6. The method for estimating the pose of an excavator bucket based on a temporal neural network according to claim 1, characterized in that The training process of the time series neural network model includes: Obtaining the visual data of the training excavator and the inertial measurement unit data of the bucket, forearm, and boom; Using the visual data, based on an improved object detection network model, obtaining the position information of the connection between the bucket and the forearm and the forearm and the boom of the training excavator, and calculating the preliminary bucket angle in the two-dimensional plane; Fusing the preliminary bucket angle in the two-dimensional plane with the inertial measurement unit data of the forearm and the boom, and constructing a mapping with the inertial measurement unit data of the bucket to obtain a time series training data set; Constructing a time series neural network model and training the time series neural network model using the time series training data set.
7. A method for estimating the pose of an excavator bucket based on a temporal neural network according to claim 6, characterized in that, The object detection network model includes a head part and a C2F module, and the improvement of the object detection network model includes: Adding at least 2 shared convolutional layers in the head part and replacing the batch normalization in the head part with group normalization; In the C2F module, a RepConv structure is used and the activation function in the convolutional layer is removed, and parameter sharing is performed among the convolutional kernels of the C2F module; After the target detection network model is trained, pruning processing is performed on the model.
8. A method for estimating the pose of an excavator bucket based on a temporal neural network according to claim 1, characterized in that, The temporal neural network model includes a fully connected neural network and a recurrent neural network.
9. The method for estimating the pose of an excavator bucket based on a temporal neural network according to claim 1, wherein The bucket pose includes the attitude angle of the bucket tooth tip and the position coordinates of the bucket tooth tip.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the excavator bucket pose estimation method based on the temporal neural network according to any one of claims 1-9 are implemented.
Citation Information
Patent Citations
Excavator posture detection method based on deep learning
CN117058619A