Unmanned clamping vehicle reinforcement learning reward function design method considering error accumulation
By designing a reinforced learning reward function for unmanned clamped car that considers error accumulation, the error accumulation problem in the existing technology is solved, and the operating performance and reliability of unmanned clamped car is significantly improved.
Patent Information
- Application Number
- CN202510103090.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
AI Technical Summary
The reinforcement learning reward function of the existing unmanned clamped car has insufficient design architecture, and the special attributes and requirements of the cotton bale handling operation are not fully considered, resulting in the accumulation of errors and affecting the quality and stability of the operation.
Design a reinforcement learning reward function for unmanned clamped car that considers error accumulation. By identifying the links that may occur during the operation, collecting and analyzing relevant data, building a multi-source error propagation model, integrating multi-source error information, accurately assessing work errors, and designing basic reward items, error cumulative reward items and long-term operation quality reward items to form a comprehensive reward function, guiding clamped car adjustment control strategy to reduce error accumulation.
Significantly improve the overall performance of unmanned clamped vehicles in cotton bag handling operations, improve the accuracy and stability of operations, enhance the reliability and adaptability of the system, and can operate accurately in different complex environments and effectively control error accumulation.
Smart Images

Figure CN120029057A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned gripping vehicle control, and in particular to a method for designing a reinforcement learning reward function for an unmanned gripping vehicle taking error accumulation into consideration. Background Art
[0002] In the modern logistics industry system, the logistics links related to the cotton textile industry occupy an important position, among which the handling of cotton bales is particularly critical. The unmanned clamping vehicle, with its automated operation characteristics, has shown significant advantages in reducing labor intensity and costs, and can maintain a certain degree of operating efficiency and accuracy, and has become one of the core equipment for cotton bale handling operations.
[0003] As the focus and hotspot of the logistics industry, many related studies continue to improve the intelligent and unmanned level of unmanned forklifts. The unmanned forklift stacking fault location system and method based on machine learning (application number: CN202410788630.3) monitors and analyzes various potential faults of unmanned forklifts in real time during use. The multi-agent unmanned stacking forklift and control method based on event trigger mechanism (application number: CN202410709055.3) can discharge goods in an orderly manner, prevent stacking when stacking goods, and cause damage to goods, which greatly improves the working efficiency of unmanned forklifts. The unmanned vehicle path planning method based on deep reinforcement learning and A-star algorithm (application number: CN202210357348.0) uses deep neural networks and integrates data enhancement and curriculum learning technology to train unmanned vehicle agents in a simulation environment, and is committed to improving the path planning ability of unmanned vehicles.
[0004] However, the current reinforcement learning reward function used in unmanned clamping vehicles has obvious deficiencies in its design architecture and fails to fully consider the special properties and requirements of cotton bale handling operations. In the specific operation scenario of clamping cotton bales, it is far from enough to only focus on whether the cotton bales are successfully clamped and placed. The potential impact of various errors gradually accumulating during the operation on the quality of the cotton bales and the stability of handling should not be ignored. The simple success-failure binary evaluation model used by traditional reward functions is too crude to cope with such complex and changing operational requirements.
[0005] In the actual operation process of the unmanned clamping vehicle, the error accumulation phenomenon presents multi-dimensional manifestations. In the driving process, the occurrence of positioning errors and speed control errors may cause the clamping vehicle to deviate from the pre-planned optimal driving path. For example, due to slight deviations in the positioning system, the clamping vehicle may deviate from the correct route when driving to the cotton bale storage area, resulting in inaccurate starting position during the clamping operation; unstable speed control, such as fluctuations during acceleration or deceleration, may make the clamping vehicle unable to accurately align when approaching the cotton bale, affecting the accuracy and stability of subsequent clamping actions. During the execution of the clamping action, slight deviations in clamping force and angle may seem insignificant, but after repeated operations, they may cause serious damage to the cotton bale. Excessive clamping force may destroy the fiber structure inside the cotton bale and affect the quality of the cotton bale; deviations in the clamping angle may cause the cotton bale to be unstable and there is a risk of slipping during transportation. In addition, the errors between different operating links do not exist in isolation, but are interrelated and influence each other. Driving path deviation may lead to an unsatisfactory clamping position, thereby increasing the probability of errors in the clamping action; errors in the clamping action may affect the stability of the subsequent handling process, further exacerbating the decline in overall operation quality.
[0006] At present, most of the research on the operation error of unmanned clamping vehicles is limited to the correction and optimization of errors in a single link, and lacks effective solutions to explore the error accumulation effect from a macro and overall operation process perspective. In particular, in the design of the reward function of reinforcement learning, the comprehensive consideration of the error accumulation factor has not been fully integrated, which has largely restricted the control algorithm of the unmanned clamping vehicle to flexibly respond to and efficiently handle complex cotton bale handling operations, and it is difficult to meet the growing demand for refined operations in the logistics link of the cotton textile industry. Summary of the invention
[0007] The purpose of the present invention is to address the technical defects of the traditional reward function in the operation process of the unmanned clamping vehicle in the prior art, and to provide a reinforcement learning reward function design method for the error accumulation in the operation process of the unmanned clamping vehicle clamping cotton bales, and to optimize the design specifically for the error accumulation in the operation process of the unmanned clamping vehicle clamping cotton bales. The reward function can accurately evaluate the pros and cons of each operation step according to the real-time operation performance of the clamping vehicle, and provide timely feedback to the control algorithm, guiding the clamping vehicle to continuously adjust the control strategy, thereby gradually reducing the error accumulation in the operation process, and significantly improving the overall performance of the unmanned clamping vehicle in the cotton bale handling operation.
[0008] The technical solution adopted to achieve the purpose of the present invention is:
[0009] A method for designing a reinforcement learning reward function for an unmanned gripping vehicle considering error accumulation includes the following steps:
[0010] Step 1: Identify the links that may cause errors in the operation process of unmanned forklifts, and provide direction for data collection and the construction of multi-source error propagation models;
[0011] Step 2: Collect and organize data related to unmanned forklift operations i ;
[0012] Step 3: Analyze and collect data D i Accuracy, consistency and reliability, identify potential error sources, and determine the contribution of different error sources to data errors x,j ;
[0013] Step 4: Construct a multi-source error propagation model for unmanned forklift operation. The system error includes the errors generated by p error sources, which are ∈ source1 ,∈ source2 ,…,∈ sourcep ;
[0014] Step 5: Based on the multi-source error propagation model and error propagation characteristics of unmanned forklift operation, data fusion technology is used to fuse multi-source error information, and the system error ∈ total , comprehensive operation error estimate ∈ fused Accurately evaluate the operation error and calculate the comprehensive operation quality index Q and the whole operation error ∈ total-process ;
[0015] Step 6: Design a reward function R for the multi-source error propagation model, R = R b +R e +R l , R b As a basic reward item, R e is the error cumulative reward and R l is the long-term work quality reward item, R b Standardize work behavior, R e Constraint error generation and accumulation, R l Encourage forklifts to maintain high-quality operations over the long term.
[0016] In the above technical solution, in step 2, the collected data D i = {O i ,E i ,G i ,R i},O i Indicates the operating parameters of the forklift, including travel speed, clamping force, and steering angle; E i Indicates the working environment parameters, including temperature, humidity, light intensity, ground flatness, etc.; G i Indicates cargo status parameters, including cargo weight, size, shape, etc.; R iIndicates the operation results, including whether the clamping is successful, whether there is a collision, and the location of the cargo.
[0017] In the above technical solution, in step 3, it is assumed that data D i One of the parameters in is x, and the true value is x true , the measured value is x measured , and thus calculate the error ∈ x , the mean of the errors and standard deviation j represents different potential error source categories, error ∈ x With each error source component ∈ x,j The relationship between:
[0018]
[0019] δ source (j) is the error source indicator function, k x,j is the coefficient associated with the jth error source, indicating the degree of influence of this error source on the error of parameter x. The coefficient k is estimated by analyzing a large amount of data and using statistical methods. x,j , thereby determining the contribution of different error sources to data errors;
[0020] In the above technical solution, in step 4, the error sources include positioning error sources, map error sources, environmental error sources, control algorithm error sources, environmental perception error sources, actuator error sources and model error sources.
[0021] In the above technical solution, in step 5, the system error ∈ total It is expressed as:
[0022]
[0023] A is the error transfer matrix:
[0024]
[0025] p represents the number of error sources, and m represents the number of types of errors in the final job results;
[0026] The comprehensive operation error estimation value ∈ is obtained by the probabilistic data fusion method based on Bayesian theory fused :
[0027]
[0028] ∈ fused =∫∈P(∈|∈ source1 ,∈ source2 ,…,∈ sourcep )d
[0029] P(∈) is the prior probability distribution, L(∈|∈ sourcei ) is the likelihood function, and the calculated ∈ total ,∈ fused Describe the relationship between the systematic error of the entire operation process and each error source.
[0030] In the above technical solution, in step 5, the comprehensive operation quality index Q is expressed as:
[0031]
[0032] w j is the weight value of the jth error indicator, ∈ j is the jth error index, and m is the number of error indexes.
[0033] In the above technical solution, in step 5, the whole operation error ∈ total-process It is expressed as:
[0034]
[0035] ∈ taski is the system error of the ith subtask, k is the number of subtasks, t i is the time proportion of the ith subtask in the entire operation process,
[0036] In the above technical solution, in step 6, the basic reward item R b Including successful clamping reward R success-grasp , Successful placement reward R success-place , collision penalty R collision And the cargo drop penalty R drop , where R success-grasp =k 1 V g D c , V g is the value of goods, D c is the clamping difficulty coefficient, k 1 is a proportionality coefficient, R success-place =k 2 V g D c P a , P a is the placement accuracy requirement coefficient, k 2 is a proportionality coefficient, R collision =-k 3 C, C is the collision severity coefficient, k 3 is the penalty coefficient, r drop =-k 4 D, D is the coefficient of cargo drop loss, k 4is the penalty coefficient.
[0037] In the above technical solution, the error accumulation reward term is R e =-k e E, E is the error accumulation, k e Is the proportionality coefficient: E = ∑ i ∈ wi Or E = ∫∈ wi (t)dt, where the weighted error for each operation The error vector generated by each operation is the weight vector,
[0038] In the above technical solution, the long-term operation quality reward item or where k l is the reward coefficient, E avg is the average error, or E i is the cumulative error of the i-th operation.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] 1. Aiming at the limitations of traditional reward functions in the operation process of unmanned clamping vehicles, this invention proposes a reinforcement learning reward function design method for error accumulation in the operation process of unmanned clamping vehicles clamping cotton bales, and specifically optimizes the design for error accumulation in the operation process of unmanned clamping vehicles clamping cotton bales. This reward function can accurately evaluate the pros and cons of each operation step according to the real-time operation performance of the clamping vehicle, and timely feedback to the control algorithm, guiding the clamping vehicle to continuously adjust the control strategy, thereby gradually reducing the error accumulation in the operation process, and significantly improving the overall performance of the unmanned clamping vehicle in cotton bale handling operations.
[0041] 2. Through comprehensive and in-depth modeling of various error sources during the operation of the unmanned clamping vehicle, analysis of the error propagation and accumulation mechanism, integration of multi-source error information and accurate evaluation of the operation quality accuracy, it can not only significantly improve the accuracy and stability of the operation, so that the unmanned clamping vehicle can operate accurately and effectively control error accumulation in different complex environments, but also enhance the reliability and adaptability of the system, and respond to environmental changes and potential failure risks in a timely manner. At the same time, it provides a more scientific basis for operation decision-making, optimizes the operation process and parameter adjustment based on precise error feedback, thereby comprehensively improving the operation efficiency and overall performance of the unmanned clamping vehicle, showing excellent application value and competitiveness in many fields such as warehousing and logistics, and effectively promoting the efficient, stable and reliable operation and wide application expansion of unmanned clamping vehicle technology in actual scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Shown is the overall framework of the reward function design method of the present invention. DETAILED DESCRIPTION
[0043] The present invention is further described in detail below in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0044] A method for designing a reinforcement learning reward function for an unmanned gripping vehicle considering error accumulation includes the following steps:
[0045] Step 1: Analysis of the unmanned forklift operation process. By deeply analyzing the operation process of the unmanned forklift in the actual operation scenario, the aim is to comprehensively identify the links where errors may occur during the unmanned forklift's execution of tasks.
[0046] This step is mainly a qualitative analysis. By observing and analyzing the operation process of the unmanned forklift in actual operation, the links that may cause errors are determined. Possible problems in the actual operation process include: the forklift may deviate from the path during driving; the clamping position may be inaccurate or the force may be improper when clamping the goods; there may be a placement deviation when placing the goods, etc. This provides direction for subsequent data collection and error modeling.
[0047] Step 2: Collect and organize data related to unmanned forklift operations. By collecting a large amount of actual operation data, including forklift operating parameters, operating environment parameters, cargo status parameters, operation results, etc., it provides rich data support for subsequent error modeling;
[0048] By performing a large number of test operations on the unmanned forklift at the actual operation site, various data of the unmanned forklift in the real operation environment are obtained. Suppose that by performing N test operations on the unmanned forklift at the actual operation site, a set of data can be obtained for each test operation i (i=1,2,…,N):
[0049] [D i = {O i ,E i ,G i ,R i},]
[0050] Among them, O i Indicates the operating parameters of the forklift, such as travel speed, clamping force, steering angle, etc. i Indicates working environment parameters, such as temperature, humidity, light intensity, ground flatness, etc. i Indicates cargo status parameters, such as cargo weight, size, shape, etc. i Indicates the operation results, such as whether the clamping is successful, whether there is a collision, the location of the cargo, etc.
[0051] Step 3: In-depth analysis of the accuracy, consistency, and reliability of the collected data to identify more potential sources of error. Based on issues such as inaccurate sensor measurements and inappropriate data collection frequency, data quality assurance is provided for subsequent error modeling to ensure that the error model built based on these data has practical application value.
[0052] Assume data D i A parameter x in (x is any specific value of the forklift's operating parameter, operating environment parameter or cargo status parameter) has a true value of x true , the measured value is x measured The error is defined as ∈ x =x measured -x true .
[0053] To evaluate the reliability and accuracy of the data, the mean error can be calculated and standard deviation
[0054]
[0055] N is the number of parameters x. By analyzing these statistics, we can identify the data source that may cause errors. If it is large, it means that there may be a large error in the measurement of the data. Further, the error source indicator function δ is introduced source (j), where j represents different categories of potential error sources, j = 1 means inaccurate sensor measurement, j = 2 means inappropriate data collection frequency, etc.
[0056] Let ∈ x,j Denotes the error component due to the jth error source, we can build a simple linear model to describe the total error ∈ x With each error source component ∈ x,j The relationship between:
[0057]
[0058] Among them, k x,j is the coefficient associated with the jth error source, indicating the degree of influence of this error source on the error of parameter x. These coefficients k can be estimated by analyzing a large amount of data and statistical methods (such as multivariate linear regression) x,j , thereby determining the contribution of different error sources to the data error, and then identifying the main potential error sources.
[0059] Specifically, for the case where the sensor measurement is inaccurate (j=1), if k x,1 The absolute value of is large, and the corresponding δ source(1) If non-zero values appear frequently, it can be determined that inaccurate sensor measurement is an important potential error source. For the case where the data acquisition frequency is inappropriate (j = 2), the appropriate acquisition frequency range can be determined by analyzing the error changes of the data under different acquisition frequencies. When the actual acquisition frequency deviates from this range, δ source (2) takes the value of 1, otherwise it is 0, and then evaluates its impact on the error.
[0060] Step 4: Construct a multi-source error propagation model for unmanned forklift operations. Based on the operation process analysis and data processing, the positioning error sources (including GPS positioning error ∈ GPS (t), drift error Δ∈ GPS , LiDAR positioning error ∈ Lidar ), map error sources (including laser-built maps ∈ map-making and the map update error ∈ map-update ), environmental error sources (including temperature error ∈ temperature (T) and the ground flatness affects the error ∈ flatness (F)), control algorithm error sources (including path tracking algorithm error RMSE control And the decision algorithm error P error (S)), environmental perception error sources (including obstacle detection errors ∈ obstacle (I, d, M) and the cargo identification error ∈ goods-recognition (P, S, L)), actuator error sources (including clamping mechanism actuator error ∈ F (W, D),∈ θ (W, D),∈ P (W, D) and the driving actuator error ∈ v (v, L),∈ θ (v, L)) and model error sources (including ∈ model (v, W)) Establish a multi-source error propagation model for unmanned forklift operations.
[0061] First, we need to construct the positioning error. Let the positioning error of GPS in outdoor environment be ∈ GPS (t), where t represents time. By recording GPS positioning data for a long time and comparing it with the actual position on the high-precision map, the error can be expressed as:
[0062] ∈ GPS (t) = x GPS (t)-x true (t)
[0063] Among them, x GPS (t) is the position measured by GPS at time t, x true(t) is the actual position. For drift error analysis of GPS positioning, the drift rate can be set to r drift , then within the time interval Δt, the drift error Δ∈ GPS It can be expressed approximately as:
[0064] Δ∈ GPS =r drift Δt
[0065] The variation of positioning accuracy over time and environment can be described by the fitting function. Assuming that the positioning accuracy P GPS (t, e) (e represents environmental factors) has a linear relationship with time and environment:
[0066] P GPS (t, e) = a 0 +a 1 t+a 2 e
[0067] Among them, a 0 , a 1 , a 2 is the coefficient obtained by fitting the experimental data. Secondly, in the indoor environment, let the laser radar positioning error be ∈ Lidar , by conducting positioning experiments under different indoor layouts L and weather conditions I, the error can be expressed as:
[0068] ∈ Lidar =f(L,I)
[0069] Among them, f is a function related to indoor layout and weather conditions. The specific form can be obtained by fitting experimental data, for example:
[0070] ∈ Lidar =b 0 +b 1 L+b 2 I+b 3 LI
[0071] Among them, b 0 , b 1 , b 2 , b 3 are the fitting coefficients.
[0072] Next, we construct the map error. Suppose that when using a laser scanner to construct a warehouse map, the map error is ∈ map-making The map data deviation caused by the laser scanning angle θ and the distance measurement error Δd can be expressed as:
[0073] ∈ map-making =g(θ, Δd)
[0074] The relationship is:
[0075] ∈ map-making =c 0 +c 1 θ+c 2 Δd+c 3 θΔd
[0076] Among them, c 0 , c 1 , c 2 , c 3 is a coefficient that needs to be determined through experiments. In addition to map production, there is also the error of map update. Suppose the error of untimely map update caused by changes in the layout of goods and shelf positions in the warehouse is ∈ map-update , which is consistent with the map update time interval ΔT update It is related to the change in cargo layout ΔG and can be expressed as:
[0077] ∈ map-update =h(ΔT update , ΔG)
[0078] Specifically, it can be expressed as the following relationship:
[0079] ∈ map-update =d 0 +d 1 ΔT update +d 2 ΔG+d 3 ΔT update ΔG
[0080] Among them, d 0 , d 1 , d 2 , d 3 is the correlation coefficient.
[0081] Then there are the errors caused by the environment. The first is the influence of ambient temperature. Let the influence of temperature T on the operation error of the unmanned forklift be ∈ temperature (T), by testing the performance of unmanned forklifts under different temperature environments, assuming that there is a linear relationship between operation error and temperature:
[0082] ∈ temperature (T) = e 0 +e 1 T
[0083] Among them, e 0 and e 1 is the coefficient obtained by fitting the experimental data, and is determined by collecting temperature data and corresponding operation error data. The second is the effect of ground flatness. Assume that the error caused by the effect of ground flatness F on the driving stability and operation accuracy of the forklift is ∈ flatness(F), the vibration and attitude change of the vehicle are measured by the accelerometer and gyroscope installed on the forklift. It is assumed that their relationship can be described by a quadratic function:
[0084] ∈ flatness (F) = f 0 +f 1 F+f 2 F 2
[0085] Among them, f 0 , f 1 , f 2 is a coefficient determined experimentally.
[0086] Next is the error of the control algorithm. For the path tracking algorithm of the forklift, let the path planned by the algorithm be P planned (s), the actual ideal path is P ideal (s) (s represents the path parameter), and its error can be quantified by a deviation function, such as the root mean square error (RMSE):
[0087]
[0088] Where M is the number of sampling points on the path. For the decision algorithm error, let the error probability distribution of the decision algorithm under different operation scenarios S (such as cargo stacking density, obstacle distribution, etc.) be P error (S) is obtained through a large number of simulation experiments and actual operation data collection. Assuming that it obeys the normal distribution:
[0089]
[0090] Among them, μ S and σ S are the mean and standard deviation of the error in this scenario respectively.
[0091] Next is the error in environmental perception. First, the error in obstacle detection. When using a camera and lidar to detect obstacles, the perception error under different lighting conditions I, distance d, and obstacle material M is ∈ obstacle (I, d, M). By collecting the detection results and comparing them with the data of the actual obstacle position and size, the relationship is:
[0092] ∈ obstacle (I, d, M) = g 0 +g 1 I+g 2 d+g 3 M+g 4 Id+g 5 IM+g 6 dM+g 7 IUsQ
[0093] Among them, g 0 -g 7 is a coefficient determined by experiment. The second is the error of cargo identification. Let the error of cargo identification algorithm under different cargo packaging P, shape S and label L be ∈ goods-recognition (P, S, L) is obtained by comparing the recognition result with the actual cargo information. The relationship can be expressed as:
[0094] ∈ goods-recognition (P, S, L) = h 0 +h 1 P+h 2 S+h 3 L+h 4 PS+h 5 PL+h 6 SL+h 7 PSL
[0095] Among them, h 0 -h 7 is the correlation coefficient.
[0096] Next is the error modeling of the actuator. The first is the error of the clamping mechanism actuator. When the clamping mechanism clamps goods of different weights W and sizes D, its actual clamping force F actual 、Clamping angle θ actual and clamping position P actual The errors with the target parameters are ∈ F (W, D),∈ θ (W, D) and ∈ P (W, D):
[0097] ∈ F (W, D) = k 0 +k 1 W+k 2 D+k 3 WD
[0098] ∈ θ (W, D) = 1 0 +l 1 W+l 2 D+l 3 WD
[0099] ∈ P (W, D) = m 0 +m 1 W+m 2 D+m 3 WD
[0100] in, and m 0 -m 3is a coefficient determined by experiment. The second is the error of the driving actuator. Assume that the driving actuator (motor, drive wheel, etc.) has a driving speed error of ∈ under different driving speeds v and load conditions L. v (v, L), the steering angle error is ∈ θ (v, L).
[0101] ∈ v (v, L) = n 0 +n 1 v+n 2 L+n 3 v
[0102] ∈ θ (v, L) = o 0 +o 1 v+o 2 L+o 3 v
[0103] Among them, n 0 -n 3 and 0 -o 3 is the correlation coefficient.
[0104] Next is the modeling of model error. Assume that the prediction error of the physical model and mathematical model of the unmanned forklift under different operating conditions (such as driving speed v, load weight W) is ∈ model (v, W). By conducting actual tests on unmanned forklifts under different operating conditions, actual motion data (such as acceleration a, speed change Δv) are collected and compared with the predicted data based on the model, for example:
[0105] ∈ model (v, W) = p 0 +p 1 v+p 2 W+p 3 vW
[0106] Among them, p 0 -p 3 is a coefficient determined experimentally.
[0107] Step 5: Construct a comprehensive evaluation of the operation quality accuracy. Based on the multi-source error propagation model and error propagation characteristics of unmanned forklift operations, data fusion technology is used to fuse multi-source error information and accurately evaluate operation errors.
[0108] Assume that there are p different error sources (such as positioning error source, map error source, environmental error source, control algorithm error source, environmental perception error source, actuator error source, model error source, etc.). In the multi-source error experiment, the data errors of the error sources are ∈ source1 ,∈ source2 , ...,∈sourcep , the system error of the whole operation process ∈ total It can be described by the error transfer matrix A:
[0109]
[0110] The data error of each error source is the sum of all errors in the error source. For example, the data error of the positioning error source ∈ source1 =∈ GPS (t)+Δ∈ GPS +∈ Lidar ; Data error of map error source ∈ source2 =∈ map-making +∈ map-update ; Data error of environmental error source ∈ source3 =∈ temperature (T)+∈ flatness (F); Data error ∈ of the control algorithm error source source4 =RMSE control +P error (S); Data error ∈ of the environmental perception error source source5 =∈ obstacle (I,d,M)+∈ goods-recognition (P, S, L), data error of actuator error source ∈ source6 =∈ F (W,D)+∈ θ (W,D)+∈ P (W,D)+∈ v (v,L)+∈ θ (v, L), data error ∈ source7 =∈ model (v, W).
[0111] Among them, the element a of the error transfer matrix A ij It represents the influence of the j-th error source on the error of the i-th final operation result, which is determined by experiments.
[0112] The error transfer matrix A is an m×p matrix:
[0113]
[0114] p represents the number of error sources, m represents the number of types of errors in the final operation results, and the types of errors in the final operation results include errors in the x, y, and z directions of the cargo placement position, errors in the cargo posture (such as pitch angle, yaw angle, roll angle), and other types.
[0115] The comprehensive operation error estimation value ∈ is obtained by the probabilistic data fusion method based on Bayesian theory fused, assuming that the prior probability distribution is P(∈) and the likelihood function is L(∈|∈ sourcei ), then the posterior probability distribution is:
[0116]
[0117] ∈ fused =∫∈P(∈|∈ source1 ,∈ source2 , ...,∈ sourcep )d
[0118] The calculated ∈ total ,∈ fused Describe the relationship between the system error of the entire operation process and each error source, analyze the contribution of different error sources to the final system error, and find out the error sources that have a greater impact on the final system error, which will help to improve these error sources in a targeted manner during system design and optimization.
[0119] When performing multi-error fusion, assume that multiple sensors measure the same physical quantity (such as the driving speed of a forklift). There are n sensors, and their measured values are v 1 , v 2 , ..., v n , the measurement errors are σ 1 , σ 2 , ..., σ n . Data fusion is performed by weighted averaging to obtain more accurate estimates.
[0120]
[0121] Suppose that the multiple error indicators for evaluating the operation quality of unmanned forklifts are ∈ 1 ,∈ 2 , ...,∈ n In this embodiment, the n error indicators are ∈ GPS (t), Δ∈ GPS ,∈ Lidar ,∈ map-making ∈ map-update ,∈ temperature (T),∈ flatness (F),∈ obstacle (I, d, M),∈ goods-recognition (P,S,L),∈ F (W, D),∈ θ (W, D),∈ P (W,D)∈ v (v, L),∈ θ (v, L),∈ model (v, W), the corresponding weights are w 1 , w2 , ..., w n , the comprehensive operation quality index Q can be expressed as:
[0122]
[0123] The calculated comprehensive operation quality index Q can quantify the overall operation quality of the unmanned forklift in a handling operation. Specifically, the quality of the unmanned forklift operation can be evaluated according to the size of Q, and can be used to optimize the operation process and control strategy of the unmanned forklift to minimize the Q value, thereby improving the operation quality.
[0124] Assume that the errors caused by each subtask of the unmanned forklift in a handling operation (such as driving to the cargo location, clamping cargo, transporting cargo, placing cargo, etc.) are ∈ task1 ,∈ task2 , ...,∈ taskk , the time proportion of each subtask in the entire operation process is t 1 , t 2 , ..., t k Overall operation error∈ total-process It can be expressed as:
[0125]
[0126] The calculated ∈ total-process Taking into account the errors of different subtasks and their time proportion in the operation process, it can more comprehensively reflect the error of the unmanned forklift in a complete handling operation.
[0127] Step 6: Design a reward function for the multi-source error propagation model. The basic reward item regulates the operation behavior, the error accumulation reward item constrains the error generation and accumulation, and the long-term operation quality reward item encourages the forklift to maintain high-quality operation for a long time.
[0128] In this step, you first need to set the basic reward items. Let R b The basic reward items include four aspects:
[0129] The first is the reward for successful clamping, assuming the value of the goods is V g , the clamping difficulty coefficient is D c , Successful hug reward R success-grasp It can be expressed as:
[0130] R success-grasp =k 1 V g D c
[0131] Among them, k 1 is a proportionality factor.
[0132] The second is the successful placement reward, assuming that the placement accuracy requirement coefficient is P a , successful placement reward R success-place It can be expressed as:
[0133] R success-place =k 2 V g D c P a
[0134] Among them, k 2 is another proportionality factor;
[0135] The third is collision penalty. Let collision penalty R collision for:
[0136] R collision =-k 3 C
[0137] Where C is the collision severity coefficient, k 3 is the penalty coefficient;
[0138] The fourth is the cargo drop penalty. Let the cargo drop penalty R drop for:
[0139] R drop =-k 4 D
[0140] Where D is the coefficient of cargo drop loss, k 4 is the penalty coefficient.
[0141] In addition to the basic reward items, it is also necessary to calculate the error value generated by each operation after each operation based on the various errors identified in the error model. Suppose the error vector generated by each operation is:
[0142]
[0143] (where i is the i-th operation and n is the number of error indicators), the weight vector is:
[0144]
[0145] The weighted error for each operation is:
[0146]
[0147] The error accumulation E during the operation can be obtained by summing or integrating the weighted errors of each operation. If the summation method is used:
[0148]
[0149] Or use the integral method, which is suitable for continuous-time systems:
[0150] E=∫∈ wi (t)dt
[0151] Error accumulation reward term R e It is inversely proportional to the error accumulation E. It can be expressed in the form of:
[0152] R e =-k e E
[0153] Among them, k e is the proportional coefficient, which is determined by experiment. When the error accumulation E is small, R e The absolute value of is small, and the impact on the reward function is relatively small; when the error accumulation increases, R e The absolute value of increases, thus significantly reducing the reward value. Finally, the long-term operation quality reward item is designed. Within a certain time window T or operation number N, the average error of the unmanned clamping forklift is calculated. Let the cumulative error of the i-th operation be E i , then the average error can be expressed as:
[0154]
[0155] Or in integral form for continuous operation:
[0156]
[0157] Long-term work quality reward item R l According to the average error E avg To determine. avg Less than a certain threshold E th (need to be set according to the job requirements) when:
[0158]
[0159] Among them, k l is the reward coefficient, determined by experiments; when E avg When the threshold is exceeded, a negative reward is given:
[0160]
[0161] In this way, the long-term operation quality reward item encourages the clamping forklift to maintain a low average error in multiple operations and improve the overall operation quality. It encourages the unmanned clamping vehicle to not only focus on the success or failure of a single operation, but also pay more attention to maintaining stable operation accuracy and reliability in the long-term operation process.
[0162] The final reward function R consists of the basic reward term, the error accumulation reward term, and the long-term job quality reward term, namely:
[0163] R=R b +R e +R l
[0164] Among them, R b The sum of all basic rewards (including the successful hug reward R success-grasp , Successful placement reward R success-place , collision penalty R collision and cargo drop penalty R drop Through such a reward function design, the unmanned gripping vehicle can comprehensively consider the operation results, error accumulation and long-term operation quality during the reinforcement learning training process, and continuously adjust the control strategy based on the feedback of the reward value, effectively deal with the error accumulation problem during the operation process, and improve the accuracy, stability and overall quality of the operation.
[0165] For example, in a continuous handling task, even if a single operation may produce a small error due to some accidental factors, if the overall average error can be controlled at a low level, a certain reward can still be obtained, otherwise it will be punished, thereby guiding the unmanned clamping vehicle to continuously optimize its control strategy and parameter adjustment mechanism to adapt to the long-term stable operation requirements under different operating conditions. In different operating scenarios, such as when carrying fragile items, the unmanned clamping vehicle will operate more cautiously due to the increased penalty for cargo falling and the higher requirements for clamping accuracy (by adjusting the correlation coefficient); in complex environments, due to the increased risk of collision, the constraint effect of collision penalties will prompt it to pay more attention to path planning and obstacle avoidance strategies, thereby achieving flexible adaptation to different operating requirements and continuous optimization of operating performance.
[0166] The above is only a preferred embodiment of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for designing a reinforcement learning reward function for an unmanned gripping vehicle considering error accumulation, characterized in that: The following steps are involved: Step 1: Identify the links that may cause errors in the operation process of unmanned forklifts, and provide direction for data collection and the construction of multi-source error propagation models; Step 2: Collect and organize data related to unmanned forklift operations i ; Step 3: Analyze and collect data D i Accuracy, consistency and reliability, identify potential error sources, and determine the contribution of different error sources to data errors x,j ; Step 4: Construct a multi-source error propagation model for unmanned forklift operation. The system error includes the errors generated by p error sources, which are ∈ source1 ,∈ source2 ,…,∈ sourcep ; Step 5: Based on the multi-source error propagation model and error propagation characteristics of unmanned forklift operation, data fusion technology is used to fuse multi-source error information, and the system error ∈ total , comprehensive operation error estimate ∈ fused Accurately evaluate the operation error and calculate the comprehensive operation quality index Q and the whole operation error ∈ total-process ; Step 6: Design a reward function R for the multi-source error propagation model, R = R b +R e +R l , R b As a basic reward item, R e is the error cumulative reward and R l is the long-term work quality reward item, R b Standardize work behavior, R e Constraint error generation and accumulation, R l Encourage forklifts to maintain high-quality operations over the long term.
2. The method for designing a reinforcement learning reward function for an unmanned gripping vehicle according to claim 1, characterized in that: In step 2, the collected data D i = {O i ,E i ,G i ,R i },O i Indicates the operating parameters of the forklift, including travel speed, clamping force, and steering angle; E i Indicates the working environment parameters, including temperature, humidity, light intensity, ground flatness, etc.; G i Indicates cargo status parameters, including cargo weight, size, shape, etc.; R i Indicates the operation results, including whether the clamping is successful, whether there is a collision, and the location of the cargo.
3. The method for designing a reinforcement learning reward function for an unmanned gripping vehicle according to claim 1, characterized in that: In step 3, let data D i One of the parameters in is x, and the true value is x true , the measured value is x measured , and thus calculate the error ∈ x , the mean error μ ∈x and standard deviation σ ∈x , j represents different potential error source categories, error ∈ x With each error source component ∈ x,j The relationship between: δ source (j) is the error source indicator function, k x,j is the coefficient associated with the jth error source, indicating the degree of influence of this error source on the error of parameter x. The coefficient k is estimated by analyzing a large amount of data and using statistical methods. x,j , thereby determining the contribution of different error sources to the data error.
4. The method for designing a reinforcement learning reward function for an unmanned gripping vehicle according to claim 1, characterized in that: In step 4, the error sources include positioning error sources, map error sources, environmental error sources, control algorithm error sources, environmental perception error sources, actuator error sources and model error sources.
5. The method for designing a reinforcement learning reward function for an unmanned gripping vehicle according to claim 1, characterized in that: In step 5, the system error ∈ total It is expressed as: A is the error transfer matrix: p represents the number of error sources, and m represents the number of types of errors in the final job results; The comprehensive operation error estimation value ∈ is obtained by the probabilistic data fusion method based on Bayesian theory fused : ∈ fused =∫∈P(∈|∈ source1 ,∈ source2 ,…,∈ sourcep )d P(∈) is the prior probability distribution, L(∈|∈ sourcei ) is the likelihood function, and the calculated ∈ total ,∈ fused Describe the relationship between the systematic error of the entire operation process and each error source.
6. The method for designing a reinforcement learning reward function for an unmanned gripping vehicle according to claim 1, characterized in that: In step 5, the comprehensive operation quality index Q is expressed as: w j is the weight value of the jth error indicator, ∈ j is the jth error index, and m is the number of error indexes.
7. The method for designing a reinforcement learning reward function for an unmanned gripping vehicle according to claim 1, characterized in that: In step 5, the whole operation error ∈ total-process It is expressed as: ∈ taski is the system error of the ith subtask, k is the number of subtasks, t i is the time proportion of the ith subtask in the entire operation process, 8. The method for designing a reinforcement learning reward function for an unmanned gripping vehicle according to claim 1, characterized in that: In step 6, the basic reward item R b Including successful clamping reward R success-grasp , Successful placement reward success-place , collision penalty r collision And the cargo drop penalty R drop , where R success-grasp =k1V g D c , V g is the value of goods, D c is the clamping difficulty coefficient, k1 is a proportional coefficient, R success-place =k2V g D c P a , P a is the placement accuracy requirement coefficient, k2 is a proportionality coefficient, R collision =-k3C, C is the collision severity coefficient, k3 is the penalty coefficient, R drop =-k4D, D is the coefficient of cargo drop loss, and k4 is the penalty coefficient.
9. The method for designing a reinforcement learning reward function for an unmanned gripping vehicle according to claim 1, characterized in that: The error accumulation reward is R e =-k e E, E is the error accumulation, k e is the proportionality factor: E=∑ i ∈ wi Or E = ∫∈ wi (t)dt, where the weighted error for each operation The error vector generated by each operation is the weight vector, 10. The method for designing a reinforcement learning reward function for an unmanned gripping vehicle according to claim 1, characterized in that: Long-term work quality reward or where k l is the reward coefficient, E avg is the average error, or E i is the cumulative error of the i-th operation.
Citation Information
Patent Citations
Unmanned vehicle path planning method based on deep reinforcement learning and A star algorithm
CN115933629A
A system and method for locating stacking faults of unmanned forklifts based on machine learning
CN118364411B
Multi-agent unmanned stacking forklift based on event triggering mechanism and control method
CN118479399A