Unmanned clamping and holding vehicle regulation and control algorithm based on human experience pre-training acceleration
Through the combination of machine learning models based on human experience and reinforcement learning algorithms, the problem of long learning cycle of unmanned clamped vehicles in complex logistics environments is solved, and more efficient, safe and stable logistics handling operations are achieved.
Patent Information
- Application Number
- CN202510103089.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
AI Technical Summary
The existing unmanned clamping vehicle control algorithm has a long learning cycle in complex logistics environments and cannot effectively utilize the experience of human operators, resulting in low operating efficiency and high safety risks.
The regulation algorithm based on human experience is adopted to generate rules-based operation data through data acquisition and preprocessing, and the machine learning model is constructed for pre-training, and its parameters are used as the initial parameters of the reinforcement learning algorithm, and the Markov decision-making process and reinforcement learning objective function are trained.
It significantly accelerates the convergence speed of the unmanned clamped vehicle operation control algorithm, improves the safety, accuracy and stability of the operation, enhances the adaptability and flexibility of the algorithm to complex logistics scenarios, and thus improves the overall efficiency and quality of logistics handling operations.
Smart Images

Figure CN120039798A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent logistics, and particularly to a motion control algorithm for an unmanned clamping vehicle accelerated by pre-training based on human experience. Background Art
[0002] The rapid development of the global economy has promoted the continuous expansion of the scale of the logistics industry, which has put forward unprecedentedly strict requirements for the efficiency and accuracy of the goods handling link. As a core device in the logistics handling system, the unmanned clamping vehicle undertakes extremely heavy tasks of clamping and handling goods in key logistics scenarios such as warehousing centers and port terminals. With its automated operation characteristics, it can not only effectively reduce the input of labor costs, but also ensure the stability and efficiency of the operation process, playing an irreplaceable and important role in improving the overall automation level of logistics.
[0003] In the current motion control algorithm system of unmanned clamping vehicles, the reinforcement learning method occupies the mainstream position. For example, in the patent "A Path Planning Method for Unmanned Vehicles Based on Deep Reinforcement Learning and A* Algorithm" with the patent number CN202210357348.0, this method uses a deep neural network and combines data augmentation and curriculum learning techniques to train the unmanned vehicle agent in a simulation environment, aiming to improve the path planning ability of the unmanned vehicle; another example is the research in "Reinforcement learning-driven dynamic obstacle avoidance for mobile robot trajectory tracking (Hanzhen Xiao and Canghao Chen and Guidong Zhang and C.L.Philip Chen)", which focuses on the avoidance technology of mobile robots based on reinforcement learning in dealing with dynamic obstacles during trajectory tracking, providing valuable reference ideas for the dynamic obstacle avoidance of unmanned clamping vehicles in complex logistics environments.
[0004] However, the characteristics of reinforcement learning itself determine that there are obvious shortcomings in its application process. Since it usually needs to go through a large number of trial-and-error attempts to gradually explore the optimal strategy, this process inevitably leads to a long learning cycle in the actual logistics scenario application. Taking a complex logistics environment as an example, when an unmanned clamping vehicle is operating, it needs to deal with a variety of different state variables, such as the goods being in different position coordinates, having different shape characteristics, and the complex and changeable distribution of surrounding obstacles. In this case, relying solely on the reinforcement learning algorithm, the agent needs to spend a huge amount of time and resources to learn operation actions such as the planning of the best driving route and the precise control of the clamping force and angle.
[0005] Meanwhile, during the long-term actual operation of the human operator in the forklift truck, extremely rich and valuable operation experience has been accumulated. These experiences contain effective strategies and techniques for dealing with various complex operation scenarios, which undoubtedly have extremely high guiding value for the precise operation of the unmanned forklift truck. Unfortunately, the traditional planning and control algorithm architecture fails to fully exploit and integrate these human experiences, resulting in the frequent repetition of inefficient or even dangerous operations that human operators already know well and can effectively avoid during the algorithm training process. This not only greatly prolongs the time required for the algorithm to converge to the optimal strategy but also increases the potential safety risks during the operation to a certain extent, seriously restricting the overall efficiency of the unmanned forklift truck in logistics operations. Summary of the Invention
[0006] The object of the present invention is to provide a planning and control algorithm for an unmanned forklift truck based on pre-training acceleration of human experience to address the technical defects existing in the prior art.
[0007] The technical solution adopted to achieve the object of the present invention is as follows:
[0008] A planning and control algorithm for an unmanned forklift truck based on pre-training acceleration of human experience, comprising the following steps:
[0009] Step 1, data collection and preprocessing: Collect the operation data during the operation of the forklift truck, and preprocess the collected operation data to generate rule-based operation data, where the operation data is the operation state s t and the operation action a t ;
[0010] Step 2, construct a machine learning model based on human experience and pre-train: Classify and analyze the rule-based operation data generated in Step 1 to obtain common operation data. After quantifying the common operation data, pre-train the machine learning model, and optimize the parameters of the machine learning model through a loss function until the loss function is minimized, and the pre-training is completed;
[0011] Step 3, use the parameters of the pre-trained machine learning model obtained in Step 2 as the initial parameters of the reinforcement learning algorithm for training; Model the operation environment of the unmanned forklift truck as a Markov decision process, define the operation state space S, the operation action space A, and the reward function r, build the policy network and the value network in the reinforcement learning algorithm, use the Adam optimizer to update the parameters of the policy network and the value network, and introduce a learning rate decay strategy to update the number of training times. During the training process, use the advantage function and the reinforcement learning objective function to monitor the parameter changes of the policy network and the value network and the total reward trend until the reinforcement learning algorithm converges, and the unmanned forklift truck uses the trained reinforcement learning algorithm to operate in different scenarios.
[0012] In the above technical solution, in step 1, the operation status s t includes the position of the forklift (x t , y t ), speed v t , the position of the goods (x g , y g , z g ), the shape of the goods S g and the position of surrounding obstacles That is the operation action a t includes the driving direction d t , speed adjustment Δv(t), clamping force F t and clamping angle θ t , that is, a t = {d t , Δv(t), F t , θ t}.
[0013] In the above technical solution, in step 1, the preprocessing is to first extract features from the operation data and then process the operation data after feature extraction by constructing a decision tree or a rule base to generate rule-based driving data.
[0014] In the above technical solution, step 2 specifically includes the following steps:
[0015] Step 2.1, classify the rule-based operation data according to the type of goods transported by the forklift and the operation scenario where the forklift is located, use statistical analysis and clustering algorithms to analyze and obtain the common operation data of humans in different operation scenarios, and quantify the common operation data;
[0016] Step 2.2, divide the quantified common data into a training set D train and a validation set D val . Input the operation status data s train in the training set D t into the machine learning model and pre-train it. The machine learning model obtains the predicted operation action At the same time, use the mean squared error MSE of the loss function to optimize the parameters of the machine learning model:
[0017]
[0018] Among them, a t is the operation action in the training set D train . Adjust the parameters of the machine learning model by stochastic gradient descent to minimize the value of the loss function. Among them, W newis the weight parameter after gradient descent update, W is the current weight, and η is the learning rate. is the gradient of the mean squared error (MSE) loss function with respect to the current weight W, representing the ascending direction of the loss function in the weight space.
[0019] Step 2.3, during the training process, regularly use the validation set D val to evaluate the predicted operation actions output by the machine learning model. When the loss function value on the validation set D val no longer decreases or the accuracy no longer improves, stop the pre-training.
[0020] In the above technical solution, the machine learning model is a neural network model. The number of neurons in the input layer of the neural network model is n input , corresponding to the dimension of the clamping vehicle operation state characteristics. There are l hidden layers. The number of neurons in the j-th hidden layer is n j , and the number of neurons in the output layer is n output , corresponding to the operation action dimension. The forward propagation process of the neural network (the process of linearly combining the input data through the weights and biases of each layer and then performing non-linear transformation through the activation function to finally obtain the output) is:
[0021]
[0022] where h j is the output of the j-th hidden layer, x i is the i-th input training set D train in the clamping vehicle operation state data, w ij is the weight matrix element from the (j - 1)-th layer to the j-th layer, b j is the bias vector of the j-th layer, σ is the activation function, u is the predicted operation action of the output layer, and σ out is the activation function of the output layer.
[0023] In the above technical solution, when the neural network model continuously outputs predicted operation actions, the activation function of the output layer is the tanh function When the neural network model discretely outputs predicted operation actions, the activation function of the output layer is the softmax function
[0024] In the above technical solution, the machine learning model is a decision tree model. The nodes of the decision tree model represent the clamping vehicle operation state characteristics, and the branches represent the values of the clamping vehicle operation state characteristics.
[0025] In the above technical solution, in Step 3, the state space S includes the clamping vehicle position (x t , y t ), speed v t, the status s of the goods g and the positions of surrounding obstacles are represented as The action space A includes the motion control quantities of the forklift and the control actions of the clamping mechanism. The motion control quantities of the forklift include the braking quantity b ∈ [0, 1], the driving quantity d ∈ [0, 1], and the rotation angle θ; the control actions of the clamping mechanism include the lifting action h, the pitching action the clamping action f, and the transverse movement action l, which are represented as
[0026] The reward function includes the target achievement reward r goal , the distance-to-target reward r dist , the safety reward r safe and the operation efficiency reward r eff . When the forklift successfully clamps and places the target goods at the specified position, a positive reward of r goal = 100 is given. According to the distance d between the forklift and the target goods g the distance reward r dist = -αd g is given, where α is the distance reward coefficient used to adjust the influence of the distance on the total reward. When the forklift avoids collisions and the goods do not fall during operation, a positive safety reward of r safe = 10 is given, otherwise a negative reward of r safe = -50 is given. According to the operation completion time t comp the efficiency reward r eff = -βt comp is given, where β is the efficiency reward coefficient used to adjust the influence of the completion time on the total reward. The total reward r = r goal + r dist + r safe + r eff .
[0027] In the above technical solution, in step 3, the learning rate decay strategy η new = η × γ t , where η new is the decayed learning rate, η is the learning rate, γ is the decay factor, t is the number of training steps, and the clipping parameter ∈ (0.1 - 0.2).
[0028] In the above technical solution, in step 3, a policy network π θ (a|s) and a value network with a multi-layer fully connected neural network architecture are built. The policy network π θ (a|s) represents the probability of taking the operation action a t in the operation space state S, and the value network V φ (s) estimates the value of the operation space state S.
[0029] In the above technical solution, in step 3, the forklift performs operating actions according to the output of the policy network, and records the operation status s at each time step t t , operating action a t , total reward r t and the next operation status s t+1 information, which is stored in buffer R. When the buffer is full, the earliest data is discarded. The exploration and exploitation are balanced through the reinforcement learning algorithm. When the experience data in the buffer reaches the specified amount, batch data is randomly sampled for training, and the advantage function is calculated (used to measure the quality of taking the operating action a t under the state s t relative to the average situation):
[0030] A t = r t + γV φ (s t+1 ) - V φ (s t )
[0031] where V φ (s t ) is the value network function, and φ is the policy.
[0032] In the above technical solution, in step 3, the loss of the policy network and the value network is calculated using the objective function of the reinforcement learning algorithm:
[0033]
[0034] where is the expected value at time step t, r t (θ) is the probability of the policy network taking the action a t under the policy θ, is the estimated value of the advantage function, and clip(r t (θ), 1 - ∈, 1 + ∈) is the clipping function, which is used to limit the amplitude of the policy update to avoid unstable training caused by excessive updates.
[0035] In the above technical solution, in step 3, during the training process, the advantage function and the reinforcement learning objective function are used to monitor the changes in the parameters of the policy network and the value network and the trend of the total reward. If the parameter changes are drastic or the reward fluctuates abnormally, the parameters of the reinforcement learning algorithm are adjusted. The convergence of the reinforcement learning algorithm is judged by observing whether the total reward converges. After convergence, the trained reinforcement learning algorithm is used for the operation of the forklift without a driver.
[0036] In the above technical solution, the hidden layers of the policy network and the value network use the ReLU activation function to increase the non-linear expression ability. For continuous output actions, the tanh function is used to map the output value range, and for discrete output, the softmax function is used to output the probability distribution.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] Effectively integrate human experience and reinforcement learning, significantly accelerate the convergence speed of the operation control algorithm of the unmanned clamping vehicle, improve the operation safety, accuracy and stability, enhance the adaptability and flexibility of the algorithm to complex and changeable logistics scenarios, thereby comprehensively improving the overall efficiency and quality of the logistics handling operation, and reducing costs and risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic flow chart of the control algorithm of the present invention.
[0040] Figure 2 It is a framework diagram of the control algorithm of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] The present invention will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] A control algorithm for an unmanned clamping vehicle based on pre-training acceleration of human experience includes the following steps:
[0043] As Figure 1 - Figure 2 shown, Step 1, data collection and preprocessing: Collect the operation data during the operation of the clamping vehicle, and perform feature extraction on the collected operation data and process the operation data by constructing a decision tree or a rule base to generate rule-based operation data. Among them, the operation data is the clamping vehicle operation state data (including the clamping vehicle driving trajectory data and the clamping vehicle action data) and the operation action data during the human operation, and the preprocessing includes the following for the operation data;
[0044] Further, the operation state s t includes the position of the clamping vehicle (x t , y t ), speed v t , the position of the goods (x g , y g , z g ), the shape of the goods S g and the position of surrounding obstacles That is The operation action a t includes the driving direction d t, speed adjustment Δv(t), clamping force F t and clamping angle θ t , that is, a t ={d t ,Δv(t),F t ,θ t}.
[0045] Extract features from the operation status s t and operation action a t : For example, at time step t, speed acceleration clamping force change rate and clamping angle θ t adjustment frequency f θ ;
[0046] Step 2, construct a machine learning model based on human experience and pre-train it: Classify and process the rule-based operation data obtained in Step 1 and analyze to obtain common operation data. After quantifying the common operation data, pre-train the machine learning model and optimize the parameters of the machine learning model through a loss function until the loss function is minimized, and the pre-training is completed; specifically, it includes the following steps:
[0047] Step 2.1, classify the rule-based operation data according to the type of goods transported by the clamping vehicle and the operation scenario where the clamping vehicle is located, and use statistical analysis and clustering algorithms to analyze and obtain the common operation data of humans in different operation scenarios, and quantify the common operation data;
[0048] Step 2.2, divide the quantified common data into a training set D train and a validation set D val , input the operation status data s train in the training set D t into the machine learning model and pre-train it. The machine learning model obtains the predicted operation action At the same time, use the mean squared error MSE of the loss function to optimize the parameters of the machine learning model:
[0049]
[0050] where, a t is the operation action in the training set D train , adjust the parameters of the machine learning model through stochastic gradient descent to minimize the loss function value, where, W new is the weight parameter after gradient descent update, W is the current weight, η is the learning rate, is the gradient of the mean squared error MSE of the loss function with respect to the current weight W, indicating the ascending direction of the loss function in the weight space;
[0051] Step 2.3, during the training process, regularly use the validation set D val to evaluate the predicted operation actions output by the machine learning model. When the loss function value on the validation set D val no longer decreases or the accuracy no longer improves, stop the pre-training.
[0052]
[0053] where a t is the actual operation action of humans. Through stochastic gradient descent where η is the learning rate, adjust the parameters of the machine learning model to minimize the loss function value. During the training process, regularly use the validation set D val to evaluate the machine learning model. When the loss function value on the validation set D val no longer decreases or the accuracy no longer improves, stop the pre-training.
[0054] Specifically, when the goods are cotton bales, classify and process them according to the cotton bale type (size L×W×H, weight m, shape S, etc.) and the operation scenario of the clamping vehicle (warehouse aisle width w, distribution of surrounding obstacles D, etc.). Use statistical analysis and clustering algorithms to deeply mine the common operation behaviors of humans in different operation scenarios. When handling multiple cotton bales, humans usually choose a relatively low driving speed v low a relatively large clamping force F high and a specific clamping angle θ spec . Determine the quantitative indicators of common operation behaviors by calculating statistics such as the mean and variance of operation data in different operation scenarios. For example, the average driving speed:
[0055]
[0056] where N is the number of operation data in this operation scenario, and v i is the speed of the i-th operation.
[0057] Specifically, when the machine learning model is a decision tree model, the nodes of the decision tree model represent the operation state characteristics of the clamping vehicle, the branches represent the values of the operation state characteristics of the clamping vehicle, and the leaf nodes represent the operation actions. For example, taking the weight of the goods as the root node, if the weight is greater than 100 kg, then for the left branch, determine the corresponding operation strategy based on the aisle width being greater than 4 m, and for the right branch, determine another operation strategy based on the aisle width being less than or equal to 4 m, etc.
[0058] Specifically, when the machine learning model is a neural network model, the number of neurons in the input layer of the neural network model is n input , corresponding to the dimension of the operation state characteristics of the clamping vehicle. There are l hidden layers, and the number of neurons in the j-th hidden layer is n j, the number of neurons in the output layer is n output , corresponding to the operation action dimension, the forward propagation process of the neural network (the process of linearly combining the input data through the weights and biases of each layer, and then performing nonlinear transformation through the activation function to finally obtain the output) is:
[0059]
[0060] Among them, h j is the output of the jth hidden layer, x i is the i-th input training set D of the input layer train Middle clamp car operation status data, w ij is the weight matrix element from the j-1th layer to the jth layer, b j is the bias vector of the jth layer, σ is the activation function, u is the predicted operation action of the output layer, σ out is the activation function of the output layer.
[0061] In the above technical solution, when the neural network model continuously outputs the predicted operation action, the activation function of the output layer is the tanh function When the neural network model predicts the operation action through discrete output, the activation function of the output layer is the softmax function.
[0062] Step 3, use the pre-trained machine learning model parameters obtained in step 2 as the initial parameters of the reinforcement learning algorithm for training; model the unmanned gripping vehicle's operating environment as a Markov decision process, define the operating state space S, the operation action space A and the reward function r, build the policy network and value network in the reinforcement learning algorithm, use the Adam optimizer to update the policy network and value network parameters, and introduce the learning rate decay strategy to update the number of training times. During the training process, the advantage function and the reinforcement learning objective function are used to monitor the policy network and value network parameter changes and the total reward trend. If the parameters change drastically or the total reward fluctuates abnormally, adjust the reinforcement learning algorithm parameters. By observing whether the total reward converges, it is judged whether the reinforcement learning algorithm converges. The unmanned gripping vehicle uses the trained reinforcement learning algorithm to operate in different scenarios.
[0063] Specifically, the state space S includes the clamping vehicle position (x t ,y t ), speed v t , Cargo status g and surrounding obstacle locations Expressed as The action space A includes the movement control quantities of the forklift and the control actions of the clamping mechanism. The movement control quantities of the forklift include the braking quantity b ∈ [0, 1], the driving quantity d ∈ [0, 1], and the rotation angle θ; the control actions of the clamping mechanism include the lifting action h, the pitching action the clamping action f, and the lateral movement action l, expressed as
[0064] The reward function includes the target achievement reward r goal , the distance-to-target reward r dist , the safety reward r safe , and the operation efficiency reward r eff . When the forklift successfully clamps and places the target goods at the designated position, a positive reward of r goal = 100 is given. According to the distance d g between the forklift and the target goods, r dist = -αd g is given as the reward. When the forklift avoids collisions and the goods do not fall during operation, r safe = 10 positive safety reward is given, otherwise r safe = -50 negative reward is given. According to the operation completion time t comp , r eff = -βt comp is given as the reward. The total reward r = r goal + r dist + r safe + r eff ;
[0065] Specifically, a policy network and a value network with a multi-layer fully connected neural network architecture are built. The policy network π θ (a|s) represents the probability of taking the operation action a t in the space state S, and the value network V φ (s) estimates the value of the space state S; the initial learning rate is 0.001, and a learning rate decay strategy is adopted:
[0066] η new = η × γ t
[0067] where γ is the decay factor, t is the number of training steps, the clipping parameter ∈ (0.1 - 0.2), the batch size B is 128, and each batch of data is trained in 20 update rounds; the forklift takes operation actions according to the output of the policy network, and records the state s t , the action a t , the total reward r t , and the next state s t+1The information is stored in buffer R. When the buffer is full, the earliest data is discarded. The exploration and exploitation are balanced through the reinforcement learning algorithm. When the buffer experience data reaches the specified amount, batch data is randomly sampled for training, and the advantage function (used to measure the goodness or badness of taking action a t in state s t relative to the average situation) is calculated:
[0068] A t = r t + γV φ (s t+1 ) - V φ (s t )
[0069] where V φ (s t ) is the value network function and φ is the policy.
[0070] The policy network loss and value network loss are calculated using the reinforcement learning algorithm:
[0071]
[0072] where is the expected value at time step t, r t (θ) is the probability of taking action a t under policy θ of the policy network, is the estimated value of the advantage function, and clip(r t (θ), 1 - ∈, 1 + ∈) is the clipping function used to limit the amplitude of policy updates to avoid unstable training caused by overly large updates.
[0073] The policy network and value network parameters are updated through the Adam optimizer, and the policy update is restricted according to the clipping parameter. The changes in the policy network and value network parameters and the reward trend are monitored.
[0074] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A regulation and control algorithm for an unmanned gripping vehicle based on human experience pre-training acceleration, characterized in that: The following steps are involved: Step 1, data collection and preprocessing: collect the operation data of the clamping vehicle during operation, and preprocess the collected operation data to generate rule-based operation data, wherein the operation data is the operation status s of the human operation process. t and operation action a t ; Step 2: Build and pre-train a machine learning model based on human experience: classify and analyze the rule-based operation data generated in step 1 to obtain common operation data. After quantifying the common operation data, pre-train the machine learning model and optimize the parameters of the machine learning model through the loss function until the loss function is minimized and the pre-training is completed. Step 3, use the pre-trained machine learning model parameters obtained in step 2 as the initial parameters of the reinforcement learning algorithm for training; model the unmanned gripping vehicle's operating environment as a Markov decision process, define the operating state space S, operation action space A and reward function r, build the policy network and value network in the reinforcement learning algorithm, use the Adam optimizer to update the policy network and value network parameters, and introduce the learning rate decay strategy to update the number of training times. During the training process, the advantage function and reinforcement learning objective function are used to monitor the changes in the policy network and value network parameters and the total reward trend until the reinforcement learning algorithm converges. The unmanned gripping vehicle uses the trained reinforcement learning algorithm to operate in different scenarios.
2. The regulation and control algorithm according to claim 1, characterized in that: In step 1, the job status s t Including the clamping car position (x t ,y t ), speed v t , cargo position (x g ,y g ,z g ), Cargo shape S g and the surrounding obstacle positions {x oi ,y oi }, that is, s t ={x t ,y t ,v t ,x g ,y g ,z g ,S g ,x oi ,y oi }; the operation action a t Including driving direction t , speed adjustment Δv(t), clamping force F t and clamping angle θ t , that is, a t ={d t ,Δv(t),F t ,θ t }.
3. The regulation and control algorithm according to claim 1, characterized in that: In the step 1, the preprocessing is to first extract features from the operation data and then process the feature-extracted operation data by building a decision tree or a rule base to generate rule-based driving data.
4. The regulation and control algorithm according to claim 1, characterized in that: The step 2 specifically includes the following steps: Step 2.1, classify the rule-based operation data according to the type of goods transported by the clamp truck and the operation scene where the clamp truck is located, use statistical analysis and clustering algorithm analysis to obtain the common operation data of humans in different operation scenes, and quantify the common operation data; Step 2.2: Divide the quantified common data into training set D train and validation set D val , the training set D train Job status data t Input the machine learning model, pre-train it, and the machine learning model obtains the predicted operation action At the same time, the loss function mean square error MSE is used to optimize the machine learning model parameters: Among them, a t The training set D train The operation actions in , through stochastic gradient descent Adjust the machine learning model parameters to minimize the loss function value, where W new is the weight parameter updated by gradient descent, W is the current weight, η is the learning rate, is the gradient of the mean square error MSE of the loss function with respect to the current weight W, indicating the rising direction of the loss function in the weight space; Step 2.3: Regularly use the validation set D during training. val Evaluate the predicted action output by the machine learning model, when the validation set D val Stop pre-training when the loss function value on the training set no longer decreases or the accuracy no longer improves.
5. The regulation and control algorithm according to claim 1, characterized in that: The machine learning model is a neural network model or a decision tree model; When it is a neural network model, the number of neurons in the input layer of the neural network model is n input , corresponding to the characteristic dimension of the clamping vehicle operation state, there are l hidden layers, and the number of neurons in the jth hidden layer is n j , the number of neurons in the output layer is n output , corresponding to the operation action dimension, the forward propagation process of the neural network (the process of linearly combining the input data through the weights and biases of each layer, and then performing nonlinear transformation through the activation function to finally obtain the output) is: Among them, h j is the output of the jth hidden layer, x i is the i-th input training set D of the input layer train Middle clamp car operation status data, w ij is the weight matrix element from the j-1th layer to the jth layer, b j is the bias vector of the jth layer, σ is the activation function, u is the predicted operation action of the output layer, σ out is the activation function of the output layer.
6. The regulation and control algorithm according to claim 1, characterized in that: When the neural network model continuously outputs predicted operation actions, the activation function of the output layer is the tanh function. When the neural network model predicts the operation action through discrete output, the activation function of the output layer is the softmax function.
7. The regulation and control algorithm according to claim 1, characterized in that: In step 3, the state space S includes the position of the clamping vehicle (x t ,y t ), speed v t , Cargo status g and surrounding obstacle locations Expressed as The motion space A includes the clamping vehicle motion control quantity and the clamping mechanism control action. The clamping vehicle motion control quantity includes the braking quantity b∈[0,1], the driving quantity d∈[0,1] and the rotation angle θ; the clamping mechanism control action includes the lifting action h, the pitching action The clamping action f and the lateral movement action l are expressed as The reward function includes the goal achievement reward r goal , distance target reward r dist 、Safety Reward safe and operational efficiency reward r eff , the clamping vehicle successfully clamps and places the target cargo at the designated location and gives r goal =100 positive reward, based on the distance d between the clamping vehicle and the target cargo g Give distance reward dist =-αd g , α is the distance reward coefficient, which is used to adjust the impact of distance on the total reward). During the clamping operation, avoid collision and do not drop the cargo to give r safe =10 positive safety reward, otherwise r safe =-50 negative reward, based on the task completion time t comp Give efficiency bonus eff =-βt comp , β is the efficiency reward coefficient, which is used to adjust the impact of completion time on the total reward. The total reward r = r goal +r dist +r safe +r eff .
8. The regulation and control algorithm according to claim 1, characterized in that: In step 3, the learning rate decay strategy η new =η×γ t , where η new is the attenuated learning rate, η is the learning rate, γ is the attenuation factor, t is the number of training steps, and the clipping parameter ∈ (0.1-0.2).
9. The regulation and control algorithm according to claim 1, characterized in that: In step 3, a strategy network π containing a multi-layer fully connected neural network architecture is constructed. θ (a|s) and the value network, the strategy network π θ (a|s) means taking operation action a in the working space state S t The probability of value network V φ (s) estimating the value of the working space state S; The gripper car outputs the operation action according to the strategy network and records the operation state s at each time step t. t 、Operation action a t , total reward r t and the next job status s t+1 Information is stored in the buffer R. When the buffer is full, the earliest data is discarded. The exploration and utilization are balanced through the reinforcement learning algorithm. When the buffer experience data reaches the specified amount, batch data is randomly selected for training and the advantage function (used to measure the state s) is calculated. t Take action a t How good is it relative to the average): From t =r t +γV φ (s t+1 )-V φ (s t ) Among them, V φ (s t ) is the value network function and φ is the strategy.
10. The regulation and control algorithm according to claim 1, characterized in that: In step 3, the objective function of the reinforcement learning algorithm is used to calculate the policy network and value network losses: in, is the expected value at time step t, r t (θ) is the policy network's response to action a under policy θ t The probability of is the estimated value of the advantage function, clip(r t (θ),1-∈,1+∈) is a clipping function, which is used to limit the amplitude of policy updates to avoid excessive updates leading to unstable training.
Citation Information
Patent Citations
Unmanned vehicle path planning method based on deep reinforcement learning and A star algorithm
CN115933629A