Efficient reinforcement learning method for multi-vehicle intersection coordination decision and control

By adopting an efficient reinforcement learning method of coordinated decision-making and control of multiple vehicle intersections in the intelligent logistics system, the problem of low traffic efficiency of multiple unmanned vehicles at intersections without signal lights is solved, and safe and efficient coordinated traffic of multiple unmanned vehicles is achieved.

CN120010267AActive Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510464332.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-16
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

In the prior art, many unmanned vehicles at intersections without signal lights in intelligent logistics systems have low traffic efficiency and lack an effective coordination mechanism.

Method used

An efficient reinforcement learning method of coordinated decision-making and control of multiple vehicle intersections is adopted, including three stages: sample data collection, offline strategy training and online deployment control. Through Markov decision-making process, nuclear sparse method and time-domain differential error of minimizing reinforcement learning, an approximation structure of the action-state value function is constructed, and a synergistic relationship is established between unmanned vehicles, a local joint action return function is introduced, and the global value function is optimized to obtain decision-making actions.

Benefits of technology

The coordinated decision-making efficiency of multi-unit vehicle systems has been improved, and the safe and efficient passage of multiple self-unit vehicles at intersections without signal lights has been achieved, which has improved the traffic efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010267A_ABST
    Figure CN120010267A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient reinforcement learning method for multi-vehicle intersection coordination decision and control, and the method comprises the steps: carrying out the offline strategy training, extracting the features of a collected high-dimensional sample, obtaining an approximately linearly independent sub-sample, constructing a primary function through the sub-sample, and obtaining an approximation structure of an action-state value function; in the on-line deployment control, a decision strategy is deployed according to the observed real-time state quantity of the unmanned vehicles to obtain a utility function after the unmanned vehicles adopt actions, and a local joint action return function is obtained according to the cooperative relationship between the unmanned vehicles, so that a corresponding global value function and a decision action are obtained. And synchronously updating the control strategy and the current state of the unmanned vehicle. The method is applied to the field of multi-unmanned vehicle intersection collaborative decision and control, has the advantages of high value function representation capability and high online calculation efficiency, and can improve the timeliness and safety of multi-unmanned vehicle intersection passing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-unmanned vehicle cooperative control, and specifically to an efficient reinforcement learning method for coordinated decision-making and control of multi-vehicle intersections. Background Art

[0002] In intelligent logistics systems, unmanned vehicles play an important role in scenarios such as material stacking, sorting and handling. Among them, logistics unmanned vehicles, also known as Automatic Guided Vehicles (AGV), are mainly responsible for material transportation. In intelligent transportation systems, most studies focus on the safe passage of a single autonomous unmanned vehicle, while the safe and efficient decision-making and control of multiple unmanned vehicle systems at intersections still need further research. In the current existing technologies, it is proposed to predict the future path of autonomous unmanned vehicles through digital maps, identify potential threats and collision areas, and use Bayesian reasoning and time window filtering methods to make action decisions. There is also a hierarchical decision and planning method based on generalized critical turning points. The upper-level planner extracts parameterized models to generate behavior-oriented paths, and the lower-level planner performs real-time two-dimensional planning. At present, most of the control of intelligent logistics systems only focuses on the safety and effectiveness of the passage of an autonomous unmanned vehicle, but the coordination mechanism of multiple unmanned vehicles passing at intersections remains to be studied. Based on the above analysis, there is little research on the decision-making problem of coordinated passage of multiple unmanned vehicles at intersections without traffic lights in logistics systems. The present invention combines collaborative graphs to study the collaborative decision-making and control mechanism of multiple unmanned vehicle systems based on reinforcement learning, so as to achieve safe and efficient passage of multiple unmanned vehicles. Summary of the invention

[0003] In view of the problem of low efficiency of multiple unmanned vehicles passing through intersections without traffic lights in the prior art of smart logistics systems, the present invention provides an efficient reinforcement learning method for coordinated decision-making and control of multiple vehicle intersections, which can enhance the representation ability of the reinforcement learning value function of multiple unmanned vehicles and improve the efficiency of collaborative decision-making, thereby realizing the safe and efficient passage of multiple unmanned vehicles at intersections without traffic lights.

[0004] To achieve the above-mentioned object, the present invention provides an efficient reinforcement learning method for coordinated decision-making and control of multi-vehicle intersections, including three stages: sample data collection, offline strategy training and online deployment control; In the sample data collection: According to the Markov decision process, a sample set of each unmanned vehicle is generated based on random sampling, where the sample set size of each unmanned vehicle is , the sample tuple for each time step contains , indicating that Current state at the moment Next, in the action space A random action is chosen , the action Will drive the unmanned vehicle to update the status , the reward function is ; In the offline policy training: Based on the sample set of each unmanned vehicle, the kernel sparsification method is used to extract the features of the collected high-dimensional samples to obtain sub-samples that are approximately linearly independent. The sub-samples are used to construct the basis functions corresponding to each original sample point to obtain the approximate structure of the action-state value function. The network weight vector in the approximate structure of the action-state value function is then updated with the goal of minimizing the time-domain differential error of reinforcement learning. In the online deployment control: According to the observed real-time state of the unmanned vehicle, the decision-making strategy is deployed to obtain the action-state value function corresponding to each action, that is, the utility function after the unmanned vehicle takes the action, and the local joint action reward function is obtained according to the collaborative relationship between the unmanned vehicles. The global value function of the unmanned vehicle is obtained based on the utility function and the local joint action reward function, and the decision action of the unmanned vehicle is obtained according to the global value function. Finally, the control strategy of the unmanned vehicle is learned according to the decision action of the unmanned vehicle, and the state of the unmanned vehicle is updated.

[0005] Compared with the prior art, the present invention has the following beneficial technical effects: The present invention includes three stages: sample data collection, offline strategy training and online deployment control. The offline strategy training adopts the least squares strategy iterative reinforcement learning method based on sparse kernel to construct high-dimensional sample features and learn approximate optimal strategies; in online deployment control, the unmanned vehicle obtains the local action-behavior value function under different decisions according to the learned strategy, and at the same time establishes collaborative edges with neighboring unmanned vehicles, introduces a reward function that characterizes the performance of joint actions, and multiple unmanned vehicles solve the optimized joint decision through iterative message propagation on the collaborative edge, and use the decision result as the expected value. Each unmanned vehicle adopts the rolling time domain reinforcement learning method for trajectory tracking control, which has obvious advantages in the efficiency and safety of multi-unmanned vehicle collaborative traffic. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.

[0007] Figure 1 A flowchart of an efficient reinforcement learning method for coordinated decision-making and control of a multi-vehicle intersection in an embodiment of the present invention; Figure 2 A schematic diagram of multiple unmanned vehicle reinforcement learning in an embodiment of the present invention; Figure 3 A diagram showing message propagation between unmanned vehicles in an embodiment of the present invention; Figure 4 A schematic diagram of a simulation example in an embodiment of the present invention; Figure 5 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 0.1s in a simulation example in an embodiment of the present invention; Figure 6 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 2.5s in a simulation example in an embodiment of the present invention; Figure 7 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 3.4s in a simulation example in an embodiment of the present invention; Figure 8 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 4.2s in a simulation example in an embodiment of the present invention; Fig. 9 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 5 seconds in a simulation example in an embodiment of the present invention; Fig.10 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 5.7s in a simulation example in an embodiment of the present invention; Fig.11 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 6.1s in a simulation example in an embodiment of the present invention; Fig.12 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 7.5s in a simulation example in an embodiment of the present invention; Fig.13 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 8.2s in a simulation example in an embodiment of the present invention.

[0008] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0009] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0010] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that ordinary technicians in the field can implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0011] like Figure 1 The figure shows an efficient reinforcement learning method for coordinated decision-making and control of a multi-vehicle intersection disclosed in this embodiment, including three stages: sample data collection, offline strategy training, and online deployment control; Sample data collection in progress: According to the Markov decision process, a sample set for each unmanned vehicle is generated based on random sampling; During offline policy training: Based on the sample set of each unmanned vehicle, the kernel sparsification method is used to extract the features of the collected high-dimensional samples to obtain sub-samples that are approximately linearly independent. The sub-samples are used to construct the basis functions corresponding to each original sample point to obtain the approximate structure of the action-state value function. The network weight vector in the approximate structure of the action-state value function is then updated with the goal of minimizing the time-domain differential error of reinforcement learning. In online deployment control: According to the observed real-time state of the unmanned vehicle, the decision-making strategy is deployed to obtain the action-state value function corresponding to each action, that is, the utility function after the unmanned vehicle takes the action, and the local joint action reward function is obtained according to the collaborative relationship between the unmanned vehicles. The global value function of the unmanned vehicle is obtained based on the utility function and the local joint action reward function, and the decision action of the unmanned vehicle is obtained according to the global value function. Finally, the control strategy of the unmanned vehicle is learned according to the decision action of the unmanned vehicle, and the state of the unmanned vehicle is updated.

[0012] refer to Figure 2 In this embodiment, the collaborative decision-making problem of multiple unmanned vehicles is modeled as a distributed partially observable Markov decision process (Decentralized-Partially Observable Markov Decision Process, Dec-POMDP). Dec-POMDP can effectively handle the serialized decision-making problem of multiple unmanned vehicles, which is specifically described by the following tuple: ; in, for The state space of the unmanned vehicle, for Moment The state of an unmanned vehicle, ; For the The action space of the unmanned vehicle, for Moment The action of the autonomous vehicle decision, ; is the probability of state transition, that is , ; is the reward function, indicating that the unmanned vehicle In Status Take joint action Afterwards in the state The reward obtained when , is the discount factor, For driverless cars Observation space.

[0013] In the specific implementation process, the reward function of each unmanned vehicle can be different, and is generally set according to the specific task requirements. When there is a cooperative relationship between multiple unmanned vehicles, the same reward function is set. When there is a competitive relationship between multiple unmanned vehicles, the competing parties have opposite reward functions. When the reward function is between the two, it indicates that there is a mixed relationship between the multiple unmanned vehicles.

[0014] In the stage of sample data collection, by Randomly initialize the distance of the unmanned vehicle within the range of The speed of the unmanned vehicle is randomly initialized. is the maximum speed of the unmanned vehicle, and each unmanned vehicle collects The data set is used as the sample set. , whose sample tuple for each time step contains , indicating that the unmanned vehicle exist Current state at the moment Next, in the action space A random action is chosen , this decision action will drive the unmanned vehicle to update the state , the reward function is Among them, unmanned vehicles exist Status of the moment , the execution strategy is , For driverless cars speed, For driverless cars The speed of the nearest unmanned vehicle in front, For driverless cars The distance to the nearest unmanned vehicle ahead, For driverless cars The speed of the nearest unmanned vehicle behind, For driverless cars The distance to the nearest unmanned vehicle behind.

[0015] The Probabilistic Graphical Model (PGM) combines probability theory with graph theory, including the representation, reasoning and learning problems of directed graph models (also known as Bayesian networks) or undirected graph models (also known as Markov random fields). From the perspective of graph theory, PGM is a graph containing nodes and edges. Nodes can be divided into two categories: implicit nodes and observation nodes, and edges can be directed or undirected. From the perspective of probability theory, PGM is a probability distribution. The nodes in the graph correspond to random variables, and the edges correspond to the relationships between random variables. In the multi-unmanned vehicle system of this embodiment, the nodes represent unmanned vehicles, and the edges with connection relationships describe the local performance of adjacent unmanned vehicles taking joint actions, that is, based on the current unmanned vehicle information, the posterior distribution of the joint action is inferred.

[0016] The multi-autonomous vehicle reinforcement learning algorithm based on value function decomposition usually adopts a sequence decision framework of centralized training and decentralized execution, which transforms the global action behavior value function into Decompose into individual driverless car value functions The sum of , but it does not take into account that neighboring unmanned vehicles in the observation space each take different actions , The performance of local collaboration is improved. In this embodiment, a collaborative graph is introduced through a collaborative reinforcement learning method to establish connecting edges between unmanned vehicles with collaborative relationships. The messages transmitted on the edges Can fully describe the coordination relationship between multiple unmanned vehicles, that is Figure 3 The combination of the collaborative graph and the variable elimination algorithm (VE) or the Max-plus algorithm can improve the representation ability of the value function and effectively solve the decision-making problem of multiple unmanned vehicle systems. At this time, the global value function It is expressed as: ; in, For driverless cars Adopt an action The utility function after For adjacent unmanned vehicles and driverless cars Take action separately , The local joint action reward function after Assemble for driverless cars; Self-driving cars To give its neighbors unmanned vehicles Send Message , the message indicates that the driverless car Adopt an action After that, driverless cars Can get the biggest return, in this embodiment Specifically: ; in, express Unmanned vehicle The neighbor set of Except for unmanned vehicles A collection of driverless cars; Self-driving cars Before the message converges, The messages received from neighboring nodes are summed, but there is no need to enumerate all the joint action spaces of neighboring nodes. After several iterations of sending and updating messages between neighboring unmanned vehicles, will converge to a fixed point. At each iteration, the unmanned vehicle The global value function of is expressed as: ; in, For driverless cars Adopt an action The global value function after For driverless cars To provide neighbors with driverless cars The message sent, For driverless cars The neighbor set of the unmanned vehicle The global value function of the unmanned vehicle The utility function and neighboring unmanned vehicles The message value composition of different subtrees as roots; Getting a driverless car After the global value function is obtained, the optimal decision action of the unmanned vehicle can be obtained: ; in, Decision-making actions for the driverless car.

[0017] In this embodiment, after obtaining a sufficient number of samples in the sample data collection stage, the kernel-based least squares policy iteration (KLSPI) reinforcement learning method is used to construct an unmanned vehicle. Adopt an action The utility function after , and its specific implementation process is: First, the kernel sparsification method is used to extract the features of the collected high-dimensional samples to obtain approximately linearly independent sub-samples. When the Mercer kernel condition is satisfied, there exists a mapping from the sample space to the Hilbert space ,Right now: ; in, , For the and training samples, is the inner product operation of the Hilbert space, All inner product operations of the surface Hilbert space can be calculated by the kernel function without knowing the specific mapping form. ; Then, the subsamples are used to construct the basis function corresponding to each original sample point, which is: ; in, is the basis function corresponding to the original sample point, is the kernel function, , , , for arrive At different times training samples of unmanned vehicles, is the transpose of the matrix, for dimensional feature space; After that, the approximate structure of the action-state value function is constructed as follows: ; in, is the approximation structure of the action-state value function, is the network weight vector, is the weight matrix; Then, the network weight vector in the approximate structure of the action-state value function is updated with the goal of minimizing the time-domain differential error of reinforcement learning, where the time-domain differential error of reinforcement learning is: ; ; in, is the time-domain difference error of reinforcement learning, is the target action-state value function, is a candidate decision action, , is a positive definite weight matrix, is the action-state value function after state transfer; Then, the approximate structure of the action-state value function in equation (8) is combined to rewrite the time domain difference error of reinforcement learning in equation (9) as follows: ; Then, multiply both sides of the time domain difference error rewritten by reinforcement learning by the basis function , and let the intermediate parameter , for: ; ; Finally, due to the intermediate parameter is full rank, thus obtaining the least squares solution of the network weight vector, namely .

[0018] Therefore, the decision-making strategy of the unmanned vehicle is pre-learned from the collected high-dimensional samples by adopting the kernel feature-based method, and then the online deployment control is deployed according to the observed real-time state quantity to obtain the action-state value function corresponding to each action, that is, the unmanned vehicle Adopt an action The utility function after .

[0019] In addition, according to the basic theory of collaborative graphs, in the collaborative decision-making problem of multiple unmanned vehicles in the intelligent logistics system, the global utility function based on collaborative graphs is expressed as: ; Formula (14) shows that the optimization decision of all unmanned vehicles can be solved by iterating the interactive local joint action rewards on the edges with collaborative relationships, without having to solve the decision actions of multiple unmanned vehicles based on the overall high-dimensional joint action space; Self-driving cars After obtaining neighbors with collaborative relationships, we further define the local joint action reward function on the edge At this moment, driverless cars The state quantity is expressed as ,in , Respectively represent unmanned vehicles The horizontal and vertical coordinates of the coordinated unmanned vehicle state , After the unmanned vehicle performs the joint action, the state quantity of the coordination will also be updated, and we can get In this embodiment, the local joint action reward function on the collaborative edge of adjacent unmanned vehicles in the current state is defined as: ; in, For adjacent unmanned vehicles and driverless cars Take action separately , The local joint action reward function after For the predicted driverless car The distance to the target location, For the predicted driverless car The distance to the target location; Specifically: ; in, , , is the location of the target point.

[0020] The impact of joint actions between collaborative unmanned vehicles is characterized. According to the constructed dynamic collaborative graph and the iterative message propagation mechanism, the unmanned vehicles will decide on a joint action that can make a large difference in the predicted distance, that is, driving the unmanned vehicle that is close to the target point at the current moment to reach the target first, while also avoiding collisions between unmanned vehicles. This is also consistent with the actual situation in the intelligent transportation system, where at intersections without traffic lights, unmanned vehicles on different lanes will decide to let the unmanned vehicle that is closer to the intersection pass first.

[0021] In the construction of unmanned vehicle using kernel-based least squares strategy iterative reinforcement learning method Adopt an action The utility function after , and the local joint action reward function is obtained according to formula (15) , we can substitute into equation (3), (4) and (5) to get the decision action of the unmanned vehicle, so that we can learn the control strategy of the unmanned vehicle system according to the decision action of the unmanned vehicle and update the state of the unmanned vehicle.

[0022] The unmanned vehicle system with nonlinear characteristics is expressed as: ; in, , Unmanned Vehicle exist The state and control input at each moment, is the nonlinear dynamic characteristic of the system state, which is only related to the state of the system related, represents the free motion of the system in the absence of external input, is the influence of input on state change, reflecting the input coupling characteristics of the system. Assume In a set containing the origin is Lipschitz continuous, and There exists a control strategy that makes the system asymptotically stable; The expected trajectory tracked by the unmanned vehicle system in formula (17) is expressed as: ; in, For driverless cars exist The state quantity of the expected trajectory at each moment.

[0023] In the collaborative decision-making and control problem of multiple unmanned vehicle systems, the unmanned vehicle is tracked and controlled with the speed obtained by collaborative decision-making as the expected value. The goal of controller design is to solve an optimized controller so that the state of the unmanned vehicle can track the reference trajectory. Therefore, this embodiment defines the tracking error as Combining equation (17) and equation (18) we can further obtain: ; For the tracking control problem of the unmanned vehicle system in equation (17), the finite time domain performance index is defined as for: ; ; in, For driverless cars The reward function is is the terminal penalty matrix, , is the preset positive definite weight matrix, is the length of the prediction time domain, , For the terminal domain, is the terminal penalty function.

[0024] The nonlinear system of equation (19) can be expressed as follows after linearization at the origin: ; Among them, the intermediate parameters , intermediate parameter ; By solving the Riccati equation of the discrete-time system in equation (22), the terminal penalty matrix can be obtained: .

[0025] The rolling horizon mechanism is used to solve the optimization strategy of unmanned vehicle tracking control. The length of the prediction time domain is The evaluator-executor reinforcement learning is used to obtain the control sequence , take the first control quantity to act on the controlled system, update the system state quantity, and use this as the new initial moment to enter the next prediction time domain. The feasible solution of the optimization problem at any moment is By repeating the above steps, online rolling time domain learning control can be realized.

[0026] At sampling time , with unmanned vehicles Current tracking error is the initial state and enters the prediction time domain Reinforce learning is performed within the prediction domain to define the value function within the prediction domain for: ; According to the Bellman optimization principle, the optimal value function satisfies the HJB equation of the discrete time system, that is: ; Defining the co-state of the value function ,in, is the target value function in the prediction time domain; The optimal cooperative state is: ; So the optimal control strategy is obtained for: ; in, To perform decision-making actions When The control input of the unmanned vehicle system calculated at each moment.

[0027] It can be seen from formula (26) that it is difficult to obtain an analytical expression for the control quantity for a complex system. Therefore, in the prediction domain, this embodiment adopts a reinforcement learning framework of multiple actuator-evaluators, based on the value iteration method, to approximate the time-varying value function and the optimized control quantity through a neural network.

[0028] In the specific implementation process, both the executor and the evaluator are constructed using a three-layer neural network, and both the executor and the evaluator include an input layer, a hidden layer and an output layer.

[0029] In the evaluator, the input layer is the tracking error of the unmanned vehicle. , the number of nodes in the hidden layer is , the output of the output layer is the co-state ; The evaluator approximates the structure of the co-state through a neural network ,for: ; in, is the weight of the evaluator network from the hidden layer to the output layer, is the weight of the evaluator network from the input layer to the hidden layer, is the activation function; In the prediction time domain, the evaluator is based on the time domain difference error To update the network, the time domain difference error for: ; in, Substituting the estimated co-state into equation (25) yields the target co-state, expressed as: ; ; The goal of the evaluator network is to minimize the error function ; Based on the gradient descent method, the weight update rule of the evaluator network is: ; ; ; ; in, is the learning rate of the evaluator network, and .

[0030] In the actuator, the input layer input is the tracking error of the unmanned vehicle , the number of nodes in the hidden layer is The output of the output layer is the control value of the unmanned vehicle ; The actuator adopts a three-layer network structure to approximate the control quantity. ,for: ; in, is the weight of the actuator network from the hidden layer to the output layer, is the weight of the actuator network from the input layer to the hidden layer; In the prediction time domain, the actuator is based on the time domain difference error To update the network, the time domain difference error for: ; in, Substituting the estimated control amount into equation (26) yields the target control amount, expressed as; ; The goal of the actuator network is to minimize the error function ; Based on the gradient descent method, the weight update rule of the actuator network is: ; ; ; ; in, is the learning rate of the actuator network, and ; Due to the real-time requirements of the solution and the limitation of actual computing resources, it is impossible to perform infinite iterations of reinforcement learning in the prediction time domain. Therefore, in this embodiment, the termination condition of reinforcement learning in the prediction time domain is: ; ; in, To predict the number of iterations of reinforcement learning in the time domain, is the threshold for the convergence of the evaluator network weights, is the threshold for the convergence of the actuator network weights; When the network weights of the previous and next two iterations meet the termination condition, it is considered that the converged strategy has been learned, and the first actuator network in the optimized solution control sequence is applied to the actual controlled unmanned vehicle system. On the other hand, the maximum number of learning times can also be given. , if in If the internal strategy has not converged, the network is reinitialized and the actuator-evaluator learning is performed in the prediction domain.

[0031] It should be noted that under the rolling horizon reinforcement learning mechanism, the prediction horizon length is ,exist Learning the optimal control sequence in the prediction time domain , the first control element acts on the controlled unmanned vehicle system. In two adjacent prediction time domains, the evaluator network and the policy network are continuous, that is, in After learning The actuator network will be passed to Before the moment The corresponding evaluator network will also be passed to the next prediction time domain. That is, the previous experience results are fully utilized in the learning process, which helps to accelerate the convergence speed of learning. The number of learning iterations in the subsequent prediction time domain is gradually reduced, which improves the real-time performance of the optimization strategy solution and facilitates transplantation to the actual system.

[0032] The following further illustrates the efficient reinforcement learning method for coordinated decision-making and control at multi-vehicle intersections in this embodiment with reference to specific simulation examples.

[0033] exist Figure 4 In the simulation scene of the intersection without traffic lights shown, there are straight-moving unmanned vehicles, right-turning unmanned vehicles, and left-turning unmanned vehicles. At the same time, there are many social unmanned vehicles in the scene. The left-turning unmanned vehicle may collide with the right-turning unmanned vehicle. It is necessary to establish a collaborative relationship so that vehicles in different lanes can make reasonable joint decisions to pass the intersection safely.

[0034] The analysis of the unmanned vehicle collaboration process in the simulation scenario is as follows: When the straight-moving unmanned vehicle enters the decision area, the running speed and relative distance of the unmanned vehicle in the left-turn lane are observed to obtain the state of the straight-moving unmanned vehicle. At the same time, a collaborative relationship is established with the unmanned vehicle in the left-turn lane, and the learned decision-making strategy is deployed to obtain the behavior-action value function under discrete actions, and the local action reward function is combined to derive the optimized decision-making action using a collaborative message passing mechanism; Similarly, when the unmanned vehicle in the left-turn lane enters the decision area, it establishes a cooperative relationship with the unmanned vehicle going straight and the unmanned vehicle turning right, observes the status of the unmanned vehicles in the relevant lanes, and uses the method of this embodiment to obtain an optimized joint decision action. When the unmanned vehicle in the right-turn lane enters the decision area, it establishes a neighbor cooperative relationship with the unmanned vehicle turning left.

[0035] Based on the above analysis of the cooperative process vehicles, the learning strategy is deployed in the intersection traffic scenario to test the performance of the method of this embodiment. Figures 5 to 13 The following diagrams show the decision-making states of the unmanned vehicle at some typical moments in a test in a simulation scenario, among which: Figure 5 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 0.1s. Figure 6 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 2.5s. Figure 7 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 3.4s. Figure 8 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 4.2s. Fig. 9 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 5 seconds. Fig.10 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 5.7s. Fig.11 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 6.1s. Fig.12 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 7.5s. Fig.13 This is a schematic diagram of the collaborative decision-making of the unmanned vehicle system at 8.2s. Unmanned vehicles c0~c11 Figures 6 to 13 The color identification in Figure 5 The color of the vehicle is the same as that of the intersection. The unmanned vehicles start from their initial positions, and the initial speeds of each unmanned vehicle are also random. Before entering the decision area, the unmanned vehicle is tracked and controlled at the expected speed of 8 m / s.

[0036] like Figure 5 As shown, when , indicating the desired direction of travel of the unmanned vehicle. The black box indicates the social unmanned vehicle on the lane, and the lane is congested due to traffic control and other reasons, so the unmanned vehicle turning left will pass to the lane with less traffic; like Figure 6 As shown, when At this time, the unmanned vehicle c6 traveling straight in the decision area observes that the front and rear vehicles are both within the safe distance range at this moment, and decides to accelerate. At this time, the unmanned vehicle c8 turning left observes that the front has just passed c0, and for safety reasons c8 decides to slow down. like Figure 7 As shown, when At this time, the unmanned vehicle c2 turning left enters the decision area, and obtains that the unmanned vehicle c6 is moving straight at a certain speed from the rear. The method of this embodiment decides a joint action of c2 decelerating and c6 accelerating. At this time, the unmanned vehicle c8 turning left also enters the decision area, obtains the status of the neighboring unmanned vehicles in front and behind, and further decides the action of c8 accelerating; like Fig. 9 As shown, when At this time, c7 going straight enters the decision area and establishes a collaborative relationship with c2 and c3 turning left. The combined strategy and current state result in c7 slowing down and waiting, while c2 accelerates to pass the lane intersection. At this time, c3 adjusts its speed according to the decision action and the unmanned vehicle following control. At this time, c1 going straight also enters the decision area. Similarly, the algorithm decides to slow down. like Fig.10 As shown, when At this time, c2, which is turning left, observes that the unmanned vehicle c11 is turning right in front of it, so c2 will slow down, and c7 will continue to slow down and wait. c1, which is going straight, has seen that the front vehicle c8 has passed, and the rear vehicle c9 is now in a state of slowing down and waiting. c1 and c9 maintain a relatively safe distance, so c1 decides to speed up and pass. c8 enters a new decision area and establishes a collaborative relationship with the right-turning unmanned vehicle. According to the learned strategy, the joint action of c5 slowing down and c8 speeding up is obtained, which is also in line with the driving rules of turning right to turning left in daily life; like Fig.11 As shown, when When , among the neighboring unmanned vehicles with a collaborative relationship, the unmanned vehicles c2, c5, and c9 located in their respective decision-making areas obtain the states of the unmanned vehicles in front and behind, and decide to continue waiting after deploying strategies; like Fig.12As shown in the figure, after a period of operation, the unmanned vehicle At that time, the previous congestion at the intersection had been greatly alleviated, and more driverless cars passed through the desired lanes; like Fig.13 As shown, when It can be seen that the unmanned vehicle has basically passed the intersection.

[0037] exist Figure 4 In the scene of crossing the intersection without traffic lights, the current unmanned vehicle establishes a collaborative relationship with the neighboring unmanned vehicle, deploys the learned strategy, observes the current state, and infers a reasonable joint decision action by passing messages on the collaborative edge to safely pass through the intersection. A total of 100 tests were conducted, and the initial position of the unmanned vehicle in each test was 100 meters away from the intersection. Random inside, the speed of the driverless car is m / s, and the travel time, collision rate, failure rate and comfort index of the unmanned vehicles at the intersection are counted. Among them, the travel time refers to the time after all unmanned vehicles pass through the intersection and reach the desired lane, the collision rate refers to the number of collisions between unmanned vehicles in 1000 tests, and the failure rate refers to the proportion of times that there are still unmanned vehicles in the intersection within 10s and have not completed the intersection passage task. Comfort is the average value of the root mean square acceleration of the unmanned vehicle. The smaller the value, the higher the comfort. The performance of the method of this embodiment is compared with the existing independent learning method. The results are shown in Table 1.

[0038]

[0039] According to Table 1, the average and standard deviation of the independent learning method in the pass time are greater than those of the method in this embodiment. The collision rate of independent learning is 26%. The failure rate refers to the proportion of the pass task that has not been completed within 13s. The failure rate of the independent learning method is 14. . 5%, the comfort of the method in this embodiment is better than that of the independent learning method. It can be seen that the decision-making control performance of the method in this embodiment in terms of traffic efficiency, safety and driving stability of multiple vehicles at the intersection is better than that of the independent learning method.

[0040] This is because in the collaborative decision-making process of the multi-unmanned vehicle system, the method of this embodiment decomposes the value function of the system as a whole into two parts: a single action-state value function and the utility of the joint action between unmanned vehicles. This is different from the independent learning method that only focuses on individual performance. While deploying strategies, the rewards of joint actions are considered. The designed joint action reward function effectively represents the collaborative goals between intelligent agents. Through iterative message propagation between neighboring intelligent agents with collaborative relationships, the optimized strategy is converged, that is, a reasonable decision action is obtained. The method of this embodiment is applied to the scene of multiple unmanned vehicles passing through intersections without traffic signal indications. It can improve traffic efficiency while avoiding collisions and ensure the safe passage of multiple unmanned vehicles.

[0041] The above description is only a preferred embodiment of the present invention, and does not limit the protection scope of the present invention. All equivalent structural changes made by using the contents of the present invention specification and drawings under the inventive concept of the present invention, or directly / indirectly applied in other related technical fields are included in the protection scope of the present invention.

Claims

1. An efficient reinforcement learning method for coordinated decision-making and control at multi-vehicle intersections, characterized in that: It includes three stages: sample data collection, offline strategy training and online deployment control; In the sample data collection: According to the Markov decision process, a sample set of each unmanned vehicle is generated based on random sampling, where the sample set size of each unmanned vehicle is , the sample tuple for each time step contains , indicating that Current state at the moment Next, in the action space A random action is chosen , the action Will drive the unmanned vehicle to update the status , the reward function is ; In the offline policy training: Based on the sample set of each unmanned vehicle, the kernel sparsification method is used to extract the features of the collected high-dimensional samples to obtain sub-samples that are approximately linearly independent. The sub-samples are used to construct the basis functions corresponding to each original sample point to obtain the approximate structure of the action-state value function. The network weight vector in the approximate structure of the action-state value function is then updated with the goal of minimizing the time-domain differential error of reinforcement learning. In the online deployment control: According to the observed real-time state of the unmanned vehicle, the decision-making strategy is deployed to obtain the action-state value function corresponding to each action, that is, the utility function after the unmanned vehicle takes the action, and the local joint action reward function is obtained according to the collaborative relationship between the unmanned vehicles. The global value function of the unmanned vehicle is obtained based on the utility function and the local joint action reward function, and the decision action of the unmanned vehicle is obtained according to the global value function. Finally, the control strategy of the unmanned vehicle is learned according to the decision action of the unmanned vehicle, and the state of the unmanned vehicle is updated.

2. The efficient reinforcement learning method for coordinated decision-making and control of multi-vehicle intersections according to claim 1 is characterized in that: In the offline strategy training, the approximate structure of the action-state value function is: in, is the basis function corresponding to the original sample point, is the kernel function, , , , for arrive At different times training samples of unmanned vehicles, is the transpose of the matrix, for dimensional feature space, is the approximation structure of the action-state value function, is the network weight vector, is the weight matrix.

3. The efficient reinforcement learning method for coordinated decision-making and control of multi-vehicle intersections according to claim 2 is characterized in that: In the offline strategy training, the updating of the network weight vector in the approximation structure of the action-state value function with the goal of minimizing the time-domain difference error of reinforcement learning includes: Get the time domain difference error of reinforcement learning: in, is the time-domain difference error of reinforcement learning, is the target action-state value function, is a candidate decision action, , is a positive definite weight matrix, is the discount factor, is the action-state value function after state transfer; Based on the approximate structure of the action-state value function, the time domain difference error of reinforcement learning is rewritten as: Multiply both sides of the time domain difference error rewritten by reinforcement learning by the basis function , and let the intermediate parameter , for: Because the intermediate parameter is full rank, thus obtaining the least squares solution of the network weight vector, namely .

4. The efficient reinforcement learning method for coordinated decision-making and control of multi-vehicle intersections according to claim 1, 2 or 3, characterized in that: In the online deployment control, the local joint action reward function is obtained as: in, For adjacent unmanned vehicles and driverless cars Take actions separately , The local joint action reward function after For the predicted driverless car The distance to the target location, For the predicted driverless car The distance to the target location.

5. The efficient reinforcement learning method for coordinated decision-making and control of multi-vehicle intersections according to claim 4 is characterized in that: In the online deployment control, the global value function of the unmanned vehicle based on the utility function and the local joint action reward function is specifically: in, For driverless cars Adopt an action The global value function after For driverless cars Adopt an action The utility function after For driverless cars To provide neighbors with driverless cars The message sent, For driverless cars The set of neighbors of ; The message sent by the unmanned vehicle to the neighboring unmanned vehicle is defined as: in, For driverless cars To provide neighbors with driverless cars The message sent, express Unmanned vehicle The neighbor set Except for unmanned vehicles A collection of driverless cars.

6. The efficient reinforcement learning method for coordinated decision-making and control of multi-vehicle intersections according to claim 5 is characterized in that: In the online deployment control, the decision action of the unmanned vehicle is: in, Decision-making actions for the driverless car.

7. The efficient reinforcement learning method for coordinated decision-making and control of multi-vehicle intersections according to claim 1, 2 or 3, characterized in that: In the online deployment control, the control strategy of the unmanned vehicle is learned according to the decision action of the unmanned vehicle, and the state of the unmanned vehicle is updated, specifically: At sampling time , with unmanned vehicles Current tracking error is the initial state and enters the prediction time domain Reinforcement learning is performed within, where is the length of the prediction time domain; Define the value function in the prediction time domain for: in, For driverless cars exist The tracking error at that moment, For driverless cars exist The control input at the moment, For driverless cars The reward function is To predict the terminal cost in the time domain, is the terminal penalty matrix, is the transpose of the matrix; According to the Bellman optimization principle, the optimal value function satisfies the HJB equation of the discrete time system, that is: in, For driverless cars Control input of Defining the co-state of the value function ,in, is the target value function; The optimal cooperative state is: in, is the preset positive definite weight matrix; Therefore, the optimal control strategy for: in, For driverless cars Decision-making actions The corresponding control input.

8. The efficient reinforcement learning method for coordinated decision-making and control of multi-vehicle intersections according to claim 7 is characterized in that: In the online deployment control, a reinforcement learning framework of multiple actuator-evaluators is adopted. Based on the value iteration method, a neural network is used to approximate the time-varying value function and the optimal control strategy. Specifically: Both the actuator and the evaluator are constructed using a three-layer neural network, and both the actuator and the evaluator include an input layer, a hidden layer, and an output layer; In the evaluator, the input layer is the tracking error of the unmanned vehicle. , the number of nodes in the hidden layer is , the output of the output layer is the co-state ; The evaluator approximates the structure of the co-state through a neural network ,for: in, is the weight of the evaluator network from the hidden layer to the output layer, is the weight of the evaluator network from the input layer to the hidden layer, is the activation function; In the prediction time domain, the evaluator is based on the time domain difference error To update the network, the time domain difference error for: in, is the target co-state obtained according to the estimated co-state, expressed as: The goal of the evaluator network is to minimize the error function ; Based on the gradient descent method, the weight update rule of the evaluator network is: in, is the learning rate of the evaluator network, and ; In the actuator, the input layer input is the tracking error of the unmanned vehicle , the number of nodes in the hidden layer is The output of the output layer is the control value of the unmanned vehicle ; The actuator adopts a three-layer network structure to approximate the control quantity. ,for: in, is the weight of the actuator network from the hidden layer to the output layer, is the weight of the actuator network from the input layer to the hidden layer; In the prediction time domain, the actuator is based on the time domain difference error To update the network, the time domain difference error for: in, is the target control quantity obtained according to the estimated control quantity, expressed as: The goal of the actuator network is to minimize the error function ; Based on the gradient descent method, the weight update rule of the actuator network is: in, is the learning rate of the actuator network, and ; The termination condition of reinforcement learning in the prediction domain is: in, To predict the number of iterations of reinforcement learning in the time domain, is the threshold for the convergence of the evaluator network weights, is the threshold for the convergence of the actuator network weights; When the network weights of the previous and next two iterations meet the termination conditions, the first actuator network in the optimized solution control sequence at this time is applied to the actual controlled unmanned vehicle system.

Citation Information

Patent Citations

  • Unmanned vehicle control method and device based on data driving and computer equipment

    CN113534669A

  • Dexterous two-hand cooperative control method based on multi-agent reinforcement learning

    CN119795175A

  • Reinforcement learning using meta-learned intrinsic rewards

    US20210089910A1