An intelligent connected vehicle collaborative control method and device in a signal-free intersection scenario
Through the GA-MADDPG algorithm and the safety inspection module, the coordinated control of vehicles without signal intersections is solved, and efficient and safe vehicle scheduling is achieved.
Patent Information
- Application Number
- CN202410522965.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-04-28
AI Technical Summary
Existing machine learning methods cannot efficiently and stably realize vehicle coordinated control in complex signal-free intersection scenarios, resulting in frequent conflicts between vehicles and reduced traffic capacity and safety.
The GA-MADDPG algorithm is used to train the signal-free vehicle collaborative control network model, combine the genetic algorithm to solve pseudo-expert data, introduce a security inspection module, establish a reasonable reward function, and optimize the scheduling strategy of the agent through the GA-MADDPG algorithm.
Improve training efficiency, ensure the safety and traffic efficiency of intersections, reduce the instability of policy updates, output approximate optimal scheduling schemes, and avoid collision accidents.
Smart Images

Figure CN118397854B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent transportation, and relates to a cooperative control method and device for intelligent connected vehicles in the scenario of unsignalized intersections. Background Technique
[0002] The cooperative control of intelligent connected vehicles is an advanced technology that integrates information and communication technology, sensor technology, Internet technology, etc. in the field of traffic scheduling and management, realizing the intelligent cooperative control of vehicles, so as to achieve the goals of making full use of traffic spatio-temporal resources, improving traffic safety levels, alleviating traffic congestion, etc., and guiding the efficient operation of the entire intelligent transportation system.
[0003] The current application of vehicle cooperative control technology in the transportation system mainly focuses on adaptive cooperative cruise control and platoon driving, cooperative merging at highway ramps, cooperative vehicle speed control on highways, cooperative green driving at signalized intersections, and cooperative control at unsignalized intersections. Unsignalized intersections are one of the most common scenarios with traffic conflicts and are the most complex part of the urban road traffic system, playing an important role in converging and dispersing vehicle flows, and realizing the turning and passing of vehicles from multiple lanes in different directions. At the same time, due to the characteristics of large vehicle flow density and many conflict points at intersections, strong coupling interference between vehicles occurs in the intersection area, making it extremely easy for intersections to become congested, greatly reducing the traffic capacity and safety of intersections. Existing research usually models the vehicle passing problem in intersections as an optimization problem, and solves for a specific sequence of vehicle allocation in the intersection and controls the vehicle driving state, so that the trajectories of all vehicles do not intersect at the same time to avoid collisions. Due to the large number of vehicles and complex conflict situations in intersections, there is often a complex and huge solution space. Therefore, traditional optimization-based methods cannot meet the vehicle scheduling requirements of actual complex unsignalized intersections.
[0004] For the problem of vehicle cooperative control at unsignalized intersections in a highly complex and highly dynamic solution space, one solution is to transform the vehicle scheduling problem into a combinatorial optimization problem and use heuristic methods to quickly find a feasible scheduling scheme, such as genetic algorithms, particle swarm algorithms, etc. However, the feasible solutions obtained cannot guarantee being the optimal solutions, and the deviation between the general feasible solutions and the optimal solutions cannot be predicted. Moreover, the control of vehicles within the intersection often requires strict accuracy, otherwise it is easy to cause collision accidents. Another solution is to regard vehicles as agents and model the sequential movement of vehicles passing through the intersection as a multi-agent Markov decision process, thus solving the centralized coordination problem as a machine learning problem. For example, reinforcement learning (RL), multi-agent reinforcement learning (MARL), deep reinforcement learning (DRL), etc. However, the efficiency and performance of current machine learning methods are often greatly affected by the environment, reward function, etc., and cannot efficiently and stably solve the vehicle cooperative control problem in complex unsignalized intersection scenarios. Summary of the Invention
[0005] The present invention provides a cooperative control method for intelligent connected vehicles in an unsignalized intersection scenario to solve the technical problem that existing machine learning methods cannot efficiently and stably achieve vehicle cooperative control in complex unsignalized intersection scenarios.
[0006] To solve the above technical problem, the present invention adopts the following technical solutions:
[0007] The present invention provides a cooperative control method for intelligent connected vehicles in an unsignalized intersection scenario, including the following steps:
[0008] Use the GA-MADDPG algorithm to train the vehicle cooperative control network model for unsignalized intersections. The GA-MADDPG algorithm adds pseudo-expert data solved by the genetic algorithm on the basis of the MADDPG algorithm; input the state information of the agents in the control area into the vehicle cooperative control network model for unsignalized intersections, and the vehicle cooperative control network model for unsignalized intersections outputs the optimal scheduling schemes for each agent ; ;
[0009] Input the optimal scheduling schemes into the safety inspection module. If there are safety risks, correct , otherwise output the final scheduling scheme ;
[0010] Output the final scheduling scheme Transmitted to each agent and executed;
[0011] A vehicle cooperative control network model for unsignalized intersections includes the following steps: setting the hyperparameters of MADDPG and genetic algorithms; initializing the Actor network parameters of the agent , Critic network parameters of the agent , the corresponding target Actor network parameters , target Critic network parameters and reference factor ; initializing the states of the unsignalized intersection agents ; after initializing the vehicles at the unsignalized intersection, obtaining the control strategies of the agents at time ;
[0012] Optimizing the Actor network parameters and Critic network parameters of the agent , specifically: constructing a value function of the MSE loss function and updating the parameters using gradient descent, constructing a reward function of the MSE loss function, calculating the gradient with respect to the parameters and updating the parameters
[0013] Updating the corresponding target Actor network parameters and target Critic network parameters of the agent , specifically: updating the parameters of the target Actor network using the moving average method , updating the target Critic network parameters using the moving average method ; when the preset number of training episodes is reached, the training terminates;
[0014] Reference factor :
[0015]
[0016] where A is the degree of change of the reference factor, k is the decay constant, is the number of training episodes, to is the imitation learning stage; to It is the imitation - multi - agent deep reinforcement learning stage; At it satisfies the first - order exponential decay; To it is the multi - agent deep reinforcement learning stage.
[0017] Furthermore, the final scheduling scheme is transmitted to each agent and executed. Specifically: the final scheduling scheme is transmitted to each agent , if the agent drives out of the control area, it drives out at the maximum speed limit of the road, otherwise it obeys the unified scheduling.
[0018] The present invention also provides an intelligent connected vehicle cooperative control device in a signal - free intersection scenario, including
[0019] An operation module, which is used to train an intelligent vehicle cooperative control network model at a signal - free intersection using the GA - MADDPG algorithm. The GA - MADDPG algorithm is based on the MADDPG algorithm and adds genetic algorithm to solve the pseudo - expert data; the state information of the agents in the control area is input into the intelligent vehicle cooperative control network model at a signal - free intersection, and the output of the intelligent vehicle cooperative control network model at a signal - free intersection is the optimal scheduling scheme for each agent ;
[0020] An inspection module, which is used to input the optimal scheduling scheme into the safety inspection module. If there are safety risks, then is corrected, otherwise the final scheduling scheme is output;
[0021] An execution module, which is used to transmit the final scheduling scheme to each agent and execute it;
[0022] The intelligent vehicle cooperative control network model at a signal - free intersection includes the following steps: setting the hyperparameters of MADDPG and genetic algorithms; initializing the Actor network parameters of the agent , the Critic network parameters , the corresponding target Actor network parameters of the agent , the target Critic network parameters and the reference factor ; initializing the agents at the signal - free intersection Initialize the state; after initializing the vehicles at the unsignalized intersection, obtain agents at the control strategy at time ;
[0023] Optimize the Actor network parameters and Critic network parameters of the agent, specifically: construct the value function of the MSE loss function and update the parameters using gradient descent ; construct the reward function of the MSE loss function, calculate the gradient with respect to the parameters and update the parameters using gradient descent ;
[0024] Update the corresponding target Actor network parameters and target Critic network parameters of the agent, specifically: update the parameters of the target Actor network using the moving average method ; update the target Critic network parameters using the moving average method ; when the preset number of training episodes is reached, the training terminates; ;
[0025] Reference factor :
[0026]
[0027] where A is the degree of change of the reference factor, k is the decay constant, is the number of training episodes, to is the imitation learning stage; to is the imitation - multi - agent deep reinforcement learning stage; At , it satisfies the first - order exponential decay; to is the multi - agent deep reinforcement learning stage.
[0028] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The special feature is that when the processor executes the computer program, it implements the steps of the method for collaborative control of intelligent connected vehicles in an unsignalized intersection scenario described in any one of the above.
[0029] The present invention also provides a computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, the steps of the intelligent connected vehicle cooperative control method in the signal-free intersection scenario described in any one of the above are realized.
[0030] Compared with the prior art, the present invention has the following beneficial technical effects:
[0031] An intelligent connected vehicle cooperative control method in a signal-free intersection scenario of the present invention is applicable to a signal-free intersection scenario. The MADDPG algorithm is used, and pseudo-expert data solved by a genetic algorithm is added to the MADDPG algorithm, which is improved to the GA-MADDPG algorithm, effectively improving the training efficiency. At the same time, considering that the cooperative control model of the signal-free intersection cannot guarantee the absolute safety of the scheduling strategy, a safety inspection module is introduced to ensure the safety of the intersection. The MADDPG algorithm has the characteristic of distributed execution, enabling each intelligent agent to learn independently and adapt to the complex signal-free intersection system. At the same time, centralized training ensures that each intelligent agent will consider the strategies of other intelligent agents during the training process, thereby reducing the instability during policy update and meeting the requirements of cooperative decision-making of intelligent agents.
[0032] An intelligent connected vehicle cooperative control method in a signal-free intersection scenario of the present invention introduces a reference factor to ensure a smooth transition from imitation learning, imitation-multi-agent deep reinforcement learning to multi-agent deep reinforcement learning. At the same time, the genetic algorithm can solve a better scheduling strategy than the MADDPG algorithm in the early stage of training. Therefore, an optimization model for vehicle cooperative control at signal-free intersections is established, and an approximate optimal solution of the model is obtained within a limited number of iterations through the genetic algorithm, thereby outputting an approximate optimal scheduling scheme to guide model learning and solving the problem of excessive trial and error of the MADDPG algorithm in the early stage of training.
[0033] An intelligent connected vehicle cooperative control method in a signal-free intersection scenario of the present invention considers intersection safety, traffic efficiency, and vehicle stability, and establishes a reasonable reward function as the benefit of the MADDPG and genetic algorithms. In addition, in order to ensure the strict safety of the final scheduling scheme, a safety inspection strategy considering the braking distance is established to check one by one the scheduling output by the cooperative control network model of the signal-free intersection, meeting the collision-free requirements within the signal-free intersection. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flow chart of the steps of an intelligent connected vehicle cooperative control method in a signal-free intersection scenario of the present invention;
[0035] Figure 2Schematic diagram of a two-way six-lane signal-free crossroads in an embodiment of the present invention;
[0036] Figure 3 Diagram of the virtual lane conflict judgment method in an embodiment of the present invention;
[0037] Figure 4 Schematic diagram of the Actor network in an embodiment of the present invention;
[0038] Figure 5 Schematic diagram of the Critic network in an embodiment of the present invention;
[0039] Figure 6 Schematic diagram of the genetic algorithm process in an embodiment of the present invention;
[0040] Figure 7 Schematic diagram of the GA-MADDPG algorithm process in an embodiment of the present invention. Detailed implementation manners
[0041] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0042] For the interaction with the environment, the present invention takes a two-way six-lane signal-free crossroads as the research scenario, as Figure 2 shown. It is assumed that the intelligent agents all have clear driving intentions when entering the signal-free intersection control area, and will not change lanes, but drive along fixed paths. In a signal-free intersection, the cross-conflict and diversion-conflict situations are the most complex. A cross-conflict refers to the situation where the vehicle flows coming from different directions cross and then drive to different directions, and a collision occurs within the intersection. Generally, the collision angle is relatively large, such as a side collision. A diversion-conflict refers to the intelligent agents coming from the same entrance of the intersection and driving out of different exits, and a collision occurs within the intersection. Generally, the collision angle is relatively small, such as a rear-end collision. These two conflict situations have the greatest impact on the traffic efficiency and safety of the intersection, and the occurrence probability is also the highest. Therefore, the present invention conducts research on cross-conflicts and diversion-conflicts. In this embodiment, the intelligent agents are the vehicles within the intersection control area, and one intelligent agent represents one vehicle.
[0043] The present invention designs an action network and a value network using a fully connected neural network (FCNN). It should be noted that the action network is the Actor network, and the value network is the Critic network. In a fully connected neural network, each neuron is connected to all neurons in the previous layer, forming a fully connected network structure. Each layer of neurons receives the output of the neurons in the previous layer, processes and extracts features from the input data through a series of non-linear transformations, and finally outputs to the neurons in the next layer or the output layer. Through multiple non-linear transformations, high-order features in the input data can be extracted, so as to more accurately complete the specified tasks. As Figure 4 shown, the Actor network consists of two fully connected layers in total, and each layer is followed by a ReLU activation function. The last layer outputs a single action value, and finally outputs through a tanh activation function. As Figure 5 shown, the Critic network also consists of two fully connected layers, and each layer is followed by a ReLU activation function. The last layer outputs a single scalar value. A double-ended queue is used to construct an experience replay buffer, and the length of the queue is set according to the maximum capacity of the experience pool settings.
[0044] The present invention provides an intelligent connected vehicle collaborative control method in a signal-free intersection scenario. As Figure 1 shown, it includes:
[0045] Input the state information of the agents in the control area into the signal-free intersection vehicle collaborative control network model trained by the GA-MADDPG algorithm. The output of the signal-free intersection vehicle collaborative control network model is the optimal scheduling scheme for each agent ; ;
[0046] Input the optimal scheduling scheme into the safety inspection module. If there is a safety risk, correct it , otherwise output the final scheduling scheme ;
[0047] Transmit the final scheduling scheme to each agent and execute it.
[0048] The GA-MADDPG algorithm is used to train a signal-free intersection collaborative control network model. First, set the hyperparameters of MADDPG and the genetic algorithm, and the Actor network parameters of the agent , Critic network parameters , Agent Corresponding target Actor network parameters , Target Critic network parameters and reference factor Initialize; initialize the state of the vehicle at the unsignalized intersection. The Actor network outputs an action vector in the continuous action space, i.e., the vehicle acceleration, according to the input state vector; the Critic network outputs the Q-value of taking the corresponding action in a certain state according to the input state and action vector, which is used to evaluate the state-action pair.
[0049] Agent The environmental information perceived by the agent and the information generated by the agent's own actions are important bases for the agent to make decisions and estimate long-term rewards. The intelligent connected agent Needs to know the speed, acceleration, position (latitude and longitude coordinates of the vehicle's location), and lane (indicating the driving intention of the vehicle) of the vehicle when driving in the unsignalized intersection. At the same time, it needs to master the speed, acceleration, position, and lane of the surrounding adjacent vehicles. This paper uses the local state representation method to use the information of the vehicle and its surrounding adjacent vehicles as the state space information.
[0050] Furthermore, the virtual lane projection method is used to select the surrounding adjacent agents , specifically:
[0051] The first step: Take the center point of the cross-shaped unsignalized intersection as the center of the circle, and first calculate the Euclidean distance from each agent in the control area to the center point : :
[0052]
[0053] Among them, represents the Euclidean distance between the agent and the center point of the intersection at time , and represents the coordinates of the vehicle at time
[0054] The second step: Rotate and project all agents with the center point as the center of the circle and the Euclidean distance between the agent and the center point of the intersection as the radius onto a one-dimensional virtual lane. For the agents on the lanes with potential cross-conflict and diversion conflict risks require each agent Maintain a safety distance constraint between agents in lanes without potential conflict risks There is no safety distance constraint, and position overlap is allowed.
[0055] Step 3: Agent The information of adjacent vehicles around the agent only focuses on the vehicles in the lanes with conflict risks. And to reduce the complexity of the input, focus on the vehicle with the closest conflict risk to the agent in the virtual lane as the local observation vector of the vehicle at time . As shown in , specifically: First, map all vehicles in the cooperative control area to the virtual lane to form a virtual vehicle fleet. According to Figure 3 the order from small to large, traverse all vehicles in the virtual vehicle fleet. If there is no conflict relationship between vehicles, they can be grouped into the same sub-vehicle fleet. For example vehicles 3 and 1, vehicles 5 and 7 in Figure 3 . If there is a conflict relationship between vehicles, they will be mapped to different sub-vehicle fleets to avoid conflicts. For example Figure 3 vehicles 6 and 3, vehicles 2 and 4 in . Vehicles outside the control area do not need to be considered, such as vehicle 8.
[0056] Furthermore, establish the longitudinal kinematic equation of the agent at time :
[0057]
[0058] where is the decision acceleration of the agent at time , is the local state of the agent at time , is the policy output by the Actor network corresponding to the agent , is the distance traveled by the agent at time , is the vehicle speed of the agent at time , is the control period , is the distance traveled by the agent at time is the vehicle speed of the agent at time
[0059] In an intersection without signals, the agent shall comply with the road restrictions, i.e.:
[0060]
[0061] wherein, and are the minimum and maximum accelerations within the intersection respectively, and are the minimum and maximum speeds within the intersection respectively.
[0062] In multi-agent deep reinforcement learning, every time the agent makes an action, the environment needs to immediately respond with a reward to motivate the agent to learn, encourage it to take correct actions to obtain higher rewards, and the ultimate goal of training is the accumulation of rewards, i.e., maximizing the benefits. In the genetic algorithm, the reward function (also known as the fitness function) is used to evaluate the quality of each individual (or solution), thereby guiding the search process of the algorithm. To ensure that the approximate optimal solution obtained by the genetic algorithm has guiding significance, the reward function of the same MADDPG algorithm and the fitness function of the genetic algorithm are established, and the reward function is designed from three aspects: safety, stability, and traffic efficiency as follows:
[0063] The agent at time has a time headway reward function as follows:
[0064]
[0065] wherein, represents the time headway at time, represents the position of the vehicle closest to the agent around it, represents at time, the reward value of the time headway of the agent represents the minimum safe time headway. is at time, the Euclidean distance from the agent to the center point is at time, the Euclidean distance from the closest agent to the center point ).
[0066] If at Time agent The time headway with the closest neighboring agent is less than , then return a penalty value; if it is not less than , then the return value is 0.
[0067] Agent at time, the stability reward function is:
[0068]
[0069] where represents the acceleration change rate of the agent at time, the reward value of the stability of the agent at is the acceleration of the agent at is the acceleration of the agent at -1 time, is the control period, is the maximum acceleration change rate.
[0070] Agent at time, the passing efficiency reward function is:
[0071]
[0072] In the formula, represents the reward value of the passing efficiency of the agent at represents the speed of the agent at represents the maximum speed of the agent allowed in the intersection, represents the minimum speed of the agent allowed in the intersection.
[0073] To ensure the smooth transition of the unsignalized intersection cooperative control network model from imitation learning, imitation-multi-agent deep reinforcement learning to multi-agent deep reinforcement learning, therefore, the reference factor Ensure that in the imitation learning stage, the multi-agent deep reinforcement learning model can accelerate convergence by referring to the approximately optimal scheduling scheme obtained by the genetic algorithm, and in the multi-agent deep reinforcement learning stage, it can self-explore better scheduling schemes to adapt to more complex environments. In addition, to ensure a smooth transition during the training process, the reward should be gradually adjusted during the imitation - multi-agent deep reinforcement learning stage. Therefore, the reference factor should also meet the requirement of smooth decay. Therefore, the reference factor is designed as follows:
[0074]
[0075] where A is the degree of change of the reference factor, k is the decay constant, is the number of training episodes, to is the imitation learning stage; to is the imitation - multi-agent deep reinforcement learning stage; At it satisfies the first-order exponential decay; to is the multi-agent deep reinforcement learning stage.
[0076] In the imitation learning and imitation - multi-agent deep reinforcement learning stages, the agent At the difference reward function between the approximately optimal scheduling scheme and the optimal scheduling scheme obtained by training at time t is:
[0077]
[0078] where represents the acceleration of the agent at time t, represents the approximately optimal acceleration obtained by solving with the genetic algorithm.
[0079] The value of the reference factor should gradually decrease linearly with the increase of the number of training episodes, and the degree of change should be stable. The degree of change should not be too large to ensure a smooth transition from imitation learning to multi-agent deep reinforcement learning, nor should it be too small, and reasonably control the number of training episodes in the transition stage.
[0080] In the imitation learning stage, the reference factor is relatively large, which means that the model will pay more attention to the feasible solutions obtained by the genetic algorithm and will minimize the difference between the "pseudo-expert data" and the output results of deep reinforcement learning as much as possible; in the imitation-multi-agent deep reinforcement learning stage, the reference factor gradually decreases, which means that the model's attention to the output results of deep reinforcement learning will gradually increase to ensure a smooth transition of the model to the multi-agent deep reinforcement learning stage; in the multi-agent deep reinforcement learning stage, the reference factor is reduced to 0, which means that the model completely relies on the multi-agent to explore autonomously to obtain a better scheduling strategy.
[0081] Furthermore, the multi-objective fitness function of the genetic algorithm is set as:
[0082]
[0083] where, represents the weight of the inter-depot time reward function, represents the weight of the stability reward function, represents the weight of the passing efficiency reward function, is the inter-depot time reward function of agent at time, is the stability reward function of agent at time, is the passing efficiency reward function of agent at time.
[0084] Furthermore, the multi-objective reward function of the multi-agent deep reinforcement learning is set as:
[0085]
[0086] The vehicle cooperative control optimization model for unsignalized intersections is used to solve a joint scheduling strategy to maximize the cooperative benefit within the entire intersection, specifically:
[0087]
[0088] where, represents the fitness function value of an individual in the population, represents the maximum number of individuals in the population for this setting.
[0089] Furthermore, the GA-MADDPG algorithm is trained into an unsignalized intersection cooperative control network model, specifically: after initializing the states of the vehicles at the unsignalized intersection, the first step: agent extracts the th group of experience data from the experience pool Training is carried out, where: the global state , the global action , the global reward , the global state at the next moment , is the number of agents.
[0090] The GA-MADDPG algorithm belongs to the "centralized training, distributed execution" type of algorithm. The Actor network is executed in a "distributed" manner. The input agents respectively at times and of the global observation state and , add noise and then output the decision-making actions and . The Critic network is trained in a "centralized" manner. The current Critic network evaluates the state-action pair of the agent based on the global observation state and the global action state to obtain . The target Critic network evaluates the state-action pair of the agent based on the global observation state and the global action state to obtain , as shown in Figure 7 .
[0091] The objective function is:
[0092]
[0093] where is the action reward value of the agent at time, is the discount factor.
[0094] The loss function of the current Critic network is defined as:
[0095]
[0096] where is the current time, and M is the number of samples drawn.
[0097] In the second step, use the stochastic gradient descent method to update the parameters of the current Critic network :
[0098]
[0099] where Represents the agent 's Critic network parameters, indicating gradient descent calculation.
[0100] In the third step, the current Actor network outputs the action state based on the local observation state . The current Critic network evaluates the state-action pair of the agent ([[]] , ) and obtains .
[0101] The loss function of the current Actor network is defined as:
[0102]
[0103] Using the gradient descent method, update the current Actor network parameters which is expressed as:
[0104]
[0105] where represents the current Actor network parameters of the agent .
[0106] In the fourth step, update the target Actor and Critic network parameters and which is expressed as:
[0107]
[0108]
[0109] where is a constant between 0 and 1, representing the learning rate.
[0110] Furthermore, the genetic algorithm solves the optimization model for vehicle cooperative control at unsignalized intersections as follows:
[0111] First, randomly generate B groups of individuals according to the acceleration constraint in the formula to form the initial population. The length of each individual is N, representing all the vehicles in the intersection at the current moment. The th gene of the individual represents the acceleration of the agent . Then the th individual can be expressed as:
[0112]
[0113] Further, enter the iterative optimization process:
[0114] As Figure 6 shown, first calculate the state of the agent according to . Calculate the fitness of each agent according to the vehicle cooperative control optimization model for unsignalized intersections. Deposit the individual with the highest fitness into the optimal pool . Select the eliminated individuals using the roulette wheel method. The roulette wheel method means that the larger the fitness value of an individual, the larger the area it occupies on the roulette wheel and the greater the probability of being selected. Then, use the multi-point crossover method to swap the genes on individuals and perform mutation, which can increase the diversity of individuals in the population, that is, the diversity of solutions, and generate new offspring.
[0115] When the iteration number is reached, select the optimal fitness value in the optimal pool . The corresponding individual is the approximate optimal scheduling scheme.
[0116] Further, use the cooperative control network model for unsignalized intersections to uniformly schedule the agents in the intersection control area. Specifically:
[0117] The first step: When the agent drives into the control area, within each control cycle , transmit the agent information to the terminal device through the OBU (On Board Unit).
[0118] The second step: The terminal device first uses the virtual lane projection method to output the local state information of each agent , and uses the local state of each vehicle as the input of the Actor network to obtain the optimal scheduling strategy .
[0119] The third step: To ensure the accuracy and safety of the decision-making actions, before outputting the decision-making actions, it is necessary to check the decision-making actions and input the optimal scheduling strategy into the safety inspection module. Calculate the minimum safety distance between the front and rear agents considering the reaction distance and braking distance as:
[0120]
[0121] Among them, and represents the speed of the agent at and moments, and represents the speed of the nearest vehicle of the agent at and moments, represents the optimal decision acceleration of the vehicle output by the Actor network at moment, represents the minimum safe distance of the vehicle represents the reaction time.
[0122] Step 4: If the minimum safe distance of the agent is greater than or equal to the minimum safe distance within the intersection when the agent takes the decision action output by the vehicle collaborative control network model , otherwise decelerates at the minimum acceleration allowed within the intersection . After being checked by the safety inspection module, the final scheduling plan is output . The vehicle scheduling plan is broadcast by the terminal device, and the vehicle receives the broadcast message through the OBU and adjusts the vehicle state after parsing.
[0123] As shown in Table 1 are the key hyperparameters for model training:
[0124] Table 1 Key hyperparameters for model training
[0125]
[0126] The training of the cooperative control network model for unsignalized intersections is divided into three stages: imitation learning, imitation-multi-agent deep reinforcement learning, and multi-agent deep reinforcement learning. In the imitation learning stage, the main goal of training is to enable the model to have basic cooperative control decision-making capabilities under the guidance of "pseudo-expert data", so as to solve the problem of slow convergence speed in the early stage of multi-agent reinforcement learning. The "pseudo-expert data" is obtained by the genetic algorithm. By converting the cooperative control problem of unsignalized intersections into a combinatorial optimization problem and establishing a fitness function with the same objective as the reward function of multi-agent deep reinforcement learning, a guiding feasible solution can be quickly obtained. In the imitation-multi-agent deep reinforcement learning stage, the main purpose is to enable the model to transition from imitation learning to multi-agent deep reinforcement learning. The main goal is to ensure that the model does not deviate from the basic cooperative control decision-making capabilities acquired in the imitation learning stage during the exploration stage of multi-agent deep reinforcement learning. In the multi-agent deep reinforcement learning stage, it mainly relies on the interaction between agents and the environment, and the main goal is to autonomously explore better cooperative control scheduling schemes.
[0127] After three stages of training, the model simultaneously possesses the advantages of heuristic algorithms and multi-agent deep reinforcement learning, that is, finally, a cooperative control network model for unsignalized intersections with a fast convergence speed and strong decision-making ability can be obtained.
[0128] The present invention also provides a cooperative control device for intelligent connected vehicles in an unsignalized intersection scenario, including:
[0129] An operation module for training a cooperative control network model of vehicles at unsignalized intersections using the GA-MADDPG algorithm. The GA-MADDPG algorithm adds pseudo-expert data solved by the genetic algorithm on the basis of the MADDPG algorithm; inputs the state information of agents in the control area into the cooperative control network model of vehicles at unsignalized intersections, and the cooperative control network model of vehicles at unsignalized intersections outputs the optimal scheduling scheme for each agent ;
[0130] An inspection module for inputting the optimal scheduling scheme into a safety inspection module. If there are safety risks, correct , otherwise output the final scheduling scheme ;
[0131] An execution module for transmitting the final scheduling scheme to each agent and executing it;
[0132] The cooperative control network model of vehicles at unsignalized intersections includes the following steps: setting the hyperparameters of MADDPG and the genetic algorithm; the agent Actor network parameters , Critic network parameters , agent corresponding target Actor network parameters , target Critic network parameters and reference factor Initialize; initialize the state of the unsignalized intersection agent ; after initializing the vehicles at the unsignalized intersection, obtain the control policy of the agent at time; ;
[0133] Optimize the Actor network parameters and Critic network parameters of the agent, specifically: construct the value function of the MSE loss function and update the parameters using gradient descent , construct the reward function of the MSE loss function, calculate the gradient with respect to the parameters , and update the parameters using gradient descent ;
[0134] Update the corresponding target Actor network parameters and target Critic network parameters of the agent, specifically: update the parameters of the target Actor network using the moving average method , update the target Critic network parameters using the moving average method ; when the preset number of training episodes is reached, the training terminates; ;
[0135] Reference factor :
[0136]
[0137] where A is the degree of change of the reference factor, k is the decay constant, is the number of training episodes, to is the imitation learning stage; to is the imitation - multi - agent deep reinforcement learning stage; At , it satisfies the first - order exponential decay; to is the multi - agent deep reinforcement learning stage.
[0138] It should be noted that the intelligent connected vehicle collaborative control device provided by the present invention in a signal-free intersection scenario can implement the method steps consistent with the above method, so it will not be elaborated here.
[0139] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The special feature is that when the processor executes the computer program, it implements the steps of the method for intelligent connected vehicle collaborative control in a signal-free intersection scenario described in any one of the above embodiments.
[0140] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.
[0141] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. The feature is that when the computer program is executed by a processor, it implements the steps of the method for intelligent connected vehicle collaborative control in a signal-free intersection scenario described in any one of the above embodiments.
[0142] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a terminal device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. And, in this storage space, there is also stored one or more instructions suitable for being loaded and executed by a processor. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the intelligent connected vehicle collaborative control method in the scenario of signal-free intersections in the above embodiments.
[0143] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0144] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0145] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions in Figure 1 one flow or multiple flows and / or blocksFigure 1 The functions specified in one or more boxes.
[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one process or more processes and / or boxes Figure 1 or more boxes.
[0147] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
Claims
1. An intelligent connected vehicle collaborative control method under a signal-free intersection scenario, characterized in that, Including: Train the vehicle cooperative control network model for unsignalized intersections using the GA-MADDPG algorithm. The GA-MADDPG algorithm adds genetic algorithm to solve pseudo-expert data on the basis of the MADDPG algorithm. Input the state information of the agents in the control area into the vehicle cooperative control network model for unsignalized intersections, and output the optimal scheduling scheme for each agent ; ; Input the optimal scheduling plan into the safety inspection module. If there are safety risks, then make corrections, otherwise output the final scheduling plan ; Transfer the final scheduling plan to each agent and execute it; Vehicle cooperative control network model for unsignalized intersections, including the following steps: setting the hyperparameters of MADDPG and genetic algorithm; initializing the Actor network parameters of the agent , Critic network parameters , the corresponding target Actor network parameters of the agent , target Critic network parameters and reference factors ; initializing the states of the unsignalized intersection agents ; after initializing the vehicles at the unsignalized intersection, obtaining control strategies of the agents at time ; For the agent 's Actor network parameters , Critic network parameters are optimized as follows: construct a value function with an MSE loss function and update the parameters using gradient descent , construct a reward function with an MSE loss function, calculate the gradient with respect to the parameters , and update the parameters using gradient descent For the agent The corresponding target Actor network parameters , and the target Critic network parameters are updated. Specifically: The parameters of the target Actor network are updated using the moving average method , and the parameters of the target Critic network are updated using the moving average method ; When the preset number of training episodes is reached, the training terminates; Reference factor : Among them, A is the degree of change of the reference factor, and k is the decay constant. is the number of training episodes, to is the imitation learning stage; to is the imitation-multi-agent deep reinforcement learning stage; At it satisfies the first-order exponential decay; to is the multi-agent deep reinforcement learning stage.
2. The intelligent connected vehicle cooperative control method in the unsignalized intersection scenario according to claim 1, characterized in that: The final scheduling plan is transmitted to each agent and executed. Specifically: The final scheduling plan is transmitted to each agent . If the agent drives out of the control area, it drives out at the maximum speed limit of the road. Otherwise, it obeys the unified scheduling.
3. An intelligent connected vehicle collaborative control device in a signal-free intersection scenario, characterized in that, Including: An operation module, which is used to train a vehicle cooperative control network model for unsignalized intersections using the GA-MADDPG algorithm. The GA-MADDPG algorithm is based on the MADDPG algorithm and adds a genetic algorithm to solve the pseudo-expert data; the intelligent agents ' state information within the control area is input into the vehicle cooperative control network model for unsignalized intersections, and the optimal scheduling plan for each intelligent agent is output ; Verification module, used to input the optimal scheduling plan into the security verification module. If there are security risks, it will be corrected, otherwise the final scheduling plan will be output ; Execution module, used to transfer the final scheduling plan to each agent and execute it; Vehicle collaborative control network model for signal - free intersections, including the following steps: setting the hyperparameters of MADDPG and genetic algorithms; initializing the Actor network parameters of the agent , Critic network parameters , the corresponding target Actor network parameters of the agent , target Critic network parameters and reference factors ; initializing the states of the signal - free intersection agents ; after initializing the vehicles at the signal - free intersection, obtaining control strategies of the agents at ; For the agent 's Actor network parameters , Critic network parameters are optimized as follows: construct a value function of the MSE loss function and update the parameters using gradient descent , construct a reward function of the MSE loss function, calculate the gradient with respect to the parameters , and update the parameters using gradient descent For the agent The corresponding target Actor network parameters , the target Critic network parameters are updated as follows: The parameters of the target Actor network are updated using the moving average method , and the parameters of the target Critic network are updated using the moving average method ; When the preset number of training episodes is reached, the training terminates; Reference factor : where A is the degree of change of the reference factor, k is the decay constant, is the number of training episodes, to is the imitation learning stage; to is the imitation-multi-agent deep reinforcement learning stage; At it satisfies the first-order exponential decay; to is the multi-agent deep reinforcement learning stage.
4. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the intelligent connected vehicle cooperative control method in the unsignalized intersection scenario as described in any one of claims 1-2 are implemented.
5. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the intelligent connected vehicle cooperative control method in the unsignalized intersection scenario as described in any one of claims 1-2 are implemented.
Citation Information
Patent Citations
Cooperative control method and system for unmanned mine cars at intersection without signal lamps
CN116935676A