Method and device for controlling automatic driving vehicle at intersection, and storage medium

By constructing conflict state diagrams and using decision-making neural network models to predict acceleration, the problem of low traffic efficiency of cross-driving vehicles in the existing technology is solved, and efficient and safe vehicle coordination is achieved.

CN120199103APending Publication Date: 2025-06-24HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311774333.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing autonomous driving vehicle control strategy has a large traffic flow, and the traffic efficiency of the intersection is reduced, and it is difficult to achieve refined coordination between vehicles.

Method used

By constructing a conflict state diagram between autonomous vehicles, combining the decision-making neural network model, the vehicle's acceleration is predicted and controlled based on this, the vehicle's refined scheduling is achieved.

Benefits of technology

It improves the traffic efficiency and safety of autonomous driving vehicles at the intersection, reduces the number of parking times, avoids collisions, and achieves efficient coordination between vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199103A_ABST
    Figure CN120199103A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intersection automatic driving vehicle control method and device and a storage medium, and the method comprises the steps: constructing a conflict state diagram corresponding to a target automatic driving vehicle at a current time step length based on the collision conflict relation between the target automatic driving vehicle and surrounding automatic driving vehicles at an intersection through an intelligent agent deployed in the target automatic driving vehicle, the conflict state diagram comprises initial driving state feature vectors corresponding to a plurality of automatic driving vehicles, and the plurality of automatic driving vehicles comprise a target automatic driving vehicle and adjacent automatic driving vehicles meeting the collision conflict relation with the target automatic driving vehicle at the intersection. And according to the conflict state diagram and the decision neural network model, obtaining a predicted acceleration corresponding to the target autonomous vehicle at the next time step, and according to the predicted acceleration, controlling the target autonomous vehicle to run at the next time step. According to the scheme, the passing efficiency and safety of the automatic driving vehicle in the intersection scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method, device and storage medium for controlling an autonomous vehicle at an intersection. Background Art

[0002] With the gradual maturity of technologies such as vehicle-road cooperation and autonomous driving, the "in-vehicle traffic lights" will potentially replace the physical traffic lights at existing intersections (i.e., crossroads or T-junctions). Roadside devices will control autonomous vehicles to pass through the intersection based on a certain control strategy according to the arrival situations of vehicles in all directions at the intersection, so as to improve the vehicle passing efficiency at the intersection. Among them, the so-called "in-vehicle traffic lights" refer to virtual traffic light signals displayed in the vehicle. For example, when it is determined based on the control strategy that a certain autonomous vehicle passes through the intersection, a green light signal is displayed on the on-vehicle terminal interface in the autonomous vehicle, so as to control the vehicle to pass through the intersection based on the green light signal.

[0003] However, currently, the above control strategy can only make high-level decision results on whether to pass through the intersection, and currently the control strategy mainly still makes release decisions according to the vehicle arrival order. In the case of large traffic flow, the passing efficiency of the intersection will decrease. Summary of the Invention

[0004] Embodiments of the present invention provide a method, device and storage medium for controlling an autonomous vehicle at an intersection, so as to improve the passing efficiency and safety of autonomous vehicles in the intersection scenario.

[0005] In a first aspect, an embodiment of the present invention provides a method for controlling an autonomous vehicle at an intersection, which is applied to an agent deployed in a target autonomous vehicle. The method includes:

[0006] At the current time step, based on the collision conflict relationship with surrounding autonomous vehicles at the intersection, construct a conflict state graph corresponding to the target autonomous vehicle. The conflict state graph includes initial driving state feature vectors respectively corresponding to multiple autonomous vehicles. The multiple autonomous vehicles include the target autonomous vehicle and neighboring autonomous vehicles that satisfy the collision conflict relationship with the target autonomous vehicle at the intersection;

[0007] According to the conflict state graph and a decision neural network model, obtain the predicted acceleration corresponding to the target autonomous vehicle at the next time step. The decision neural network model is trained by means of deep reinforcement learning;

[0008] Control the driving of the target autonomous vehicle at the next time step according to the predicted acceleration.

[0009] Second aspect, an embodiment of the present invention provides a control device for an autonomous driving vehicle at an intersection, which is applied to an agent deployed in a target autonomous driving vehicle. The device includes:

[0010] A composition module, configured to construct a conflict state graph corresponding to the target autonomous driving vehicle at the current time step based on the collision conflict relationship with surrounding autonomous driving vehicles at the intersection. The conflict state graph includes initial driving state feature vectors corresponding to multiple autonomous driving vehicles, and the multiple autonomous driving vehicles include the target autonomous driving vehicle and neighboring autonomous driving vehicles that satisfy the collision conflict relationship with the target autonomous driving vehicle at the intersection;

[0011] A decision-making module, configured to obtain a predicted acceleration corresponding to the target autonomous driving vehicle at the next time step according to the conflict state graph and a decision neural network model, and the decision neural network model is trained by means of deep reinforcement learning;

[0012] A control module, configured to control the driving of the target autonomous driving vehicle at the next time step according to the predicted acceleration.

[0013] Third aspect, an embodiment of the present invention provides a control method for an autonomous driving vehicle at an intersection, which is applied to an agent corresponding to the intersection deployed in the cloud. The method includes:

[0014] At the current time step, construct conflict state graphs corresponding to the multiple autonomous driving vehicles based on the collision conflict relationship of the multiple autonomous driving vehicles at the intersection. Among them, the conflict state graph corresponding to the target autonomous driving vehicle includes the initial driving state feature vectors corresponding to the target autonomous driving vehicle and multiple neighboring autonomous driving vehicles of the target autonomous driving vehicle, and the multiple neighboring autonomous driving vehicles satisfy the collision conflict relationship with the target autonomous driving vehicle at the intersection;

[0015] According to the conflict state graphs corresponding to the multiple autonomous driving vehicles and a decision neural network model, obtain the predicted accelerations corresponding to the multiple autonomous driving vehicles at the next time step, and the decision neural network model is trained by means of deep reinforcement learning;

[0016] Send the predicted accelerations corresponding to the multiple autonomous driving vehicles at the next time step to the multiple autonomous driving vehicles.

[0017] Fourth aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a communication interface; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor can at least implement the control method for an autonomous driving vehicle at an intersection as described in the first aspect or the third aspect.

[0018] In a fifth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the method for controlling an autonomous vehicle at an intersection as described in the first aspect or the third aspect.

[0019] The solution provided by the embodiment of the present invention is applicable to the scenario of controlling autonomous vehicles near intersections. A single intelligent agent (AI-agent) can be deployed in each autonomous vehicle, and multiple intelligent agents can share the same decision neural network model (which can also be called the Actor network model) trained by deep reinforcement learning. The goal of this decision neural network model is to enable the vehicle to drive through the intersection in the fastest and smoothest possible way, minimizing the number of stops and avoiding collisions.

[0020] Specifically, taking any target autonomous vehicle as an example, first, at the current time step, the intelligent agent in the target autonomous vehicle constructs a conflict state graph corresponding to the target autonomous vehicle based on the collision conflict relationship with surrounding autonomous vehicles at the intersection. The conflict state graph includes multiple autonomous vehicles and the initial driving state feature vectors corresponding to each of the multiple autonomous vehicles. The multiple autonomous vehicles include the target autonomous vehicle and neighboring autonomous vehicles that satisfy the collision conflict relationship with the target autonomous vehicle at the intersection. Then, based on the conflict state graph and the decision neural network model, the predicted acceleration corresponding to the target autonomous vehicle at the next time step is obtained, and the driving of the target autonomous vehicle at the next time step is controlled according to this predicted acceleration.

[0021] The scheduling control of vehicles at intersections is carried out through the decision neural network model. The output of the model is an action in the continuous action space - acceleration, which is a lower-level control instruction compared to traffic light signals and can achieve more refined vehicle operation control. Because traffic light signals only give an indication of whether an autonomous vehicle can pass through the intersection, while acceleration gives the specific speed at which to move. By constructing a conflict relationship graph between the vehicle itself (the target autonomous vehicle) and surrounding vehicles, the driving state information of the vehicle itself and surrounding vehicles with collision conflict relationships can be jointly encoded, so that the decision neural network model can better make accurate decisions based on the intersection environment and avoid collisions at the same time. The combination of multiple intelligent agents (the intelligent agents corresponding to multiple autonomous vehicles near the intersection) with their respective conflict state graphs makes the decision neural network model exhibit the characteristics of distributed deployment and can be applied to any road network scale. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0023] Figure 1 Schematic diagram of the vehicle-end execution environment of a method for controlling an autonomous vehicle at an intersection provided by an embodiment of the present invention;

[0024] Figure 2 Overall architecture diagram of a method for controlling an autonomous vehicle at an intersection provided by an embodiment of the present invention;

[0025] Figure 3 Flowchart of a method for controlling an autonomous vehicle at an intersection provided by an embodiment of the present invention;

[0026] Figure 4 Schematic diagram of an intersection scenario provided by an embodiment of the present invention;

[0027] Figure 5 Schematic diagram of the working process of a graph attention network model provided by an embodiment of the present invention;

[0028] Figure 6 Schematic diagram of the composition of a decision neural network model provided by an embodiment of the present invention;

[0029] Figure 7 Flowchart of a method for controlling an autonomous vehicle at an intersection provided by an embodiment of the present invention;

[0030] Figure 8 Schematic diagram of the composition of an evaluation neural network model provided by an embodiment of the present invention;

[0031] Figure 9 Schematic diagram of the cloud computing environment of a method for controlling an autonomous vehicle at an intersection provided by an embodiment of the present invention;

[0032] Figure 10 Application schematic diagram of a method for controlling an autonomous vehicle at an intersection provided by an embodiment of the present invention;

[0033] Figure 11 Schematic diagram of the structure of a device for controlling an autonomous vehicle at an intersection provided by an embodiment of the present invention;

[0034] Figure 12 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0036] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0037] The following will describe in detail some embodiments of the present invention with reference to the accompanying drawings. Without conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other. In addition, the step timings in the following method embodiments are only examples and are not strictly limited.

[0038] First, the terms or concepts involved in the embodiments of the present invention will be explained:

[0039] Vehicle-road cooperation: mainly refers to the realization of dynamic real-time information interaction between vehicles and between vehicles and roads through the cross-integration of multiple technologies and the use of technologies such as wireless communication and the Internet, and to carry out active vehicle safety control and road coordination management on the basis of the collection and integration of full-time and full-space dynamic traffic information, so as to fully achieve the effective coordination of people, vehicles, and roads, thereby forming a safe, efficient, and environmentally friendly road traffic system.

[0040] V2X: The in-vehicle unit communicates with other devices, including but not limited to communication between in-vehicle units (V2V), communication between the in-vehicle unit and the roadside unit (V2I), communication between the in-vehicle unit and pedestrian devices (V2P), and communication between the in-vehicle unit and the network (V2N).

[0041] Deep Deterministic Policy Gradient (DDPG for short) algorithm: is a reinforcement learning method that can solve continuous control problems. The model structure corresponding to the DDPG algorithm is similar to the neural network model structure of the decision maker (Actor) - evaluator (Critic), that is, the DDPG algorithm involves two models: the decision-making neural network model and the evaluation neural network model.

[0042] Graph Attention Networks (GAT) model: It is a deep learning model for graph data, capable of learning the relationships between nodes and the importance of nodes. In traditional graph convolutional neural networks, the information of the target node is updated by aggregating the information of neighboring nodes. However, this aggregation method ignores the fact that the importance of each neighboring node to the target node is different. The GAT model assigns a weight to each neighboring node by introducing an attention mechanism to reflect its importance to the target node. The main idea of the GAT model is to calculate the weights between nodes through self-attention mechanism and then use these weights to aggregate the features of neighboring nodes.

[0043] Graph Attention Layer (GAL): It is a neural network layer in the GAT model for processing graph data, specifically used to calculate the respective attention weights of neighboring nodes for the target node.

[0044] Multi-Agent Deep Reinforcement Learning (MADRL): It is a method that uses deep reinforcement learning algorithms to solve multi-agent decision-making problems. In traditional reinforcement learning, an agent usually faces a single-agent environment where the decisions of each agent do not affect the states or rewards of other agents. However, in a multi-agent environment, the interactions and dependencies between agents make the problem more complex. In this case, each agent needs to make decisions based on the behaviors of other agents and the feedback from the environment to optimize the overall common goal. MADRL approximates and learns the policy function (or decision function) of agents, as well as the value function or advantage function to evaluate the quality of actions, by introducing deep neural networks. These neural networks can be trained using gradient-based methods, such as Deep Q-Networks or policy gradient methods.

[0045] Markov Decision Process (MDP): It is a mathematical model used to describe sequential decision-making problems with the Markov property. It is a commonly used tool in reinforcement learning for modeling the decision-making process of an agent in an uncertain environment. An MDP consists of five elements: State Space: The set of all possible states that the environment can be in; Action Space: The set of all actions that the agent can take; Transition Probability: Describes the probability that the environment transitions from one state to the next under a combination of state and action; Reward Function: Describes the immediate reward obtained by the agent under different combinations of state and action; Discount Factor: Describes the degree of discount for future rewards, used to balance the importance of immediate rewards and future rewards. The agent selects an action based on the current state in the MDP. The environment transitions to the next state according to the transition probability and gives the corresponding immediate reward. This process continues, and the agent continuously interacts between states and actions. The goal is to maximize the cumulative reward by choosing a good sequence of actions.

[0046] Multi-Layer Perceptron (MLP): It is a common neural network model. An MLP consists of multiple neurons, divided into an input layer, hidden layers, and an output layer. Each neuron receives inputs from the previous layer, performs calculations through weights and activation functions, and then passes the results to the neurons in the next layer. There can be multiple hidden layers between the input layer and the output layer, and each hidden layer has multiple neurons. The role of the hidden layers is to perform non-linear transformations on the inputs to extract higher-level and more abstract features.

[0047] Intelligent Driver Model (IDM): It is a model used to describe the car-following behavior between vehicles. It is a microscopic traffic flow model used to study the stability of traffic flow, the generation of congestion, and the dynamic behavior of traffic flow. IDM is based on the behavioral rules and physical principles of car drivers and simulates the acceleration and deceleration behavior of vehicles during car-following by considering the interaction and relative motion between vehicles. Its main idea is to calculate the acceleration of a vehicle based on the distance, speed, and desired speed between vehicles.

[0048] Autonomous Intersection Management (AIM): It refers to the use of autonomous driving vehicles and intelligent transportation system technologies to achieve autonomous control and optimization of intersections. Traditional intersection control methods usually rely on traffic lights or the command of traffic police, but these methods have some problems, such as high latency, low efficiency, and easy to cause congestion. The goal of AIM is to improve the traffic flow efficiency and safety of intersections by utilizing the communication and coordination between autonomous driving vehicles and using algorithms and strategies in intelligent transportation systems. In the AIM system, autonomous driving vehicles can obtain real-time traffic information and road conditions through communication with traffic infrastructure and other vehicles. Then, based on this information, the control algorithm at the intersection can determine the driving order, speed, and intersection passing time of each vehicle to improve the overall passing efficiency and safety of the intersection. The goal of AIM is to provide a more efficient and safer intersection management method, reduce traffic congestion and accident rates, and lay a foundation for the future development of the traffic system.

[0049] The problem to be solved by the solution provided in the embodiment of the present invention is the control problem of autonomous driving vehicles in the AIM environment. Specifically, by controlling the longitudinal movement of autonomous driving vehicles, that is, acceleration control, the coordination between vehicles is implicitly achieved. For a certain intersection, it is assumed that the passing routes of all autonomous driving vehicles at this intersection are fixed, and the lane changes of all autonomous driving vehicles are controlled by a rule-based model. The speed and position of autonomous driving vehicles will be controlled by the decision neural network model provided in the embodiment of the present invention. The control goal of the decision neural network model is to improve the passing efficiency of the intersection while ensuring collision-free driving of vehicles. In other words, the goal is to make the vehicles drive as fast and smoothly as possible and minimize the number of stops. However, increasing the vehicle speed will also increase the risk of collision, so a more intelligent automatic control method is needed. Therefore, the goal is to achieve a balance between passing efficiency and safety through the speed control of autonomous driving vehicles.

[0050] To achieve this goal, the embodiment of the present invention proposes a control method based on graph structure and deep reinforcement learning to more intelligently control the speed of autonomous driving vehicles near intersections, improve passing efficiency, and ensure safety, which can be applied to multi-intersection scenarios.

[0051] The following introduces and explains the control scheme of autonomous driving vehicles at intersections provided by the embodiment of the present invention.

[0052] Figure 1 Schematic diagram of the vehicle-side execution environment of a control method for autonomous driving vehicles at intersections provided by the embodiment of the present invention, as Figure 1As shown, the vehicle - end execution environment of the automatic driving vehicle control method at this intersection may include Figure 1 several automatic driving vehicles (such as vehicle 1, 2, 3, 4) as shown in

[0053] Briefly understood, the agent in the embodiment of the present invention can be understood as an intelligent entity that can perceive the surrounding road traffic environment, make vehicle control decisions, and execute the decision results.

[0054] Each of the above - mentioned agents will include a decision neural network model, and the decision neural network models included in each agent are the same, that is, multiple agents share one decision neural network model.

[0055] The combination of multiple agents and the local conflict status maps constructed by each agent based on the perception of the surrounding environment enable the decision neural network model to present the characteristics of distributed deployment and can be applied to any road network scale.

[0056] In fact, each automatic driving vehicle will also include functional units such as communication, cameras, and various types of sensors to implement communication functions such as vehicle - to - vehicle communication and vehicle - to - roadside device communication, as well as functions such as perceiving the environment and collecting vehicle - specific data.

[0057] As Figure 1 shown, in fact, there may also be some roadside devices deployed near the intersection. One or more sensors may be provided on the roadside device to observe nearby vehicles, and a communication unit may be provided to achieve communication with the vehicle.

[0058] In the embodiment of the present invention, each agent can make a decision on how to control the vehicle at each time step with a set time step. For example, the time step can be set to values such as 1 second, 500 millimeters, etc. Make a decision on how much acceleration the automatic driving vehicle should move at the next time step at the previous time step.

[0059] In an alternative embodiment, when each automatic driving vehicle determines that it has moved near the intersection based on the set moving route and map, it can execute the solution provided by the embodiment of the present invention.

[0060] First, the overall architecture diagram of the automatic driving vehicle control method at the intersection provided by the embodiment of the present invention will be described in combination with Figure 2 as shown in Figure 2As shown in [description], taking an agent in any autonomous vehicle as an example, the environmental information that the agent can perceive mainly includes the driving state information of other autonomous vehicles and electronic map information. In this agent, there is at least a reinforcement learning model. Optionally, it can also include a car-following model IDM as a physical model. Among them, the reinforcement learning model includes a decision neural network model (schematically shown as an actor network model in the figure), and can also include a graph attention network model (such as a GAT or GAL network model). In an alternative embodiment, the graph attention network model can be included in the decision neural network model or exist independently of the decision neural network model. Figure 2 In [description], in order to more directly illustrate the execution of the graph attention mechanism, it is schematically shown separately from the decision neural network model. Among them, the decision neural network model is obtained through offline training based on the method of deep reinforcement learning. During the training process, an evaluation neural network model ( Figure 2 the critic network model schematically shown in [description]) needs to be used. It should be noted that in the inference stage, the critic network model is not used.

[0061] Combined with Figure 2 , the vehicle control process of an agent in a certain autonomous vehicle is generally described as follows:

[0062] The agent can obtain the driving state information of other vehicles from the environment every second (assuming a time step of 1 second), and then combine its own vehicle driving state and electronic map information to construct a conflict state graph reflecting the environmental information of the vehicle itself. Through the graph attention mechanism, the conflict state graph is embedded into a high-dimensional state vector, and through the decision neural network model, the acceleration that the vehicle needs to execute in the next second is output as a vehicle control instruction. The car-following model is used to judge whether the acceleration output by the decision neural network model is safe according to the current driving states of the vehicle itself and surrounding vehicles. If there is a risk of vehicle collision, the vehicle control instruction output by the decision neural network model is corrected to ensure the safety of the instruction. Finally, the vehicle is controlled to drive with the corrected vehicle control instruction.

[0063] Next, taking the agent deployed in the target autonomous vehicle as an example, the following embodiments are combined to introduce how to implement the control method for autonomous vehicles at intersections. Among them, the target autonomous vehicle can be Figure 1 any vehicle schematically shown in [description].

[0064] Figure 3 is a flowchart of a control method for autonomous vehicles at intersections provided by an embodiment of the present invention. As Figure 3 shown, the method includes the following steps:

[0065] 301. At the current time step, based on the collision conflict relationship with surrounding autonomous vehicles at the intersection, construct a conflict state graph corresponding to the target autonomous vehicle. The conflict state graph contains initial driving state feature vectors corresponding to multiple autonomous vehicles, and the multiple autonomous vehicles include the target autonomous vehicle and neighboring autonomous vehicles that satisfy the collision conflict relationship with the target autonomous vehicle at the intersection.

[0066] 302. According to the conflict state graph and the decision neural network model, obtain the predicted acceleration corresponding to the target autonomous vehicle at the next time step. The decision neural network model is trained by means of deep reinforcement learning.

[0067] 303. Control the driving of the target autonomous vehicle at the next time step according to the predicted acceleration.

[0068] In this embodiment, in the conflict state graph corresponding to the target autonomous vehicle at the current time step, on the one hand, it can reflect the autonomous vehicles that may have a collision conflict with the target autonomous vehicle at the intersection, and on the other hand, it can also reflect the real-time vehicle driving state of each autonomous vehicle in the graph at the current time step. Therefore, it is called a conflict state graph. Moreover, the conflict state graph can be used to describe the local traffic environment from the perspective of the target autonomous vehicle (i.e., the traffic environment near the intersection).

[0069] Generally speaking, if the moving routes of the target autonomous vehicle and another surrounding autonomous vehicle will intersect at the intersection, and the arrival times of the two vehicles at the intersection are relatively close, then it can be considered that there is a risk of collision conflict between the two vehicles at the intersection. Based on this idea, optionally, the process of constructing the conflict state graph corresponding to the target autonomous vehicle at the current time step may include:

[0070] According to the moving routes of the target autonomous vehicle and surrounding autonomous vehicles at the intersection respectively, determine the collision conflict points existing between the target autonomous vehicle and surrounding autonomous vehicles at the intersection;

[0071] Determine the time difference between the target autonomous vehicle reaching the collision conflict point and the surrounding autonomous vehicle reaching the collision conflict point (the time difference between the time when the target autonomous vehicle reaches the collision conflict point and the time when the surrounding autonomous vehicle reaches the collision conflict point);

[0072] If the time difference is less than the set threshold, determine the surrounding autonomous vehicle as the neighboring autonomous vehicle of the target autonomous vehicle, so as to add the initial driving state feature vector corresponding to the neighboring autonomous vehicle to the conflict state graph.

[0073] As described above, to construct the conflict state graph corresponding to the target autonomous vehicle, it is necessary to first determine the neighborhood set of the target autonomous vehicle, which includes the neighboring autonomous vehicles of the target autonomous vehicle. In addition, as described above, the autonomous vehicles located around the target autonomous vehicle are not necessarily its neighboring autonomous vehicles, and the condition that the above time difference is less than the set threshold needs to be met.

[0074] The conflict state graph of the target autonomous vehicle can be used to describe the local traffic environment of the target autonomous vehicle, and those vehicles that are too far away from the target autonomous vehicle can be safely ignored.

[0075] For the target autonomous vehicle i, its neighborhood set can be defined as N i ={i∪j|TTC ij <TTC t}, where TTC represents the time to collision, and TTC ij represents the time difference between the target autonomous vehicle i and the autonomous vehicle j to reach the collision conflict point, which can be estimated based on the moving routes, current speeds and accelerations of the two vehicles. TTC t is a hyperparameter of the TTC threshold, which is used to determine the scale of the graph. A larger TTC t results in a larger neighborhood set, which will increase the information acquisition cost and calculation time. A smaller TTC t increases the risk of collision. Based on the above definition of the neighborhood set, the conflict state graph G i can be defined as: G i =(N i ,E i ), E i ={e(ij)}, where j represents the autonomous vehicle j located in N i . N i is the node set of the conflict state graph, and E i is the edge set of the conflict state graph.

[0076] As introduced above, to determine the conflict state graph of the target autonomous vehicle, it is first necessary to determine the collision conflict points existing between the target autonomous vehicle and the surrounding autonomous vehicles at the intersection. Specifically, the collision conflict points existing between the target autonomous vehicle and the surrounding autonomous vehicles at the intersection can be determined according to the moving routes of the target autonomous vehicle and the surrounding autonomous vehicles at the intersection respectively. Simply put, the intersection point where the moving routes of the two vehicles exist at the intersection is the collision conflict point.

[0077] For a target autonomous vehicle, it can know its moving route based on navigation information and know that the moving route will pass through an intersection based on map information. The target autonomous vehicle can know the moving routes of surrounding autonomous vehicles at the intersection through the active sharing of surrounding autonomous vehicles or the notification of roadside units. Thus, the collision conflict points between the target autonomous vehicle and surrounding autonomous vehicles at the intersection can be determined. Furthermore, the time for itself and surrounding autonomous vehicles to reach the collision conflict points can be determined through a set TTC estimation algorithm, and whether the surrounding autonomous vehicle belongs to the neighboring autonomous vehicles that can be added to its neighborhood set can be determined based on the time difference.

[0078] For ease of understanding, it is exemplified and described in combination with Figure 4 the schematic intersection scenario shown. In Figure 4 it, assume that the target autonomous vehicle is Vehicle 1, and the surrounding autonomous vehicles include Vehicle 2, Vehicle 3, and Vehicle 4. Figure 4 The situation of two intersections is schematically shown in Figure 4 where the arrowed lines represent the moving routes of the corresponding vehicles, and Figure 4 the roads shown are in a two-lane situation.

[0079] For Vehicle 1, according to the moving routes of each vehicle at the intersection schematically shown in the figure above, three collision conflict points shown in Figure 4 can be determined: the collision conflict point P12 with Vehicle 2, the collision conflict point P13 with Vehicle 3, and the collision conflict point P14 with Vehicle 4. And assume that it is finally determined that the time differences between Vehicle 1 and the other three vehicles reaching the corresponding collision conflict points are all less than the set threshold, then it is determined that Vehicle 2, Vehicle 3, and Vehicle 4 are all its neighboring autonomous vehicles, and thus the following conflict state graph of Vehicle 1 is formed:

[0080] G1=(N1,E1), N1={1,2,3,4}, indicating that Vehicle 1-4 form a node set, and E i ={e(12),e(13),e(14)}, indicating the edge set.

[0081] In this example, the observation range of Vehicle 1 covers 2 consecutive intersections (that is, it is found that the vehicles near these two intersections all meet the neighborhood conditions), which enables implicit vehicle coordination between closely adjacent intersections. Therefore, the vehicle can pass through multiple intersections without stopping.

[0082] The process of constructing the conflict state graph corresponding to the target autonomous vehicle at the current time step is introduced above. However, in this process, mainly the neighborhood set of the target autonomous vehicle, that is, each neighboring autonomous vehicle, is determined. It is also necessary to determine the real-time driving state information of each autonomous vehicle in the conflict state graph at the current time step.

[0083] First, define a feature vector h to represent the real-time driving state of an autonomous vehicle: h = {len, px, py, v, acc, heading, lane, d}, where len is the length of the vehicle, (px, py) is the position of the vehicle in a two-dimensional coordinate system, (v, acc, heading) are the current speed, acceleration, and heading angle of the vehicle respectively, lane is the lane number where the vehicle is currently located, and d is the distance between the vehicle and the downstream intersection. These features help the autonomous vehicle better understand its position, driving state, and distance from the intersection.

[0084] It can be understood that the information of (len, v, acc, heading) can be obtained based on the inherent attribute information of the autonomous vehicle and relevant detection devices inside the vehicle, and (px, py, lane, d) can be determined based on positioning technology in combination with an electronic map.

[0085] For a target autonomous vehicle, it can determine the driving state feature vector corresponding to the current time step of the vehicle itself through the above method (to distinguish from the feature vector encoded by the graph attention mechanism in the following text, this driving state feature vector is called the initial driving state feature vector). For the initial driving state feature vectors corresponding to other neighboring autonomous vehicles in the neighborhood set, they can be obtained from neighboring autonomous vehicles through vehicle-to-vehicle communication, or obtained from roadside devices based on the observation of the real-time driving states of each autonomous vehicle.

[0086] So far, a conflict state graph corresponding to the target autonomous vehicle has been obtained, which contains the initial driving state feature vectors corresponding to multiple autonomous vehicles respectively.

[0087] After that, according to this conflict state graph and the decision neural network model, the predicted acceleration corresponding to the target autonomous vehicle at the next time step is obtained.

[0088] In an optional embodiment, the conflict state graph can be encoded by a graph attention network model to obtain the target driving state feature vector corresponding to the target autonomous vehicle, and then the target driving state feature vector corresponding to the target autonomous vehicle is input into the decision neural network model to obtain the predicted acceleration corresponding to the target autonomous vehicle at the next time step.

[0089] When specifically implemented, if the above graph attention network model is inside the decision neural network model, it can be understood that the input of the target driving state feature vector into the decision neural network model means inputting it into other neural network layers connected after the graph attention network model in the decision neural network model.

[0090] In another alternative embodiment, the conflict state graph can be encoded by a graph attention network model to obtain the target driving state feature vector corresponding to the target autonomous vehicle. Then, the concatenation result of the target driving state feature vector corresponding to the target autonomous vehicle and the initial driving state feature vector is input into the decision neural network model to obtain the predicted acceleration corresponding to the target autonomous vehicle at the next time step. Then, the target autonomous vehicle is controlled to drive at the next time step according to the predicted acceleration.

[0091] It can be understood that after the predicted acceleration is executed, the intelligent body in the target autonomous vehicle will execute the process of steps 301-303 again to complete the prediction of the acceleration at subsequent other time steps.

[0092] Since the decision neural network model is trained with the control objective of improving the traffic efficiency at intersections while ensuring collision-free driving of vehicles, by constructing the conflict relationship graph between the vehicle itself (the target autonomous vehicle) and surrounding vehicles, the driving state information of the vehicle itself and the surrounding vehicles with collision conflict relationships can be jointly encoded, so that the decision neural network model can make more accurate decisions based on the intersection environment and avoid collisions. Moreover, by using the decision neural network model to perform the scheduling control of vehicles at intersections, the output of the model is an action - acceleration in the continuous action space, which is a lower-level control instruction compared to traffic light signals and can achieve more refined vehicle operation control. Because traffic light signals only give an indication of whether the autonomous vehicle can pass through the intersection, while acceleration gives the specific speed at which to move.

[0093] The following combines Figure 5 to specifically illustrate the working process of the graph attention network model.

[0094] In Figure 5 is shown the encoding process of the conflict state graph of vehicle 1 determined based on Figure 4 where h1 - h4 represent the initial driving state feature vectors of vehicles 1 - 4.

[0095] In practical applications, the graph attention mechanism introduced by the GAT model can be used to encode the conflict state graph of vehicle 1 into a higher-dimensional feature vector h1'. The graph attention layer (GAL) of GAT is selected as the encoding layer for the conflict state graph because it adapts to the flexibility of the graph structure in dynamic traffic control problems.

[0096] Figure 5An example of the working principle of the GAL is given. In the first step, the feature vectors of each initial driving state are enhanced through a linear transformation of the weight matrix w and converted into feature vectors of a higher dimension. In the second step, for vehicle 1 and each vehicle in its neighborhood, the normalized attention coefficients are obtained through a set formula: a 11 、a 12 、a 13 、a 14 。Among them, a ij reflects the importance of the features of vehicle j to vehicle i. In the third step, the output feature vector of one head (head) is obtained through a set formula: h1”, and finally, the multi-head attention is applied by connecting the self-attention mechanisms of K heads to stabilize the learning process, and the target driving state feature vector of vehicle 1 finally output: h1’.

[0097] The relevant formulas in the above graph attention mechanism can refer to the existing related technologies and will not be elaborated in the embodiments of the present invention.

[0098] Next, the working process of the decision neural network model will be described in conjunction with Figure 6 .

[0099] The decision neural network model is responsible for mapping the observations of the vehicle to an action and following the decision μ, which is parameterized by μ θ . That is, the decision neural network model is the decision μ, and its model parameters are μ θ .

[0100] In the embodiments of the present invention, the observation of the vehicle is the conflict state graph corresponding to the target autonomous vehicle at the current time step, and the action is the predicted acceleration of the target autonomous vehicle at the next time step. The predicted acceleration corresponds to upper and lower threshold values.

[0101] Figure 6 schematically shows the composition structure of a decision neural network model, which includes three modules: GAL, multi-layer perceptron (MLP) and a scaling function (f_scaling).

[0102] As described above, GAL can be set independently of the decision neural network model or can be included in the decision neural network model. In this embodiment, still taking Figure 4 as an example of the conflict state graph of vehicle 1 shown, and assuming that GAL is included in the decision neural network model and is used to output h1’ in the above text. It can be understood that if GAL is not included in the decision neural network model, the above h1’ output by GAL can be directly input into the decision neural network model.

[0103] In practical applications, the GAL module consists of two GAL layers and is used to encode the conflict status graph of vehicle 1 into a higher-dimensional target driving status feature vector h1'. The MLP module concatenates the encoded target driving status feature vector h1' with the initial driving status feature vector h1 and maps it to a normalized acceleration a1' whose value range is [-1, 1]. The MLP can consist of three fully connected (FC) layers. After the third layer, layer normalization operations and the tanh activation function can be performed to generate a1'. Finally, the scaling function converts the normalized acceleration a1' into a real executable acceleration a1 considering the acceleration boundary and speed limit. The specific form of the scaling function is not limited in the embodiments of the present invention.

[0104] Figure 7 The flowchart of an intersection autonomous vehicle control method provided by an embodiment of the present invention is as Figure 7 shown, and the method includes the following steps:

[0105] 701. At the current time step, based on the collision conflict relationship with surrounding autonomous vehicles at the intersection, construct a conflict status graph corresponding to the target autonomous vehicle. The conflict status graph contains initial driving status feature vectors corresponding to multiple autonomous vehicles, and the multiple autonomous vehicles include the target autonomous vehicle and neighboring autonomous vehicles that satisfy the collision conflict relationship with the target autonomous vehicle at the intersection.

[0106] 702. According to the conflict status graph and the decision neural network model, obtain the predicted acceleration corresponding to the target autonomous vehicle at the next time step.

[0107] 703. Use the car-following model to determine the safe acceleration corresponding to the target autonomous vehicle at the next time step.

[0108] 704. Determine the target acceleration corresponding to the target autonomous vehicle at the next time step according to the predicted acceleration and the safe acceleration, and control the driving of the target autonomous vehicle at the next time step according to the target acceleration.

[0109] In this embodiment, by combining the decision neural network model with the physical model - the car-following model, the anti-collision safety can be further improved.

[0110] Based on the introduction of the foregoing embodiments, the predicted acceleration corresponding to the target autonomous vehicle at the next time step can be obtained. However, this predicted acceleration may not be safe, that is, driving at this predicted acceleration may cause the target autonomous vehicle to collide with other autonomous vehicles at the intersection. Therefore, a following vehicle model is introduced to determine the safe acceleration corresponding to the target autonomous vehicle at the next time step, and the target acceleration that the target autonomous vehicle needs to execute at the next time step is finally determined by combining this safe acceleration.

[0111] Specifically, if the predicted acceleration is much greater than the safe acceleration (such as greater than a set multiple), it is considered that the predicted acceleration is not safe, and the safe acceleration is used as the target acceleration. Otherwise, the target acceleration is determined to be the predicted acceleration.

[0112] In addition, optionally, if the predicted acceleration is much greater than the safe acceleration, the difference between the predicted acceleration and the safe acceleration can also be used as one item in the reward value when optimizing the training of the decision neural network model, so as to optimize the decision neural network model.

[0113] Optionally, using the following vehicle model to determine the safe acceleration corresponding to the target autonomous vehicle at the next time step can be as follows: First, determine at least one collision conflict group corresponding to the target autonomous vehicle at the intersection. Then, based on the at least one collision conflict group, use the following vehicle model to determine the safe acceleration corresponding to the target autonomous vehicle at the next time step.

[0114] Among them, different collision conflict groups indicate that the target autonomous vehicle has collision conflicts with different neighboring autonomous vehicles at different collision conflict points at the intersection.

[0115] Among them, determining at least one collision conflict group corresponding to the target autonomous vehicle at the intersection specifically includes: determining the approach lane segment within a preset distance from the intersection as the control area; determining the virtual lane corresponding to the target collision conflict point, where the target collision conflict point is any one of the above different collision conflict points; determining the autonomous vehicles driving in the control area and towards the target collision conflict point to form the target collision conflict group corresponding to the target collision conflict point, and associating the target collision conflict group with the virtual lane corresponding to the target collision conflict point.

[0116] Among them, based on at least one collision conflict group, using the following vehicle model to determine the safe acceleration corresponding to the target autonomous vehicle at the next time step can specifically be: using the following vehicle model to respectively determine the safe acceleration corresponding to the target autonomous vehicle under at least one collision conflict group, and determining the safe acceleration corresponding to the target autonomous vehicle at the next time step according to the safe acceleration corresponding to the target autonomous vehicle under at least one collision conflict group.

[0117] In the embodiments of the present invention, a car-following model, a physical model, is introduced to obtain a safe acceleration that ensures no collision. The original car-following model is used for car-following control within the same lane, and now it is extended to handle the scenario of passing through intersections.

[0118] First, define the control area: It is the lane segment of the approach lane within a set distance range from the center of the intersection. Among them, the approach lane refers to the lane that drives into the intersection.

[0119] Secondly, determine the virtual lane (vlane): As mentioned above, the target autonomous vehicle can determine the collision conflict points with each surrounding autonomous vehicle at the intersection based on its own movement route at the intersection and the movement routes of the surrounding autonomous vehicles at the intersection (reference can be made to Figure 4 ). A virtual lane can be defined for each collision conflict point within the intersection.

[0120] After that, map the autonomous vehicles driving towards the collision conflict points within the control area to the corresponding virtual lanes. Apply the car-following model to each virtual lane to determine the safe accelerations of each autonomous vehicle on the virtual lane.

[0121] For easy understanding, an example is given in combination with the situation shown in Figure 4 . Referring to Figure 4 , there are collision conflict points P12 between vehicle 1 and vehicle 2, P13 between vehicle 1 and vehicle 3, and P14 between vehicle 1 and vehicle 4 at the intersection. Then, these three collision conflict points will correspond to three virtual lanes: vlane1, vlane2, vlane3, and the mapping results of the vehicles on each virtual lane are as follows:

[0122] vlane1: vehicle 1, vehicle 2;

[0123] vlane2: vehicle 1, vehicle 3;

[0124] vlane3: vehicle 1, vehicle 4.

[0125] That is to say, three groups of collision conflict groups corresponding to the three collision conflict points will be formed.

[0126] Apply the car-following model to each group of collision conflict groups on the above three virtual lanes respectively, and the safe accelerations corresponding to vehicle 1 and vehicle 2 under vlane1, vehicle 1 and vehicle 3 under vlane2, and vehicle 1 and vehicle 4 under vlane2 can be obtained.

[0127] After that, for the target autonomous vehicle, i.e., vehicle 1, the minimum or average value of the corresponding safety accelerations of vehicle 1 under three virtual lanes can be taken as the safety acceleration corresponding to vehicle 1 at the next time step finally determined.

[0128] In practical applications, taking the vehicle following model applied to vehicle 1 and vehicle 2 under the above-mentioned vlane1 as an example, parameters such as the relative distance and relative speed between vehicle 1 and vehicle 2, as well as the desired speed, safety time interval, maximum acceleration, and comfort deceleration of the following vehicle (assumed to be vehicle 1) can be input into the vehicle following model. The vehicle following model will output the safety accelerations of vehicle 1 and vehicle 2, and such safety accelerations will prevent the following vehicle from colliding with the preceding vehicle.

[0129] In the above example, the order of vehicle 1 and vehicle 2 on vlane1 can be determined according to the time sequence of their respective arrivals at the corresponding collision conflict points. The specific working process of the vehicle following model can be implemented with reference to existing related technologies and will not be elaborated here.

[0130] In summary, by combining the conflict status diagram, the decision neural network model, and the vehicle following model, an acceleration decision result that ensures traffic efficiency and safety can be obtained.

[0131] The following briefly introduces the process of deep reinforcement learning training for the decision neural network model by combining the evaluation neural network model (critic model).

[0132] Generally speaking, the decision neural network model is used to output the action - acceleration that needs to be executed at the next time step based on the conflict status diagram of the autonomous vehicle at the current time step. The evaluation neural network model is used to output the value function value (i.e., q value) that evaluates the quality of the decision result given by the decision neural network model based on the conflict status diagram of the autonomous vehicle at the current time step and the above-mentioned acceleration.

[0133] The evaluation neural network model and the decision neural network model are obtained through offline training. For the application scenario of autonomous vehicle control in the intersection scenario, experimental data of the intersection can be generated through a certain traffic simulation tool. Based on this experimental data and the evaluation neural network model, deep reinforcement learning training is performed on the decision neural network model. Among them, the experimental data includes the driving state feature vectors of multiple virtual autonomous vehicles at the first time step, the accelerations that multiple virtual autonomous vehicles need to execute at the next second time step determined by the decision neural network model to be trained, and the reward values after executing the accelerations.

[0134] Among them, the virtual autonomous vehicle refers to an autonomous vehicle added in the intersection traffic scenario built by a traffic simulation tool.

[0135] In the traffic scenario of this intersection, a simulation tool is used to observe the driving states of each virtual autonomous driving vehicle at each time step in the intersection traffic scenario, and to call the decision neural network model to be trained to execute the following logic: determining the accelerations that each virtual autonomous driving vehicle needs to execute at the second time step based on the driving state feature vectors of each virtual autonomous driving vehicle at the first time step.

[0136] It can be understood that, similar to the working process of the decision neural network model in the inference stage in the foregoing embodiment, in the training stage, for any virtual autonomous driving vehicle, it is also necessary to generate its corresponding conflict state graph based on the driving state feature vectors of the virtual autonomous driving vehicles within its neighborhood set for training.

[0137] The following combines Figure 8 a schematic composition structure of an evaluation neural network model to exemplify and illustrate the training process.

[0138] The purpose of the evaluation neural network model in the training stage is to train the decision neural network model and it is not required in the execution stage. Generally speaking, its input is the state and action, and the output is the value function value (i.e., the q value).

[0139] The evaluation neural network model may include four modules: GAL, MLP1, MLP2, and MLP3. Still taking the foregoing Figure 4 schematic situation as an example, assuming that the intersection scenario during the training process is the Figure 4 situation shown in. Taking vehicle 1 as an example, at the current time step, generate the conflict state graph shown in the foregoing embodiment, and obtain the acceleration a1 corresponding to vehicle 1 based on the processing of the decision neural network model. Similarly, assuming that the conflict state graphs corresponding to vehicle 2, vehicle 3, and vehicle 4 at the current time step are respectively input into the decision neural network model, the acceleration a2 corresponding to vehicle 2, the acceleration a3 corresponding to vehicle 3, and the acceleration a4 corresponding to vehicle 4 are obtained.

[0140] Based on the acceleration decision results of the above four vehicles at the next time step, a Figure 8 state-action graph corresponding to vehicle 1 as shown in can be formed. Among them, each node in the state-action graph is the concatenation result of the driving state feature vector hi of the corresponding vehicle and the acceleration ai.

[0141] The state-action graph is input into GAL to encode it into a hidden higher-dimensional feature vector H1', and the encoded driving state feature vector H1' corresponding to vehicle 1 is concatenated with the initial driving state feature vector h1 and then input into MLP1. MLP1 has no normalization operation and uses sigmoid as the activation function in the output layer.

[0142] The input vectors of MLP2 and MLP3 are {v1, d1, a1}, representing the speed of vehicle 1 at the current time step, the distance to the next intersection, and the acceleration of vehicle 1 corresponding to the next time step output by the decision neural network model.

[0143] As Figure 8 shown, the evaluation function value q1 corresponding to vehicle 1 is q1 = λ d q d + λ e q e + λ c q c , where λ d , λ e , and λ c are weight parameters used to weigh the moving forward value q d , the value of leaving the intersection q e , and the collision value q c .

[0144] In practical applications, MLP3 can include, for example, 3 fully connected (FC) layers, and its output layer has no activation function. MLP2 can include, for example, 2 FC layers, and its output layer activation function is sigmoid.

[0145] For the current time step, in the above description, only vehicle 1 is taken as an example to illustrate the calculation process of q1 corresponding to vehicle 1. Similar processes are performed for vehicle 2, vehicle 3, and vehicle 4, and q2, q3, and q4 corresponding to vehicle 2, vehicle 3, and vehicle 4 can be obtained respectively. Finally, the average value of q1 - q4 is calculated to obtain the average value q as the value function value of the acceleration decision results (a1, a2, a3, a4) of each vehicle at the current time step.

[0146] After obtaining the average value q, generally speaking, the gradient of the evaluation neural network model is updated by minimizing the TD-error of the average value q, and the gradient of the decision neural network model is updated by maximizing the average value q to train the evaluation neural network model and the decision neural network model.

[0147] In addition, during the calculation process of the value function, a reward value is actually used. From the MDP, the initial experimental data generated by the traffic simulation tool during the training process can be represented as the following sequence:

[0148] S0, A0, R1, S1, A1, R2, …

[0149] Among them, S t represents the state, A t represents the action, and R t is the reward value. Taking S0, A0, R1, S1 as an example, it means that from state S0 by performing action AO The reward value R1 given after jumping to the next state S1.

[0150] Corresponding to the control scenario of the intersection autonomous driving vehicle in the embodiment of the present invention, S t is the state conflict graph of the vehicle (or the driving state feature vector of each vehicle), A t is the acceleration of the vehicle determined by the decision.

[0151] In the embodiment of the present invention, optionally, the reward value can be defined as:

[0152] R(S t , A t , S t+1 ) = λ d d t + λ c c t + λ s stop t , where λ d , λ e , λ s are weight parameters. d t represents the average distance traveled by all vehicles (for vehicle i, it refers to all vehicles in its neighborhood set) in time step t. If vehicle i leaves the intersection at the end of this time step, its travel distance in this time step is set to the default value. c t and stop t are respectively the total number of collisions and the total number of stops that occur in the intersection in time step t.

[0153] The intersection autonomous driving vehicle control method provided by the embodiment of the present invention can be executed not only by an agent deployed on the autonomous driving vehicle, but also by deploying an agent in the cloud. After the cloud agent outputs the vehicle control instruction for each vehicle according to the environmental state of each vehicle, it is sent to the corresponding vehicle.

[0154] The cloud service provider maintains several cloud servers in the cloud - called computing nodes. As Figure 9 shown in the cloud computing environment, it can include several ( Figure 9 901-1, 901-2,... shown in Figure 9Services A, B, C, and D as shown in the figure. The way to provide these services in the cloud computing environment can be to provide a service interface 902 externally, and the client device calls this service interface 902 to use the corresponding service. The service interface 902 includes forms such as a Software Development Kit (SDK) and an Application Programming Interface (API).

[0155] The above services are deployed according to various virtualization technologies supported by the cloud computing environment, such as virtualization technologies based on virtual machines and containers. Taking the virtualization technology based on containers as an example, several containers corresponding to a service can be assembled into a container group (pod). For example Figure 9 Service B as shown in the figure can be configured with one or more pods, and each pod can include a proxy and one or more containers. One or more containers in the pod are used to process requests related to one or more corresponding functions of the service, and the proxy in the pod is used to control network functions related to the service, such as routing, load balancing, etc.

[0156] During the operation process, when executing requests from the client device, it may be necessary to call one or more services in the cloud computing environment, and when executing one or more functions of a service, it may be necessary to call one or more functions of another service. As Figure 9 shown, after Service A receives a request sent by the client device, it can call Service B, and Service B can request Service D to execute one or more functions.

[0157] Under the above cloud computing environment, an embodiment of the present invention provides an application schematic diagram of a method for controlling an autonomous vehicle at an intersection as shown in Figure 10 the figure.

[0158] In Figure 10 the figure, an intersection vehicle control service and a corresponding service interface are provided in the cloud computing environment. The in-vehicle terminal device of the autonomous vehicle calls this service interface and can report its own driving status information to the intersection vehicle control service in real time.

[0159] Among them, an agent and related models (such as the decision neural network model, graph attention network model, and following vehicle model in the above text) are implemented in the intersection vehicle control service, and this service can call the map service.

[0160] Moreover, in an optional embodiment, in a cloud computing environment, the above intersection vehicle control service can be deployed distributively. For example, an instance of the intersection vehicle control service is responsible for controlling vehicle passage at several consecutive intersections. At this time, each intersection vehicle control service instance has an agent, and different agents share the same decision neural network model.

[0161] The agent in the intersection vehicle control service can specifically perform the following steps:

[0162] At the current time step, based on the collision conflict relationships of multiple autonomous vehicles at the intersection, construct conflict state diagrams corresponding to each of the multiple autonomous vehicles. Among them, the conflict state diagram corresponding to the target autonomous vehicle includes the target autonomous vehicle and the initial driving state feature vectors corresponding to each of the multiple neighboring autonomous vehicles of the target autonomous vehicle. The multiple neighboring autonomous vehicles and the target autonomous vehicle satisfy collision conflict relationships at the intersection.

[0163] According to the conflict state diagrams corresponding to each of the multiple autonomous vehicles and the decision neural network model, obtain the predicted accelerations corresponding to each of the multiple autonomous vehicles at the next time step.

[0164] Send the predicted accelerations corresponding to the multiple autonomous vehicles at the next time step to the multiple autonomous vehicles.

[0165] In addition, the agent in the intersection vehicle control service can also determine the safe accelerations corresponding to the multiple autonomous vehicles at the next time step based on a following model, correct the predicted accelerations based on the safe accelerations, and send the corrected accelerations to the multiple autonomous vehicles correspondingly so that they drive at the corresponding accelerations.

[0166] The following will describe in detail the autonomous vehicle control device at the intersection of one or more embodiments of the present invention. Those skilled in the art can understand that these devices can all be configured by using commercially available hardware components through the steps taught by this solution.

[0167] Figure 11 The structural schematic diagram of an autonomous vehicle control device at the intersection provided for the embodiments of the present invention is as Figure 11 shown. The device includes: a composition module 11, a decision module 12, and a control module 13.

[0168] A composition module 11, configured to construct a conflict status graph corresponding to the target autonomous vehicle at the current time step based on the collision conflict relationship with surrounding autonomous vehicles at an intersection. The conflict status graph includes initial driving state feature vectors corresponding to multiple autonomous vehicles, and the multiple autonomous vehicles include the target autonomous vehicle and neighboring autonomous vehicles that satisfy the collision conflict relationship with the target autonomous vehicle at the intersection.

[0169] A decision-making module 12, configured to obtain a predicted acceleration corresponding to the next time step of the target autonomous vehicle at the current time step according to the conflict status graph and a decision neural network model, where the decision neural network model is trained by means of deep reinforcement learning.

[0170] A control module 13, configured to control the driving of the target autonomous vehicle at the next time step according to the predicted acceleration.

[0171] Optionally, the apparatus further includes: a determination module, configured to determine a safety acceleration corresponding to the next time step of the target autonomous vehicle by using a car-following model. Thus, optionally, the control module 13 is further configured to: determine a target acceleration corresponding to the next time step of the target autonomous vehicle according to the predicted acceleration and the safety acceleration, and control the driving of the target autonomous vehicle at the next time step according to the target acceleration.

[0172] Optionally, the determination module is specifically configured to: determine at least one collision conflict group corresponding to the target autonomous vehicle at the intersection, where different collision conflict groups indicate that the target autonomous vehicle has collision conflicts with different neighboring autonomous vehicles at different collision conflict points at the intersection; and determine the safety acceleration corresponding to the next time step of the target autonomous vehicle based on the at least one collision conflict group by using a car-following model.

[0173] Optionally, the determination module is specifically configured to: use a car-following model to respectively determine the safety accelerations corresponding to the target autonomous vehicle under the at least one collision conflict group; and determine the safety acceleration corresponding to the next time step of the target autonomous vehicle according to the safety accelerations corresponding to the target autonomous vehicle under the at least one collision conflict group.

[0174] Optionally, the determination module is specifically configured to: determine an approach lane segment within a preset distance from the intersection as a control area; determine a virtual lane corresponding to a target collision conflict point, where the target collision conflict point is any one of the different collision conflict points; determine an autonomous vehicle traveling within the control area and toward the target collision conflict point, form a target collision conflict group corresponding to the target collision conflict point, and associate the target collision conflict group with the virtual lane corresponding to the target collision conflict point.

[0175] Optionally, the decision-making module 12 is specifically configured to: encode the conflict status graph through a graph attention network model to obtain a target driving status feature vector corresponding to the target autonomous vehicle; input the concatenation result of the target driving status feature vector corresponding to the target autonomous vehicle and the initial driving status feature vector into the decision neural network model to obtain a predicted acceleration corresponding to the target autonomous vehicle at the next time step.

[0176] Optionally, the graph construction module 11 is specifically configured to: determine a collision conflict point existing between the target autonomous vehicle and the surrounding autonomous vehicles at the intersection according to the moving routes of the target autonomous vehicle and the surrounding autonomous vehicles at the intersection; determine the time difference between the target autonomous vehicle reaching the collision conflict point and the surrounding autonomous vehicles reaching the collision conflict point; if the time difference is less than a set threshold, determine the surrounding autonomous vehicles as the neighboring autonomous vehicles of the target autonomous vehicle, and add the initial driving status feature vectors corresponding to the neighboring autonomous vehicles to the conflict status graph.

[0177] Optionally, the device further includes: a training module, configured to generate experimental data of the intersection based on a traffic simulation tool, where the experimental data includes driving status feature vectors of multiple virtual autonomous vehicles at a first time step, accelerations that need to be executed by the multiple virtual autonomous vehicles at a second time step determined through a decision neural network model to be trained, and reward values after executing the accelerations, and the second time step is the next time step of the first time step; perform deep reinforcement learning training on the decision neural network model based on the experimental data and an evaluation neural network model.

[0178] Figure 11 The shown device can execute the steps provided in the foregoing embodiments. For the detailed execution process and technical effects, refer to the descriptions in the foregoing embodiments, which will not be elaborated herein.

[0179] In a possible design, the above Figure 11 The structure of the autonomous vehicle control device at the shown intersection can be implemented as an electronic device. As Figure 12As shown, the electronic device may include: a processor 21, a memory 22, and a communication interface 23. Among them, executable code is stored on the memory 22. When the executable code is executed by the processor 21, the processor 21 can at least implement the method for controlling an autonomous vehicle at an intersection provided in the foregoing embodiments.

[0180] In an alternative embodiment, the electronic device for executing the method for controlling an autonomous vehicle at an intersection provided in the embodiments of the present invention may be any user terminal, such as a mobile phone, a laptop computer, a PC, or may also be an Extended Reality (XR) device. XR is a collective term for various forms such as virtual reality and augmented reality.

[0181] In addition, embodiments of the present invention provide a non-transitory machine-readable storage medium. Executable code is stored on the non-transitory machine-readable storage medium. When the executable code is executed by a processor of an electronic device, the processor can at least implement the method for controlling an autonomous vehicle at an intersection provided in the foregoing embodiments.

[0182] The device embodiments described above are merely illustrative. The network elements described as separate components may or may not be physically separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0183] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, can also be implemented by a combination of hardware and software. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A control method for autonomous vehicles at intersections, characterized in that, Applied to an agent deployed in a target autonomous vehicle, including: At the current time step, based on the collision conflict relationship with surrounding autonomous vehicles at an intersection, construct a conflict state graph corresponding to the target autonomous vehicle, where the conflict state graph contains initial driving state feature vectors corresponding to multiple autonomous vehicles respectively, and the multiple autonomous vehicles include the target autonomous vehicle and neighboring autonomous vehicles that satisfy the collision conflict relationship with the target autonomous vehicle at the intersection; According to the conflict state graph and the decision neural network model, obtain the predicted acceleration corresponding to the next time step of the target autonomous vehicle at the current time step, and the decision neural network model is trained by means of deep reinforcement learning; Control the driving of the target autonomous vehicle at the next time step according to the predicted acceleration.

2. The method according to claim 1, wherein The controlling the driving of the target autonomous vehicle at the next time step according to the predicted acceleration includes: Use a car-following model to determine the safe acceleration corresponding to the target autonomous vehicle at the next time step; Determine the target acceleration corresponding to the target autonomous vehicle at the next time step according to the predicted acceleration and the safe acceleration; Control the driving of the target autonomous vehicle at the next time step according to the target acceleration.

3. The method according to claim 2, wherein The using a car-following model to determine the safe acceleration corresponding to the target autonomous vehicle at the next time step includes: Determine at least one collision conflict group corresponding to the target autonomous vehicle at the intersection, and different collision conflict groups indicate that the target autonomous vehicle has collision conflicts with different neighboring autonomous vehicles at different collision conflict points at the intersection; Based on the at least one collision conflict group, use the car-following model to determine the safe acceleration corresponding to the target autonomous vehicle at the next time step.

4. The method according to claim 3, wherein The based on the at least one collision conflict group, using a car-following model to determine the safe acceleration corresponding to the target autonomous vehicle at the next time step includes: Use the car-following model to respectively determine the safe accelerations corresponding to the target autonomous vehicle under the at least one collision conflict group; Determine the safe acceleration corresponding to the target autonomous vehicle at the next time step according to the safe accelerations corresponding to the target autonomous vehicle under the at least one collision conflict group.

5. The method according to claim 3, characterized in that, The determining at least one collision conflict group corresponding to the target autonomous vehicle at the intersection includes: Determine the approach lane segment within a preset distance from the intersection as the control area; Determine the virtual lane corresponding to the target collision conflict point, where the target collision conflict point is any one of the different collision conflict points; Determine the autonomous vehicles driving within the control area and towards the target collision conflict point to form the target collision conflict group corresponding to the target collision conflict point, and associate the target collision conflict group with the virtual lane corresponding to the target collision conflict point.

6. The method according to claim 1, wherein Obtaining the predicted acceleration corresponding to the target autonomous driving vehicle at the next time step according to the conflict status graph and the decision neural network model includes: Encoding the conflict status graph through a graph attention network model to obtain a target driving state feature vector corresponding to the target autonomous driving vehicle; Inputting the concatenation result of the target driving state feature vector corresponding to the target autonomous driving vehicle and the initial driving state feature vector into the decision neural network model to obtain the predicted acceleration corresponding to the target autonomous driving vehicle at the next time step.

7. The method according to claim 1, characterized in that, Constructing the conflict status graph corresponding to the target autonomous driving vehicle based on the collision conflict relationship with surrounding autonomous driving vehicles at the intersection, including: Determining the collision conflict points existing between the target autonomous driving vehicle and the surrounding autonomous driving vehicles at the intersection according to the moving routes of the target autonomous driving vehicle and the surrounding autonomous driving vehicles at the intersection respectively; Determining the time difference between the target autonomous driving vehicle reaching the collision conflict point and the surrounding autonomous driving vehicle reaching the collision conflict point; If the time difference is less than a set threshold, determining the surrounding autonomous driving vehicle as the neighboring autonomous driving vehicle of the target autonomous driving vehicle, and adding the initial driving state feature vector corresponding to the neighboring autonomous driving vehicle to the conflict status graph.

8. The method according to claim 1, wherein The method further includes: Generating experimental data of the intersection based on a traffic simulation tool, where the experimental data includes the driving state feature vectors of multiple virtual autonomous driving vehicles at the first time step, the accelerations that multiple virtual autonomous driving vehicles need to execute at the second time step determined by a decision neural network model to be trained, and the reward values after executing the accelerations, and the second time step is the next time step of the first time step; Performing deep reinforcement learning training on the decision neural network model based on the experimental data and an evaluation neural network model.

9. A method for controlling an autonomous vehicle at an intersection, characterized in that, Applied to an intelligent agent corresponding to the intersection deployed in the cloud, including: At the current time step, constructing the conflict status graph corresponding to each of the multiple autonomous driving vehicles based on the collision conflict relationship of the multiple autonomous driving vehicles at the intersection, where the conflict status graph corresponding to the target autonomous driving vehicle includes the target autonomous driving vehicle and the initial driving state feature vectors corresponding to multiple neighboring autonomous driving vehicles of the target autonomous driving vehicle respectively, and the multiple neighboring autonomous driving vehicles satisfy the collision conflict relationship with the target autonomous driving vehicle at the intersection; Obtaining the predicted accelerations corresponding to each of the multiple autonomous driving vehicles at the next time step of the current time step according to the conflict status graphs corresponding to each of the multiple autonomous driving vehicles and the decision neural network model, and the decision neural network model is trained by means of deep reinforcement learning; Sending the predicted accelerations corresponding to the multiple autonomous driving vehicles at the next time step to the multiple autonomous driving vehicles.

10. An electronic device, characterized in that, Including: A memory, a processor, and a communication interface; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor is caused to execute the method for controlling an autonomous vehicle at an intersection according to any one of claims 1 to 8 or claim 9.

11. A non-transitory machine-readable storage medium, characterized in that, An executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the method for controlling an autonomous vehicle at an intersection according to any one of claims 1 to 8 or claim 9.

Citation Information

Cited By

  • Vehicle lane changing planning method and system based on graph neural network and multiple agents

    CN120756486A

  • Internet of vehicles cooperative driving control method and system based on deep reinforcement learning

    CN121375850A