Non-rational driving behavior decision-making method based on QGNN and HRL
By introducing quantum graph neural networks and hierarchical reinforcement learning in the autonomous driving system, the uncertainty of irrational driving behavior is used to model the uncertainty of irrational driving behavior, and the problem of insufficient decision-making safety and efficiency of autonomous driving systems in hybrid driving environments is solved, achieving higher safety and driving efficiency.
Patent Information
- Application Number
- CN202510223197.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
AI Technical Summary
Existing autonomous driving decision-making methods are difficult to effectively deal with uncertainty under the irrational behavior of human drivers, resulting in insufficient safety and efficiency of decision-making in hybrid driving environments.
Using a method based on quantum graph neural network (QGNN) and hierarchical reinforcement learning (HRL), we model the uncertainty of driving behavior through quantum superposition states and quantum interference mechanisms to achieve global information sharing and dynamic decision-making.
This method can comprehensively capture the characteristics of irrational driving behavior, dynamically adapt to complex environments, improve the safety and driving efficiency of autonomous driving systems, and reduce the safety risks caused by irrational driving.
Smart Images

Figure CN120106131A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving decision-making methods, and relates to an irrational driving behavior decision-making method based on quantum graph neural network (QGNN) and hierarchical reinforcement learning (HRL). Background Art
[0002] With the rapid development of science and technology, especially the booming of new technologies such as artificial intelligence, big data and cloud computing, the progress of autonomous driving technology has shown an unprecedented speed and scale. Due to the limitations of current technical conditions and the constraints of relevant legal responsibilities, manual and autonomous driving vehicles will operate together in a mixed driving environment for a long time in the future. In this process, autonomous driving will inevitably interact with human drivers, and the decision-making problem of autonomous driving in a mixed human-machine driving environment will be a key issue.
[0003] Among them, human drivers have irrational driving problems, such as being forced to cut in or overtaken by other vehicles. The uncertainty of the human driver's state in this process is difficult to characterize and has not received sufficient attention.
[0004] Traditional methods are based on rule- and model-driven methods (such as the Gipps model and the IDM model), which fail to fully consider uncertainty and have single decision-making results. Data-driven deep learning methods (LSTM, Transformer, DRL, etc.) have also made certain progress in recent years. Although they can solve the above problems to a certain extent, irrational behavior is scarce and difficult to digitize, and irrational driving is essentially a small sample and data is difficult to collect, so this method also has certain limitations. Most of the above studies are based on the assumption of rational behavior.
[0005] Quantum theory can well explain human uncertainty due to its inherent characteristics (superposition, probability, etc.). Therefore, this paper combines quantum cognitive decision theory with graph neural networks based on deep learning, and proposes a quantum neural network fusion hierarchical reinforcement learning architecture. This architecture uses the uncertainty modeling of quantum characteristics to solve the decision-making problem of autonomous driving systems under irrational driving conditions of human drivers. By combining quantum neural networks with quantum superposition and probabilistic characteristics, multiple potential results can be comprehensively considered, thereby comprehensively characterizing the uncertainty of human driver behavior. This method provides a safe and reliable decision-making mechanism for autonomous driving systems, aiming to improve the safety and driving efficiency of the system. Summary of the invention
[0006] The purpose of the present invention is to provide an irrational driving behavior decision-making method based on QGNN and HRL, which can comprehensively deal with the uncertainty in driving behavior and effectively reduce the safety risks caused by irrational driving.
[0007] The technical solution adopted by the present invention is an irrational driving behavior decision-making method based on QGNN and HRL, comprising the following steps: Step 1: Build a main strategy model with quantum graph neural network as the core, where each node represents a vehicle, the node features include speed, acceleration, position and heading angle, and the edge features represent the relative distance and relative speed between vehicles; sub-strategies include acceleration, deceleration and lane change; Step 2: Use quantum superposition states to represent node features, update the quantum states of nodes through the quantum graph convolution mechanism, and combine quantum interference to optimize information flow; Step 3: Use quantum entanglement mechanism to realize global information sharing in graph network and capture the complex dynamic interaction relationship between vehicles; Step 4: The main strategy determines the sub-strategy selection under the current driving situation through the QGNN model; Step 5: Generate specific driving instructions based on the sub-strategy, and convert them into vehicle control signals through the actuator to achieve vehicle driving.
[0008] The present invention is also characterized in that: In step 2, the quantum superposition state is expressed as: (1) in, For Node The probability magnitude of the feature; is the quantum ground state, representing different decision states of the node.
[0009] In step 2, the node state update formula of quantum graph convolution is: (2) in, For aggregation operations, quantum superposition and interference mechanisms are combined to optimize the information flow between nodes. Representation Node The quantum state characteristics of Representation Node The quantum state characteristics of The feature of the edge.
[0010] In step 2, quantum graph convolution combines the quantum state of each node with the quantum state of neighboring nodes and the quantum features of edges to achieve information aggregation through quantum superposition and interference. The specific implementation steps are as follows: (3) in, Representation Node exist The quantum state when Is a node The neighbor set of Representation Node exist The quantum state when Representation Node With Node The edge feature of the node The impact strength of the status update, is the edge feature, For Node The probability magnitude of the feature.
[0011] In step 2, quantum interference optimizes the information flow between nodes by adjusting the phase to enhance or suppress the information of a specific path. The update formula is: (4) in, represents the complex phase factor.
[0012] In step 3, the entangled state of the quantum entanglement mechanism is expressed as: (5) in, Representation Node and nodes The joint quantum state between For Node The probability amplitude of the feature, Representation Node exist The quantum state when Representation Node exist The quantum state when Represents the tensor product of quantum states.
[0013] In step 4, QGNN determines the choice of sub-strategy by measuring the quantum state. The quantum measurement formula is: (6) in, is the quantum state of the node, Is with sub-strategy The corresponding quantum state, Is the selection strategy probability.
[0014] In step 5, the sub-strategy adopts the reinforcement learning method. The specific input is the observable data outside the vehicle, including the vehicle's speed, acceleration, position and heading angle, and the output is the driving command, including the accelerator pedal opening, brake pedal force and steering wheel angle.
[0015] The beneficial effects of the present invention are: (1) The method of the present invention introduces quantum graph neural network to replace traditional neural network, and uses superposition and probability in quantum theory to represent the driver's various possible behavior states through quantum superposition state, which can fully capture the characteristics of irrational driving behavior. Through quantum interference, the information flow path is adjusted, specific decision paths are strengthened or suppressed, and dynamic adaptation to complex environments is achieved; (2) The method of the present invention uses quantum graph neural networks to quantize the dynamic relationship between vehicles. The characteristics of nodes and edges can efficiently model the interaction and complex relationship between vehicles. Through quantum entanglement, global information sharing is achieved, and the collaborative decision-making ability between remote nodes is enhanced; (3) Compared with traditional methods based on rational assumptions, the method of the present invention can comprehensively deal with the uncertainty in driving behavior, effectively reduce the safety risks caused by irrational driving, and improve the reliability and driving efficiency of the system in dynamic scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flow chart of the method of the present invention; Figure 2 This is a diagram of the decision framework for irrational driving behavior in the method of the present invention. DETAILED DESCRIPTION
[0017] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] Embodiment 1: This paper is based on the irrational driving behavior decision-making method of quantum graph neural network (QGNN) and hierarchical reinforcement learning (HRL). Aiming at the characteristics of irrational driver behavior in human-machine mixed driving environment, the uncertainty of driving behavior is modeled by introducing quantum superposition and quantum interference mechanism. This method uses quantum graph convolution to efficiently capture the dynamic interaction relationship between vehicles. The main strategy dynamically selects sub-strategies (acceleration, deceleration, lane change) in complex scenarios through QGNN. The sub-strategies generate specific driving instructions based on external observable data to ensure the safety and adaptability of real-time decision-making.
[0019] The present invention is based on the irrational driving behavior decision-making method of QGNN and HRL, such as Figure 1 As shown, the specific implementation steps are as follows: Step 1: Build a main strategy model with quantum graph neural network as the core, where each node represents a vehicle, the node features include speed, acceleration, position and heading angle, and the edge features represent the relative distance and relative speed between vehicles. In addition, sub-strategies include acceleration, deceleration and lane change.
[0020] The decision framework of the method of the present invention is as follows Figure 2 As shown in the figure, the main strategy uses a quantum graph neural network, where each node represents a car , the characteristics of the nodes include the speed, acceleration, position and heading angle of the vehicle. The edges in the graph represent the relationship between vehicles, and their characteristics include: relative distance and relative speed. If the vehicle and If there is an edge between them, then the characteristics of the edge It is expressed as: ,Through the above quantization, QGNN can efficiently capture the mutual ,influence and complex relationships among vehicles.
[0021] Sub-strategies are selected based on the main strategy, including acceleration, deceleration and lane changing.
[0022] Step 2: Use quantum superposition states to represent node characteristics, update the quantum states of nodes through the quantum graph convolution mechanism, and combine quantum interference to optimize information flow.
[0023] The node features in the method of the present invention are represented by quantum superposition states, which enables the vehicle to consider multiple possibilities simultaneously during the decision-making process and to handle dynamic driving situations. The quantum superposition state is represented as: (1) in, For Node The probability magnitude of the feature; is the quantum ground state, representing different decision states of the node.
[0024] In quantum graph neural networks, node updates are implemented through quantum graph convolution, and the node state update formula of quantum graph convolution is: (2) in, For aggregation operations, quantum superposition and interference mechanisms are combined to optimize the information flow between nodes. Representation Node The quantum state characteristics of Representation Node quantum state characteristics.
[0025] Quantum graph convolution combines the quantum state of each node with the quantum state of neighboring nodes and the quantum features of edges, and realizes information aggregation through quantum superposition and interference. The specific implementation steps are as follows: (3) in, Representation Node exist The quantum state when Is a node The neighbor set of Representation Node exist The quantum state when Representation Node With Node The edge feature of the node The impact strength of the status update.
[0026] Quantum interference optimizes the information flow between nodes by adjusting the phase to strengthen or suppress the information of a specific path. The update formula is: (4) in, represents the complex phase factor.
[0027] Through intervention, the model can adjust the decision path according to changes in the environment and situation, simulating the driver's rapid reactions and irrational decisions in complex environments.
[0028] Step 3: Realize global information sharing in the graph network through the quantum entanglement mechanism to capture the complex dynamic interaction relationship between vehicles.
[0029] The quantum entanglement mechanism allows all nodes in the graph network to share global information, and its entangled state is represented as: (5) in, Representation Node and nodes The joint quantum state between Represents the tensor product of quantum states.
[0030] Through quantum entanglement, nodes can share information across the entire range.
[0031] Step 4: The main strategy determines the sub-strategy selection under the current driving situation through the QGNN model.
[0032] QGNN determines the choice of sub-strategy by measuring the quantum state. The quantum measurement formula is: (6) in, is the quantum state of the node, Is with sub-strategy The corresponding quantum state, Is the selection strategy probability.
[0033] Through the measurement of quantum states and the calculation of inner products, the system can select the most suitable sub-strategy based on the similarity between the current node quantum state and the quantum states of each sub-strategy. The main strategy ensures that the best decision is made under the current state by selecting a sub-strategy with a higher probability of selection. For example, if the current node quantum state matches the quantum state of the "acceleration" sub-strategy well (large inner product), then "acceleration" will be the sub-strategy selected by the main strategy.
[0034] Step 5: Generate specific driving instructions based on the sub-strategy, and convert them into vehicle control signals through the actuator to achieve vehicle driving.
[0035] The sub-strategy adopts the reinforcement learning method, and the specific input is the external observable data of the vehicle , including the vehicle's speed , acceleration ,Location and heading angle , the output is the driving command , including accelerator pedal opening, brake pedal force and steering wheel angle.
[0036] Table 1 shows the sub-strategy state space settings, and Table 2 shows the sub-strategy action space settings, where the accelerator pedal opening range is , brake pedal force range , the vehicle's steering wheel angle range The actuator converts the control instructions output by the sub-strategy into specific vehicle control signals, thereby completing the vehicle's driving.
[0037] Table 1 Sub-strategy state space settings
[0038] Table 2 Sub-strategy action space settings
[0039] The role of the reward function is to evaluate the quality of the actions taken by the sub-strategy, and give rewards or penalties based on the evaluation, thereby guiding the sub-strategy to learn the optimal strategy. For the setting of the reward function, the purpose of the acceleration sub-strategy is to reach the target speed and ensure driving safety. Its reward function is designed as follows: (7) in, and represents the weight coefficient related to speed reward and distance reward; , which represents the absolute difference between the current vehicle speed and the target vehicle speed.
[0040] (8) in, Indicates the safe distance reward from other cars. Indicates the distance between the vehicle and the vehicle in front. Indicates a safe distance.
[0041] The purpose of the deceleration sub-strategy is to decelerate the vehicle smoothly and maintain safety. Its reward function is as follows: (9) in, and Represents the weight coefficient related to safety and stability; , represents the smooth deceleration reward, represents the acceleration of the vehicle; Represents a security bonus.
[0042] The lane-changing sub-strategy aims to maintain the safety and efficiency of the lane-changing process in order to successfully change lanes. Its reward function is as follows: (10) in, , , They represent the weight coefficients related to lane change success reward, collision risk reward and lane change efficiency reward respectively; Represents the reward for successful lane change, if the lane change is successful, it means +1, otherwise it means 0; Represents the collision risk penalty, if there is a collision risk, it means -1, otherwise it means 0; stands for Lane Changing Efficiency Reward, , Represents the time required to complete the lane change.
[0043] Embodiment 2: The irrational driving behavior decision-making method based on QGNN and HRL in this embodiment includes the following steps: Step 1: Build a main strategy model with quantum graph neural network as the core, where each node represents a vehicle, the node features include speed, acceleration, position and heading angle, and the edge features represent the relative distance and relative speed between vehicles; sub-strategies include acceleration, deceleration and lane change; Step 2: Use quantum superposition states to represent node features, update the quantum states of nodes through the quantum graph convolution mechanism, and combine quantum interference to optimize information flow; Step 3: Use quantum entanglement mechanism to realize global information sharing in graph network and capture the complex dynamic interaction relationship between vehicles; Step 4: The main strategy determines the sub-strategy selection under the current driving situation through the QGNN model; Step 5: Generate specific driving instructions based on the sub-strategy, and convert them into vehicle control signals through the actuator to achieve vehicle driving.
[0044] Embodiment 3: Based on Example 2, in step 2, the quantum superposition state is expressed as: (1) in, For Node The probability magnitude of the feature; is the quantum ground state, representing different decision states of the node.
[0045] Embodiment 4: Based on Example 3, in step 2, the node state update formula of quantum graph convolution is: (2) in, For aggregation operations, quantum superposition and interference mechanisms are combined to optimize the information flow between nodes. Representation Node The quantum state characteristics of Representation Node The quantum state characteristics of The feature of the edge.
[0046] Embodiment 5: On the basis of Example 4, in step 2, quantum graph convolution combines the quantum state of each node with the quantum state of neighboring nodes and the quantum features of edges, and realizes information aggregation through quantum superposition and interference. The specific implementation steps are as follows: (3) in, Representation Node exist The quantum state when Is a node The neighbor set of Representation Node exist The quantum state when Representation Node With Node The edge feature of the node The impact strength of the status update, is the edge feature, For Node The probability magnitude of the feature.
[0047] In step 2, quantum interference optimizes the information flow between nodes by adjusting the phase to enhance or suppress the information of a specific path. The update formula is: (4) in, represents the complex phase factor.
[0048] Embodiment 6: Based on Example 5, in step 3, the entangled state of the quantum entanglement mechanism is expressed as: (5) in, Representation Node and nodes The joint quantum state between For Node The probability amplitude of the feature, Representation Node exist The quantum state when Representation Node exist The quantum state when Represents the tensor product of quantum states.
Claims
1. The irrational driving behavior decision-making method based on QGNN and HRL is characterized by: The following steps are involved: Step 1: Build a main strategy model with quantum graph neural network as the core, where each node represents a vehicle, the node features include speed, acceleration, position and heading angle, and the edge features represent the relative distance and relative speed between vehicles; sub-strategies include acceleration, deceleration and lane change; Step 2: Use quantum superposition states to represent node features, update the quantum states of nodes through the quantum graph convolution mechanism, and combine quantum interference to optimize information flow; Step 3: Use quantum entanglement mechanism to realize global information sharing in graph network and capture the complex dynamic interaction relationship between vehicles; Step 4: The main strategy determines the sub-strategy selection under the current driving situation through the QGNN model; Step 5: Generate specific driving instructions based on the sub-strategy, and convert them into vehicle control signals through the actuator to achieve vehicle driving.
2. The irrational driving behavior decision-making method based on QGNN and HRL according to claim 1, characterized in that: In step 2, the quantum superposition state is expressed as: (1) in, For Node The probability magnitude of the feature; is the quantum ground state, representing different decision states of the node.
3. The irrational driving behavior decision-making method based on QGNN and HRL according to claim 1, characterized in that: In step 2, the node state update formula of quantum graph convolution is: (2) in, For aggregation operations, quantum superposition and interference mechanisms are combined to optimize the information flow between nodes. Representation Node The quantum state characteristics of Representation Node The quantum state characteristics of The feature of the edge.
4. The irrational driving behavior decision-making method based on QGNN and HRL according to claim 1, characterized in that: In step 2, quantum graph convolution combines the quantum state of each node with the quantum state of neighboring nodes and the quantum features of edges to achieve information aggregation through quantum superposition and interference. The specific implementation steps are as follows: (3) in, Representation Node exist The quantum state when Is a node The neighbor set of Representation Node exist The quantum state when Representation Node With Node The edge feature of the node The impact strength of the status update, is the edge feature, For Node The probability magnitude of the feature.
5. The irrational driving behavior decision-making method based on QGNN and HRL according to claim 4 is characterized in that: In step 2, quantum interference optimizes the information flow between nodes by adjusting the phase to enhance or suppress the information of a specific path. The update formula is: (4) in, represents the complex phase factor.
6. The irrational driving behavior decision-making method based on QGNN and HRL according to claim 1, characterized in that: In step 3, the entangled state of the quantum entanglement mechanism is expressed as: (5) in, Representation Node and nodes The joint quantum state between For Node The probability amplitude of the feature, Representation Node exist The quantum state when Representation Node exist The quantum state when Represents the tensor product of quantum states.
7. The irrational driving behavior decision-making method based on QGNN and HRL according to claim 1, characterized in that: In step 4, QGNN determines the choice of sub-strategy by measuring the quantum state. The quantum measurement formula is: (6) in, is the quantum state of the node, Is with sub-strategy The corresponding quantum state, Is the selection strategy probability.
8. The irrational driving behavior decision-making method based on QGNN and HRL according to claim 1, characterized in that: In step 5, the sub-strategy adopts the reinforcement learning method. The specific input is the observable data outside the vehicle, including the vehicle's speed, acceleration, position and heading angle, and the output is the driving command, including the accelerator pedal opening, brake pedal force and steering wheel angle.