Automatic driving decision-making method based on space-time diagram network and risk perception reinforcement learning

By employing spatiotemporal graph networks and risk perception reinforcement learning, a structured traffic state representation is constructed and a risk situation description is performed. This addresses the issue of decision-making insecurity for autonomous vehicles in complex traffic scenarios, enabling safe, efficient, and robust driving decisions.

CN121997461APending Publication Date: 2026-05-08SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEST JIAOTONG UNIV
Filing Date
2026-01-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing autonomous driving decision-making methods struggle to effectively characterize the interaction relationships between traffic participants, the evolution characteristics of their states over time, and potential behavioral risks in complex traffic scenarios. This results in insufficient response to potentially dangerous scenarios during the decision-making process, making it difficult to balance driving safety, traffic efficiency, and driving comfort.

Method used

We employ a method based on spatiotemporal graph networks and risk perception reinforcement learning. By constructing a structured traffic state representation, we perform information dissemination and clustering, combine it with Bayesian inference to obtain a risk situation description, and use it as the state input for reinforcement learning. The goal is to optimize decision-making by maximizing the cumulative expectation.

Benefits of technology

It enables safe, efficient and robust driving decision control for autonomous vehicles in complex traffic scenarios, significantly improves the ability to proactively constrain and assess risks in potentially dangerous scenarios, and enhances the safety and stability of the decision-making process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997461A_ABST
    Figure CN121997461A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic driving and intelligent traffic, and discloses an automatic driving decision-making method based on a space-time diagram network and risk perception reinforcement learning, which comprises the following steps of: acquiring states of an automatic driving vehicle and surrounding traffic participants, abstracting the automatic driving vehicle and the surrounding traffic participants into traffic map nodes, traffic state representation containing node time sequence characteristics and connection relations is generated; information spreading and clustering are carried out on the node time sequence features, and automatic driving vehicle node representation is generated; constructing an observation state, performing Bayesian inference by taking potential behaviors of surrounding traffic participants as hidden variables, and outputting risk situation description; and finally, fusing automatic driving vehicle node representation and risk situation description, generating an enhanced state, taking the enhanced state as reinforcement learning input, and realizing an optimal driving decision by taking the risk situation description as a constraint and maximizing cumulative expectation as a target. According to the invention, safe, efficient and robust decision control of the autonomous vehicle in a complex traffic scene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of autonomous driving and intelligent transportation technology, specifically to an autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning. Background Technology

[0002] With the continuous development of autonomous driving technology, the autonomous decision-making ability of autonomous vehicles in complex traffic environments has gradually become a key focus of research and application. In real-world traffic scenarios, autonomous vehicles typically need to interact simultaneously with multiple traffic participants, such as other vehicles, pedestrians, or non-motorized vehicles, and make driving decisions in dynamic environments such as multi-lane and dense traffic flow. These traffic scenarios are often characterized by a large number of traffic participants, complex interrelationships, and constantly changing states over time, placing high demands on the safety and stability of autonomous driving decision-making methods.

[0003] In existing technologies, autonomous driving decisions are typically based on modeling and analyzing traffic scene states. This involves processing information about the vehicle's own state and the surrounding environment to generate corresponding driving decisions. However, in practical applications, it has been found that traffic participants not only have directly observable relationships such as position and speed, but also complex interactive relationships formed by lane structure, traffic rules, and mutual behavioral influences. Furthermore, the driving behavior of traffic participants has a degree of uncertainty, and their future behavioral changes are often difficult to accurately describe solely based on the deterministic state at the current moment.

[0004] To address the aforementioned issues, some existing methods attempt to improve the decision-making performance of autonomous driving by introducing mechanisms such as structured modeling, behavior prediction, or risk assessment. However, in complex and dynamic traffic scenarios, simultaneously characterizing the interaction relationships between traffic participants, the evolutionary characteristics of states over time, and potential behavioral risks within a unified decision-making framework remains challenging. Especially when balancing driving safety, traffic efficiency, and driving comfort, existing decision-making methods often struggle to effectively constrain traffic risks, easily leading to insufficient response to potentially dangerous scenarios. Summary of the Invention

[0005] To address the aforementioned shortcomings in existing technologies, this invention provides an autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning. By constructing a structured traffic state representation with time awareness capabilities and explicitly modeling the uncertainty of the behavior of surrounding traffic participants, this method solves the safety issues of autonomous vehicles making decisions in real traffic environments that exist in existing technologies, and achieves safe, efficient, and robust driving decision control for autonomous vehicles in complex traffic scenarios.

[0006] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: An autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning includes the following steps: S1. Obtain the state information of autonomous vehicles and their corresponding surrounding traffic participants, abstract autonomous vehicles and their corresponding surrounding traffic participants into traffic graph nodes in the current traffic scenario, obtain the connection relationships between each node, and generate a traffic state representation that includes the temporal features of the nodes and the connection relationships. S2. Based on traffic state representation, information propagation and clustering are performed on the temporal features of nodes to generate autonomous vehicle node representations that integrate spatial interactions with surrounding traffic participants. S3. Construct the observed traffic state, and at the same time, use the potential behavior or behavioral intentions of surrounding traffic participants as latent variables, perform Bayesian inference, calculate the posterior probability distribution of the latent variables, and obtain a risk situation description of the potential risk level under the current traffic scenario. S4. The node representation of the autonomous vehicle is fused with the risk situation description to generate an enhanced state representation that includes traffic interaction situation information and risk uncertainty information. This enhanced state representation is then used as the state input for reinforcement learning. With the risk situation description as a constraint and the goal of maximizing the cumulative expectation, the optimal driving decision of the autonomous vehicle is achieved.

[0007] The present invention has the following beneficial effects: The autonomous driving decision-making method proposed in this invention, based on spatiotemporal graph networks and risk perception reinforcement learning, constructs a structured traffic state representation with time-aware capabilities and explicitly models the uncertainty of the behavior of surrounding traffic participants. By introducing the risk situation description obtained based on Bayesian inference into the traffic state representation, a reinforcement learning strategy optimization process is constructed, which unifies the risk uncertainty with traffic interaction modeling and decision optimization. This not only achieves forward-looking constraints on potentially dangerous behaviors but also enables safe, efficient, and robust driving decision control for autonomous vehicles in complex traffic scenarios. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating the autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning proposed in this invention. Detailed Implementation

[0009] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0010] like Figure 1As shown, the autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning includes the following steps S1-S4: S1. Obtain the state information of the autonomous vehicle and the corresponding surrounding traffic participants, abstract the autonomous vehicle and the corresponding surrounding traffic participants into traffic graph nodes in the current traffic scenario, obtain the connection relationship between each node, and generate a traffic state representation containing the temporal features of the nodes and the connection relationship.

[0011] Specifically, the connection relationship between nodes is at least one of the following: relative distance relationship, relative position relationship, and lane topology relationship.

[0012] In this embodiment, the traffic scenario is a multi-lane road environment. The autonomous vehicle can interact longitudinally and laterally with multiple surrounding traffic participants within its own lane and adjacent lanes. The autonomous vehicle obtains its own state information and the state information of surrounding traffic participants through the onboard perception system. The state information of the autonomous vehicle may include the vehicle's current position, speed, acceleration, heading angle, etc., while the state information of surrounding traffic participants may include the relative position, relative speed, and lane occupancy information of other vehicles, pedestrians, or non-motorized vehicles.

[0013] Meanwhile, the current traffic scenario is modeled using a local traffic scenario centered on the autonomous vehicle. Specifically, the current traffic scenario covers the lane where the autonomous vehicle is located and its adjacent lanes in the horizontal direction, and covers the surrounding traffic participants within a preset distance range in front of and behind the autonomous vehicle in the vertical direction, thus forming an approximately rectangular local traffic area. This modeling method can limit the surrounding traffic participants related to autonomous driving decisions to a limited spatial range, thereby reducing the complexity of state modeling and improving decision-making efficiency. Furthermore, the vertical distance range can be set according to the vehicle's driving speed or application requirements, and this invention does not impose specific limitations on this.

[0014] Then, based on the obtained state information of the current traffic scene, a unified model is performed on the motion state of surrounding traffic participants, road semantic information, and driving constraints. Specifically, the autonomous vehicle and its surrounding traffic participants are abstracted as nodes in a traffic graph, and the temporal features of each node are represented as state vectors to serve as inputs for subsequent graph interaction modeling and reinforcement learning decisions. The temporal features of each node at the current moment include the temporal features of the previous moment and the observation features of the current moment. Furthermore, the state information of each node at a certain moment is the observation feature of the node, which is represented as follows:

[0015] in, Represents a node At any moment Status information, Represents a node At any moment The longitudinal position of the lane at that time. Represents a node At any moment speed, Represents a node At any moment acceleration, Represents a node At any moment The lane markings at the time. Represents a node At any moment Lateral offset from the lane centerline Represents a node At any moment The constraints on whether lane changing is permitted refer to the physical feasibility constraints determined by objective conditions such as road structure, traffic rules, or environmental factors. Represents a node At any moment The lateral position of the lane at that time. This refers to autonomous vehicles; the remaining cases involve surrounding traffic participants. This indicates the transpose operation.

[0016] To address the characteristics of the evolution of the state of surrounding traffic participants over time, this invention introduces a temporal feature recursive update mechanism to encode the historical information of each node, ultimately generating the temporal features of each node at a certain moment. The temporal feature recursive update mechanism employs a gated loop unit; therefore, the node temporal features are recursively updated using a gated loop unit, which can be expressed as:

[0017] in, Represents a node At any moment The temporal characteristics, This represents a nonlinear state update function. Represents a node At any moment The temporal characteristics.

[0018] Furthermore, the aforementioned recursive update mechanism for time-series features can also be implemented using other equivalent time-series modeling structures. These structures can adaptively retain historical information that has a significant impact on decision-making during continuous state updates, while gradually attenuating or forgetting historical information that has a smaller impact on current decisions. Therefore, the length of effective historical information contained in the time-series features is not fixed to a single moment, but is adaptively determined by the model structure and its parameters. Its time span can cover multiple recent historical moments, or even longer time ranges.

[0019] Therefore, the above recursive form does not mean that information from only a single historical moment is used; on the contrary, temporal characteristics... It already encodes historical observation information obtained at multiple previous moments; through a recursive update method, the historical observation state is gradually compressed and integrated into the temporal features at each moment, thereby forming a comprehensive representation of the historical movement state of the surrounding traffic participants.

[0020] S2. Based on traffic state representation, information propagation and clustering are performed on the temporal features of nodes to generate autonomous vehicle node representations that integrate spatial interactions with surrounding traffic participants.

[0021] In this embodiment, by performing information propagation and clustering on the traffic map, the temporal features of the nodes of surrounding traffic participants are transmitted to the nodes corresponding to the autonomous vehicle, thereby obtaining the autonomous vehicle node representation that integrates the spatiotemporal interaction relationships of the surrounding traffic participants, which can be used to reflect the overall interaction status of the autonomous vehicle in the current traffic scenario; the operation is as follows: First, to reduce computational complexity, we construct a set of candidate neighbors for each node. , Represents a node The candidate neighbor set is specifically determined by using a distance threshold or a Top-K method to identify candidate neighbors for each node, thus generating a candidate neighbor set for each node. .

[0022] Specifically, the formula for determining the candidate neighbors of each node using a distance threshold is as follows:

[0023] in, Represents a node Candidate neighbors, Represents a node At any moment The longitudinal position of the lane at that time. Represents a node At any moment The longitudinal position of the lane at that time. This represents the distance threshold.

[0024] Specifically, using the Top-K method, the formula for determining the candidate neighbors of each node is as follows:

[0025] in, Represents a node Candidate neighbors, This indicates the upper limit of the number of candidate neighbors. Represents a node At any moment The longitudinal position of the lane at that time. Represents a node At any moment The longitudinal position of the lane at that time.

[0026] Secondly, before information propagation between nodes, each node adaptively selects its neighbor set. After obtaining the candidate neighbor sets of each node, a multilayer perceptron (MLP) is used to implement scoring. Therefore, the interaction scoring between nodes can be expressed as:

[0027]

[0028] in, Indicates at time node With nodes Interactive ratings between them This represents a function that characterizes the correlation or influence strength of interactions between nodes. Represents a node At any moment The temporal characteristics, Represents a node At any moment The temporal characteristics, Indicates at time node With nodes Relative state information between them Represents a node With nodes The longitudinal relative position between them Represents a node With nodes The horizontal relative position between them Represents a node With nodes The relative velocity between them Represents a node With nodes The relative acceleration between them Represents a node With nodes The spatial relationship of the lanes This indicates the transpose operation. Represents the weight vector. Represents a non-linear activation function. This indicates a splicing operation. These represent the weight matrix and the bias vector, respectively.

[0029] Furthermore, the interaction scoring between nodes can also be achieved by other linear mappings, attention networks, or other equivalent nonlinear mapping structures besides multilayer perceptrons (MLPs), without affecting the characterization of the interaction correlation between nodes.

[0030] Then, based on the interaction scores between nodes, the candidate neighbor set is selected. The Top-K2 neighbors are selected as the final neighbor set. , Represents a node The final neighbor set, where K2 represents the number of final neighbor sets participating in spatial interaction aggregation; and satisfies K2 can be configured based on road density, computing power budget, and interaction complexity.

[0031] Based on this, and using the final neighbor set, the Softmax function is used to map the interaction scores between nodes into interaction weights, which are expressed as follows:

[0032] in, Represents a node With nodes Interaction weights between them Represents an exponential function. Indicates at time node With nodes Interactive ratings between users.

[0033] In addition, interaction weights can also be obtained using methods other than the Softmax function, such as Sigmoid normalization, linear normalization, or sparse normalization.

[0034] Finally, the spatial interaction features of nodes can be obtained by weighted clustering of the temporal features of the final neighbor nodes, ultimately yielding an autonomous vehicle node representation that integrates the spatial interactions of surrounding traffic participants, which is represented as follows:

[0035]

[0036] in, Indicates at time node Spatial interaction characteristics, Indicates at time The set of spatial interaction features of nodes Indicates at time The spatial interaction characteristics of autonomous vehicle nodes are represented by the spatial interactions of surrounding traffic participants. Indicates at time node Spatial interaction characteristics.

[0037] In summary, the autonomous vehicle node representation is used to reflect the overall interactive situation of autonomous vehicles in the current traffic scenario and serves as one of the inputs for subsequent risk situation modeling and decision-making.

[0038] S3. Construct the observed traffic state, and at the same time, use the potential behavior or behavioral intentions of surrounding traffic participants as latent variables, perform Bayesian inference, calculate the posterior probability distribution of the latent variables, and obtain a risk situation description of the potential risk level under the current traffic scenario.

[0039] In this embodiment, this step involves modeling the uncertainty of the potential behavioral states or intentions of surrounding traffic participants. This is achieved by introducing a risk situation assessment mechanism based on probability inference, utilizing latent variables to characterize some surrounding traffic participants at time [time value missing]. The potential behavioral states, including but not limited to acceleration, deceleration, lane changing, or lateral intrusion, are used to ultimately obtain a risk situation description of the potential risk level in the current traffic scenario. The operation is as follows: First, Bayesian inference is used to observe traffic conditions. As an observation, traffic conditions are observed. The state representation constructed from a risk assessment perspective partially overlaps with the traffic state representation used for graph interaction modeling in terms of information sources, but their functions differ; the observed traffic state representation is as follows:

[0040] in, Indicates the observed traffic conditions. Indicates the time of the autonomous vehicle node Status information, Indicates at time The representation of autonomous vehicle nodes that integrates spatial interactions with surrounding traffic participants. This represents the set pooling operator, which can take either mean pooling or max pooling. This represents the final set of neighbors for the autonomous vehicle node. Represents a node At any moment Status information, Indicates at time Autonomous vehicle nodes and nodes Relative state information between them Indicates at time Environmental semantic information, such as speed limits, road types, lane line types, and traffic light status.

[0041] Secondly, given the observed traffic conditions The potential behaviors or intentions of surrounding traffic participants are then used as latent variables. Bayes' theorem is used to calculate the posterior probability distribution of these latent variables, which is expressed as follows:

[0042]

[0043] in, Representing latent variables The prior probability distribution function, Indicates that given a latent variable Traffic conditions observed under certain conditions The likelihood function, Indicates the traffic conditions observed. Hidden variables The posterior probability distribution function, Indicates the observed traffic conditions The probability distribution function.

[0044] Then, based on the posterior probability distribution of the latent variables, a risk situation description is further constructed. Specifically, to simultaneously consider the proximity risk between the autonomous vehicle and surrounding vehicles in both longitudinal and lateral directions, for the autonomous vehicle node and any target node (target vehicle), based on the posterior probability distribution... In the prediction time domain The relative displacements of the two are calculated, and a joint safety factor is constructed using the longitudinal and lateral safety envelope half-thresholds. Collision events are determined, and the most dangerous value of a collision between the autonomous vehicle and the target node in the prediction time domain is obtained, further forming an overall risk assessment. This overall risk assessment will be used as risk situation description information to participate in subsequent state enhancement and security constraint modeling, specifically: First, calculate the predicted vertical and horizontal positions of the target node at future times, given the latent variables:

[0045]

[0046]

[0047] in, Indicates that given a latent variable target node under conditions In the future The predicted vertical position, Represents the target node At any moment The vertical position, Represents the target node At any moment speed, Indicates that given a latent variable target node under conditions acceleration, Indicates that given a latent variable target node under conditions In the future The predicted lateral position, Represents the target node At any moment The lateral position of the lane at that time. Represents the target node The lateral coordinates of the center line of the lane in question. This represents a continuous-time variable within the prediction time domain. This indicates the total length of the prediction time domain.

[0048] Secondly, based on the predicted longitudinal and lateral positions of the target node at future times under given latent variables, the relative longitudinal and lateral displacements between the autonomous vehicle node and the target node at future times are calculated under given latent variables, i.e.:

[0049]

[0050] in, Indicates that given a latent variable Under certain conditions, autonomous vehicle nodes and target nodes In the future The relative longitudinal displacement, Indicating autonomous vehicles in the future The vertical position, Indicates that given a latent variable Under certain conditions, autonomous vehicles and target nodes In the future The relative lateral displacement, Indicates the node of the autonomous vehicle at a future time. The lateral position of the lane at that time.

[0051] Then, after calculating the longitudinal and lateral relative positions of the autonomous vehicle node and the target node in the prediction time domain, to achieve concise and engineering-feasible safety constraint modeling, a safety envelope and safety metric are further constructed to form a continuous risk quantity, specifically: Obtain vehicle geometry and construct a collision envelope half-threshold, including parameters for determining the relationship between autonomous vehicle nodes and target nodes. Envelope half threshold for collisions occurring in the longitudinal direction Used to determine the relationship between autonomous vehicle nodes and target nodes. Envelope half threshold for collisions occurring in the lateral direction ,Right now:

[0052]

[0053] in, , Representing the autonomous vehicle node and the target node respectively. The length of the car body, , Representing the autonomous vehicle node and the target node respectively. The width of the vehicle body; Determine relative longitudinal displacement Is it less than or equal to the envelope half threshold? And relative lateral displacement Less than or equal to the envelope half threshold If so, then the autonomous vehicle node and the target node If a collision occurs, proceed to the next step: calculate the overall risk assessment; otherwise, do not engage in a collision.

[0054] Introducing a safety margin, we calculate the longitudinal safety envelope half-threshold and the lateral safety envelope half-threshold, namely:

[0055]

[0056] in, Represents the autonomous vehicle node and the target node The longitudinal safety envelope half threshold at which a collision occurs. Represents the autonomous vehicle node and the target node The lateral safety envelope half threshold at which a collision occurs. , Both represent safety margins.

[0057] In this step, the safety margin is used to reflect safety redundancy (e.g., response time, sensing error, policy conservatism, etc.), and the safety margin... and It can be set to a constant based on the road scene level, vehicle dynamics characteristics, or safety strategy.

[0058] Based on the vertical security envelope half-threshold and the horizontal security envelope half-threshold, the joint security degree is calculated, i.e.:

[0059] in, Indicates the relationship between autonomous vehicles and target nodes In the future The combined safety level in the event of a collision. This indicates taking the maximum value.

[0060] Based on the joint safety degree, the most dangerous value in the prediction time domain is calculated, namely:

[0061] in, Represents the autonomous vehicle node and the target node At any moment The most dangerous value at that time This indicates taking the minimum value.

[0062] Based on the most dangerous value in the prediction time domain, a segmented truncation method is used to construct the continuous risk quantity corresponding to the target node, i.e.:

[0063] in, Represents the target node At any moment Continuous risk quantity This represents the truncation operator, which truncates the target value to the interval [0, 1], thus making the autonomous vehicle node and the target node (target vehicle) closer together in the prediction time domain (i.e., ...). The smaller the value, the greater the risk; conversely, the greater the value, the closer the risk approaches zero.

[0064] Furthermore, when there are multiple target nodes, the continuous risk quantities corresponding to each target node can be aggregated to obtain the maximum continuous risk quantity, which can then be used as the overall risk assessment quantity.

[0065] in, Indicates at time The overall risk assessment quantity, i.e., the risk situation description of the potential risk level.

[0066] Finally, the overall risk assessment is used as a risk situation description of the potential risk level in the current traffic scenario.

[0067] S4. The node representation of the autonomous vehicle is fused with the risk situation description to generate an enhanced state representation that includes traffic interaction situation information and risk uncertainty information. This enhanced state representation is then used as the state input for reinforcement learning. With the risk situation description as a constraint and the goal of maximizing the cumulative expectation, the optimal driving decision of the autonomous vehicle is achieved.

[0068] In this embodiment, the node representation of the autonomous vehicle is fused with the risk situation description to construct an enhanced state representation that simultaneously includes traffic interaction situation information and risk uncertainty information. This enhanced state representation is used as the state input of the reinforcement learning decision model, which is modeled as a constrained Markov decision process. The risk situation description not only participates in policy learning as state information representing the level of traffic risk but also serves as risk constraint information limiting the risk level of driving decisions. By constraining the highest-risk factor, driving behavior during policy optimization is constrained, thereby preventing the autonomous vehicle from outputting driving decisions that do not meet safety requirements in high-risk traffic situations. This driving decision behavior includes at least longitudinal, lateral, and lane-changing decisions of the autonomous vehicle; its operation is as follows: First, the node representation of autonomous vehicles is fused with the risk situation description to generate an enhanced state representation that includes traffic interaction situation information and risk uncertainty information, namely:

[0069] in, Indicates at time The enhanced state representation, This represents the fusion operation function. Indicates at time The spatial interaction characteristics of autonomous vehicle nodes are represented by the spatial interactions of surrounding traffic participants. Indicates at time The overall risk assessment measure, i.e., the risk situation description of the potential risk level, , All of these represent learnable parameters. This indicates a splicing operation.

[0070] In this step, the representation of autonomous vehicle nodes is concatenated with the risk situation description before mapping; in other embodiments, weighted fusion or attention mechanisms can also be used, but this does not affect the core process of the present invention.

[0071] Therefore, by constructing the above-mentioned enhanced state representation, the autonomous driving decision-making model (reinforcement learning decision-making model, which is adopted in this invention) can not only consider the spatial interaction relationship and temporal evolution characteristics between surrounding traffic participants when optimizing the strategy, but also comprehensively consider the risk impact brought about by potential behavioral uncertainties, thereby providing a unified and complete state input for the subsequent reinforcement learning decision-making process that introduces safety constraints.

[0072] Secondly, enhance the state representation As a reinforcement learning decision model at time... Given a state representation as input, the reinforcement learning decision model outputs the corresponding driving behavior action based on the reinforced state representation. Driving actions do not directly correspond to the underlying control commands of autonomous vehicles, but are used to characterize the driving actions of autonomous vehicles in the current traffic scenario. These driving actions include at least longitudinal and lateral driving actions, as well as lane-changing actions, and are represented as follows:

[0073] in, Indicates at time Driving behavior and actions, Indicates at time Time-discrete longitudinal driving behavior Indicates at time Time-discrete lateral driving behavior Indicates at time Time-discrete lane-changing behavior.

[0074] In this embodiment, the optimization objective of autonomous driving decision-making can be viewed as a multi-objective optimization problem, where different objectives are collaboratively modeled through reward and constraint terms. In one implementation of the invention, a reinforcement learning reward function is used to characterize optimization objectives related to vehicle driving performance, which may include, but are not limited to, traffic efficiency, driving smoothness, or energy consumption level. Safety-related objectives are reflected through safety constraints constructed based on risk situation description information. Through the above methods, the reinforcement learning decision-making model can comprehensively weigh multiple driving performance objectives during strategy optimization, while meeting preset safety constraints, thereby achieving multi-objective driving decision optimization. Therefore, to achieve comprehensive optimization of safety, traffic efficiency, and comfort, the reward function is constructed as follows:

[0075] in, Let represent the reward function, which comprehensively characterizes driving safety, driving comfort, and traffic efficiency. , , , All of these represent non-negative weighting coefficients used to balance traffic efficiency, comfort, and safety goals. The weights can be calibrated offline or through scenario playback to ensure a safety-first approach. Indicates at time The gain term is close to the desired speed. , They represent the times respectively. The longitudinal acceleration and impact amplitude are related to comfort and energy consumption. Indicates at time Lane change penalty Indicates at time Lane center deviation penalty.

[0076] At the same time, the overall risk assessment is used as an important safety constraint, thereby constructing a single-step safety cost constraint. Furthermore, it also considers the degree of speeding and illegal lane changing, which is expressed as follows:

[0077] in, Indicates at time The security cost, This means taking the maximum value, i.e., the safety cost is determined by the most dangerous object. Indicates at time Overspeed constraint, this parameter constraint is used to characterize the vehicle at time... The speeding constraint is defined as the degree to which the driving speed exceeds the corresponding road safety speed threshold. In one implementation, this speeding constraint is expressed as a non-negative function that monotonically increases with the degree of speeding. Indicates at time Lane change constraint, which is used to characterize the vehicle's lane change at the current moment. The physical and regulatory feasibility of executing lane-changing behavior is determined by objective conditions such as road structure, traffic rules, or environmental factors. In one implementation method, when lane-changing behavior is not feasible under the current conditions, the lane-changing constraint generates a non-zero constraint quantity to suppress lane-changing behavior that does not conform to the objective conditions.

[0078] Then, the optimization objective of the reinforcement learning decision model can be expressed as maximizing the cumulative expected reward during policy optimization, including the reward function. and safety cost constraints Therefore, it can be represented as:

[0079] in, This represents the optimization objective of maximizing cumulative expectation. Indicates driving decision-making strategy, Represents the expectation operator. This represents the safety cost penalty coefficient, and it increases with increasing... This will increase the conservatism of the strategy.

[0080] Therefore, through the aforementioned reinforcement learning decision-making process based on safety cost constraints, autonomous vehicles can learn and output driving decision strategies that meet safety constraints in dynamic traffic environments, based on the traffic interaction situation and risk uncertainty information reflected in the augmented state representation. This effectively reduces the probability of potential risk events and improves the safety and robustness of autonomous driving decisions.

[0081] Furthermore, reinforcement learning decision-making models can be pre-trained using historical traffic data and perform online reasoning during vehicle operation, thus eliminating the need for complex model updates during vehicle operation and meeting the real-time and stability requirements of practical autonomous driving systems.

[0082] It should be noted that the symbols, variables, and parameters involved in the above embodiments (including but not limited to) , , , , (etc.) are only used to illustrate one specific implementation of the technical solution of the present invention. Their specific physical meaning can be understood by those skilled in the art in combination with the context, and can be equivalently replaced or adjusted according to the specific application scenario without affecting the implementation of the technical solution of the present invention.

[0083] In summary, the autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning proposed in this invention achieves the following technical effects: 1. This invention introduces a spatiotemporal graph network to perform structured modeling of traffic scenarios, and expresses the spatial interaction relationships and temporal evolution characteristics among surrounding traffic participants in a unified manner. This effectively overcomes the problem of information redundancy or missing key information in complex traffic scenarios by traditional vectorized state description methods, and significantly improves the ability of reinforcement learning decision models to perceive complex traffic environments. 2. This invention employs a Bayesian state estimation method based on latent variables to explicitly model the potential behavioral states and uncertainties of surrounding traffic participants. This method can fully characterize the uncertainty features of traffic risks during the decision-making process. Compared with methods that make decisions based solely on deterministic states, this method significantly enhances the ability of autonomous vehicles to perceive and assess potential dangerous scenarios in advance. 3. This invention introduces risk information obtained from risk assessment into the reinforcement learning decision-making process, and imposes risk constraints on driving decisions in the form of constraints or safety costs, so that the reinforcement learning decision-making model can effectively meet the preset safety constraints while optimizing driving efficiency and driving comfort, thereby improving the safety and robustness of the autonomous driving decision-making process. 4. This invention achieves a close coupling between risk perception and decision optimization by constructing an enhanced state representation that integrates traffic interaction situation and risk uncertainty. This avoids the problem of the separation between risk information and decision-making process in existing methods, improves the comprehensive performance and stability of decision-making strategies in complex urban traffic scenarios, and thus forms a closed-loop optimization mechanism from traffic risk perception, risk expression to risk-constrained decision-making. 5. The method of the present invention also has good versatility and scalability, and can be adapted to various complex traffic scenarios such as multi-lane and heterogeneous traffic participants. It can also be combined with different forms of perception systems, graph network structures and reinforcement learning algorithms, and has high engineering application value.

[0084] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

[0085] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. An autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning, characterized in that, Includes the following steps: S1. Obtain the state information of autonomous vehicles and their corresponding surrounding traffic participants, abstract autonomous vehicles and their corresponding surrounding traffic participants into traffic graph nodes in the current traffic scenario, obtain the connection relationships between each node, and generate a traffic state representation that includes the temporal features of the nodes and the connection relationships. S2. Based on traffic state representation, information propagation and clustering are performed on the temporal features of nodes to generate autonomous vehicle node representations that integrate spatial interactions with surrounding traffic participants. S3. Construct the observed traffic state, and at the same time, use the potential behavior or behavioral intentions of surrounding traffic participants as latent variables, perform Bayesian inference, calculate the posterior probability distribution of the latent variables, and obtain a risk situation description of the potential risk level under the current traffic scenario. S4. The node representation of the autonomous vehicle is fused with the risk situation description to generate an enhanced state representation that includes traffic interaction situation information and risk uncertainty information. This enhanced state representation is then used as the state input for reinforcement learning. With the risk situation description as a constraint and the goal of maximizing the cumulative expectation, the optimal driving decision of the autonomous vehicle is achieved.

2. The autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning according to claim 1, characterized in that, The node temporal features are updated recursively using gated cyclic units and are represented as follows: in, Represents a node At any moment The temporal characteristics, This represents a nonlinear state update function. Represents a node At any moment The temporal characteristics, Represents a node At any moment Status information, Represents a node At any moment The longitudinal position of the lane at that time. Represents a node At any moment speed, Represents a node At any moment acceleration, Represents a node At any moment The lane markings at the time. Represents a node At any moment Lateral offset from the lane centerline Represents a node At any moment The constraints on whether lane-changing behavior is allowed. Represents a node At any moment The lateral position of the lane at that time. This refers to autonomous vehicles; the remaining cases involve surrounding traffic participants. This indicates the transpose operation.

3. The autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning according to claim 1, characterized in that, The connection relationship between nodes is at least one of the following: relative distance relationship, relative position relationship, and lane topology relationship.

4. The autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning according to claim 1, characterized in that, Step S2 specifically includes: S21. Using a distance threshold or Top-K method, determine the candidate neighbors of each node and generate a candidate neighbor set for each node. , Represents a node The set of candidate neighbors; S22. Based on the candidate neighbor set, calculate the interaction score between nodes, i.e.: in, Indicates at time node With nodes Interactive ratings between them This represents a function that characterizes the correlation or influence strength of interactions between nodes. Represents a node At any moment The temporal characteristics, Represents a node At any moment The temporal characteristics, Indicates at time node With nodes Relative state information between them Represents a node With nodes The longitudinal relative position between them Represents a node With nodes The horizontal relative position between them Represents a node With nodes The relative velocity between them Represents a node With nodes The relative acceleration between them Represents a node With nodes The spatial relationship of the lanes This indicates the transpose operation. Represents the weight vector. Represents a non-linear activation function. This indicates a splicing operation. These represent the weight matrix and the bias vector, respectively. S23. Based on the interaction score between nodes, select from the candidate neighbor set. The Top-K2 neighbors are selected as the final neighbor set. , Represents a node The final neighbor set, where K2 represents the number of final neighbor sets participating in spatial interaction aggregation; S24. Based on the final neighbor set, map the interaction scores between nodes to interaction weights, that is: in, Represents a node With nodes Interaction weights between them Represents an exponential function. Indicates at time node With nodes Interactive ratings between users; S25. Based on the final neighbor set and combined with interaction weights, the temporal features of all nodes are weighted and aggregated to generate spatial interaction features of the nodes, thus obtaining the autonomous vehicle node representation that integrates the spatial interactions of surrounding traffic participants, i.e.: in, Indicates at time node Spatial interaction characteristics, Indicates at time The set of spatial interaction features of nodes Indicates at time The spatial interaction characteristics of autonomous vehicle nodes are represented by the spatial interactions of surrounding traffic participants. Indicates at time node Spatial interaction characteristics.

5. The autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning according to claim 4, characterized in that, The formula for determining candidate neighbors for each node using a distance threshold is as follows: in, Represents a node Candidate neighbors, Represents a node At any moment The longitudinal position of the lane at that time. Represents a node At any moment The longitudinal position of the lane at that time. This represents the distance threshold.

6. The autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning according to claim 4, characterized in that, Using the Top-K method, the formula for determining the candidate neighbors of each node is as follows: in, Represents a node Candidate neighbors, This indicates the upper limit of the number of candidate neighbors. Represents a node At any moment The longitudinal position of the lane at that time. Represents a node At any moment The longitudinal position of the lane at that time.

7. The autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning according to claim 1, characterized in that, Step S3 specifically includes: S31. Construct observational traffic conditions, namely: in, Indicates the observed traffic conditions. Indicates the time of the autonomous vehicle node Status information, This indicates a splicing operation. Indicates at time The representation of autonomous vehicle nodes that integrates spatial interactions with surrounding traffic participants. This represents the set pooling operator. This represents the final set of neighbors for the autonomous vehicle node. Represents a node At any moment Status information, Indicates at time Autonomous vehicle nodes and nodes Relative state information between them Indicates at time Environmental semantic information; S32. Using the potential behaviors or intentions of surrounding traffic participants as latent variables, and combining this with observed traffic conditions, perform Bayesian inference to calculate the posterior probability distribution of the latent variables, i.e.: in, Representing latent variables The prior probability distribution function, Indicates that given a latent variable Traffic conditions observed under certain conditions The likelihood function, Indicates the traffic conditions observed. Hidden variables The posterior probability distribution function, Indicates the observed traffic conditions The probability distribution function; S33. Based on the posterior probability distribution of latent variables, the neighboring nodes of the autonomous vehicle node are taken as the target node. In the prediction time domain, the relative longitudinal and lateral displacements between the autonomous vehicle node and the target node are calculated. At the same time, combined with the longitudinal and lateral safety envelope half thresholds, a joint safety degree is constructed to obtain the most dangerous value of the collision between the autonomous vehicle and the target node in the prediction time domain, generate the overall risk assessment quantity, and use it as a risk situation description of the potential risk level in the current traffic scenario.

8. The autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning according to claim 7, characterized in that, Step S33 specifically includes: S331. Calculate the predicted longitudinal and lateral positions of the target node at future times, given the latent variables: in, Indicates that given a latent variable target node under conditions In the future The predicted vertical position, Represents the target node At any moment The vertical position, Represents the target node At any moment speed, This represents a continuous-time variable within the prediction time domain. Indicates that given a latent variable target node under conditions acceleration, Indicates that given a latent variable target node under conditions In the future The predicted lateral position, Represents the target node At any moment The lateral position of the lane at that time. Represents the target node The lateral coordinates of the center line of the lane in question. Indicates the total length of the prediction time domain; S332. Based on the predicted longitudinal and lateral positions of the target node at future times under given latent variables, calculate the relative longitudinal and lateral displacements between the autonomous vehicle node and the target node at future times under given latent variables, i.e.: in, Indicates that given a latent variable Under certain conditions, autonomous vehicle nodes and target nodes In the future The relative longitudinal displacement, Indicating autonomous vehicles in the future The vertical position, Indicates that given a latent variable Under certain conditions, autonomous vehicles and target nodes In the future The relative lateral displacement, Indicates the node of the autonomous vehicle at a future time. The lateral position of the lane at that time; S333. Obtain vehicle geometry and construct a collision envelope half-threshold, including parameters for determining the relationship between autonomous vehicle nodes and target nodes. Envelope half threshold for collisions occurring in the longitudinal direction Used to determine the relationship between autonomous vehicle nodes and target nodes. Envelope half threshold for collisions occurring in the lateral direction ,Right now: in, , Representing the autonomous vehicle node and the target node respectively. The length of the car body, , Representing the autonomous vehicle node and the target node respectively. The width of the vehicle body; S334. Determine the relative longitudinal displacement Is it less than or equal to the envelope half threshold? And relative lateral displacement Less than or equal to the envelope half threshold If so, then the autonomous vehicle node and the target node If a collision occurs, proceed to step S335; otherwise, do not cause a collision. S335. Introduce a safety margin and calculate the longitudinal safety envelope half-threshold and the lateral safety envelope half-threshold, i.e.: in, Represents the autonomous vehicle node and the target node The longitudinal safety envelope half threshold at which a collision occurs. Represents the autonomous vehicle node and the target node The lateral safety envelope half threshold at which a collision occurs. , Both represent safety margins; S336. Based on the longitudinal security envelope half-threshold and the lateral security envelope half-threshold, calculate the joint security degree, i.e.: in, Indicates the relationship between autonomous vehicles and target nodes In the future The combined safety level in the event of a collision. This indicates taking the maximum value; S337. Based on the joint security degree, calculate the most dangerous value in the prediction time domain, i.e.: in, Represents the autonomous vehicle node and the target node At any moment The most dangerous value at that time This indicates taking the minimum value; S338. Based on the most dangerous value in the prediction time domain, construct the continuous risk quantity corresponding to the target node, i.e.: in, Represents the target node At any moment Continuous risk quantity Indicates the truncation operator; S339. Based on the continuous risk quantity corresponding to the target node, obtain the maximum continuous risk quantity and generate the overall risk assessment quantity, that is: in, Indicates at time The overall risk assessment quantity, i.e., the risk situation description of the potential risk level.

9. The autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning according to claim 1, characterized in that, Step S4 specifically includes: S41. Integrate the autonomous vehicle node representation with the risk situation description to generate an enhanced state representation that includes traffic interaction situation information and risk uncertainty information, namely: in, Indicates at time The enhanced state representation, This represents the fusion operation function. Indicates at time The spatial interaction characteristics of autonomous vehicle nodes are represented by the spatial interactions of surrounding traffic participants. Indicates at time The overall risk assessment measure, i.e., the risk situation description of the potential risk level, , All of these represent learnable parameters. Indicates a splicing operation; S42. Construct driving behavior actions, namely: in, Indicates at time Driving behavior and actions, Indicates at time Time-discrete longitudinal driving behavior Indicates at time Time-discrete lateral driving behavior Indicates at time Time-discrete lane-changing behavior; S43. Based on the enhanced state representation and driving behavior actions, construct a reward function, namely: in, Represents the reward function, , , , All represent non-negative weighting coefficients. Indicates at time The gain term is close to the desired speed. , They represent the times respectively. longitudinal acceleration and impact amplitude Indicates at time Lane change penalty Indicates at time Lane center deviation penalty; S44. Based on the risk situation description, construct security cost constraints, namely: in, Indicates at time The security cost, This indicates taking the maximum value. Indicates at time Overspeed constraints Indicates at time Lane changing constraints; S45. Based on the reward function and safety cost constraints, construct an optimization objective that maximizes the cumulative expected value; S46. When the accumulated expectation is maximized, the corresponding driving behavior action is taken as the optimal driving decision strategy for the autonomous vehicle, so as to achieve the optimal driving decision for the autonomous vehicle.

10. The autonomous driving decision-making method based on spatiotemporal graph networks and risk perception reinforcement learning according to claim 9, characterized in that, The optimization objective for maximizing cumulative expectation is: in, This represents the optimization objective of maximizing cumulative expectation. Indicates driving decision-making strategy, Represents the expectation operator. This represents the penalty coefficient for safety costs.