An implicit state representation method for an autonomous driving system at signal-free intersections

By introducing an implicit state representation method in the autonomous driving system, using lane keeping signs and target speed to simplify state information, the problem of low decision efficiency in complex traffic environments without signal intersections is solved, and more efficient calculation and training is achieved.

CN119808880BActive Publication Date: 2025-06-17CHANGCHUN UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510279839.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-17
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

When dealing with complex traffic environments such as signalless intersections, autonomous driving technology faces the challenge of efficient decision-making, resulting in high consumption of computing resources and affecting training efficiency and response speed.

Method used

An implicit state representation method based on a signalless intersection automatic driving system is proposed, which implicitly implicits the state information of surrounding traffic participants and road environments by introducing discrete lane-keeping signs and continuous target speeds, and simplifies policy network input.

Benefits of technology

It significantly reduces the state space dimension, improves computing efficiency and training speed, enhances the robustness of the system in dynamic traffic environments, and ensures efficient decision-making of autonomous vehicles in complex traffic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808880B_ABST
    Figure CN119808880B_ABST
Patent Text Reader

Abstract

The present invention relates to an implicit state representation method for an autonomous driving system at a signal-free intersection, aiming to simplify the state space representation and improve the system training efficiency and robustness. This method introduces a lane-keeping flag and a target speed as implicit state variables to represent environmental information independent of the ego vehicle, thereby reducing the state space dimension and avoiding the computational complexity caused by considering too many surrounding traffic participants and complex environmental information. In the absence of surrounding traffic participants, the target speed is set according to traffic rules and road conditions, while in the presence of surrounding traffic participants, the target speed is dynamically adjusted based on the relative position and speed with respect to the surrounding vehicles. The lane-keeping flag determines whether to change lanes or maintain the lane based on the current position of the ego vehicle and road information. This method effectively simplifies the state space, improves the reinforcement learning training efficiency, and can better handle autonomous driving tasks in complex environments such as signal-free intersections.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention proposes an implicit state representation method for an autonomous driving system at unsignalized intersections, which relates to the technical field of simplifying the input of a policy network for autonomous driving vehicles based on deep reinforcement learning. Background Art

[0002] Autonomous driving technology faces many challenges when dealing with complex traffic environments such as unsignalized intersections, and efficient decision-making is required to ensure driving safety and stability. Traditional methods usually rely on vehicle dynamic information and surrounding traffic states for decision-making. Although multiple factors are considered, as the environmental complexity increases, the state space dimension continuously expands, resulting in high consumption of computing resources, which in turn affects the training efficiency and response speed. To solve this problem, the present invention proposes an implicit state representation method for an autonomous driving system at unsignalized intersections. By introducing discrete lane-keeping signs and continuous target speeds, state information such as surrounding traffic participants and road environments is implicitized, reducing the direct dependence on the states of complex traffic participants. Compared with traditional methods, the present invention significantly reduces the state space dimension while ensuring the decision-making effect, simplifies the input of the policy network, and improves the computing efficiency and training speed. This method is particularly applicable to complex traffic scenarios such as unsignalized intersections and is of great significance for promoting the application and development of autonomous driving technology. Summary of the Invention

[0003] To solve the above problems, the present invention proposes an autonomous driving system method based on implicit state representation. This method transforms the state information of the vehicle and the complex information of the surrounding environment into a simplified implicit state representation, reducing the dimension of the state space, thereby simplifying the input of the policy network and improving the computing and training efficiency. Especially in a changeable and uncertain traffic environment, the implicit state representation method demonstrates its high robustness and superiority.

[0004] An implicit state representation method for an autonomous driving system at unsignalized intersections according to the present invention includes the following specific related steps:

[0005] Step S1: Define the implicit state space;

[0006] Step S2: Calculate two variable parameters of the implicit state representation: the lane-keeping sign and the target speed;

[0007] Step S3: Optimize the driving policy based on the implicit state representation and design a reward function.

[0008] Step S1 specifically includes the following content:

[0009] Step S11: Define a series of state variables to describe the state of the autonomous vehicle itself, the road environment, and the information of surrounding traffic participants. The state variables are divided into three different state spaces: the ego-vehicle state space , the environmental state space and the implicit state representation space ;

[0010] Step S12: Represents the ego-vehicle state space, which is used to describe the state information related to the ego-vehicle itself, including speed, position, and attitude parameters, that is, the longitudinal speed of the vehicle and the lateral speed , the yaw angle of the vehicle , the slip rates of the four tires in the front, rear, left, and right directions ;

[0011] Step S13: Represents the environmental state space, which is used to describe the parameters included in the state space of surrounding traffic participants and the road environment, including lane markings The identifier indicating lane information, the lateral offset Indicates the vertical distance of the ego-vehicle from the center line of the lane in the horizontal direction, and the lateral offset change rate Indicates the speed at which the vehicle deviates from the lane;

[0012] Step S14: Represents the implicit state representation space, which includes two implicit state variable parameters, the lane keeping flag and the target speed , where is used to implicitly represent whether the vehicle should maintain the current lane or perform a lane change operation during autonomous driving, is used to implicitly represent the environmental information related to the vehicle position, surrounding traffic participants, and road environment, etc.;

[0013] Step S15: Represents the implicit state space, which is used to calculate the implicit state variables: and , which describe the relative position and speed between the ego-vehicle and surrounding traffic participants, including: the longitudinal distances between the ego-vehicle and the vehicles in front and behind in the same lane and the left and right adjacent lanes are respectively and the corresponding speeds are respectively , the distance between the ego-vehicle and the nearest traffic participant within the signal-free intersection and the corresponding speed , the distance between the entrance lane of the signal-free intersection and the crosswalk , the longitudinal safety distance , the lane width , the lateral distance from the vehicle to the center line of the target lane .

[0014] Step S2 specifically includes the following content:

[0015] Step S21: The implicit state representation determines the behavior of the vehicle through and and is a discrete value, taking values of 0 or 1, representing the following two situations: when , it means the vehicle stays in the current lane, and when , it means the vehicle changes lanes. is a continuous value, which is set based on different scenarios and conditions in the autonomous driving task at an intersection without signals;

[0016] Step S22: In the scenario without surrounding traffic participants, changes as follows: When the lateral offset of the vehicle , in meters, and the yaw angle , and the lateral distance from the vehicle to the center line of the target lane is greater than half of the lane width, that is , at this time , it means the vehicle is changing lanes. When the vehicle crosses the boundary of the current lane, it means the lane change is completed. At this time , until the next lane change task starts. When inside the intersection without signals, the vehicle needs to drive along the desired path generated by the Bezier curve and enter the left lane. At this time , it means the vehicle follows the target path and does not change lanes;

[0017] Step S23: In the scenario without surrounding traffic participants, is mainly determined by factors such as road environment, lane position, and the position of the intersection without signals. The calculation rule of is as follows: On the straight lane, the vehicle in the outermost lane is 10m / s, and the vehicle in other lanes is 15m / s. Inside the intersection without signals, to ensure safety, the vehicle will be reduced to 10m / s. When the distance from the vehicle to the intersection without signals , in meters, at this time ;

[0018] Step S24: In the scenario with surrounding traffic participants, changes as follows: When the longitudinal relative distance between the vehicle and the vehicle in front in the same lane is greater than the longitudinal safety distance, that is , and the current speed of the vehicle is less than the target speed, that is , at this time indicates that the vehicle changes lanes. When the vehicle is in the left-turn lane or at an intersection without signals, at this time indicates that the vehicle must stay in the current lane and drive along the predetermined route;

[0019] Step S25: In the scenario with surrounding traffic participants, is determined by the relative distance, speed between the host vehicle and surrounding traffic participants, and traffic rules. The calculation rule is as follows: When the longitudinal relative distance between the host vehicle and the vehicle in front in the same lane is greater than the safety distance, that is , where When the longitudinal relative distance between the host vehicle and the vehicle behind in the left lane is less than the safety distance, that is , When the vehicle is in the left-turn lane, at this time When the vehicle is inside the intersection without signals, at this time .

[0020] Step S3 specifically includes the following content:

[0021] Step S31: Design a state-related reward function which is composed of five parts: speed reward , lane-keeping reward , lane-changing reward , yaw angle reward and steering reward ;

[0022] Step S32: is a penalty term related to and . By measuring the gap between and and to encourage the vehicle to stay within the desired speed range. The expression of is as follows:

[0023]

[0024] Among them, is the longitudinal speed of the current vehicle, is the absolute difference between the two. When is closer to , the reward value is greater. When the reward approaches 0, it means there is no penalty. When is farther from , the reward value is smaller, and the greater the absolute difference, the greater the penalty. The system guides the vehicle to adjust its speed through penalties, thereby reducing the speed deviation;

[0025] Step S33: For evaluating whether the vehicle stays in the center of the lane and preventing the vehicle from deviating from the lane, the expression is as follows:

[0026]

[0027] where, is the lateral offset of the vehicle from the center line of the lane, is the change rate of the lateral offset, is the width of the lane. When it gives the maximum reward , indicating that the vehicle is at the center of the lane. When it is , indicating that the vehicle is reducing the lateral offset and returning to the center of the lane. When the lateral offset is large, gradually decreases. When the reward is negative, the vehicle is encouraged to immediately take measures to reduce its lateral offset from the center line of the lane to correct the lane departure;

[0028] Step S34: For encouraging the vehicle to perform a lane change operation under appropriate circumstances, the expression is as follows:

[0029]

[0030] When it gives the maximum reward , indicating that the vehicle is changing lanes. When the during the lane change process is large, then changes according to to encourage the vehicle to complete the lane change operation quickly and accurately;

[0031] Step S35: For evaluating the steering stability of the vehicle, the expression is as follows:

[0032]

[0033] where, is the yaw angle of the vehicle, representing the steering angle of the vehicle;

[0034] Step S36: Is the sum of the steering reward for the straight lane and the steering reward for the intersection without signals , that is . On the straight lane, to ensure driving comfort and conform to actual driving habits, the vehicle should maintain a small steering wheel angle , at an intersection without signals, when performing a left-turn task, the vehicle is required to reach a yaw angle when passing through the intersection without signals to successfully complete the turn, and The expression is as follows:

[0035]

[0036]

[0037] The total state-related reward function is the sum of all sub-rewards , , , , , and its expression is as follows:

[0038]

[0039] Among them, the coefficient represents the weight coefficient, which is used to determine the relative weight of each sub-reward item in the total reward calculation;

[0040] Step S37: Initialize a policy network Actor and a value network Critic. The policy network is a deep neural network that can decide the best action according to the current implicit state representation. Through continuous training and learning, its decision-making ability in different states is improved and optimized, so that the policy network can select those actions that can bring higher cumulative rewards. The value network is used to evaluate the long-term expected return of the current policy. The value network helps to adjust the policy network by selecting actions that may lead to higher cumulative rewards to adjust the policy;

[0041] Step S38: Execute the actions of the policy network in the simulation environment or the actual vehicle, and collect the implicit state representation and the corresponding rewards , , , and , and use this data to train and optimize the policy and value networks. The Advantage Actor-Critic algorithm is used to train the policy network. The A2C algorithm uses the advantage function to accelerate the convergence speed of the network, so that it can select the optimal action according to the implicit state representation. A2C is a parallel synchronous method based on the advantage function, which is used to evaluate the output of the network and reduce the estimation variance under the policy gradient. The expression of its advantage function is as follows:

[0042]

[0043] Among them, is under the policy Under the state Take an action The action-value function and the state-value function Indicates that under the policy From the state The expected cumulative return starting from Indicates the expected cumulative return starting from the state Starting from Is the immediate reward obtained after taking the action Under the state After that Is the discount factor;

[0044] The expression of the policy gradient is shown as follows:

[0045] J ( ) = E ~ [ t = 0 T A ? ( s t , a t ) ] ? log ( a t | s t )

[0046] Among them, Indicates the expectation of the trajectory Sampled according to the policy Of Indicates the length of the trajectory, Indicates the probability of taking an action Under the given state Using the parameter To represent the policy, Indicates the state Take an action Of the policy The logarithmic probability of With respect to the parameter Gradient, indicating the rate of change of the policy with respect to the parameter Indicates the state Take an action Of the advantage function, which depends on the parameter And adjusts the policy by evaluating the greater cumulative reward of the action in the current state.

[0047] The beneficial effects of the present invention are:

[0048] 1. A method for implicit state representation based on an autonomous driving system at unsignalized intersections is proposed, specifically targeting scenarios with complex state spaces at unsignalized intersections. This method converts the vehicle's position, road environment, and the state information of surrounding traffic participants into two variable parameters of the implicit state: a discrete lane-keeping flag and a continuous target speed. Such a conversion not only simplifies the dimensionality of the state space but also enhances the system's robustness in a dynamic traffic environment by reducing the direct dependence on the complex state of traffic participants. In addition, the simplified state representation also reduces the computational complexity, thereby improving the training efficiency of the autonomous driving system when dealing with complex traffic scenarios such as unsignalized intersections.

[0049] 2. A state-dependent reward function is proposed, which can guide the vehicle to optimize its actions according to the vehicle's motion state. This reward mechanism not only improves the stability and safety of the autonomous driving system but also enhances the driving efficiency, enabling the autonomous driving vehicle to make more accurate and efficient decisions when facing complex traffic challenges. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 Scenarios of implicit state representation in the absence of surrounding traffic participants;

[0051] Figure 2 Scenarios of implicit state representation on a straight lane in the presence of surrounding traffic participants;

[0052] Figure 3 Scenarios of implicit state representation within an unsignalized intersection in the presence of surrounding traffic participants;

[0053] Figure 4 For the yaw angle of the vehicle when making a steering decision and the steering wheel angle Schematic diagram of the relationship between;

[0054] Figure 5 Frame diagram of the autonomous driving strategy at unsignalized intersections. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] To further understand the content of the present invention, the present invention will be described in detail in combination with the drawings and embodiments. It should be understood that the embodiments are only for explaining the present invention and not for limiting it.

[0056] The specific implementation steps are as follows:

[0057] Step S1: Define the implicit state space;

[0058] Step S2: Calculate the two variable parameters of the implicit state representation: the lane-keeping flag and the target speed;

[0059] Step S3: Optimize the driving strategy based on the implicit state representation and design the reward function.

[0060] In step S1, in order to define the implicit state space, the following steps are required:

[0061] Step S11: Define a series of state variables to describe the state of the autonomous vehicle itself, the road environment, and the information of surrounding traffic participants, and divide the state variables into three different state spaces: the ego-vehicle state space , the environmental state space and the implicit state representation space ;

[0062] Step S12: represents the ego-vehicle state space, which is used to describe the state information related to the ego-vehicle itself, including speed, position, and attitude parameters, that is, the longitudinal speed and the lateral speed , the yaw angle of the vehicle, and the slip rates of the four tires, front, rear, left, and right ;

[0063] Step S13: represents the environmental state space, which is used to describe the parameters included in the state space of surrounding traffic participants and the road environment, including lane markings an identifier indicating lane information, the lateral offset indicating the vertical distance of the ego-vehicle from the center line of the lane in the horizontal direction, and the lateral offset rate of change indicating the speed at which the vehicle deviates from the lane;

[0064] Step S14: represents the implicit state representation space, which includes two implicit state variable parameters, the lane keeping flag and the target speed , where is used to implicitly represent whether the vehicle should maintain the current lane or perform a lane change operation during autonomous driving, is used to implicitly represent the environmental information related to the vehicle's position, surrounding traffic participants, and road environment, etc.;

[0065] Step S15: represents the implicit state space, which is used to calculate the implicit state variables: and , which describe the relative position and speed of the ego-vehicle and surrounding traffic participants, including: the longitudinal distances between the ego-vehicle and the vehicles in front and behind in the same lane and the left and right adjacent lanes are and the corresponding speeds are , and the distance between the ego-vehicle and the nearest traffic participant within the signal-free intersection and the corresponding speed is , the distance between the entrance lane of the intersection without signal and the crosswalk , the longitudinal safety distance , the lane width , the lateral distance from the host vehicle to the center line of the target lane .

[0066] In step S2, calculate the lane keeping flag and the target speed, which specifically includes the following steps:

[0067] Step S21: The implicit state representation determines the behavior of the host vehicle through and . is a discrete value, taking values of 0 or 1, representing the following two situations: when , it means that the host vehicle stays in the current lane; when , it means that the host vehicle changes lanes. is a continuous value, which is set based on different scenarios and conditions of the host vehicle in the autonomous driving task at the intersection without signal;

[0068] Step S22: The scenario without surrounding traffic participants, changes as follows: when the lateral offset of the vehicle , in meters, the yaw angle , and the lateral distance from the host vehicle to the center line of the target lane is greater than half of the lane width, that is , at this time , it means that the vehicle is changing lanes. When the vehicle crosses the boundary of the current lane, it means that the lane change is completed. At this time , until the next lane change task starts. When inside the intersection without signal, the vehicle needs to drive along the desired path generated by the Bezier curve and enter the left lane. At this time , it means that the vehicle follows the target path and does not change lanes;

[0069] Step S23: The scenario without surrounding traffic participants, is mainly determined by factors such as the road environment, lane position, and the position of the intersection without signal. The calculation rule of is as follows: on the straight lane, for the vehicle in the outermost lane is 10m / s, and for the vehicles in other lanes is 15m / s. Inside the intersection without signal, to ensure safety, the vehicle will be reduced to 10m / s. When the distance of the vehicle from the intersection without signal Figure 1 , in meters, as shows, at this time ;

[0070] Step S24: In the scenario with surrounding traffic participants, the changes are as follows: When the longitudinal relative distance between the host vehicle and the vehicle in front in the same lane is greater than the longitudinal safety distance, i.e., , and the current vehicle speed is less than the target speed, i.e., , at this time , it indicates that the vehicle changes lanes. When the vehicle is in the left-turn lane or at an intersection without signals, at this time , it indicates that the vehicle must stay in the current lane and drive along the predetermined route;

[0071] Step S25: In the scenario with surrounding traffic participants, is determined by the relative distance, speed between the host vehicle and surrounding traffic participants, and traffic rules. The calculation rules are as follows: When the longitudinal relative distance between the host vehicle and the vehicle in front in the same lane is greater than the safety distance, i.e., , , where , when the longitudinal relative distance between the host vehicle and the vehicle behind in the left lane is less than the safety distance, i.e., , , when the vehicle is in the left-turn lane, at this time , as Figure 2 shows, in the straight section, Figure 2 the number 1 in represents the vehicle in front of the host vehicle in the same lane, the number 2 represents the vehicle in front of the host vehicle in the left lane, the number 3 represents the vehicle behind the host vehicle in the left lane, the number 4 represents the vehicle in front of the host vehicle in the right lane, the number 5 represents the vehicle behind the host vehicle in the right lane. When the vehicle is in the intersection without signals, at this time Figure 3 , as Figure 3 shows,

[0072] In step S3, a state-related reward function is designed, which specifically includes the following steps:

[0073] Step S31: Design a state-related reward function , which is composed of five parts: speed reward , lane-keeping reward , lane-changing reward , yaw angle reward and steering reward ;

[0074] Step S32: is a penalty term related to and , by measuring and The gap between them is used to encourage the vehicle to maintain within the desired speed range. The expression of

[0075]

[0076] is as follows: where is the longitudinal speed of the current vehicle, is the absolute difference between the two. When is closer to the reward value is larger. When the reward approaches 0, it means there is no penalty. When is farther from

[0077] Step S33: is used to evaluate whether the vehicle maintains at the center of the lane and avoid the vehicle deviating from the lane. The expression of

[0078]

[0079] is as follows: where is the lateral offset of the vehicle from the center line of the lane, is the change rate of the lateral offset, is the width of the lane. When the maximum reward is given indicating that the vehicle is at the center of the lane. When it means that the vehicle is reducing the lateral offset and returning to the center of the lane. When the lateral offset is large, gradually decreases. When the reward is negative, the vehicle is encouraged to immediately take measures to reduce its lateral offset from the center line of the lane to correct the lane deviation.

[0080] Step S34: is used to encourage the vehicle to perform a lane change operation under appropriate circumstances. The expression of

[0081]

[0082] is as follows: When the maximum reward is given indicating that the vehicle is performing a lane change. When during the lane change process is large, then varies according to the change of to encourage the vehicle to complete the lane change operation quickly and accurately.

[0083] Step S35: For evaluating the steering stability of the vehicle, the expression is as follows:

[0084]

[0085] wherein, is the yaw angle of the vehicle, representing the steering angle of the vehicle;

[0086] Step S36: is the steering reward for the straight lane and the steering reward for the intersection without signals The sum, that is On the straight lane, in order to ensure driving comfort and conform to actual driving habits, the vehicle should maintain a small steering wheel angle At the intersection without signals, when performing a left-turn task, it is required that the vehicle reach the yaw angle when passing through the intersection without signals to complete the turn smoothly. The relationship diagram between the yaw angle of the vehicle and the steering wheel angle is as shown in Figure 4 as follows, and The expressions of are as follows respectively:

[0087]

[0088]

[0089] The total state-related reward function is the sum of all sub-rewards , , , , The sum of, and its expression is as follows:

[0090]

[0091] wherein, the coefficient represents the weight coefficient, which is used to determine the relative weight of each sub-reward item in the total reward calculation;

[0092] Step S37: Initialize a policy network Actor and a value network Critic. The policy network is a deep neural network that can determine the best action based on the current implicit state representation. Through continuous training and learning, its decision-making ability in different states is improved and optimized, enabling the policy network to select actions that can bring higher cumulative rewards. The value network is used to evaluate the long-term expected return of the current policy. The value network helps adjust the policy network by selecting actions that may lead to higher cumulative rewards to adjust the policy, as Figure 5 shown in the self-driving policy framework diagram for signal-free intersections, which is used to describe the structure and learning process of the network, Figure 5 where the action input in includes four-wheel torque and steering wheel angle , which directly act on the vehicle model. SUMO is used as a traffic simulation environment to generate the state information of other vehicles and traffic participants. The gradient is calculated through the reward function to update the policy network and the value network, represents the state space, including the ego-vehicle state space , the environmental state space , and the implicit state representation space

[0093] Step S38: Execute the actions of the policy network in the simulation environment or on the actual vehicle, and collect the implicit state representation and the corresponding rewards , , , , and . Use this data to train and optimize the policy and value networks. The Advantage Actor-Critic algorithm is used to train the policy network. The A2C algorithm uses the advantage function to accelerate the convergence speed of the network, enabling it to select the optimal action based on the implicit state representation. A2C is a parallel synchronous method based on the advantage function, which is used to evaluate the output of the network and reduce the estimation variance under the policy gradient. The expression of its advantage function is shown as follows:

[0094]

[0095] where is the action value function of taking action in state under policy . The state value function represents the expected cumulative return starting from state under policy . represents the expected cumulative return starting from state . is the immediate reward obtained after taking an action in state ; is the discount factor;

[0096] The expression of the policy gradient is shown as follows:

[0097] J ( ) = E ~ [ t = 0 T A ? ( s t , a t ) ] ? log ( a t | s t )

[0098] where represents the expectation of the trajectory sampled according to the policy ; represents the length of the trajectory, represents the probability of taking an action in the given state , using the parameter to represent the policy, represents the logarithm probability of the policy of taking an action in the state with respect to the parameter , representing the rate of change of the policy with respect to the parameter ; represents the advantage function of taking an action in the state , depending on the parameter , and adjusting the policy by evaluating the greater cumulative reward of the action in the current state.

Claims

1. An implicit state representation method based on an automatic driving system at an unsignalized intersection, characterized in that: The method comprises the following steps: Step S1: define the implicit state space; Step S2: Calculate two variable parameters of implicit state representation: lane keeping sign and target speed; Step S3: Optimize the driving strategy based on the implicit state representation and design the reward function; In step S1, the state space is defined and the following processing is required: Step S11: A series of state variables are defined to describe the state of the autonomous vehicle itself, the road environment, and the information of surrounding traffic participants. The state variables are divided into three different state spaces: the vehicle state space S ego , environmental state space S else and the implicit state representation space S rep ; Step S12: ego Represents the state space of the vehicle, which is used to describe the state information related to the vehicle itself, including speed, position and posture parameters, that is, the longitudinal speed V of the vehicle x and the lateral velocity V y , the vehicle's yaw angle r, the slip rate of the four tires λ fl , fr , rl , rr ; Step S13: else Represents the environment state space, which is used to describe the parameters contained in the surrounding traffic participants and road environment state space, including lane markings S l Identifier indicating lane information, lateral offset Y l Indicates the vertical distance from the vehicle to the lane centerline in the horizontal direction and the lateral offset change rate Indicates the speed at which the vehicle deviates from the lane; Step S14: rep represents the implicit state representation space, including two implicit state variable parameters, lane keeping sign S lk and target speed V * , where S lk It is used to implicitly indicate whether the vehicle should maintain the current lane or perform lane change during autonomous driving. * Used to implicitly represent information related to the vehicle's position, surrounding traffic participants, and road environment; Step S15: imp Represents the implicit state space, which is used to calculate the implicit state variables: S lk and V * , describes the relative position and speed of the vehicle and the surrounding traffic participants, including: the longitudinal distances between the vehicle and the vehicles in the same lane and the left and right lanes are d1, d2, d3, d4, d5 and the corresponding speeds are V1, V2, V3, V4, V5, the distance d6 between the vehicle and the nearest traffic participant in the unsignalized intersection and the corresponding speed V6, the distance d7 between the entrance lane of the unsignalized intersection and the crosswalk, the longitudinal safety distance d8, the lane width d9, the lateral distance d from the vehicle to the center line of the target lane 10 ; Step S2 calculates the two variable parameters S represented by the implicit state lk and V * , the specific process is as follows: Step S21: Implicit state representation through S lk and V * To determine the behavior of the vehicle, S lk is a discrete value, which takes the value of 0 or 1, representing the following two situations: lk =1, indicating that the vehicle stays in the current lane. lk = 0, indicating that the vehicle is changing lanes, V * It is a continuous value, which is set based on different scenarios and conditions of the vehicle in the unsignalized intersection autonomous driving task; Step S22: Scene without surrounding traffic participants, S lk The changes are as follows: When the vehicle's lateral displacement Y l <0.2, in meters, the yaw angle r<2°, and the lateral distance from the vehicle to the center line of the target lane is greater than half the lane width, that is, At this time S lk = 0, indicating that the vehicle is changing lanes. When the vehicle crosses the boundary of the current lane, it indicates that the lane change is completed. At this time, S lk = 1 until the next lane change task starts. When in an unsignaled intersection, the vehicle needs to follow the expected path generated by the Bezier curve to enter the left lane. At this time, S lk =1, indicating that the vehicle follows the target path and does not change lanes; Step S23: Scene without surrounding traffic participants, V * Determined by the road environment, lane position and unsignaled intersection location factors, V * The calculation rules are as follows: On the straight lane, the vehicle V in the outermost lane * is 10m / s, and the vehicles in other lanes V * is 15m / s. In order to ensure safety at non-signaled intersections, the vehicle V * It will drop to 10m / s. When the distance d7 between the vehicle and the non-signaled intersection is less than 50 meters, V * is a linear function, that is, V * =0.2d7+5; Step S24: Scene with surrounding traffic participants, S lk The changes of are as follows: When the longitudinal relative distance between the vehicle and the front vehicle in the same lane is greater than the longitudinal safety distance, that is, d1>d8, and the current speed of the vehicle is less than the target speed, that is, V x <V * , at this time S lk =0, indicating that the vehicle is changing lanes. When the vehicle is in the left turn lane or at an intersection without a signal, S lk =1, indicating that the vehicle must stay in the current lane and drive along the predetermined route; Step S25: Scene with surrounding traffic participants, V * Determined by the relative distance between the vehicle and surrounding traffic participants, speed, and traffic rules, V * The calculation rule is as follows: When the longitudinal relative distance between the vehicle and the vehicle in front of it in the same lane is greater than the safety distance, that is, d1>d8, V * =V m -5, where V m =min(V1,V2), when the longitudinal relative distance between the vehicle and the vehicle behind in the left lane is less than the safety distance, that is, d3<d8, V * = V3-5, when the vehicle is in the left turn lane, V * =V1+d1-2, when the vehicle is in a non-signal intersection, V * =d6-2.

2. The implicit state representation method based on an automatic driving system at an unsignalized intersection according to claim 1, characterized in that: The step S3 designs a state-dependent reward function R SRR And through the reinforcement learning algorithm, the vehicle's lane keeping, lane changing and unprotected left turn behaviors in complex traffic environments are optimized. The specific implementation process is as follows: Step S31: Design a state-dependent reward function R SRR , the reward function consists of five parts: speed reward R v , Lane Keeping Reward R lk , Lane Change Reward R lc , yaw angle reward R yaw and the steering reward R st ; Step S32: R v Is with V x and V * The relevant penalty term, R v By measuring V x and V * The gap between the two encourages the vehicle to stay within the desired speed range, R v The expression is as follows: R v =-|V x -V * | Among them, V x is the current longitudinal velocity of the vehicle, |V x -V * | is the absolute difference between the two, when V x The closer to V * , the larger the reward value, the closer the reward is to 0, indicating no penalty. x V * The farther away, the smaller the reward value, the larger the absolute difference, and the greater the penalty. The system guides the vehicle to adjust its speed through penalties, thereby reducing speed deviation; Step S33: R lk Used to assess whether the vehicle is staying in the center of the lane and prevent the vehicle from deviating from the lane. lk The expression is as follows: Among them, Y l is the lateral offset of the vehicle from the centerline of the lane, is the rate of change of lateral offset, d9 is the width of the lane, when |Y l When |<0.1, the maximum reward R is given lk =1, indicating that the vehicle is in the center of the lane. When R lk = 0.9, indicating that the vehicle is reducing the lateral offset and returning to the center of the lane. l When R is larger, lk Gradually decreases, and when the reward is negative, the vehicle is encouraged to take immediate measures to reduce its lateral deviation from the lane centerline to correct the lane departure; Step S34: R lc Used to encourage vehicles to perform lane change maneuvers under appropriate circumstances, R lc The expression is as follows: when When , the maximum reward R is given lk =1, indicating that the vehicle is changing lanes. l When it is larger, R lc According to Y l Changes according to the changes in the lane, encouraging vehicles to complete lane changes quickly and accurately; Step S35: R yaw Used to evaluate the vehicle's steering stability, R yaw The expression is as follows: R yaw =-|r| Among them, r is the yaw angle of the vehicle, indicating the steering angle of the vehicle; Step S36: R st The turning reward R for the straight lane st1 Turn reward R at unsignalized intersection st2 The sum of R st =R st1 +R st2 , in order to ensure driving comfort and conform to actual driving habits, the vehicle should maintain a small steering wheel angle δ f , at an unsignalized intersection, executing a left turn task requires the vehicle to reach a yaw angle of r = -90 when passing through the unsignalized intersection to successfully complete the turn, R st1 and R st2 The expression is as follows: R st1 =-abs(δ f ) / 180 The total state-dependent reward function R SRR is all sub-rewards R v , R lk , R lc , R yaw , R st The sum of is expressed as follows: R SRR =c1*R v +c2*R lk +c3*R lc +c4*R yaw +c5*R st Among them, coefficients c1, c2, c3, c4, and c5 represent weight coefficients, which are used to determine the relative weight of each sub-reward item in the calculation of the total reward; Step S37: Initialize a policy network Actor and a value network Critic. The policy network is a deep neural network that can determine the best action based on the current implicit state representation. Through continuous training and learning, it improves and optimizes its decision-making ability in different states, so that the policy network can select actions that can bring higher cumulative rewards. The value network is used to evaluate the long-term expected return of the current strategy. The value network helps adjust the policy network and adjusts the strategy by selecting actions that lead to higher cumulative rewards. Step S38: Execute the actions of the policy network in the simulation environment or in the actual vehicle, and collect the implicit state representation and the corresponding reward R v , R lk , R lc , R yaw and R st , use these data to train and optimize the policy and value networks, use the AdvantageActor-Critic algorithm to train the policy network, and the A2C algorithm uses the advantage function to speed up the convergence of the network so that it can select the optimal action based on the implicit state representation. A2C is a parallel synchronization method based on the advantage function, which is used to evaluate the output of the network and reduce the estimated variance under the policy gradient. The expression of its advantage function is shown in the following formula: Among them, Q π (s t ,a t ) is under policy π, state s t Take action a t The action value function and state value function V π (s t ) means that under the strategy π, from state s t The expected cumulative return from departure, V π (s t+1 ) indicates that from state s t+1 The expected cumulative return from departure, R t is in state s t Take action a t The immediate reward obtained after , γ is the discount factor; The expression of policy gradient is as follows: in, According to the strategy π θ The expectation of the sampled trajectory τ, T represents the length of the trajectory, π θ (a t |s t ) means that in a given state s t Take action a t The probability of using the parameter θ to represent the strategy, Indicates that in state s t Take action a t The gradient of the logarithmic probability of the strategy π with respect to the parameter θ represents the rate of change of the strategy with respect to the parameter θ, A φ (s t ,a t ) means in state s t Take action a t The advantage function of , which depends on the parameter φ, adjusts the policy by evaluating the larger cumulative reward of the action in the current state.

Citation Information

Patent Citations

  • Automatic driving planning control method based on world model

    CN117872909A

  • Automatic driving non-signalized intersection decision generation method based on hierarchical reinforcement learning

    CN118790290A

  • Automatic driving training method, apparatus and device, and medium

    WO2022052406A1