Intersection driving decision-making method, related device, equipment and storage medium

By using covariance matrix in the intelligent driving system to model the future position uncertainty of traffic participants and analyze it in combination with the planned path of the bicycle, the problem of difficulty in dealing with uncertainty in the unprotected intersection scenario is solved, and a more reasonable and safe driving decision is achieved.

CN119992821APending Publication Date: 2025-05-13ZHEJIANG LEAPMOTOR TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411996590.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing intelligent driving technology is difficult to effectively model and deal with uncertainty in unprotected intersection scenarios, resulting in unreasonable actions in dynamic environments and it is difficult to provide guidance information for motion planning.

Method used

By obtaining the covariance matrix of traffic participants outside the bicycle at the intersection to be decided at the current moment, the prediction error of the traffic participant's future position is characterized, and based on this prediction, the position range of traffic participants at the next moment is obtained. Based on the planned path of the bicycle and the location range of traffic participants at the next moment, we analyze it to determine whether there is a collision risk, so as to determine the decision information of the bicycle, including the expected acceleration.

Benefits of technology

Through uncertainty modeling, the rationality of driving decisions at intersections to be decided is improved, the collision risk can be judged more accurately and appropriate decisions can be made, and the operational safety and efficiency of smart cars in dynamic environments is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992821A_ABST
    Figure CN119992821A_ABST
Patent Text Reader

Abstract

The invention discloses an intersection driving decision method, a related device, equipment and a storage medium. The method comprises the following steps: acquiring a covariance matrix of traffic participants except an own vehicle at a to-be-decided intersection at the current moment; wherein the covariance matrix represents prediction errors of future positions of the traffic participants; performing prediction based on the covariance matrix of the traffic participant at the current moment to obtain a position range of the traffic participant at the next moment; based on the planned path of the vehicle and the position range of the traffic participant at the next moment, analysis is carried out to obtain an analysis result, and the analysis result comprises whether there is a collision risk between the vehicle and the traffic participant at the next moment at the intersection to be decided; based on whether the analysis result includes the collision risk, obtaining decision information of the own vehicle; wherein the decision information at least comprises the expected acceleration. According to the scheme, uncertainty modeling can be carried out at the to-be-decided intersection, so that the reasonability of driving decision making at the to-be-decided intersection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent driving technology, and in particular to a method for making decisions on driving at intersections and related devices, equipment and storage media. Background Art

[0002] With the development of intelligent driving technology, AVP (Automated Valet Parking) has attracted more and more attention. In the actual parking environment, smart cars will encounter various driving scenarios. Among them, one of the most typical scenarios is the intersection scenario, which is an unprotected type without traffic light guidance, resulting in a lot of uncertainty in this scenario.

[0003] Existing decision-making technologies ignore the uncertainty in actual driving scenarios, causing the intelligent car to make unreasonable movements in a dynamic environment and making it difficult to provide guidance information for motion planning. In view of this, how to model uncertainty at the intersection to be decided in order to improve the rationality of driving decisions at the intersection to be decided has become an urgent problem to be solved. Summary of the invention

[0004] The main technical problem solved by the present application is to provide a method for making driving decisions at intersections and related devices, equipment and storage media, which can perform uncertainty modeling at intersections to be decided, so as to improve the rationality of driving decisions at intersections to be decided.

[0005] In order to solve the above problems, the first aspect of the present application provides a method for driving decision-making at an intersection, including: obtaining the covariance matrix of traffic participants other than the own vehicle at the intersection to be decided at the current moment; wherein the covariance matrix represents the prediction error of the future position of the traffic participants; making predictions based on the covariance matrix of the traffic participants at the current moment, and obtaining the position range of the traffic participants at the next moment; performing analysis based on the planned path of the own vehicle and the position range of the traffic participants at the next moment, and obtaining analysis results, wherein the analysis results include whether there is a collision risk between the own vehicle and the traffic participants at the intersection to be decided at the next moment; obtaining decision information of the own vehicle based on whether the analysis result includes the existence of a collision risk; wherein the decision information at least includes the expected acceleration.

[0006] Among them, obtaining the covariance matrix of traffic participants other than the vehicle at the intersection to be decided at the current moment includes: obtaining the covariance matrix of traffic participants when making driving decisions at the previous moment, and obtaining the system state matrix, control input matrix and state feedback matrix; based on the system state matrix, control input matrix and state feedback matrix, updating the covariance matrix of traffic participants at the previous moment, and obtaining the covariance matrix of traffic participants at the current moment.

[0007] Among them, when the previous moment is the initial moment, the covariance matrix of the traffic participants when making driving decisions at the previous moment is obtained, including: obtaining the historical collected information of the traffic participants before the initial moment; based on the historical collected information of the traffic participants, predicting the covariance matrix of the traffic participants at the initial moment.

[0008] Among them, the covariance matrix of the traffic participant at the previous moment is updated based on the system state matrix, the control input matrix and the state feedback matrix to obtain the covariance matrix of the traffic participant at the current moment, including: obtaining the variance update matrix based on the sum of the product of the control input matrix and the state feedback matrix and the system state matrix; obtaining the covariance matrix of the traffic participant at the current moment based on the product of the variance update matrix, the covariance matrix of the traffic participant at the previous moment and the transposed matrix of the variance update matrix.

[0009] Among them, the system state matrix is ​​used to characterize the changes of longitudinal position, longitudinal velocity, lateral position and lateral velocity over time, the control input matrix characterizes the changes of lateral acceleration and longitudinal acceleration over time, the state feedback matrix includes feedback coefficients of longitudinal position, longitudinal velocity, lateral position and lateral velocity respectively, and the future position characterized by the covariance matrix includes at least longitudinal position and lateral position.

[0010] Among them, the prediction error is expressed as a variance value, and the prediction is made based on the covariance matrix of the traffic participants at the current moment to obtain the position range of the traffic participants at the next moment, including: obtaining the target distribution value of the chi-square distribution with a degree of freedom of 2 at the target quantile; based on the variance value and the target distribution value, obtaining the position range of the traffic participants at the next moment.

[0011] The target distribution number is determined by the maximum tolerable probability of the prediction error.

[0012] Among them, the covariance matrix includes a first variance value and a second variance value, the first variance value represents the prediction error of the future longitudinal position, and the second variance value represents the prediction error of the future lateral position. Based on the variance value and the target distribution value, the position range of the traffic participant at the next moment is obtained, including: based on the first variance value and the target distribution value, the major semi-axis of the position range at the next moment when it is represented by an elliptical area is obtained, and based on the second variance value and the target distribution value, the minor semi-axis of the position range at the next moment when it is represented by an elliptical area is obtained; based on the major semi-axis and the minor semi-axis, the position range at the first moment when represented by an elliptical area is obtained.

[0013] Among them, an analysis is performed based on the planned path of the vehicle and the position range of the traffic participant at the next moment to obtain an analysis result, including: determining the target position where the vehicle and the traffic participant may collide based on the intersection between the planned path and the position range of the traffic participant at the next moment; obtaining a first time taken for the vehicle to reach the target position based on a first distance from the vehicle to the target position at the current moment and the extreme speed of the vehicle through the intersection to be decided; and obtaining a second time taken for the traffic participant to reach the target position based on a second distance from the traffic participant to the target position at the current moment and the extreme speed of the traffic participant through the intersection to be decided; obtaining an analysis result based on the relationship between the time difference between the first time and the second time and the time threshold.

[0014] Among them, the speed extreme value includes the minimum speed value and the maximum speed value, and either the first time consumption or the second time consumption is the average time consumption between the longest time consumption calculated by the minimum speed value and the shortest time consumption calculated by the maximum speed value; and / or, the time consumption threshold is the smaller one of the following: a preset multiple of the ratio of the first length of the vehicle and its speed extreme value, or a preset multiple of the ratio of the second length of the traffic participant and its speed extreme value.

[0015] Among them, based on the relationship between the time consumption difference between the first time consumption and the second time consumption and the time consumption threshold, an analysis result is obtained, including: in response to the time consumption difference being less than the time consumption threshold, determining that the analysis result includes the existence of a collision risk; in response to the time consumption difference being not less than the time consumption threshold, determining that the analysis result includes the absence of a collision risk.

[0016] Among them, when the analysis result includes whether there is a collision risk, decision information of the vehicle is obtained based on whether the analysis result includes the existence of a collision risk, including: obtaining a first speed ratio between the current speed of the vehicle and the expected speed of the current road section; expanding the first speed ratio based on the speed expansion coefficient to obtain a first expansion ratio, and obtaining the product of the difference between 1 and the first expansion ratio and the maximum comfortable acceleration to obtain the expected acceleration.

[0017] Among them, when the analysis result includes the existence of a collision risk, decision information of the vehicle is obtained based on whether the analysis result includes the existence of a collision risk, including: obtaining a second speed ratio between the current speed of the vehicle and the expected speed of the current road section, and obtaining a driving distance ratio between the expected distance and the actual distance between the vehicle and the traffic participants; expanding the second speed ratio based on the speed expansion coefficient to obtain a second expansion ratio, and obtaining a third expansion ratio based on a preset power of the driving distance ratio; obtaining the expected acceleration by obtaining the product of the difference between 1 and the second expansion ratio, the third expansion ratio and the maximum comfortable acceleration.

[0018] Among them, the steps of obtaining the expected distance include: obtaining a preset expected following distance, a preset expected headway, and a relative speed between the vehicle and the traffic participants; and obtaining the expected distance based on the preset expected following distance, the preset expected headway, the current speed and the relative speed.

[0019] Among them, the decision information is that the decision model selects the driving strategy with the largest cumulative reward based on the belief space at the current moment, and the decision model includes the state space, action space, state transition model and observation space under the working conditions of the intersection to be decided, and the cumulative reward is calculated by a reward function composed of at least one of the safety reward function, efficiency reward function, comfort reward function, and action continuity reward function.

[0020] Among them, the state space includes the first state of the ego vehicle and the second state of each traffic participant, the first state includes the first position and the first speed of the ego vehicle in the longitudinal direction of the map coordinate system, the second state includes the second position and the second speed of the traffic participant in the longitudinal direction of the map coordinate system and the predicted path of the traffic participant; and / or, the action space includes a number of semantic instructions, and the several semantic instructions at least include acceleration, deceleration, and constant speed; and / or, the observation space includes the first posture of the ego vehicle and the second posture of the traffic participant, the first posture includes the observed position and the observed speed of the ego vehicle, the second posture includes the observed position, the observed speed, the observed heading angle and the predicted probability of several candidate paths of the traffic participant in the global coordinate system, and the several candidate paths include going straight, turning left, and turning right; and / or, the state space includes a number of semantic instructions, and the action space includes a number of semantic instructions, and the several semantic instructions include acceleration, deceleration, and constant speed; and / or, the observation space includes the first posture of the ego vehicle and the second posture of the traffic participant, the first posture includes the observed position and the observed speed of the ego vehicle, and the second posture includes the observed position, the observed speed, the observed heading angle and the predicted probability of several candidate paths of the traffic participant in the global coordinate system, and the several candidate paths include going straight, turning left, and turning right; The state transfer model is defined as the product of the first transfer probability of the vehicle and the second transfer probability of each traffic participant. The first transfer probability represents the state information of the vehicle transferred to the next moment under the state information of the current moment, and the second transfer probability represents the state information of the traffic participant transferred to the next moment under the state information of the current moment. The second transfer probability is defined as the product of the first transfer sub-probability and the second transfer sub-probability. The first transfer sub-probability represents the candidate path of the traffic participant transferred to the next moment under the state information of the current moment, and the second transfer sub-probability represents the system state of the traffic participant transferred to the next moment under the state information of the current moment. The system state includes at least one of the longitudinal position, longitudinal speed, lateral position, and lateral speed.

[0021] Among them, the safety reward function represents a first penalty value determined based on the collision probability evaluated based on the future state of the vehicle and the future state of the traffic participants, and when the collision probability is lower than the probability threshold, the first penalty value is the product of the difference between 1 and the collision probability and the safety weight coefficient, and when the collision probability is not lower than the probability threshold, the first penalty value is a preset value, and the absolute value of the preset value is greater than the maximum first penalty value when the collision probability is lower than the probability threshold; and / or, the efficiency reward function determines a second penalty value based on the speed difference between the current speed of the vehicle and the map speed limit, and when the current speed is greater than the map speed limit, the second penalty value is the square of the speed difference. The product of the value and the first efficiency weight coefficient, when the current speed is not greater than the map speed limit, the second penalty value is the product of the absolute value of the speed difference and the second efficiency weight coefficient; and / or, the comfort reward function determines a third penalty value based on the change value of the longitudinal acceleration of the ego vehicle, and the third penalty value is the product of the square of the change value and the comfort weight coefficient; and / or, the action continuity reward function determines a reward value based on whether the longitudinal acceleration of the ego vehicle is the same at adjacent moments, and when the longitudinal acceleration of the ego vehicle is the same at adjacent moments, the reward value is the preset cost weight coefficient, and when the longitudinal acceleration of the ego vehicle is different at adjacent moments, the reward value is zero.

[0022] In order to solve the above problems, the second aspect of the present application provides a driving decision-making device at an intersection, including: an error prediction module, a range modeling module, a risk analysis module, and a self-vehicle decision module. The error prediction module is used to obtain the covariance matrix of traffic participants other than the self-vehicle at the intersection to be decided at the current moment; wherein the covariance matrix represents the prediction error of the future position of the traffic participant; the range modeling module is used to make predictions based on the covariance matrix of the traffic participant at the current moment to obtain the position range of the traffic participant at the next moment; the risk analysis module is used to make analysis based on the planned path of the self-vehicle and the position range of the traffic participant at the next moment to obtain an analysis result, wherein the analysis result includes whether there is a collision risk between the self-vehicle and the traffic participant at the intersection to be decided at the next moment; the self-vehicle decision module is used to obtain the decision information of the self-vehicle based on whether the analysis result includes the existence of a collision risk; wherein the decision information at least includes the expected acceleration.

[0023] In order to solve the above-mentioned problem, the third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other, the memory storing program instructions, and the processor being used to execute the program instructions to implement the intersection driving decision-making method in the above-mentioned first aspect.

[0024] In order to solve the above-mentioned problem, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be executed by a processor, and the program instructions are used for the intersection driving decision method in the above-mentioned first aspect.

[0025] The above scheme obtains the covariance matrix of traffic participants other than the vehicle at the intersection to be decided at the current moment, and the covariance matrix represents the prediction error of the future position of the traffic participants, and predicts based on the covariance matrix of the traffic participants at the current moment to obtain the position range of the traffic participants at the next moment, and then analyzes based on the planned path of the vehicle and the position range of the traffic participants at the next moment to obtain the analysis result, and the analysis result includes whether there is a collision risk between the vehicle and the traffic participants at the intersection to be decided at the next moment, so as to obtain the decision information of the vehicle based on whether the analysis result includes the existence of the collision risk, and the decision information at least includes the expected acceleration, and then on the one hand, the position range of the traffic participants at the next moment is predicted through the covariance matrix of the traffic participants, and the uncertainty of each traffic participant other than the vehicle at the current moment can be modeled, and on the other hand, the planned path of the vehicle and the position range of the traffic participants at the next moment are analyzed, and the dynamics of the traffic participants can be considered in combination with the uncertainty modeled. Therefore, uncertainty modeling can be performed at the intersection to be decided to improve the rationality of driving decisions at the intersection to be decided. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flow chart of an embodiment of the method for making driving decisions at intersections of the present application;

[0027] Figure 2a This is a schematic diagram of an embodiment of the present application in which a vehicle turns left at an intersection to be decided;

[0028] Figure 2b This is a schematic diagram of an embodiment of the present application in which a vehicle goes straight at a road intersection to be decided;

[0029] Figure 2c This is a schematic diagram of an embodiment of the present application of a vehicle turning right at an intersection to be decided;

[0030] Figure 3 is a schematic diagram of an embodiment of the position range of a traffic participant at the next moment of the present application;

[0031] Figure 4a is a schematic diagram of an embodiment of the time interval to the collision point of the present application;

[0032] Figure 4b It is a schematic diagram of an embodiment of the time difference of the present application;

[0033] Figure 5 This is a schematic diagram of the process of an embodiment of the intersection driving decision method of the present application;

[0034] Figure 6 It is a schematic diagram of the framework of an embodiment of the intersection driving decision-making device of the present application;

[0035] Figure 7 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0036] Figure 8 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0037] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.

[0038] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0039] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.

[0040] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the method for making decisions at intersections of the present application. Specifically, the following steps may be included:

[0041] Step S11: Obtain the covariance matrix of traffic participants other than the vehicle at the intersection to be decided at the current moment.

[0042] In the disclosed embodiment, the intersection to be decided may include but is not limited to a crossroad, a T-junction, etc., which is not limited here. For ease of understanding, taking the intersection to be decided as a crossroad as an example, the following introduces different possible situations such as the vehicle turning left, going straight, and turning right. Figure 2a , Figure 2a This is a schematic diagram of an embodiment of the present application in which a vehicle turns left at a road intersection to be decided. Figure 2aAs shown in the figure, there are three traffic participants (Vehicle 2, Vehicle 3 and Vehicle 4) at the intersection to be decided. For the vehicle, the movement routes of the other three traffic participants are uncertain. They may turn left, go straight or turn right. These states are hidden states. The possible collision situations are: (1) Vehicle 2 turns left or goes straight, (2) Vehicle 3 goes straight, turns left or turns right, (3) Vehicle 4 turns left or goes straight. Therefore, it is necessary to observe the information of other vehicles to determine the most likely movement route of other vehicles and choose the appropriate turning time to find a gap in the oncoming traffic flow and pass the intersection quickly. This unprotected left turn is one of the most challenging scenarios in autonomous driving scenarios. Please continue to read Figure 2b , Figure 2b This is a schematic diagram of an embodiment of the present application in which a vehicle goes straight at a road intersection to be decided. Figure 2b As shown in the figure, there is a self-vehicle (vehicle 1) and three other traffic participants (vehicles 2, 3, and 4) at the intersection to be decided. The possible collision situations are: (1) vehicle 2 turns left or goes straight, (2) vehicle 3 turns left, and (3) vehicle 4 turns left, turns right, or goes straight. Therefore, it is necessary to observe the information of other vehicles to determine the most likely movement route of other vehicles and choose the appropriate time to pass. Please continue to refer to Figure 2c , Figure 2c This is a schematic diagram of an embodiment of the present application in which a vehicle turns right at an intersection to be decided. Figure 2c As shown in the figure, there is a self-vehicle (vehicle 1) and three other traffic participants (vehicles 2, 3, and 4) at the intersection to be decided. The possible collision situations are: (1) vehicle 2 goes straight, (2) vehicle 3 turns left, so the smart car needs to turn right safely and efficiently. In addition, the above Figure 2a to Figure 2c The black solid dots in the figure indicate the possible collision locations. It can be seen that through collision analysis, there are three possible scenarios where smart cars may collide at intersections. If there is a conflict between two cars, the smart car needs to judge the movement intention of the other car when passing through the intersection and make a decision to give way or speed up. Of course, Figure 2a to Figure 2c The examples shown are just a few possible examples of driving decisions in a parking lot intersection scenario in actual application, and other possible situations will not be given one by one here.

[0043] In the disclosed embodiment, the covariance matrix represents the prediction error of the future position of the traffic participant. It should be noted that a covariance matrix can be maintained for each moment of each traffic participant at the intersection to be decided to model the prediction error of its future position at the next moment. In addition, the future position may specifically include a longitudinal position and a lateral position. For example, the longitudinal position and the lateral position can be represented with reference to the map coordinate system of the intersection to be decided. Traffic participants may include but are not limited to: cars, electric bicycles, etc., and traffic participants are not limited here.

[0044] In an implementation scenario, when the current moment is the initial moment, the historical collection information of the traffic participants before the initial moment can be obtained first. It should be noted that the historical collection information can be obtained by collecting information on the traffic participants through sensors such as the laser radar and camera of the vehicle. For example, when collecting through the laser radar of the vehicle, the historical collection information may include the point cloud information collected from the traffic participants, and when collecting through the camera of the vehicle, the historical collection information may include the image information of the traffic participants. The specific type of historical collection information is not limited here. On this basis, the covariance matrix of the traffic participants at the initial moment can be predicted based on the historical collection information of the traffic participants. Exemplarily, the historical collection information can be predicted by an error prediction model (such as a convolutional neural network, a recurrent neural network, etc.) to obtain the covariance matrix at the initial moment. Still taking the future position including the longitudinal position and the lateral position as an example, the covariance matrix of the traffic participants at the initial moment may include the prediction error of the traffic participants in the longitudinal position (such as the variance of the longitudinal position) and the prediction error in the lateral position (such as the variance of the lateral position). Exemplarily, the covariance matrix of the traffic participants at the initial moment can be expressed as:

[0045]

[0046] In the above formula (1), Respectively represent the variance of the longitudinal position and the lateral position at the initial moment. In the above method, at the initial moment, the historical collected information of the traffic participants before the initial moment is obtained, and then based on the historical collected information of the traffic participants, the covariance matrix of the traffic participants at the initial moment is predicted, and the covariance matrix can be obtained through the historical collected information at the initial moment to start the driving decision at the initial moment.

[0047] In an implementation scenario, as a possible example, when the current moment is after the initial moment, the covariance matrix of the current moment can be obtained by referring to the specific process of obtaining the covariance matrix at the initial moment. For example, when the current moment is after the initial moment, the historical collected information of the traffic participants before the current moment can be obtained, and then the covariance matrix of the traffic participants at the initial moment can be predicted based on the historical collected information of the traffic participants before the current moment. For details, please refer to the aforementioned process steps for obtaining the covariance matrix at the initial moment, which will not be repeated here. Alternatively, as another possible example, when the current moment is after the initial moment, in order to reduce the computational load, the covariance matrix of the traffic participants when making driving decisions at the previous moment can also be obtained, and the system state matrix, control input matrix and state feedback matrix can be obtained. The covariance matrix of the traffic participants at the previous moment is then updated based on the system state matrix, control input matrix and state feedback matrix to obtain the covariance matrix of the traffic participants at the current moment. In other words, the covariance matrix of the traffic participants at moment k+1 can be updated by updating the covariance matrix of the traffic participants at moment k, without repeatedly predicting the covariance matrix based on historical collected information, which helps to reduce the computational load of obtaining the covariance matrix as much as possible.

[0048] In a specific implementation scenario, when the current moment is the initial moment, since there is no previous moment before this, the specific process of obtaining the covariance matrix at the initial moment can be referred to, so as to obtain the covariance matrix of the traffic participants at the current moment when the current moment is the initial moment. For details, please refer to the above-mentioned related description, which will not be repeated here.

[0049] In a specific implementation scenario, the system state matrix can be used to characterize the changes in longitudinal position, longitudinal velocity, lateral position, and lateral velocity over time. As a possible example, the system state matrix A obj It can be expressed as:

[0050]

[0051] In the above formula (2), Δt represents the time interval of discrete sampling. In addition, the control input matrix represents the change of lateral acceleration and longitudinal acceleration over time. As a possible example, the control input matrix B obh It can be expressed as:

[0052]

[0053] In addition, the state feedback matrix includes feedback coefficients for longitudinal position, longitudinal velocity, lateral position and lateral velocity, respectively. As a possible example, the state feedback matrix K can be expressed as:

[0054]

[0055] In the above formula (4), k s represents the feedback factor for longitudinal position consideration, k vs represents the feedback factor for longitudinal velocity consideration, k l Denotes the feedback factor for lateral position consideration, k vl Represents the feedback factor taken into account for the lateral velocity.

[0056] In a specific implementation scenario, after obtaining the system state matrix, the control input matrix and the state feedback matrix, the variance update matrix can be obtained based on the product of the control input matrix and the state feedback matrix and the sum of the system state matrix:

[0057] A obj +B obk K……(5)

[0058] On this basis, it is assumed that the forecast error follows a normal distribution Then, the covariance matrix of the traffic participants at the current moment can be obtained based on the product of the variance update matrix, the covariance matrix of the traffic participants at the previous moment, and the transposed matrix of the variance update matrix:

[0059]

[0060] In the above formula (6), represents the covariance matrix of traffic participants at the previous moment (i.e., moment k), Represents the covariance matrix of the traffic participant at the current moment (i.e., moment k+1), and T represents the transpose operation of the matrix. In the above method, the variance update matrix is ​​obtained based on the sum of the product of the control input matrix and the state feedback matrix and the system state matrix, and then the covariance matrix of the traffic participant at the current moment is obtained based on the product of the variance update matrix, the covariance matrix of the traffic participant at the previous moment, and the transpose matrix of the variance update matrix. Therefore, the covariance matrix at the current moment can be updated through matrix operations, which helps to reduce the computational load of obtaining the covariance matrix as much as possible. For ease of understanding, the state update of the prediction error is theoretically explained below:

[0061] For traffic participants other than the ego vehicle, their future motion state can be described by a point mass model. Such a linear model can propagate uncertainty over time and is defined as:

[0062] x obj,k+1 =A obk x obk,k +B obj u obj,k ……(7)

[0063] In the above formula (7), Respectively represent the longitudinal position s in the road coordinate system obj,k , longitudinal speed Horizontal position obj,k , and the lateral velocity In addition, the subscript k represents the time k, and the subscript k+1 represents the next time after time k. obj represents the system state matrix, B obj represents the control input matrix, which can be found in the above related description and will not be described here. obj,k represents the control input, which can include lateral acceleration and longitudinal acceleration. Assume that the predicted true trajectory is This means that the surrounding dynamic traffic participants follow the given reference state with allowable deviations, and a feedback controller is defined as follows:

[0064]

[0065] In the above formula (8), K represents the state feedback matrix. For details, please refer to the above related description, which will not be repeated here. Based on this, formula (8) can be substituted into formula (7) to obtain:

[0066]

[0067] Similarly, the predicted motion state of the surrounding dynamic traffic participants is:

[0068]

[0069] The resulting prediction error can be defined as:

[0070]

[0071] Furthermore, the state update variance of the prediction error can be obtained as follows:

[0072] e k+1 =(A obj +B obj K)e k ……(12)

[0073] As time goes by, the accuracy of trajectory prediction will decrease, and the uncertainty of the corresponding motion position will increase. The uncertainty of the vehicle's future state over time, that is, the probability distribution of the future state, can be calculated by formula (12).

[0074] Step S12: Predict the position range of the traffic participant at the next moment based on the covariance matrix of the traffic participant at the current moment.

[0075] Specifically, as mentioned above, the prediction error can be expressed as a variance value. For details, please refer to the above-mentioned related description, which will not be repeated here. On this basis, the target distribution value of the chi-square distribution with a degree of freedom of 2 at the target quantile can be obtained, and then based on the variance value and the target distribution value, the position range of the traffic participants at the next moment can be obtained. In the above method, the target distribution value of the chi-square distribution with a degree of freedom of 2 at the target quantile is combined with the variance value representing the prediction error to predict the position range at the next moment, which can reduce the complexity of uncertainty modeling as much as possible.

[0076] In an implementation scenario, for ease of understanding, the prediction error is first theoretically explained. As mentioned above, it can be assumed that only future positions are considered in uncertainty modeling, that is, the error vector can be defined They represent the longitudinal position error and the lateral position error respectively, and the corresponding covariance matrix can be expressed as:

[0077]

[0078] In the above formula (13), represents the variance of the longitudinal position error, represents the variance of the lateral position error. For ease of distinction, we can It is called the first variance value. It is called the second variance. Assume that the longitudinal position error and the lateral position error form a bivariate normal distribution with a mean of μ = [μ s ,μ l ] T =0, and the longitudinal position error and the lateral position error are independent (unrelated). Therefore, two independent random variables that obey the standard normal distribution can be obtained respectively:

[0079]

[0080] Considering the influence of the above two variables, the sum of squares of random variables in formula (14) can be defined as 2 (2) (Chi-square) distribution:

[0081]

[0082] The chi-square distribution commonly used in mathematical statistics is used to describe the confidence region (i.e., the position range at the next moment, i.e., the uncertainty range) of dynamic obstacle vehicles (i.e., traffic participants). This confidence region can be enclosed by an elliptical contour line. The area inside the ellipse represents the spatial location area that may appear in the future. On this basis, the risk parameter (i.e., the maximum tolerable probability p) can be defined. risk ), so that the probability of the true prediction state considering the prediction error in this area is not less than the maximum tolerable probability prisk When the maximum tolerable probability p risk As χ increases, the range of the confidence region will also expand. This is due to 2 (2) The cumulative distribution function of the distribution is determined by its properties.

[0083] In one implementation scenario, the target distribution number is determined by the maximum tolerable probability of the prediction error. For example, the target distribution number can be set to the maximum tolerable probability, and the target distribution value can be expressed as

[0084] In one implementation scenario, as described above, the probability that the true predicted state is within the region taking into account the prediction error is not less than the maximum tolerable probability p risk , which can be expressed as:

[0085]

[0086] On this basis, based on the first variance value and the target distribution value Get the major semi-axis of the position range at the next moment when it is represented by an elliptical area

[0087]

[0088] At the same time, based on the second variance value and the target distribution value Get the short semi-axis when the position range at the next moment is represented by an elliptical area

[0089]

[0090] On this basis, please refer to Figure 3 , Figure 3 Schematic diagram of an embodiment of the position range of the traffic participants at the next moment in the present application. Figure 3As shown, the position range at the first moment when represented by an elliptical area can be obtained based on the major semi-axis and the minor semi-axis. In the above method, the covariance matrix includes a first variance value and a second variance value, the first variance value represents the prediction error of the future longitudinal position, and the second variance value represents the prediction error of the future lateral position. Based on this, based on the first variance value and the target distribution value, the major semi-axis of the position range at the next moment when represented by an elliptical area is obtained, and based on the second variance value and the target distribution value, the minor semi-axis of the position range at the next moment when represented by an elliptical area is obtained, and then based on the major semi-axis and the minor semi-axis, the position range at the first moment when represented by an elliptical area is obtained, which can reduce the complexity of uncertainty modeling as much as possible. In addition, it should be noted that, for each traffic participant other than the self-vehicle at the intersection to be decided, the self-vehicle can determine its position range at the next moment through the above process steps, that is, the self-vehicle can perform uncertainty modeling for each traffic participant other than itself at the intersection to be decided through the above process steps.

[0091] Step S13: Analyze the planned path of the vehicle and the position range of the traffic participants at the next moment to obtain the analysis result.

[0092] In the disclosed embodiment, the analysis result includes whether there is a risk of collision between the vehicle and the traffic participant at the intersection to be decided at the next moment. Exemplarily, the analysis result may include that there is a risk of collision between the vehicle and the traffic participant at the intersection to be decided at the next moment, or the analysis result may include that there is no risk of collision between the vehicle and the traffic participant at the intersection to be decided at the next moment. Specifically, the target position where the vehicle and the traffic participant may collide can be determined based on the intersection between the planned path and the position range of the traffic participant at the next moment, and the first time taken for the vehicle to reach the target position can be obtained based on the first distance from the vehicle to the target position at the current moment and the speed extreme value of the vehicle passing through the intersection to be decided, and the second time taken for the traffic participant to reach the target position can be obtained based on the second distance from the traffic participant to the target position at the current moment and the speed extreme value of the traffic participant passing through the intersection to be decided, and then the analysis result is obtained based on the relationship between the time difference between the first time and the second time and the time threshold. In the above method, by first determining the target position where the collision may occur, and then combining the vehicle and the traffic participant to determine the target position based on the first distance from the vehicle to the target position at the current moment and the speed extreme value of the traffic participant passing through the intersection to be decided, the analysis result is obtained based on the relationship between the time difference between the first time and the second time and the time threshold.

[0093] The difference in the time it takes for the vehicle and traffic participants to reach the target location is used to determine whether there is a collision risk, which helps to improve the accuracy of collision risk detection.

[0094] In one implementation scenario, for each traffic participant other than the vehicle at the intersection to be decided, an analysis can be performed based on the planned path of the vehicle (e.g., left turn, straight ahead, right turn, etc.) and the position range of the traffic participant at the next moment to determine the target position where the two may collide. Exemplarily, when there is an intersection between the planned path of the vehicle and the position range of the traffic participant at the next moment, the position point on the intersection can be used as the target position where the vehicle and the traffic participant may collide. It should be noted that for different situations such as the vehicle turning left, turning right, and going straight at the intersection to be decided, please refer to the following table respectively. Figure 2a to Figure 2c , determine the target position that may collide with traffic participants (such as Figure 2a to Figure 2c Indicated by black solid dots).

[0095] In one implementation scenario, the speed extreme value includes a minimum speed value and a maximum speed value, and either the first time consumption or the second time consumption is the average time consumption between the longest time consumption calculated by the minimum speed value and the shortest time consumption calculated by the maximum speed value. For the convenience of description, for the i-th traffic participant, its minimum speed value can be recorded as Its minimum speed can be recorded as The corresponding longest time can be recorded as The shortest time can be recorded as The calculation formula can be expressed as:

[0096]

[0097] In the above formula (19), s i represents the second distance from the i-th traffic participant to the target location at the current moment, v i represents the extreme speed of the ith traffic participant, t i Represents the relevant time consumption of the i-th traffic participant (e.g., shortest time consumption, longest time consumption). Figure 4a , Figure 4a Schematic diagram of an embodiment of the time interval to the collision point of the present application. Figure 4a As shown, on this basis, for any traffic participant (or vehicle), the average time of the longest time and the shortest time can be obtained, that is, the first time for the vehicle to reach the target location and the second time for the traffic participant to reach the target location can be obtained. Figure 4b , Figure 4b Schematic diagram of an embodiment of the time difference of the present application. Figure 4b As shown, on this basis, for any traffic participant, the time difference between the first time of the vehicle and the second time of the traffic participant can be obtained, which can be recorded as t for the convenience of description. di ff .

[0098] In one implementation scenario, the time-consuming threshold is the smaller of the following: a preset multiple of the ratio of the first length of the vehicle and its speed extreme value, and a preset multiple of the ratio of the second length of the traffic participant and its speed extreme value. It should be noted that the preset multiple can be set according to actual application needs, such as being set to 2, and the specific value of the preset multiple is not limited here. Taking the vehicle as A and the traffic participant as B as an example, the time-consuming threshold can be expressed as:

[0099]

[0100] In the above formula (20), t threshold Indicates the time-consuming threshold, min means taking the smaller one, L i represents the first length of the vehicle A or the second length of the traffic participant B, v i Indicates the speed extreme of the ego vehicle or traffic participant. In this example, the preset multiple is set to 2. On this basis, when the time difference satisfies the following formula, it can be considered that there is no collision risk between the ego vehicle and the traffic participant:

[0101] t diff =|t A -t B | <t threshold ……(twenty one)

[0102] In the above formula (21), t A represents the first time taken by vehicle A, t B Indicates the second time taken by traffic participant B. That is, in response to the time difference being less than the time threshold, it can be determined that the analysis result includes the presence of a collision risk, whereas in response to the time difference being not less than the time threshold, it can be determined that the analysis result includes the absence of a collision risk.

[0103] Step S14: Based on whether the analysis result includes the existence of a collision risk, the decision information of the vehicle is obtained.

[0104] In the disclosed embodiment, the decision information includes at least the expected acceleration. Of course, according to the specific value of the expected acceleration (e.g., positive, negative, zero), the decision information may further include the target action determined in several semantic actions (e.g., acceleration, deceleration, uniform speed). For example, when the expected acceleration is zero, the target action can be determined to be uniform speed; when the expected acceleration is a positive number, the target action can be determined to be acceleration; when the expected acceleration is a negative number, the target action can be determined to be deceleration. The above example is only a possible example of decision information in actual application, and the specific content of the decision information will not be given one by one here.

[0105] In one implementation scenario, as a possible example, when the analysis result includes that there is no collision risk, a first speed ratio between the current speed of the vehicle and the expected speed of the current section can be obtained, and then the first speed ratio is expanded based on the speed expansion coefficient to obtain a first expansion ratio, and the product of the difference between 1 and the first expansion ratio and the maximum comfortable acceleration is obtained to obtain the expected acceleration. For ease of description, the current speed of the vehicle can be recorded as v, the expected speed of the current section can be recorded as v0, the expansion coefficient can be recorded as σ, and the maximum comfortable acceleration can be recorded as a. Then, when the analysis result includes that there is no collision risk, the expected acceleration a des It can be expressed as:

[0106]

[0107] It can be seen from the above formula (22) that the smaller the ratio between the current speed of the vehicle and the expected speed, the larger the calculated acceleration will be, which can reduce the speed error.

[0108] In another implementation scenario, as another possible example, different from the aforementioned situation, when the analysis results include the existence of a collision risk, a second speed ratio between the current speed of the vehicle and the expected speed of the current road section can be obtained, and the driving distance ratio between the expected distance and the actual distance between the vehicle and the traffic participants can be obtained, so that the second speed ratio can be expanded based on the speed expansion coefficient to obtain a second expansion ratio, and based on the preset power of the driving distance ratio, a third expansion ratio can be obtained, and then the product of the difference between 1 and the second expansion ratio, the third expansion ratio and the maximum comfortable acceleration can be obtained to obtain the expected acceleration, so that a certain safe distance between the vehicle and the traffic participants can be ensured as much as possible.

[0109] In a specific implementation scenario, in order to obtain the expected distance, the preset expected following distance, the preset expected headway, and the relative speed between the vehicle and the traffic participant can be obtained, and then the expected distance is obtained based on the preset expected following distance, the preset expected headway, the current speed and the relative speed. For example, the expected distance s * (v,Δv) can be expressed as:

[0110]

[0111] In the above formula (23), s0 represents the preset expected following distance, T hrepresents the preset expected headway, Δv represents the relative speed between the vehicle and the traffic participant, and b is the maximum comfortable deceleration. The specific meanings of other parameters in formula (23) can be found in the relevant description of the aforementioned formula (22), which will not be repeated here. In other words, the first product between the current speed of the vehicle and the preset expected headway, the second product between the maximum comfortable acceleration and the maximum comfortable deceleration, and the third product between the current speed of the vehicle and the aforementioned relative speed can be obtained, and then the first ratio between the third product and the root value of the second product can be obtained, and then the sum of the preset expected following distance, the first product, and the first ratio can be obtained to obtain the expected spacing. Of course, formula (23) is only one possible calculation method for the expected spacing, and other possible calculation methods are not limited here.

[0112] In a specific scenario, after obtaining the expected distance, the expected acceleration a can be calculated when the analysis results include the risk of collision. des :

[0113]

[0114] In the above formula (24), v represents the current speed of the vehicle, v0 represents the expected speed of the current road section, Δs represents the actual distance between the vehicle and the traffic participants, σ ​​represents the speed expansion coefficient, and a represents the maximum comfortable acceleration. In addition, in this example, the preset power can be set to 2.

[0115] In an implementation scenario, the decision information can be used for the decision model to select a driving strategy with the largest cumulative reward based on the belief space at the current moment, and the decision model can include the state space, action space, state transition model and observation space under the working conditions of the intersection to be decided, and the cumulative reward is calculated by a reward function composed of at least one of the safety reward function, efficiency reward function, comfort reward function and action continuity reward function. The above method combines the state space, action space, state transition model and observation space under the working conditions of the intersection to be decided and the cumulative reward determined by the reward function composed of at least one of the safety reward function, efficiency reward function, comfort reward function and action continuity reward function to make driving decisions, which can fully understand the surrounding environment through the decision model and reward function, and help to maximize the rationality of driving decisions.

[0116] In a specific implementation scenario, the state space may include the first state of the vehicle and the second state of each traffic participant. It should be noted that in order to model the interactive behavior under the decision-making intersection conditions, all traffic participants (including the vehicle and traffic participants other than the vehicle) must be represented in the state space. For the state space composed of the entire environment, information such as road topology and geometry can be assumed to be static information of the model (that is, it can be obtained through the map), so it may not be included in the state. For ease of expression, the state space can be expressed as:

[0117] χ={χ ego ,χ1,…,χ n}……(25)

[0118] In the above formula (25), χ ego represents the first state of the vehicle, χ i (i∈1,…,n) represents the second state of the i-th traffic participant. For the ego vehicle, the position and posture can be observed through its own high-precision positioning and navigation system (e.g., the accuracy can reach the centimeter level). In addition, since the smart car has obtained the global path (i.e., the aforementioned planned path) at the intersection to be decided, it can only consider the movement of the ego vehicle in the longitudinal direction, and then the map coordinate system can still be used to describe the movement state of the ego vehicle in the longitudinal direction, including the longitudinal position and longitudinal speed, such as the first state χ ego It can be defined as:

[0119]

[0120] In the above formula (26), χ ego represents the first state, which includes the first position s of the vehicle in the longitudinal direction of the map coordinates ego and the first velocity v ego In addition, for traffic participants other than the vehicle, in addition to the above states, the movement route also needs to be considered, that is, there may be left turns, straight driving, and right turns. This state cannot be directly observed, but is a partially observable state, such as the location information of the traffic participants and the information of the prediction model. For example, the second state can be defined as:

[0121]

[0122] In the above formula (27), χ i represents the second state of the i-th traffic participant, which includes the second position s of the traffic participant in the longitudinal direction of the map coordinate system i and the second speed v i and the predicted path r of the traffic participant i , such as turn left, go straight, turn right, etc.

[0123] In a specific implementation scenario, the action space may include several semantic instructions, and the several semantic instructions include at least acceleration, deceleration, and uniform speed. For example, the longitudinal decision-making system at the intersection only needs to generate a reasonable speed curve on a given reference path. At this time, the longitudinal decision-making system only needs to calculate the expected acceleration instruction and send it to the lower layer. Therefore, several semantic instructions can be expressed as:

[0124] A lon =[ADD,DEC,MS]……(28)

[0125] In the above formula (28), ACC means acceleration, DEC means deceleration, and MSA means constant speed. These semantic instructions are discrete instructions, and the acceleration of the ego vehicle is determined by the generated strategy and the current belief state. However, these discrete instructions are difficult to provide detailed guidance for the lower-level motion planning. Finally, the future speed curve of the ego vehicle in the future prediction period can be obtained through closed-loop simulation. Although it is a rough speed curve, it can restrict the convex space of the motion planning, which is convenient for the optimization solution of the lower-level speed planning.

[0126] In a specific implementation scenario, the observation space may include the first posture of the vehicle and the second posture of the traffic participant. It should be noted that the observation space is the belief space distribution that takes into account the current driving behavior and driving action estimation, and contains the information that can be observed by the intelligent vehicle environment perception and positioning system. Corresponding to the state space, it can be defined as:

[0127] O=[o ego ,o1,o2,…,o n ]……(29)

[0128] In the above formula (29), O represents the observation space, o ego represents the first position of the vehicle, o i represents the second position of the i-th traffic participant. In addition, the observation space can be decomposed into the observation model of the smart car itself (i.e., the ego car) and other surrounding vehicles (i.e., traffic participants):

[0129]

[0130] In the above formula (30), since the state of the vehicle itself is completely observable, P(o ego |χ′ ego )=1. It should be noted that the first pose can specifically include the observation position and observation speed of the vehicle, so the first pose can be expressed as:

[0131]

[0132] In addition, for traffic participants other than the vehicle, the second position includes the observed position, observed speed, observed heading angle and predicted probabilities of several candidate paths in the global coordinate system, and the candidate paths include going straight, turning left and turning right. It should be noted that the candidate path r selected by the i-th traffic participant can be obtained through the upper prediction model. i The predicted probability of , so the second posture can be expressed as:

[0133]

[0134] In a specific implementation scenario, the state transition model can be defined as the product of the first transition probability of the vehicle and the product of the second transition probabilities of each traffic participant. As a possible example, the state transition model can be expressed as:

[0135]

[0136] In the above formula (33), P(χ′ ego | ego ,a) represents the first transition probability, which represents the state information of the vehicle at the current moment ego The state information x′ transferred to the next moment under the condition ego It should be noted that the state update of the vehicle can be expressed as:

[0137]

[0138] In the above formula (34), t is the sampling step length, s ego,k represents the longitudinal position of the vehicle at time k, v ego,k represents the longitudinal velocity of the vehicle at time k, a ego,k represents the longitudinal acceleration of the vehicle at time k, s ego,k+1 represents the longitudinal position of the vehicle at time k+1, v ego,k+1 represents the longitudinal velocity of the vehicle at time k+1. In addition, P(χ′ i | i ) represents the second transition probability, which represents the state information x of the traffic participant at the current moment i The condition is transferred to the state information X' at the next moment i. It should be noted that, as mentioned above, in conditions such as waiting for a decision at an intersection, the movement uncertainty of traffic participants other than the vehicle specifically includes the uncertainty of the movement route, such as left turn, straight and right turn, and the uncertainty of the movement intention. The state transition probability of traffic participants cannot be directly obtained through environmental perception. Therefore, the second transition probability can be defined as the product of the first transfer sub-probability and the second transfer sub-probability, that is, the second transition probability can be split into two parts: driving intention transfer and movement state transfer:

[0139] P(χ′ i |X i ) = P(r′ i | i )P(x′ i | i )……(35)

[0140] In the above formula (35), P(r′ i | i ) represents the first transfer sub-probability, which represents the state information x of the traffic participant at the current moment i The candidate path r′ that is transferred to the next moment under the condition i , P(x′ i |X i ) represents the second transfer probability, which represents the state information X of the traffic participant at the current moment i The system state x′ at the next moment is transferred to i As mentioned above, the system state may include at least one of the longitudinal position, longitudinal speed, lateral position, and lateral speed. Of course, in addition to the uncertainty caused by observing traffic participants, the movement of traffic participants will also affect the decision-making process of the vehicle. As a possible example, the uncertainty distribution of each decision compensation (i.e., each moment) can be obtained according to the process steps of the aforementioned uncertainty modeling. For details, please refer to the aforementioned related description, which will not be repeated here.

[0141] In a specific implementation scenario, the decision model can be a POMDP model, which selects the strategy with the largest cumulative reward based on the current belief space. The setting of the reward function is crucial to the decision-making process. The decision-making actions of smart cars at intersections need to meet the criteria of safety, comfort, and efficiency, while also considering the continuity of driving actions. Therefore, the reward function under the intersection condition is composed of a linear combination of the costs defined above, defined as:

[0142] R i =R s +R e +R c +R a ……(36)

[0143] In the above formula (36), R s represents the security reward function, R e represents the efficiency reward function, R c represents the comfort reward function, R a represents the action continuity reward function.

[0144] In a specific implementation scenario, the safety reward function represents a first penalty value determined based on the collision probability assessed based on the future state of the vehicle and the future state of the traffic participant, and when the collision probability is lower than the probability threshold, the first penalty value is the product of the difference between 1 and the collision probability and the safety weight coefficient, and when the collision probability is not lower than the probability threshold, the first penalty value is a preset value, and the absolute value of the preset value is greater than the maximum first penalty value when the collision probability is lower than the probability threshold. In other words, when there is a high probability of a collision, a huge cost penalty can be given. Due to the uncertainty of the movement of surrounding vehicles (i.e., traffic participants), it is difficult to always accurately locate at the predicted trajectory position, making the collision detection of the trajectory unreliable. The degree of danger of a collision between the vehicle and other vehicles (i.e., traffic participants) can also be presented in terms of the size of the collision probability. As a possible example, the collision probability can be calculated by Monte Carlo sampling, and a large penalty is given if a collision occurs. Exemplarily, the safety reward function can be expressed as:

[0145]

[0146] In the above formula (37), R s represents the security reward function, P colLision represents the collision probability, w collision,lon Represents the safety weight coefficient, P safe Represents the probability threshold.

[0147] In a specific implementation scenario, the efficiency reward function determines the second penalty value based on the speed difference between the current speed of the vehicle and the speed limit on the map, and when the current speed is greater than the speed limit on the map, the second penalty value is the product of the square of the speed difference and the first efficiency weight coefficient, and when the current speed is not greater than the speed limit on the map, the second penalty value is the product of the absolute value of the speed difference and the second efficiency weight coefficient. In other words, for the efficiency reward function, it means that the smart car (i.e., the vehicle) wants to pass the intersection (i.e., the intersection to be decided) as quickly as possible, and specifically the difference between the speed limit on the vehicle and the map can be considered, so that the vehicle can be as close to the speed limit on the map as possible while meeting safety requirements. ref . It can be mainly divided into two situations: low speed and speed greater than the map speed limit, which can be expressed as:

[0148]

[0149] In the above formula (38), R e represents the efficiency reward function, v ego represents the current speed of the vehicle, v ref Indicates map speed limit, w v1 represents the first efficiency weight coefficient, w v2 Represents the second efficiency weight coefficient.

[0150] In a specific implementation scenario, the comfort reward function can be based on the third penalty value determined by the change value of the longitudinal acceleration of the vehicle, and the third penalty value is the product of the square of the change value and the comfort weight coefficient. In other words, the comfort reward function mainly considers the change of longitudinal acceleration, which can be expressed as:

[0151]

[0152] In the above formula (39), R c represents the comfort reward function, Δa ego represents the change in longitudinal acceleration of the vehicle, w c,lon Represents the comfort weight coefficient.

[0153] In a specific implementation scenario, the action continuity reward function can be based on whether the longitudinal acceleration of the ego vehicle is the same at adjacent moments to determine the reward value, and when the longitudinal acceleration of the ego vehicle is the same at adjacent moments, the reward value is a preset cost weight coefficient, and when the longitudinal acceleration of the ego vehicle is different at adjacent moments, the reward value is zero, which can be expressed as:

[0154]

[0155] In the above formula (40), w a,lon Represents the preset cost weight coefficient, a′ ego =a ego It means that the longitudinal acceleration of the vehicle is the same at adjacent moments.

[0156] In a specific implementation scenario, since a single driving action is difficult to meet the needs of motion, it is necessary to provide a series of future state sequences and ensure the non-convexity of the optimization problem. In order to improve the search efficiency and obtain the future state trajectory of the self-vehicle, the future motion state of the self-vehicle and the surrounding vehicles can be deduced through forward simulation, and the uncertainty of the trajectory at the corresponding moment can be evaluated to obtain the possible future trajectories of the self-vehicle and the surrounding vehicles. The optimal driving strategy can be evaluated according to the designed reward function, which can convert the complex POMDP solution problem into the problem of selecting the best strategy with the highest return from a limited number of strategies, which can effectively improve the real-time performance of the operation. At the same time, the future state sequence of the self-vehicle can be obtained through forward simulation, which can effectively provide guidance for the lower-level motion planning. The framework of the entire algorithm can specifically simulate and evaluate multiple strategies starting from the initial belief space, and finally pass the safety check. For example, the multi-strategy closed-loop simulation solution algorithm can refer to Table 1:

[0157] Table 1 Multi-strategy closed-loop simulation solution algorithm

[0158]

[0159] As shown in Table 1, the input is the uncertainty distribution of the ego vehicle state and the surrounding dynamic vehicles, the semantic action set A predefined by the ego vehicle, and the semantic action of the previous moment, which is used to ensure the continuity of the action. Assuming that the driving strategy of the ego vehicle will not change in the current decision time domain, we start from the initial belief space, where the initial state estimate is the semantic action probability P(b′) that other obstacle vehicles may perform. i |state i ). First, the strategy action set of the ego vehicle can be traversed, considering the surrounding n traffic participants and the corresponding j-th predicted trajectory (i.e., predicted path), starting from the initial moment to the end of the decision time domain, and collecting the uncertainty distribution of the traffic participants at the kth moment (corresponding to the 5th line of the algorithm), and using the corresponding strategy action to generate lateral control and longitudinal control instructions (corresponding to the 6th line of the algorithm), and the state model of the ego vehicle is updated to obtain the state of the next ego vehicle (corresponding to the 7th line of the algorithm), and the reward function of each step is calculated (corresponding to the 8th line of the algorithm). Secondly, traverse all possible reward functions, select the state sequence with a larger reward function as the decision output, and obtain a rough decision state sequence (corresponding to the 9th line of the algorithm). Finally, check whether the semantic action is safe through a safety check model (such as RSS, i.e., Responsibility Sensitive Safety). For details, please refer to technical details such as RSS, which will not be repeated here. Therefore, the method of multi-strategy closed-loop simulation evaluation can ensure real-time performance while considering the uncertainty of surrounding vehicle perception. Please refer to Figure 5 , Figure 5Schematic diagram of the process of an embodiment of the intersection driving decision method of the present application. Figure 5 As shown in the figure, the sensor model is firstly modeled, and the behavioral decision problem under the intersection working condition is modeled using the POMDP model, including the state space, action space, state transfer function and reward function, and the multi-strategy closed-loop simulation is used to solve it. In the intersection working condition, when the collision time interval is less than the safety threshold, the decision action is generated in the longitudinal self-vehicle intelligent model, and a safety check is performed at the back end, and finally a rough decision trajectory is output.

[0160] It should be noted that the above process steps are used to make driving decisions at the current moment to determine the decision information for the next moment. After the vehicle controls its movement based on the determined decision information, it comes to a new current moment, and the last current moment will be changed to the previous moment. At this time, the above process steps can be used again to make driving decisions at the new current moment to determine the decision information for the new next moment. By repeating this cycle, driving decisions can be made in real time and efficiently to drive the vehicle safely through the intersection to be decided.

[0161] The above scheme obtains the covariance matrix of traffic participants other than the vehicle at the intersection to be decided at the current moment, and the covariance matrix represents the prediction error of the future position of the traffic participants, and predicts based on the covariance matrix of the traffic participants at the current moment to obtain the position range of the traffic participants at the next moment, and then analyzes based on the planned path of the vehicle and the position range of the traffic participants at the next moment to obtain the analysis result, and the analysis result includes whether there is a collision risk between the vehicle and the traffic participants at the intersection to be decided at the next moment, so as to obtain the decision information of the vehicle based on whether the analysis result includes the existence of the collision risk, and the decision information at least includes the expected acceleration, and then on the one hand, the position range of the traffic participants at the next moment is predicted through the covariance matrix of the traffic participants, and the uncertainty of each traffic participant other than the vehicle at the current moment can be modeled, and on the other hand, the planned path of the vehicle and the position range of the traffic participants at the next moment are analyzed, and the dynamics of the traffic participants can be considered in combination with the uncertainty modeled. Therefore, uncertainty modeling can be performed at the intersection to be decided to improve the rationality of driving decisions at the intersection to be decided.

[0162] See also Figure 6 , Figure 6It is a schematic diagram of the framework of an embodiment of the intersection driving decision device of the present application. The intersection driving decision device 60 includes: an error prediction module 61, a range modeling module 62, a risk analysis module 63, and a self-vehicle decision module 64. The error prediction module 61 is used to obtain the covariance matrix of traffic participants other than the self-vehicle at the intersection to be decided at the current moment; wherein the covariance matrix represents the prediction error of the future position of the traffic participant; the range modeling module 62 is used to make a prediction based on the covariance matrix of the traffic participant at the current moment to obtain the position range of the traffic participant at the next moment; the risk analysis module 63 is used to make an analysis based on the planned path of the self-vehicle and the position range of the traffic participant at the next moment to obtain an analysis result, wherein the analysis result includes whether there is a collision risk between the self-vehicle and the traffic participant at the intersection to be decided at the next moment; the self-vehicle decision module 64 is used to obtain the decision information of the self-vehicle based on whether the analysis result includes the existence of a collision risk; wherein the decision information at least includes the expected acceleration.

[0163] In the above scheme, the intersection driving decision device 60 obtains the covariance matrix of traffic participants other than the vehicle at the intersection to be decided at the current moment, and the covariance matrix represents the prediction error of the future position of the traffic participants, and predicts based on the covariance matrix of the traffic participants at the current moment to obtain the position range of the traffic participants at the next moment, and then analyzes based on the planned path of the vehicle and the position range of the traffic participants at the next moment to obtain the analysis result, and the analysis result includes whether there is a collision risk between the vehicle and the traffic participants at the intersection to be decided at the next moment, so as to obtain the decision information of the vehicle based on whether the analysis result includes the existence of the collision risk, and the decision information at least includes the expected acceleration, and then on the one hand, the position range of the traffic participants at the next moment is predicted through the covariance matrix of the traffic participants, and the uncertainty of each traffic participant other than the vehicle at the current moment can be modeled, and on the other hand, the planned path of the vehicle and the position range of the traffic participants at the next moment are analyzed, and the dynamics of the traffic participants can be considered in combination with the uncertainty modeled. Therefore, uncertainty modeling can be performed at the intersection to be decided to improve the rationality of driving decisions at the intersection to be decided.

[0164] In some disclosed embodiments, the error prediction module 61 includes a matrix acquisition submodule, which is used to obtain the covariance matrix of the traffic participant when making driving decisions at the previous moment, and obtain the system state matrix, control input matrix and state feedback matrix; the error prediction module 61 includes a matrix update submodule, which is used to update the covariance matrix of the traffic participant at the previous moment based on the system state matrix, control input matrix and state feedback matrix, and obtain the covariance matrix of the traffic participant at the current moment.

[0165] Therefore, the covariance matrix of the traffic participant at time k+1 is obtained by updating the covariance matrix of the traffic participant at time k, without repeatedly predicting the covariance matrix based on historical collected information, which helps to reduce the computational load of obtaining the covariance matrix as much as possible.

[0166] In some disclosed embodiments, the matrix acquisition submodule includes an information acquisition unit, which is specifically used to obtain historical collection information of traffic participants before the initial moment when the previous moment is the initial moment; the matrix acquisition submodule includes a matrix prediction unit, which is used to predict the covariance matrix of the traffic participants at the initial moment based on the historical collection information of the traffic participants.

[0167] Therefore, at the initial moment, the historical collected information of the traffic participants before the initial moment is obtained, and then based on the historical collected information of the traffic participants, the covariance matrix of the traffic participants at the initial moment is predicted. The covariance matrix can be obtained through the historical collected information at the initial moment to start the driving decision at the initial moment.

[0168] In some disclosed embodiments, the matrix update submodule includes an update matrix acquisition unit, which is used to obtain a variance update matrix based on the product of a control input matrix and a state feedback matrix and the sum of a system state matrix; the matrix update submodule includes a variance matrix update unit, which is used to obtain a covariance matrix of a traffic participant at a current moment based on the product of the variance update matrix, a covariance matrix of the traffic participant at a previous moment, and a transposed matrix of the variance update matrix.

[0169] Therefore, based on the product of the control input matrix and the state feedback matrix and the sum of the system state matrix, the variance update matrix is ​​obtained, and then based on the product of the variance update matrix, the covariance matrix of the traffic participants at the previous moment and the transposed matrix of the variance update matrix, the covariance matrix of the traffic participants at the current moment is obtained. Therefore, the covariance matrix at the current moment can be updated through matrix operations, which helps to reduce the computational load of obtaining the covariance matrix as much as possible.

[0170] In some disclosed embodiments, the system state matrix is ​​used to characterize the changes in longitudinal position, longitudinal velocity, lateral position and lateral velocity over time, the control input matrix characterizes the changes in lateral acceleration and longitudinal acceleration over time, the state feedback matrix includes feedback coefficients for the longitudinal position, longitudinal velocity, lateral position and lateral velocity respectively, and the future position represented by the covariance matrix includes at least the longitudinal position and the lateral position.

[0171] In some disclosed embodiments, the prediction error is expressed as a variance value, and the range modeling module 62 includes a distribution value acquisition submodule for obtaining a target distribution value of a chi-square distribution with 2 degrees of freedom at a target quantile; the range modeling module 62 includes a range determination submodule for obtaining the position range of the traffic participant at the next moment based on the variance value and the target distribution value.

[0172] Therefore, by using the target distribution value of the chi-square distribution with a degree of freedom of 2 at the target quantile, combined with the variance value representing the prediction error, the position range at the next moment is predicted, which can reduce the complexity of uncertainty modeling as much as possible.

[0173] In some disclosed embodiments, the target distribution number is determined by a maximum tolerable probability of prediction error.

[0174] In some disclosed embodiments, the covariance matrix includes a first variance value and a second variance value, the first variance value represents a prediction error of a future longitudinal position, the second variance value represents a prediction error of a future lateral position, the range determination submodule includes a semi-axis calculation unit, for obtaining the major semi-axis of the position range at the next moment when it is represented by an elliptical area based on the first variance value and the target distribution value, and for obtaining the minor semi-axis of the position range at the next moment when it is represented by an elliptical area based on the second variance value and the target distribution value; the range determination submodule includes an ellipse determination unit, for obtaining the position range at the first moment when it is represented by an elliptical area based on the major semi-axis and the minor semi-axis.

[0175] Therefore, the covariance matrix includes a first variance value and a second variance value, the first variance value represents the prediction error of the future longitudinal position, and the second variance value represents the prediction error of the future lateral position. Based on the first variance value and the target distribution value, the major semi-axis of the position range at the next moment when it is represented by an elliptical area is obtained, and based on the second variance value and the target distribution value, the minor semi-axis of the position range at the next moment when it is represented by an elliptical area is obtained. Based on the major semi-axis and the minor semi-axis, the position range at the first moment when represented by an elliptical area is obtained, which can reduce the complexity of uncertainty modeling as much as possible.

[0176] In some disclosed embodiments, the risk analysis module 63 includes a collision position prediction submodule, which is used to determine the target position where the vehicle may collide with the traffic participant based on the intersection between the planned path and the position range of the traffic participant at the next moment; the risk analysis module 63 includes a first time calculation submodule, which is used to obtain the first time for the vehicle to reach the target position based on the first distance from the vehicle to the target position at the current moment and the extreme speed of the vehicle through the intersection to be decided; the risk analysis module 63 includes a second time calculation submodule, which is used to obtain the second time for the traffic participant to reach the target position based on the second distance from the traffic participant to the target position at the current moment and the extreme speed of the traffic participant through the intersection to be decided; the risk analysis module 63 includes an analysis result acquisition submodule, which is used to obtain the analysis result based on the relationship between the time difference between the first time and the second time and the time threshold.

[0177] Therefore, by first determining the target location where a collision may occur, and then determining whether there is a collision risk based on the difference in time taken for the vehicle and the traffic participants to reach the target location, it helps to improve the accuracy of collision risk detection.

[0178] In some disclosed embodiments, the speed extremes include a minimum speed value and a maximum speed value, and either the first time consumption or the second time consumption is the average time consumption between the longest time consumption calculated by the minimum speed value and the shortest time consumption calculated by the maximum speed value; and / or, the time consumption threshold is the smaller of the following: a preset multiple of the ratio of the first length of the vehicle and its speed extreme value, or a preset multiple of the ratio of the second length of the traffic participant and its speed extreme value.

[0179] In some disclosed embodiments, the analysis result acquisition submodule includes a first response unit for determining that the analysis result includes the presence of a collision risk in response to a time difference being less than a time threshold; the analysis result acquisition submodule includes a second response unit for determining that the analysis result includes the absence of a collision risk in response to a time difference being not less than a time threshold.

[0180] In some disclosed embodiments, when the analysis result includes that there is no collision risk, the self-vehicle decision module 64 includes a first speed ratio module, which is used to obtain a first speed ratio between the current speed of the self-vehicle and the expected speed of the current road section; the self-vehicle decision module 64 includes a first expectation calculation module, which is used to expand the first speed ratio based on the speed expansion coefficient to obtain a first expansion ratio, and obtain the product of the difference between 1 and the first expansion ratio and the maximum comfortable acceleration to obtain the expected acceleration.

[0181] Therefore, the smaller the ratio between the current speed of the vehicle and the expected speed, the larger the calculated acceleration will be, which can reduce the speed error.

[0182] In some disclosed embodiments, when the analysis result includes the existence of a collision risk, the self-vehicle decision module 64 includes a second speed ratio module, which is used to obtain a second speed ratio between the current speed of the self-vehicle and the expected speed of the current road section; the self-vehicle decision module 64 includes a driving distance ratio module, which is used to obtain the driving distance ratio between the expected distance and the actual distance between the self-vehicle and the traffic participants; the self-vehicle decision module 64 includes a proportional collision module, which is used to expand the second speed ratio based on the speed expansion coefficient to obtain a second expansion ratio, and obtain a third expansion ratio based on a preset power of the driving distance ratio; the self-vehicle decision module 64 includes a second expectation calculation module, which is used to obtain the product of the difference between 1 and the second expansion ratio, the third expansion ratio and the maximum comfortable acceleration to obtain the expected acceleration.

[0183] Therefore, the second speed ratio is expanded based on the speed expansion coefficient to obtain the second expansion ratio, and the third expansion ratio is obtained based on the preset power of the driving distance ratio. Then, the product of the difference between 1 and the second expansion ratio, the third expansion ratio and the maximum comfortable acceleration can be obtained to obtain the expected acceleration, so as to ensure that a certain safe distance is maintained between the vehicle and traffic participants as much as possible.

[0184] In some disclosed embodiments, the vehicle spacing ratio module includes a parameter acquisition submodule for acquiring a preset expected following distance, a preset expected headway, and a relative speed between the vehicle and traffic participants; the vehicle spacing ratio module includes an expected spacing submodule for obtaining an expected spacing based on the preset expected following distance, the preset expected headway, the current speed and the relative speed.

[0185] In some public embodiments, the decision information is that the decision model selects the driving strategy with the largest cumulative reward based on the belief space at the current moment, and the decision model includes the state space, action space, state transition model and observation space under the working conditions of the intersection to be decided, and the cumulative reward is calculated by a reward function composed of at least one of the safety reward function, efficiency reward function, comfort reward function, and action continuity reward function.

[0186] Therefore, driving decisions are made by combining the state space, action space, state transition model and observation space under the working conditions of the intersection to be decided, as well as the cumulative reward determined by the reward function composed of at least one of the safety reward function, efficiency reward function, comfort reward function and action continuity reward function. This can enable a comprehensive understanding of the surrounding environment through the decision-making model and the reward function, which helps to maximize the rationality of driving decisions.

[0187] In some disclosed embodiments, the state space includes a first state of the vehicle and a second state of each traffic participant, the first state includes a first position and a first speed of the vehicle in the longitudinal direction of the map coordinate system, the second state includes a second position and a second speed of the traffic participant in the longitudinal direction of the map coordinate system and a predicted path of the traffic participant.

[0188] In some disclosed embodiments, the action space includes a number of semantic instructions, and the number of semantic instructions includes at least acceleration, deceleration, and constant speed.

[0189] In some disclosed embodiments, the observation space includes a first pose of the ego-vehicle and a second pose of the traffic participant, the first pose includes the observed position and observed speed of the ego-vehicle, the second pose includes the observed position, observed speed, observed heading angle and predicted probability of several candidate paths of the traffic participant in the global coordinate system, and the several candidate paths include going straight, turning left and turning right.

[0190] In some disclosed embodiments, the state transition model is defined as the product of the first transition probability of the vehicle and the cumulative product of the second transition probabilities of each traffic participant, the first transition probability represents the state information of the vehicle transferred to the next moment under the state information condition of the current moment, the second transition probability represents the state information of the traffic participant transferred to the next moment under the state information condition of the current moment, and the second transition probability is defined as the product of the first transfer sub-probability and the second transfer sub-probability, the first transfer sub-probability represents the candidate path of the traffic participant transferred to the next moment under the state information condition of the current moment, the second transfer sub-probability represents the system state of the traffic participant transferred to the next moment under the state information condition of the current moment, and the system state includes at least one of the longitudinal position, longitudinal speed, lateral position, and lateral speed.

[0191] Therefore, by defining the second transfer probability as the product of the first transfer sub-probability and the second transfer sub-probability, the second transfer probability can be split into two parts: driving intention transfer and motion state transfer, thereby improving the expression accuracy of the second transfer probability.

[0192] In some disclosed embodiments, the safety reward function represents a first penalty value determined based on an assessed collision probability based on a future state of the vehicle and a future state of a traffic participant, and when the collision probability is lower than a probability threshold, the first penalty value is the product of a difference between 1 and the collision probability and a safety weight coefficient, and when the collision probability is not lower than the probability threshold, the first penalty value is a preset value, and the absolute value of the preset value is greater than the maximum first penalty value when the collision probability is lower than the probability threshold.

[0193] Therefore, when there is a high probability of a collision, a corresponding cost penalty can be given, which can force the decision model to avoid choosing decision information with a high probability of a collision as much as possible.

[0194] In some disclosed embodiments, the efficiency reward function determines a second penalty value based on the speed difference between the current speed of the vehicle and the map speed limit, and when the current speed is greater than the map speed limit, the second penalty value is the product of the square of the speed difference and the first efficiency weight coefficient, and when the current speed is not greater than the map speed limit, the second penalty value is the product of the absolute value of the speed difference and the second efficiency weight coefficient.

[0195] Therefore, the efficiency reward function represents that the smart car (i.e., the ego car) wants to pass the intersection (i.e., the decision-making intersection) as quickly as possible. By considering the difference between the ego car speed limit and the map speed limit, the ego car can stay as close to the map limit as possible while meeting safety requirements.

[0196] In some disclosed embodiments, the comfort reward function determines a third penalty value based on a change value of the longitudinal acceleration of the vehicle, and the third penalty value is a product of a square of the change value and a comfort weight coefficient.

[0197] Therefore, the penalty value is determined based on the change value of the longitudinal acceleration of the vehicle to obtain the comfort reward function, and the penalty value is specifically the product of the square of the change value and the comfort weight coefficient. Therefore, it can force the decision model to select decision information with small longitudinal acceleration as much as possible while taking other indicators into consideration, which helps to improve the comfort of passing the intersection to be decided.

[0198] In some disclosed embodiments, the action continuity reward function determines a reward value based on whether the longitudinal acceleration of the self-vehicle is the same at adjacent moments, and when the longitudinal acceleration of the self-vehicle is the same at adjacent moments, the reward value is a preset cost weight coefficient, and when the longitudinal acceleration of the self-vehicle is different at adjacent moments, the reward value is zero.

[0199] Therefore, the reward value is determined based on whether the longitudinal acceleration of the ego vehicle is the same at adjacent moments to obtain the action continuity reward function, and when the longitudinal acceleration of the ego vehicle is the same at adjacent moments, the reward value is the preset cost weight coefficient, and when the longitudinal acceleration of the ego vehicle is different at adjacent moments, the reward value is zero. Therefore, the decision model can be forced to select decision information with a larger reward value as much as possible while taking other indicators into account, that is, decision information that can make the longitudinal acceleration the same at adjacent moments, which helps to improve the continuity of the semantic actions of the ego vehicle at adjacent moments.

[0200] See also Figure 7 , Figure 7: is a schematic diagram of a framework of an embodiment of an electronic device 70 of the present application. The electronic device 70 includes: a memory 71 and a processor 72 coupled to each other, the memory 71 stores program instructions, and the processor 72 is used to execute the program instructions to implement the steps in any of the above-mentioned intersection driving decision-making method embodiments. Specifically, the electronic device 70 may include but is not limited to: a desktop computer, a laptop computer, a server, a mobile phone, a tablet computer, a car computer, etc., which are not limited here.

[0201] Specifically, the processor 72 is used to control itself and the memory 71 to implement the steps in any of the above-mentioned intersection driving decision-making method embodiments. The processor 72 can also be called a CPU (Central Processing Unit). The processor 72 may be an integrated circuit chip with signal processing capabilities. The processor 72 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 72 can be implemented by an integrated circuit chip.

[0202] In the above scheme, the electronic device 70 obtains the covariance matrix of the traffic participants other than the vehicle at the intersection to be decided at the current moment, and the covariance matrix represents the prediction error of the future position of the traffic participants, and predicts based on the covariance matrix of the traffic participants at the current moment to obtain the position range of the traffic participants at the next moment, and then analyzes based on the planned path of the vehicle and the position range of the traffic participants at the next moment to obtain the analysis result, and the analysis result includes whether there is a collision risk between the vehicle and the traffic participants at the intersection to be decided at the next moment, so as to obtain the decision information of the vehicle based on whether the analysis result includes the existence of the collision risk, and the decision information at least includes the expected acceleration, and then on the one hand, the position range of the traffic participants at the next moment is predicted through the covariance matrix of the traffic participants, and the uncertainty of each traffic participant other than the vehicle at the current moment can be modeled, and on the other hand, the planned path of the vehicle and the position range of the traffic participants at the next moment are analyzed, and the dynamics of the traffic participants can be considered in combination with the uncertainty modeled. Therefore, uncertainty modeling can be performed at the intersection to be decided to improve the rationality of driving decisions at the intersection to be decided.

[0203] See also Figure 8 , Figure 8The schematic diagram of the framework of an embodiment of a computer-readable storage medium 80 of the present application. The computer-readable storage medium 80 stores program instructions 81 that can be executed by a processor, and the program instructions 81 are used to implement the steps in any of the above-mentioned intersection driving decision-making method embodiments.

[0204] In the above scheme, the computer-readable storage medium 80 obtains the covariance matrix of traffic participants other than the vehicle at the intersection to be decided at the current moment, and the covariance matrix represents the prediction error of the future position of the traffic participants, and predicts based on the covariance matrix of the traffic participants at the current moment to obtain the position range of the traffic participants at the next moment, and then analyzes based on the planned path of the vehicle and the position range of the traffic participants at the next moment to obtain the analysis result, and the analysis result includes whether there is a collision risk between the vehicle and the traffic participants at the intersection to be decided at the next moment, so as to obtain the decision information of the vehicle based on whether the analysis result includes the existence of the collision risk, and the decision information at least includes the expected acceleration, and then on the one hand, the position range of the traffic participants at the next moment is predicted through the covariance matrix of the traffic participants, and the uncertainty of each traffic participant other than the vehicle at the current moment can be modeled, and on the other hand, the planned path of the vehicle and the position range of the traffic participants at the next moment are analyzed, and the dynamics of the traffic participants can be considered in combination with the uncertainty modeled. Therefore, uncertainty modeling can be performed at the intersection to be decided to improve the rationality of driving decisions at the intersection to be decided.

[0205] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0206] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0207] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0208] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

[0209] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A method for making decisions on driving at an intersection, characterized in that: include: Obtaining a covariance matrix of traffic participants other than the vehicle at the intersection to be decided at the current moment; wherein the covariance matrix represents the prediction error of the future position of the traffic participant; Predicting based on the covariance matrix of the traffic participant at the current moment, obtaining the position range of the traffic participant at the next moment; Analyze the planned path of the vehicle and the position range of the traffic participant at the next moment to obtain an analysis result, wherein the analysis result includes whether there is a collision risk between the vehicle and the traffic participant at the intersection to be decided at the next moment; Based on whether the analysis result includes the existence of a collision risk, the decision information of the vehicle is obtained; wherein the decision information at least includes an expected acceleration.

2. The method according to claim 1, characterized in that The obtaining of the covariance matrix of traffic participants other than the vehicle at the intersection to be decided at the current moment includes: Obtaining the covariance matrix of the traffic participant when making a driving decision at the last moment, and obtaining the system state matrix, control input matrix and state feedback matrix; The covariance matrix of the traffic participant at the previous moment is updated based on the system state matrix, the control input matrix and the state feedback matrix to obtain the covariance matrix of the traffic participant at the current moment.

3. The method according to claim 2, characterized in that In the case where the previous moment is the initial moment, the obtaining of the covariance matrix of the traffic participant when making a driving decision at the previous moment includes: Acquire historical collected information of the traffic participant before the initial moment; Based on the historically collected information of the traffic participants, the covariance matrix of the traffic participants at the initial moment is predicted.

4. The method according to claim 2, characterized in that: The updating of the covariance matrix of the traffic participant at the previous moment based on the system state matrix, the control input matrix and the state feedback matrix to obtain the covariance matrix of the traffic participant at the current moment includes: Obtaining a variance update matrix based on the product of the control input matrix and the state feedback matrix and the sum of the system state matrix; The covariance matrix of the traffic participant at the current moment is obtained based on the product of the variance update matrix, the covariance matrix of the traffic participant at the previous moment, and the transposed matrix of the variance update matrix.

5. The method according to claim 2, characterized in that: The system state matrix is ​​used to characterize the changes of longitudinal position, longitudinal velocity, lateral position and lateral velocity over time, the control input matrix characterizes the changes of lateral acceleration and longitudinal acceleration over time, the state feedback matrix includes feedback coefficients of longitudinal position, longitudinal velocity, lateral position and lateral velocity respectively, and the future position characterized by the covariance matrix includes at least longitudinal position and lateral position.

6. The method according to claim 1, characterized in that The prediction error is represented by a variance value, and the prediction based on the covariance matrix of the traffic participant at the current moment to obtain the position range of the traffic participant at the next moment includes: Get the target distribution value of the chi-square distribution with 2 degrees of freedom at the target quantile; Based on the variance value and the target distribution value, the position range of the traffic participant at the next moment is obtained.

7. The method according to claim 6, characterized in that The target distribution number is determined by the maximum tolerable probability of the prediction error.

8. The method according to claim 6, characterized in that The covariance matrix includes a first variance value and a second variance value, wherein the first variance value represents a prediction error of a future longitudinal position, and the second variance value represents a prediction error of a future lateral position. The position range of the traffic participant at the next moment is obtained based on the variance value and the target distribution value, including: Based on the first variance value and the target distribution value, obtaining the major semi-axis when the position range at the next moment is represented by an elliptical area, and based on the second variance value and the target distribution value, obtaining the minor semi-axis when the position range at the next moment is represented by an elliptical area; Based on the major semi-axis and the minor semi-axis, the position range at the first moment when represented by an elliptical area is obtained.

9. The method according to claim 1, characterized in that: The analysis is performed based on the planned path of the vehicle and the position range of the traffic participant at the next moment to obtain the analysis result, including: Determine a target position where the ego vehicle and the traffic participant may collide based on an intersection between the planned path and the position range of the traffic participant at the next moment; Based on the first distance from the vehicle to the target location at the current moment and the extreme speed of the vehicle passing through the intersection to be decided, a first time taken for the vehicle to reach the target location is obtained; and Based on a second distance from the traffic participant to the target location at the current moment and an extreme value of the speed of the traffic participant passing through the to-be-decided intersection, a second time taken for the traffic participant to reach the target location is obtained; The analysis result is obtained based on the relationship between the time consumption difference between the first time consumption and the second time consumption and the time consumption threshold.

10. The method according to claim 9, characterized in that The speed extreme value includes a minimum speed value and a maximum speed value, and either the first time consumption or the second time consumption is an average time consumption between the longest time consumption calculated by the minimum speed value and the shortest time consumption calculated by the maximum speed value; And / or, the time-consuming threshold is the smaller one of the following: a preset multiple of the ratio of the first length of the vehicle and its speed limit, and a preset multiple of the ratio of the second length of the traffic participant and its speed limit.

11. The method according to claim 9, characterized in that The obtaining of the analysis result based on the relationship between the time consumption difference between the first time consumption and the second time consumption and the time consumption threshold value includes: In response to the time-consuming difference being less than the time-consuming threshold, determining that the analysis result includes the presence of a collision risk; In response to the time-consuming difference being not less than the time-consuming threshold, determining that the analysis result includes no collision risk.

12. The method according to claim 1, characterized in that In the case that the analysis result includes that there is no collision risk, obtaining the decision information of the vehicle based on whether the analysis result includes that there is a collision risk includes: Obtaining a first speed ratio between the current speed of the vehicle and the expected speed of the current road section; The first speed ratio is expanded based on a speed expansion coefficient to obtain a first expansion ratio, and the product of a difference between 1 and the first expansion ratio and a maximum comfortable acceleration is obtained to obtain the expected acceleration.

13. The method according to claim 1, characterized in that In the case where the analysis result includes the existence of a collision risk, obtaining the decision information of the vehicle based on whether the analysis result includes the existence of a collision risk includes: Obtaining a second speed ratio between the current speed of the vehicle and the expected speed of the current road section, and obtaining a driving distance ratio between the expected distance and the actual distance between the vehicle and the traffic participant; Expanding the second speed ratio based on a speed expansion coefficient to obtain a second expansion ratio, and obtaining a third expansion ratio based on a preset power of the vehicle spacing ratio; The expected acceleration is obtained by obtaining the product of the difference between 1 and the second expansion ratio, the third expansion ratio and the maximum comfortable acceleration.

14. The method according to claim 13, characterized in that The step of obtaining the desired spacing includes: Obtaining a preset expected following distance, a preset expected headway, and a relative speed between the vehicle and the traffic participant; The expected spacing is obtained based on the preset expected following distance, the preset expected headway, the current speed and the relative speed.

15. The method according to any one of claims 1 to 14, characterized in that The decision information is that the decision model selects the driving strategy with the largest cumulative reward based on the belief space at the current moment, and the decision model includes the state space, action space, state transition model and observation space under the working conditions of the intersection to be decided, and the cumulative reward is calculated by a reward function composed of at least one of a safety reward function, an efficiency reward function, a comfort reward function, and an action continuity reward function.

16. The method according to claim 15, characterized in that The state space includes a first state of the vehicle and a second state of each of the traffic participants, the first state includes a first position and a first speed of the vehicle in the longitudinal direction of the map coordinate system, the second state includes a second position and a second speed of the traffic participant in the longitudinal direction of the map coordinate system and a predicted path of the traffic participant; And / or, the action space includes a plurality of semantic instructions, and the plurality of semantic instructions at least include acceleration, deceleration, and uniform speed; And / or, the observation space includes a first posture of the ego vehicle and a second posture of the traffic participant, the first posture includes the observed position and observed speed of the ego vehicle, the second posture includes the observed position, observed speed, observed heading angle and predicted probabilities of several candidate paths of the traffic participant in a global coordinate system, and the several candidate paths include going straight, turning left and turning right; And / or, the state transition model is defined as the product of the first transition probability of the vehicle and the cumulative product of the second transition probabilities of each of the traffic participants, the first transition probability represents the state information of the vehicle transferred to the next moment under the state information of the current moment, the second transition probability represents the state information of the traffic participant transferred to the next moment under the state information of the current moment, and the second transition probability is defined as the product of the first transfer sub-probability and the second transfer sub-probability, the first transfer sub-probability represents the candidate path of the traffic participant transferred to the next moment under the state information of the current moment, the second transfer sub-probability represents the system state of the traffic participant transferred to the next moment under the state information of the current moment, and the system state includes at least one of the longitudinal position, longitudinal speed, lateral position, and lateral speed.

17. The method according to claim 15, characterized in that The safety reward function represents a first penalty value determined based on the collision probability evaluated based on the future state of the vehicle and the future state of the traffic participant, and when the collision probability is lower than a probability threshold, the first penalty value is the product of a difference between 1 and the collision probability and a safety weight coefficient, and when the collision probability is not lower than the probability threshold, the first penalty value is a preset value, and the absolute value of the preset value is greater than the maximum first penalty value when the collision probability is lower than the probability threshold; And / or, the efficiency reward function determines a second penalty value based on a speed difference between the current speed of the vehicle and a map speed limit, and when the current speed is greater than the map speed limit, the second penalty value is the product of the square of the speed difference and a first efficiency weight coefficient, and when the current speed is not greater than the map speed limit, the second penalty value is the product of the absolute value of the speed difference and a second efficiency weight coefficient; And / or, the comfort reward function is based on a third penalty value determined based on a change value of the longitudinal acceleration of the vehicle, and the third penalty value is a product of a square value of the change value and a comfort weight coefficient; And / or, the action continuity reward function determines a reward value based on whether the longitudinal acceleration of the self-vehicle is the same at adjacent moments, and when the longitudinal acceleration of the self-vehicle is the same at adjacent moments, the reward value is a preset cost weight coefficient, and when the longitudinal acceleration of the self-vehicle is different at adjacent moments, the reward value is zero.

18. A driving decision device at an intersection, characterized in that: include: The error prediction module is used to obtain the covariance matrix of traffic participants other than the vehicle at the intersection to be decided at the current moment; wherein the covariance matrix represents the prediction error of the future position of the traffic participant; A range modeling module, used for predicting based on the covariance matrix of the traffic participant at the current moment to obtain the position range of the traffic participant at the next moment; A risk analysis module, configured to analyze the planned path of the vehicle and the position range of the traffic participant at the next moment to obtain an analysis result, wherein the analysis result includes whether there is a collision risk between the vehicle and the traffic participant at the intersection to be decided at the next moment; The ego vehicle decision module is used to obtain decision information of the ego vehicle based on whether the analysis result includes the existence of a collision risk; wherein the decision information at least includes an expected acceleration.

19. An electronic device, characterized in that: It comprises a memory and a processor coupled to each other, the memory stores program instructions, and the processor is used to execute the program instructions to implement the intersection driving decision method as described in any one of claims 1 to 17.

20. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the intersection driving decision method described in any one of claims 1 to 17.