Decision-making method, device, electronic device and medium for interacting with reverse obstacles
Through the game theory analysis method of the Level-k framework, the optimal behavior and interaction rewards of reverse obstacles at different decision levels are evaluated, which solves the problem of stable decision-making of autonomous vehicles when encountering reverse obstacles, realizes safer and more flexible vehicle decision-making, and is suitable for the simultaneous processing of multiple reverse obstacles.
Patent Information
- Application Number
- CN202311628071.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-11-30
AI Technical Summary
Existing autonomous driving technology has difficulty making stable and continuous decisions when oncoming obstacles appear, especially the inability to effectively handle longitudinal and lateral avoidance behaviors of oncoming obstacles, resulting in overly conservative or inappropriate vehicle behavior.
The game theory analysis method of the Level-k framework is adopted to determine the decision-making behavior of the vehicle and the reverse obstacle by evaluating the optimal behavior and interaction rewards of the vehicle and the reverse obstacle at different decision levels. The multiple possibilities of the reverse obstacle and the behavioral rewards of the vehicle are considered to achieve stable decision-making.
It improves the decision-making stability and safety of autonomous vehicles when interacting with adverse obstacles, can adapt to the simultaneous processing of multiple adverse obstacles, meets real-time requirements, and the decision-making behavior is closer to the intention of human drivers.
Smart Images

Figure CN117622215B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of autonomous driving technology, and in particular to a decision-making method, device, electronic device, and medium for interacting with a reverse obstacle. Background Art
[0002] With the development of autonomous driving technology, people's requirements for the comprehensive performance of autonomous vehicles are gradually increasing. As a key process that directly affects the safety and comfort of autonomous vehicles, the study of obstacle interaction has gradually become a key part of the autonomous driving field.
[0003] During normal driving, autonomous vehicles (hereinafter referred to as vehicles) often encounter interactions with other traffic participants, in addition to following and overtaking other vehicles. These other traffic participants can be aggressive or conservative, and their behavior is highly uncertain. Therefore, it's impossible to handle interactions in various scenarios simply by assuming the behavior of obstacles.
[0004] Furthermore, existing solutions struggle to achieve continuous decision-making interactions. For the continuous presence of traffic participants, it's impossible to balance stable, continuous, and decisive decision-making during the interaction process. Compared to other situations, the higher relative speed of oncoming obstacles poses a greater risk to vehicles. Furthermore, when oncoming obstacles appear, vehicles can't simply yield or cut in, and they can't simply consider the acceleration and deceleration of the obstacle; lateral avoidance must also be considered. Therefore, the aforementioned issues are even more pronounced for oncoming obstacles. Summary of the Invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present disclosure provide a decision-making method, device, electronic device and medium for interacting with an adverse obstacle, which realizes the decision-making when the vehicle interacts with the adverse obstacle, and can achieve the final stable decision through continuous interaction in the process of the vehicle gradually approaching the adverse obstacle.
[0006] In a first aspect, an embodiment of the present disclosure provides a decision-making method for interacting with an inverse obstacle, the method comprising:
[0007] For each oncoming obstacle corresponding to the vehicle, determining an optimal obstacle behavior for the oncoming obstacle at each decision level based on a current vehicle state of the vehicle and a current obstacle state of the oncoming obstacle;
[0008] Determining a reward matrix between the vehicle and the oncoming obstacle based on the optimal obstacle behavior at each decision level and the current vehicle state, wherein the reward matrix includes vehicle interaction rewards between each to-be-evaluated behavior of the vehicle and the optimal obstacle behavior at each decision level;
[0009] Based on a reward matrix between the vehicle and each oncoming obstacle, a vehicle decision behavior of the vehicle is determined among all behaviors to be evaluated, and an obstacle decision behavior of each oncoming obstacle is determined according to the vehicle decision behavior and each reward matrix.
[0010] In a second aspect, an embodiment of the present disclosure further provides a decision-making device for interacting with a reverse obstacle, the device comprising:
[0011] a level behavior determination module, configured to determine, for each oncoming obstacle corresponding to the vehicle, an optimal obstacle behavior for the oncoming obstacle at each decision level based on a current vehicle state of the vehicle and a current obstacle state of the oncoming obstacle;
[0012] a reward matrix determination module, configured to determine a reward matrix between the vehicle and the oncoming obstacle based on the optimal obstacle behavior at each decision level and the current vehicle state, wherein the reward matrix includes vehicle interaction rewards between each to-be-evaluated vehicle behavior and the optimal obstacle behavior at each decision level;
[0013] A decision module is configured to determine a vehicle decision behavior of the vehicle among all behaviors to be evaluated based on a reward matrix between the vehicle and each oncoming obstacle, and to determine an obstacle decision behavior of each oncoming obstacle based on the vehicle decision behavior and each reward matrix.
[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; and when the one or more programs are executed by the one or more processors, the one or more processors implement the decision-making method for interacting with a reverse obstacle as described above.
[0015] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the decision-making method for interacting with a reverse obstacle as described above.
[0016] The embodiment of the present disclosure provides a decision-making method for interacting with a reverse obstacle. For each reverse obstacle corresponding to a vehicle, the optimal obstacle behavior of the reverse obstacle at each decision level is determined according to the current vehicle state and the current obstacle state. Then, based on the optimal obstacle behavior at each decision level and the current vehicle state, the vehicle interaction reward between each to-be-evaluated behavior of the vehicle and the optimal obstacle behavior at each decision level is determined. In this way, the optimal behavior of the reverse obstacle at different decision levels is considered, and the behavior of the vehicle is evaluated respectively. Finally, the vehicle decision behavior is determined from all to-be-evaluated behaviors according to the reward matrix between the vehicle and all reverse obstacles, so as to determine the vehicle decision behavior according to the vehicle decision behavior and all reward matrix. The matrix is reported to obtain the obstacle decision behavior of each adverse obstacle, and the decision-making when the vehicle interacts with the adverse obstacle is realized. By analyzing the optimal behavior of the adverse obstacle at different decision levels, the intention of the adverse obstacle can be gradually deduced, making the intention of the adverse obstacle clearer, thereby making the decision-making behavior of the vehicle more stable and safe. The intention of the adverse obstacle is variable and more flexible. Compared with the method of only changing the longitudinal or lateral intention, it can be closer to the intention of a real human driver. The method provided in the embodiment of the present disclosure can be applied to the synchronous processing of multiple adverse obstacles, meet real-time requirements, and can ensure the output of reasonable decision-making behavior when the vehicle interacts with multiple adverse obstacles. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0018] Figure 1 is a flowchart of a decision-making method for interacting with a reverse obstacle in an embodiment of the present disclosure;
[0019] Figure 2 is a schematic diagram of a reverse obstacle in an embodiment of the present disclosure;
[0020] Figure 3 A schematic diagram of a confidence region provided in an embodiment of the present disclosure;
[0021] Figure 4 Schematic diagram of a decision-making process in an embodiment of the present disclosure;
[0022] Figure 5 Schematic diagram of the structure of a decision-making device for interacting with a reverse obstacle in an embodiment of the present disclosure;
[0023] Figure 6 Schematic diagram of the structure of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0025] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0026] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0027] Before introducing in detail the decision-making method for interacting with reverse obstacles provided by the embodiments of the present disclosure, the technical problems solved by the method are first described.
[0028] In existing technologies, autonomous vehicles often adopt conservative behavior when dealing with oncoming obstacles due to unclear intentions. For example, they design their behavior based on the assumption that the oncoming obstacle is moving toward the vehicle at a constant speed. However, this often results in overly conservative behavior, making it impossible for the vehicle to operate normally in complex traffic scenarios.
[0029] To address the above-mentioned issues, the disclosed embodiments provide a decision-making method for interacting with an oncoming obstacle. This method takes into account that it is difficult for human drivers to immediately identify the intention of the obstacle and then take action when interacting with the oncoming obstacle. Human drivers usually guess the intention of the obstacle while driving and select the optimal action based on the guessed intention of the obstacle. In the process of guessing intention -> selecting action, the oncoming obstacle is also unable to confirm the intention of the human driver. Therefore, the oncoming obstacle also adopts the method of guessing intention -> selecting action. As the vehicle and the obstacle gradually approach, the two achieve a final stable decision through continuous interaction.
[0030] Therefore, the method provided by the embodiment of the present disclosure can use the Level-k framework to make decisions and interact with obstacles. Level-k is a game theory analysis method whose basic assumption is that players have different levels of thinking. Based on this assumption, the thinking levels of players are divided into different levels, represented as k levels. In the Level-k game theory, each player is assigned to a level, from level 1 to level k. Level 1 players are the simplest players, they only consider their own actions and do not consider the reactions of other players. Level 2 players consider the reactions of other players, but only consider the level 1 reactions of other players. Level 3 players consider the level 2 reactions of other players, and so on. The highest level players are called perfectly rational players, they consider the reactions of all other players, as well as the reactions of all reactions of other players.
[0031] In autonomous driving interactions, the Level-K concept can be applied to interactive scenarios. Assuming that the reverse obstacle has k levels of possibility, that is, k decision-making levels, a comprehensive approach is used to evaluate the rewards of each vehicle behavior in the current environment. The rewards are then used to further determine the final vehicle and obstacle decision-making behaviors.
[0032] Figure 1 This is a flow chart of a decision-making method for interacting with a reverse obstacle in an embodiment of the present disclosure. This method can be executed by a decision-making device for interacting with a reverse obstacle, which can be implemented in software and / or hardware and can be configured in an electronic device. Figure 1 As shown, the method may specifically include the following steps:
[0033] S110 . For each oncoming obstacle corresponding to the vehicle, determine an optimal obstacle behavior for the oncoming obstacle at each decision level based on the current vehicle state of the vehicle and the current obstacle state of the oncoming obstacle.
[0034] Among them, the reverse obstacle can be an obstacle with a collision point between the predicted trajectory and the vehicle trajectory, and the collision angle at the collision point is within a set angle range. The set angle range can be set according to a collision angle of 180°, such as 170° to 190°.
[0035] Figure 2 is a schematic diagram of a reverse obstacle in an embodiment of the present disclosure, such as Figure 2 As shown in Figure 2, the predicted trajectories of the obstacles in (a) to (e) all have collision points with the vehicle's trajectory, and the collision angle is close to 180°. Therefore, the obstacles in (a) to (e) are all reverse obstacles.
[0036] Specifically, the reverse obstacles can be screened from all obstacles based on the predicted trajectories of the obstacles and the trajectory of the vehicle. It should be noted that the number of the screened reverse obstacles can be one or more.
[0037] In the embodiment of the present disclosure, for an oncoming obstacle, its behavior at the lowest decision level can be defined as ignoring the other party's state, that is, at the lowest decision level, the oncoming obstacle can consider the vehicle to be absolutely stationary or that the vehicle does not exist, without considering the impact of the vehicle's behavior changes on the oncoming obstacle.
[0038] For a vehicle, its behavior at the lowest decision level can be defined as the behavior of the vehicle considering that the oncoming obstacle is in a uniform deceleration state. Considering that the oncoming obstacle is in a uniform deceleration state, the theoretical stopping time and parking position of the oncoming obstacle can be determined (assuming that the oncoming obstacle stops at the moment of collision with the vehicle).
[0039] Specifically, for each oncoming obstacle, the current vehicle state of the vehicle and the current obstacle state of the oncoming obstacle can be combined, starting from the lowest decision level (Level-0), and the optimal obstacle behavior of the oncoming obstacle at each decision level can be derived until the optimal obstacle behavior at the highest decision level is derived.
[0040] The current vehicle state may include the current position and current speed of the vehicle, and the current obstacle state may include the current position and current speed of the oncoming obstacle.
[0041] In a specific embodiment, determining the optimal obstacle behavior of the oncoming obstacle at each decision level based on the current vehicle state of the vehicle and the current obstacle state of the oncoming obstacle includes the following steps:
[0042] Step 111: Obtaining the optimal obstacle behavior of the reverse obstacle at the lowest decision level and the optimal vehicle behavior of the vehicle at the lowest decision level;
[0043] Step 112: The next level after the lowest decision level is used as the current level. Based on the current obstacle state and the optimal vehicle behavior at the level before the current level, the behavior of the oncoming obstacle at the current level is sampled and evaluated, and the optimal obstacle behavior at the current level is obtained.
[0044] Step 113: Based on the current vehicle state and the optimal obstacle behavior of the oncoming obstacle at the previous level of the current level, the vehicle behavior at the current level is sampled and reported for evaluation to obtain the optimal vehicle behavior at the current level.
[0045] Step 114 : Set the next level after the current level as the new current level, and return to the step of sampling and reporting the behavior of the reverse obstacle at the current level until the current level reaches the highest decision level.
[0046] Taking any oncoming obstacle as an example, specifically, in step 111, it can be assumed that the oncoming obstacle is at the lowest decision level (Level-0), and the vehicle is at the next level after the lowest decision level (Level-1). At this time, the optimal obstacle behavior of the oncoming obstacle at the lowest decision level can be obtained, wherein the optimal obstacle behavior at Level-0 can be pre-set.
[0047] Furthermore, in step 111 , it may be assumed that the vehicle is at the lowest decision level (Level-0), at which point the optimal vehicle behavior at the lowest decision level may be obtained, wherein the optimal vehicle behavior at Level-0 may be pre-set.
[0048] Furthermore, in step 112, the next level (Level-1) after the lowest decision level (Level-0) can be used as the current level, and then assuming that the reverse obstacle is at the current level, the behavior of the reverse obstacle at Level-1 is sampled and reported for evaluation.
[0049] When the reverse obstacle is at Level-1, the vehicle is considered to be at the previous level (Level-0). Therefore, it is necessary to sample and report the behavior of the reverse obstacle at Level-1 based on the vehicle's optimal vehicle behavior at Level-0 and the current obstacle status to obtain the optimal obstacle behavior at Level-1.
[0050] Among them, the reward evaluation of the sampling behavior of the reverse obstacle can be understood as calculating the benefits of the sampling behavior of the reverse obstacle. For example, the reward of the sampling behavior of the reverse obstacle can be analyzed from the perspectives of safety, time utility, and whether it deviates from the road network.
[0051] Optionally, based on the current obstacle state and the vehicle's optimal vehicle behavior at the previous level, the behavior of the reverse obstacle at the current level is sampled and evaluated, including:
[0052] The behavior of the oncoming obstacle at the current level is sampled to obtain multiple obstacle sample behaviors. Each obstacle sample behavior is evaluated based on the current obstacle state, the obstacle evaluation function, and the optimal vehicle behavior at the previous level to obtain an obstacle interaction reward for each obstacle sample behavior. Based on each obstacle interaction reward, the optimal obstacle behavior at the current level is determined among all obstacle sample behaviors. The obstacle evaluation function includes a safety evaluation function, a time utility evaluation function, and a road network deviation evaluation function.
[0053] Specifically, the obstacle sampling behavior may include the longitudinal acceleration and yaw rate of the oncoming obstacle. To ensure the rationality of the sampling, the oncoming obstacle behavior may be sampled within the confidence region of the predicted trajectory of the oncoming obstacle.
[0054] In one example, the method provided by an embodiment of the present disclosure further includes: obtaining a predicted trajectory of an oncoming obstacle, wherein the predicted trajectory includes trajectory points corresponding to a plurality of time instants; determining an uncertainty of each trajectory point on the predicted trajectory, and determining a confidence region for each trajectory point based on the uncertainty of each trajectory point;
[0055] Accordingly, the behavior of the reverse obstacle at the current level is sampled to obtain a plurality of obstacle sampling behaviors, including: sampling the behavior of the reverse obstacle within the corresponding confidence region to obtain a plurality of obstacle sampling behaviors.
[0056] Specifically, the predicted trajectory of the reverse obstacle may be generated based on the current obstacle state of the reverse obstacle, the relationship between the reverse obstacle and other obstacles, and the relationship between the reverse obstacle and the road network.
[0057] Furthermore, a corresponding confidence region can be generated based on the predicted trajectory, wherein the confidence region describes the possible deviation range of the trajectory point on the predicted trajectory, and the center of the confidence region is the position with the highest probability of the trajectory point. Figure 3 A schematic diagram of a confidence region provided in an embodiment of the present disclosure, such as Figure 3 As shown, the confidence region can be defined as an elliptical region. In the embodiment of the present disclosure, the confidence region can be generated according to the uncertainty of the trajectory points on the predicted trajectory.
[0058] Assume that each trajectory point on the predicted trajectory obeys a two-dimensional Gaussian distribution:
[0059]
[0060] Where p(x) represents the probability distribution of the trajectory point x on the predicted trajectory, the mean μ can be the position of the trajectory point on the predicted trajectory at the same time, and the goal of generating uncertainty is to determine the covariance matrix ∑ in the above formula. Let the above formula be equal to the constant c, then:
[0061]
[0062] Based on the above formula, the covariance matrix ∑ can be calculated according to the probability distribution of the trajectory points. The uncertainty of the trajectory points is determined by the covariance matrix ∑, and then the confidence region of the trajectory points is generated according to the uncertainty of the trajectory points.
[0063] When generating the uncertainty and confidence region of trajectory points, the following principles must also be met: 1. Uncertainty increases with time, because the longer the prediction time, the less accurate the prediction result; 2. The uncertainty and confidence region of adjacent moments must be greater than the previous moment; 3. Without violating principles 1 and 2, the confidence region should not include static obstacles as much as possible; 4. Without violating principles 1, 2, and 3, the greater the uncertainty in the event of collision with other obstacles or vehicles, the larger the confidence region.
[0064] By generating a confidence region for each trajectory point in the predicted trajectory and then sampling the longitudinal acceleration and yaw angular velocity within the confidence region, the rationality of the sampling behavior of the adverse obstacle can be ensured, thereby further ensuring the reliability of the decision.
[0065] After sampling the behaviors of the reverse obstacles, the safety evaluation function, the time utility evaluation function, and the road network deviation evaluation function can be used to evaluate each sampled obstacle behavior, so as to select the obstacle sampling behavior with the highest interaction return as the optimal obstacle behavior at the current level.
[0066] Furthermore, in step 113, it is assumed that the vehicle is at the current level (Level-1) and the vehicle's behavior at Level-1 is sampled and reported. When the vehicle is at Level-1, the oncoming obstacle is assumed to be at the level before the current level (Level-0). Therefore, based on the optimal obstacle behavior for the oncoming obstacle at Level-0 and the current vehicle state, the vehicle's behavior at Level-1 is sampled, reported, and evaluated to determine the optimal vehicle behavior at Level-1.
[0067] Among them, the reward evaluation of the vehicle's sampling behavior can be understood as calculating the benefits of the vehicle's sampling behavior. For example, the reward of the vehicle's sampling behavior can be analyzed from the perspectives of safety, time utility, and whether it is driving in the wrong direction.
[0068] Optionally, based on the current vehicle state and the optimal obstacle behavior of the opposite obstacle at the previous level of the current level, the vehicle behavior at the current level is sampled and reported back to evaluate, and the optimal vehicle behavior at the current level is obtained, including:
[0069] The vehicle's behavior at the current level is sampled to obtain multiple vehicle sample behaviors; each vehicle sample behavior is evaluated based on the current vehicle state, the vehicle evaluation function, and the optimal obstacle behavior at the previous level to obtain a vehicle interaction reward for each vehicle sample behavior; based on each vehicle interaction reward, the optimal vehicle behavior at the current level is determined among all vehicle sample behaviors; wherein the vehicle evaluation function includes a safety evaluation function, a time utility evaluation function, and a retrograde evaluation function.
[0070] Specifically, the sampled vehicle behaviors may include the vehicle's longitudinal acceleration and yaw rate. After sampling the vehicle behaviors, each sampled vehicle behavior can be evaluated using a safety evaluation function, a time utility evaluation function, and a retrograde evaluation function. The vehicle behavior with the highest interaction reward is selected as the optimal vehicle behavior for the current level.
[0071] It should be noted that in the disclosed embodiments, different considerations are applied to oncoming obstacles and vehicles, and therefore the evaluation functions used for their reward evaluation are not identical. For oncoming obstacles, the evaluation function cannot assume compliance with traffic regulations. Therefore, the obstacle evaluation function for oncoming obstacles primarily focuses on safety and time utility. Furthermore, oncoming obstacles should also consider whether they are outside the road network. For vehicles, the vehicle evaluation function primarily focuses on safety, time utility, and compliance with traffic regulations (i.e., whether they are traveling against traffic).
[0072] The safety evaluation function can be used to evaluate whether a collision occurs between a target object (a vehicle or an oncoming obstacle) and other objects, or to evaluate the collision risk between a target object and other objects. For example, the safety evaluation function can satisfy the following formula:
[0073]
[0074] Where R s is the safety evaluation value. In the event of a collision between the target object and other objects, R s Take negative infinity; when there is no collision between the target object and other objects, R s Take the non-collision safety value R snc , R snc It can be calculated by the following formula:
[0075]
[0076]
[0077] Where d′ lon and d′ latThey are the reference longitudinal distance and the reference lateral distance, which can be obtained through the RSS (Responsibility Sensitive Safety) model; d lon and d lat They are the predicted longitudinal distance (the longitudinal distance between the vehicle and the oncoming obstacle) and the predicted lateral distance (the lateral distance between the vehicle and the oncoming obstacle) at the same future time point; a, b, c, and d are preset parameters.
[0078] Among them, time utility can refer to the expectation that the target object (vehicle or oncoming obstacle) can quickly end the interaction process and try not to deviate from the original state and behavior during the interaction process. Therefore, the time utility evaluation function can be used to evaluate the efficiency of the interaction and the degree of deviation from the original state and behavior during the interaction process. It is designed based on the deviation between the current speed and the target speed, and the deviation between the current angular velocity and the target angular velocity, as shown in the following formula:
[0079]
[0080]
[0081] Where R t is the time utility evaluation value, is the speed term, is the angular velocity term, v, v des are the current speed and target speed respectively, ω, ω des are the current angular velocity and target angular velocity respectively.
[0082] Among them, the reverse evaluation function can be used to evaluate the degree to which a vehicle occupies the reverse lane. The higher the degree of occupying the reverse lane, the lower the reverse evaluation value, as shown in the following formula:
[0083]
[0084] Where R l is the retrograde evaluation value, s occupy Indicates the area occupied by the vehicle in the opposite lane, s agent Indicates the area when the vehicle completely occupies the oncoming lane.
[0085] It should be noted that in the process of calculating the interactive feedback of the sampling behavior of the reverse obstacle or vehicle, a single-step simulation can be performed according to the sampling behavior to generate the predicted states of multiple future time points in a discrete time manner; then, the predicted state of each future time point is evaluated separately, and the interactive feedback of the sampling behavior is obtained by combining the evaluation results of all predicted states.
[0086] Taking the reverse obstacle as an example, in one example, each obstacle sample behavior is evaluated based on the current obstacle state, the obstacle evaluation function, and the optimal vehicle behavior at the previous level. The obstacle interaction reward for each obstacle sample behavior is obtained, including:
[0087] For each obstacle sampling behavior, the predicted state of the inverse obstacle at each future time point is determined based on the current obstacle state, resulting in a state sequence consisting of the current obstacle state and all predicted states. Each state in the state sequence is evaluated based on the obstacle evaluation function and the optimal vehicle behavior at the previous level, resulting in an obstacle interaction reward for each state in the state sequence. The obstacle interaction reward for the obstacle sampling behavior is determined based on the obstacle interaction reward for each state in the state sequence and the various time discount factors.
[0088] Specifically, the predicted state (including predicted speed and predicted position) of the reverse obstacle at each future time point can be derived through simulation based on the current obstacle state (including current speed and current position) and combined with the longitudinal acceleration and yaw angular velocity in the obstacle sampling behavior, to obtain a state sequence, such as {s0, s1, s2, ...s m}, m is the maximum simulation step size. For each state in the state sequence, assuming that the vehicle adopts the optimal vehicle behavior at the previous level, the obstacle evaluation function is used to calculate the corresponding obstacle interaction reward.
[0089] For example, for each state in the state sequence, the safety evaluation value, time utility evaluation value, and road network deviation evaluation value can be weighted averaged to obtain the corresponding obstacle interaction reward:
[0090]
[0091] Where w1, w2, and w3 are the weights corresponding to the safety evaluation function, time utility evaluation function, and road network deviation evaluation function, respectively; R1, R2, and R3 are the safety evaluation value, time utility evaluation value, and road network deviation evaluation value, respectively; R(s i ) represents the state s in the state sequence i Corresponding obstacle interaction feedback.
[0092] Furthermore, we can perform state backtracking and use the time discount factors corresponding to the current time point and each future time point to fuse the obstacle interaction returns of all states in the state sequence to obtain the obstacle interaction return of the obstacle sampling behavior. This is shown in the following formula:
[0093]
[0094] Where, γ iis the time discount factor corresponding to time point i, and R(s0,a) represents the obstacle interaction reward of obstacle sampling behavior a under the current obstacle state s0.
[0095] In the disclosed embodiment, considering that the uncertainty of the predicted state at future time points gradually increases over time, a time discount factor is introduced. When integrating the obstacle interaction reports from all time points, a larger time discount factor is applied to the obstacle interaction reports for states closer to the current time point, while a smaller time discount factor is applied to the obstacle interaction reports for states farther from the current time point. This further improves the reliability of the calculated interaction reports. For vehicles, the vehicle interaction reports for sampled vehicle behaviors can also be calculated in the same manner as described above. The difference lies in the vehicle evaluation function used, which is different from the obstacle evaluation function. This is not discussed here.
[0096] Furthermore, in step 114, the next level after the current level can be used as the new current level, such as Level-2, and the derivation process in steps 112 and 113 can be repeated. For example, since human drivers typically do not think beyond Level-2 in traffic scenarios, the highest decision-making level can be assumed to be Level-2. Based on the current obstacle state and the vehicle's optimal behavior at Level-1, the oncoming obstacle's behavior at Level-2 can be sampled and evaluated, resulting in the optimal obstacle behavior at Level-2.
[0097] Through the above steps 111 to 114, interactive reasoning based on the Level-k framework is implemented, which makes the consideration of the vehicle and the reverse obstacle more comprehensive and the reasoning of the reverse obstacle intention more accurate.
[0098] S120. Determine a reward matrix between the vehicle and the oncoming obstacle based on the optimal obstacle behavior at each decision level and the current vehicle state, wherein the reward matrix includes vehicle interaction rewards between each to-be-evaluated vehicle behavior and the optimal obstacle behavior at each decision level.
[0099] Specifically, after deriving the optimal obstacle behavior of the reverse obstacle at each decision level, for each behavior to be evaluated by the vehicle (which can be obtained through sampling, including longitudinal acceleration and yaw angular velocity), it can be assumed that the reverse obstacle adopts the optimal obstacle behavior at each decision level. Combined with the current vehicle state, the behavior to be evaluated is evaluated, and the vehicle interaction reward between the behavior to be evaluated and the optimal obstacle behavior at each decision level is obtained. Then, the vehicle interaction reward between all the behaviors to be evaluated and the optimal obstacle behavior at each decision level is constructed to form a reward matrix.
[0100] Among them, the vehicle evaluation function can be used to evaluate the behavior to be evaluated and obtain the vehicle interaction reward. The specific process can be found in the previous discussion.
[0101] It should be noted that the number of reward matrices is the same as the number of reverse obstacles. For each reverse obstacle, a reward matrix can be generated between the vehicle and it. For example, see Table 1, which shows a reward matrix for the highest decision level, Level-2.
[0102] Table 1 A return matrix
[0103]
[0104] Taking R(a,0) in Table 1 as an example, R(a,0) represents the vehicle interaction reward between the behavior a to be evaluated and the optimal obstacle behavior at Level-0.
[0105] S130 , based on the reward matrix between the vehicle and each oncoming obstacle, determining the vehicle decision behavior of the vehicle among all behaviors to be evaluated, and determining the obstacle decision behavior of each oncoming obstacle according to the vehicle decision behavior and each reward matrix.
[0106] In the embodiment of the present disclosure, considering that when multiple reverse obstacles appear, how to interact with multiple reverse obstacles at the same time is a very complex problem, based on the Level-k reasoning framework, the object currently at Level-k will think that other objects are at Level-(k-1).
[0107] Therefore, for interactions with multiple opposing obstacles, we can evaluate each obstacle individually. For each obstacle, a reward matrix is generated between the vehicle and the obstacle. In scenarios with multiple opposing obstacles, the goal of selecting the optimal behavior to be evaluated is to achieve overall optimization, rather than optimization for each individual obstacle. Therefore, after obtaining the reward matrix between the vehicle and each obstacle, the vehicle's decision-making behavior can be determined based on all the reward matrices.
[0108] In a specific embodiment, based on the reward matrix between the vehicle and each oncoming obstacle, determining the vehicle decision behavior of the vehicle among all behaviors to be evaluated includes the following steps:
[0109] Step 131: For each reverse obstacle, based on the vehicle interaction rewards at each decision level in the corresponding reward matrix and the level adoption probability corresponding to each decision level, determine the multi-level fusion reward between the vehicle and the reverse obstacle for each behavior to be evaluated;
[0110] Step 132: For each behavior to be evaluated, determine the cumulative reward corresponding to the behavior to be evaluated based on the multi-level fusion rewards between the vehicle and all oncoming obstacles under the behavior to be evaluated;
[0111] Step 133 : Determine the vehicle decision behavior of the vehicle among all behaviors to be evaluated based on the cumulative rewards corresponding to all behaviors to be evaluated.
[0112] In step 131, the level-adoption probability corresponding to the decision level can be the probability that the oncoming obstacle adopts that decision level. Specifically, the interaction feedback between the behavior to be evaluated and the vehicle at all decision levels can be fused based on the level-adoption probability corresponding to each decision level to obtain a multi-level fusion feedback between the vehicle and the oncoming obstacle for the behavior to be evaluated. Taking the highest decision level of Level-2 as an example, the following formula is shown:
[0113]
[0114] Where R p (a) is the multi-level fusion return between the vehicle and the opposite obstacle under the behavior to be evaluated a, P(k) is the probability of taking the level corresponding to the decision level k, is the optimal obstacle behavior under the behavior a to be evaluated and the decision level k in the reward matrix Vehicle interaction returns between them.
[0115] In the embodiment of the present disclosure, the probability of taking the level corresponding to each decision level can be obtained by analyzing the historical driving data of the reverse obstacle in combination with the current observation state.
[0116] Optionally, the method provided in the embodiment of the present disclosure further includes:
[0117] For each reverse obstacle, the current observation state of the reverse obstacle is obtained, where the current observation state includes the obstacle speed, obstacle acceleration, and obstacle trajectory deviation of the reverse obstacle; according to the speed level probability table, acceleration level probability table, deviation level probability table, and level transition probability table corresponding to the reverse obstacle, the level adoption probability corresponding to each decision level under the current observation state is determined.
[0118] The obstacle trajectory deviation can be the difference between the actual historical trajectory at m time points before the current time point t and the intended trajectory output at time point (tm). The intended trajectory can be the trajectory determined based on the obstacle decision behavior output at time point (tm).
[0119] In the embodiment of the present disclosure, for example, the real historical trajectory of m time points before the current time point t can be represented by Indicates that the intention trajectory output at time (tm) can be used The obstacle trajectory deviation can be and The distance difference between the two at equal times.
[0120] In the embodiment of the present disclosure, considering that the impact of trajectory deviation at a future time point is relatively small, the obstacle trajectory deviation can also be calculated in combination with a preset discount factor, as shown in the following formula:
[0121]
[0122] Where Dist represents the trajectory point in the obstacle intention trajectory recorded in the history The corresponding trajectory point in the real historical trajectory of the obstacle The straight-line distance between is the preset discount factor corresponding to the i-th time point.
[0123] Specifically, the corresponding probability value can be queried from the speed level probability table, the acceleration level probability table, and the deviation level probability table according to the current observation state of the reverse obstacle, and the corresponding probability value can be queried from the level transfer probability table according to the decision level of the reverse obstacle at the previous time point.
[0124] For example, taking Level-2 as the highest decision level, obstacle acceleration is represented by a0, obstacle speed by a1, and obstacle trajectory deviation by a2. Level-0, Level-1, and Level-2 are represented by b0, b1, and b2, respectively. Tables 2, 3, 4, and 5 show a speed level probability table, an acceleration level probability table, a deviation level probability table, and a level transition probability table, respectively.
[0125] Table 2 A speed level probability table
[0126] Speed Grade Probability Table <![CDATA[P(b0|a1)]]> <![CDATA[P(b1|a1)]]> <![CDATA[P(b2|a1)]]> <![CDATA[a1<10]]> 0.05 0.2 0.05 <![CDATA[10<=a1<20]]> 0.2 0.3 0.05 <![CDATA[20<=a1<30]]> 0.25 0.2 0.25 <![CDATA[30<=a1<40]]> 0.3 0.1 0.4 <![CDATA[40<=a1<50]]> 0.15 0.05 0.15 <![CDATA[a1>=50]]> 0.05 0.05 0.1
[0127] Table 3 Acceleration level probability table
[0128] Acceleration level probability table <![CDATA[P(b0|a0)]]> <![CDATA[P(b1|a0)]]> <![CDATA[P(b2|a0)]]> <![CDATA[a0<-1.0]]> 0.05 0.2 0.05 <![CDATA[-1.0<=a0<-0.6]]> 0.05 0.4 0.1 <![CDATA[-0.6<=a0<0.2]]> 0.6 0.25 0.5 <![CDATA[0.6<=a0<1.0]]> 0.2 0.1 0.25 <![CDATA[a0>1.0]]> 0.1 0.05 0.1
[0129] Table 4 A deviation level probability table
[0130] Deviation level probability table <![CDATA[P(b0|a2)]]> <![CDATA[P(b1|a2)]]> <![CDATA[P(b2|a2)]]> <![CDATA[a2<1.0]]> 0.4 0.4 0.4 <![CDATA[1.0<=a2<2.0]]> 0.2 0.3 0.3 <![CDATA[2.0<=a2<3.0]]> 0.1 0.1 0.1 <![CDATA[3.0<=a2<4.0]]> 0.1 0.1 0.1 <![CDATA[4.0<=a2<5.0]]> 0.05 0.05 0.05 <![CDATA[a2>=5.0]]> 0.05 0.05 0.05
[0131] Table 5 A level transfer probability table
[0132] Level transition probability table <![CDATA[b0]]> <![CDATA[b1]]> <![CDATA[b2]]> <![CDATA[b0]]> 0.7 0.3 0.1 <![CDATA[b1]]> 0.2 0.4 0.3 <![CDATA[b2]]> 0.1 0.3 0.6
[0133] In one example, before determining the level adoption probability corresponding to each decision level under the current observation state according to the speed level probability table, acceleration level probability table, deviation level probability table, and level conversion probability table corresponding to the reverse obstacle, the method further includes:
[0134] Obtain historical driving data of oncoming obstacles, determine the statistical probability of transition between decision levels based on the historical driving data, and obtain a level transition probability table; for each preset speed interval, determine the statistical probability corresponding to each decision level in the preset speed interval based on the historical driving data corresponding to the preset speed interval, and obtain a speed level probability table; for each preset acceleration interval, determine the statistical probability corresponding to each decision level in the preset acceleration interval based on the historical driving data corresponding to the preset acceleration interval, and obtain an acceleration level probability table; for each preset trajectory deviation interval, determine the statistical probability corresponding to each decision level in the preset trajectory deviation interval based on the historical driving data corresponding to the preset trajectory deviation interval, and obtain a deviation level probability table.
[0135] Specifically, the historical driving data of the reverse obstacle before the current time point can be obtained, wherein the historical driving data includes the obstacle speed, obstacle acceleration, obstacle trajectory deviation and decision level at multiple historical time points.
[0136] Furthermore, the historical driving data can be statistically analyzed to obtain the statistical probabilities of various transitions between decision levels at adjacent time points, that is, to obtain the statistical probabilities of transitions between decision levels, and to construct a level transition probability table, such as the statistical probabilities of transitions between the three decision levels shown in Table 5.
[0137] Furthermore, the historical driving data corresponding to each preset speed range can be obtained from the historical driving data by clustering or data screening. Then, the historical driving data corresponding to each preset speed range can be statistically analyzed to obtain the statistical probability corresponding to each decision level in each preset speed range, such as the statistical probability corresponding to each decision level in the six preset speed ranges shown in Table 2.
[0138] Furthermore, the historical driving data corresponding to each preset acceleration interval can be obtained from the historical driving data by clustering or data screening. Statistical analysis can then be performed on the historical driving data corresponding to each preset acceleration interval to obtain the statistical probability corresponding to each decision level under each preset acceleration interval, such as the statistical probability corresponding to each decision level under the five preset acceleration intervals shown in Table 3.
[0139] Furthermore, the historical driving data corresponding to each preset trajectory deviation interval can be obtained from the historical driving data by clustering or data screening. Statistical analysis can then be performed on the historical driving data corresponding to each preset trajectory deviation interval to obtain the statistical probability corresponding to each decision level under each preset trajectory deviation interval, as shown in Table 4.
[0140] After querying the corresponding probability values from each probability table according to the current observation state of the reverse obstacle, the Bayesian formula can be used to calculate the level-taking probability corresponding to each decision level under the current observation state based on the query probability values:
[0141]
[0142] Where, Indicates the current observation state Lower decision-making level The probability of taking the level, is the observed state at the previous time point.
[0143] It can be calculated by looking up the probability values from the speed level probability table, acceleration level probability table, deviation level probability table and level transition probability table.
[0144] In the above embodiment, the speed, acceleration, and trajectory deviation of the oncoming obstacle are selected as features to determine the level adoption probability of each decision level of the oncoming obstacle at the current time point, which can ensure the reliability of the level adoption probability of each decision level.
[0145] After obtaining the multi-level fusion reward between the vehicle and the oncoming obstacle for each behavior to be evaluated based on the level probability, in step 132, the multi-level fusion reward between the vehicle and all oncoming obstacles for each behavior to be evaluated can be added together to obtain the cumulative reward corresponding to the behavior to be evaluated. For example:
[0146]
[0147] Where R p (a) is the multi-level fusion return between the vehicle and the reverse obstacle under the behavior to be evaluated a, n is the number of reverse obstacles, R Δ (a) is the cumulative return corresponding to the behavior a to be evaluated.
[0148] Furthermore, in step 133 , the behavior to be evaluated with the largest cumulative reward may be selected as the vehicle decision-making behavior.
[0149] Figure 4FIG. 1 is a schematic diagram of a decision-making process in an embodiment of the present disclosure, such as Figure 4 As shown in the figure, taking decision levels 0 to 2 (i.e., Level-0 to Level-2) as an example, the cumulative reward of behavior a to be evaluated is the sum of the multi-level fusion reward Ra_1 with opposite obstacle 1, the multi-level fusion reward Ra_2 with opposite obstacle 2, and the multi-level fusion reward Ra_3 with opposite obstacle 3. Assuming that the cumulative reward of behavior c to be evaluated is the largest, behavior c to be evaluated is regarded as the vehicle decision behavior (including longitudinal acceleration and yaw angular velocity).
[0150] In the disclosed embodiment, the multi-level fusion return between the vehicle and the oncoming obstacle under the behavior to be evaluated is calculated by taking the probability of each decision level. The intention of the oncoming obstacle can be deduced in combination with the historical driving data of the oncoming obstacle, making the intention reasoning of the oncoming obstacle more reliable, thereby making the vehicle decision-making behavior more stable and safer.
[0151] After obtaining the vehicle decision behavior, for each reverse obstacle, the obstacle decision behavior (including longitudinal acceleration and yaw angular velocity) of the reverse obstacle can be determined based on the vehicle interaction reward between the vehicle decision behavior and the optimal behavior of the obstacle at each decision level in the reward matrix, as well as the level adoption probability corresponding to each decision level.
[0152] In a specific embodiment, determining the obstacle decision behavior of each reverse obstacle based on the vehicle decision behavior and each reward matrix includes:
[0153] For each reverse obstacle, the obstacle behavior level is determined among all decision levels based on the vehicle interaction rewards between the vehicle decision behavior in the corresponding reward matrix and the optimal obstacle behavior at each decision level, as well as the level adoption probability corresponding to each decision level. The optimal obstacle behavior at each obstacle behavior level is determined as the obstacle decision behavior for the reverse obstacle.
[0154] Specifically, each vehicle interaction reward under the vehicle decision behavior in the reward matrix can be multiplied by the probability of the corresponding decision level, that is, is the optimal obstacle behavior under the behavior a to be evaluated and the decision level k in the reward matrix The vehicle interaction returns between Figure 4 As shown in Figure 1, taking the reverse obstacle 1 as an example, the vehicle interaction rewards R(c,0), R(c,1), and R(c,2) under the vehicle decision behavior c are multiplied by the level adoption probabilities corresponding to the decision levels 0 to 2 of the reverse obstacle 1.
[0155] Furthermore, the maximum multiplication result is determined to obtain the obstacle behavior level, such as, Figure 4 The R(c,0) of the reverse obstacle 1 is the largest after multiplying the corresponding level adoption probability, and the corresponding decision level 0 is determined as the obstacle behavior level; the R(c,1) of the reverse obstacle 2 is the largest after multiplying the corresponding level adoption probability, and the corresponding decision level 1 is determined as the obstacle behavior level; the R(c,2) of the reverse obstacle 3 is the largest after multiplying the corresponding level adoption probability, and the corresponding decision level 2 is determined as the obstacle behavior level.
[0156] Furthermore, the optimal obstacle behavior under the obstacle behavior level is used as the obstacle decision behavior to realize the decision-making of each reverse obstacle.
[0157] The decision-making method for interacting with reverse obstacles provided in this embodiment determines the optimal obstacle behavior of the reverse obstacle at each decision level for each reverse obstacle corresponding to the vehicle based on the current vehicle state and the current obstacle state. Then, based on the optimal obstacle behavior at each decision level and the current vehicle state, the vehicle interaction reward between each to-be-evaluated behavior of the vehicle and the optimal obstacle behavior at each decision level is determined. This evaluates the vehicle's behavior separately by considering the optimal behavior of the reverse obstacle at different decision levels. Finally, based on the reward matrix between the vehicle and all reverse obstacles, the vehicle's decision behavior is determined from all to-be-evaluated behaviors, thereby determining the vehicle's decision behavior based on the vehicle's decision behavior and all reward matrices. The obstacle decision behavior of each oncoming obstacle is obtained, and the decision-making process when the vehicle interacts with the oncoming obstacle is realized. By analyzing the optimal behavior of the oncoming obstacle at different decision levels, the intention of the oncoming obstacle can be gradually deduced, making the intention of the oncoming obstacle clearer, thereby making the decision-making behavior of the vehicle more stable and safer. The intention of the oncoming obstacle is variable and more flexible. Compared with the method of only changing the longitudinal or lateral intention, it can be closer to the intention of a real human driver. The method provided in the embodiment of the present disclosure can be applied to the synchronous processing of multiple oncoming obstacles, meeting real-time requirements, and ensuring the output of reasonable decision-making behavior when the vehicle interacts with multiple oncoming obstacles.
[0158] Figure 5 FIG. 1 is a schematic diagram of a structure of a decision-making device for interacting with a reverse obstacle in an embodiment of the present disclosure. Figure 5 As shown, the device includes: a level behavior determination module 510, a reward matrix determination module 520 and a decision module 530.
[0159] a level behavior determination module 510 for determining, for each oncoming obstacle corresponding to the vehicle, an optimal obstacle behavior for the oncoming obstacle at each decision level based on the current vehicle state of the vehicle and the current obstacle state of the oncoming obstacle;
[0160] a reward matrix determination module 520 for determining a reward matrix between the vehicle and the oncoming obstacle based on the optimal obstacle behavior at each decision level and the current vehicle state, wherein the reward matrix includes vehicle interaction rewards between each to-be-evaluated vehicle behavior and the optimal obstacle behavior at each decision level;
[0161] The decision module 530 is used to determine the vehicle decision behavior of the vehicle among all behaviors to be evaluated based on the reward matrix between the vehicle and each oncoming obstacle, and to determine the obstacle decision behavior of each oncoming obstacle based on the vehicle decision behavior and each reward matrix.
[0162] Optionally, the grade behavior determination module 510 is specifically configured to:
[0163] Obtain the optimal obstacle behavior of the reverse obstacle at the lowest decision level, and the optimal vehicle behavior of the vehicle at the lowest decision level; take the level after the lowest decision level as the current level, and based on the current obstacle state and the optimal vehicle behavior of the vehicle at the level before the current level, sample and report back the behavior of the reverse obstacle at the current level to obtain the optimal obstacle behavior at the current level; based on the current vehicle state and the optimal obstacle behavior of the reverse obstacle at the level before the current level, sample and report back the behavior of the vehicle at the current level to obtain the optimal vehicle behavior at the current level; take the level after the current level as the new current level, and return to the step of sampling and reporting back the behavior of the reverse obstacle at the current level until the current level reaches the highest decision level.
[0164] Optionally, the level behavior determination module 510 is further configured to sample the behavior of the oncoming obstacle at the current level to obtain a plurality of obstacle sampling behaviors; evaluate each obstacle sampling behavior based on the current obstacle state, the obstacle evaluation function, and the optimal vehicle behavior at the previous level to obtain an obstacle interaction return for each obstacle sampling behavior; and determine the optimal obstacle behavior at the current level among all obstacle sampling behaviors based on the obstacle interaction returns; wherein the obstacle evaluation function includes a safety evaluation function, a time utility evaluation function, and a road network deviation evaluation function.
[0165] Optionally, the level behavior determination module 510 is further used to sample the behavior of the vehicle at the current level to obtain multiple vehicle sampling behaviors; evaluate each vehicle sampling behavior according to the current vehicle state, the vehicle evaluation function and the optimal obstacle behavior at the previous level to obtain the vehicle interaction return of each vehicle sampling behavior; based on each vehicle interaction return, determine the optimal vehicle behavior at the current level among all vehicle sampling behaviors; wherein the vehicle evaluation function includes a safety evaluation function, a time utility evaluation function and a retrograde evaluation function.
[0166] Optionally, the level behavior determination module 510 is further configured to determine, for each obstacle sampling behavior, a predicted state of the reverse obstacle at each future time point based on the current obstacle state, thereby obtaining a state sequence consisting of the current obstacle state and all predicted states; evaluate each state in the state sequence according to the obstacle evaluation function and the optimal vehicle behavior at the previous level, thereby obtaining an obstacle interaction reward for each state in the state sequence; and determine an obstacle interaction reward for the obstacle sampling behavior based on the obstacle interaction reward for each state in the state sequence and various time discount factors.
[0167] Optionally, the device further includes a confidence determination module, the confidence determination module being configured to obtain a predicted trajectory of the oncoming obstacle, wherein the predicted trajectory includes trajectory points corresponding to a plurality of moments; determine an uncertainty of each trajectory point on the predicted trajectory, and determine a confidence region for each trajectory point based on the uncertainty of each trajectory point;
[0168] Correspondingly, the level behavior determination module 510 is further configured to sample the behavior of the reverse obstacle within the corresponding confidence region to obtain a plurality of obstacle sample behaviors.
[0169] Optionally, the decision module 530 is specifically configured to:
[0170] For each oncoming obstacle, based on the vehicle interaction rewards at each decision level in the corresponding reward matrix and the level adoption probability corresponding to each decision level, the multi-level fusion reward between the vehicle and the oncoming obstacle under each behavior to be evaluated is determined; for each behavior to be evaluated, based on the multi-level fusion rewards between the vehicle and all oncoming obstacles under the behavior to be evaluated, the cumulative reward corresponding to the behavior to be evaluated is determined; based on the cumulative rewards corresponding to all behaviors to be evaluated, the vehicle decision behavior of the vehicle is determined among all behaviors to be evaluated.
[0171] Optionally, the decision module 530 is further used to obtain the current observation state of each reverse obstacle, wherein the current observation state includes the obstacle speed, obstacle acceleration and obstacle trajectory deviation of the reverse obstacle; and determine the level adoption probability corresponding to each decision level under the current observation state according to the speed level probability table, acceleration level probability table, deviation level probability table and level transfer probability table corresponding to the reverse obstacle.
[0172] Optionally, the device also includes a level probability update module, which is used to obtain the historical driving data of the reverse obstacle, determine the transfer statistical probability between each decision level based on the historical driving data, and obtain a level transfer probability table; for each preset speed interval, based on the historical driving data corresponding to the preset speed interval, determine the statistical probability corresponding to each decision level in the preset speed interval, and obtain a speed level probability table; for each preset acceleration interval, based on the historical driving data corresponding to the preset acceleration interval, determine the statistical probability corresponding to each decision level in the preset acceleration interval, and obtain an acceleration level probability table; for each preset trajectory deviation interval, based on the historical driving data corresponding to the preset trajectory deviation interval, determine the statistical probability corresponding to each decision level in the preset trajectory deviation interval, and obtain a deviation level probability table.
[0173] Optionally, the decision module 530 is further used to determine, for each reverse obstacle, an obstacle behavior level among all decision levels based on the vehicle interaction reward between the vehicle decision behavior and the optimal obstacle behavior at each decision level in the corresponding reward matrix, and the level adoption probability corresponding to each decision level; and determine the optimal obstacle behavior at the obstacle behavior level as the obstacle decision behavior of the reverse obstacle.
[0174] The decision-making device for interacting with a reverse obstacle provided in the embodiment of the present disclosure can execute the steps of the decision-making method for interacting with a reverse obstacle provided in the method embodiment of the present disclosure, and the execution steps and beneficial effects are not repeated here.
[0175] Figure 6 This is a schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 6 , which shows a structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0176] like Figure 6As shown, the electronic device 500 may include a processing device 501, a ROM 502, a RAM 503, a bus 504, an input / output (I / O) interface 505, an input device 506, an output device 507, a storage device 508, and a communication device 509. The processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501 can perform various appropriate actions and processes to implement the method of the embodiment as described in the present disclosure according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage device 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via the bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0177] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart, thereby implementing the decision-making method for interacting with the reverse obstacle as described above. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0178] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0179] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method provided in any embodiment of the present disclosure.
[0180] Optionally, when the above one or more programs are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.
[0181] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0182] Solution 1: A decision-making method for interacting with an inverse obstacle, the method comprising:
[0183] For each oncoming obstacle corresponding to the vehicle, determining an optimal obstacle behavior for the oncoming obstacle at each decision level based on a current vehicle state of the vehicle and a current obstacle state of the oncoming obstacle;
[0184] Determining a reward matrix between the vehicle and the oncoming obstacle based on the optimal obstacle behavior at each decision level and the current vehicle state, wherein the reward matrix includes vehicle interaction rewards between each to-be-evaluated behavior of the vehicle and the optimal obstacle behavior at each decision level;
[0185] Based on a reward matrix between the vehicle and each oncoming obstacle, a vehicle decision behavior of the vehicle is determined among all behaviors to be evaluated, and an obstacle decision behavior of each oncoming obstacle is determined according to the vehicle decision behavior and each reward matrix.
[0186] Solution 2: According to the method of Solution 1, determining the optimal obstacle behavior of the oncoming obstacle at each decision level based on the current vehicle state of the vehicle and the current obstacle state of the oncoming obstacle includes:
[0187] Obtaining an optimal obstacle behavior of the oncoming obstacle at the lowest decision level, and an optimal vehicle behavior of the vehicle at the lowest decision level;
[0188] Taking the level after the lowest decision level as the current level, sampling and reporting the behavior of the oncoming obstacle at the current level based on the current obstacle state and the optimal vehicle behavior of the vehicle at the level before the current level to obtain the optimal obstacle behavior at the current level;
[0189] Based on the current vehicle state and the optimal obstacle behavior of the oncoming obstacle at the previous level of the current level, sampling and reporting the behavior of the vehicle at the current level to obtain the optimal vehicle behavior at the current level;
[0190] The next level after the current level is used as the new current level, and the steps of sampling and reporting the behavior of the reverse obstacle at the current level are returned to be executed until the current level reaches the highest decision level.
[0191] Solution 3: The method according to Solution 2, wherein based on the current obstacle state and the optimal vehicle behavior of the vehicle at the level before the current level, sampling and reporting evaluation of the behavior of the oncoming obstacle at the current level to obtain the optimal obstacle behavior at the current level, includes:
[0192] Sampling the behavior of the reverse obstacle at the current level to obtain a plurality of obstacle sampling behaviors;
[0193] Evaluate each obstacle sampling behavior based on the current obstacle state, the obstacle evaluation function, and the optimal vehicle behavior at the previous level to obtain an obstacle interaction reward for each obstacle sampling behavior;
[0194] Based on the interaction feedback of each obstacle, determine the optimal obstacle behavior at the current level among all obstacle sampled behaviors;
[0195] The obstacle evaluation function includes a safety evaluation function, a time utility evaluation function and a road network deviation evaluation function.
[0196] Solution 4: The method according to Solution 2, wherein based on the current vehicle state and the optimal obstacle behavior of the oncoming obstacle at the level before the current level, sampling and reporting the vehicle behavior at the current level to obtain the optimal vehicle behavior at the current level, includes:
[0197] Sampling the behavior of the vehicle at the current level to obtain a plurality of vehicle sample behaviors;
[0198] Evaluate each vehicle sample behavior based on the current vehicle state, the vehicle evaluation function, and the optimal obstacle behavior at the previous level to obtain a vehicle interaction reward for each vehicle sample behavior;
[0199] Determine the optimal vehicle behavior at the current level among all sampled vehicle behaviors based on the interaction feedback of each vehicle;
[0200] The vehicle evaluation function includes a safety evaluation function, a time utility evaluation function and a retrograde evaluation function.
[0201] Solution 5: According to the method of Solution 3, evaluating each obstacle sampling behavior based on the current obstacle state, the obstacle evaluation function, and the optimal vehicle behavior at the previous level to obtain an obstacle interaction reward for each obstacle sampling behavior includes:
[0202] For each obstacle sampling behavior, determining a predicted state of the reverse obstacle at each future time point based on the current obstacle state, and obtaining a state sequence consisting of the current obstacle state and all predicted states;
[0203] Evaluate each state in the state sequence according to the obstacle evaluation function and the optimal vehicle behavior at the previous level to obtain an obstacle interaction reward for each state in the state sequence;
[0204] An obstacle interaction reward of the obstacle sampling behavior is determined based on the obstacle interaction reward of each state in the state sequence and each time discount factor.
[0205] Solution 6: The method according to Solution 3, further comprising:
[0206] Obtaining a predicted trajectory of the reverse obstacle, wherein the predicted trajectory includes trajectory points corresponding to multiple moments;
[0207] Determining the uncertainty of each trajectory point on the predicted trajectory, and determining a confidence region for each trajectory point based on the uncertainty of each trajectory point;
[0208] Accordingly, the behavior of the reverse obstacle at the current level is sampled to obtain multiple obstacle sampling behaviors, including:
[0209] The behavior of the reverse obstacle is sampled within the corresponding confidence region to obtain a plurality of obstacle sampling behaviors.
[0210] Solution 7: According to the method of Solution 1, determining the vehicle decision behavior of the vehicle among all behaviors to be evaluated based on the reward matrix between the vehicle and each oncoming obstacle includes:
[0211] For each reverse obstacle, based on the vehicle interaction rewards at each decision level in the corresponding reward matrix and the level adoption probability corresponding to each decision level, determine the multi-level fusion reward between the vehicle and the reverse obstacle for each behavior to be evaluated;
[0212] For each behavior to be evaluated, determining a cumulative reward corresponding to the behavior to be evaluated based on a multi-level fusion reward between the vehicle and all oncoming obstacles under the behavior to be evaluated;
[0213] According to the accumulated rewards corresponding to all the behaviors to be evaluated, the vehicle decision behavior of the vehicle is determined among all the behaviors to be evaluated.
[0214] Solution 8: The method according to Solution 7, further comprising:
[0215] For each reverse obstacle, obtaining a current observation state of the reverse obstacle, wherein the current observation state includes an obstacle speed, an obstacle acceleration, and an obstacle trajectory deviation of the reverse obstacle;
[0216] According to the speed level probability table, acceleration level probability table, deviation level probability table and level transition probability table corresponding to the reverse obstacle, the level adoption probability corresponding to each decision level under the current observation state is determined.
[0217] Solution 9. The method according to Solution 8, before determining the level adoption probability corresponding to each decision level under the current observation state based on the speed level probability table, acceleration level probability table, deviation level probability table, and level conversion probability table corresponding to the oncoming obstacle, further comprising:
[0218] Acquiring historical driving data of the oncoming obstacle, determining the statistical probability of transition between decision levels based on the historical driving data, and obtaining a level transition probability table;
[0219] For each preset speed interval, based on the historical driving data corresponding to the preset speed interval, determining the statistical probability corresponding to each decision level in the preset speed interval, and obtaining a speed level probability table;
[0220] For each preset acceleration interval, based on historical driving data corresponding to the preset acceleration interval, determining the statistical probability corresponding to each decision level in the preset acceleration interval, and obtaining an acceleration level probability table;
[0221] For each preset trajectory deviation interval, based on the historical driving data corresponding to the preset trajectory deviation interval, the statistical probability corresponding to each decision level in the preset trajectory deviation interval is determined to obtain a deviation level probability table.
[0222] Solution 10: According to the method of Solution 7, determining the obstacle decision behavior of each reverse obstacle based on the vehicle decision behavior and each reward matrix includes:
[0223] For each reverse obstacle, the obstacle behavior level is determined across all decision levels based on the vehicle interaction rewards between the vehicle decision behavior described in the corresponding reward matrix and the optimal obstacle behavior at each decision level, as well as the level adoption probability corresponding to each decision level.
[0224] The optimal obstacle behavior under the obstacle behavior level is determined as the obstacle decision behavior of the reverse obstacle.
[0225] Solution 11: A decision-making device for interacting with a reverse obstacle, comprising:
[0226] a level behavior determination module, configured to determine, for each oncoming obstacle corresponding to the vehicle, an optimal obstacle behavior for the oncoming obstacle at each decision level based on a current vehicle state of the vehicle and a current obstacle state of the oncoming obstacle;
[0227] a reward matrix determination module, configured to determine a reward matrix between the vehicle and the oncoming obstacle based on the optimal obstacle behavior at each decision level and the current vehicle state, wherein the reward matrix includes vehicle interaction rewards between each to-be-evaluated vehicle behavior and the optimal obstacle behavior at each decision level;
[0228] A decision module is configured to determine a vehicle decision behavior of the vehicle among all behaviors to be evaluated based on a reward matrix between the vehicle and each oncoming obstacle, and to determine an obstacle decision behavior of each oncoming obstacle based on the vehicle decision behavior and each reward matrix.
[0229] Solution 12. An electronic device, comprising:
[0230] one or more processors;
[0231] a storage device for storing one or more programs;
[0232] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of solutions 1-10.
[0233] Solution 13: A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of Solutions 1-10.
[0234] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
Claims
1. A decision-making method for interacting with reverse obstacles, characterized in that: The method comprises: For each oncoming obstacle corresponding to the vehicle, determining an optimal obstacle behavior for the oncoming obstacle at each decision level based on a current vehicle state of the vehicle and a current obstacle state of the oncoming obstacle; Determining a reward matrix between the vehicle and the oncoming obstacle based on the optimal obstacle behavior at each decision level and the current vehicle state, wherein the reward matrix includes vehicle interaction rewards between each to-be-evaluated behavior of the vehicle and the optimal obstacle behavior at each decision level; Determining a vehicle decision behavior of the vehicle among all behaviors to be evaluated based on a reward matrix between the vehicle and each oncoming obstacle, and determining an obstacle decision behavior for each oncoming obstacle based on the vehicle decision behavior and each reward matrix; The determining, based on the current vehicle state of the vehicle and the current obstacle state of the oncoming obstacle, an optimal obstacle behavior of the oncoming obstacle at each decision level includes: Obtaining an optimal obstacle behavior of the oncoming obstacle at the lowest decision level, and an optimal vehicle behavior of the vehicle at the lowest decision level; Taking the level after the lowest decision level as the current level, sampling and reporting the behavior of the oncoming obstacle at the current level based on the current obstacle state and the optimal vehicle behavior of the vehicle at the level before the current level to obtain the optimal obstacle behavior at the current level; Based on the current vehicle state and the optimal obstacle behavior of the oncoming obstacle at the previous level of the current level, sampling and reporting the behavior of the vehicle at the current level to obtain the optimal vehicle behavior at the current level; The next level after the current level is used as the new current level, and the steps of sampling and reporting the behavior of the reverse obstacle at the current level are returned to be executed until the current level reaches the highest decision level.
2. The method according to claim 1, characterized in that The sampling and reporting evaluation of the behavior of the reverse obstacle at the current level based on the current obstacle state and the optimal vehicle behavior of the vehicle at the level before the current level to obtain the optimal obstacle behavior at the current level includes: Sampling the behavior of the reverse obstacle at the current level to obtain a plurality of obstacle sampling behaviors; Evaluate each obstacle sampling behavior based on the current obstacle state, the obstacle evaluation function, and the optimal vehicle behavior at the previous level to obtain an obstacle interaction reward for each obstacle sampling behavior; Based on the interaction feedback of each obstacle, determine the optimal obstacle behavior at the current level among all obstacle sampled behaviors; The obstacle evaluation function includes a safety evaluation function, a time utility evaluation function and a road network deviation evaluation function.
3. The method according to claim 1, characterized in that The sampling and reporting evaluation of the vehicle behavior at the current level based on the current vehicle state and the optimal obstacle behavior of the oncoming obstacle at the previous level of the current level to obtain the optimal vehicle behavior at the current level includes: Sampling the behavior of the vehicle at the current level to obtain a plurality of vehicle sample behaviors; Evaluate each vehicle sample behavior based on the current vehicle state, the vehicle evaluation function, and the optimal obstacle behavior at the previous level to obtain a vehicle interaction reward for each vehicle sample behavior; Determine the optimal vehicle behavior at the current level among all sampled vehicle behaviors based on the interaction feedback of each vehicle; The vehicle evaluation function includes a safety evaluation function, a time utility evaluation function and a retrograde evaluation function.
4. The method according to claim 2, characterized in that The step of evaluating each obstacle sampling behavior based on the current obstacle state, the obstacle evaluation function, and the optimal vehicle behavior at the previous level to obtain an obstacle interaction reward for each obstacle sampling behavior includes: For each obstacle sampling behavior, determining a predicted state of the reverse obstacle at each future time point based on the current obstacle state, and obtaining a state sequence consisting of the current obstacle state and all predicted states; Evaluate each state in the state sequence according to the obstacle evaluation function and the optimal vehicle behavior at the previous level to obtain an obstacle interaction reward for each state in the state sequence; An obstacle interaction reward of the obstacle sampling behavior is determined based on the obstacle interaction reward of each state in the state sequence and each time discount factor.
5. The method according to claim 2, characterized in that The method further comprises: Obtaining a predicted trajectory of the reverse obstacle, wherein the predicted trajectory includes trajectory points corresponding to multiple moments; determining an uncertainty for each trajectory point on the predicted trajectory, and determining a confidence region for each trajectory point based on the uncertainty for each trajectory point; Accordingly, the behavior of the reverse obstacle at the current level is sampled to obtain multiple obstacle sampling behaviors, including: The behavior of the reverse obstacle is sampled within the corresponding confidence region to obtain a plurality of obstacle sampling behaviors.
6. The method according to claim 1, characterized in that The step of determining a vehicle decision behavior of the vehicle among all behaviors to be evaluated based on a reward matrix between the vehicle and each of the oncoming obstacles includes: For each reverse obstacle, based on the vehicle interaction rewards at each decision level in the corresponding reward matrix and the level adoption probability corresponding to each decision level, determine the multi-level fusion reward between the vehicle and the reverse obstacle for each behavior to be evaluated; For each behavior to be evaluated, determining a cumulative reward corresponding to the behavior to be evaluated based on a multi-level fusion reward between the vehicle and all oncoming obstacles under the behavior to be evaluated; According to the accumulated rewards corresponding to all the behaviors to be evaluated, the vehicle decision behavior of the vehicle is determined among all the behaviors to be evaluated.
7. The method according to claim 6, characterized in that The method further comprises: For each reverse obstacle, obtaining a current observation state of the reverse obstacle, wherein the current observation state includes an obstacle speed, an obstacle acceleration, and an obstacle trajectory deviation of the reverse obstacle; According to the speed level probability table, acceleration level probability table, deviation level probability table and level transition probability table corresponding to the reverse obstacle, the level adoption probability corresponding to each decision level under the current observation state is determined.
8. The method according to claim 7, characterized in that Before determining the level adoption probability corresponding to each decision level under the current observation state according to the speed level probability table, acceleration level probability table, deviation level probability table, and level conversion probability table corresponding to the oncoming obstacle, the method further includes: Acquiring historical driving data of the oncoming obstacle, determining the statistical probability of transition between decision levels based on the historical driving data, and obtaining a level transition probability table; For each preset speed interval, based on the historical driving data corresponding to the preset speed interval, determining the statistical probability corresponding to each decision level in the preset speed interval, and obtaining a speed level probability table; For each preset acceleration interval, based on historical driving data corresponding to the preset acceleration interval, determining the statistical probability corresponding to each decision level in the preset acceleration interval, and obtaining an acceleration level probability table; For each preset trajectory deviation interval, based on the historical driving data corresponding to the preset trajectory deviation interval, the statistical probability corresponding to each decision level in the preset trajectory deviation interval is determined to obtain a deviation level probability table.
9. The method according to claim 6, characterized in that Determining the obstacle decision behavior of each reverse obstacle based on the vehicle decision behavior and each reward matrix includes: For each reverse obstacle, the obstacle behavior level is determined across all decision levels based on the vehicle interaction rewards between the vehicle decision behavior described in the corresponding reward matrix and the optimal obstacle behavior at each decision level, as well as the level adoption probability corresponding to each decision level. The optimal obstacle behavior under the obstacle behavior level is determined as the obstacle decision behavior of the reverse obstacle.
10. A decision-making device for interacting with reverse obstacles, characterized in that: The device comprises: The level behavior determination module is used to perform, for each oncoming obstacle corresponding to the vehicle: Obtaining an optimal obstacle behavior of the oncoming obstacle at the lowest decision level, and an optimal vehicle behavior of the vehicle at the lowest decision level; Taking the level after the lowest decision level as the current level, sampling and reporting the behavior of the oncoming obstacle at the current level based on the current obstacle state and the optimal vehicle behavior of the vehicle at the level before the current level, and obtaining the optimal obstacle behavior at the current level; Based on the current vehicle state and the optimal obstacle behavior of the oncoming obstacle at the previous level of the current level, sampling and reporting the behavior of the vehicle at the current level are performed to obtain the optimal vehicle behavior at the current level; Taking the next level after the current level as the new current level, returning to the steps of sampling and reporting the behavior of the reverse obstacle at the current level until the current level reaches the highest decision level; a reward matrix determination module, configured to determine a reward matrix between the vehicle and the oncoming obstacle based on the optimal obstacle behavior at each decision level and the current vehicle state, wherein the reward matrix includes vehicle interaction rewards between each to-be-evaluated vehicle behavior and the optimal obstacle behavior at each decision level; A decision module is configured to determine a vehicle decision behavior of the vehicle among all behaviors to be evaluated based on a reward matrix between the vehicle and each oncoming obstacle, and to determine an obstacle decision behavior of each oncoming obstacle based on the vehicle decision behavior and each reward matrix.
11. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Vehicle anti-collision control method, system and equipment and storage medium
CN115891990A
Automatic avoidance decision-making method and system, electronic equipment, vehicle and storage medium
CN116080639A