Automatic driving decision method, device and storage medium

CN117452934BActive Publication Date: 2026-10-09UISEE TECH BEIJING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311385296.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-24
Publication Date
2026-10-09
Estimated Expiration
2043-10-24

AI Technical Summary

Technical Problem

[0005]为了解决上述技术问题或者至少部分地解决上述技术问题,本公开实施例提供了一种自动驾驶决策方法、装置和存储介质,解决现有技术中仅考虑单个障碍物的最佳决策导致车辆决策准确性低的问题

Benefits of technology

[0016] This disclosure provides an autonomous driving decision-making method that acquires a sequence of obstacles corresponding to a vehicle. Starting from the first obstacle and working backwards, based on the currently derived interaction benefit matrix between the obstacle and the vehicle, it determines the interaction benefit matrix between the next obstacle and the vehicle, until the final interaction benefit matrix between the obstacle and the vehicle is obtained. Then, starting from the final obstacle in the obstacle sequence and working backwards, it updates the interaction benefit matrix between the previous obstacle and the vehicle based on the benefits of each game strategy in the currently updated interaction benefit matrix between the obstacle and the vehicle, until the interaction benefit matrix between the first obstacle and the vehicle is updated. Finally, based on the benefits of each... The interaction payoff matrix between obstacles and vehicles is used to obtain the globally optimal game strategy for each obstacle, thus completing the vehicle's decision-making. This method considers multiple obstacles corresponding to the vehicle and derives the interaction payoff matrix of each obstacle from the first obstacle backwards. It takes into account the payoff of the vehicle's interaction with each obstacle, solving the problem of low vehicle decision-making accuracy caused by considering only the optimal decision of a single obstacle in the prior art. Furthermore, updating the interaction payoff matrix backwards from the final obstacle can ensure that the payoff of the vehicle and each obstacle is globally optimal without limiting the flexibility of the vehicle's decision-making strategy, avoiding the problem of conflicting decision results when making decisions on multiple obstacles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117452934B_ABST
    Figure CN117452934B_ABST
Patent Text Reader

Abstract

The embodiment of the present disclosure discloses an automatic driving decision method and device and a storage medium. The method obtains an obstacle sequence corresponding to a vehicle, determines each interaction benefit matrix of a subsequent obstacle based on an interaction benefit matrix of a current deduced obstacle from the first obstacle, until each interaction benefit matrix of a final obstacle is obtained, then updates the interaction benefit matrix of a previous obstacle based on the benefits of each game strategy in the interaction benefit matrix of the current updated obstacle from the final obstacle, until the interaction benefit matrix of the first obstacle is updated, finally, obtains a global optimal game strategy corresponding to each obstacle according to the interaction benefit matrix of each obstacle, completes vehicle decision, and solves the problem of low vehicle decision accuracy caused by only considering the optimal decision of a single obstacle, which can ensure the global optimality of the benefits of the vehicle and each obstacle without limiting the flexibility of the vehicle decision strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of autonomous driving technology, and in particular to an autonomous driving decision-making method, apparatus, and storage medium. Background Technology

[0002] During the movement of an autonomous vehicle, the planned path of the vehicle will inevitably intersect with the movement paths of surrounding traffic participants (which can be understood as a collision). When the planned path of the vehicle is fixed and cannot be avoided laterally, the vehicle needs to make behavioral decisions on whether to cut in front of other traffic participants or to give way to other traffic participants who are likely to collide with it.

[0003] Reasonable decision-making can improve a vehicle's maneuverability and ride comfort, while unreasonable conservative decisions may lead to constant braking and eventual entrapment in traffic. Unreasonable aggressive decisions can result in collisions with other road users. Existing decision-making solutions primarily address the interaction between the vehicle and individual obstacles. These solutions acquire current state information about both the vehicle and obstacles, simulate their interaction, and use Nash equilibrium to determine the equilibrium game strategy between the two parties.

[0004] However, autonomous vehicle behavior is actually affected by the combined effects of multiple obstacles. The best decision considering only a single obstacle may become a suboptimal or even unacceptable decision in a multi-obstacle decision environment, resulting in poor applicability. Summary of the Invention

[0005] To address the aforementioned technical problems, or at least partially address them, this disclosure provides an autonomous driving decision-making method, apparatus, and storage medium, which solves the problem of low vehicle decision-making accuracy caused by considering only a single obstacle in the prior art.

[0006] In a first aspect, embodiments of this disclosure provide an autonomous driving decision-making method, the method comprising:

[0007] Obtain the obstacle sequence corresponding to the vehicle. Starting from the first obstacle in the obstacle sequence and proceeding backward, based on the currently derived interaction benefit matrix between the obstacle and the vehicle, determine the interaction benefit matrix between the next obstacle and the vehicle, until the final obstacle in the obstacle sequence and the vehicle are obtained.

[0008] Starting from the last obstacle in the obstacle sequence and moving forward, based on the payoff of each game strategy in the interaction payoff matrix between the currently updated obstacle and the vehicle, the local optimal game strategy corresponding to the currently updated obstacle is determined, and the interaction payoff matrix between the previous obstacle and the vehicle is updated based on the payoff of the local optimal game strategy, until the interaction payoff matrix between the first obstacle and the vehicle is updated. The game strategy includes a first strategy describing the longitudinal acceleration of the vehicle, a second strategy describing the longitudinal acceleration of the obstacle, and the vehicle's yielding behavior.

[0009] Based on the interaction benefit matrix between each obstacle and the vehicle, the global optimal game strategy corresponding to each obstacle is determined, and the target decision result of the vehicle is obtained.

[0010] Secondly, embodiments of this disclosure also provide an autonomous driving decision-making device, the device comprising:

[0011] The revenue derivation module is used to obtain the obstacle sequence corresponding to the vehicle. Starting from the first obstacle in the obstacle sequence and proceeding backward, based on the currently derived interaction revenue matrix between the obstacle and the vehicle, the interaction revenue matrix between the next obstacle and the vehicle is determined, until the interaction revenue matrix between the final obstacle and the vehicle in the obstacle sequence is obtained.

[0012] The payout update module is used to start from the last obstacle in the obstacle sequence and move forward, based on the payout of each game strategy in the interaction payout matrix between the currently updated obstacle and the vehicle, determine the local optimal game strategy corresponding to the currently updated obstacle, and update the interaction payout matrix between the previous obstacle and the vehicle based on the payout of the local optimal game strategy, until the interaction payout matrix between the first obstacle and the vehicle is updated. The game strategy includes a first strategy describing the longitudinal acceleration of the vehicle, a second strategy describing the longitudinal acceleration of the obstacle, and the vehicle's yielding behavior.

[0013] The decision-making module is used to determine the globally optimal game strategy for each obstacle based on the interaction benefit matrix between each obstacle and the vehicle, and to obtain the target decision result of the vehicle.

[0014] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising: one or more processors; a storage device for storing one or more programs; and when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the autonomous driving decision-making method as described above.

[0015] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the autonomous driving decision-making method as described above.

[0016] This disclosure provides an autonomous driving decision-making method that acquires a sequence of obstacles corresponding to a vehicle. Starting from the first obstacle and working backwards, based on the currently derived interaction benefit matrix between the obstacle and the vehicle, it determines the interaction benefit matrix between the next obstacle and the vehicle, until the final interaction benefit matrix between the obstacle and the vehicle is obtained. Then, starting from the final obstacle in the obstacle sequence and working backwards, it updates the interaction benefit matrix between the previous obstacle and the vehicle based on the benefits of each game strategy in the currently updated interaction benefit matrix between the obstacle and the vehicle, until the interaction benefit matrix between the first obstacle and the vehicle is updated. Finally, based on the benefits of each... The interaction payoff matrix between obstacles and vehicles is used to obtain the globally optimal game strategy for each obstacle, thus completing the vehicle's decision-making. This method considers multiple obstacles corresponding to the vehicle and derives the interaction payoff matrix of each obstacle from the first obstacle backwards. It takes into account the payoff of the vehicle's interaction with each obstacle, solving the problem of low vehicle decision-making accuracy caused by considering only the optimal decision of a single obstacle in the prior art. Furthermore, updating the interaction payoff matrix backwards from the final obstacle can ensure that the payoff of the vehicle and each obstacle is globally optimal without limiting the flexibility of the vehicle's decision-making strategy, avoiding the problem of conflicting decision results when making decisions on multiple obstacles. Attached Figure Description

[0017] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0018] Figure 1 This is a flowchart of an autonomous driving decision-making method according to an embodiment of this disclosure;

[0019] Figure 2 This is a schematic diagram illustrating the derivation of benefits in one embodiment of this disclosure;

[0020] Figure 3 This is a schematic diagram of a pruning strategy in one embodiment of this disclosure;

[0021] Figure 4 This is a schematic diagram illustrating a revenue update in one embodiment of this disclosure;

[0022] Figure 5 This is a schematic diagram illustrating the search for a globally optimal game strategy in an embodiment of this disclosure;

[0023] Figure 6 This is a schematic diagram of the structure of an autonomous driving decision-making device according to an embodiment of the present disclosure;

[0024] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0027] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0028] Before providing a detailed description of the autonomous driving decision-making method provided in the embodiments of this disclosure, the technical problem solved by this method will be explained first.

[0029] In existing technologies, when a target autonomous vehicle (hereinafter referred to as "vehicle") makes decisions with surrounding traffic participants (hereinafter referred to as "agents" or "obstacles"), there are currently two main approaches. One is a model-based approach, which quantifies the interaction between traffic participants and the vehicle by constructing a mathematical model to find the optimal decision-making strategy. The mainstream approaches in this category are game theory approaches and Markov decision / partially observable Markov decision approaches. The other approach is a data-driven approach, which directly encodes high-dimensional perception environment data and trains a network to directly output low-dimensional decision results.

[0030] Among these approaches, data-driven decision-making does not rely on developers' understanding of decision-making behavior or the design of reward functions, and can directly leverage massive amounts of data to continuously improve decision-making capabilities. However, this type of method suffers from the problem that vehicle behavior patterns are unknown, making it impossible to guarantee the safety and reliability of autonomous vehicle behavior in real-world projects without safety drivers.

[0031] Among them, the intelligent agent decision-making scheme based on game theory is able to find an equilibrium game decision result by taking advantage of the "selfishness" of vehicles and traffic participants. Therefore, this scheme naturally has the characteristic of eliminating the uncertainty of traffic participants' behavior, and the decision result obtained is more executable, thus possessing better safety and rationality.

[0032] Current game theory-based agent decision-making schemes mostly address the decision-making problem between the vehicle and a single surrounding agent. This involves acquiring the current state information of both the vehicle and the surrounding agents, simulating their interaction process, and finally obtaining the equilibrium game strategy through Nash equilibrium. The drawback of selecting only a single agent for decision-making is that in the real world, the vehicle's behavior is influenced by multiple obstacles. A decision optimal for only a single obstacle may become suboptimal or even unacceptable in a multi-obstacle environment, resulting in poor applicability.

[0033] There are two approaches to directly considering game decision-making involving multiple agents. The first approach is to directly perform a single-round multiplayer game. Because of the large number of participants, this type of game compresses the number of strategies available to each participant, resulting in a coarse solution and reduced solution quality. Furthermore, since it's a single-round game, the strategies of the participating agents are fixed and continue forward in the time dimension, which limits the flexibility of the solution strategy.

[0034] Another approach involves multi-round two-player game theory with multiple agents. Each round, the vehicle plays against only a single agent. Using a greedy strategy, it seeks the optimal strategy for each round. If an optimal strategy cannot be found (e.g., avoiding collisions with obstacles is not necessary), it backtracks, removing the previously found optimal strategy from its options. It then solves for the optimal strategy for the next obstacle and continues the two-player game until the optimal strategy for the last obstacle is found. Finally, it determines its strategy for all obstacles and infers the overall decision outcome for both the vehicle and all agents. This approach breaks down multi-player games into multi-round two-player games, reducing the strategy dimension of a single round and increasing the flexibility of the vehicle's strategy through its adjustability across multiple rounds. However, the method of obtaining a definite game result in each round and making decisions backward in a greedy manner cannot guarantee the optimality of the vehicle's strategy in multiple rounds of the game (that is, the local optimality of each game cannot guarantee the global optimality of all strategies of the vehicle).

[0035] Therefore, in order to solve the problem that the vehicle decision accuracy is low due to the best decision considering only a single obstacle in the prior art, the embodiments of this application provide an autonomous driving decision-making method. By deriving the benefits between the vehicle and each obstacle in sequence, a two-player game is played multiple times, and by updating the benefits backward from the final obstacle, the decision result is guaranteed to be globally optimal.

[0036] Figure 1 This is a flowchart illustrating an autonomous driving decision-making method according to an embodiment of this disclosure. The method can be executed by an autonomous driving decision-making device, which can be implemented in software and / or hardware and can be configured in an electronic device. Figure 1 As shown, the method may specifically include the following steps:

[0037] S110. Obtain the obstacle sequence corresponding to the vehicle. Starting from the first obstacle in the obstacle sequence and working backwards, determine the interaction benefit matrix between the next obstacle and the vehicle based on the currently derived interaction benefit matrix between the obstacle and the vehicle, until the final interaction benefit matrix between the obstacle and the vehicle in the obstacle sequence is obtained.

[0038] The obstacle sequence can consist of various obstacles located along the vehicle's planned path, such as other vehicles and pedestrians. Specifically, the obstacles in the obstacle sequence can be ordered in ascending order of their distance from the vehicle.

[0039] Optionally, obtain the obstacle sequence corresponding to the vehicle, including:

[0040] Obtain the predicted driving trajectory of all obstacles and the planned path of the vehicle, and remove obstacles that do not intersect with the planned path; for each obstacle, determine the intersection point between the predicted driving trajectory of the obstacle and the planned path; sort all obstacles in order of distance from the intersection point to the vehicle from near to far to obtain the obstacle sequence corresponding to the vehicle.

[0041] The obstacle can be any other object detected by the vehicle's perception module. For each obstacle detected by the vehicle's perception module, the vehicle's prediction module can determine the predicted trajectory of that obstacle.

[0042] Specifically, the predicted trajectory of each obstacle can be compared with the planned path of the vehicle. If there is no intersection between the predicted trajectory and the planned path, the corresponding obstacle can be ignored, while obstacles that intersect between the predicted trajectory and the planned path are retained.

[0043] Furthermore, all obstacles can be sorted in order of proximity to the vehicle's front based on the distance between the initial intersection point and the vehicle's front end. This ensures that obstacles whose intersection points are closer to the vehicle's current position are decided first, and obstacles whose intersection points are farther from the vehicle's current position are decided later. The sorting results form an obstacle sequence corresponding to the vehicle.

[0044] Using the above method, the obstacle sequence can be obtained based on the predicted driving trajectory of the obstacles, ensuring that subsequent decisions are made on all obstacles that intersect with the vehicle. Furthermore, it can ensure that each obstacle is decided in the order in which it intersects with the vehicle, taking into account the impact of the currently decided obstacle on the next obstacle, thus ensuring the accuracy of the decision.

[0045] Specifically, after obtaining the sequence of obstacles corresponding to the vehicle, starting from the first obstacle in the sequence, the interaction benefit matrix between the second obstacle and the vehicle is calculated based on the interaction benefit matrix between the first obstacle and the vehicle. Then, based on the interaction benefit matrix between the second obstacle and the vehicle, the interaction benefit matrix between the third obstacle and the vehicle is calculated, and so on, until the final interaction benefit matrix between the obstacles and the vehicle is obtained.

[0046] The interaction payoff matrix can include the payoffs of each game strategy. A game strategy can consist of the vehicle's first strategy (the vehicle's longitudinal acceleration) and the obstacle's second strategy (the obstacle's longitudinal acceleration); the payoffs of each game strategy include the vehicle's payoff and the obstacle's payoff.

[0047] In this embodiment, at least one of the vehicle's safety, ride comfort, and efficiency can be evaluated according to the game strategy to obtain at least one vehicle-specific benefit. Similarly, at least one of the obstacle's safety, ride comfort, and efficiency can be evaluated according to the game strategy to obtain at least one obstacle-specific benefit.

[0048] For example, Figure 2This is a schematic diagram illustrating the derivation of payoffs in one embodiment of this disclosure. Matrices 1_1, 2_1, etc., can be understood as interaction payoff matrices. Taking matrix 1_1 as an example, the horizontal axis represents the vehicle's first strategy, and the vertical axis represents the obstacle's second strategy. There are 2 first strategies and 3 second strategies. Combining all first strategies and all second strategies in matrix 1_1 yields 6 game strategies. Assuming the vehicle's individual payoff is 2 and the obstacle's individual payoff is 2, the payoff for each game strategy can be represented as [obj_cost1, obj_cost2, ego_cost1, ego_cost2], where obj_cost1 and obj_cost2 are the obstacle's individual payoffs, and ego_cost1 and ego_cost2 are the vehicle's individual payoffs.

[0049] In one specific implementation, starting from the first obstacle in the obstacle sequence and proceeding backwards, based on the currently derived interaction benefit matrix between the obstacle and the vehicle, the interaction benefit matrix between the next obstacle and the vehicle is determined, until the final interaction benefit matrix between the obstacle and the vehicle in the obstacle sequence is obtained, including the following steps:

[0050] Step 1: Take the first obstacle in the obstacle sequence as the current deduced obstacle and obtain the decision state information corresponding to the current deduced obstacle. The decision state information is used to describe the state of the current deduced obstacle and the vehicle at the corresponding decision time.

[0051] Step 2: Determine the first strategy set of the vehicle and the second strategy set of the current inferred obstacle based on the decision state information corresponding to the current inferred obstacle, and determine the interaction benefit matrix between the current inferred obstacle and the vehicle according to the first strategy set and the second strategy set;

[0052] Step 3: Based on each of the first strategies in the first strategy set, determine the decision state information corresponding to the next obstacle after the current deduced obstacle, and take the next obstacle as the new current deduced obstacle. Return to the step of determining the first strategy set and the second strategy set based on the decision state information corresponding to the current deduced obstacle, until the interaction benefit matrix between the final obstacle and the vehicle in the obstacle sequence is obtained.

[0053] In step 1 above, when the first obstacle is taken as the current inference obstacle, the current moment can be taken as the decision moment, and the corresponding decision state information can be obtained based on the current state of the vehicle and the obstacle at the current moment.

[0054] The decision status information may include the distance between the vehicle and the junction point, the distance the vehicle has traveled from the junction point, the vehicle's speed, the vehicle's acceleration, the speed limit for entering the junction point, the speed limit for leaving the junction point, the distance between the obstacle and the junction point, the distance the obstacle has traveled from the junction point, the obstacle's speed, the obstacle's acceleration, the obstacle's speed limit, and the obstacle's decision result in the previous decision cycle. For example, the decision status information is defined as follows:

[0055] typedef struct Interaction_Model

[0056] {

[0057] float32_t ego_enter_intersec_s; / * Distance between the vehicle and the intersection point * /

[0058] float32_t ego_leave_intersec_s; / * Distance the vehicle has traveled from the intersection point * /

[0059] float32_t ego_vel; / * Vehicle speed * /

[0060] float32_t ego_acc; / * Vehicle acceleration * /

[0061] float32_t ego_start_max_vel; / * Speed ​​limit for vehicles entering the junction * /

[0062] float32_t ego_end_max_vel; / * Speed ​​limit for vehicles leaving the intersection * /

[0063] float32_t obj_enter_intersec_s; / * Distance between the obstacle and the intersection point * /

[0064] float32_t obj_leave_intersec_s; / * The distance the obstacle has traveled from the intersection point * /

[0065] float32_t obj_vel; / * The speed of the obstacle * /

[0066] float32_t obj_acc; / * Acceleration of the obstacle * /

[0067] float32_t obj_max_vel; / * Speed ​​limiter for obstacles * /

[0068] lon_decision_t history_decision; / * The decision result of the obstacle in the previous decision cycle * /

[0069] }interaction_model_t;

[0070] Furthermore, in step 2 above, the first set of strategies for the vehicle and the second set of strategies for the current obstacle can be obtained based on the decision state information corresponding to the current obstacle.

[0071] For example, based on the vehicle's current speed (ego_vel), the speed limit at the intersection (ego_start_max_vel and / or ego_end_max_ve) from the decision state information, and the pre-configured preset acceleration range [-max_dec1, max_acc1] and preset sampling interval (step_acc1), a first set of strategies that can be adopted when the vehicle plays against the obstacle can be obtained. Furthermore, based on the obstacle's preset acceleration range [-max_dec2, max_acc2] and preset sampling interval (step_acc2), a second set of strategies that can be adopted when the obstacle plays against the sum can be obtained.

[0072] Furthermore, the first strategy in the first strategy set and the second strategy in the second strategy set can be combined to obtain various game strategies, and then the payoff of each game strategy can be determined to obtain the interaction payoff matrix between the vehicle and the currently derived obstacle. For example... Figure 2 As shown, the current state can be used as the decision state information corresponding to the first obstacle, and then the interaction benefit matrix between the first obstacle and the vehicle can be calculated.

[0073] In this embodiment of the disclosure, since the benefits of each obstacle are derived sequentially, all feasible strategies for an obstacle need to be downsampled, which may lead to dimensionality explosion. Therefore, in step 2 above, considering that there may be multiple first strategies that are similar, similar strategies can also be pruned in order to solve the dimensionality explosion problem.

[0074] For example, regarding step 2 above, optionally, determining the vehicle's first strategy set based on the current deduced decision state information corresponding to the obstacle includes the following steps:

[0075] Step 21: Based on the current vehicle speed, the speed limit of the intersection point between the vehicle and the current obstacle in the decision state information corresponding to the current inferred obstacle, as well as the vehicle's preset acceleration range and preset sampling interval, determine the vehicle's candidate strategy set.

[0076] Step 22: For each first strategy in the candidate strategy set, determine the estimated time for the vehicle to reach the intersection using the first strategy;

[0077] Step 23: Based on the time difference between the expected times corresponding to each first strategy, some first strategies in the candidate strategy set are eliminated to obtain the first strategy set for the vehicle.

[0078] That is, multiple accelerations can be sampled in a preset acceleration range based on the preset sampling interval, the vehicle's current speed, and the speed limit at the intersection point, to obtain a set of candidate strategies for the vehicle. For each first strategy in the candidate strategy set, the time when the vehicle arrives at the intersection point, i.e. the estimated time, is calculated based on the first strategy, the current speed, and the distance between the vehicle and the intersection point.

[0079] For example, the estimated time can be calculated using the following formula:

[0080] ego_vel·t+0.5·stragety_acct 2 =ego_enter_intersec_s;

[0081] In the formula, ego_vel is the vehicle's current speed, stragety_acc is the acceleration corresponding to the first strategy, t is the estimated time, and ego_enter_intersec_s is the distance between the vehicle and the intersection point.

[0082] Furthermore, by calculating the time difference between the expected times corresponding to each first strategy, for example, by calculating the time difference between any two expected times, similar first strategies are determined, and then similar first strategies are removed from the candidate strategy set to obtain the first strategy set.

[0083] In one example, based on the time difference between the expected times corresponding to each first strategy, a portion of the first strategies in the candidate strategy set are eliminated, including the following steps:

[0084] Step 231: Among all the estimated times corresponding to the first strategy, the estimated time corresponding to the first strategy that satisfies the current acceleration of the vehicle is determined as the reference time for the vehicle to arrive at the intersection point, wherein the current acceleration is the acceleration of the vehicle in the decision state information corresponding to the obstacle currently derived.

[0085] Step 232: Select the estimated time adjacent to the reference time as the time to be judged;

[0086] Step 233: For each time to be judged, determine whether the time difference between the time to be judged and the reference time is less than a preset time threshold;

[0087] Step 234: If yes, remove the first strategy corresponding to the time to be judged, and select a new time to be judged from other estimated times adjacent to the time to be judged. Then return to the step of determining whether the time difference between the time to be judged and the base time is less than the preset time threshold.

[0088] Step 235: If not, then take the time to be judged as the new base time, and return to execute the step of taking the expected time adjacent to the base time in the sorting result as the time to be judged, until the judgment of all expected times is completed.

[0089] That is, the estimated time to reach the intersection point under the first strategy corresponding to the vehicle's current acceleration can be used as the base time to find similar first strategies starting from the first strategy corresponding to the current acceleration.

[0090] Specifically, the estimated time adjacent to the reference time can be used as the time to be judged. Further, it is judged whether the time difference between the time to be judged and the reference time is less than a preset time threshold. If so, it means that the first strategy corresponding to the time to be judged is similar to the first strategy corresponding to the reference time. The first strategy corresponding to the time to be judged can be eliminated. Then, a new time to be judged is selected from other estimated times adjacent to the time to be judged and the execution step 233 is returned until the time difference between the time to be judged and the reference time is not less than the preset time threshold. At this time, the time to be judged can be used as the new reference time and the execution step 232 is returned until all estimated times are judged.

[0091] For example, Figure 3 This is a schematic diagram of a pruning strategy in an embodiment of this disclosure. Figure 3 As shown, the estimated time corresponding to the current acceleration is first used as the base time, and the first strategy is filtered from both sides. Then, after the first strategy with a time difference greater than the preset time threshold is found, the time corresponding to the first strategy is used as the new base time, and the filtering continues.

[0092] In order to adapt to the increase in sampling caused by the sequential derivation of obstacles and avoid the problem of decision dimension explosion, the above implementation method proposes a strategy merging and pruning process for vehicles. While ensuring the sampling space, it can reasonably limit the sampling density and improve the decision efficiency of vehicles.

[0093] After obtaining the interaction benefit matrix between the current inferred obstacle and the vehicle, further, based on each of the first strategies in the first strategy set, through dynamic simulation, the decision state information corresponding to the next obstacle when the vehicle travels to the next obstacle according to each of the first strategies can be derived. The decision time of the next obstacle can be the time of the intersection between the vehicle overtaking or yielding and the current inferred obstacle. Then, the next obstacle is taken as the new current inferred obstacle, and step 2 above is executed again to obtain the interaction benefit matrix between the new current inferred obstacle and the vehicle.

[0094] like Figure 2 As shown, when making a decision regarding obstacle_1, the vehicle has two first strategies, corresponding to the two columns in matrix 1_1. Each first strategy can derive a decision state information corresponding to obstacle_2. Based on these two decision state information, matrices 2_1 and 2_2 under obstacle_2 can be obtained respectively. Similarly, the two first strategies in matrix 2_1 can further derive matrices 3_1 and 3_2, and the two first strategies in matrix 2_2 can further derive matrices 3_3 and 3_4. From... Figure 2 As can be seen, the whole process can be understood as calculating the payoff of the game strategy and generating the decision tree from the bottom up.

[0095] It should be noted that the interaction benefit matrix of each obstacle calculated at this time is only related to the historical state (the decision state of each obstacle before the corresponding obstacle), and is not related to the subsequent decision results.

[0096] Through steps 1-3 above, the impact of the decision on the previous obstacle on the decision on the next obstacle is considered. The benefits are derived sequentially from the first obstacle, thus avoiding the problem of conflicting decision results when making decisions on multiple obstacles.

[0097] S120. Starting from the last obstacle in the obstacle sequence and moving forward, based on the payoff of each game strategy in the interaction payoff matrix between the current obstacle and the vehicle, determine the local optimal game strategy corresponding to the current obstacle, and update the interaction payoff matrix between the previous obstacle and the vehicle based on the payoff of the local optimal game strategy, until the interaction payoff matrix between the first obstacle and the vehicle is updated.

[0098] The game strategies include a first strategy describing the longitudinal acceleration of the vehicle, a second strategy describing the longitudinal acceleration of the obstacle, and the vehicle's right-of-way behavior. Right-of-way behavior can involve either overtaking an obstacle or yielding to an obstacle.

[0099] Specifically, after obtaining the interaction payoff matrices between the final obstacle and the vehicle in the obstacle sequence, the first strategy that the vehicle can adopt in each interaction payoff matrix, apart from affecting the current two parties in the game (i.e., the vehicle and the final obstacle), will not have any additional impact on the vehicle's future payoffs. Therefore, based on the interaction payoff matrices between the final obstacle and the vehicle, combined with methods such as Nash equilibrium, the locally optimal game strategy corresponding to the final obstacle can be obtained. It should be noted that since the decision has no subsequent impact, the locally optimal game strategy corresponding to the final obstacle obtained at this time has global optimality.

[0100] Furthermore, based on the local optimal game strategy corresponding to the final obstacle, the interaction payoff matrix between the previous obstacle and the vehicle can be updated in reverse, and the local optimal game strategy can be determined in the interaction payoff matrix of the previous obstacle. This process continues until the interaction payoff matrix of the first obstacle is updated.

[0101] In one specific implementation, starting from the final obstacle in the obstacle sequence and moving forward, based on the payoffs of each game strategy in the interaction payoff matrix between the currently updated obstacle and the vehicle, the locally optimal game strategy corresponding to the currently updated obstacle is determined, and the interaction payoff matrix between the previous obstacle and the vehicle is updated based on the payoff of the locally optimal game strategy, until the interaction payoff matrix between the first obstacle and the vehicle is updated, including the following steps:

[0102] Step 4: Take the final obstacle in the obstacle sequence as the current updated obstacle. Based on the payoff of each game strategy in the interaction payoff matrix between the current updated obstacle and the vehicle, determine the local optimal game strategy corresponding to the current updated obstacle among each game strategy.

[0103] Step 5: Based on the vehicle's individual payoff in the local optimal game strategy, update the vehicle's individual payoff in the game strategy corresponding to the interaction payoff matrix between the previous obstacle and the vehicle.

[0104] Step 6: Take the previous obstacle as the new current updated obstacle, return to the step of determining the local optimal game strategy corresponding to the current updated obstacle in each game strategy, until the interaction payoff matrix between the first obstacle and the vehicle is updated.

[0105] That is, the final obstacle can be used as the current updated obstacle, and for each interaction payoff matrix of the current updated obstacle, the locally optimal game strategy can be determined based on the payoff of each game strategy. For example, Figure 4 This is a schematic diagram illustrating a revenue update in one embodiment of this disclosure, such as... Figure 4 As shown, matrices 3_1 to 3_4 of obstacle_3 can each determine a local optimal game strategy.

[0106] Furthermore, the vehicle's individual payoff in the interaction payoff matrix of the previous obstacle can be updated based on the vehicle's individual payoff (which could be smoothness payoff, safety payoff, or efficiency payoff) of the local optimal game strategy.

[0107] like Figure 4 As shown, taking the locally optimal game strategy in matrix 3_1 as an example, since matrix 3_1 is derived from the first strategy of the vehicle corresponding to the first column of matrix 2_1, the vehicle's individual payoff in the first column of matrix 2_1 can be updated based on the individual payoff of the vehicle in the locally optimal game strategy in matrix 3_1. Similarly, matrix 3_4 is derived from the first strategy of the vehicle corresponding to the second column of matrix 2_2, so the vehicle's individual payoff in the second column of matrix 2_2 can be updated based on the individual payoff of the vehicle in the locally optimal game strategy in matrix 3_4. After updating the first and second columns of matrix 2_1, the locally optimal game strategy can be determined in the entire matrix 2_1, that is, the game strategy with the best payoff among the 6 game strategies is determined as the locally optimal game strategy, and matrix 1_1 is updated based on the locally optimal game strategy of matrix 2_1.

[0108] After updating the individual vehicle payoffs for the obstacle preceding the final obstacle, the preceding obstacle can be used as the new currently updated obstacle. The process returns to the step of determining the locally optimal game strategy in the interaction payoff matrix to continue updating the interaction payoff matrix of the preceding obstacles until the interaction payoff matrix of the first obstacle is updated. See also Figure 4 Based on the vehicle payoff of the local optimal game strategy in matrix 2_1, the vehicle payoff of the first game strategy in matrix 1_1 can be updated. Based on the vehicle payoff of the local optimal game strategy in matrix 2_2, the vehicle payoff of the second game strategy in matrix 1_1 can be updated.

[0109] Through steps 4-6 above, the vehicle's individual payoff for each first strategy in the first strategy set of each decision state information is updated by superimposing the vehicle's individual payoff for the corresponding local optimal game strategy of the next obstacle. Then, based on the vehicle's individual payoff and the obstacle's individual payoff, the local optimal game strategy is solved using a method such as Nash equilibrium.

[0110] In other words, after the payoff of the final obstacle is derived, the globally optimal game strategy corresponding to the final obstacle can be obtained. The individual payoffs of each vehicle can be directly retrieved. These individual payoffs are obtained from the upper-level vehicle strategies that generate their decision state information. Therefore, the individual payoffs of the corresponding vehicles can be superimposed on the individual payoffs of the corresponding upper-level vehicle strategies. In this way, the subsequent impact of the upper-level vehicle strategies is directly reflected in the individual payoffs of the current vehicle's corresponding strategy through superposition, thereby ensuring that in multi-round two-player games, the vehicle's payoffs can reflect the game results of subsequent obstacles.

[0111] In this embodiment of the disclosure, the interaction benefit matrix of the vehicle's individual benefit based on the local optimal game strategy is updated to the previous obstacle. This can be done by accumulating the benefits or by summing the worst benefit.

[0112] Optionally, for step 5 above, based on the vehicle's individual payoff in the local optimal game strategy, the vehicle's individual payoff in the interaction payoff matrix between the previous obstacle and the vehicle is updated, including:

[0113] For each type of vehicle benefit, determine the minimum vehicle benefit among the vehicle benefit of the local best game strategy corresponding to the current updated obstacle and the vehicle benefit of the game strategy corresponding to the interaction benefit matrix of the previous obstacle; update the vehicle benefit of the game strategy corresponding to the interaction benefit matrix of the previous obstacle based on the minimum vehicle benefit.

[0114] That is, the vehicle's individual payoff for the local optimal game strategy can be compared with the vehicle's individual payoff for the corresponding game strategy in the interaction payoff matrix of the previous obstacle to obtain the minimum individual payoff, and the vehicle's individual payoff for the corresponding game strategy in the interaction payoff matrix of the previous obstacle can be updated to the minimum individual payoff.

[0115] For example, it can be represented by the following formula:

[0116] ego_costi=min(ego_costi,ego_cost_i ′ );

[0117] In the formula, ego_costi on the left side of the equal sign represents the individual revenue of the i-th vehicle after the previous obstacle update, and ego_costi on the right side of the equal sign represents the individual revenue of the i-th vehicle before the previous obstacle update. ′ Let i be the vehicle's single-item payoff under the current locally optimal game strategy of updating obstacles.

[0118] The above method enables payout updates based on the worst-case payout, avoiding payout inflation that could affect the solution of the global optimal game strategy.

[0119] S130. Based on the interaction benefit matrix between each obstacle and the vehicle, determine the global optimal game strategy corresponding to each obstacle, and obtain the target decision result of the vehicle.

[0120] Specifically, after updating the interaction payoff matrix of the first obstacle, the global optimal game strategy of the first obstacle can be obtained. Then, the global optimal game strategy of the next obstacle derived under the corresponding first strategy can be found, until the global optimal game strategy of the final obstacle is obtained.

[0121] It should be noted that once the globally optimal strategy for the first obstacle is determined, the globally optimal strategy for the vehicle against all obstacles in the obstacle sequence can be obtained by sequentially querying the globally optimal strategy for the next obstacle. Since each round of the two-player game considers the future payoff impact of the vehicle's first strategy, the overall vehicle decision strategy obtained through the optimal strategy of a single round possesses globally optimal characteristics.

[0122] In one specific implementation, the globally optimal game strategy for each obstacle is determined based on the interaction payoff matrix between each obstacle and the vehicle, thus obtaining the vehicle's target decision result. This includes the following steps:

[0123] Step 7: Based on the payoff of each game strategy in the interaction payoff matrix between the first obstacle and the vehicle, determine the global optimal game strategy corresponding to the first obstacle, and take the next obstacle after the first obstacle as the current decision obstacle.

[0124] Step 8: Among all the interaction payoff matrices of the current decision obstacle, determine the interaction payoff matrix corresponding to the global best game strategy of the previous obstacle as the target matrix, determine the local best game strategy in the target matrix as the global best game strategy corresponding to the current decision obstacle, and take the next obstacle of the current decision obstacle as the new current decision obstacle. Return to the step of determining the target matrix among all the interaction payoff matrices of the current decision obstacle, until the global best game strategy corresponding to the final obstacle is obtained.

[0125] That is, the best-paying game strategy can be determined by Nash equilibrium in the interaction payoff matrix of the first obstacle, and this strategy can be used as the global best game strategy corresponding to the first obstacle.

[0126] Furthermore, taking the latter obstacle as the current decision obstacle, from all the interaction payoff matrices of the current decision obstacle, the interaction payoff matrix corresponding to the global best game strategy of the former obstacle is taken as the target matrix. Then, the local best game strategy in the target matrix (i.e. the game strategy with the best payoff in the target matrix) is determined as the global best game strategy corresponding to the current decision obstacle. This process is repeated until the global best game strategy corresponding to the final obstacle is obtained.

[0127] For example, Figure 5 This is a schematic diagram illustrating the search for a globally optimal game strategy in an embodiment of this disclosure. For example... Figure 5 As shown, the global optimal strategy is first queried from matrix 1_1 of obstacle_1. Since matrix 2_1 is derived from the first strategy in the first column of matrix 1_1, under obstacle_2, matrix 2_1 can be used as the target matrix, and the local optimal strategy in matrix 2_1 can be used as the global optimal strategy. Since matrix 3_1 is derived from the first strategy in the first column of matrix 2_1, under obstacle_3, matrix 3_1 can be used as the target matrix, and the local optimal strategy in matrix 3_1 can be used as the global optimal strategy.

[0128] After obtaining the global optimal game strategy corresponding to each obstacle in the obstacle sequence, the vehicle's yielding behavior to each obstacle, the longitudinal acceleration taken by the vehicle when interacting with each obstacle, and the longitudinal acceleration taken by each obstacle when interacting with the vehicle can be determined based on the global optimal game strategy corresponding to each obstacle. The vehicle's target decision result can then be obtained. Furthermore, the target decision result can be output to the vehicle's planning module for decision response.

[0129] In this embodiment, by using backpropagation of benefits, the global optimality of multi-obstacle decision-making is guaranteed without changing the core algorithm of single-round two-person decision-making, thus avoiding the problem of conflicting decision results when making decisions on multiple obstacles. While ensuring global optimality of multi-obstacle decision-making, the overall algorithm complexity is not significantly increased, making the solution highly feasible.

[0130] The autonomous driving decision-making method provided in this embodiment obtains a sequence of obstacles corresponding to the vehicle. Starting from the first obstacle and working backwards, based on the currently derived interaction benefit matrix between the obstacle and the vehicle, it determines the interaction benefit matrix between the next obstacle and the vehicle, until the final interaction benefit matrix between the obstacle and the vehicle is obtained. Then, starting from the final obstacle in the obstacle sequence and working backwards, it updates the interaction benefit matrix between the previous obstacle and the vehicle based on the benefits of each game strategy in the currently updated interaction benefit matrix between the obstacle and the vehicle, until the interaction benefit matrix between the first obstacle and the vehicle is updated. Finally, based on the obstacles... The interaction payoff matrix between the vehicle and the obstacle is used to obtain the globally optimal game strategy for each obstacle, thus completing the vehicle's decision-making. This method considers multiple obstacles corresponding to the vehicle and derives the interaction payoff matrix of each obstacle from the first obstacle backwards. It takes into account the payoff of the vehicle's interaction with each obstacle, solving the problem of low vehicle decision-making accuracy caused by considering only the optimal decision of a single obstacle in the existing technology. Furthermore, updating the interaction payoff matrix backwards from the final obstacle can ensure that the payoff of the vehicle and each obstacle is globally optimal without limiting the flexibility of the vehicle's decision-making strategy, avoiding the problem of conflicting decision results when making decisions on multiple obstacles.

[0131] Figure 6 This is a schematic diagram of the structure of an autonomous driving decision-making device according to an embodiment of this disclosure. Figure 6 As shown: The device includes: a revenue derivation module 610, a revenue update module 620, and a decision module 630.

[0132] The revenue derivation module 610 is used to obtain the obstacle sequence corresponding to the vehicle, starting from the first obstacle in the obstacle sequence and proceeding backwards, based on the currently derived interaction revenue matrix between the obstacle and the vehicle, to determine the interaction revenue matrix between the next obstacle and the vehicle, until the final obstacle in the obstacle sequence and the vehicle are obtained.

[0133] The revenue update module 620 is used to start from the last obstacle in the obstacle sequence and move forward, based on the revenue of each game strategy in the interaction revenue matrix between the currently updated obstacle and the vehicle, determine the local optimal game strategy corresponding to the currently updated obstacle, and update the interaction revenue matrix between the previous obstacle and the vehicle based on the revenue of the local optimal game strategy, until the interaction revenue matrix between the first obstacle and the vehicle is updated. The game strategy includes a first strategy describing the longitudinal acceleration of the vehicle, a second strategy describing the longitudinal acceleration of the obstacle, and the vehicle's yielding behavior.

[0134] The decision module 630 is used to determine the global optimal game strategy corresponding to each obstacle based on the interaction benefit matrix between each obstacle and the vehicle, and to obtain the target decision result of the vehicle.

[0135] Optionally, the revenue derivation module 610 includes a sequence acquisition unit; the sequence acquisition unit is used to acquire the predicted driving trajectories of all obstacles and the planned path of the vehicle, and to remove obstacles that do not intersect with the planned path; for each obstacle, to determine the intersection point between the predicted driving trajectory of the obstacle and the planned path; and to sort all obstacles in order of distance from the intersection point to the vehicle from closest to farthest to obtain the obstacle sequence corresponding to the vehicle.

[0136] Optionally, the benefit derivation module 610 includes a derivation unit; the derivation unit is configured to take the first obstacle in the obstacle sequence as the current derivation obstacle, obtain the decision state information corresponding to the current derivation obstacle, wherein the decision state information is used to describe the state of the current derivation obstacle and the vehicle at the corresponding decision time; determine the first strategy set of the vehicle and the second strategy set of the current derivation obstacle based on the decision state information corresponding to the current derivation obstacle, and determine the interaction benefit matrix between the current derivation obstacle and the vehicle according to the first strategy set and the second strategy set; determine the decision state information corresponding to the next obstacle of the current derivation obstacle based on each first strategy in the first strategy set, and take the next obstacle as the new current derivation obstacle, and return to execute the step of determining the first strategy set and the second strategy set based on the decision state information corresponding to the current derivation obstacle, until the interaction benefit matrix between the final obstacle in the obstacle sequence and the vehicle is obtained.

[0137] Optionally, the derivation unit is further configured to determine a candidate strategy set for the vehicle based on the decision state information corresponding to the currently derived obstacle, including the vehicle's current speed, the speed limit of the vehicle passing through the intersection point with the currently derived obstacle, and the vehicle's preset acceleration range and preset sampling interval; for each first strategy in the candidate strategy set, determine the estimated time for the vehicle to reach the intersection point using the first strategy; and based on the time difference between the estimated times corresponding to each first strategy, eliminate some first strategies in the candidate strategy set to obtain the vehicle's first strategy set.

[0138] Optionally, the derivation unit is further configured to, among all the estimated times corresponding to the first strategies, determine the estimated time corresponding to the first strategy that satisfies the current acceleration of the vehicle as the reference time for the vehicle to arrive at the intersection point, wherein the current acceleration is the vehicle's acceleration in the decision state information corresponding to the currently derivation obstacle; take the estimated time adjacent to the reference time as the time to be judged; for each time to be judged, determine whether the time difference between the time to be judged and the reference time is less than a preset time threshold; if so, remove the first strategy corresponding to the time to be judged, and reselect a new time to be judged from other estimated times adjacent to the time to be judged, and return to execute the step of determining whether the time difference between the time to be judged and the reference time is less than the preset time threshold; if not, take the time to be judged as the new reference time, and return to execute the step of taking the estimated time adjacent to the reference time in the sorting result as the time to be judged, until the judgment of all estimated times is completed.

[0139] Optionally, the revenue update module 620 includes an update unit; the update unit is configured to take the final obstacle in the obstacle sequence as the current updated obstacle, and based on the revenue of each game strategy in the interaction revenue matrix between the current updated obstacle and the vehicle, determine the local optimal game strategy corresponding to the current updated obstacle in each game strategy; based on the vehicle's individual revenue of the local optimal game strategy, update the vehicle's individual revenue of the game strategy corresponding to the interaction revenue matrix between the previous obstacle and the vehicle; take the previous obstacle as the new current updated obstacle, and return to the step of determining the local optimal game strategy corresponding to the current updated obstacle in each game strategy, until the interaction revenue matrix between the first obstacle and the vehicle is updated.

[0140] Optionally, the updating unit is further configured to, for each type of vehicle individual benefit, determine the minimum value of the vehicle individual benefit of the local best game strategy corresponding to the currently updated obstacle and the vehicle individual benefit of the game strategy corresponding to the interaction benefit matrix of the previous obstacle; and update the vehicle individual benefit of the game strategy corresponding to the interaction benefit matrix of the previous obstacle based on the minimum value of the individual benefit.

[0141] Optionally, the decision module 630 is specifically used for:

[0142] Based on the payoff of each game strategy in the interaction payoff matrix between the first obstacle and the vehicle, the global optimal game strategy corresponding to the first obstacle is determined, and the next obstacle after the first obstacle is taken as the current decision obstacle.

[0143] In the interaction payoff matrices of the current decision obstacle, the interaction payoff matrix corresponding to the global best game strategy of the previous obstacle is determined as the target matrix. The local best game strategy in the target matrix is ​​determined as the global best game strategy corresponding to the current decision obstacle. The next obstacle of the current decision obstacle is taken as the new current decision obstacle. The step of determining the target matrix in the interaction payoff matrices of the current decision obstacle is returned to until the global best game strategy corresponding to the final obstacle is obtained.

[0144] The autonomous driving decision-making device provided in this disclosure can execute the steps in the autonomous driving decision-making method provided in this disclosure, and has the execution steps and beneficial effects, which will not be elaborated here.

[0145] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. See below for details. Figure 7 It shows a schematic diagram of a structure suitable for implementing the electronic device 500 in the embodiments of this disclosure. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0146] like Figure 7 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 501, which can perform various appropriate actions and processes to implement the methods of the embodiments described herein, based on a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing device 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0147] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts, thereby implementing the autonomous driving decision-making method as described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0148] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0149] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the autonomous driving decision-making method provided in any of the above embodiments.

[0150] Optionally, when one or more of the above-described procedures are executed by the electronic device, the electronic device may also perform other steps described in the above embodiments.

[0151] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0152] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

Claims

1. An autonomous driving decision-making method, characterized in that, The method includes: Obtain the obstacle sequence corresponding to the vehicle. Starting from the first obstacle in the obstacle sequence and proceeding backward, based on the currently derived interaction benefit matrix between the obstacle and the vehicle, determine the interaction benefit matrix between the next obstacle and the vehicle, until the final obstacle in the obstacle sequence and the vehicle are obtained. Starting from the final obstacle in the obstacle sequence and moving forward, based on the payoffs of each game strategy in the interaction payoff matrix between the currently updated obstacle and the vehicle, a locally optimal game strategy corresponding to the currently updated obstacle is determined. The interaction payoff matrix between the previous obstacle and the vehicle is then updated based on the payoff of the locally optimal game strategy, until the interaction payoff matrix between the first obstacle and the vehicle is updated. The game strategy includes a first strategy describing the longitudinal acceleration of the vehicle, a second strategy describing the longitudinal acceleration of the obstacle, and the vehicle's yielding behavior. The vehicle's individual payoff of the locally optimal game strategy is compared with the vehicle's individual payoff of the corresponding game strategy in the interaction payoff matrix of the previous obstacle to obtain the minimum individual payoff. The vehicle's individual payoff of the corresponding game strategy in the interaction payoff matrix of the previous obstacle is then updated to the minimum individual payoff. Based on the interaction benefit matrix between each obstacle and the vehicle, the global optimal game strategy corresponding to each obstacle is determined, and the target decision result of the vehicle is obtained.

2. The method according to claim 1, characterized in that, The acquisition of the obstacle sequence corresponding to the vehicle includes: Obtain the predicted driving trajectory of all obstacles and the planned path of the vehicle, and remove obstacles that do not intersect with the planned path; For each obstacle, determine the intersection point between the predicted trajectory of the obstacle and the planned path; All obstacles are sorted in order of their distance from the intersection point to the vehicle, from closest to furthest, to obtain the obstacle sequence corresponding to the vehicle.

3. The method according to claim 1, characterized in that, Starting from the first obstacle in the obstacle sequence and proceeding backwards, based on the currently derived interaction benefit matrix between the obstacle and the vehicle, determining the interaction benefit matrix between the next obstacle and the vehicle, until obtaining the interaction benefit matrix between the final obstacle and the vehicle in the obstacle sequence, includes: The first obstacle in the obstacle sequence is taken as the current deduced obstacle, and the decision state information corresponding to the current deduced obstacle is obtained. The decision state information is used to describe the state of the current deduced obstacle and the vehicle at the corresponding decision time. Based on the decision state information corresponding to the currently derived obstacle, a first strategy set for the vehicle and a second strategy set for the currently derived obstacle are determined, and an interaction benefit matrix between the currently derived obstacle and the vehicle is determined based on the first strategy set and the second strategy set. Based on each first strategy in the first strategy set, determine the decision state information corresponding to the next obstacle of the current deduced obstacle, and take the next obstacle as the new current deduced obstacle. Then return to the step of determining the first strategy set and the second strategy set based on the decision state information corresponding to the current deduced obstacle, until the interaction benefit matrix between the final obstacle in the obstacle sequence and the vehicle is obtained.

4. The method according to claim 3, characterized in that, The step of determining the first set of strategies for the vehicle based on the decision state information corresponding to the currently derived obstacle includes: Based on the current vehicle speed, the speed limit of the vehicle passing through the intersection point with the current obstacle, the vehicle's preset acceleration range and preset sampling interval in the decision state information corresponding to the current inferred obstacle, the candidate strategy set of the vehicle is determined. For each first strategy in the candidate strategy set, determine the estimated time for the vehicle to reach the intersection using the first strategy; Based on the time difference between the expected times corresponding to each first strategy, some first strategies in the candidate strategy set are eliminated to obtain the first strategy set of the vehicle.

5. The method according to claim 4, characterized in that, The step of eliminating a portion of the first strategies in the candidate strategy set based on the time difference between the expected times corresponding to each first strategy includes: Among all the estimated times corresponding to the first strategy, the estimated time corresponding to the first strategy that satisfies the current acceleration of the vehicle is determined as the reference time for the vehicle to reach the intersection point, wherein the current acceleration is the acceleration of the vehicle in the decision state information corresponding to the currently derived obstacle; The estimated time adjacent to the reference time is taken as the time to be judged; For each time to be judged, determine whether the time difference between the time to be judged and the reference time is less than a preset time threshold; If so, the first strategy corresponding to the time to be judged is removed, and a new time to be judged is selected from other estimated times adjacent to the time to be judged. Then, the process returns to the step of determining whether the time difference between the time to be judged and the reference time is less than the preset time threshold. If not, the time to be judged is taken as the new base time, and the process returns to the step of taking the estimated time adjacent to the base time in the sorting result as the time to be judged, until the judgment of all estimated times is completed.

6. The method according to claim 3, characterized in that, Starting from the final obstacle in the obstacle sequence and moving forward, based on the payoffs of each game strategy in the interaction payoff matrix between the currently updated obstacle and the vehicle, a locally optimal game strategy corresponding to the currently updated obstacle is determined, and the interaction payoff matrix between the previous obstacle and the vehicle is updated based on the payoffs of the locally optimal game strategy, until the interaction payoff matrix between the first obstacle and the vehicle is updated, including: The final obstacle in the obstacle sequence is taken as the current updated obstacle. Based on the payoff of each game strategy in the interaction payoff matrix between the current updated obstacle and the vehicle, the local optimal game strategy corresponding to the current updated obstacle is determined among each game strategy. Based on the vehicle's individual payoff from the local optimal game strategy, the vehicle's individual payoff from the game strategy corresponding to the interaction payoff matrix between the previous obstacle and the vehicle is updated. The previous obstacle is taken as the new current updated obstacle. The process returns to the step of determining the local optimal game strategy corresponding to the current updated obstacle in each game strategy, until the interaction payoff matrix between the first obstacle and the vehicle is updated.

7. The method according to claim 6, characterized in that, The vehicle's individual payoff based on the locally optimal game strategy is updated by updating the vehicle's individual payoff for the corresponding game strategy in the interaction payoff matrix between the previous obstacle and the vehicle, including: For each type of vehicle benefit, the minimum value of the vehicle benefit is determined from the vehicle benefit of the local best game strategy corresponding to the currently updated obstacle and the vehicle benefit of the game strategy corresponding to the interaction benefit matrix of the previous obstacle. The vehicle's individual payoff is updated based on the minimum value of the individual payoff, corresponding to the game strategy in the interaction payoff matrix of the previous obstacle.

8. The method according to claim 1, characterized in that, The process of determining the globally optimal game strategy for each obstacle based on the interaction benefit matrix between each obstacle and the vehicle, and obtaining the target decision result for the vehicle, includes: Based on the payoff of each game strategy in the interaction payoff matrix between the first obstacle and the vehicle, the global optimal game strategy corresponding to the first obstacle is determined, and the next obstacle after the first obstacle is taken as the current decision obstacle. In the interaction payoff matrices of the current decision obstacle, the interaction payoff matrix corresponding to the global best game strategy of the previous obstacle is determined as the target matrix. The local best game strategy in the target matrix is ​​determined as the global best game strategy corresponding to the current decision obstacle. The next obstacle of the current decision obstacle is taken as the new current decision obstacle. The step of determining the target matrix in the interaction payoff matrices of the current decision obstacle is returned to until the global best game strategy corresponding to the final obstacle is obtained.

9. An autonomous driving decision-making device, characterized in that, include: The revenue derivation module is used to obtain the obstacle sequence corresponding to the vehicle. Starting from the first obstacle in the obstacle sequence and proceeding backward, based on the currently derived interaction revenue matrix between the obstacle and the vehicle, the interaction revenue matrix between the next obstacle and the vehicle is determined, until the interaction revenue matrix between the final obstacle and the vehicle in the obstacle sequence is obtained. The payout update module is used to start from the final obstacle in the obstacle sequence and move forward, determining the local optimal game strategy corresponding to the currently updated obstacle based on the payout of each game strategy in the interaction payout matrix between the currently updated obstacle and the vehicle, and updating the interaction payout matrix between the previous obstacle and the vehicle based on the payout of the local optimal game strategy, until the interaction payout matrix between the first obstacle and the vehicle is updated. The game strategy includes a first strategy describing the longitudinal acceleration of the vehicle, a second strategy describing the longitudinal acceleration of the obstacle, and the vehicle's yielding behavior. The vehicle's individual payout of the local optimal game strategy is compared with the vehicle's individual payout of the corresponding game strategy in the interaction payout matrix of the previous obstacle to obtain the minimum individual payout, and the vehicle's individual payout of the corresponding game strategy in the interaction payout matrix of the previous obstacle is updated to the minimum individual payout. The decision-making module is used to determine the globally optimal game strategy for each obstacle based on the interaction benefit matrix between each obstacle and the vehicle, and to obtain the target decision result of the vehicle.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Decision-making method and device and vehicle

    CN115246415A

  • System and method for ride order dispatching and vehicle repositioning

    US20200074353A1