Vehicle interactive decision-making method, device, electronic device and storage medium

By determining the lateral behavior semantics and obstacle prediction trajectory in autonomous driving vehicles and optimizing the longitudinal acceleration sequence in combination with the interaction model, the problem of decision-making accuracy caused by the unobservable intention of obstacles is solved, and the decision reliability of autonomous driving is improved.

CN117002536BActive Publication Date: 2025-09-16UISEE TECH BEIJING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311201666.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2025-09-16
Estimated Expiration
2043-09-18

AI Technical Summary

Technical Problem

In existing technologies, the intention of obstacles cannot be fully observed in autonomous driving, which reduces the accuracy of the vehicle's decision-making. In particular, when the vehicle's behavior changes under different conditions, it is difficult to accurately predict the intention of obstacles.

Method used

By determining multiple lateral behavior semantics in the pre-decision trajectory, the predicted trajectory of the obstacle is obtained, and the longitudinal behavior solution is gradually predicted based on the initial state information. Combined with the interaction model of the obstacle, the longitudinal acceleration sequence is calculated to achieve step-by-step state prediction and decision optimization.

Benefits of technology

It improves the accuracy of obstacle intention prediction, enhances the decision-making reliability of autonomous vehicles, and ensures safe driving in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117002536B_ABST
    Figure CN117002536B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a method, device, electronic device and storage medium for interactive decision-making of a vehicle. The method determines multiple lateral behavior semantics in a pre-decision trajectory of the current vehicle, starts from the first lateral behavior semantic, and sequentially determines each predicted state information under each lateral behavior semantic. Finally, based on the state benefits of each predicted state information under all lateral behavior semantics, the first longitudinal acceleration sequence of the current vehicle in the pre-decision trajectory, as well as the decision trajectory and corresponding second longitudinal acceleration sequence of each obstacle are obtained, thereby realizing the longitudinal behavior decision of the current vehicle and the lateral and longitudinal behavior decision of the obstacle. The method sequentially predicts the next state based on the longitudinal behavior interactively solved in one state, making the obstacle intention prediction in each state more accurate, solving the problem of low decision accuracy caused by changes in the vehicle behavior and the obstacle intention in different states, and improving the decision reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of autonomous driving technology, and in particular to a vehicle interactive decision-making method, device, electronic device, and storage medium. Background Art

[0002] For autonomous vehicles, the most challenging aspect of decision-making is handling interactions with obstacles. If the vehicle mispredicts the intentions of obstacles, it can make irrational decisions. Furthermore, the intentions of obstacles are constantly changing over time and are therefore not directly observable.

[0003] The interactive decision-making between the ego vehicle and the obstacle is actually a POMDP (Partially Observable Markov Decision Process) problem. Existing technologies usually use POMDP to predict the intention of the obstacle in the entire process at one time.

[0004] However, the behavior of the ego vehicle changes in different states, and the changes in the intentions of obstacles cannot be fully observed, which leads to reduced decision accuracy. Summary of the Invention

[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present disclosure provide a vehicle interactive decision-making method, device, electronic device and storage medium to achieve gradual prediction of the state, make the obstacle intention prediction more accurate, and improve decision reliability.

[0006] In a first aspect, an embodiment of the present disclosure provides a vehicle interactive decision-making method, the method comprising:

[0007] Determining a plurality of lateral behavior semantics in a pre-determined trajectory of a current vehicle and obtaining a predicted trajectory of an obstacle corresponding to the current vehicle, the lateral behavior semantics being used to describe the lateral behavior of the current vehicle;

[0008] Determining, based on the initial state information, multiple target longitudinal behavior solutions under a first lateral behavior semantics, determining, based on each target longitudinal behavior solution, each predicted state information under the first lateral behavior semantics, and determining, based on each predicted state information, multiple target longitudinal behavior solutions under a next lateral behavior semantics, until each predicted state information under a final lateral behavior semantics is obtained, the target longitudinal behavior solution including the longitudinal acceleration of the current vehicle and the longitudinal acceleration of the obstacle;

[0009] Determining, based on the state benefits of each predicted state information under all lateral behavior semantics, a first longitudinal acceleration sequence of the current vehicle in the pre-decision trajectory, and a decision trajectory of the obstacle and a corresponding second longitudinal acceleration sequence;

[0010] The state information is used to describe the speed and position of the current vehicle and the obstacle, the initial state information is the state information at the current moment, and the predicted state information is the state information at the end moment in the corresponding lateral behavior semantics.

[0011] In a second aspect, an embodiment of the present disclosure further provides a vehicle interactive decision-making device, the device comprising:

[0012] an acquisition module, configured to determine a plurality of lateral behavior semantics in a pre-determined trajectory of a current vehicle and acquire a predicted trajectory of an obstacle corresponding to the current vehicle, the lateral behavior semantics being used to describe the lateral behavior of the current vehicle;

[0013] a state determination module, configured to determine, based on the initial state information, a plurality of target longitudinal behavior solutions under a first lateral behavior semantics, determine, based on each target longitudinal behavior solution, respective predicted state information under the first lateral behavior semantics, and determine, based on each predicted state information, a plurality of target longitudinal behavior solutions under a next lateral behavior semantics, until respective predicted state information under a final lateral behavior semantics is obtained, the target longitudinal behavior solutions including the longitudinal acceleration of the current vehicle and the longitudinal acceleration of the obstacle;

[0014] a decision module, configured to determine, based on state benefits of each predicted state information under all lateral behavior semantics, a first longitudinal acceleration sequence of the current vehicle in the pre-decision trajectory, and a decision trajectory of the obstacle and a corresponding second longitudinal acceleration sequence;

[0015] The state information is used to describe the speed and position of the current vehicle and the obstacle, the initial state information is the state information at the current moment, and the predicted state information is the state information at the end moment in the corresponding lateral behavior semantics.

[0016] In a third aspect, an embodiment of the present disclosure also provides an electronic device, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the interactive decision-making method for the vehicle as described above.

[0017] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored, which, when executed by a processor, implements the interactive decision-making method for a vehicle as described above.

[0018] An interactive vehicle decision-making method provided by the present disclosure determines multiple lateral behavior semantics in a pre-decision trajectory of the current vehicle and obtains a predicted trajectory of each obstacle. Multiple target longitudinal behavior solutions are then determined under the first lateral behavior semantic based on initial state information. Predicted state information under each target longitudinal behavior solution is then determined. Multiple target longitudinal behavior solutions under the next lateral behavior semantic are then determined using each predicted state information until the predicted state information under the last lateral behavior semantic is obtained. Finally, based on the state benefits of each predicted state information under all lateral behavior semantics, a first longitudinal acceleration sequence of the current vehicle in the pre-decision trajectory, as well as the decision trajectory and corresponding second longitudinal acceleration sequence of each obstacle are obtained. This method implements longitudinal behavior decisions for the current vehicle and lateral and longitudinal behaviors of the obstacle. The method continues to predict the next state based on the target longitudinal behavior solution solved for one state, sequentially predicting the next state based on the longitudinal behavior interactively solved for one state. This makes obstacle intention predictions more accurate in each state, solves the problem of low decision accuracy caused by changes in vehicle behavior and obstacle intentions in different states, and improves decision reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0020] Figure 1 is a flow chart of an interactive decision-making method for a vehicle in an embodiment of the present disclosure;

[0021] Figure 2 A schematic diagram of lateral behavior semantics of a pre-decision trajectory in an embodiment of the present disclosure;

[0022] Figure 3 This is a schematic diagram of a key obstacle screening in an embodiment of the present disclosure;

[0023] Figure 4 This is a driving scenario in an embodiment of the present disclosure;

[0024] Figure 5 This is another driving scenario in the embodiment of the present disclosure;

[0025] Figure 6 This is a lane borrowing scenario in an embodiment of the present disclosure;

[0026] Figure 7 This is a lane keeping scenario in an embodiment of the present disclosure;

[0027] Figure 8 is a schematic diagram of a profit matrix in an embodiment of the present disclosure;

[0028] Figure 9 A schematic diagram of the fusion of a local optimal behavior solution set in an embodiment of the present disclosure;

[0029] Figure 10 A schematic diagram of a decision tree determination process in an embodiment of the present application;

[0030] Figure 11 Schematic diagram of the structure of an interactive decision-making device for a vehicle in an embodiment of the present disclosure;

[0031] Figure 12 Schematic diagram of the structure of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0034] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0035] Figure 1 This is a flow chart of an interactive decision-making method for a vehicle in an embodiment of the present disclosure. The method can be executed by an interactive decision-making device of a vehicle, which can be implemented in software and / or hardware. The device can be configured in an electronic device, such as being integrated into a decision-making module of the vehicle. Figure 1 As shown, the method may specifically include the following steps:

[0036] S110 , determining a plurality of lateral behavior semantics in a pre-determined trajectory of the current vehicle, and obtaining a predicted trajectory of an obstacle corresponding to the current vehicle, wherein the lateral behavior semantics are used to describe the lateral behavior of the current vehicle.

[0037] The current vehicle can be understood as the ego vehicle. The pre-decision trajectory can be the predicted driving trajectory output by the pre-decision module of the current vehicle. It should be noted that there can be one or more pre-decision trajectories. For example, the pre-decision module can output the pre-decision trajectory of the current vehicle based on the predicted trajectories of each obstacle provided by the prediction module.

[0038] Specifically, multiple lateral behavior semantics can be divided in the pre-decision trajectory according to time, where each lateral behavior semantic is used to describe the lateral behavior of the current vehicle at a corresponding time.

[0039] For example, taking the pre-decision trajectory of the lane change type as an example, the multiple lateral behavior semantics are arranged in chronological order as follows: lane keeping, lane change, lane keeping, or, arranged in chronological order as follows: lane change, lane keeping.

[0040] Figure 2 is a semantic diagram of a lateral behavior of a pre-decision trajectory in an embodiment of the present disclosure, such as Figure 2 As shown in Figure 1, the lateral behavior semantics of the current vehicle in the pre-decision trajectory are divided into: lane keeping, lane change, and lane keeping.

[0041] It should be noted that in the disclosed embodiments, the purpose of determining multiple lateral behavior semantics for the current vehicle is to assign corresponding lateral semantic information to the pre-decision trajectory, thereby facilitating subsequent interaction with each obstacle and sampling longitudinal velocity based on this lateral semantic information. This allows for sampling only longitudinal behavior, thus eliminating the need for lateral behavior sampling. Assuming that within a lateral behavior semantic, the lateral and longitudinal behaviors of the current vehicle and each obstacle remain unchanged, this ensures compliance with the principle of only allowing one longitudinal behavior sample per lateral behavior semantic, while also avoiding reliance on trajectory planning in decision-making.

[0042] In the disclosed embodiment, the obstacle corresponding to the current vehicle may be an obstacle in the current lane of the current vehicle, or an obstacle in the lane to which the current vehicle is heading, or an obstacle predicted to reach the lane to which the current vehicle is heading in the future, etc. The number of obstacles is at least one.

[0043] Specifically, the predicted trajectory of each obstacle can be provided by the current vehicle's prediction module. For example, the prediction module can predict the predicted trajectory of each obstacle based on information such as the obstacle's current speed and position. For each obstacle, there can be one or more predicted trajectories.

[0044] Considering that the method provided in this embodiment is time-consuming when there are dense obstacles, and has high requirements on the performance of the system implementing the method, therefore, in the embodiment of the present disclosure, obstacles can also be screened purposefully, and the predicted trajectories of the obstacles can be screened to improve the efficiency of interactive decision-making.

[0045] For example, key obstacles can be identified based on the location of each obstacle, and then the predicted trajectory with the highest prediction probability can be obtained for each key obstacle. Key obstacles can be obstacles within the corresponding range of the current vehicle. The faster the current vehicle's speed, the wider the forward range of key obstacles is. Obstacles outside the road network, certain types of obstacles (such as birds and dust), and static obstacles behind the current vehicle are not considered key obstacles.

[0046] In a specific embodiment, obtaining a predicted trajectory of an obstacle corresponding to the current vehicle includes the following steps:

[0047] Step 111: In the current lane of the current vehicle, the closest obstacle in front and behind the current vehicle are considered as key obstacles.

[0048] Step 112: If the pre-decision trajectory includes a following trajectory segment, the left and right obstacles closest to the current vehicle are considered key obstacles. If the pre-decision trajectory includes a lane-changing trajectory segment, the front and rear obstacles closest to the current vehicle in the target lane corresponding to the lane-changing trajectory segment are considered key obstacles.

[0049] Step 113: For each key obstacle, sort the predicted trajectories corresponding to the key obstacle in descending order of predicted probability, and select the top M predicted trajectories from the sorted results;

[0050] Step 114 : Determine obstacles other than the key obstacles as non-key obstacles. For each non-key obstacle, perform collision detection on all predicted trajectories of the non-key obstacle to obtain a portion or all of the predicted trajectories where collision occurs.

[0051] Specifically, for the screening of key obstacles, a certain range can be selected along the front, back, left, and right directions with the current position of the current vehicle as the center, such as Figure 3 As shown, Figure 3 This is a schematic diagram of a key obstacle screening in an embodiment of the present disclosure; further, key obstacles are screened according to the selected range.

[0052] For the screening rules of key obstacles in the front and rear directions, the front obstacles and rear obstacles closest to the current vehicle in the current lane of the current vehicle can be used as key obstacles; for the screening rules of key obstacles in the left and right directions, the screening rules in different scenarios can be different. For example, in the following scenario, that is, when the pre-decision trajectory includes the following trajectory segment, the left obstacles and right obstacles closest to the current vehicle can be used as key obstacles. In the lane changing scenario, that is, when the pre-decision trajectory includes the lane changing trajectory segment, in the target lane that the current vehicle expects to switch to, the front obstacles and rear obstacles closest to the current vehicle can be used as key obstacles.

[0053] After the key obstacles are screened through steps 111 and 112, the predicted trajectories of the key obstacles can be further screened. Considering that the predicted trajectories of some key obstacles may not collide with the current vehicle, screening the predicted trajectories based on whether or not they will collide will result in low accuracy in the interaction prediction between the current vehicle and the key obstacles.

[0054] For example, Figure 4 This is a driving scenario in an embodiment of the present disclosure, in which the current vehicle is traveling in a straight line in the current lane. Vehicle A in the adjacent lane will not collide with the current vehicle, but will affect the safety of the current vehicle while traveling parallel to the current vehicle. Figure 5 This is another driving scenario in the embodiment of the present disclosure, in which the current vehicle expects to switch to the right lane and does not collide with vehicle B during the lane change, but the behavior of vehicle B will affect the safety of the current vehicle.

[0055] Therefore, in the disclosed embodiment, the predicted trajectories can be sorted by predicted probability to obtain the top M predicted trajectories for each key obstacle. The predicted probability can be the driving probability corresponding to the predicted trajectory. For example, M can be 2, and the disclosed embodiment does not restrict the value of M.

[0056] In addition, for non-critical obstacles other than critical obstacles, the predicted trajectories that do not collide with the current vehicle can be discarded, and all or part of the predicted trajectories that collide with the current vehicle can be obtained.

[0057] In one example, obtaining a portion of or all of the predicted trajectory in which a collision occurs includes:

[0058] All predicted trajectories in which a collision occurs are obtained; or, based on the predicted probabilities of all predicted trajectories in which a collision occurs, a portion of the predicted trajectories in which a collision occurs is obtained; or, based on the collision positions of all predicted trajectories in which a collision occurs, a portion of the predicted trajectories in which a collision occurs is obtained.

[0059] That is, all predicted trajectories corresponding to non-critical obstacles that collide with the current vehicle can be obtained. Alternatively, some predicted trajectories can be ignored according to the predicted probability or collision location to ensure that at least one predicted trajectory of a non-critical obstacle is obtained.

[0060] For example, the predicted trajectory with the highest prediction probability may be obtained from all predicted trajectories where a collision may occur, or the predicted trajectory where the collision position is in front of the current vehicle may be obtained.

[0061] Through the above example, it is possible to screen the trajectories of non-critical obstacles, further consider the impact of each obstacle on the safety of the current vehicle, and improve the accuracy of the interaction prediction between the current vehicle and the obstacle.

[0062] The above steps 111 to 114 implement the screening of key obstacles. By using different prediction trajectory acquisition methods for key obstacles and non-key obstacles, various predicted trajectories that may affect the safety of the current vehicle are obtained as comprehensively as possible, ensuring the accuracy of the interactive prediction.

[0063] It should be noted that the predicted trajectory of an obstacle can be used to describe the obstacle's lateral behavior. Specifically, after obtaining the predicted trajectory of each obstacle, if the lateral behavior of each obstacle cannot be determined, for example, due to multiple predicted trajectories, the predicted state information under all lateral behavior semantics can be determined for each obstacle lateral behavior. The final longitudinal acceleration of the current vehicle and obstacle under each obstacle lateral behavior can then be determined from all obstacle lateral behaviors.

[0064] For example, a trajectory set can be determined based on all predicted trajectories of all obstacles. Then, for each trajectory set, the predicted state information for all lateral behavior semantics can be determined. For example, if the predicted trajectory of obstacle 1 includes trajectory A and trajectory B, and the predicted trajectory of obstacle 2 includes trajectory C and trajectory D, then the trajectory sets AC, BC, AD, and BD can be obtained.

[0065] S120. Determine multiple target longitudinal behavior solutions under the first lateral behavior semantics based on the initial state information, determine each predicted state information under the first lateral behavior semantics based on each target longitudinal behavior solution, and determine multiple target longitudinal behavior solutions under the next lateral behavior semantics based on each predicted state information, until each predicted state information under the last lateral behavior semantics is obtained, the target longitudinal behavior solution includes the longitudinal acceleration of the current vehicle and the longitudinal acceleration of the obstacle.

[0066] The state information describes the current speed and position of the vehicle and each obstacle. The initial state information is the current state information, and the predicted state information is the state information at the end time of the corresponding lateral behavior semantics. It should be understood that the state in this embodiment refers to the current state of the vehicle and the obstacle.

[0067] In the disclosed embodiment, starting from the first horizontal behavior semantics, multiple target vertical behavior solutions under the first horizontal behavior semantics can be determined based on the initial state information. Furthermore, for each target vertical behavior solution under the first horizontal behavior semantics, corresponding predicted state information can be determined.

[0068] The initial state information is the current speed and position of the vehicle and each obstacle. The predicted state information under the first lateral behavior semantics is the speed and position of the vehicle and each obstacle at the end of the first lateral behavior semantics. The target longitudinal behavior solution is the final solved longitudinal behavior of the vehicle and obstacles under the corresponding lateral behavior semantics, including the longitudinal acceleration of the vehicle and each obstacle.

[0069] In a specific embodiment, determining multiple target vertical behavior solutions under the first horizontal behavior semantics based on the initial state information includes the following steps:

[0070] Step 121: Determine the local optimal behavior solution set corresponding to each obstacle based on the initial state information and the interaction model corresponding to each obstacle;

[0071] Step 122: fuse all local optimal behavior solution sets to obtain multiple target vertical behavior solutions under the first horizontal behavior semantics.

[0072] Considering that the vehicle's intentions are constantly changing over time, different interaction models can be configured for different driving scenarios to more accurately describe interaction issues. Specifically, corresponding interaction models can be pre-configured for actions in different driving scenarios. Furthermore, the vehicle's driving scenario can be determined based on its pre-decision trajectory. Based on the lateral behavior semantics in the pre-decision trajectory, the interaction model for the action corresponding to that driving scenario can be invoked.

[0073] The driving scenario may be a lane keeping scenario, a lane borrowing scenario, a left turn scenario, a right turn scenario, or a lane changing scenario. For example, Figure 6 This is a lane borrowing scenario in an embodiment of the present disclosure, in which the execution actions corresponding to the semantics of each lateral behavior are: approaching the target lane, borrowing the lane, and approaching the initial lane. Figure 7This is a lane keeping scenario in an embodiment of the present disclosure, wherein the execution actions corresponding to the semantics of each lateral behavior are: go straight, go straight, and go straight.

[0074] In the disclosed embodiment, each obstacle may correspond to an interaction model. The interaction model may be used to predict the benefits under various interaction situations between the current vehicle and the obstacle. Specifically, the interaction model may be used to calculate the benefits under different longitudinal accelerations of the current vehicle and the obstacle, thereby obtaining a benefit matrix between the current vehicle and the obstacle.

[0075] Specifically, each obstacle interaction model can output a payoff matrix, from which a local optimal behavior solution set can be derived. The local optimal behavior solution set can include at least one local optimal behavior solution, each of which has a different longitudinal acceleration for the current vehicle.

[0076] It should be noted that if there's no interaction between the obstacle and the current vehicle, the Intelligent Driver Model (IDM) can intervene. For example, in a following vehicle scenario, where the current vehicle can typically only perform reactive maneuvers, to maintain the intent of the interaction model, intervention is only made when the longitudinal acceleration sampled by the interaction model differs from the IDM's longitudinal acceleration by more than a threshold. If there's no interactive obstacle, intervention is normally enabled at a constant vehicle speed.

[0077] With respect to the above step 121, optionally, determining the local optimal behavior solution set corresponding to each obstacle based on the initial state information and the interaction model corresponding to each obstacle includes the following steps:

[0078] Step 1210: For each obstacle, based on the initial state information, call the interaction model corresponding to the obstacle to calculate the benefit matrix between the first sampled accelerations of the current vehicle and the second sampled accelerations of the obstacle.

[0079] Step 1211: Determine the local optimal behavior solution set corresponding to the obstacle based on the profit matrix.

[0080] Specifically, at least one first sampled acceleration of the current vehicle and a second sampled acceleration of the obstacle can be determined first, wherein the first sampled acceleration is a value obtained by sampling the longitudinal acceleration of the current vehicle, and the second sampled acceleration is a value obtained by sampling the longitudinal acceleration of the obstacle. In the embodiment of the present disclosure, the acceleration can be 0.2 m / s. 2 As the resolution of the current vehicle's longitudinal behavior, different interaction models are used to explore obstacle behavior, that is, 0.2m / s2 The longitudinal acceleration is sampled as the step size. To prevent sampling explosion, the longitudinal acceleration of the current vehicle is fixed within a lateral behavior, that is, within a lateral behavior semantics, the longitudinal acceleration of the current vehicle at each moment remains consistent.

[0081] Given the importance of interaction for decision-making, mispredicting an obstacle's intention can lead to irrational decisions by the vehicle, potentially creating danger and confusing the obstacle. Therefore, we can combine the current speed and position of the vehicle and the obstacle, invoke an interaction model (such as a game model) corresponding to each obstacle, and derive the payoff matrix between each first-sampled acceleration and each second-sampled acceleration.

[0082] For example, Figure 8 Schematic diagram of a benefit matrix in an embodiment of the present disclosure. Each obstacle can correspond to a benefit matrix. Figure 8 Taking two obstacles as an example, the payoff matrices corresponding to obstacle 1 and obstacle 2 are shown. The payoff matrix includes each behavioral solution and the corresponding payoff.

[0083] See also Figure 8 , in the payoff matrix corresponding to obstacle 1, when the first sampling acceleration of the current vehicle is -2m / s 2 , the second sampled acceleration of the obstacle is -2m / s 2 , that is, when the behavior solution is (-2,-2), the corresponding benefit is (-2,1), where -2 represents the current vehicle's benefit, 1 represents the obstacle benefit, and over (overtaking) represents the current vehicle's decision-making behavior. Benefits can be composed of early stopping costs, safety costs, efficiency costs, and consistency costs.

[0084] After obtaining the payoff matrix corresponding to each obstacle, the set of locally optimal behavioral solutions corresponding to the obstacle can be determined based on the payoff matrix. For example, for each behavioral solution with the same first sampled acceleration in the payoff matrix, the one with the highest payoff and the decision behavior of overtaking can be identified as the locally optimal behavioral solution. If the decision behavior for the behavior with the highest payoff is not overtaking, then there is no locally optimal behavioral solution under that first sampled acceleration.

[0085] by Figure 8 The payoff matrix of obstacle 1 is shown as an example, where the first sample acceleration is -2 m / s 2 Among the various behavior solutions, the behavior solution (-2,2) has the best benefit. However, since its corresponding decision behavior is follow (yield), there is no local optimal behavior solution under the first sampling acceleration. 2Among the various behavior solutions, the behavior solutions (-1,1) and (-1,2) have the best benefits and can be determined as local optimal behavior solutions. 2 Among all the behavior solutions, the behavior solution (0, -2) has the best benefit and can be determined as the local optimal behavior solution. 2 Among the various behavior solutions, the behavior solution (1,1) has the best benefit and can be determined as the local optimal behavior solution. 2 Among all the behavioral solutions, the behavioral solution (2,2) has the best benefit and can be determined as the local optimal behavioral solution.

[0086] Furthermore, after determining each local optimal behavior solution from the payoff matrix, all local optimal behavior solutions can constitute the corresponding local optimal behavior solution set. Figure 9 As shown, Figure 9 : is a fusion diagram of a local optimal behavior solution set in an embodiment of the present disclosure, which shows the local optimal behavior solution sets corresponding to obstacle 1 and obstacle 2 respectively.

[0087] Through steps 1210 to 1212 above, the local optimal behavior solution set corresponding to each obstacle is accurately determined. By sampling the longitudinal acceleration for each longitudinal speed change trend and calling the interaction model of each obstacle to perform benefit analysis, the comprehensiveness of the interaction prediction is ensured.

[0088] Since the interaction model corresponding to each obstacle can modify the longitudinal behavior of the obstacle, in order to select the optimal longitudinal behavior under each lateral behavior semantics, after obtaining the local optimal behavior solution set corresponding to each obstacle, all local optimal behavior solution sets can be fused to obtain multiple target longitudinal behavior solutions under the first lateral behavior semantics.

[0089] Optionally, the local optimal behavior solution sets are fused to obtain multiple target vertical behavior solutions under the first horizontal behavior semantics, including:

[0090] For each local optimal behavior solution set corresponding to an obstacle, determine the behavior solution with the same first sampled acceleration as other local optimal behavior solution sets to obtain a solution intersection. Based on the behavior solutions in all solution intersections, determine multiple target longitudinal behavior solutions. Based on the individual benefits of the behavior solutions in each solution intersection, determine the cumulative benefits of each target longitudinal behavior solution. The individual benefit is the benefit of the current vehicle interacting with a single obstacle, and the cumulative benefit is the benefit of the current vehicle interacting with all obstacles.

[0091] Correspondingly, each predicted state information under the first horizontal behavior semantics is determined based on each target vertical behavior solution, including: for each target vertical behavior solution, the corresponding predicted state information is determined based on the target vertical behavior solution and the initial state information, and the state benefit of the predicted state information is determined according to the cumulative benefit of the target vertical behavior solution.

[0092] Specifically, the method can first identify behavioral solutions with the same first sampled acceleration as other locally optimal behavioral solutions in the set of locally optimal solutions, thereby obtaining a solution intersection. Furthermore, multiple target longitudinal behavioral solutions can be determined based on the behavioral solutions in each solution intersection. For example, the behavioral solutions with the same first sampled acceleration in the solution intersection can be fused to obtain the target longitudinal behavioral solution. Alternatively, the fused solutions can be sorted according to the current vehicle's benefit, and some of the behavioral solutions selected based on the sorting results can be used as the target longitudinal behavioral solution.

[0093] Furthermore, for the solutions of each behavior with the same first sampled acceleration in different solution intersections, the individual benefits of the current vehicle can be added together to obtain the cumulative benefit of the current vehicle.

[0094] like Figure 9 As shown, the intersection of the solutions with the same first sampled acceleration can be selected from the local optimal behavior solution set for obstacle 1 and the local optimal behavior solution set for obstacle 2. Based on the intersection of the two solutions, the target longitudinal behavior solutions (0, -2, -2), (1, 1, 1), and (2, 2, 2) can be obtained. The individual benefits of the current vehicle corresponding to the behavior solutions with the same first sampled acceleration are summed together to obtain the cumulative benefits of 6, 14, and 16, respectively, for the three target longitudinal behavior solutions. Of course, during the benefit fusion process, a corresponding weight can also be set for each obstacle based on the driving scenario. For example, obstacles that interact with the current vehicle have a higher weight. The cumulative benefit can then be obtained by combining the weights and individual benefits.

[0095] For example, three target longitudinal behavior solutions may be determined, and the first sampled accelerations in the three target longitudinal behavior solutions may represent the acceleration trend, uniform velocity trend, and deceleration trend, respectively, under the first lateral behavior semantics. For example, the smallest first sampled acceleration represents a deceleration trend, the middle first sampled acceleration represents a uniform velocity trend, and the largest first sampled acceleration represents an acceleration trend.

[0096] In the process of further determining the predicted state information under the first lateral behavior semantics based on the target longitudinal behavior solution, for each target longitudinal behavior solution, the position and speed of the current vehicle and obstacle at the end time of the first lateral behavior semantics can be deduced from the speed and position of the current vehicle and obstacle described in the initial state information, as well as the longitudinal acceleration of the current vehicle and obstacle described in the target longitudinal behavior solution, thereby obtaining the predicted state information under the first lateral behavior semantics. Furthermore, the state benefit of the predicted state information can be determined based on this accumulated benefit.

[0097] The state payoff can be the interaction payoff between the current vehicle and the obstacle when the corresponding predicted state is met. After deriving the predicted state information for all lateral behavior semantics (i.e., obtaining a decision tree), the state payoff can be used to search the decision tree for the optimal longitudinal behavior of the current vehicle and the obstacle, as well as the optimal lateral behavior of the obstacle. By determining the cumulative payoff of the target longitudinal behavior solution, the current vehicle's payoff is integrated and used as the state payoff to ensure the accuracy of the optimal longitudinal and lateral behaviors determined during the search phase.

[0098] In the above embodiment, by calling the local optimal behavior solution sets corresponding to different obstacles and fusing the local optimal behavior solution sets for each obstacle, it is avoided to fall into the local optimum when predicting the interaction between the current vehicle and the obstacle, and the accuracy of the decision is further guaranteed.

[0099] In the embodiment of the present disclosure, after obtaining the various predicted state information under the first horizontal behavior semantics, multiple target vertical behavior solutions under the next horizontal behavior semantics can be determined based on the predicted state information for each predicted state information under the first horizontal behavior semantics, and each predicted state information can be obtained based on each target vertical behavior solution. This process is repeated until the various predicted state information under the last horizontal behavior semantics is obtained, completing the derivation of the decision tree, that is, the various predicted state information under all horizontal behavior semantics constitute a decision tree.

[0100] For example, Figure 10 FIG. 1 is a schematic diagram of a decision tree determination process in an embodiment of the present application. Figure 10As shown, the initial state can be taken as State1, and the three target longitudinal behavior solutions under the lateral behavior semantics 1 can be obtained, which are represented by ACC (acceleration trend), MAIN (uniform speed trend), and DEC (deceleration trend). Then, through each target longitudinal behavior solution, combined with State1, State2, State3, and State4 under the lateral behavior semantics 1 can be derived respectively. Furthermore, taking State2 as an example, the three target longitudinal behavior solutions under the lateral behavior semantics 2 can be obtained. Then, through each target longitudinal behavior solution, combined with State2, State5, State6, and State7 under the lateral behavior semantics 2 can be derived respectively. Similarly, State3 and State4 can also be derived from the three predicted states under the lateral behavior semantics 2 by solving the target longitudinal behavior solution. Furthermore, taking State6 as an example, the three target longitudinal behavior solutions under the lateral behavior semantics 3 can be obtained. Then, repeating the above steps, each predicted state under the lateral behavior semantics 3 can be obtained, thereby completing the construction of the decision tree.

[0101] It should be noted that in the process of determining the decision tree, after completing the derivation of multiple prediction states under a horizontal behavior semantics, it is necessary to return and update the state benefits of other prediction states on the same branch as each prediction state through the derived state benefits of each prediction state. Figure 10 Taking State6 in the example, the state benefit of State6 can be added to the state benefit of State2, and the state benefit of State1 can be updated.

[0102] In summary, in the embodiment of the present disclosure, we can start from the first horizontal behavior semantics, obtain the target vertical behavior solution under the horizontal behavior semantics by sampling and calling the interaction model, and then obtain the predicted state information under the horizontal behavior semantics by deduction, and repeat this process until the predicted state information under the last horizontal behavior semantics is obtained.

[0103] Considering that the sampling process may overlook the potential response of the following vehicle after the current vehicle takes action (although this response may not necessarily occur), this may affect the prediction result and further affect the next action of the current vehicle. This continuous interaction process in time series cannot be handled during the current sampling process because the prediction information received by the decision module cannot take into account the future action of the current vehicle.

[0104] Therefore, before deriving the predicted state information, a single-step simulation can be performed. A single-step simulation refers to simulating from one state to the next. Taking the first lateral behavior semantics as an example, after obtaining the target longitudinal behavior solution for this lateral behavior semantics, simulation can be performed to deduce the target longitudinal behavior solution. After completing the simulation, the predicted state information can be derived to address the issue of ignoring the potential response of obstacles during the sampling process.

[0105] In a specific embodiment, before determining each predicted state information under the first lateral behavior semantics based on each target longitudinal behavior solution, it also includes: for each target longitudinal behavior solution, simulating the pre-decision trajectory based on the target longitudinal behavior solution, and judging whether there is a collision in the simulation result of the pre-decision trajectory; if so, determining the collision responsible object in the current vehicle and each obstacle, and updating the longitudinal acceleration of the collision responsible object in the target longitudinal behavior solution.

[0106] Specifically, to ensure the accuracy of interactions between multiple obstacles and the current vehicle and prevent collisions or vehicles with obstacles from crossing each other, the pre-determined trajectory is simulated according to the predicted intent, i.e., the target longitudinal behavior solution. If a collision is detected in the simulation results, the responsible party is determined, and the longitudinal acceleration of the responsible party in the target longitudinal behavior solution is backtracked to update the longitudinal acceleration. If a collision is detected in the simulation results, and the collision specifically involves the rear end of the current vehicle, this collision can be ignored to account for the possibility of the following vehicle yielding.

[0107] By simulating the target longitudinal behavior solution, simulation deduction from one state to another can be achieved to further consider the rationality of the prediction intention of the interaction model and ensure the accuracy of the interaction decision.

[0108] S130 , determining a first longitudinal acceleration sequence of the current vehicle in the pre-decision trajectory, and a decision trajectory of the obstacle and a corresponding second longitudinal acceleration sequence based on the state benefits of each predicted state information under all lateral behavior semantics.

[0109] Specifically, for each trajectory set, after obtaining the predicted state information under all lateral behavior semantics, we can start from the first lateral behavior semantic and select the target longitudinal behavior solution corresponding to the predicted state information with the best state benefit. Then, in the next lateral behavior semantic, from the predicted state information derived from the target longitudinal behavior solution selected by the first lateral behavior semantic, select the target longitudinal behavior solution corresponding to the predicted state information with the best state benefit. Repeat this process until the optimal target longitudinal behavior solution is selected under the last lateral behavior semantic, thereby obtaining the first longitudinal acceleration sequence of the current vehicle and the second longitudinal acceleration sequence of the obstacle under the trajectory set.

[0110] Furthermore, a target set can be selected from all trajectory sets based on the state benefits of each predicted state information in all trajectory sets, and the decision trajectory of the obstacle can be obtained based on the target set to determine the lateral behavior of the obstacle, and output the first longitudinal acceleration sequence and the second longitudinal acceleration sequence under the target set.

[0111] For example, the decision module may output the decision trajectory of the obstacle, the first longitudinal acceleration sequence, and the second longitudinal acceleration sequence to the planning module of the current vehicle.

[0112] It is understandable that if the current vehicle has multiple pre-decision trajectories, the above S110-S130 can be executed separately for each pre-decision trajectory to obtain the first longitudinal acceleration sequence of the current vehicle, the decision trajectory of the obstacle, and the second longitudinal acceleration sequence of the obstacle in each pre-decision trajectory.

[0113] In the disclosed embodiment, in order to reflect the real-time nature of the interaction, an interaction model is used to predict the intention of the obstacle, and the solution of the intention model is introduced as a heuristic into the architecture of the decision tree. A simulator is used as a transition between states. The simulator will consider the interaction between obstacles through single-step simulation deduction and follow the strong traffic rules. In order to explore the diversity of ego vehicle behavior, the expansion of the decision tree will select the ego vehicle with the best benefit and its corresponding obstacle behavior from the three ego vehicle longitudinal semantics of {acceleration, constant speed, deceleration} in each ego vehicle's lateral behavior semantics, and update them to solve the problems of decision tree dimension and history explosion.

[0114] Among them, by attaching semantic segments to the pre-decision trajectory of the current vehicle, the intention of the obstacle is inferred and backtracked in each semantic segment, and the various obstacle intentions are simplified into the longitudinal optimal semantic behaviors of acceleration, constant speed and deceleration through fusion, realizing a three-dimensional decision-making that integrates time, horizontal and vertical directions. Through the Monte Carlo tree (decision tree) and deduction method, the interaction information under different states can be obtained more accurately, solving the problem of inaccurate decision-making of autonomous driving vehicles due to the change of the semantic behavior of the vehicle due to obstacles while the intention remains unchanged.

[0115] When transitioning between different states, a single-step deduction method is employed, along with processing of traffic rules and topological relationships, to ensure the rationality and real-time nature of algorithmic interaction. By integrating multiple interaction models, the problem of a single interaction model being inadequate for all scenarios, as well as the inability to integrate and verify multiple interaction models during interaction with obstacles, is resolved.

[0116] The interactive vehicle decision-making method provided in this embodiment determines multiple lateral behavior semantics in the pre-decision trajectory of the current vehicle and obtains the predicted trajectory of each obstacle. Based on the initial state information, multiple target longitudinal behavior solutions are determined for the first lateral behavior semantic. Based on each target longitudinal behavior solution, each predicted state information for the first lateral behavior semantic is determined. Multiple target longitudinal behavior solutions for the next lateral behavior semantic are determined using each predicted state information until each predicted state information for the last lateral behavior semantic is obtained. Finally, based on the state benefits of each predicted state information for all lateral behavior semantics, the first longitudinal acceleration sequence of the current vehicle in the pre-decision trajectory, as well as the decision trajectory and corresponding second longitudinal acceleration sequence for each obstacle are obtained. This method implements longitudinal behavior decisions for the current vehicle and the lateral and longitudinal behaviors of the obstacles. Based on the target longitudinal behavior solution solved for one state, the method continues to predict the next state, sequentially predicting the next state based on the longitudinal behavior interactively solved for one state. This makes obstacle intention predictions more accurate in each state, addresses the issue of low decision accuracy caused by changes in vehicle behavior and obstacle intentions in different states, and improves decision reliability.

[0117] Figure 11 FIG. 1 is a schematic diagram of the structure of an interactive decision-making device for a vehicle in an embodiment of the present disclosure. Figure 11 As shown: the device includes: an acquisition module 1110, a state determination module 1120 and a decision module 1130.

[0118] an acquisition module 1110 for determining a plurality of lateral behavior semantics in a pre-determined trajectory of a current vehicle and acquiring a predicted trajectory of an obstacle corresponding to the current vehicle, the lateral behavior semantics being used to describe the lateral behavior of the current vehicle;

[0119] a state determination module 1120 for determining, based on the initial state information, multiple target longitudinal behavior solutions for a first lateral behavior semantics, determining, based on each target longitudinal behavior solution, each predicted state information for the first lateral behavior semantics, and determining, based on each predicted state information, multiple target longitudinal behavior solutions for a next lateral behavior semantics, until each predicted state information for a final lateral behavior semantics is obtained, wherein the target longitudinal behavior solution includes the longitudinal acceleration of the current vehicle and the longitudinal acceleration of the obstacle;

[0120] A decision module 1130 is configured to determine a first longitudinal acceleration sequence of the current vehicle in the pre-decision trajectory, and a decision trajectory of the obstacle and a corresponding second longitudinal acceleration sequence based on the state benefits of each predicted state information under all lateral behavior semantics;

[0121] The state information is used to describe the speed and position of the current vehicle and the obstacle, the initial state information is the state information at the current moment, and the predicted state information is the state information at the end moment in the corresponding lateral behavior semantics.

[0122] The interactive decision-making device for a vehicle provided in the embodiment of the present disclosure can execute the steps of the interactive decision-making method for a vehicle provided in the embodiment of the method of the present disclosure. The execution steps and beneficial effects are not repeated here.

[0123] Figure 12 This is a schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 12 , which shows a structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. Figure 12 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0124] like Figure 12 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes to implement the methods of the embodiments described in the present disclosure according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage device 508 into the random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0125] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart, thereby implementing the interactive decision-making method of the vehicle as described above. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0126] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0127] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device:

[0128] Determining a plurality of lateral behavior semantics in a pre-determined trajectory of a current vehicle and obtaining a predicted trajectory of an obstacle corresponding to the current vehicle, the lateral behavior semantics being used to describe the lateral behavior of the current vehicle;

[0129] Determining, based on the initial state information, multiple target longitudinal behavior solutions under a first lateral behavior semantics, determining, based on each target longitudinal behavior solution, each predicted state information under the first lateral behavior semantics, and determining, based on each predicted state information, multiple target longitudinal behavior solutions under a next lateral behavior semantics, until each predicted state information under a final lateral behavior semantics is obtained, the target longitudinal behavior solution including the longitudinal acceleration of the current vehicle and the longitudinal acceleration of the obstacle;

[0130] Determining, based on the state benefits of each predicted state information under all lateral behavior semantics, a first longitudinal acceleration sequence of the current vehicle in the pre-decision trajectory, and a decision trajectory of the obstacle and a corresponding second longitudinal acceleration sequence;

[0131] The state information is used to describe the speed and position of the current vehicle and the obstacle, the initial state information is the state information at the current moment, and the predicted state information is the state information at the end moment in the corresponding lateral behavior semantics.

[0132] Optionally, when the above one or more programs are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.

[0133] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

Claims

1. A vehicle interactive decision-making method, characterized in that: The method comprises: Determining a plurality of lateral behavior semantics in a pre-determined trajectory of a current vehicle and obtaining a predicted trajectory of an obstacle corresponding to the current vehicle, the lateral behavior semantics being used to describe the lateral behavior of the current vehicle; Determining, based on the initial state information, multiple target longitudinal behavior solutions under a first lateral behavior semantics, determining, based on each target longitudinal behavior solution, each predicted state information under the first lateral behavior semantics, and determining, based on each predicted state information, multiple target longitudinal behavior solutions under a next lateral behavior semantics, until each predicted state information under a final lateral behavior semantics is obtained, the target longitudinal behavior solution including the longitudinal acceleration of the current vehicle and the longitudinal acceleration of the obstacle; Determining, based on the state benefits of each predicted state information under all lateral behavior semantics, a first longitudinal acceleration sequence of the current vehicle in the pre-decision trajectory, and a decision trajectory of the obstacle and a corresponding second longitudinal acceleration sequence; The state information is used to describe the speed and position of the current vehicle and the obstacle. The initial state information is the state information at the current moment, and the predicted state information is the state information at the end moment of the corresponding lateral behavior semantics. The step of determining multiple target longitudinal behavior solutions under the first transverse behavior semantics according to the initial state information includes: According to the initial state information and the interaction model corresponding to each obstacle, the local optimal behavior solution set corresponding to each obstacle is determined; All local optimal behavior solution sets are integrated to obtain multiple target vertical behavior solutions under the first horizontal behavior semantics; The above-mentioned fusion of all local optimal behavior solution sets obtains multiple target vertical behavior solutions under the first horizontal behavior semantics, including: For each local optimal behavior solution set corresponding to the obstacle, determine a behavior solution having the same first sampled acceleration as other local optimal behavior solution sets to obtain a solution intersection; Determine multiple target longitudinal behavior solutions based on the behavior solutions in all solution intersections, and determine the cumulative benefits of each target longitudinal behavior solution based on the individual benefits of the behavior solutions in each solution intersection, wherein the individual benefit is the benefit of the current vehicle interacting with a single obstacle, and the cumulative benefit is the benefit of the current vehicle interacting with all obstacles; Accordingly, the step of determining each predicted state information under the first horizontal behavior semantics based on each target vertical behavior solution includes: For each of the target longitudinal behavior solutions, corresponding predicted state information is determined based on the target longitudinal behavior solution and the initial state information, and the state benefit of the predicted state information is determined according to the accumulated benefit of the target longitudinal behavior solution.

2. The method according to claim 1, characterized in that Determining the local optimal behavior solution set corresponding to each obstacle based on the initial state information and the interaction model corresponding to each obstacle includes: For each obstacle, based on the initial state information, calling the interaction model corresponding to the obstacle, and calculating a benefit matrix between each first sampled acceleration of the current vehicle and each second sampled acceleration of the obstacle; A local optimal behavior solution set corresponding to the obstacle is determined based on the benefit matrix.

3. The method according to claim 1, characterized in that The obtaining of a predicted trajectory of an obstacle corresponding to the current vehicle includes: In the current lane of the current vehicle, the front obstacle and the rear obstacle closest to the current vehicle are regarded as key obstacles; If the pre-decision trajectory includes a following trajectory segment, the left and right obstacles closest to the current vehicle are used as key obstacles. If the pre-decision trajectory includes a lane-changing trajectory segment, the front and rear obstacles closest to the current vehicle in the target lane corresponding to the lane-changing trajectory segment are used as key obstacles. For each of the key obstacles, sort the predicted trajectories corresponding to the key obstacle in descending order of predicted probability, and select the top M predicted trajectories from the sorted results; Obstacles other than the key obstacles are determined as non-key obstacles. For each of the non-key obstacles, collision detection is performed on all predicted trajectories of the non-key obstacle to obtain a portion or all of the predicted trajectories where a collision occurs.

4. The method according to claim 3, characterized in that The obtaining of a portion of the predicted trajectory or the entire predicted trajectory in which the collision occurs includes: Get all predicted trajectories in which collisions occur; or, Based on the prediction probabilities of all predicted trajectories that collide, a portion of the predicted trajectory that collides is obtained, or based on the collision positions of all predicted trajectories that collide, a portion of the predicted trajectory that collides is obtained.

5. The method according to claim 1, characterized in that Before determining each predicted state information under the first horizontal behavior semantics based on each target vertical behavior solution, the method further includes: For each target longitudinal behavior solution, simulating the pre-decision trajectory based on the target longitudinal behavior solution, and determining whether a collision condition exists in the simulation result of the pre-decision trajectory; If so, a collision-responsible object is determined among the current vehicle and each obstacle, and the longitudinal acceleration of the collision-responsible object in the target longitudinal behavior solution is updated.

6. An interactive decision-making device for a vehicle, characterized in that: include: an acquisition module, configured to determine a plurality of lateral behavior semantics in a pre-determined trajectory of a current vehicle and acquire a predicted trajectory of an obstacle corresponding to the current vehicle, the lateral behavior semantics being used to describe the lateral behavior of the current vehicle; a state determination module, configured to determine, based on the initial state information, a plurality of target longitudinal behavior solutions under a first lateral behavior semantics, determine, based on each target longitudinal behavior solution, respective predicted state information under the first lateral behavior semantics, and determine, based on each predicted state information, a plurality of target longitudinal behavior solutions under a next lateral behavior semantics, until respective predicted state information under a final lateral behavior semantics is obtained, the target longitudinal behavior solutions including the longitudinal acceleration of the current vehicle and the longitudinal acceleration of the obstacle; a decision module, configured to determine, based on state benefits of each predicted state information under all lateral behavior semantics, a first longitudinal acceleration sequence of the current vehicle in the pre-decision trajectory, and a decision trajectory of the obstacle and a corresponding second longitudinal acceleration sequence; The state information is used to describe the speed and position of the current vehicle and the obstacle. The initial state information is the state information at the current moment, and the predicted state information is the state information at the end moment of the corresponding lateral behavior semantics. The method of determining multiple target longitudinal behavior solutions under the first lateral behavior semantics based on the initial state information includes: According to the initial state information and the interaction model corresponding to each obstacle, the local optimal behavior solution set corresponding to each obstacle is determined; All local optimal behavior solution sets are integrated to obtain multiple target vertical behavior solutions under the first horizontal behavior semantics; The above-mentioned fusion of all local optimal behavior solution sets obtains multiple target vertical behavior solutions under the first horizontal behavior semantics, including: For each local optimal behavior solution set corresponding to the obstacle, determine a behavior solution having the same first sampled acceleration as other local optimal behavior solution sets to obtain a solution intersection; Determine multiple target longitudinal behavior solutions based on the behavior solutions in all solution intersections, and determine the cumulative benefits of each target longitudinal behavior solution based on the individual benefits of the behavior solutions in each solution intersection, wherein the individual benefit is the benefit of the current vehicle interacting with a single obstacle, and the cumulative benefit is the benefit of the current vehicle interacting with all obstacles; Accordingly, the step of determining each predicted state information under the first horizontal behavior semantics based on each target vertical behavior solution includes: For each of the target longitudinal behavior solutions, corresponding predicted state information is determined based on the target longitudinal behavior solution and the initial state information, and the state benefit of the predicted state information is determined according to the accumulated benefit of the target longitudinal behavior solution.

7. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Variable-speed dynamic lane changing track planning method based on vehicle driving rule

    CN111806467A

  • Intelligent decision-making and local trajectory planning method for autonomous vehicle and decision-making system thereof

    CN113386795A