Decision-making method, device and vehicle

By constructing sampling game space and calculating strategy cost methods, bicycles can make reasonable decisions with obstacles in various scenarios, solving the problem of poor scenario generalization ability in the prior art, and improving the adaptability and efficiency of decisions.

CN115246415BActive Publication Date: 2025-08-08YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110454337.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-26
Publication Date
2025-08-08
Estimated Expiration
2041-04-26

AI Technical Summary

Technical Problem

The existing intelligent driving system cannot adapt to multiple scenarios in the decision-making of obstacles, resulting in poor generalization ability of traffic scenarios and inability to effectively deal with different obstacle environments.

Method used

By obtaining the predicted motion trajectories of bicycles and obstacles, a sampling game space is constructed, the strategic cost of each game strategy is calculated, the game strategy with the minimum cost is selected as the decision result, adapt to all scenarios, and simultaneous gameplay is performed when facing multiple obstacles.

Benefits of technology

Reasonable decision-making between bicycles and obstacles in various scenarios is achieved, the adaptability and rationality of decision-making in traffic scenarios is improved, decision-making time is reduced, and user experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115246415B_ABST
    Figure CN115246415B_ABST
Patent Text Reader

Abstract

This application provides a decision-making method, device, and vehicle related to the field of intelligent driving technology. The method comprises obtaining predicted motion trajectories of the vehicle and each obstacle surrounding it, determining whether the predicted motion trajectories intersect or whether the distance between the two vehicles is less than a set threshold, and then determining the game object. Then, a sampling game space is constructed between the vehicle and each obstacle, and the strategy cost of each game strategy in each sampling game space is calculated. The strategy with the lowest strategy cost among the same game strategies is selected as the game result by solving the shared game strategies in the sampling game space for each obstacle. Because this solution is scenario-independent, it is adaptable to all scenarios. Furthermore, during the game process, when faced with multiple game objects, the vehicle can simultaneously play games with multiple game objects by solving the shared game strategies in each sampling game space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving technology, and in particular to a decision-making method, device and vehicle. Background Art

[0002] With the development and widespread adoption of intelligent vehicles, intelligent driving has become a hot research topic. Intelligent driving systems can be divided into four key functional modules based on functional requirements: positioning, environmental perception, path planning, and decision-making and control. Within the decision-making and control module, various manufacturers have proposed decision-making and planning methods for different scenarios. These methods primarily encompass high-level semantic decision-making (such as lane change and lane keeping) and obstacle-related decisions (such as avoidance, following, overtaking, and yielding).

[0003] Existing methods for determining obstacle avoidance by detecting obstacle types and planning vehicle routes are limited to specific scenarios. These methods typically quantitatively describe the specific scenario and then extract key information about key obstacles to make decisions. Consequently, they have poor generalization capabilities across traffic scenarios and are unable to handle obstacle environments in other scenarios. Summary of the Invention

[0004] In order to solve the above problems, embodiments of the present application provide a decision-making method, device and vehicle.

[0005] In a first aspect, the present application provides a decision-making method, including: obtaining the predicted motion trajectory of the ego vehicle and each obstacle around the ego vehicle; determining the game object, which is an obstacle among the obstacles around the ego vehicle that intersects with the predicted motion trajectory of the ego vehicle or whose distance to the ego vehicle is less than a set threshold; constructing a sampling game space for each game object based on the vehicle information of the ego vehicle collected by the sensor system, the obstacle information and road condition information of the game object, each sampling game space includes at least one game strategy; calculating the strategy cost of each game strategy, which is a value obtained by weighting the weights of each factor of the strategy cost; determining the decision result of the ego vehicle, which is the game strategy with the smallest strategy cost in the shared sampling game space, the shared sampling game space includes at least one game strategy, and each sampling game space includes the game strategies in the shared sampling game space.

[0006] In this implementation, the predicted motion trajectories of the ego vehicle and each obstacle around it are obtained. The predicted motion trajectories are then determined to determine whether they intersect or whether the distance between the two vehicles is less than a set threshold. A sampling game space is then constructed between the ego vehicle and each obstacle, and the strategy cost of each game strategy in each sampling game space is calculated. The strategy with the lowest strategy cost among the identical game strategies in each sampling game space is then selected as the game outcome. Because this solution is scenario-independent, it is adaptable to all scenarios. Furthermore, when faced with multiple game targets during the game, the ego vehicle can simultaneously engage in game play with multiple targets by solving the identical game strategy in each sampling game space.

[0007] In one embodiment, determining the decision result of the vehicle includes: constructing a feasible domain of each sampled game space, where the feasible domain of each sampled game space is at least one game strategy corresponding to a strategy cost that meets the set requirements; and determining the game strategy with the smallest strategy cost among the same game strategies in the intersection of the feasible domains of all sampled game spaces.

[0008] In this embodiment, unlike the prior art which directly obtains the result of the optimal game strategy, the present application constructs a feasible domain between the vehicle and each obstacle by outputting each game strategy that meets the requirements, so that the present application can judge the conflict between multiple game objects based on the feasible domain, making the output game result more reasonable.

[0009] In one embodiment, the method further includes: determining a non-game object, wherein the non-game object is an obstacle among various obstacles around the ego vehicle that does not intersect with the predicted motion trajectory of the ego vehicle or is at a distance from the ego vehicle that is not less than a set threshold; constructing a feasible domain of the ego vehicle based on vehicle information of the ego vehicle collected by a sensor system, obstacle information of the non-game object, and road condition information, wherein the feasible domain of the ego vehicle is at least one strategy for the ego vehicle to take different decisions without colliding with the non-game object; and outputting the decision result of the ego vehicle when detecting that the decision result of the ego vehicle is within the feasible domain of the ego vehicle.

[0010] In this embodiment, by constructing the feasible domain between the self-vehicle and non-game objects, the feasible domain between the self-vehicle and each game object and the feasible domain between the self-vehicle and non-game objects are intersected, and the game strategy with the minimum game cost is selected from the intersection as the decision result, thereby ensuring that the selected decision result can be applied to scenarios including game objects and non-game objects.

[0011] In one embodiment, based on the vehicle information of the ego vehicle, the obstacle information and the road condition information of the game object collected by the sensor system, a sampling game space is constructed for each of the game objects, including: determining the decision upper limit and decision lower limit of the ego vehicle and each obstacle in the game object based on the vehicle information of the ego vehicle, the obstacle information and the road condition information of the game object; obtaining the decision strategy of the ego vehicle and each obstacle in the game object from the decision upper limit and decision lower limit of each obstacle in the ego vehicle and the game object according to the set rules; combining the decision strategy of the ego vehicle with the decision strategy of each obstacle in the game object to obtain at least one game strategy for the ego vehicle and each obstacle in the game object.

[0012] In this embodiment, the game strategies of the self-vehicle and each game object are obtained by determining the selection range and selection method of the game strategies of the self-vehicle and each game object, and then the game strategies of the self-vehicle are combined with the game strategies of each game object to obtain a set of game strategies of the self-vehicle and each game object, thereby ensuring the rationality of the game strategies in each sampled game space.

[0013] In one embodiment, the method further includes: determining a behavior label for each game strategy based on the distance between the ego vehicle and the game object and the conflict point, and the at least one game strategy of the ego vehicle and each obstacle in the game object, the conflict point being the position where the predicted motion trajectory of the ego vehicle and the obstacle intersects or the position where the distance between the ego vehicle and the obstacle is less than a set threshold, and the behavior label includes at least one of the ego vehicle giving way, the ego vehicle overtaking, and both the ego vehicle and the obstacle giving way.

[0014] In this embodiment, by labeling each game strategy, after the game result is selected later, the game strategy label can be directly sent to the execution unit of the next layer. There is no need to analyze the game methods adopted by both parties in the game strategy to determine whether the vehicle should give way, overtake the obstacle, or both give way to the obstacle during the game, thereby greatly reducing decision-making time and improving user experience.

[0015] In one embodiment, the calculation of the strategy cost of each game strategy includes: determining the various factors of the strategy cost, the various factors of the strategy cost including safety, comfort, passing efficiency, right of way, prior probability of obstacles and at least one of historical decision correlation; calculating the factor cost of each factor in each strategy cost; weighting the factor cost of each factor in each strategy cost to obtain the strategy cost of each game strategy.

[0016] In this embodiment, when calculating the strategy cost of each game strategy, the cost of each factor can be calculated and then weighted to obtain the cost of each game strategy, so as to determine the rationality of each game strategy.

[0017] In one embodiment, after calculating the strategy cost of each game strategy, the method further includes: comparing whether each factor in the strategy cost is within a set range; and deleting the game strategy corresponding to the strategy cost including any factor that is not within the set range.

[0018] In this embodiment, by deleting unreasonable game strategies, it is avoided that the decision results cannot be executed or are wrong due to the subsequent selection of unreasonable game strategies, thereby reducing the reliability of the decision-making method.

[0019] In one embodiment, the method further includes: detecting that the decision result of the ego vehicle is not within the feasible domain of the ego vehicle, and outputting the decision result of the ego vehicle yielding the right of way.

[0020] In this embodiment, if the output decision result is not within the feasible domain of the vehicle, it indicates that the decision result does not meet the conditions, and the vehicle will not output the decision result. This situation is equivalent to the vehicle not participating in the game process, which has serious defects. Therefore, when the decision result cannot be determined, according to the principle of "safety", "the vehicle gives way" is selected as the decision result, thereby ensuring that the decision result selected by the vehicle can make the vehicle safe during driving.

[0021] In a second aspect, the present application provides a decision-making device, comprising: a transceiver unit for obtaining the predicted motion trajectory of the ego vehicle and each obstacle around the ego vehicle; a processing unit for determining a game object, which is an obstacle among the obstacles around the ego vehicle that intersects with the predicted motion trajectory of the ego vehicle or whose distance from the ego vehicle is less than a set threshold; constructing a sampling game space for each game object based on the vehicle information of the ego vehicle collected by the sensor system, the obstacle information and road condition information of the game object, each sampling game space includes at least one game strategy; calculating the strategy cost of each game strategy, which is a value obtained by weighting the weights of each factor of the strategy cost; and determining the decision result of the ego vehicle, which is the game strategy with the smallest strategy cost in the shared sampling game space, the shared sampling game space includes at least one game strategy, and each sampling game space includes the game strategy in the shared sampling game space.

[0022] In one embodiment, the processing unit is specifically used to construct a feasible domain for each sampled game space, where the feasible domain of each sampled game space is at least one game strategy corresponding to a strategy cost that meets the set requirements; and in the intersection of the feasible domains of all sampled game spaces, determine the game strategy with the smallest strategy cost among the same game strategies.

[0023] In one embodiment, the processing unit is further used to determine a non-game object, which is an obstacle among various obstacles around the ego vehicle that does not intersect with the predicted motion trajectory of the ego vehicle or is at a distance from the ego vehicle that is not less than a set threshold; construct a feasible domain of the ego vehicle based on the vehicle information of the ego vehicle collected by the sensor system, the obstacle information of the non-game object, and the road condition information, where the feasible domain of the ego vehicle is at least one strategy for the ego vehicle to take different decisions without colliding with the non-game object; and output the decision result of the ego vehicle when it is detected that the decision result of the ego vehicle is within the feasible domain of the ego vehicle.

[0024] In one embodiment, the processing unit is specifically used to determine the decision upper limit and decision lower limit of the self-vehicle and each obstacle in the game object based on the vehicle information of the self-vehicle, the obstacle information of the game object and the road condition information; according to the set rules, obtain the decision strategy of the self-vehicle and each obstacle in the game object from the decision upper limit and decision lower limit of each obstacle in the game object; combine the decision strategy of the self-vehicle with the decision strategy of each obstacle in the game object to obtain the at least one game strategy of the self-vehicle and each obstacle in the game object.

[0025] In one embodiment, the processing unit is further used to determine the behavior label of each game strategy based on the distance between the ego vehicle and the game object and the conflict point, and the at least one game strategy of the ego vehicle and each obstacle in the game object. The conflict point is the position where the predicted motion trajectory of the ego vehicle and the obstacle intersects or the position where the distance between the ego vehicle and the obstacle is less than a set threshold. The behavior label includes at least one of the ego vehicle giving way, the ego vehicle overtaking, and both the ego vehicle and the obstacle giving way.

[0026] In one embodiment, the processing unit is specifically used to determine various factors of the strategy cost, and the various factors of the strategy cost include at least one of safety, comfort, passing efficiency, road right, prior probability of obstacles and historical decision correlation; calculate the factor cost of each factor in each strategy cost; and weight the factor cost of each factor in each strategy cost to obtain the strategy cost of each game strategy.

[0027] In one embodiment, the processing unit is further configured to compare whether each factor in the strategy cost is within a set range; and delete the game strategy corresponding to the strategy cost including any factor that is not within the set range.

[0028] In one embodiment, the processing unit is further configured to detect that the decision result of the ego vehicle is not within the feasible domain of the ego vehicle, and output the decision result of the ego vehicle yielding the right of way.

[0029] In a third aspect, the present application provides an intelligent driving system, comprising at least one processor, which is used to execute instructions stored in a memory and execute various possible implementation embodiments of the first aspect.

[0030] In a fourth aspect, the present application provides a vehicle comprising at least one processor, wherein the processor is configured to execute various possible implementations of the first aspect.

[0031] In a fifth aspect, the present application provides an intelligent driving system, comprising a sensor system and a processor, wherein the processor is configured to execute various possible implementations of the first aspect.

[0032] In a sixth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute various possible implementations of the first aspect.

[0033] In a seventh aspect, the present application provides a computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the various possible implementation embodiments of the first aspect.

[0034] In an eighth aspect, the present application provides a computer program product, which stores instructions. When the instructions are executed by a computer, the computer implements the various possible implementation embodiments of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The following is a brief introduction to the drawings required for describing the embodiments or prior art.

[0036] Figure 1 A schematic diagram of the architecture of an intelligent driving system provided in an embodiment of the present application;

[0037] Figure 2 A schematic diagram of the architecture of a decision module provided in an embodiment of the present application;

[0038] Figure 3 Schematic diagrams of four common scenarios between the vehicle and obstacles provided in the embodiments of this application;

[0039] Figure 4 A schematic diagram of a scenario between a vehicle and a non-game object provided in an embodiment of the present application;

[0040] Figure 5 A schematic diagram of a scenario of trajectory conflict between the vehicle and the game object provided in Example 1 of the present application;

[0041] Figure 6 A schematic diagram of the functional relationship between the time domain security cost and the absolute value of TDTC provided in the first embodiment of the present application;

[0042] Figure 7 A schematic diagram showing the functional relationship between the spatial domain safety cost and the minimum distance between two vehicles provided in Example 1 of the present application;

[0043] Figure 8 A schematic diagram of the functional relationship between the comfort cost and the acceleration change provided in the first embodiment of the present application;

[0044] Figure 9 A schematic diagram of the functional relationship between the passability cost and the time to pass the collision point provided in Example 1 of the present application;

[0045] Figure 10 Schematic diagram of the functional relationship between the prior probability cost of the game object to overtake the vehicle and the probability of the game object letting the vehicle pass provided in Example 1 of the present application;

[0046] Figure 11 A schematic diagram showing the functional relationship between the right-of-way ratio and the distance from the social vehicle to the conflict point provided in the first embodiment of the present application;

[0047] Figure 12 A schematic diagram of the functional relationship between the historical decision result associated cost provided in the first embodiment of the present application and the cost of overtaking or yielding corresponding to each frame of image;

[0048] Figure 13 An encroachment relationship diagram of longitudinal distance and time for longitudinal planning by the motion planning module provided in Example 1 of the present application;

[0049] Figure 14 A schematic diagram of multi-vehicle conflict resolution provided in Example 2 of this application;

[0050] Figure 15 A flowchart of a decision-making method provided in an embodiment of the present application;

[0051] Figure 16 A structural block diagram of a decision-making device method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0053] An intelligent driving system uses sensors to detect the surrounding environment and its own status, such as navigation positioning information, road information, information about obstacles such as other vehicles and pedestrians, its own position information and motion status information, etc. After a certain decision-making and planning algorithm, it accurately controls the vehicle's speed and steering, thus realizing autonomous driving. Figure 1 As shown, according to the functional requirements of the intelligent driving system 100 , the system 100 can be divided into a prediction module 10 , a navigation module 20 , a decision module 30 , a planning module 40 and a control module 50 .

[0054] The prediction module 10 is used to obtain information such as the vehicle's position, the environment around the vehicle, and the vehicle's status through data collected by sensors such as the global positioning system (GPS) unit, inertial navigation system (INS) unit, odometer, camera, radar, etc. in the sensor system, and predict the path that the vehicle and various obstacles around the vehicle will travel in a certain period of time in the future.

[0055] The navigation module 20 may be an in-vehicle navigation system, a navigation application (APP) on an external terminal, etc., and is used to provide road condition information such as the navigation route of the vehicle, lane lines on the route, traffic lights, and forks.

[0056] The decision module 30 receives the predicted routes of the vehicle and surrounding vehicles over a future period of time from the prediction module 10, as well as the navigation route, lane markings, traffic lights, intersections, and other road condition information provided by the navigation module 20. It then determines whether the vehicle will collide with obstacles while traveling along the predicted (or navigation) route. If the vehicle does not collide with an obstacle, no negotiation occurs between the vehicle and the obstacle, and the vehicle's movement mode and trajectory are determined according to the pre-set rules. If the vehicle does collide with an obstacle, the negotiation results between the vehicle and each obstacle are calculated based on the input data, and each obstacle is labeled with a behavior label such as yield / cut over, avoid / follow.

[0057] The planning module 40 is used to receive the decision results output by the decision module 30, and determine the actions to be taken for each obstacle, such as giving way / overtaking, obstacle avoidance / following, etc., based on the behavior labels of the obstacles, such as which lane the vehicle should choose, whether to change lanes, whether to follow the vehicle, whether to detour, or whether to stop.

[0058] The control module 50 is used to follow the planning results issued by the planning module 40 and control the vehicle to achieve the desired speed and steering angle.

[0059] The following section of this application will use the decision module 30 as an example to specifically describe the technical solution of this application. Object decision-making refers to the process in which the autonomous vehicle, during autonomous navigation, needs to make decisions about obstacles in the environment and assign behavioral labels to them. For example, if an obstacle is determined to be circumvented, followed, or overtaken, the obstacle will be labeled as such.

[0060] Figure 2 This is a schematic diagram of the architecture of a decision module provided in an embodiment of the present application. Figure 2 As shown, the decision module 30 includes a game object screening unit 301 , a game decision unit 302 , a rule decision unit 303 and a conflict processing unit 304 .

[0061] The game object screening unit 301 determines whether the ego vehicle will collide with an obstacle while traveling along the reference path based on vehicle information, obstacle information, and road condition information input from other upper-level modules such as the sensor system, positioning module, and environmental perception module. Obstacles can be classified as game objects and non-game objects. Game objects are obstacles that may collide with the ego vehicle, while non-game objects are obstacles that are unlikely to collide with the ego vehicle.

[0062] In the present application, the vehicle information of the own vehicle includes the navigation route provided by the navigation module 20 of the own vehicle or the navigation device on the external terminal, as well as the speed, acceleration, heading angle, position and other data of the own vehicle detected by various sensors in the vehicle; the obstacle information includes the position of each obstacle, the distance between obstacles, the distance between each obstacle and the own vehicle, the type of each obstacle, the status of each obstacle, the historical trajectory of the own vehicle and each obstacle, the predicted driving trajectory and motion status in the future, as well as the speed, acceleration, heading angle and other data of each obstacle; the road condition information includes traffic light information, the indication information of road signs, etc.

[0063] Exemplarily, after obtaining the vehicle information, obstacle information, and road condition information of the ego vehicle, the game object screening unit 301 determines whether any obstacle is located at the same position as the ego vehicle (or the distance between the two positions is less than a set threshold) based on whether the ego vehicle's and each obstacle's trajectories intersect, or based on the ego vehicle's and each obstacle's trajectories, speed, acceleration, and other data. If an obstacle's trajectory intersects with the ego vehicle's trajectory or is located at the same position as the ego vehicle's, the obstacle is classified as a game object and its obstacle information is sent to the game decision unit 302. All other obstacles are classified as non-game objects, and their obstacle information is sent to the rule decision unit 303.

[0064] Generally speaking, the scenarios constructed between the self-vehicle and the game object can be roughly divided into four scenarios, such as Figure 3 As shown, specifically (taking the vehicle moving straight as an example):

[0065] Scenario 1: Single obstacle decision (tracks have intersections): If the vehicle is moving straight and an obstacle crosses it;

[0066] Scenario 2: Single obstacle decision (no track intersection, potential conflict): This scenario involves scenarios such as the vehicle traveling straight ahead and an obstacle merging into an adjacent lane or in the same lane ahead of the vehicle.

[0067] Scenario 3: Multiple obstacle decision-making (multiple game objects): For example, the vehicle is moving straight, and the planned path of the vehicle passes through multiple obstacles;

[0068] Scenario 4: Multiple Obstacle Decision-Making (with Gaming Objects and Non-Gaming Objects): For example, the vehicle is moving straight and following the vehicle in front (non-gaming object), while there is an obstacle on the side.

[0069] In this application, when processing multiple obstacle decisions, the game decision unit 302 will split the multiple obstacles into multiple single-obstacle decisions. The decision is made from the feasible domain between the self-vehicle and each obstacle. Then, the shared game strategy is extracted. Each feasible domain includes the shared game strategy to obtain the intersection of these feasible domains. If the intersection exists, the optimal game strategy for each game object in the intersection is calculated; if the intersection does not exist, the most conservative decision result for the current situation is output, such as the self-vehicle outputting the "yield" strategy. Therefore, this application only needs to have the game decision unit 302 handle the decision between the self-vehicle and a single game object. The specific implementation process is as follows:

[0070] 1. Sampling Strategy Space Generation: Game decision unit 302 determines the upper and lower limits of the game strategies for both players based on the predefined game mode, road conditions, and the mobility of the ego vehicle and obstacles, thereby obtaining a reasonable game decision range for both players. Within this game strategy range, feasible game strategies are sampled for both the ego vehicle and the obstacle to determine the number of feasible game strategies for both players. These feasible game strategies are then combined to create a game strategy space with multiple combinations.

[0071] If both the ego vehicle and the game object are vehicles, the corresponding actions of the game strategy can be turning the steering wheel and increasing or decreasing the throttle. Turning the steering wheel changes the vehicle's steering angle, causing the vehicle to perform actions such as overtaking or avoiding obstacles by changing its lateral displacement. Accelerating or decreasing the throttle changes the vehicle's acceleration and speed, causing the vehicle to perform actions such as overtaking or avoiding obstacles by changing its longitudinal displacement.

[0072] For example, using the game strategy of varying acceleration as an example, the game decision unit 302 determines the different game strategies that can be achieved by varying the acceleration values of the ego vehicle and the opponent based on received data such as the distance between the ego vehicle and the opponent from the theoretical collision point, the maximum and minimum acceleration values of the vehicles, the ego vehicle's speed, and the maximum speed limit of the road. This set of game strategies is then used as the game strategy range. Then, using a predetermined sampling method, n ego vehicle acceleration values and m opponent acceleration values are selected to obtain a space of n×m possible game strategies for both parties.

[0073] 2. Strategic Cost Evaluation: The strategic cost of each game strategy calculated by the game decision unit 302 is related to factors such as safety, comfort, passing efficiency, right of way, probability of passing obstacles, and historical decision-making methods. Therefore, when calculating the strategic cost of each game strategy, the cost of each factor can be calculated and then weighted to obtain the strategic cost of each game strategy. This application analyzes the strategic cost of each game strategy based on six factors: safety cost, comfort cost, passing efficiency cost, right of way cost, obstacle prior probability cost, and historical decision-related cost. The details are as follows:

[0074] (1) Safety cost: During the game, the two parties should maintain a reasonable safety distance. When the distance is less than the safety threshold or a collision occurs, a large safety cost should be incurred. The safety cost is inversely proportional to the distance between the two parties.

[0075] (2) Comfort cost: During the game, both parties tend to maintain their current motion state as long as they do not collide. When there are large changes in the motion state (such as acceleration, lateral acceleration, etc.), it will affect the passenger experience and generate a large experience cost. Among them, the comfort cost is inversely proportional to the degree of change in the motion state.

[0076] (3) Passing efficiency cost: During the game, both parties tend to pass through the current traffic scene as quickly as possible to complete the game. If both parties spend more time to complete the game, a larger passing efficiency cost will be incurred. The passing efficiency cost is inversely proportional to the time to complete the game.

[0077] (4) Road right cost: During the game, both parties tend to follow the driving order specified in the traffic rules. If the game strategy differs significantly from the driving rules specified in the road right information, a larger road right cost will be incurred. The road right cost is proportional to the degree of violation of the regulations.

[0078] (5) Prior probability cost of obstacles: During the game process, the decision results of obstacles tend to approach the prior probability of the corresponding behavior obtained from observation. If the game strategy deviates greatly from the prior probability, a larger prior probability cost of obstacles will be generated. Among them, the prior probability of obstacles is related to the game scenario. If the game scenario is a game decision of overtaking / yielding, the prior probability of obstacles is the prior probability of overtaking; if the game scenario is a game decision of avoiding / not avoiding obstacles, the prior probability of obstacles is the prior probability of avoiding obstacles.

[0079] (6) Historical decision-related costs: During the game, both parties tend to maintain the decision results obtained in the previous frame of the game. When the game results change, a larger historical decision-related cost is generated.

[0080] 3. Generation of strategy feasible domain: The game decision unit 302 weights the costs of the above six factors according to certain rules to obtain the strategy cost of each game strategy. It then evaluates and screens the rationality of each factor weighted to the strategy cost of each game strategy, deletes the strategy costs of game strategies that include unreasonable factors, and thus screens out the strategy costs of reasonable game strategies as the feasible domain of the vehicle and the game object.

[0081] The rule decision unit 303 is used to estimate the feasible domain of non-game objects. In this application, in order to deal with the conflict of decision results between non-game objects and game objects, the feasible domain of the self-vehicle for the constraint area should be estimated based on the constraint area constituted by the non-game objects. For example, for longitudinal (referring to the direction along the road where the self-vehicle is traveling) action games (such as overtaking / yielding), a virtual wall will be virtualized in front of the self-vehicle as the upper limit constraint of acceleration; for lateral (perpendicular to the direction of the road where the self-vehicle is traveling) action games, the self-vehicle will use the lateral maximum offset range of the non-game objects as a constraint, thereby constructing the feasible domain of the self-vehicle and the non-game objects. Among them, the virtual wall refers to the longitudinal constraint generated by decision / planning, which usually refers to the speed of the self-vehicle passing a certain position point.

[0082] For example, Figure 4 In the scenario shown, the ego vehicle (black box) and the game objects (A and B) are making a decision to cut in / yield. The optimal game strategy for the ego vehicle may be to cut in, but due to the obstacle in front of the ego vehicle, the ego vehicle cannot cut in. Therefore, the upper limit of the ego vehicle's acceleration in this scenario should be generated based on the target vehicle it is following, which serves as the feasible domain for the ego vehicle and the non-game objects.

[0083] After obtaining the feasible domains of the self-vehicle and each game object sent by the game decision unit 302 and the feasible domains of the self-vehicle and each non-game object sent by the rule decision unit 303, the conflict handling unit 304 calculates the intersection of the received feasible domains. If an intersection exists, the optimal game strategy of the self-vehicle for each game object in the intersection is calculated; if an intersection does not exist, the most conservative decision result under the current situation is output, such as the decision result of giving way to each game object.

[0084] The present application provides an object decision-making scheme for autonomous driving vehicles based on a sampled game space. By receiving data sent by a prediction module, a navigation module, and a sensor system, a sampled game space is constructed between the vehicle and each game object. The cost of each factor affecting the vehicle game is then calculated, and the strategy cost of each game strategy is obtained by weighting. The strategy cost of the game strategy including unreasonable factors is eliminated to obtain the feasible domain between the vehicle and each game object. The feasible domain between the vehicle and non-game objects is then combined to calculate the optimal game strategy of the vehicle for each game object. Since the scheme does not rely on the regulations of the scene, it can be adapted to all scenes. At the same time, during the game process, when facing multiple game objects, the vehicle can simultaneously play games with multiple game objects by finding the intersection of the feasible domains of the vehicle and each game object.

[0085] The following will describe in detail how the game decision unit 302 determines the passable domain through two embodiments.

[0086] Example 1

[0087] like Figure 5 The traffic scenario shown shows a trajectory conflict. In this scenario, the planned path of the ego vehicle (black) conflicts with the predicted path of the social vehicle (gray), potentially leading to a collision at the intersection. The upper-level module provides the planned reference path of the ego vehicle, the predicted trajectory of the social vehicle, and the current speed, acceleration, and distance to the collision point of the ego vehicle and the social vehicle. The ego vehicle uses this information to conduct longitudinal maneuvers, such as cutting in or yielding to an obstacle.

[0088] 1. Sampling Strategy Space Generation

[0089] The longitudinal game strategy between the ego vehicle and the other player can be characterized by the magnitude of acceleration (cutting in / yielding). First, the upper and lower decision limits for the game strategy (acceleration) are generated, taking into account the longitudinal dynamics of the vehicle, kinematic constraints, and the relative position and velocity relationship between the ego vehicle and the other player. Figure 5 In the scenario, at the current moment, the ego vehicle's speed is 17 km / h, the oncoming straight vehicle's speed is 15 km / h, the distance between the ego vehicle and point X is 20.11 m, the distance between the oncoming game vehicle and point X is 35.92 m, and the ego vehicle's feedforward acceleration (planned acceleration) is 0.5 m / s2 , the observed acceleration of the vehicle is -0.67m / s 2 , the observed acceleration of the social car is 0.0m / s 2 , the static speed limit on the road is 60km / h and the curvature speed limit on the path is 30km / h. The allowed acceleration range of the vehicle is [-4.0, 2.0]m / s 2 The allowed acceleration range for social vehicles is [-3.55, 3.0] m / s 2 Considering the balance between computational complexity and spatial accuracy of sampling strategy, the acceleration interval is set to 1m / s. 2 , and finally the sampling strategy space shown in Table 1 can be generated.

[0090] Table 1 Example 1 Constructing the game strategy space

[0091]

[0092] 2. Strategic Cost Evaluation of Game Strategies

[0093] The strategic cost of each game strategy is quantitatively described by the cost of each design factor. The costs mentioned in this application include safety cost, comfort cost, passability cost, prior probability cost of the game object, road right cost and historical decision result association cost. The total cost is the weighted sum of the six costs. For each decision strategy pair in the strategy space, calculate the corresponding total benefit. Here, [1.0, -1.45] (the vehicle sampling acceleration is 1m / s 2 , the social vehicle sampling acceleration is 1.45m / s 2 ) This strategy is described in detail as an example.

[0094] 1. Safety costs can be divided into temporal safety costs and spatial safety costs. The temporal safety cost is based on the time difference to collision (TDTC) between the ego vehicle and the game vehicle. The larger the TDTC, the safer it is and the smaller the temporal safety cost. The quantitative relationship is as follows: Figure 6 As shown. Taking the [1.0,1.45] strategy pair as an example, the time it takes for the ego car to reach the collision point is eTTC = 3.41s, and the time it takes for the game car to reach the collision point is oTTC = 5.02s. Then the |TDTC| of this strategy pair is |eTTC–oTTC| = |-1.65| = 1.65s. Figure 6 Based on the relationship shown, the time domain security cost can be calculated as 100000*1.0=100000, where 100000 is the time domain security cost weight.

[0095] Spatial domain safety is based on the sampled acceleration corresponding to the game strategy of the two vehicles along the planned path of the self-vehicle and the predicted path of the obstacle, and recursively predicts the future movement of the two parties in the game to obtain the minimum distance between the two vehicles in the next 10 seconds (with a recursive step of 0.2 seconds). The larger the minimum distance, the safer it is, and the lower the cost of spatial domain safety. The quantitative relationship is as follows: Figure 7 As shown in Figure 2, taking the minimum distance between two vehicles as 0.77m as an example, the spatial domain safety cost is 10000*0.1756=1756, where 10000 is the spatial domain safety cost weight.

[0096] 2. The comfort cost is related to the acceleration change rate (jerk) of the vehicle / game object. The smaller the jerk, the better the comfort and the smaller the comfort cost. The quantitative relationship is as follows: Figure 8 As shown. Taking the [1.0,1.45] strategy pair as an example, for the ego car, its acceleration change is eDeltaAcc=1-(-0.67)=1.67m / s2, and the comfort cost is eComfcost=100*0.315=31.5. Among them, 100 is the weight of the ego car's comfort cost. For the game car, its acceleration change is oDeltaAcc=1.45-(0)=1.45m / s 2 The comfort cost is oComf cost = 300 * 0.286 = 85.94, where 300 is the weight of the comfort cost of the social vehicle.

[0097] 3. The passability cost is related to the time it takes to pass the collision point. If the time it takes to slow down and give way to the collision point is longer, the passability cost will increase. On the contrary, if the time it takes to speed up and pass the collision point is shorter, the passability cost will decrease. The quantitative relationship is as follows: Figure 9 As shown. Among them, Figure 9 The horizontal axis is the difference deltaPassTime between the time samplePassTime to pass the collision point calculated using the acceleration in the strategy space and the time realPassTime to pass the collision point using the currently observed speed and acceleration, and the vertical axis is the passability cost.

[0098] For example, the time it takes for the ego vehicle to pass the collision point at its current acceleration and velocity is eRealPassTime = 4.47 seconds, and the time it takes for the game car to pass the collision point at its current acceleration and velocity is oRealPassTime = 7.2 seconds, using the [1.0, 1.45] strategy. For the ego vehicle, its time to pass the collision point is eSamplePassTime = 3.41 seconds. The difference between this and the observed pass time is eDeltaPassTime = eSamplePassTime – eRealPassTime = 3.41 - 4.47 = -1.06. The passability cost is ePassCost = 100 * 4.1150 = 411.5. 100 is the passability cost weight for the ego vehicle.

[0099] For the game car, the time it takes to pass the collision point is oSamplePassTime = 5.02 seconds. The difference between this time and the observed time is oDeltaPassTime = oSamplePassTime – oRealPassTime = 5.02 - 7.2 = -2.18 seconds. The passability cost is oPassCost = 100 * 2.3324 * 233.24, where 100 is the passability cost weight of the social car.

[0100] 4. The prior probability cost of the game object is related to the probability of the game object giving way to the car. The greater the probability of giving way, the smaller the prior probability cost of the game object. The quantitative relationship is as follows: Figure 10 As shown in the figure, the yield probability of the game player reflects its individual driving style. It is a dynamic factor that depends on its historical speed, acceleration, position, and other information. It is the input of the game module. In the scenario described above, the yield probability of the game player is 0.2, and the corresponding eProb cost of the ego car cutting in is (1-0.2)*1000=800. 1000 is the prior probability cost weight.

[0101] 5. The right-of-way cost describes the degree to which both parties in the game comply with traffic rules. The vehicle with the right of way has an advantage in overtaking, so its overtaking cost should be reduced, while the cost of yielding to the own vehicle should be increased. The right-of-way cost depends on the current traffic right-of-way relationship between the social vehicle and the own vehicle, as well as the objective distance from the social vehicle to the conflict point. The formula for calculating the right-of-way cost is:

[0102] Road right cost = dynamic road right ratio of the scene * dynamic road right weight

[0103] The right of way ratio of the scenario = f(distance between the social vehicle and the conflict point), a nonlinear function in the range of [0.0, 1.0]. The right of way ratio of the scenario refers to the concept of the social vehicle obtaining the right of way, which is related to the distance between the social vehicle and the conflict point. The specific quantitative relationship is as follows: Figure 11As shown in the figure, if the distance between the social car and the conflict point is less than 5m, the right of way ratio of the scene is 1; if the distance between the social car and the conflict point is greater than 5m, the right of way ratio of the scene decreases as the distance between the social car and the conflict point increases; if the distance between the social car and the conflict point is greater than 100m, the right of way ratio of the scene is 0.

[0104] Dynamic road right weight = road right weight base value + 1000 * road right value of the scene.

[0105] For example, in a scenario where the opponent is going straight and the ego vehicle is turning left, the opponent has the right of way. The road weight for this scenario is 0.4, and the distance from the opponent to the conflict point is 35.92 meters. The cost of the ego vehicle taking the right of way increases, i.e., eGWRoadRightCost = f(35.92)*(5000+1000*0.4)=0.61*5400=3305.67. Here, 500 is the base right of way weight.

[0106] 6. The introduction of the historical decision result association cost is to prevent the decision jump between two consecutive frames. If the previous frame is a rushing time, the rushing time cost of the current frame is reduced, making it easier to output the rushing time decision. If the previous frame is a yielding time, the yielding time cost of the current frame is reduced, making it easier to output the yielding time decision. The quantitative relationship between the historical decision result association cost and the rushing time cost or yielding time cost corresponding to each frame image is as follows: Figure 12 As shown, if the K-th frame image is represented as grabbing the line, after obtaining the K+1-th frame image, if the cost of grabbing the line is calculated, the association cost between the K-th frame image and the K+1-th frame image will decrease; if the cost of giving way is calculated, the association cost between the K-th frame image and the K+1-th frame image will increase.

[0107] The hysteresis value between the ego vehicle's decision to yield and then to overtake is 50. This means that if the previous frame was YD, the YD cost in the current frame is reduced by 50. Conversely, the hysteresis value between the ego vehicle's decision to overtake and then to yield is 20. This means that if the previous frame was GW, the GW cost in the current frame is reduced by 20. If the previous frame decision was yielding, the associated cost for this historical decision is 50. As shown in the next section, the optimal YD cost for this frame is 11087.57. Therefore, the final optimal YD cost is 11087.57 – 50 = 11037.57.

[0108] By summing the above six costs, we can get the final cost Total corresponding to the strategy space point {1,1.45}. That is:

[0109] Total cost = 100,000 + 3,305.67 + 1,756 + 31.50 + 85.94 - 50 + 800 + 411.5 + 233.24 = 106,573.85.

[0110] Among them, each item in the formula is the time domain safety cost, road right cost, spatial domain safety cost, vehicle comfort cost, game vehicle comfort cost, inter-frame correlation cost, prior probability cost and passability cost.

[0111] At the same time, by determining the time difference between the two vehicles reaching the collision point (TDTC = -1.65s), we can conclude that the ego vehicle arrived at the collision point first. This strategic point corresponds to the decision to cut in. Then, performing the above calculations for each action combination in Table 1, we can obtain the total cost for all action combinations, as shown in Table 2.

[0112] Table 2 Cost of each game strategy pair in Example 1

[0113]

[0114] The cost values in Table 2 are divided into normal font, bold font, and italic font. Among them, when the ego car arrives at the conflict point before the game car, it indicates that the ego car's behavior is to cut in (the cost value corresponding to the normal font represents the strategic cost of the game strategy of cutting in); when the ego car arrives at the conflict point after the game car, it indicates that the ego car's behavior is to yield (the cost value corresponding to the italic font represents the strategic cost of the game strategy of yielding); when the ego car and the game car both brake before reaching the conflict point, it indicates that the behavior of the ego car and the game car is to yield (the cost value corresponding to the bold font represents the strategic cost of the game strategy of yielding).

[0115] 3. Generation of Strategy Feasible Domain

[0116] After step 2, all possible action combinations within the valid action space between the ego vehicle and the game vehicle are generated, along with the total cost associated with each action pair. For these action combinations, the sub-costs are evaluated and screened, and all reasonable alternative combinations are selected to form the feasible domain for the ego vehicle for the game object. Sub-costs such as comfort and right-of-way are considered valid within their specified ranges.

[0117] For safety costs, validity assessment is required, and unreasonable strategy pairs are directly deleted. For temporal safety costs, if the TDTC is less than a certain threshold (1s), the ego vehicle and the social vehicle are considered to have a collision risk. This action combination cannot be output as a valid action, so this action pair is deleted. Similarly, for spatial safety costs, if the recursive minimum distance is less than a certain threshold (0.01m), the ego vehicle and the social vehicle are considered to have a collision risk, and this action pair is deleted. For passability costs, if both vehicles stop before the conflict point, traffic efficiency decreases, making this action pair infeasible, and thus deleted. The remaining valid actions constitute the feasible domain of the strategy space, as shown in Table 3.

[0118] Table 3. Feasible strategy sets (feasible domains) and their costs in Example 1

[0119]

[0120] Among them, "Delete-1" is the invalid action output of the time domain security cost, "Delete-2" is the invalid action output of the space domain security cost, and "Delete-3" is the invalid action output of the pass cost.

[0121] From the above analysis, the feasible domain for the ego vehicle to target this game object is the valid acceleration combinations listed in the table. Within this feasible domain, the game strategy with the lowest total cost is selected as the final decision. This strategy, in other words, maximizes the global payoff and serves as the optimal action combination for both the ego vehicle and the game vehicle, ensuring adequate safety, passability, and comfort. When selecting this optimal strategy, the game decision module transmits the optimal acceleration for the game object to the downstream motion planning module, which then performs planning based on this acceleration value.

[0122] Figure 5 In the scenario shown, the ego vehicle does not have the right of way (traffic regulations require turning vehicles to yield to those going straight). Throughout the process, obstacles are labeled "yield." While determining whether the opponent will cut in, the system calculates the acceleration at which the opponent will cut in. This acceleration value is sent to the motion planner, which then selects the corresponding ego vehicle acceleration value based on the corresponding acceleration combination within the feasible domain, enabling it to accurately plan the yielding maneuver.

[0123] When the ego vehicle and the game object just enter the intersection, the ego vehicle has a larger feasible domain for the game object and uses the yield (-4,0) m / s 2 , rush [0,2]m / s 2 The interaction can be completed with accelerations within the range. However, according to the results of the weighted summation of various costs, it can be seen that when selecting (-3, 0.45) m / s 2 When the acceleration strategy is right, the game pair composed of the vehicle and the game object has an optimal solution. The main advantage of this game result compared to other solutions is that it follows the road right description while ensuring sufficient safety and comfort. In this optimal game strategy, the obstacle will be (0.45) m / s 2 The acceleration value is sent to the motion planning layer. The motion planning module will correct the obstacle prediction trajectory based on the acceleration value, which is specifically reflected in the translation of the obstacle's occupied area on the T axis. Figure 13 The speed planning is performed based on the station-time (ST) relationship diagram of the longitudinal distance and time occupied by obstacles on the vehicle path to achieve safe yielding.

[0124] The above process primarily describes how the game proceeds in each frame. Regarding the overall object decision-making process, as the object approaches, in the first frame of the game, the ego vehicle yields to the object. As the object approaches the trajectory intersection, the safety cost of cutting in increases. Therefore, in selecting the optimal game outcome, the ego vehicle must continue to yield to the object until it passes the trajectory intersection, ending the game.

[0125] The obstacle-based overtaking / yielding decision-making scheme of Example 1 of this application does not rely on specific obstacle interactions or trajectory intersection characteristics. Instead, it utilizes a sensor system to acquire traffic scene information, rationally abstracting the traffic scene and achieving generalization of application scenarios. Simultaneously, it determines the acceleration at which the ego vehicle should overtake / yield, and the acceleration at which the obstacle will overtake / yield. These values are then used to influence motion planning, ensuring the correct execution of decision instructions.

[0126] Example 2

[0127] by Figure 4 Taking the scenario shown as an example, the ego vehicle turns left at an unprotected intersection, while social vehicles A and B are driving straight in the opposite lane. There is a social vehicle C driving in the same direction in front of the ego vehicle. The ego vehicle needs to follow social vehicle C, and a game relationship is formed between the ego vehicle and social vehicles A and B.

[0128] Assume that at the current moment, the vehicle's speed is 8 km / h, the oncoming straight vehicle A's speed is 14 km / h, the oncoming straight vehicle B's speed is 14 km / h, the distance between the vehicle and the intersection with vehicle A is 21 meters, the distance between the vehicle and the intersection with vehicle B is 25 meters, the distance between vehicle A and the vehicle's intersection is 35 meters, and the distance between vehicle B and the vehicle's intersection is 24 meters. The feedforward acceleration of the vehicle is 0.22 m / s. 2 , the observed acceleration of the vehicle is 0.0m / s 2 , the observed acceleration of the social car is 0.0m / s 2 The static speed limit on the road is 40 km / h and the curvature speed limit is 25 km / h. The speed of vehicle C is 10 km / h and the acceleration is 0.0 m / s 2 The distance from the rear of the vehicle to the front of the ego vehicle is 15m. The acceleration sampling interval allowed by the ego vehicle is [-4, 2] m / s 2 , the acceleration sampling interval allowed for social car A and social car B is [-3.55, 3.0] m / s 2 Considering the balance between computational complexity and spatial accuracy of sampling strategy, the acceleration interval is set to 1m / s. 2 .

[0129] A single-car game is performed for social car A and social car B, respectively. The cost function design, weight distribution, and feasible region selection methods are consistent with those in Example 1. The feasible regions corresponding to social cars A and B are obtained. The feasible regions of the self-car and social car A are shown in Table 4.

[0130] Table 4 The feasible domain of the self-car and social car A in Example 2

[0131]

[0132] The costs in Table 2 are categorized into normal, bold, and italic fonts. Normal font corresponds to the ego car cutting in, italic font corresponds to the ego car yielding, and bold font corresponds to the ego car and the game car braking before the conflict point. The feasible regions [0.45, -1] and [-1.55, -2] for the ego car and the social car A are the optimal costs in the set of cutting in or yielding costs.

[0133] The feasible domains of the self-car and the social car B are shown in Table 5.

[0134] Table 5 Feasible regions of the self-car and social car A in Example 2

[0135]

[0136]

[0137] The cost values in Table 2 are categorized into normal, bold, and italic fonts. Normal font corresponds to the ego car cutting in, italic font corresponds to the ego car yielding, and bold font corresponds to the ego car and the game car braking before the conflict point. The feasible regions [0.45, -1] and [-3.55, -2] for the ego car and the social car A are the optimal costs in the set of cutting in or yielding costs.

[0138] For social car C (not a game player), the ego car's decision-making must not create risks with it. The ego car's acceleration feasible domain must be estimated based on its speed, acceleration, speed, acceleration of social car C, and the distance from the ego car to social car C. This is implemented by the longitudinal planning module, and its calculation model is:

[0139] accUpLimit = speedGain*(objV-egoV)+distTimeGain*(distance to the vehicle ahead-minimum following distance) / egoV.

[0140] Among them, accUpLimit is the decision upper limit of acceleration, objV is the obstacle speed, egoV is the ego vehicle speed, speedGain and distTimeGain are adjustable parameters.

[0141] Then directly use its output value, that is, for the social car C, substitute this scenario parameter into the above calculation model, and we can get:

[0142] 0.85*(10 / 3.6-8 / 3.6)+0.014*(15-4.56) / (8 / 3.6)=0.8

[0143] That is, the upper limit of the vehicle's acceleration is 0.8m / s 2 , the feasible region is [-4,0.8]m / s 2 .

[0144] When using this embodiment to play a separate game with social cars A and B, the result is that A takes the right of way and B gives way. However, in reality, such actions cannot be completed at the same time. Therefore, multi-car conflict resolution is required. Combining the feasible domains of the above three social cars, the conflict resolution diagram is as follows: Figure 14 The data in the "Cost" and "Decision Tag" columns are the costs corresponding to the optimal decision of a single vehicle, and the data in the "Ego Vehicle Acceleration Sampling Space" column are the public feasible domains for the three social vehicles and the ego vehicle.

[0145] For the feasible domain formed by the social game car A, B and the non-game car (social car C), first find their intersection, and we can get the common feasible domain of the car to these three objects as [-4.0, -1.0] m / s 2 , so in this public domain [-4.0,-1.0]m / s 2 Find the optimal policy cost within the public feasible domain. Calculate the sum of the policy costs of the ego car for social cars A and B in this common feasible domain. The optimal solution is found when the ego car's acceleration is -1.0. At this point, for social car A, the ego car's optimal cost is 11135.31, and social car A's expected acceleration is 0.45, resulting in the decision to yield. For social car B, the ego car's optimal cost is 11067.41, and social car B's expected acceleration is 0.45, resulting in the decision to yield.

[0146] Therefore, the final optimal solution for both vehicles is: yield to the social vehicle A (expected acceleration is 0.45), and then yield to the social vehicle B (expected acceleration is 0.45). The optimal expected acceleration of the ego vehicle is -1.0. This comprehensive conflict resolution strategy maximizes overall benefits and ensures that the decision is feasible for all obstacles.

[0147] Figure 4In the scenario shown, the ego vehicle makes a comprehensive decision based on all considered obstacles, achieving an optimal game solution that simultaneously satisfies multiple obstacles. When the vehicles on the ego vehicle's planned path impose a virtual wall constraint on the ego vehicle, this corresponds to a feasible domain with an acceleration range of [-4.0, 0.8]. During the game, the ego vehicle makes game decisions for both social vehicles A and B, obtaining a feasible domain for each vehicle. The intersection of these three feasible domains yields a feasible domain that satisfies all obstacles in the scenario. Within this feasible domain, the ego vehicle's optimal solution for all game obstacles is determined, resulting in the ego vehicle yielding to both social vehicles A and B.

[0148] The second embodiment primarily addresses the multi-objective game scenario. For non-game vehicles, the corresponding feasible domain is estimated based on the game type. For game vehicles, the feasible domain is directly derived. Within the feasible domain, the optimal solution for all game objects is found to achieve consistency in multi-objective decision-making.

[0149] In the multi-car game method proposed in Example 2 of the present application, the feasible domain of the sampling space of each game car is first obtained, and then the feasible domain of the self-car is estimated for non-game cars. Finally, the public feasible domain is taken from the feasible domains of the game car and non-game car, and the optimal solution therein is calculated, and finally the global optimal solution of the self-car and multiple social cars is obtained.

[0150] Figure 15 A flowchart of a decision-making method provided in an embodiment of the present application. Figure 15 As shown, the embodiment of the present application provides a decision-making method, and the specific implementation process is as follows:

[0151] Step S1501: Obtain the predicted motion trajectories of the vehicle and obstacles around the vehicle.

[0152] Among them, the predicted motion trajectory can obtain information such as the vehicle's location, the environment around the vehicle, and the vehicle status through data collected by sensors such as the GPS unit, INS unit, odometer, camera, and radar in the sensor system. Then, the obtained information is processed to predict the path that the vehicle and the obstacles around the vehicle will take in the future.

[0153] Step S1503: Determine the game object, wherein the game object is an obstacle among the obstacles around the vehicle that intersects with the predicted motion trajectory of the vehicle or whose distance to the vehicle is less than a set threshold.

[0154] Specifically, after obtaining the predicted motion trajectories of the ego vehicle and each obstacle, the present application determines whether the predicted motion trajectories of the ego vehicle and each obstacle intersect, or, based on the predicted motion trajectories, driving trajectories, speed, acceleration, and other data of the ego vehicle and each obstacle, determines whether the distance between the position of the obstacle and the ego vehicle is less than a set threshold. If it is detected that the predicted motion trajectory of an obstacle intersects the predicted motion trajectory of the ego vehicle, or the distance between the two vehicles is less than a set threshold, then this type of obstacle is classified as a game object; other obstacles are classified as non-game objects.

[0155] In step S1505, based on the vehicle information of the ego vehicle, the obstacle information of the game object, and the road condition information collected by the sensor system, a sampling game space is constructed for each game object. Each sampling game space is a set of different game strategies between the ego vehicle and an obstacle in the game object.

[0156] This application determines the game strategy range for the ego vehicle and each game object, such as the ego vehicle's acceleration range and speed range, based on predefined game modes, road conditions, the mobility of the ego vehicle and each obstacle, and other factors. Within this game strategy range, feasible game strategies are sampled for the ego vehicle and each game object to obtain the number of feasible game strategies for the ego vehicle and each game object. The feasible game strategies for the ego vehicle and each game object are then combined to obtain a variety of different game strategy spaces.

[0157] For example, using acceleration as a game mode, based on received data such as the distance between the ego vehicle and a player from the theoretical collision point, the vehicle's maximum and minimum acceleration values, the ego vehicle's speed, and the road's maximum speed limit, the game strategy types that can be achieved by varying the ego vehicle and the player are determined. This set of game strategies is then used as the game strategy range. Then, using a predetermined sampling method, n ego vehicle acceleration values and m player acceleration values are selected, resulting in a game strategy space of n×m possible combinations of the two players.

[0158] Step S1507: Calculate the strategy cost of each game strategy, wherein the strategy cost is a value obtained by weighting the weights of various factors that affect the strategy cost.

[0159] Among them, factors that affect the strategic cost include safety, comfort, passing efficiency, right of way, the probability of giving way to obstacles, historical decision-making methods, etc. Therefore, when calculating the strategic cost of each game strategy, we can calculate the cost of each factor and then perform a weighted calculation on the cost of each factor to obtain the cost of each game strategy.

[0160] Optionally, the present application determines whether the ego vehicle or each obstacle reaches the conflict point first based on the distance between the ego vehicle and the game object and the conflict point, and the set of game strategies between the ego vehicle and each obstacle in the game object. When the decision-making strategy of the ego vehicle and the obstacle in a game strategy determines that the ego vehicle arrives at the conflict point before the obstacle, indicating that the ego vehicle's behavior is to cut in, the game strategy is labeled as "ego vehicle cuts in"; when the decision-making strategy of the ego vehicle and the obstacle in a game strategy determines that the ego vehicle arrives at the conflict point after the obstacle, indicating that the ego vehicle's behavior is to yield, the game strategy is labeled as "ego vehicle yields"; when the decision-making strategy of the ego vehicle and the obstacle in a game strategy determines that both the ego vehicle and the obstacle stop before the conflict point, indicating that the ego vehicle and the obstacle's behavior is to yield, the game strategy is labeled as "ego vehicle and obstacle both give way".

[0161] Step S1509: Determine the decision result of the vehicle, wherein the decision result is the game strategy with the minimum strategy cost among the same game strategies in each sampled game space.

[0162] Specifically, the cost of each factor is weighted according to a certain rule to obtain the cost of each game strategy. The factors weighted into each game strategy cost are then evaluated for rationality and screened. Game strategy costs that include unreasonable factors are removed, thereby screening out reasonable game strategy costs as the feasible domain for the ego vehicle and the game object. After obtaining the feasible domains for each ego vehicle and each game object, the intersection of these feasible domains is calculated to obtain a common feasible domain that satisfies the current scenario when the ego vehicle encounters multiple game objects. The game strategy with the lowest game cost is then selected from this common feasible domain as the decision result.

[0163] Optionally, to address conflicting decision outcomes between non-game objects and game objects, the feasible domain of the ego vehicle within the constraint region created by the non-game objects is estimated. For example, for longitudinal (i.e., along the direction of the ego vehicle's travel) action games (e.g., overtaking / yielding), a virtual wall is created in front of the ego vehicle to serve as an upper limit on acceleration. For lateral (perpendicular to the direction of the road) action games, the ego vehicle uses the maximum lateral offset range of the non-game objects as a constraint, thereby constructing a feasible domain for the ego vehicle and the non-game objects. The intersection of the common feasible domain between the ego vehicle and each game object and the feasible domain between the ego vehicle and the non-game objects is then calculated. The game strategy with the lowest game cost is selected from the intersection as the decision outcome. If no game strategy is found in the intersection, the decision outcome of "yielding" is selected based on the principle of "safety."

[0164] In the embodiment of the present application, the predicted motion trajectories of the ego vehicle and the obstacles around it are obtained, and the game object is determined by determining whether the predicted motion trajectories intersect or whether the distance between the two vehicles is less than a set threshold. Then, a sampling game space is constructed between the ego vehicle and each obstacle, and the strategy cost of each game strategy in each sampling game space is calculated. By solving the same game strategy in each sampling game space, the game strategy with the lowest strategy cost among the same game strategies is selected as the game result. Since this solution is independent of the scenario, it can be adapted to all scenarios. At the same time, during the game process, when facing multiple game objects, by solving the same game strategy in each sampling game space, the ego vehicle can play games with multiple game objects simultaneously.

[0165] Figure 16 This is a schematic diagram of the architecture of a decision-making device provided in an embodiment of the present application. Figure 16 The device 1600 shown includes a transceiver unit 1601 and a processing unit 1602. Specifically, it performs the following functions:

[0166] The transceiver unit 1601 is configured to obtain predicted motion trajectories of the ego vehicle and each obstacle surrounding the ego vehicle. The processing unit 1702 is configured to determine a game object, which is an obstacle among the obstacles surrounding the ego vehicle that intersects with the predicted motion trajectory of the ego vehicle or is located at a distance from the ego vehicle that is less than a set threshold. A sampling game space is constructed for each game object based on vehicle information collected by the sensor system, obstacle information of the game object, and road condition information. Each sampling game space is a set of different game strategies adopted between the ego vehicle and an obstacle in the game object. The strategy cost of each game strategy is calculated, where the strategy cost is a value obtained by weighting the weights of various factors affecting the strategy cost. A decision result of the ego vehicle is determined, where the decision result is the game strategy with the smallest strategy cost in a shared sampling game space. The shared sampling game space includes at least one game strategy, and each sampling game space includes a game strategy in the shared sampling game space.

[0167] In one embodiment, the processing unit 1602 is specifically used to construct a feasible domain for each sampled game space, where the feasible domain of each sampled game space is at least one game strategy corresponding to a strategy cost that meets the set requirements; in the intersection of the feasible domains of all sampled game spaces, the game strategy with the smallest strategy cost among the same game strategies is determined.

[0168] In one embodiment, the processing unit 1602 is further configured to determine a non-game object, which is an obstacle among various obstacles around the ego vehicle that does not intersect with the predicted motion trajectory of the ego vehicle or is at a distance from the ego vehicle that is not less than a set threshold; construct a feasible domain of the ego vehicle based on the vehicle information of the ego vehicle collected by the sensor system, the obstacle information of the non-game object, and the road condition information, where the feasible domain of the ego vehicle is at least one strategy for the ego vehicle to take different decisions without colliding with the non-game object; and output the decision result of the ego vehicle when it is detected that the decision result of the ego vehicle is within the feasible domain of the ego vehicle.

[0169] In one embodiment, the processing unit 1602 is specifically used to determine the decision upper limit and decision lower limit of the self-vehicle and each obstacle in the game object based on the vehicle information of the self-vehicle, the obstacle information of the game object and the road condition information; according to the set rules, obtain the decision strategy of the self-vehicle and each obstacle in the game object from the decision upper limit and decision lower limit of each obstacle in the game object; combine the decision strategy of the self-vehicle with the decision strategy of each obstacle in the game object to obtain the at least one game strategy of the self-vehicle and each obstacle in the game object.

[0170] In one embodiment, the processing unit 1602 is also used to determine the behavior label of each game strategy based on the distance between the ego vehicle and the game object and the conflict point, and the at least one game strategy of the ego vehicle and each obstacle in the game object. The conflict point is the position where the predicted motion trajectory of the ego vehicle and the obstacle intersects or the position where the distance between the ego vehicle and the obstacle is less than a set threshold. The behavior label includes at least one of the ego vehicle giving way, the ego vehicle overtaking, and the ego vehicle and the obstacle both giving way.

[0171] In one embodiment, the processing unit 1602 is specifically used to determine various factors of the strategy cost, and the various factors of the strategy cost include at least one of safety, comfort, passing efficiency, road right, prior probability of obstacles and historical decision correlation; calculate the factor cost of each factor in each strategy cost; and weight the factor cost of each factor in each strategy cost to obtain the strategy cost of each game strategy.

[0172] In one embodiment, the processing unit 1602 is further configured to compare whether each factor in the strategy cost is within a set range; and delete the game strategy corresponding to the strategy cost including any factor that is not within the set range.

[0173] In one embodiment, the processing unit 1602 is further configured to detect that the decision result of the ego vehicle is not within the feasible domain of the ego vehicle, and output the decision result of the ego vehicle yielding the right of way.

[0174] The present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute any one of the above methods.

[0175] The present invention provides a computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, any one of the above methods is implemented.

[0176] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.

[0177] In addition, various aspects or features of the embodiments of the present application can be implemented as methods, devices or products using standard programming and / or engineering techniques. The term "product" used in this application covers computer programs that can be accessed from any computer-readable device, carrier or medium. For example, computer-readable media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks or tapes, etc.), optical disks (e.g., compact discs (CDs), digital versatile discs (DVDs), etc.), smart cards and flash memory devices (e.g., erasable programmable read-only memories (EPROMs), cards, sticks or key drives, etc.). In addition, the various storage media described herein may represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing and / or carrying instructions and / or data.

[0178] In the above embodiment, Figure 16The decision-making device 1600 in the embodiment can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).

[0179] It should be understood that in various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0180] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0181] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the unit is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0182] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of the solution of this embodiment according to actual needs.

[0183] If this function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or an access network device, etc.) to execute all or part of the steps of the method of each embodiment of the embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0184] The above is only a specific implementation method of the embodiment of the present application, but the protection scope of the embodiment of the present application is not limited to this. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in the embodiment of the present application, which should be covered by the protection scope of the embodiment of the present application.

Claims

1. A decision-making method, characterized in that: include: Obtain the predicted motion trajectory of the ego vehicle and the obstacles around it; Determining a game object, where the game object is an obstacle among the obstacles around the ego vehicle whose predicted motion trajectory intersects with the predicted motion trajectory of the ego vehicle, or an obstacle among the obstacles around the ego vehicle whose real-time distance to the ego vehicle is less than a set threshold; Constructing a sampling game space for each of the game objects based on the vehicle information of the vehicle, obstacle information and road condition information of the game objects collected by the sensor system, each sampling game space including at least one game strategy; Calculate the strategic cost of each game strategy, where the value of the strategic cost is determined by weighting each factor of the strategic cost; Determine a decision result of the vehicle, where the decision result is a game strategy with the smallest strategy cost in a shared sampling game space, where the shared sampling game space includes at least one game strategy, and each of the shared sampling game spaces includes a game strategy in the shared sampling game space.

2. The method according to claim 1, characterized in that Determining the decision result of the vehicle includes: A feasible domain of each sampled game space is constructed, where the feasible domain of each sampled game space is at least one game strategy corresponding to a strategy cost that meets set requirements.

3. The method according to claim 1, characterized in that The method further comprises: Determining non-game objects, where the non-game objects are obstacles among the obstacles around the ego vehicle whose predicted motion trajectories do not intersect with the predicted motion trajectory of the ego vehicle, or obstacles among the obstacles around the ego vehicle whose real-time distance to the ego vehicle is not less than a set threshold; Constructing a feasible domain of the ego vehicle based on the vehicle information of the ego vehicle, the obstacle information of the non-game object, and the road condition information collected by the sensor system, wherein the feasible domain of the ego vehicle is at least one strategy for the ego vehicle to adopt different decisions without colliding with the non-game object; It is detected that the decision result of the vehicle is within the feasible region of the vehicle, and the decision result of the vehicle is output.

4. The method according to claim 1, wherein The method constructs a sampling game space for each of the game objects based on the vehicle information of the vehicle, the obstacle information and the road condition information of the game objects collected by the sensor system, including: Determining a decision upper limit and a decision lower limit for each obstacle in the self-vehicle and the game object based on the vehicle information of the self-vehicle, the obstacle information of the game object, and the road condition information; According to the set rules, the decision strategy of each obstacle in the self-vehicle and the game object is obtained from the decision upper limit and the decision lower limit of each obstacle in the game object; The decision-making strategy of the ego vehicle is combined with the decision-making strategy of each obstacle in the game object to obtain the at least one game strategy of the ego vehicle and each obstacle in the game object.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: A behavior label for each game strategy is determined based on the distance between the ego vehicle and the game object and the conflict point, and the at least one game strategy of the ego vehicle and each obstacle in the game object. The conflict point is the location where the predicted motion trajectory of the ego vehicle and the obstacle intersects or the location where the distance between the ego vehicle and the obstacle is less than a set threshold. The behavior label includes at least one of the following: the ego vehicle yields, the ego vehicle overtakes, and both the ego vehicle and the obstacle yield.

6. The method according to claim 1, characterized in that The calculation of the strategy cost of each game strategy includes: Determining various factors of the strategic cost, wherein the various factors of the strategic cost include at least one of safety, comfort, passing efficiency, right of way, prior probability of obstacles, and historical decision relevance; Calculate the factor cost of each factor in each strategy cost; The factor cost of each factor in each strategy cost is weighted to obtain the strategy cost of each game strategy.

7. The method according to any one of claims 1 to 4 or 6, characterized in that After calculating the strategy cost of each game strategy, the method further includes: Compare each factor in the strategy cost to see if it is within the set range; Delete the game strategy corresponding to the strategy cost of any factor that is not within the set range.

8. The method according to any one of claims 1 to 4 or 6, characterized in that: The method further comprises: It is detected that the decision result of the own vehicle is not within the feasible region of the own vehicle, and the decision result of the own vehicle yielding is output.

9. A decision-making device, characterized in that: include: A transceiver unit is used to obtain the predicted motion trajectory of the ego vehicle and the obstacles around it; a processing unit configured to determine a game object, wherein the game object is an obstacle among the obstacles around the ego vehicle whose predicted motion trajectory intersects with the predicted motion trajectory of the ego vehicle, or an obstacle among the obstacles around the ego vehicle whose real-time distance to the ego vehicle is less than a set threshold; Constructing a sampling game space for each of the game objects based on the vehicle information of the vehicle, obstacle information and road condition information of the game objects collected by the sensor system, each sampling game space including at least one game strategy; Calculating the strategic cost of each gaming strategy, where the value of the strategic cost is determined by weighting each factor of the strategic cost; and Determine a decision result of the vehicle, where the decision result is a game strategy with the smallest strategy cost in a shared sampling game space, where the shared sampling game space includes at least one game strategy, and each of the shared sampling game spaces includes a game strategy in the shared sampling game space.

10. The device according to claim 9, characterized in that The processing unit is specifically used for A feasible domain of each sampled game space is constructed, where the feasible domain of each sampled game space is at least one game strategy corresponding to a strategy cost that meets set requirements.

11. The device according to claim 9, characterized in that The processing unit is also used to Determining non-game objects, where the non-game objects are obstacles among the obstacles around the ego vehicle whose predicted motion trajectories do not intersect with the predicted motion trajectory of the ego vehicle, or obstacles among the obstacles around the ego vehicle whose real-time distance to the ego vehicle is not less than a set threshold; Constructing a feasible domain for the ego vehicle based on the vehicle information, the obstacle information of the non-game object, and the road condition information collected by the sensor system, wherein the feasible domain for the ego vehicle is at least one strategy for the ego vehicle to adopt different decisions without colliding with the non-game object; It is detected that the decision result of the vehicle is within the feasible region of the vehicle, and the decision result of the vehicle is output.

12. The device according to claim 9, characterized in that The processing unit is specifically used for Determining a decision upper limit and a decision lower limit for each obstacle in the self-vehicle and the game object based on the vehicle information of the self-vehicle, the obstacle information of the game object, and the road condition information; According to the set rules, the decision strategy of each obstacle in the self-vehicle and the game object is obtained from the decision upper limit and the decision lower limit of each obstacle in the game object; The decision-making strategy of the ego vehicle is combined with the decision-making strategy of each obstacle in the game object to obtain the at least one game strategy of the ego vehicle and each obstacle in the game object.

13. The device according to any one of claims 9 to 12, characterized in that: The processing unit is also used to A behavior label for each game strategy is determined based on the distance between the ego vehicle and the game object and the conflict point, and the at least one game strategy of the ego vehicle and each obstacle in the game object. The conflict point is the location where the predicted motion trajectory of the ego vehicle and the obstacle intersects or the location where the distance between the ego vehicle and the obstacle is less than a set threshold. The behavior label includes at least one of the following: the ego vehicle yields, the ego vehicle overtakes, and both the ego vehicle and the obstacle yield.

14. The device according to claim 9, characterized in that The processing unit is specifically used for Determining various factors of the strategic cost, wherein the various factors of the strategic cost include at least one of safety, comfort, passing efficiency, right of way, prior probability of obstacles, and historical decision relevance; Calculate the factor cost of each factor in each strategy cost; The factor cost of each factor in each strategy cost is weighted to obtain the strategy cost of each game strategy.

15. The device according to any one of claims 9 to 12 or 14, characterized in that: The processing unit is further configured to compare whether each factor in the strategy cost is within a set range; Delete the game strategy corresponding to the strategy cost of any factor that is not within the set range.

16. The device according to any one of claims 9 to 12 or 14, characterized in that: The processing unit is further configured to detect that the decision result of the ego vehicle is not within the feasible domain of the ego vehicle, and output the decision result of the ego vehicle yielding the right of way.

17. A vehicle, comprising at least one processor, wherein the processor is configured to execute instructions stored in a memory to perform the method according to any one of claims 1 to 8.

18. An intelligent driving system, comprising a sensor system and a processor, wherein the processor is configured to execute the method according to any one of claims 1 to 8.

19. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 8.

20. A computing device comprising a memory and a processor, characterized in that: The memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 8 is implemented.

21. A computer program product, characterized in that The computer program product stores instructions, and when the instructions are executed by a computer, the computer is caused to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for establishing automatic driving lane changing decision model in hybrid driving environment

    CN110298131A

  • Automatic driving vehicle decision planning method considering interactive game

    CN112373485A