Intelligent driving decision method, vehicle driving control method, device and vehicle

By repeatedly releasing the strategy space of the vehicle and the game opponent in intelligent driving, the feasible domain of the strategy is determined, which solves the problem of excessive computing power consumption in the existing technology, realizes efficient and flexible intelligent driving decision-making, and is applicable to autonomous vehicles at Level 2 and above.

CN115943354BActive Publication Date: 2026-05-29YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YINWANG INTELLIGENT TECHNOLOGIES CO LTD
Filing Date
2021-07-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies for autonomous driving at Level 2 and above struggle to effectively reduce computing power consumption while ensuring decision-making accuracy. In particular, the computational burden is too heavy in multi-dimensional game spaces, resulting in insufficient hardware computing power and making it difficult to commercialize.

Method used

By repeatedly releasing the strategy space of the vehicle and the game opponent, the feasible domain of the strategy is determined, reducing the number of releases and computational load. The strategy space is released sequentially in the vertical, horizontal, and time dimensions. Combining the cost values ​​such as safety, right-of-way, and passability, the final decision result is determined.

Benefits of technology

While ensuring decision-making accuracy, the number of policy space releases and computational loads are reduced, lowering the requirements for hardware computing power and improving the flexibility and safety of intelligent driving decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115943354B_ABST
    Figure CN115943354B_ABST
Patent Text Reader

Abstract

The application relates to an intelligent driving decision method, which relates to intelligent driving technology, and comprises the following steps: firstly, determining a game object of a self vehicle; then, performing multiple releases of each strategy space from multiple strategy spaces of the self vehicle and the game object, determining a strategy feasible region of the self vehicle and the game object according to the released strategy spaces, and determining a decision result of self vehicle driving according to the strategy feasible region, wherein the decision result is an executable behavior action of the self vehicle. Through multiple releases of the strategy space, the decision result can be obtained under the condition of releasing as few strategy spaces as possible while keeping the decision accuracy, the calculation amount is reduced, and the demand for hardware computing power is lowered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to intelligent driving technology, and in particular to intelligent driving decision-making methods, vehicle driving control methods, devices, and vehicles. Background Technology

[0002] With the development of artificial intelligence technology, autonomous driving technology is being widely applied, thereby reducing the driving burden on drivers. Regarding autonomous driving, the Society of Automotive Engineers (SAE International) has proposed five levels: L1-L5. Level 1 is driver assistance, capable of helping the driver complete certain driving tasks, but only one at a time. Level 2 is partial automation, capable of automatically accelerating, decelerating, and steering simultaneously. Level 3 is conditional automation, where the vehicle can automatically accelerate, decelerate, and steer in specific environments without driver intervention. Level 4 is high automation, allowing for fully driverless driving, but with limitations such as speed limits and a relatively fixed driving area. Level 5 is full automation, fully adaptive driving, adapting to any driving scenario. Higher levels indicate more advanced autonomous driving capabilities.

[0003] Currently, autonomous driving technologies at Level 2 and above that require human driver intervention as needed are generally considered to fall under the category of intelligent driving. When a vehicle is in intelligent driving mode, it needs to be able to perceive surrounding obstacles in a timely and accurate manner, such as oncoming vehicles, vehicles crossing the road, stationary vehicles, and pedestrians, and make decisions based on its own driving behavior and trajectory, such as accelerating, decelerating, and changing lanes. Summary of the Invention

[0004] This application provides an intelligent driving decision-making method, a vehicle driving control method, a device, and a vehicle, which can make decisions about their own driving while consuming as little computing power as possible, while ensuring decision-making accuracy.

[0005] The first aspect of this application provides an intelligent driving decision-making method, including: obtaining the game object of the vehicle; performing multiple releases of multiple strategy spaces from multiple strategy spaces of the vehicle and the game object; after one of the multiple releases is executed, determining the policy feasible region of the vehicle and the game object based on the released strategy spaces; and determining the decision result of the vehicle driving based on the policy feasible region.

[0006] The feasible policy domain for the vehicle and the non-player object includes the actions that the vehicle can perform relative to the non-player object. Therefore, by releasing multiple policy spaces multiple times, while ensuring decision accuracy (which can be, for example, the probability of executing the decision result), the feasible policy domain is obtained with minimal release of policy space. A pair of actions is then selected from the feasible policy domain as the decision result, thus minimizing the number of policy space releases and computations, and reducing the requirements for hardware computing power.

[0007] As one possible implementation of the first aspect, the dimensions of the multiple policy spaces include at least one of the following: longitudinal sampling dimension, lateral sampling dimension, or temporal sampling dimension.

[0008] Based on the longitudinal sampling dimension, lateral sampling dimension, or temporal sampling dimension, multiple policy spaces are spanned. These multiple policy spaces include a longitudinal sampling policy space spanned by the longitudinal sampling dimension of the player and / or the game objects, a lateral sampling policy space spanned by the lateral sampling dimension of the player and / or the game objects, a temporal policy space spanned by the player and / or the game objects in the temporal sampling dimension, or a policy space formed by any pairwise or triple combination of the longitudinal, lateral, or temporal sampling dimensions. The temporal policy space corresponds to the policy spaces spanned in multiple single-frame deductions included in a single decision step, and the policy space spanned in each single-frame deduction can include both the longitudinal sampling policy space and / or the lateral sampling policy space.

[0009] Based on the above, a corresponding policy space can be spanned in at least one sampling dimension according to the traffic scenario, and the policy space can be released.

[0010] As one possible implementation of the first aspect, performing multiple releases of multiple policy spaces includes performing the releases in the following order: longitudinal sampling dimension, lateral sampling dimension, and temporal sampling dimension.

[0011] The above, following the order of releasing the vertical sampling dimension, releasing the horizontal sampling dimension, and releasing the temporal sampling dimension, can result in multiple releases of the policy space, which may include the following policy spaces:

[0012] The vertical sampling strategy space spanned by a set of values ​​of the vehicle's vertical sampling dimension; the vertical sampling strategy space spanned by another set of values ​​of the vehicle's vertical sampling dimension; the vertical sampling strategy space spanned by a set of values ​​of the vehicle's vertical sampling dimension and a set of values ​​of the game object's vertical sampling dimension; the vertical sampling strategy space spanned by another set of values ​​of the vehicle's vertical sampling dimension and a set of values ​​of the game object's vertical sampling dimension; the vertical sampling strategy space spanned by another set of values ​​of the vehicle's vertical sampling dimension and another set of values ​​of the game object's vertical sampling dimension; the horizontal sampling strategy space spanned by a set of values ​​of the vehicle's horizontal sampling dimension, and the strategy space spanned by the vertical sampling strategy space spanned by the vehicle's vertical sampling dimension and / or the game object's vertical sampling dimension; the horizontal sampling strategy space spanned by another set of values ​​of the vehicle's horizontal sampling dimension, and the strategy space spanned by the vehicle's vertical sampling dimension and / or the game object's vertical sampling dimension. The strategy space spanned by the vertical sampling dimensions of the game and the vertical sampling strategy space spanned by ... vertical sampling strategy space spanned by the game and the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the game and the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by the vertical sampling strategy space spanned by Furthermore, after determining the policy feasible domains of the vehicle and the game object based on the released policy spaces, and determining the decision result of the vehicle's driving based on the policy feasible domains, the released time-dimensional policy space includes: the policy spaces spanned in multiple single-frame deductions included in a single decision step. In each single-frame deduction, the spanned policy space may include the aforementioned vertical sampling policy spaces, horizontal sampling policy spaces, and the policy space spanned by the vertical sampling policy spaces and the horizontal sampling policy spaces.

[0013] As described above, sequentially executing multiple strategy space releases—that is, first changing the vehicle's acceleration longitudinally, then adjusting its offset laterally—is more in line with driving habits and better meets driving safety requirements. Finally, based on multi-frame deductions across the time dimension, a decision with better temporal consistency can be further determined from multiple feasible domains.

[0014] As one possible implementation of the first aspect, when determining the policy feasible region of the vehicle and the game object, the total cost of action pairs in the policy feasible region is determined according to one or more of the following: the safety cost of the vehicle or the game object, the right-of-way cost, the lateral offset cost, the passability cost, the comfort cost, the inter-frame correlation cost, and the risk area cost.

[0015] Based on the above, one or more substitution values ​​can be selected as needed to calculate the total substitution value, which is used to determine the feasible region.

[0016] As one possible implementation of the first aspect, when the total value of a pair of actions is determined based on two or more values, each value has a different weight.

[0017] As described above, these different weights can respectively focus on driving safety, right-of-way, passability, comfort, and risk. By flexibly setting each weight, the flexibility of intelligent driving decision-making is increased. In some possible implementations, the weight allocation can be as follows: safety weight > right-of-way weight > lateral offset weight > passability weight > comfort weight > risk area weight > inter-frame correlation weight.

[0018] As one possible implementation of the first aspect, when there are two or more game objects, the decision outcome of the vehicle's driving is determined based on the feasible domains of each strategy of the vehicle and each game object.

[0019] Therefore, when there are multiple game objects, the final feasible strategy region is determined by obtaining the feasible strategies of each object separately and then finding the intersection of these feasible strategies. Here, the intersection refers to actions that all include the same action of the same vehicle.

[0020] As one possible implementation of the first aspect, it also includes: obtaining the non-game object of the vehicle; determining the policy feasible region between the vehicle and the non-game object; the policy feasible region between the vehicle and the non-game object includes the actions that the vehicle can perform relative to the non-game object; and determining the decision result of the vehicle's driving based at least on the policy feasible region between the vehicle and the non-game object.

[0021] Therefore, when there are non-game objects, the final decision result is related to the non-game objects.

[0022] As one possible implementation of the first aspect, the strategic feasible region of the vehicle's driving decision is determined based on the intersection of the strategic feasible regions of the vehicle and each game player, or the strategic feasible region of the vehicle's driving decision is determined based on the intersection of the strategic feasible regions of the vehicle and each game player, as well as the strategic feasible regions of the vehicle and each non-game player.

[0023] Therefore, when there are multiple game partners and the vehicle itself, the final feasible strategy domain and the vehicle's driving decision can be obtained by finding the intersection of the feasible strategies of the vehicle and the multiple game partners. When there are multiple game partners and non-game partners, the final feasible strategy domain and the vehicle's driving decision can also be obtained by finding the intersection of the feasible strategies of the vehicle and the multiple game partners and non-game partners.

[0024] As a possible implementation of the first aspect, it also includes: obtaining the non-game object of the vehicle; constraining the longitudinal sampling strategy space corresponding to the vehicle, or constraining the lateral sampling strategy space corresponding to the vehicle, based on the motion state of the non-game object.

[0025] The above defines the constraints on the longitudinal sampling strategy space corresponding to the vehicle, which is the range of values ​​used in the longitudinal sampling dimension of the vehicle when the constraints are expanded into the longitudinal sampling strategy space; and the constraints on the lateral sampling strategy space corresponding to the vehicle, which is the range of values ​​used in the lateral sampling dimension of the vehicle when the constraints are expanded into the lateral sampling strategy space.

[0026] As shown above, the range of longitudinal acceleration or lateral offset values ​​of the vehicle in Zhang's strategy space can be constrained by the motion state of non-game objects, such as position and speed, which reduces the number of actions in the strategy space and can further reduce the amount of computation.

[0027] As a possible implementation of the first aspect, it also includes: obtaining the non-game objects of the vehicle's game objects; and constraining the longitudinal sampling strategy space corresponding to the vehicle's game objects, or constraining the lateral sampling strategy space corresponding to the vehicle's game objects, based on the motion state of the non-game objects.

[0028] The above defines the range of values ​​in the vertical sampling strategy space corresponding to the game object of the self-vehicle, which is the range of values ​​used in the vertical sampling dimension corresponding to the game object of the self-vehicle when constraining the vertical sampling strategy space; and the range of values ​​in the horizontal sampling strategy space corresponding to the game object of the self-vehicle, which is the range of values ​​used in the horizontal sampling dimension of the game object of the self-vehicle when constraining the horizontal sampling strategy space.

[0029] As described above, the range of longitudinal acceleration or lateral offset values ​​of the self-vehicle game objects in Zhangcheng's strategy space can be constrained by the motion states of non-game objects, such as position and speed, thereby reducing the number of actions in the strategy space and further reducing the computational load.

[0030] As one possible implementation of the first aspect, when the intersection is an empty set, a conservative decision is made regarding the vehicle's movement. The conservative decision includes actions that allow the vehicle to stop safely or actions that allow the vehicle to decelerate safely.

[0031] From the above, it can be achieved that the vehicle can drive safely when the feasible domain of the strategy is empty.

[0032] As one possible implementation of the first aspect, the game object or non-game object is determined based on the attention method.

[0033] Based on the above, the game partners and non-game partners can be determined according to the attention allocated to the vehicle by each obstacle. This attention method can be implemented through algorithms or through neural network inference.

[0034] As one possible implementation of the first aspect, it also includes: displaying at least one of the following through a human-computer interaction interface: the decision result of the vehicle's driving, the policy feasible domain of the decision result, the driving trajectory of the vehicle corresponding to the decision result of the vehicle's driving, or the driving trajectory of the game object corresponding to the decision result of the vehicle's driving.

[0035] Therefore, the decision-making results of the vehicle or the corresponding game can be displayed with rich content on the human-computer interaction interface, making the interaction with the user more user-friendly.

[0036] The second aspect of this application provides an intelligent driving decision-making device, comprising: an acquisition module for acquiring the game object of the vehicle; and a processing module for executing multiple releases of multiple strategy spaces from multiple strategy spaces of the vehicle and the game object, wherein after one of the multiple releases is executed, the feasible domain of the strategies of the vehicle and the game object is determined based on the released strategy spaces, and the decision result of the vehicle driving is determined based on the feasible domain of the strategies.

[0037] As one possible implementation of the second aspect, the dimensions of the multiple policy spaces include at least one of the following: longitudinal sampling dimension, lateral sampling dimension, or temporal sampling dimension.

[0038] As one possible implementation of the second aspect, performing multiple releases of multiple policy spaces includes performing the releases in the following order: longitudinal sampling dimension, lateral sampling dimension, and temporal sampling dimension.

[0039] As a possible implementation of the second aspect, when determining the policy feasible region of the vehicle and the game object, the total cost of action pairs in the policy feasible region is determined according to one or more of the following: safety cost of the vehicle or the game object, right-of-way cost, lateral offset cost, passability cost, comfort cost, inter-frame correlation cost, and risk area cost.

[0040] As one possible implementation of the second aspect, when the total value of a behavior-action pair is determined based on two or more value pairs, each value pair has a different weight.

[0041] As one possible implementation of the second aspect, when there are two or more game objects, the decision result of the vehicle's driving is determined according to the feasible domain of each strategy of the vehicle and each game object.

[0042] As a possible implementation of the second aspect, the acquisition module is also used to acquire the non-game object of the vehicle; the processing module is also used to determine the policy feasible region between the vehicle and the non-game object; the policy feasible region between the vehicle and the non-game object includes the actions that the vehicle can perform relative to the non-game object; and the decision result of the vehicle's driving is determined at least based on the policy feasible region between the vehicle and the non-game object.

[0043] As a possible implementation of the second aspect, the processing module is also used to determine the policy feasible region of the decision result of the self-vehicle's driving based on the intersection of the policy feasible regions of the self-vehicle and each policy feasible region of each game object, or to determine the policy feasible region of the decision result of the self-vehicle's driving based on the intersection of the policy feasible regions of the self-vehicle and each policy feasible region of the self-vehicle and each non-game object.

[0044] As a possible implementation of the second aspect, the acquisition module is also used to acquire the non-game objects of the vehicle; the processing module is also used to constrain the longitudinal sampling strategy space corresponding to the vehicle, or constrain the lateral sampling strategy space corresponding to the vehicle, based on the motion state of the non-game objects.

[0045] As a possible implementation of the second aspect, the acquisition module is also used to acquire the non-game objects of the game objects of the vehicle; the processing module is also used to constrain the longitudinal sampling strategy space corresponding to the game objects of the vehicle, or constrain the lateral sampling strategy space corresponding to the game objects of the vehicle, based on the motion state of the non-game objects.

[0046] As one possible implementation of the second aspect, when the intersection is an empty set, a conservative decision is made regarding the vehicle's movement. The conservative decision includes actions that allow the vehicle to stop safely or actions that allow the vehicle to decelerate safely.

[0047] As one possible implementation of the second aspect, the game object or non-game object is determined based on the attention method.

[0048] As a possible implementation of the second aspect, the processing module is also used to display at least one of the following through a human-computer interaction interface: the decision result of the vehicle's driving, the policy feasible domain of the decision result, the driving trajectory of the vehicle corresponding to the decision result of the vehicle's driving, or the driving trajectory of the game object corresponding to the decision result of the vehicle's driving.

[0049] A third aspect of this application provides a vehicle driving control method, comprising: acquiring external obstacles; determining a vehicle driving decision result based on any method of the first aspect, in relation to the obstacles; and controlling the vehicle driving based on the decision result.

[0050] The fourth aspect of this application provides a vehicle driving control device, comprising: an acquisition module for acquiring obstacles outside the vehicle; a processing module for determining a decision result for vehicle driving based on any method of the first aspect, in response to the obstacles; and the processing module is further configured to control the driving of the vehicle based on the decision result.

[0051] The fifth aspect of this application provides a vehicle, including: the vehicle driving control device of the fourth aspect, and a driving system; the vehicle driving control device controls the driving system.

[0052] The sixth aspect of this application provides a computing device, including: a processor and a memory storing program instructions thereon, wherein when executed by the processor, the program instructions cause the processor to implement any intelligent driving decision method of the first aspect, or when executed by the processor, the program instructions cause the processor to implement a vehicle driving control method of the third aspect.

[0053] The seventh aspect of this application provides a computer-readable storage medium having program instructions stored thereon, wherein when executed by a processor, the program instructions cause the processor to implement any intelligent driving decision method of the first aspect, or when executed by a processor, the program instructions cause the processor to implement a vehicle driving control method of the third aspect.

[0054] These and other aspects of this application will become more apparent in the description of the following embodiments(s). Attached Figure Description

[0055] The following description, with reference to the accompanying drawings, further illustrates the various features of this application and the relationships between them. The drawings are exemplary; some features are not shown to scale, and some drawings may omit conventional features in the field of this application that are not essential to it, or additional features that are not essential to this application may be shown. The combination of features shown in the drawings is not intended to limit this application. Furthermore, throughout this specification, the same reference numerals refer to the same things. The specific description of the drawings is as follows:

[0056] Figure 1 A schematic diagram of a traffic scenario of vehicles traveling on a road, provided in an embodiment of this application;

[0057] Figure 2 This is a schematic diagram illustrating the application of this application's embodiments to a vehicle;

[0058] Figures 3A-3E Schematic diagrams of game objects and non-game objects in different traffic scenarios provided in the embodiments of this application;

[0059] Figure 4 A flowchart of the intelligent driving decision-making method provided in the embodiments of this application;

[0060] Figure 5 for Figure 4 The flowchart for obtaining the game object;

[0061] Figure 6 for Figure 4 The flowchart for obtaining the decision results;

[0062] Figure 7 This is a schematic diagram of multi-frame inference provided in the embodiments of this application;

[0063] Figures 8A-8F A schematic diagram of the cost function provided in an embodiment of this application;

[0064] Figure 9 A flowchart of driving control provided in another embodiment of this application;

[0065] Figure 10 This is a schematic diagram of a traffic scenario in the embodiments of this application;

[0066] Figure 11 A flowchart of driving control provided for a specific embodiment of this application;

[0067] Figure 12 A schematic diagram of the intelligent driving decision-making device provided in the embodiments of this application;

[0068] Figure 13 A flowchart of a vehicle driving control method provided in an embodiment of this application;

[0069] Figure 14 A schematic diagram of a vehicle driving control device provided in an embodiment of this application;

[0070] Figure 15 This is a schematic diagram of a vehicle provided in an embodiment of this application;

[0071] Figure 16 This is a schematic diagram of one embodiment of the computing device of this application. Detailed Implementation

[0072] The technical solutions provided in this application will be further described below with reference to the accompanying drawings and embodiments. It should be understood that the system architecture and business scenarios provided in the embodiments of this application are mainly for illustrating possible implementations of the technical solutions of this application and should not be construed as the sole limitation on the technical solutions of this application. Those skilled in the art will recognize that the technical solutions provided in this application are equally applicable to similar technical problems as system architectures evolve and new business scenarios emerge.

[0073] It should be understood that the intelligent driving decision-making schemes provided in the embodiments of this application include intelligent driving decision-making methods, devices, vehicle driving control methods and devices, vehicles, electronic devices, computing devices, computer-readable storage media, and computer program products. Since these technical solutions solve problems based on the same or similar principles, some repetitive details may not be repeated in the following descriptions of specific embodiments. However, it should be considered that these specific embodiments have mutual references and can be combined with each other.

[0074] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. In case of any inconsistency, the meaning set forth in this specification or derived from the content described herein shall prevail. Furthermore, the terminology used herein is intended to describe the purposes of embodiments of this application and is not intended to limit the application.

[0075] Figure 1 The image shows a traffic scene with vehicles moving on a road. Figure 1 As shown, in this traffic scenario, a north-south road and an east-west road form intersection A. Vehicle 901 is located south of intersection A and traveling from south to north; vehicle 902 is located north of intersection A and traveling from north to south; vehicle 903 is located east of intersection A and traveling from east to south, meaning it will turn left at intersection A to merge into the north-south road; behind vehicle 901 is vehicle 904, also traveling from south to north; near the southeast corner of intersection A, on the side of the north-south road, is vehicle 905, located in front of vehicle 901. Assuming vehicle 901 has its intelligent driving function activated, it can detect the current traffic scenario and make decisions regarding its driving strategy. Based on these decisions, it can control the vehicle's movement, such as accelerating, decelerating, or changing lanes according to the chosen strategy of overtaking, yielding, or avoiding obstacles. One intelligent driving decision-making scheme is based on game theory to determine driving strategies. For example, a first vehicle 901 uses game theory to decide its driving strategy in a traffic scenario where a second vehicle 902 is traveling in the opposite direction. However, game theory-based driving strategy decisions struggle to handle complex traffic scenarios, such as... Figure 1The traffic scenario illustrated involves a first vehicle 901 facing both a second vehicle 902 traveling in the opposite direction and a third vehicle 903 crossing intersection A from the side. A fifth vehicle 905 is also parked on the roadside ahead. When making a driving strategy decision based on a game theory approach, the first vehicle 901's game partners are simultaneously the second and third vehicles 902 and 903. Therefore, a multi-dimensional game space is required for decision-making, such as a multi-dimensional game space spanned by the lateral and longitudinal driving dimensions of each vehicle. Using a multi-dimensional game space leads to an explosive increase in the number of solutions, resulting in a geometric increase in computational burden, posing a significant challenge to existing hardware computing power. Therefore, currently, due to hardware computing power constraints, using a multi-dimensional game space for decision-making is difficult to commercialize in intelligent driving scenarios.

[0076] This application provides an improved intelligent driving decision-making scheme. When applied to intelligent driving of a vehicle, the basic principle of this scheme includes: for the vehicle, identifying obstacles in the current traffic scene, which may include the vehicle's game partners and non-game partners. For a single game partner, releasing a single sampling dimension or a strategy space spanned by multiple sampling dimensions from the multi-dimensional game space between the vehicle and the single game partner, and searching for a solution between the vehicle and the single game partner within the strategy space after each release. When a solution exists, i.e., when there is a game result, that is, when the vehicle and the single game partner have a policy feasible region within the strategy space, determining the vehicle's driving decision based on the game result, and then controlling the vehicle's driving based on the decision result. At this point, it is no longer necessary to release the remaining strategy space from the multi-dimensional game space. For multiple game opponents of the vehicle, the feasible strategy domains of the vehicle against each game opponent can be determined as described above. Using the vehicle's actions as indices, the driving strategy of the vehicle is obtained from the intersection of the feasible strategy domains of the vehicle and each game opponent (the intersection refers to all domains including the same action of the vehicle). This method, while ensuring decision accuracy (which can be, for example, the execution probability of the decided outcome), can obtain the optimal driving decision in a multi-dimensional game space with the fewest search iterations. It also minimizes the use of the strategy space, thus reducing the hardware computing power requirements and making it easier to commercialize in vehicles.

[0077] The implementing entity of the intelligent driving decision-making scheme in this application embodiment can be a powered, autonomously mobile intelligent agent. The intelligent agent can use the intelligent driving decision-making scheme provided in this application embodiment to engage in game-theoretic decision-making with other objects in the traffic scene, generating semantic-level decision labels and the agent's desired driving trajectory. This allows the intelligent agent to perform reasonable lateral and longitudinal motion planning. The intelligent agent can be, for example, a vehicle with autonomous driving capabilities, or an autonomously mobile robot. Vehicles here include general motor vehicles, such as land transportation devices including cars, sport utility vehicles (SUVs), multi-purpose vehicles (MPVs), automated guided vehicles (AGVs), buses, trucks, and other freight or passenger vehicles, as well as water transportation devices including various ships and boats, and aircraft. Motor vehicles also include hybrid vehicles, electric vehicles, gasoline vehicles, plug-in hybrid vehicles, fuel cell vehicles, and other alternative fuel vehicles. Hybrid vehicles refer to vehicles with two or more power sources, and electric vehicles include pure electric vehicles and range-extended electric vehicles. In some embodiments, the aforementioned autonomously mobile robot can also be classified as a type of vehicle.

[0078] The following description uses the intelligent driving decision-making scheme provided in the embodiments of this application applied to a vehicle as an example. Figure 2 As shown, when applied to a vehicle, the vehicle 10 may include an environmental information acquisition device 11, a control device 12, and a driving system 13. In some embodiments, it may also include a communication device 14, a navigation device 15, or a display device 16.

[0079] In this embodiment, the environmental information acquisition device 11 can be used to acquire information about the vehicle's external environment. In some embodiments, the environmental information acquisition device 11 may include a camera, lidar, millimeter-wave radar, ultrasonic radar, or a Global Navigation Satellite System (GNSS) as described later, and the number may be one or more. The camera may include a conventional RGB (Red, Green, Blue) three-primary-color camera sensor, an infrared camera sensor, etc. The acquired external environment includes road surface information and objects on the road surface. Objects on the road surface include surrounding vehicles, pedestrians, etc. Specifically, it may include the vehicle's motion state information, which may include vehicle speed, acceleration, heading angle information, trajectory information, etc. In some embodiments, the motion state information of surrounding vehicles may also be acquired through the vehicle 10's communication device 14. The external environment information acquired by the environmental information acquisition device 11 can be used to form a world model constructed from roads (corresponding to road surface information) and obstacles (corresponding to objects on the road surface).

[0080] In some other embodiments, the environmental information acquisition device 11 may also be an electronic device that receives vehicle external environment information transmitted by camera sensors, infrared night vision camera sensors, lidar, millimeter-wave radar, ultrasonic radar, etc., such as a data transmission chip. The data transmission chip may be a bus data transceiver chip, a network interface chip, etc., or it may be a wireless transmission chip, such as a Bluetooth chip or a Wi-Fi chip. In other embodiments, the environmental information acquisition device 11 may also be integrated into the control device 12, becoming an interface circuit or data transmission module integrated into the processor.

[0081] In this embodiment, the control device 12 can be used to make intelligent driving strategy decisions based on the acquired vehicle external environment information (including the constructed world model) and generate decision results. For example, the decision results may include acceleration, braking, steering (including lane changing or steering), and also the vehicle's short-term (e.g., within a few seconds) desired driving trajectory. In some embodiments, the control device 12 can further generate corresponding instructions based on the decision results to control the driving system 13, so that the driving system 13 can perform driving control of the vehicle and control the vehicle to achieve the desired driving trajectory according to the decision results. In this embodiment, the control device 12 can be an electronic device, such as a processor of an in-vehicle processing device such as a vehicle infotainment system, domain controller, mobile data center (MDC), or vehicle computer, or a conventional chip such as a central processing unit (CPU) or microcontroller (MCU).

[0082] In this embodiment, the driving system 13 may include a power system 131, a steering system 132, and a braking system 133, which will be described below:

[0083] The powertrain 131 may include an Electrical Control Unit (ECU) and a drive source. The ECU controls the driving force (such as torque) of the vehicle 10 by controlling the drive source. Examples of drive sources include an engine, a drive motor, etc. The ECU can control the drive source based on the driver's operation of the accelerator pedal or based on commands sent from the control device 12, thereby controlling the driving force. The driving force of the drive source is transmitted to the wheels via a transmission or the like, thereby driving the vehicle 10.

[0084] The steering system 132 may include an electronic steering control unit (ECU) and an electric power steering (EPS) system. The steering ECU can control the EPS motor according to the driver's operation of the steering wheel, or it can control the EPS motor according to the instructions sent from the control device 12, thereby controlling the direction of the wheels (specifically the steering wheels). In addition, steering can also be performed by changing the torque distribution or braking force distribution to the left and right wheels.

[0085] The braking system 133 may include a brake electronic control unit (ECU) and a braking mechanism. The braking mechanism operates the braking components through a brake motor, hydraulic mechanism, etc. The brake ECU can control the braking mechanism according to the driver's operation of the brake pedal, or according to the instructions sent from the control device 12, thereby controlling the braking force. If the vehicle 10 is an electric vehicle or a hybrid vehicle, the braking system 133 may also include an energy recovery braking mechanism.

[0086] In this embodiment, a communication device 14 may also be included. The communication device 14 is capable of interacting with external objects wirelessly to obtain the data required for the vehicle 10 to make intelligent driving decisions. In some embodiments, the external objects that can be communicated may include cloud servers, mobile terminals (such as mobile phones, laptops, tablets, etc.), roadside equipment, or surrounding vehicles. In some embodiments, the data required for decision-making includes user profiles of vehicles surrounding the vehicle 10 (i.e., other vehicles). These user profiles reflect the driving habits of the drivers of other vehicles and may also include the location and motion status information of other vehicles.

[0087] In this embodiment, a navigation device 15 may also be included. The navigation device 15 may include a Global Navigation Satellite System (GNSS) receiver and a map database. The navigation device 15 can determine the position of the vehicle 10 through satellite signals received by the GNSS receiver, and can generate a path to the destination based on map information in the map database, and provide information about the path (including the position of the vehicle 10) to the control device 12. The navigation device 15 may also have an inertial measurement unit (IMU), which can perform more accurate positioning of the vehicle 10 by fusing information from the GNSS receiver and information from the IMU.

[0088] In this embodiment, a display device 16 may also be included, such as a display screen installed in the center console of the vehicle cabin, or a head-up display (HUD). In some embodiments, the control device 12 can display the decision results in a user-understandable manner, such as the desired driving trajectory, arrows, text, etc., on the display device 16 in the vehicle cabin. In some embodiments, when displaying the desired driving trajectory, the current traffic scene of the vehicle (such as a graphical traffic scene) can also be combined and displayed on the display device in the vehicle cabin in the form of a partially magnified view. The control device 12 can also display information about the path to the destination provided by the navigation device 15.

[0089] In some embodiments, a voice playback system may also be included to prompt the user with the decision made in the current traffic scenario by playing voice prompts.

[0090] The intelligent driving decision-making method provided in the embodiments of this application will be described below. For ease of description, in the embodiments of this application, the intelligent driving vehicle that is in a traffic scene and executes the intelligent driving decision-making method provided in the embodiments of this application is referred to as the "automobile". From the perspective of the automobile, other objects in the traffic scene that affect or may affect the automobile's driving are referred to as obstacles of the automobile.

[0091] In this embodiment of the application, the vehicle has a certain behavioral decision-making ability and can generate driving strategies to change its own motion state. The driving strategies include acceleration, braking, and steering (including lane changing or steering). The vehicle also has the ability to execute driving behaviors, including executing the driving strategies and driving according to the desired driving trajectory determined by the decision.

[0092] In some embodiments, obstacles to a vehicle may also possess behavioral decision-making capabilities to change their own state of motion; for example, obstacles may be autonomously moving vehicles, pedestrians, etc. Obstacles to a vehicle may also lack behavioral decision-making capabilities or not change their state of motion; for example, obstacles may be vehicles parked on the roadside (inactive) or width-limiting barriers on the road. In summary, obstacles to a vehicle may include: pedestrians, bicycles, motor vehicles (such as motorcycles, cars, trucks, buses, etc.), wherein motor vehicles may include intelligent driving vehicles capable of executing intelligent driving decision-making methods.

[0093] Based on whether or not a game-theoretic interaction relationship is established with the vehicle, obstacles can be further categorized into game objects, non-game objects, or irrelevant obstacles. Specifically, the interaction intensity between game objects, non-game objects, and irrelevant obstacles and the vehicle gradually decreases from strong interaction to no interaction. It should be understood that in multiple traffic scenarios corresponding to different decision moments, game objects, non-game objects, and irrelevant obstacles may transform into each other. The position or movement state of irrelevant obstacles makes them completely unrelated to the vehicle's future behavior. There is no trajectory conflict or intention conflict between the vehicle and irrelevant obstacles in the future. Therefore, unless otherwise specified, obstacles in this embodiment refer to game objects and non-game objects of the vehicle.

[0094] If a non-game element of the vehicle has a potential trajectory conflict or intention conflict with the vehicle in the future, then the non-game element will constrain the vehicle's future behavior. However, the non-game element will not respond to any potential trajectory or intention conflicts with the vehicle in the future. Instead, the vehicle needs to unilaterally adjust its motion state to resolve any potential trajectory or intention conflicts with its non-game element. In other words, the non-game element does not establish a game-theoretic interaction relationship with the vehicle. That is, the non-game element will not be affected by the vehicle's driving behavior and will maintain its predetermined motion state without adjusting its motion state to resolve any potential trajectory or intention conflicts with the vehicle in the future.

[0095] The self-driving vehicle establishes a game-theoretic interaction relationship with its game partners. These game partners will respond to potential future trajectory or intention conflicts with the self-driving vehicle. At the start of the game decision, a trajectory or intention conflict exists between the self-driving vehicle and its game partners. During the game, both the self-driving vehicle and its game partners can adjust their respective states of motion to gradually resolve potential trajectory or intention conflicts while ensuring safety. When the self-driving vehicle, as a game partner, adjusts its state of motion, this can include automatic adjustments through its intelligent driving functions or manual adjustments by its driver.

[0096] To further understand the roles of game participants and non-game participants, the following will combine... Figures 3A-3E The diagrams illustrate several traffic scenarios, providing examples of the game players and non-game players for each vehicle.

[0097] like Figure 3A As shown, vehicle 101 proceeds straight through an unprotected intersection. Oncoming vehicle 102 (located to the left front of vehicle 101) turns left through the same unprotected intersection. At this point, there is a trajectory conflict or intention conflict between oncoming vehicle 102 and vehicle 101, and oncoming vehicle 102 becomes the game-playing partner of vehicle 101.

[0098] like Figure 3B Car 101 is traveling straight. Car 102 from the left crosses the lane where Car 101 is located and passes by. At this time, there is a trajectory conflict or intention conflict between Car 102 from the left and Car 101. Car 102 from the left is the game partner of Car 101.

[0099] like Figure 3C Vehicle 101 is traveling straight. Vehicle 102 traveling in the same direction (to the right front of vehicle 101) merges into vehicle 101's lane or its adjacent lane. At this point, vehicle 102 and vehicle 101 have a trajectory conflict or intention conflict, and will establish a game interaction relationship with vehicle 101. Vehicle 102 is the game object of vehicle 101.

[0100] like Figure 3D Vehicle 101 is traveling straight. Oncoming vehicle 103 is traveling straight in the lane adjacent to the left of vehicle 101. There is a stationary vehicle 102 in the lane adjacent to the right of vehicle 101 (located to the right front of vehicle 101). At this point, oncoming vehicle 103 and vehicle 101 have a trajectory conflict or intention conflict, and will establish a game-theoretic interaction relationship. Oncoming vehicle 103 is the game-playing partner of vehicle 101. The location of stationary vehicle 102 may conflict with the trajectory of vehicle 101 in the future. However, based on the acquired external environmental information, it can be confirmed that during the interactive game decision-making process, stationary vehicle 102 will not switch to a moving state, or even if it does switch to a moving state, it will have higher right-of-way and will not establish a game-theoretic interaction relationship with vehicle 101. Therefore, stationary vehicle 102 is a non-game-playing partner of vehicle 101, and vehicle 101 will adjust its driving behavior and movement state independently to resolve the trajectory conflict between the two.

[0101] like Figure 3EVehicle 101 changes lanes to the right from its current lane to merge into the adjacent lane on its right. In the adjacent lane on the right, there is a first straight-going vehicle 103 (to the right front of vehicle 101) and a second straight-going vehicle 102 (to the right rear of vehicle 101). The first straight-going vehicle 103 has higher right-of-way than vehicle 101 and will not establish a game-theoretic interaction with vehicle 101; therefore, it is a non-game-theoretic partner of vehicle 101. The second straight-going vehicle 102, located to the right rear of vehicle 101, may have a trajectory conflict with vehicle 101 in the future and will establish a game-theoretic interaction with vehicle 101; therefore, the second straight-going vehicle 102 is a game-theoretic partner of vehicle 101.

[0102] Below, for reference Figure 1 and Figure 2 and combined Figure 4 The flowchart shown illustrates the intelligent driving decision-making method provided in this application embodiment, including the following steps:

[0103] S10: Obtain the game object of the self-vehicle.

[0104] In some embodiments, such as Figure 5 The flowchart shown may include the following sub-steps:

[0105] S11: The vehicle acquires information about its external environment, including the motion state and relative position information of the vehicle and obstacles in the road scene.

[0106] In this embodiment, the vehicle acquires information about its external environment through its environmental information acquisition devices, such as camera sensors, infrared night vision camera sensors, lidar, millimeter-wave radar, ultrasonic radar, GNSS, etc. In some embodiments, the vehicle acquires this information by communicating with roadside devices or a cloud server. The roadside devices may have cameras or communication devices that can acquire information about surrounding vehicles, while the cloud server can receive and store information reported by various roadside devices. In some embodiments, a combination of these two methods can be used to acquire information about the vehicle's external environment.

[0107] S12: Based on the obtained motion state of the obstacle, or the motion state of the obstacle over a period of time, or the trajectory of the obstacle, and the relative position information with respect to the obstacle, the vehicle identifies the game object from the obstacle.

[0108] In this embodiment, in step S12, non-game objects of the vehicle can be identified from the obstacles at the same time, or non-game objects of the vehicle's game objects can be identified from the obstacles.

[0109] In some embodiments, the aforementioned game-playing or non-game-playing objects can be identified according to pre-set judgment rules. For example, if an obstacle's trajectory or intention conflicts with the vehicle's trajectory or intention, and the obstacle has the ability to make behavioral decisions and change its own motion state, then it is a game-playing object for the vehicle. If an obstacle's trajectory or intention conflicts with the vehicle's trajectory or intention, but it does not actively change its own motion state to avoid the conflict, then it is a non-game-playing object for the vehicle. In some embodiments, the obstacle's trajectory or intention can be determined based on its lane (straight or turning lane), whether its turn signal is on, and the direction the vehicle is facing.

[0110] For example Figures 3A-3C In the case of obstacles crossing the road or merging into the lane, the obstacles intersect the vehicle's trajectory at a large angle, thus causing a conflict with the vehicle's trajectory, and are therefore classified as "game-playing vehicles". Figure 3D Oncoming vehicles 103 and 104 passing through the narrow passage Figure 3E Vehicle 104, which merged into the adjacent lane, was considered to have a conflict of intent and was therefore classified as a game-theoretic vehicle. Figure 3D Narrow passage for stationary vehicles 102 on the right front Figure 3E If a vehicle 103 merges into the traffic flow in the adjacent lane, and although there is a conflict in trajectory or intention, the other vehicle will not take any action to resolve the conflict because the vehicle has a lower right-of-way relative to the other vehicle, and the vehicle's actions cannot change the other vehicle's actions. Therefore, the other vehicle is considered a non-game vehicle.

[0111] In some embodiments, the vehicle obtains obstacle information from the external environment information it perceives or acquires based on some known algorithms, and identifies the vehicle's game object, non-game object, or non-game object of the game object from the obstacles.

[0112] In some embodiments, the above algorithm can be, for example, a classification neural network based on deep learning. Since it identifies the type of obstacle, which is equivalent to classification, the classification model can be used for inference and determination. This classification neural network can employ Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Bidirectional Encoder Representations from Transformers (BERT), etc. When training the classification neural network, sample data can be used. The sample data can be images or video clips of vehicle driving scenes labeled with classification tags. The classification tags can include game objects, non-game objects, and non-game objects of game objects.

[0113] In some embodiments, the algorithm described above can also utilize attention-related algorithms, such as a modeled attention model. The attention model outputs the attention value assigned to the vehicle by each obstacle, which is related to the degree of intent conflict or trajectory conflict between the obstacle and the vehicle. For example, obstacles with intent or trajectory conflicts with the vehicle will assign more attention; obstacles without intent or trajectory conflicts will assign less or no attention; obstacles with higher right-of-way than the vehicle may also assign less or no attention. If an obstacle assigns enough attention to the vehicle (e.g., above a certain threshold), the obstacle can be identified as a game opponent. If an obstacle assigns enough attention to the vehicle (e.g., below a certain threshold), the obstacle can be identified as a non-game opponent.

[0114] In some embodiments, the attention model can be constructed using a mathematical model such as y = softmax(a1x1 + a2x2 + a3x3…), where softmax represents normalization, a1, a2, a3… are weight coefficients, and x1, x2, x3… are parameters related to the obstacle and the vehicle, such as longitudinal distance, lateral distance, speed difference, acceleration difference, and vehicle position (front, rear, left, right, etc.). x1, x2, x3… can also be normalized values, i.e., values ​​between 0 and 1. In some embodiments, the attention model can also be implemented using a neural network, where the output of the neural network is the attention value assigned to the vehicle corresponding to the identified obstacle.

[0115] Step S20: For the interactive game task between the vehicle and the game object, the vehicle performs multiple releases of the multiple strategy spaces between the vehicle and the game object. After one of the multiple releases is executed, the feasible domain of the strategies between the vehicle and the game object is determined according to the released strategy spaces, and the decision result of the vehicle's driving is determined according to the feasible domain of the strategies.

[0116] In some embodiments, the decision result refers to the executable action pairs between the vehicle and the game opponent within the strategy feasible domain. At this point, step S20 completes the single-vehicle interactive game decision-making process between the vehicle and any game opponent, and determines the strategy feasible domain between the vehicle and that game opponent. In some embodiments, the execution of multiple releases of the multiple strategy spaces includes executing the successive release of each strategy space, that is, releasing only one strategy space each time, and cumulatively releasing multiple strategy spaces after executing the successive releases. In this case, each strategy space is spanned by at least one sampling dimension.

[0117] In some embodiments, based on typical vehicle control safety or driving habits, the vehicle typically accelerates or decelerates within the current lane before changing lanes. Therefore, when releasing the aforementioned multiple strategy spaces, optionally, they can be released sequentially in the following dimensional order: longitudinal sampling dimension, lateral sampling dimension, and temporal sampling dimension. Different strategy spaces can be spanned by different dimensions. For example, releasing the longitudinal sampling dimension can span a longitudinal sampling strategy space, releasing the lateral sampling dimension can span a lateral sampling strategy space, or the longitudinal sampling dimension and the lateral sampling strategy space can be combined to form a strategy space. Releasing the temporal sampling dimension can form multiple strategy spaces composed of multiple frames of inference.

[0118] In some embodiments, the release can also be a sequential release of different combinations of space components of different strategy spaces. For example, during the first release, the local space of the vertical sampling strategy space and the local space of the horizontal sampling strategy space are released sequentially. During the second release, the remaining space of the vertical sampling strategy space and the remaining local space of the horizontal sampling strategy space are released sequentially.

[0119] In some embodiments, the cumulatively released multiple policy spaces may include: a vertical sampling policy space, a horizontal sampling policy space, a policy space spanned by the combination of the vertical sampling policy space and the horizontal sampling policy space, a policy space spanned by the combination of the vertical sampling policy space and the horizontal sampling policy space with the time sampling dimension, and a policy space spanned by the combination of the vertical sampling dimension, the horizontal sampling dimension, and the time sampling dimension.

[0120] In some embodiments, as described above, the dimensions constituting the policy space may include a longitudinal sampling dimension, a lateral sampling dimension, or a temporal sampling dimension. In the context of vehicle driving scenarios, this translates to the longitudinal acceleration dimension, the lateral offset dimension, and the deduction depth corresponding to the multiple single-frame deductions included in a single decision step. Correspondingly, the longitudinal sampling dimension used when constructing the longitudinal sampling policy space includes at least one of the following: the longitudinal acceleration of the vehicle and the longitudinal acceleration of the game object; the lateral sampling dimension used when constructing the lateral sampling policy space includes at least one of the following: the lateral offset of the vehicle and the lateral offset of the game object; the temporal sampling dimension includes multiple policy spaces composed of multiple consecutive frames of deduction corresponding to consecutive time points (i.e., successively increasing deduction depth). The combination of these three dimensions can constitute the aforementioned multiple policy spaces.

[0121] At this point, the values ​​taken in each horizontal or vertical sampling dimension when the strategy space is stretched correspond to the sampling actions of the vehicle or the game object, that is, the behavioral actions.

[0122] In some embodiments, such as Figure 6 The flowchart shown indicates that step S20 may include the following sub-steps S21-S26:

[0123] S21: Perform the first release of the strategy space, release the strategy space of the vehicle and the game object, and according to the released strategy space, extract the action pairs formed by multiple values ​​on at least one sampling dimension of the vehicle and multiple values ​​on at least one sampling dimension of the game object.

[0124] In some embodiments, when the first policy space is released, the released dimension is the longitudinal sampling dimension, including the longitudinal acceleration of the vehicle and the longitudinal acceleration of the game object. The policy space spanned by the released longitudinal sampling dimension is the longitudinal sampling policy space, which will be referred to as the first released longitudinal sampling policy space for simplicity. At this time, the policy space contains action pairs consisting of the longitudinal acceleration of the vehicle to be evaluated and the longitudinal acceleration of its game object (i.e., other vehicles). Multiple sampling values ​​can be set on each sampling dimension. These sampling values ​​can be composed of multiple sampling values ​​with uniform and continuous sampling intervals. Multiple sampling values ​​dispersed across the sampling dimensions are discrete sampling points. For example, by uniformly sampling at a predetermined sampling interval on the longitudinal acceleration dimension of the vehicle, multiple sampling values ​​of the vehicle on the longitudinal acceleration dimension can be obtained, denoted as M1, which are the M1 longitudinal acceleration sampling actions of the vehicle. By uniformly sampling along the longitudinal acceleration dimension of the game object at a predetermined sampling interval, multiple sampled values ​​of the game object in the longitudinal acceleration dimension can be obtained, denoted as N1, which are the N1 longitudinal acceleration sampling actions of the game object. The first released longitudinal sampling strategy space then includes M1*N1 action pairs between the game object and its own vehicle, obtained by combining the longitudinal acceleration sampling actions of the game object with those of the player vehicle. A specific example of this strategy space can be seen in Table 1 or Table 2 below. The first row and first column of Table 1 or Table 2 are the longitudinal acceleration sampling values ​​of the game object and its player vehicle (i.e., the other vehicle O in the table), respectively. In Table 1, the player vehicle is the vehicle crossing the lane; in Table 2, the player vehicle is the vehicle facing the opponent.

[0125] S22: Extrapolate each action pair in the released strategy space to the sub-traffic scenario currently constructed by the vehicle and the game object, and determine the corresponding cost of each action pair.

[0126] At this point, the vehicle and each other in the game can construct their own sub-traffic scenarios, each of which is a subset of the road scenario in which the vehicle is located.

[0127] In some embodiments, the cost value corresponding to each action pair in the strategy space is determined based on at least one of the following: the safety cost value, comfort cost value, lateral offset cost value, passability cost value, right-of-way cost value, risk area cost value, and inter-frame correlation cost value corresponding to the action pair performed by the vehicle and the game player.

[0128] In some embodiments, a weighted sum of the aforementioned cost values ​​can be used. In this case, to distinguish between the individual cost values, the calculated weighted sum can be called the total cost value. The smaller the total cost value, the greater the decision benefit for the vehicle and the game opponent to perform the action, and the greater the likelihood that the action will be the decision outcome. The aforementioned cost values ​​will be further described later.

[0129] S23: Add the action pairs with a cost value not greater than the cost threshold to the policy feasible region of the vehicle and the game object. The policy feasible region is the game result of the vehicle and the game object when the policy space is released for the first time.

[0130] In this context, the policy feasible region refers to the set of action pairs that can be executed. For example, in Table 1 below, the table items with the content Cy or Cg constitute the policy feasible region.

[0131] S24: When the feasible domain of the current strategy space (i.e. the game outcome) is not empty, at least one executable action pair between the vehicle and the game object in the feasible domain can be taken as the decision result of the vehicle and the game object, and the release of the current strategy space ends.

[0132] When the feasible domain of the strategy is empty, it means that there is no solution in the current strategy space. At this time, the second release of the strategy space is performed, that is, the next strategy space among multiple strategy spaces is released. In this embodiment, the second release is the horizontal sampling dimension, so as to form a horizontal sampling strategy space by horizontal offset. The horizontal sampling strategy space released this time and the vertical sampling strategy space released in the first time are used together as the current strategy space. At this time, the strategy space currently used for the interaction between the vehicle and the game object is the strategy space formed by the combination of the vertical sampling strategy space and the horizontal sampling strategy space.

[0133] The lateral sampling strategy space spans along both the lateral offset dimension of the vehicle and the lateral offset dimension of the game opponent. For example, uniform sampling at predetermined intervals along the lateral offset dimension of the vehicle yields multiple sampled values, denoted as Q, representing the vehicle's Q lateral offset sampling actions. Similarly, uniform sampling at predetermined intervals along the lateral offset dimension of the game opponent yields multiple sampled values, denoted as R, representing the game opponent's R lateral offset sampling actions.

[0134] At this point, in the strategy space currently used for the interaction between the vehicle and the game opponent, each action pair between the vehicle and the game opponent is composed of the vehicle's lateral offset sampling action, the game opponent's lateral offset sampling action, the vehicle's longitudinal acceleration sampling action, and the game opponent's longitudinal acceleration sampling action.

[0135] Assuming the current strategy space consists of Q values ​​for the lateral offset of the vehicle, R values ​​for the lateral offset of the player, M² values ​​for the longitudinal acceleration of the vehicle, and N² values ​​for the longitudinal acceleration of the player, the strategy space of the second released vehicle and player includes M²*N²*Q*R action pairs. A concrete example can be seen in Table 3 below, where each table entry in the upper lateral sampling strategy space is associated with a table entry in the lower longitudinal sampling strategy space. In the upper lateral sampling strategy space table in Table 3, the player (i.e., the other vehicle O in the table) is the opposing vehicle.

[0136] S25: After the second release of the execution strategy space, based on the current strategy space, the released action pairs are extrapolated to the sub-traffic scenario currently constructed by the vehicle and the game object, the cost value corresponding to each action pair is determined, and then the feasible domain of the strategy is determined to determine the game result. This step can be referred to in steps S22-S23.

[0137] S26: If the feasible domain of the strategy (i.e. the game outcome) in step S25 is not empty, then an action pair can be selected as the decision outcome, and the release of the strategy space ends.

[0138] When the feasible region of the strategy is empty, it indicates that there is no solution in the current strategy space. At this time, the third release of the strategy space is performed, that is, the next strategy space among multiple strategy spaces is released. In this way, other strategy spaces can continue to be released in sequence to continue to determine the game result and the decision result.

[0139] In some embodiments, the strategy space released multiple times can be achieved by first releasing the strategy space spanned by the i-th set of values ​​in the longitudinal acceleration dimension and the i-th set of values ​​in the lateral offset dimension of the vehicle and / or the game object. If no feasible strategy domain exists within this strategy space, then the strategy space spanned by the (i+1)-th set of values ​​in the longitudinal acceleration dimension and the (i+1)-th set of values ​​in the lateral offset dimension of the vehicle and / or the game object is released. That is, the strategy space released multiple times moves the local position of the vehicle and / or the game object within the game space spanned by all values ​​in the longitudinal acceleration dimension and all values ​​in the lateral offset dimension, where i is a positive integer. The above example, using the i-th set of values, illustrates the sequential release of partial values ​​in each sampling dimension to sequentially search for feasible strategy domains and determine decision results within different local strategy spaces of the game space. By releasing the policy space corresponding to different localities in this way, the optimal decision result can be obtained in the multi-dimensional game space with the fewest number of searches, and the use of policy space can be minimized, thus reducing the requirements for hardware computing power.

[0140] For example, first release the strategy space spanned by the lateral offset value 0 of the opposing rook, the lateral offset value 1 of the own rook, and all longitudinal acceleration values ​​of the own rook and the opposing rook in Table 3; and if there is no feasible strategy region in this strategy space, then release the strategy space spanned by the lateral offset value 0 of the opposing rook, the lateral offset value 2 or 3 of the own rook, and all longitudinal acceleration values ​​of the own rook and the opposing rook.

[0141] In some embodiments, if the policy feasible domain of the vehicle and the game opponent is still empty after the above steps are performed multiple times to release the policy space, that is, there is still no solution, then a conservative decision can be made for the vehicle to drive. The conservative decision includes actions that make the vehicle stop safely, actions that make the vehicle decelerate safely, or giving prompts or warnings so that the driver can take over control of the vehicle.

[0142] The above steps S10-S20 complete one single-frame deduction. In some embodiments, if the policy feasible region is not empty after steps S10-S20, the process may further include: performing multiple releases in the time sampling dimension according to the development of the deduction time (i.e., sequentially increasing the deduction depth, for multiple consecutive moments), and performing multi-frame deduction. In this case, after a release is performed at a deduction moment (or time point) during the multiple releases, one single-frame deduction is completed to determine the policy feasible region of the vehicle and the game object. If the policy feasible region of the vehicle and the game object determined in this single-frame deduction is not empty, the release at the next deduced moment is performed to execute the next single-frame deduction, until the multiple releases in the time sampling dimension end or the continuous multi-frame deduction ends.

[0143] At this point, multiple single-frame deductions within a single decision-making step are achieved. For example... Figure 7 As shown, T1 is used to indicate the initial motion state of the vehicle and the game object, T2 is used to indicate the motion state of the vehicle and the game object after the first frame of deduction, that is, the deduction result of the first frame, and Tn is used to indicate the motion state of the vehicle and the game object after the (n-1)th frame of deduction.

[0144] In some embodiments, each time the release time sampling dimension is executed, the deduction time is shifted backward by a predetermined time interval (e.g., 2 seconds or 5 seconds), i.e., moved to the next moment (or time point). Accordingly, the deduction result of the current frame is used as the initial condition for the deduction of the next frame to deduce the motion state of the vehicle and the game object at the next moment; in this way, the time sampling dimension can continue to be released in subsequent moments to continue to execute the deduction of subsequent consecutive frames, so as to continue to determine the game result and the decision result.

[0145] In the release time sampling dimension, it is necessary to evaluate the decision-making results of the vehicle and the game object's behavior determined by the extrapolation of two adjacent single frames, and determine the cost of inter-frame correlation, which will be explained in detail later. The release time sampling dimension helps improve the consistency of vehicle behavior. For example, when the intention decisions corresponding to the motion states or decision results of the vehicle and the game object in extrapolation of multiple consecutive frames are the same or similar, the driving behavior of the intelligent driving vehicle executing the intelligent driving decision-making method is more stable in the time domain, the driving trajectory has less fluctuation, and the vehicle driving comfort is better.

[0146] The aforementioned release time sampling dimension, that is, at multiple consecutive decision moments, utilizes multiple released strategy spaces to obtain the policy feasible regions corresponding to the vehicle and the game opponent, and deduces the motion state of the vehicle and the game opponent after they sequentially execute the executable action pairs corresponding to these policy feasible regions in chronological order. By performing long-term deductions on the motion states and / or expected driving trajectories of the vehicle and the game opponent, temporal consistency of decision results can be achieved.

[0147] After the above multi-frame simulation is completed, if the overall payoff of the multi-frame simulation meets the decision requirements, then it can be determined that the game results of each frame gradually converge to the Nash equilibrium state. At this point, the decision result of the first frame in the multi-frame simulation can be used as the decision result for the vehicle's driving.

[0148] In some embodiments, if the overall benefit of multi-frame deduction does not meet the decision requirements, the decision result of the first frame can be reselected. The decision result of the first frame corresponds to the decision result of the single-frame deduction. Reselecting the decision result of the first frame means selecting another action pair from the policy feasible region of the first frame's decision result as the final decision result. For the reselected decision result, multi-frame deduction can be performed again to determine if it can be used as the final decision result.

[0149] In some embodiments, when making a decision on the first frame and when reselecting or reselecting the first frame multiple times, the decision can be made according to the ranking of the cost values ​​of each action, prioritizing the action with the lowest total cost value.

[0150] In some embodiments, the different cost values ​​mentioned above can have different weights, which can be respectively referred to as safety weight, comfort weight, lateral offset weight, throughput weight, right-of-way weight, risk area weight, and inter-frame correlation weight. Furthermore, in some embodiments, the weight allocation can be as follows: safety weight > right-of-way weight > lateral offset weight > throughput weight > comfort weight > risk area weight > inter-frame correlation weight. In some embodiments, the above cost values ​​can be normalized, with a value range of [0,1].

[0151] In some embodiments, the aforementioned costs can be calculated based on different cost functions, which can be referred to as safety cost function, comfort cost function, passability cost function, lateral offset cost function, and right-of-way cost function, respectively.

[0152] In some embodiments, the safety cost can be calculated based on a safety cost function with the relative distance between the vehicle and other vehicles (i.e., the game partners) as the independent variable, and the safety cost is negatively correlated with the relative distance. For example, the greater the relative distance between the two vehicles, the smaller the safety cost. Figure 8A As shown, a security cost function after homogenization is a piecewise function as follows, where C dist For the safety cost, dist is the relative distance between the vehicle and the game object. For example, the minimum distance is defined as the minimum polygon distance between the vehicle and the game object:

[0153]

[0154] Where, threLow is the lower limit threshold of the distance, such as Figure 8A The value is 0.2, and `threHigh` is the upper limit threshold for distance, such as... Figure 8A The value is 1.2. Optionally, the lower distance threshold (threLow) and the upper distance threshold (threHigh) can be dynamically adjusted according to the interaction between the vehicle and other vehicles, such as the relative speed, relative distance, and relative angle between the vehicle and other vehicles.

[0155] In some embodiments, the safety cost value defined by the safety cost function is positively correlated with relative speed or relative angle. For example, when two vehicles meet in opposite or lateral directions (lateral means that another vehicle is crossing over the vehicle), the greater the relative speed or relative angle between the two vehicles, the greater the corresponding safety cost value.

[0156] In some embodiments, the comfort cost of a vehicle (either the vehicle itself or the game object) can be calculated using a comfort cost function with the absolute value of the change in acceleration (i.e., jerk) as the independent variable. For example... Figure 8BAs shown, a uniformized comfort cost function is a piecewise function as follows, where C comf For the value of comfort, jerk represents the change in acceleration of the vehicle or the game player.

[0157]

[0158] Where, threMiddle is the jerk threshold for the middle point, such as... Figure 8B In the example, 2 is used, where threHigh is the upper limit threshold of jerk, such as... Figure 8B The value is 4. C middle Let represent the jerk cost slope. That is, the greater the change in vehicle acceleration, the worse the comfort and the greater the cost of comfort. Furthermore, once the change in vehicle acceleration exceeds the midpoint threshold, the cost of comfort increases even faster.

[0159] In some embodiments, the change in vehicle acceleration may be the change in longitudinal acceleration, the change in lateral acceleration, or a weighted sum of the two. In some embodiments, the comfort cost may be the comfort cost of the vehicle itself, the comfort cost of the game object, or a weighted sum of the comfort costs of the two.

[0160] In some embodiments, the passability cost can be calculated based on a passability cost function with the change in speed of the vehicle or the game player as the independent variable. For example, if a vehicle yields with a larger deceleration, resulting in a larger speed loss (the difference between the current speed and the future speed, i.e., acceleration) or a longer waiting time, then the vehicle's passability cost increases. Conversely, if a vehicle rushes past with a larger acceleration, resulting in a larger speed increase (the difference between the current speed and the future speed, i.e., acceleration) or a shorter waiting time, then the vehicle's passability cost decreases.

[0161] In some embodiments, the passability cost can also be calculated based on a passability cost function with the relative speed ratio of the vehicle and the game opponent as the independent variable. For example, before executing the action pair, the absolute speed of the vehicle accounts for a larger proportion of the sum of the absolute speeds of the vehicle and the game opponent, while the absolute speed of the game opponent accounts for a smaller proportion. After executing the action pair, if the vehicle yields with a large deceleration, its speed loss increases, its speed ratio decreases, and therefore the passability cost corresponding to the vehicle executing this action is larger. Conversely, if after executing the action pair, the game opponent rushes past with a large acceleration, its speed increases, its speed ratio increases, and therefore the passability cost corresponding to the game opponent executing this action is smaller.

[0162] In some embodiments, the passability cost is the passability cost of the vehicle performing the action for the corresponding vehicle, or the passability cost of the game object performing the action for the corresponding game object, or the weighted sum of the passability costs of both.

[0163] In some embodiments, such as Figure 8C As shown, the passability cost function after homogenization is a piecewise function as follows, where C pass For the value of passability, speed is the absolute value of the vehicle's speed:

[0164]

[0165] Wherein, the absolute value of the vehicle's velocity at its midpoint is speed0, and the maximum absolute value of the vehicle's velocity is speed1. middle This represents the slope of the speed cost. In other words, the greater the absolute value of the vehicle's speed, the better the passability and the smaller the passability cost. Furthermore, once the absolute value of the vehicle's speed exceeds the midpoint threshold, the passability cost decreases even faster.

[0166] In some embodiments, the right-of-way information corresponding to a vehicle can be determined based on the obtained user profile of the vehicle or the game player. For example, if the game player's driving behavior is aggressive and they tend to make overtaking decisions, they have high right-of-way; if the game player's driving behavior is conservative and they tend to use yielding strategies, they have low right-of-way. High right-of-way tends to maintain a fixed state of motion or driving behavior, while low right-of-way tends to change the fixed state of motion or driving behavior.

[0167] In some embodiments, user profiles can be determined based on a user's gender, age, or completion status of historical behavioral actions. In some embodiments, the data required to determine the user profile can be obtained and the user profile can be determined by a cloud server. If the actions performed by the vehicle and / or the game player cause a change in the motion state of the vehicle with higher right-of-way, then the action has a higher cost to the corresponding right-of-way and a lower benefit.

[0168] In some embodiments, a higher right-of-way cost is assigned to the behavioral decision that causes the high-right-of-way vehicle to change its motion state, thereby increasing the penalty. That is, through this feedback mechanism, the behavioral action of the high-right-of-way vehicle to maintain its current motion state has a greater right-of-way benefit, i.e., a smaller right-of-way cost.

[0169] In some embodiments, such as Figure 8D As shown, the normalized right-of-way cost function is a piecewise function as follows, where C roadRight Where acc is the absolute value of the vehicle's acceleration, and right-of-way is the value of the right-of-way.

[0170]

[0171] Where, threHigh is the upper limit threshold for acceleration, such as... Figure 8D The value is 1. That is, the greater the vehicle's acceleration, the greater the value of the right-of-way.

[0172] In other words, the right-of-way cost function makes the actions of high-right-way vehicles in maintaining their current state of motion have a smaller right-of-way cost value, thereby preventing the actions of high-right-way vehicles in changing their current state of motion from becoming the decision outcome.

[0173] In some embodiments, the vehicle's acceleration can be longitudinal or lateral. That is, a larger lateral change in the lateral offset dimension will result in a larger right-of-way cost. In some embodiments, the right-of-way cost can be the cost of the vehicle performing the action, the cost of the other player performing the action, or a weighted sum of the two.

[0174] In some embodiments, such as when a vehicle is in a risk zone within the road (where the vehicle faces significant driving risks and needs to leave the risk zone as soon as possible), a higher risk zone cost needs to be applied to the vehicle yielding strategy to increase the penalty. By choosing to overtake instead of yielding as the decision outcome, the vehicle is made to leave the risk zone as soon as possible. In other words, the decision to make the vehicle leave the risk zone as soon as possible is made to ensure that the vehicle leaves the risk zone as soon as possible without causing serious impact on traffic.

[0175] In other words, through the feedback mechanism that the greater the risk area's cost value, the smaller the strategy's benefit, vehicles in the risk area within the road do not yield. That is, they abandon the behavioral decision that would cause vehicles in the risk area within the road to yield (this behavioral decision has a large risk area cost value), and instead choose the decision of vehicles in the risk area within the road to rush out of the risk area as quickly as possible (which has a small risk area cost value), thereby avoiding vehicles in the risk area within the road from being stuck and having a serious impact on traffic.

[0176] In some embodiments, the risk area cost can be the cost of the action performed by the vehicle within the road risk area, or the cost of the action performed by the game player within the road risk area, or a weighted sum of the risk area costs of both. In some embodiments, the lateral offset cost can be calculated based on the lateral offset of the vehicle or the game player. Figure 8E As shown, the lateral offset cost function after homogenization is a piecewise function in the right half space as follows, where C offsetLet be the cost of the lateral offset, and offset be the lateral offset of the vehicle, in meters. Then, the equation for the left half-space can be obtained by inverting the equation for the right half-space of the coordinate plane:

[0177]

[0178] Where, threMiddle is the middle value of the lateral offset, e.g., represents the soft boundary of the road; C middle Here, is the slope of the first lateral offset cost; ...

[0179] In some embodiments, the lateral offset cost can be the lateral offset cost of the vehicle performing the action, or the lateral offset cost of the game player performing the action, or a weighted sum of the lateral offset costs of both.

[0180] In the aforementioned multi-frame deduction steps, it is necessary to evaluate the decision-making results of the vehicle and the game object determined by the deduction of two adjacent single frames, and determine the inter-frame correlation value.

[0181] In some embodiments, such as Figure 8F As shown, if the vehicle's intention decision in the previous frame K is to be the object of the preemptive game, then if the vehicle's intention decision in the current frame K+1 is to be the object of the preemptive game, the corresponding inter-frame correlation cost will be relatively small, such as 0.3, while the default value is 0.5, thus it is a reward. However, if the vehicle's intention decision in the current frame K+1 is to be the object of the yielding game, the corresponding inter-frame correlation cost will be relatively large, such as 0.8, while the default value is 0.5, thus it is a penalty. Therefore, choosing a strategy that makes the vehicle's intention decision in the current frame the object of the preemptive game becomes a feasible solution for the current frame. Through the above penalty or reward for the inter-frame correlation cost, it can be ensured that the vehicle's intention decision in the current frame is consistent with the intention decision in the previous frame, thus ensuring that the vehicle's motion state in the current frame is consistent with the motion state in the previous frame, stabilizing the vehicle's behavior decision in the time domain.

[0182] Inter-frame correlation value: In some embodiments, the inter-frame correlation value can be calculated based on the intention decision of the vehicle in the previous frame and the intention decision of the vehicle in the current frame, or it can be calculated based on the intention decision of the game object in the previous frame and the intention decision of the game object in the current frame, or it can be obtained by weighting the vehicle and the game object.

[0183] In some embodiments, such as Figure 4 As shown, after determining the decision result in step S20, the following steps S30 and / or S40 may also be included:

[0184] Step S30: Based on the decision result, the vehicle generates longitudinal / lateral control quantities, which are then executed by the vehicle's driving system to achieve the vehicle's desired driving trajectory.

[0185] In some embodiments, the vehicle's control device generates longitudinal / lateral control quantities based on the decision results and sends the longitudinal / lateral control quantities to the driving system 13, so that the driving system 13 can perform driving control of the vehicle, including power control, steering control and braking control, so that the vehicle can execute the desired driving trajectory of the vehicle according to the decision results.

[0186] Step S40: Display the decision result on the display device in a way that is understandable to the user.

[0187] The decision-making results for vehicle driving include the vehicle's actions. Based on the vehicle's actions and the vehicle's motion state acquired at the start of the decision-making process in the current frame, the vehicle's intended decision, such as cutting in, yielding, or swerving, can be predicted, as can the vehicle's desired driving trajectory. In some embodiments, the decision results are displayed on a display device within the vehicle's cabin in a user-understandable manner, such as the desired driving trajectory, arrows indicating the decision, or text indicating the decision. In some embodiments, when displaying the desired driving trajectory, the current traffic scene (such as a graphical traffic scene) can be combined and displayed on the display device within the vehicle's cabin in a partially magnified view. In some embodiments, a voice playback system may also be included, which can prompt the user with the decided intention or strategy label by playing voice prompts.

[0188] In some embodiments, considering one-way interactive decision-making between the vehicle and its non-game objects, or one-way interactive decision-making between the vehicle's game objects and its non-game objects, such as... Figure 9 In another embodiment shown, the following step is further included between step S10 and step S20:

[0189] S15: Constrain the strategy space of the vehicle or the game object by the motion state of the non-game object.

[0190] In some embodiments, the range of values ​​for the vehicle in each sampling dimension can be constrained based on the motion state of the non-game object of the vehicle.

[0191] In some embodiments, the range of values ​​for the game object of the vehicle in each sampling dimension can be constrained based on the motion state of the non-game object of the game object.

[0192] In some embodiments, the range of values ​​can be one or more sampling intervals along the sampling dimension, or it can be multiple discrete sampling points. The range of values ​​after constraint can be a partial range of values.

[0193] In some embodiments, step S15 includes: determining the range of values ​​for the vehicle in each sampling dimension under the constraint of the motion state of the non-game object in the one-way interaction process between the vehicle and its non-game object; or determining the range of values ​​for the vehicle's game object in each sampling dimension under the constraint of the motion state of the non-game object of the vehicle's game object in the one-way interaction process between the vehicle's game object and the non-game object of the vehicle's game object.

[0194] Since non-game objects do not participate in interactive games and their motion states remain unchanged, constraining the range of values ​​of the self-vehicle in each sampling dimension by the motion states of the self-vehicle's game objects, and constraining the range of values ​​of the self-vehicle's game objects in each sampling dimension by the motion states of the non-game objects of the self-vehicle's game objects, before proceeding to step S20, helps to reduce the game space and strategy space in the single-vehicle interactive game decision-making process between the self-vehicle and its game objects, and reduces the computing power used in the interactive game decision-making process.

[0195] In some embodiments, when constraining the vehicle, step S15 may include: first, receiving the motion state of a non-game object of the vehicle and observing the characteristic quantity of the non-game object, such as; then, calculating the conflict region between the vehicle and the non-game object, and determining the characteristic quantity of the vehicle, i.e., the critical action. As above, based on the position, speed, acceleration and / or driving trajectory of the non-game object, the corresponding critical action for the vehicle to make an intentional decision such as avoiding, overtaking or yielding is calculated, and the feasible interval of the vehicle for the non-game object in each sampling dimension is generated, i.e., the value range of the vehicle after being constrained by the non-game object in each sampling dimension.

[0196] In some embodiments, constraints can also be applied to the game object of the vehicle itself. Instead of applying constraints to the vehicle itself as described above, the vehicle itself can be changed to the game object of the vehicle itself to generate the range of values ​​of the game object of the vehicle itself in each sampling dimension after being constrained by non-game object constraints.

[0197] In some embodiments, when there is a non-game object C for the vehicle, the following steps can be taken: First, process the interactive game decision between vehicle A and its game object B to determine the corresponding strategy feasible region AB. Then, introduce the non-game feasible region AC between the vehicle and the non-game object C. Then, take the intersection of the strategy feasible region AB and the non-game feasible region AC to obtain the final strategy feasible region ABC. Based on the final strategy feasible region, determine the decision result for the vehicle's driving.

[0198] In some embodiments, when the game object B of the vehicle exists as a non-game object D, the interactive game decision between the vehicle A and the game object B of the vehicle can be processed first to determine the corresponding strategy feasible region AB. Then, the non-game feasible region BD of the game object B of the vehicle and its non-game object D is introduced. Then, the intersection of the strategy feasible region AB and the non-game feasible region BD is taken to obtain the final feasible region ABD. The decision result for the vehicle's driving is determined based on the final strategy feasible region.

[0199] In some embodiments, when the vehicle has a non-game object C and the vehicle's game object B has a non-game object D, the interactive game decision between the vehicle A and its game object B can be processed first to determine the corresponding strategy feasible region AB. Then, the non-game feasible region AC between the vehicle A and the non-game object C, and the non-game feasible region BD between the vehicle's game object B and its non-game object D are introduced. Then, the intersection of the strategy feasible region AB, the non-game feasible region AC, and the non-game feasible region BD is taken to obtain the final feasible region ABCD. The decision result for the vehicle's driving is determined based on the final strategy feasible region.

[0200] The above examples illustrate the steps of determining executable action pairs between the vehicle and a single game opponent by successively releasing multiple policy spaces of the vehicle and the game opponent. In some embodiments, when the vehicle has two or more game opponents, for example, taking two game opponents, including a first game opponent and a second game opponent, the intelligent driving decision-making method provided in this application includes:

[0201] Step 1: From the multiple strategy spaces of the self-vehicle and the first game opponent, execute the successive release of each strategy space to determine the feasible strategy domain for the self-vehicle's movement against the first game opponent. The determination of the feasible strategy domain for the self-vehicle's movement against the first game opponent is similar to the aforementioned step S20, and will not be repeated here.

[0202] Step 2: From the multiple strategy spaces of the self-vehicle and the second game opponent, execute the successive release of each strategy space to determine the feasible strategy domain for the self-vehicle's movement against the second game opponent. Determining the feasible strategy domain for the self-vehicle's movement against the second game opponent is similar to the aforementioned step S20, and will not be repeated here.

[0203] Step 3: Determine the decision outcome for the vehicle's driving based on the feasible regions of each strategy of the vehicle and each of the game's opponents. In some embodiments, the final feasible region is obtained by taking the intersection of the feasible regions, and the decision outcome is then determined from this feasible region. In some embodiments, the decision outcome may be the action pair with the lowest cost in the feasible region.

[0204] The following describes a specific implementation of the intelligent driving decision-making method provided in this application. This specific implementation will still be illustrated using a traffic scenario applied to vehicle driving on a road as an example, such as... Figure 10 The scenario illustrated in this specific embodiment is as follows: Vehicle 101 is traveling on a highway, which is a two-way single-lane road. Oncoming vehicle 103, i.e., an oncoming game vehicle, is approaching. A vehicle 102, i.e., a crossing game vehicle, is positioned in front of Vehicle 101 and is about to cross the highway. See below. Figure 11 The flowchart shown describes in detail the driving control method provided in the specific embodiments of this application, including the following steps:

[0205] S110: The vehicle obtains external environmental information through the environmental information acquisition device 11.

[0206] This step is the same as step S11 mentioned above, and will not be repeated here.

[0207] S120: The vehicle identifies the game partner and the non-game partner.

[0208] This step can be referred to in step S12 above, and will not be repeated here. In this step, one game object from the car is determined to be the cross-traversing car, and the other game object is the opposing car.

[0209] S130: From the multiple strategy spaces of the self-car and the crossing car, execute the successive release of each strategy space, and determine the game outcome of the self-car and the crossing car. Specifically, this may include the following steps S131-S132:

[0210] S131: Following the principle of releasing the longitudinal sampling dimension first and then the lateral sampling dimension, release the longitudinal acceleration dimension of the vehicle and the vehicle crossing the game.

[0211] From the multidimensional game space of the self-driving car and the crossing car, the longitudinal acceleration dimension of the self-driving car and the crossing car is released to form the first longitudinal sampling strategy space of the self-driving car and the crossing car. Considering the longitudinal / lateral dynamics, kinematic constraints, relative position and relative velocity relationships of the self-driving car and the crossing car, and assuming that the two cars have the same maneuverability, the longitudinal acceleration values ​​of the self-driving car and the crossing car are determined to be in the range of [-4, 3], with units of m / s. 2 Where m represents meters and s represents seconds. Based on the vehicle's computing power and the pre-set decision accuracy, the sampling interval for both the vehicle and the traversing game vehicle is determined to be 1 m / s. 2 .

[0212] Table 1. Vertical sampling strategy space of the first released self-car and cross-traversing car.

[0213]

[0214] Zhang Cheng's strategy space, when displayed as a two-dimensional table, is shown in Table 1. The first row of Table 1 lists all values ​​of the longitudinal acceleration of the self-car (Ae), and the first column lists all values ​​of the longitudinal acceleration of the traversing car (Ao1). That is, the longitudinal sampling strategy space of the self-car and the traversing car released this time includes 8 by 8, or 64, pairs of longitudinal acceleration action actions of the self-car and the traversing car.

[0215] S132: Based on the predefined methods for determining the value of each generation, such as each cost function, calculate the value of each action pair in the longitudinal sampling strategy space of the self-car and the cross-traversing car, and determine the policy feasible region.

[0216] Of the 64 action pairs in Table 1, after the self-driving car and the cross-traffic vehicle performed 9 of the sampled actions, the passability in the sub-traffic scenario constructed by the self-driving car and the cross-traffic vehicle was too poor (such as braking), and they were considered infeasible solutions. In Table 1, these action pairs are marked with the label "0".

[0217] Of the 64 action pairs released, after the self-driving car and the cross-traffic vehicle performed 39 of the sampled actions, the safety was too poor (such as collision) in the sub-traffic scenario constructed by the self-driving car and the cross-traffic vehicle, which is an infeasible solution. In Table 1, these action pairs are marked with the label "-1".

[0218] Of the 64 action pairs released, after the self-driving vehicle and the traversing game vehicle execute 3 plus 13, or 16 sampling actions, in the sub-traffic scenario constructed by the self-driving vehicle and the traversing game vehicle, the weighted sum of the safety cost, comfort cost, passability cost, lateral offset cost, right-of-way cost, risk area cost, and inter-frame correlation cost is greater than the pre-set cost threshold. This constitutes a feasible solution in the strategy space and forms the strategy feasible region of the self-driving vehicle and the traversing game vehicle.

[0219] Since the vertical sampling strategy space is being released, and no horizontal offset is involved, the horizontal offset cost is zero. Because the decision is made in the current frame and does not involve the decision result of the previous frame, the inter-frame correlation cost is zero.

[0220] At this point, the interaction between the self-driving car and the cross-traversing car has found a sufficiently large number of feasible solutions in the vertical sampling strategy space, and it is no longer necessary to continue searching for solutions in the horizontal sampling dimension. At this point, the number of search actions is 64, and this round of the game has consumed relatively little computing power and computation time.

[0221] In addition, decision labels can be added to each action pair based on the cost value corresponding to each action pair within the feasible domain of the strategy.

[0222] After the vehicle and the crossing vehicle perform three of the sampled actions, their behavioral decisions are: the vehicle accelerates while the crossing vehicle decelerates. Based on the motion states of the vehicle and the crossing vehicle obtained at the start of the decision-making process in the current frame, it can be deduced that after performing any one of these three sampled actions, the vehicle passes through the conflict area before the crossing vehicle. Therefore, the intention decision corresponding to these action pairs is determined to be the vehicle overtaking the crossing vehicle. Accordingly, a "overtaking decision" label is assigned to these three action pairs, namely "Cg" in Table 1.

[0223] After the vehicle and the vehicle crossing the game execute 13 sampled actions, their behavioral decisions are: the vehicle crossing the game accelerates, while the vehicle decelerates. Based on the motion states of the vehicle and the vehicle crossing the game at the start of the decision-making process in the current frame, it can be deduced that after any one of these 13 sampled actions, the vehicle crossing the game passes through the conflict area before the vehicle. Therefore, the intention decision corresponding to these action pairs is determined to be that the vehicle crossing the game overtakes the vehicle. Accordingly, a "Cy" label is assigned to these 13 action pairs as the "Cy" decision for the vehicle crossing the game.

[0224] S140: From the multiple strategy spaces of the player's own vehicle and the opposing vehicle, execute the successive release of each strategy space, and determine the game outcome between the player's own vehicle and the opposing vehicle. Specifically, this may include the following steps S141-S144:

[0225] S141: Following the principle of releasing the vertical sampling dimension first and then the horizontal sampling dimension, release the vertical sampling dimension of the self-car and the opposing car, forming the first vertical sampling strategy space of the self-car and the opposing car.

[0226] From the multidimensional game space of the self-vehicle and the opposing vehicle, the longitudinal acceleration dimension of the self-vehicle and the opposing vehicle is released, forming the first longitudinal sampling strategy space of the self-vehicle and the opposing vehicle. Considering the longitudinal / lateral dynamics, kinematic constraints, relative positional relationships, and relative velocity relationships of the self-vehicle and the opposing vehicle, the longitudinal acceleration values ​​of the self-vehicle and the opposing vehicle are determined to be in the range of [-4, 3], with units of m / s. 2 The sampling interval for both the self-driving vehicle and the opposing vehicle in the game was determined to be 1 m / s. 2 .

[0227] Zhang Cheng's strategy space is presented in a two-dimensional table as shown in Table 2. The first row of Table 2 lists all possible longitudinal acceleration values ​​for the self-vehicle, Ae; the first column lists all possible longitudinal acceleration values ​​for the opposing vehicle, Ao2. That is, the longitudinal sampling strategy space of the self-vehicle and the opposing vehicle released this time includes 8 by 8, a total of 64 pairs of longitudinal acceleration action pairs for the self-vehicle and the opposing vehicle.

[0228] Table 2. Vertical sampling strategy space of released self-player and opposing game vehicular vehicles

[0229]

[0230] S142: Based on the predefined methods for determining the value of each generation, such as each cost function, calculate the value of each action pair in the longitudinal sampling strategy space of the self-vehicle and the game-playing vehicle, and determine the policy feasible region.

[0231] Of the 64 action pairs released, after the self-vehicle and the opposing game vehicle executed 9 of the sampled actions, the passability in the sub-traffic scenario constructed by the self-vehicle and the opposing game vehicle was too poor (such as braking), which is an infeasible solution. In Table 2, these action pairs are marked with the label "0".

[0232] Of the 64 action pairs released, after the autonomous vehicle and the opposing vehicle performed 55 of the sampled actions, the safety was too poor (e.g., collision) in the sub-traffic scenario constructed by the autonomous vehicle and the opposing vehicle, making it an infeasible solution. In Table 2, these action pairs are marked with the label "-1".

[0233] That is, after the self-vehicle and the opposing vehicle execute the 64 action pairs released, in the sub-traffic scenario constructed by the self-vehicle and the opposing vehicle, the safety cost or the passability cost is greater than the pre-set cost threshold. There is no feasible solution in the strategy space of the first release, and the strategy feasible domain of the self-vehicle and the opposing vehicle is empty.

[0234] S143: Release the lateral offset dimension of the self-vehicle, which spans the longitudinal acceleration dimension of the self-vehicle and the opposing game vehicle to form a second strategy space for the self-vehicle and the opposing game vehicle.

[0235] Specifically, by releasing some values ​​of the self-vehicle in the lateral offset dimension and some values ​​of the self-vehicle and the opposing game vehicle in the longitudinal acceleration dimension, a second strategy space for the self-vehicle and the opposing game vehicle is formed.

[0236] First, from the multi-dimensional game space of the self-vehicle and the opposing game vehicle, determine the maximum lateral sampling strategy space spanned by the self-vehicle and the opposing game vehicle in the lateral offset dimension. Figure 10 The diagram illustrates the lateral sampling actions determined for each of the two vehicles based on lateral offset sampling. That is, multiple lateral offset actions correspond to multiple parallel lateral offset trajectories that the vehicle can execute.

[0237] Considering the longitudinal / lateral dynamics, kinematic constraints, relative positional relationships, and relative velocity relationships of the self-vehicle and the opposing vehicle, the lateral offset values ​​for both vehicles are determined to be within the range of [-3, 3], in meters (m). During sampling, based on the self-vehicle's computational capabilities and a pre-set decision precision, the sampling interval for both vehicles is determined to be 1 meter. The released lateral sampling strategy space for both vehicles is then displayed in a two-dimensional table as shown in the upper sub-table of Table 3. The first row of the upper sub-table of Table 3 lists all possible lateral offset values ​​for the self-vehicle (Oe), and the first column lists all possible lateral offset values ​​for the opposing vehicle (Oo2). Therefore, the lateral sampling strategy space spanned by the self-vehicle and the opposing vehicle in the lateral offset dimension includes at most 7 x 7, or 49, lateral offset action pairs for both vehicles.

[0238] When a vehicle is moving, it cannot independently deviate laterally without engaging in longitudinal actions. Therefore, while releasing the lateral sampling strategy space of the self-vehicle and the opposing vehicle, it is also necessary to release multiple action pairs of the self-vehicle and the opposing vehicle in the longitudinal acceleration dimension.

[0239] To reduce computational effort and conserve computing resources, only a portion of the lateral offset value of the self-vehicle is released in this release, and this value, along with a portion of the longitudinal acceleration value of the self-vehicle and the opposing vehicle, forms the strategy space for the second release. At this point, the opposing vehicle's lateral offset value is zero. As shown in the upper sub-table of Table 3, from the lateral sampling strategy space, seven lateral offset action pairs are formed by selecting the opposing vehicle's lateral offset value as 0 and the self-vehicle's lateral offset values ​​as -3, -2, -1, 0, 1, 2, or 3. These seven lateral offset action pairs are then combined with the 64 longitudinal acceleration action pairs of the self-vehicle and opposing vehicle released in the previous release (as shown in Table 2), resulting in 7 multiplied by 64, or 448 action pairs. In each action pair, the opposing vehicle's lateral offset value is 0. Compared to the strategy space corresponding to the longitudinal acceleration sampling of the self-vehicle and the opposing game vehicle, which can release up to 64 action pairs, the number of action pairs released at this time has increased by 6 times, and is 7 times that of the first release.

[0240] S144: Based on the value determination methods of each generation, such as each cost function, calculate the cost value of each action pair in the strategy space of the second release by the partial values ​​of the self-vehicle in the lateral offset dimension and the partial values ​​of the self-vehicle and the opposing game vehicle in the longitudinal acceleration dimension, and determine the policy feasible region.

[0241] As shown in the lower sub-table of Table 3, when the lateral offset value of the self-vehicle is 1, among the 64 longitudinal acceleration action pairs of the self-vehicle and the opposing game vehicle released, after the self-vehicle and the opposing game vehicle execute 16 of the sampled actions, the passability in the sub-traffic scenario constructed by the self-vehicle and the opposing game vehicle is too poor (such as braking), which is an infeasible solution. In Table 3, these action pairs are marked with the label "0".

[0242] When the lateral offset value of the self-vehicle is 1, among the 64 longitudinal acceleration action pairs released between the self-vehicle and the opposing vehicle, after the self-vehicle and the opposing vehicle execute 48 of these sampled actions, in the sub-traffic scenario constructed by the self-vehicle and the opposing vehicle, the weighted sum of safety, comfort, passability, lateral offset cost, right-of-way cost, risk area cost, and inter-frame correlation cost is greater than a pre-set cost threshold. This constitutes a feasible solution within the strategy space, forming the policy feasible region for the self-vehicle and the opposing vehicle. In Table 3, these 48 action pairs are labeled "1". At this point, because the interaction is occurring in the current frame and does not involve the decision results of the previous frame, the inter-frame correlation cost is zero.

[0243] At this point, 48 feasible solutions have been found in the interactive game between the self-playing vehicle and the opposing vehicle, and it is no longer necessary to continue searching for solutions within the game space between the self-playing vehicle and the opposing vehicle. The total number of action pairs searched is now 64, and this round of the game consumed relatively little computing power and computation time.

[0244] That is, by combining the lateral offset action pairs with the lateral offset values ​​of 0 for the opposing car and 1 for the own car with the longitudinal acceleration action pairs of the 64 cars released previously, 48 feasible solutions are obtained from the 64 action pairs (the feasible solutions are shown in the shaded area in Table 3). These feasible solutions can be added to the policy feasible domain of the own car and the opposing car.

[0245] This is because after the car shifts 1m to the right (with the car as a reference, shifting to the right is positive and shifting to the left is negative), since the car and the opposing car have already shifted laterally, the feasible domain of the strategy in the longitudinal sampling strategy space of the car and the opposing car covers all cases except when both cars stop (the action pair is shown with background in Table 3).

[0246] Furthermore, compared to Table 2, the labels for the action pairs corresponding to the longitudinal acceleration of the voluntary vehicle and the oncoming vehicle in the lower sub-table of Table 3 have been changed from "-1" to "0". This is because when the voluntary vehicle shifts 1m to the right, it is already possible for the voluntary vehicle and the oncoming vehicle to no longer pose a collision risk. These action pairs, mapped to the traffic scenario constructed by the voluntary vehicle and the oncoming vehicle, have poor passability (i.e., they are unable to stop), and are still considered infeasible solutions, but the labels have been changed from "-1" to "0".

[0247] Furthermore, at this point, the intention decision can be determined for both the player's own vehicle and the opposing vehicle, and a strategy label can be set. Please refer to step S132, which will not be repeated here.

[0248] Table 3. The strategy space spanned by the lateral sampling strategy spaces of the player and the opposing player, and the longitudinal sampling strategy spaces of the player and the opposing player.

[0249]

[0250] The above describes the release of the strategy space for the self-playing vehicle and the opposing vehicle. In addition, multiple sampling values ​​can be selected on the lateral offset dimension of the self-playing vehicle. For example, the lateral offset values ​​of the self-playing vehicle can be 2 or 3, and together with the longitudinal acceleration sampling strategy space of the self-playing vehicle and the opposing vehicle, more strategy spaces can be formed.

[0251] In this embodiment, the lateral offset action of the opposing game car with a lateral offset value of 0 and the lateral offset action of the own car with a lateral offset value of 1, along with the longitudinal acceleration action of the 64 own cars and the opposing game car released in the previous release, form the strategy space for the second release. 48 feasible solutions are found from this strategy space. Therefore, it is no longer necessary to release other strategy spaces. At this time, the interactive game consumes less computing power and less computation time.

[0252] Table 4. Feasible solutions for the player's own vehicle and opposing vehicles, as well as for vehicles playing alongside and crossing other vehicles.

[0253]

[0254] S150: Find the intersection of the feasible regions of the strategies of the self-vehicle and the opposing game vehicle with the feasible regions of the strategies of the self-vehicle and the cross-game vehicle to determine the game outcome of the self-vehicle.

[0255] For a given feasible region of strategy for the self-car and the cross-traversing car, and the feasible region of strategy for the self-car and the opposing car, find the common feasible region of the two, and find the feasible solution with the minimum cost (i.e. the best profit) from the common feasible region.

[0256] Table 4 shows the feasible solution with the minimum cost (i.e., the best payoff) found in the common feasible region of the strategy feasible regions of the self-vehicle and the opposing vehicle in Table 3 and the strategy feasible regions of the self-vehicle and the vehicle crossing the road in Table 1. This feasible solution is a pair of game decision actions of the self-vehicle, the opposing vehicle, and the vehicle crossing the road, which is a multi-dimensional action pair composed of the longitudinal acceleration of the self-vehicle, the longitudinal acceleration of the opposing vehicle, the lateral deviation of the self-vehicle, and the longitudinal acceleration of the vehicle crossing the road.

[0257] That is, the vehicle travels at -2m / s 2 The vehicle accelerates longitudinally to decelerate and yield, then shifts 1 meter to the right to avoid the oncoming vehicle; to ensure passage, the vehicle crossing the lane moves at 1 m / s. 2 Longitudinal acceleration accelerates through the conflict zone; opposing vehicles move at 1 m / s 2 Longitudinal acceleration speeds the vehicle through the conflict zone.

[0258] After the vehicle, the oncoming vehicle, and the vehicle crossing the road perform this action, their intended decisions are respectively: the vehicle crossing the road cuts in front of the vehicle, the oncoming vehicle cuts in front of the vehicle, the vehicle moves laterally to the right to avoid the oncoming vehicle, and the vehicle yields to the vehicle crossing the road.

[0259] S160: From the game results of the vehicle, select the decision result. You can choose the action pair with the minimum cost value. Based on this, determine the executable action of the vehicle, which can be used to control the vehicle to execute the action.

[0260] In some embodiments, for multiple policy feasible domains in the game outcome, an action pair can be selected as the decision outcome based on the cost value.

[0261] In some embodiments, for multiple solutions (i.e., action pairs) in the policy feasible region of the game outcome, further multi-frame deduction can be performed on each solution. That is, the time sampling dimension is released to select action pairs with good consistency in the time dimension as the decision result for the vehicle's driving. For details, please refer to the aforementioned... Figure 7 The description.

[0262] like Figure 12 As shown, this application also provides an embodiment of a corresponding intelligent driving decision-making device. For the beneficial effects of the device or the technical problems it solves, please refer to the description in the method corresponding to each device, or to the description in the invention content, which will not be repeated here.

[0263] In this embodiment of the intelligent driving decision-making device, the intelligent driving decision-making device 100 includes:

[0264] The acquisition module 110 is used to acquire the game object with the vehicle. Specifically, it is used to execute the above steps S10 or S110-S120, or various optional embodiments corresponding to these steps.

[0265] Processing module 120 is used to perform multiple releases of multiple strategy spaces from the vehicle and the game opponent's multiple strategy spaces. After one of the multiple releases is executed, the feasible strategy domain of the vehicle and the game opponent is determined based on the released strategy spaces, and the decision result of the vehicle's movement is determined based on the feasible strategy domain. Specifically, it is used to execute the above steps S20-S40, or the various optional embodiments corresponding to these steps.

[0266] In some embodiments, the dimensions of the multiple policy spaces include at least one of the following: longitudinal sampling dimension, lateral sampling dimension, or time sampling dimension.

[0267] In some embodiments, performing multiple releases of multiple policy spaces includes performing the releases in the following order: longitudinal sampling dimension, lateral sampling dimension, and temporal sampling dimension.

[0268] In some embodiments, when determining the policy feasible region of the vehicle and the game object, the total cost of action pairs in the policy feasible region is determined according to one or more of the following: safety cost of the vehicle or the game object, right-of-way cost, lateral offset cost, passability cost, comfort cost, inter-frame correlation cost, and risk area cost.

[0269] In some embodiments, when the total value of a pair of actions is determined based on two or more values, each value has a different weight.

[0270] In some embodiments, when there are two or more game objects, the decision outcome of the vehicle's driving is determined based on the feasible domains of each strategy of the vehicle and each game object.

[0271] In some embodiments, the acquisition module 110 is further configured to acquire non-game objects of the vehicle; the processing module 120 is further configured to determine the policy feasible region between the vehicle and the non-game objects; the policy feasible region between the vehicle and the non-game objects includes the actions that the vehicle can perform relative to the non-game objects; and the decision result of the vehicle's driving is determined at least based on the policy feasible region between the vehicle and the non-game objects.

[0272] In some embodiments, the processing module 120 is further configured to determine the strategy feasible region of the decision result of the vehicle driving based on the intersection of the strategy feasible regions of the vehicle and each game object, or to determine the strategy feasible region of the decision result of the vehicle driving based on the intersection of the strategy feasible regions of the vehicle and each game object and the strategy feasible regions of the vehicle and each non-game object.

[0273] In some embodiments, the acquisition module 110 is further configured to acquire non-game objects of the vehicle; the processing module 120 is further configured to constrain the longitudinal sampling strategy space corresponding to the vehicle or constrain the lateral sampling strategy space corresponding to the vehicle based on the motion state of the non-game objects.

[0274] In some embodiments, the acquisition module 110 is further configured to acquire non-game objects of the game objects of the vehicle; the processing module 120 is further configured to constrain the longitudinal sampling strategy space corresponding to the game objects of the vehicle, or constrain the lateral sampling strategy space corresponding to the game objects of the vehicle, based on the motion state of the non-game objects.

[0275] In some embodiments, when the intersection is an empty set, a conservative decision is made regarding the vehicle's movement. The conservative decision includes actions to safely stop the vehicle or actions to safely decelerate the vehicle.

[0276] In some embodiments, the game object or non-game object is determined based on the attention method.

[0277] In some embodiments, the processing module 120 is further configured to display at least one of the following through a human-computer interaction interface: the decision result of the vehicle's driving, the policy feasible domain of the decision result, the driving trajectory of the vehicle corresponding to the decision result of the vehicle's driving, or the driving trajectory of the game object corresponding to the decision result of the vehicle's driving.

[0278] The decision result of the vehicle's driving can be the decision result of the current single frame simulation, or the decision result corresponding to the multiple single frame simulations that have been executed. The decision result can be the action that the vehicle can perform, or the action that the game object can perform, or the intention decision corresponding to the vehicle performing the action, such as Cg or Cy in Table 1, such as cutting in, yielding, or avoiding.

[0279] In the above, the policy feasible region of the decision result can be the policy feasible region of the current single-frame inference, or it can be the policy feasible region corresponding to the multiple single-frame inferences that have been executed.

[0280] The above-mentioned vehicle trajectory corresponding to the decision result of the vehicle's driving can be the vehicle trajectory corresponding to the first single frame in the decision-making process, such as... Figure 7 T1 in the equation can also be the vehicle's trajectory, formed by sequentially connecting multiple single-frame deductions executed in a single decision step, such as... Figure 7 T1, T2, and Tn in the example.

[0281] The above suggests that the trajectory of the game object corresponding to the decision result of the vehicle's movement can be the trajectory of the game object corresponding to the first single-frame deduction in a decision step, such as... Figure 7 T1 in the context can also be the trajectory of the game object, formed by sequentially connecting multiple single-frame deductions executed in a single decision step, such as... Figure 7 T1, T2, and Tn in the example.

[0282] like Figure 13 As shown in the embodiments of this application, a vehicle driving control method is also provided, including:

[0283] S210: Obtain information about obstacles outside the vehicle;

[0284] S220: Based on the obstacle information, determine the decision result for vehicle driving according to any of the above intelligent driving decision-making methods;

[0285] S230: Control the vehicle's movement based on the decision result.

[0286] like Figure 14 As shown in the figure, this application embodiment also provides a vehicle driving control device 200, including: an acquisition module 210 for acquiring obstacles outside the vehicle; a processing module 220 for determining a vehicle driving decision result based on any of the above intelligent driving decision methods for the obstacles; the processing module is also used to control the vehicle driving based on the decision result.

[0287] like Figure 15As shown, this application embodiment also provides a vehicle 300, including: the aforementioned vehicle driving control device 200, and a driving system 250; the vehicle driving control device 200 controls the driving system 250. In some embodiments, the driving system 250 may include the aforementioned... Figure 2 The driving system 13 in the middle.

[0288] Figure 16 This is a schematic structural diagram of a computing device 400 provided in an embodiment of this application. The computing device 400 includes a processor 410, a memory 420, and may also include a communication interface 430.

[0289] It should be understood that Figure 16 The communication interface 430 in the computing device 400 shown can be used to communicate with other devices.

[0290] The processor 410 can be connected to the memory 420. The memory 420 can be used to store the program code and data. Therefore, the memory 420 can be a storage unit inside the processor 410, an external storage unit independent of the processor 410, or a component that includes both the storage unit inside the processor 410 and the external storage unit independent of the processor 410.

[0291] Optionally, the computing device 400 may also include a bus. The memory 420 and communication interface 430 can be connected to the processor 410 via the bus. The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0292] It should be understood that in the embodiments of this application, the processor 410 may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor. Alternatively, the processor 410 may employ one or more integrated circuits to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0293] The memory 420 may include read-only memory and random access memory, and provides instructions and data to the processor 410. A portion of the processor 410 may also include non-volatile random access memory. For example, the processor 410 may also store device type information.

[0294] When the computing device 400 is running, the processor 410 executes the computer execution instructions in the memory 420 to perform the operation steps of the above method.

[0295] It should be understood that the computing device 400 according to the embodiments of this application can correspond to the corresponding subject in executing the methods according to the various embodiments of this application, and the above and other operations and / or functions of each module in the computing device 400 are respectively for implementing the corresponding processes of the methods of this embodiment. For the sake of brevity, they will not be described in detail here.

[0296] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0297] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0298] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0299] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0300] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0301] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0302] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is used to perform the above-described method, which includes at least one of the schemes described in the above embodiments.

[0303] The computer storage medium in this application embodiment can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0304] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0305] The program code contained on a computer-readable medium may be transmitted using any suitable medium, including, but not limited to, wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0306] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0307] The terms "first, second, third, etc." or similar terms such as module A, module B, and module C used in the specification and claims are only used to distinguish similar objects and do not represent a specific ordering of objects. It is understood that a specific order or sequence may be interchanged where permitted so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0308] In the above description, the labels of the steps, such as S116, S124, etc., do not mean that the steps will always be executed. The order of the steps can be interchanged or executed simultaneously if permitted.

[0309] The term "comprising" as used in the specification and claims should not be construed as limiting itself to what follows; it does not exclude other elements or steps. Therefore, it should be interpreted as specifying the presence of the mentioned feature, integral, step, or component, but does not exclude the presence or addition of one or more other features, integrals, steps, or components, or groups thereof. Thus, the statement "device comprising means A and B" should not be limited to a device consisting solely of components A and B.

[0310] The terms "an embodiment" or "an embodiment" as used in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in at least one embodiment of this application. Therefore, the terms "in an embodiment" or "in an embodiment" appearing throughout this specification do not necessarily refer to the same embodiment, but may refer to the same embodiment. Furthermore, in one or more embodiments, the specific features, structures, or characteristics can be combined in any suitable manner, as will be apparent to those skilled in the art from this disclosure. Note that the above are merely preferred embodiments of this application and the technical principles employed. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of this application, all of which fall within the scope of protection of this application.

Claims

1. An intelligent driving decision-making method, characterized in that, include: Obtain the game object of the vehicle; From the multiple strategy spaces of the vehicle and the game object, multiple releases of the multiple strategy spaces are performed. After one of the multiple releases is performed, the strategy feasible region of the vehicle and the game object is determined according to the released strategy spaces. If the strategy feasible region is not empty, the decision result of the vehicle's driving is determined according to the strategy feasible region, and the release of the strategy spaces that have not yet been released in the multiple strategy spaces is terminated. The strategy space refers to the set of all possible action pairs that the vehicular and the game player might produce.

2. The method according to claim 1, characterized in that, The dimensions of the multiple policy spaces include at least one of the following: vertical sampling dimension, horizontal sampling dimension, or time sampling dimension.

3. The method according to claim 2, characterized in that, The execution of multiple releases of the multiple policy spaces includes executing the releases in the following order: vertical sampling dimension, horizontal sampling dimension, and time sampling dimension.

4. The method according to any one of claims 1-3, characterized in that, When determining the policy feasible region between the vehicle and the game opponent, the total cost of action pairs within the policy feasible region is determined according to one or more of the following: The value of safety, right-of-way, lateral offset, passability, comfort, inter-frame correlation, and risk area for the vehicle or the other party in the game.

5. The method according to claim 4, characterized in that, When the total value of the action pair is determined based on two or more values, each value has a different weight.

6. The method according to claim 1, characterized in that, When the game involves two or more players, the decision result for the vehicle's driving is determined based on the feasible domains of each strategy of the vehicle and each of the game players.

7. The method according to any one of claims 1-3, characterized in that, Also includes: Obtain the non-game object of the vehicle; Determine the feasible strategy domain between the vehicle and the non-game object; The decision outcome for the autonomous vehicle's driving is determined at least based on the policy feasibility domain of the autonomous vehicle and the non-game object.

8. The method according to claim 7, characterized in that, The feasible region of the vehicle's driving decision is determined by the intersection of the feasible regions of the vehicle and each of the aforementioned game opponents. The policy feasible region of the vehicle's driving decision is determined by the intersection of the policy feasible regions of the vehicle and each of the game objects, as well as the policy feasible regions of the vehicle and each of the non-game objects.

9. The method according to any one of claims 2, 3, and 6, characterized in that, Also includes: Obtain the non-game object of the vehicle; Based on the motion state of the non-game object, constrain the longitudinal sampling strategy space corresponding to the vehicle, or constrain the lateral sampling strategy space corresponding to the vehicle.

10. The method according to any one of claims 2, 3, and 6, characterized in that, Also includes: Obtain the non-game objects of the game objects of the vehicle; Based on the motion state of the non-game object, constrain the longitudinal sampling strategy space corresponding to the game object of the self-vehicle, or constrain the lateral sampling strategy space corresponding to the game object of the self-vehicle.

11. The method according to claim 8, characterized in that, When the intersection is an empty set, a conservative decision is made regarding the vehicle's movement. The conservative decision includes actions that allow the vehicle to stop safely or actions that allow the vehicle to decelerate safely.

12. The method according to claim 1, characterized in that, The game object or non-game object is determined based on the attention method.

13. The method according to any one of claims 1, 2, 3, 6, and 12, characterized in that, Also includes: Display at least one of the following through a human-computer interaction interface: The decision result of the vehicle's driving, the policy feasible region of the decision result, the driving trajectory of the vehicle corresponding to the decision result of the vehicle's driving, or the driving trajectory of the game object corresponding to the decision result of the vehicle's driving.

14. An intelligent driving decision-making device, characterized in that, include: The acquisition module is used to acquire the game object of the vehicle. The processing module is configured to perform multiple releases of the multiple strategy spaces between the vehicle and the game opponent. After one of the multiple releases is executed, the module determines the feasible strategy domain for the vehicle and the game opponent based on the released strategy spaces. If the feasible strategy domain is not empty, the module determines the decision result for the vehicle's movement based on the feasible strategy domain, and terminates the release of the remaining unreleased strategy spaces. The strategy space refers to the set of all possible action pairs that the vehicular and the game player might produce.

15. The apparatus according to claim 14, characterized in that, The dimensions of the multiple policy spaces include at least one of the following: vertical sampling dimension, horizontal sampling dimension, or time sampling dimension.

16. The apparatus according to claim 15, characterized in that, The execution of multiple releases of the multiple policy spaces includes executing the releases in the following order: vertical sampling dimension, horizontal sampling dimension, and time sampling dimension.

17. The apparatus according to any one of claims 14-16, characterized in that, When determining the policy feasible region between the vehicle and the game opponent, the total cost of action pairs within the policy feasible region is determined according to one or more of the following: The value of safety, right-of-way, lateral offset, passability, comfort, inter-frame correlation, and risk area for the vehicle or the other party in the game.

18. The apparatus according to claim 17, characterized in that, When the total value of the action pair is determined based on two or more values, each value has a different weight.

19. The apparatus according to claim 14, characterized in that, When the game involves two or more players, the decision result for the vehicle's driving is determined based on the feasible domains of each strategy of the vehicle and each of the game players.

20. The apparatus according to any one of claims 14-16, characterized in that, The acquisition module is also used to acquire non-game objects of the vehicle; The processing module is further configured to determine the policy feasible region between the vehicle and the non-game object; and to determine the decision result of the vehicle's driving based at least on the policy feasible region between the vehicle and the non-game object.

21. The apparatus according to claim 20, characterized in that, The processing module is also used for: The feasible region of the vehicle's driving decision is determined by the intersection of the feasible regions of the vehicle and each of the aforementioned game opponents. The policy feasible region of the vehicle's driving decision is determined by the intersection of the policy feasible regions of the vehicle and each of the game objects, as well as the policy feasible regions of the vehicle and each of the non-game objects.

22. The apparatus according to any one of claims 15, 16, and 19, characterized in that, The acquisition module is also used to acquire non-game objects of the vehicle; The processing module is also used to constrain the longitudinal sampling strategy space corresponding to the vehicle or the lateral sampling strategy space corresponding to the vehicle based on the motion state of the non-game object.

23. The apparatus according to any one of claims 15, 16, and 19, characterized in that, The acquisition module is also used to acquire the non-game objects of the game object of the vehicle; The processing module is further configured to constrain the longitudinal sampling strategy space corresponding to the game object of the vehicle, or constrain the lateral sampling strategy space corresponding to the game object of the vehicle, based on the motion state of the non-game object.

24. The apparatus according to claim 21, characterized in that, When the intersection is an empty set, a conservative decision is made regarding the vehicle's movement. The conservative decision includes actions that allow the vehicle to stop safely or actions that allow the vehicle to decelerate safely.

25. The apparatus according to claim 14, characterized in that, The game object or non-game object is determined based on the attention method.

26. The apparatus according to claim 14, characterized in that, The processing module is also configured to display at least one of the following through a human-computer interaction interface: The decision result of the vehicle's driving, the policy feasible region of the decision result, the driving trajectory of the vehicle corresponding to the decision result of the vehicle's driving, or the driving trajectory of the game object corresponding to the decision result of the vehicle's driving.

27. A vehicle driving control method, characterized in that, include: Obtain obstacles outside the vehicle; Regarding the obstacle, the decision result for vehicle driving is determined by any one of the methods described in claims 1-13; The vehicle's movement is controlled based on the decision results.

28. A vehicle driving control device, characterized in that, include: The acquisition module is used to acquire obstacles outside the vehicle; A processing module is configured to determine a decision result for vehicle driving in response to the obstacle, according to any one of claims 1-13; The processing module is also used to control the vehicle's movement based on the decision result.

29. A vehicle, characterized in that, include: The vehicle driving control device and driving system as described in claim 28; The vehicle driving control device controls the driving system.

30. A computing device, characterized in that, include: processor, and A memory storing program instructions that, when executed by the processor, cause the processor to implement the intelligent driving decision-making method according to any one of claims 1-13, or cause the processor to implement the vehicle driving control method according to claim 27.

31. A computer-readable storage medium, characterized in that, It stores program instructions, which, when executed by a processor, cause the processor to implement the intelligent driving decision-making method according to any one of claims 1-13, or cause the processor to implement the vehicle driving control method according to claim 27.