An intelligent automobile human-like decision-making method based on rapid optimization of sequential game
By establishing a multi-stage sequential game decision-making model and particle swarm optimization algorithm for intelligent vehicles, the problem of human-like decision-making in complex traffic environments for intelligent vehicles has been solved, achieving efficient and accurate multi-vehicle interactive decision-making and improving the reliability and computational efficiency of intelligent vehicle driving behavior.
Patent Information
- Application Number
- CN202310047881.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-31
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2043-01-31
AI Technical Summary
Existing technologies are insufficient to effectively address the human-like decision-making challenges faced by intelligent vehicles in complex traffic environments. In particular, issues such as hierarchical decision-making, information mismatch, and low optimization efficiency exist in multi-stage sequential games, making it difficult to achieve highly reliable and safe driving behavior.
By establishing a multi-stage sequential game decision-making model for intelligent vehicles, dynamic information interaction is achieved, the strategies of other vehicles are predicted, the expected utility induced functional is optimized, and the acceleration strategy is quickly solved by combining the particle swarm optimization algorithm, thus realizing the integration and synergy of multi-vehicle interactive decision-making.
It improves the human-likeness and decision-making accuracy of intelligent vehicles in complex traffic environments, enhances computational efficiency, achieves a trade-off between solution accuracy and computational efficiency, and strengthens the similarity between the driving behavior of intelligent vehicles and human-driven vehicles.
Smart Images

Figure CN116070671B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of intelligent car human-like decision-making, and particularly relates to an intelligent car human-like decision-making method based on rapid optimization of sequential games. BACKGROUND
[0002] Intelligent cars are the frontier of future car development. In order to be widely popularized and applied, intelligent cars must adapt to a traffic ecology composed of various intelligent cars of different levels and manually driven cars. Humanization is the only way for cars to naturally integrate into the traffic ecology. Only when the driving behavior of an intelligent car is similar to that of a driver and can be understood by other traffic participants can the intelligent car truly integrate into a complex traffic environment. Studies have shown that decision-making is the core of an intelligent driving system and is the key to completing various driving tasks, improving car driving safety, improving driving and riding experience, optimizing travel efficiency, and improving car intelligence. The driving environment in the real world is a complex dynamic environment with multiple intelligent agents. The human-vehicle-environment is a coupled whole. Intelligent cars do not exist independently on the road. The driving behavior of all traffic participants, including intelligent cars, will be affected by the real-time environment, and the behavior results will also have a counteraction on the overall environment. Therefore, interaction is a key element for strengthening the integration of human-vehicle-environment.
[0003] Therefore, the first task of human-like decision-making considering strong interaction is to meet the extremely high reliability and safety required by intelligent driving, effectively cope with scenario complexity, environmental variability, traffic dynamics, game interaction, and information incompleteness, and generate human-like and personalized driving behavior. The trust, acceptance, and adaptability of drivers and other traffic participants also put forward higher requirements for decision-making. In view of the current development trend, human-like decision-making is a difficult and important problem in driving decision-making technology and a problem that needs to be solved before truly realizing full intelligent driving. Multi-stage sequential games provide a new idea for solving the above problems, but the hierarchical characteristics of decision-making, the asynchronization of information stamps, and other problems bring challenges to interactive decision-making and optimization solving. In addition, due to the diversified driving needs of intelligent cars and the rich performance requirements of the traffic system, the increasingly complex optimization objective function may face challenges such as nonlinearity, discreteness, non-differentiability, and non-convexity. In addition, it involves kinematic and dynamic constraints of intelligent cars, and the classical optimization method needs to make a series of simplifying assumptions, so it is difficult to effectively handle such challenges alone.
[0004] Therefore, a decision-making method that can effectively adapt to hierarchical sequential driving behavior and balance optimization solving accuracy and computational efficiency is a key technical problem that needs to be solved to truly realize human-like driving and has very important value for accelerating the deployment of computing and complex decision-making methods. SUMMARY
[0005] In view of the above problems, the present application provides an intelligent vehicle human-like decision-making method based on rapid optimization of sequential game, which fully considers the horizontal evolution characteristics and vertical layering characteristics of multi-stage sequential game, improves the integrity, collaboration and human-likeness of multi-vehicle interactive decision-making, overcomes the problem of rough solution in discrete domain, accelerates the optimization solving process, and realizes the trade-off between solving accuracy and computing efficiency.
[0006] The embodiment of the present application provides an intelligent vehicle human-like decision-making method based on rapid optimization of sequential game, comprising:
[0007] S1, the intelligent vehicle and other vehicles in the traffic environment realize dynamic information interaction;
[0008] S2, based on the dynamic information interaction, an intelligent vehicle multi-stage sequential game decision-making model is established;
[0009] S3, each vehicle participating in the dynamic information interaction predicts the strategy to be taken by other vehicles except itself in the future, establishes an expected utility induced functional of the leading vehicle and an expected utility induced functional of the following vehicle, and optimizes the expected utility induced functional of itself in the multi-stage sequential game according to the prediction result;
[0010] S4, based on the expected utility induced functional of the leading vehicle and the expected utility induced functional of the following vehicle, the intelligent vehicle multi-stage sequential game decision-making model is solved, and the acceleration of each vehicle at H instant is output;
[0011] S5, the intelligent vehicle executes the acceleration of itself at H instant;
[0012] S6, the motion information, prediction information and cognitive information of the intelligent vehicle and other vehicles in the traffic environment are synchronously updated at H+h instant;
[0013] S7, steps S1-S6 are repeatedly executed until the intelligent vehicle autonomously exits the game system and restores the free driving mode.
[0014] Further, in the S1, the dynamic information interaction includes driving state information interaction, driving style interaction, expected utility induced index interaction and environmental awareness interaction.
[0015] Further, in the S1, when the interaction information is incomplete, the intelligent vehicle estimates the default information according to the belief and cognition of the traffic environment, and predicts the evolution of the future traffic situation.
[0016] Further, in the S2, the intelligent vehicle multi-stage sequential game decision-making model integrates the multi-objective collaborative driving demand of individual driving characteristics and the traffic system demand;
[0017] The multi-objective cooperative driving requirements include the risk characteristics, impact characteristics, and following characteristics of intelligent vehicles;
[0018] The requirements for the transportation system include its ergonomic and safety characteristics.
[0019] Furthermore, S4 specifically includes:
[0020] S41. Generate initial acceleration strategy particles that satisfy vehicle kinematics and dynamic constraints within the continuous action domain of the leading car in the upper space, and form an initial particle swarm.
[0021] S42, based on the initial particle swarm of the leading car in the upper space. Obtain the reactive particle groups that follow the car
[0022] S43, Based on the initial particle swarm and reactive particle groups Constructing a hierarchical particle swarm
[0023] S44. Measure the hierarchical particle swarm using the expected utility induced functional of the lead car and the expected utility induced functional of the following cars. Initial particle swarm and reactive particle groups The utility of the initial particle swarm. and reactive particle groups The utility values are individually sorted in descending order;
[0024] S45. Using the probability disk replication method, respectively from the sorted initial particle swarm... and reactive particle groups Copy from These particles are denoted as the leading automotive core particle group. and following the core particle group of automobiles
[0025] S46, Regarding the leading automotive core particle group and following the core particle group of automobiles Encoding, swapping, and mutation operations are performed sequentially to generate new particles corresponding to the upper-space leading car acceleration strategy. New particles that follow the car's acceleration strategy in the lower space
[0026] S47, New particles based on the aforementioned upper-space leading car acceleration strategy and the new particle of the lower space following car acceleration strategy calculate the optimal acceleration strategy of the corresponding lower space following car;
[0027] S48, according to the optimal acceleration strategy of the lower space following car, deduce the optimal acceleration strategy of the upper space leading car;
[0028] S49, calculate the satisfaction degree of the optimal acceleration strategy of the lower space following car and the optimal acceleration strategy of the upper space leading car, and output the acceleration of each car at H moment.
[0029] Further, the S42 specifically comprises:
[0030] Take each particle in the initial particle population as a fixed parameter of the action of the leading car, find the slice of the expected utility space of the following car corresponding to each parameter;
[0031] Solve the optimal response acceleration strategy of the following car and the action of the leading car in each space slice in turn, and obtain the reactive particle population of each following car
[0032] Further, in the S47, according to the new particle of the upper space leading car acceleration strategy calculate the optimal acceleration strategy of the corresponding lower space following car, specifically comprising:
[0033] obtain the dimension of the particle population acceptance of the upper space leading car determine the relationship between the new particle and the dimension of the particle population acceptance ;
[0034] (1) when the size of the new particle is greater than the dimension of the particle population acceptance :
[0035] for each particle in the new particle , find the initial particle in the initial particle population closest to it by vector norm;
[0036] and dig out the optimal acceleration strategy of the lower space following car corresponding to the initial particle in the hierarchical particle population ;
[0037] Take the particle corresponding to the optimal acceleration strategy of the following car as a reference to construct a linear function near this particle;
[0038] approximate the new particle each particle in the new particle corresponds to the optimal acceleration strategy of the following car in the lower space;
[0039] (2) when the size of the new particle is less than the dimension of the particle population acceptance , the following is performed:
[0040] for each particle in the new particle , find the initial particle closest to it in the initial particle population by vector norm;
[0041] and dig out the lower space following the optimal acceleration strategy of the car corresponding to the initial particle in the hierarchical particle population ;
[0042] Take the particle corresponding to the optimal acceleration strategy of the following car in the lower space as the starting point, construct the quadratic approximation function of the expected utility induced functional of the following car in the lower space as the quasi-expected utility induced functional, and construct the linear approximation function of its constraint as the quasi-constraint function;
[0043] Convert the problem into a quadratic programming problem based on the quasi-expected utility induced functional and the quasi-constraint function, determine the new search direction and the multiplier vector;
[0044] Based on the particle point corresponding to the optimal acceleration strategy of the following car in the lower space dug out from the hierarchical particle population , perform one-dimensional search along the search direction to obtain the search step, and then calculate the new approximate optimal acceleration strategy;
[0045] Iterate the above process until the convergence precision meets the preset condition to stop iteration, and output the optimal acceleration strategy of the following car in the lower space corresponding to each particle in the new particle .
[0046] Further, in the S47, the optimal acceleration strategy of the following car in the lower space corresponding to the new particle is calculated, which specifically includes:
[0047] Measure the utility degree of the new particle by the expected utility induced functional of the following car one by one, and select the single particle with the optimal utility degree as the optimal acceleration strategy of the following car in the lower space from the utility degrees of the reactive particle population and the new particle .
[0048] Further, the S48 specifically includes:
[0049] The lower space particle corresponding to the optimal acceleration strategy of the following car is taken as an optimal particle The comprehensive utility degree of the optimal particle and the new particle is measured by the expected utility induced functional of the leading car The optimal utility degree particle is selected as the optimal acceleration strategy of the upper space leading car by comprehensively considering the comprehensive utility degree and the utility degree of the initial particle group The optimal utility degree particle is selected as the optimal acceleration strategy of the upper space leading car by comprehensively considering the comprehensive utility degree and the utility degree of the initial particle group
[0050] Further, in the S9, the satisfaction condition includes
[0051]
[0052]
[0053]
[0054]
[0055]
[0056] wherein, and respectively represent the acceleration strategy of the upper space leading car in the (ρ+1) th iteration and the acceleration strategy of the lower space following car; and respectively represent the acceleration strategy of the upper space leading car in the (ρ+1) th iteration and the acceleration strategy of the lower space following car; γ1, γ2 and γ3 respectively represent the threshold values of the satisfaction conditions of the three types; represent the norm between the subgame perfect Nash equilibrium solutions obtained by two rounds of calculation. If the above five conditions are met at the same time, it indicates that the acceleration strategy of the upper space leading car and the acceleration strategy of the lower space following car have gradually converged to the subgame perfect Nash equilibrium state, and the solving task of the intelligent car sequential game decision model in the instantaneous H continuous action domain is completed. If the above five conditions are not met at the same time, new particle and new particle The above all processes are repeatedly executed.
[0057] Compared with the prior art, the intelligent car human-like decision method based on sequential game rapid optimization has the following beneficial effects: the method fully considers the horizontal evolution characteristics and vertical layering characteristics of the multi-stage sequential game, improves the integrity, collaboration and human-likeness of the multi-car interactive decision, overcomes the problem of rough solution in discrete domain, accelerates the optimization solving process, and realizes the trade-off between solving accuracy and calculation efficiency.
[0058] Other features and advantages of the present application will be set forth in the description that follows, and in part will be apparent from the description, or can be learned by practice of the application. The purposes and other advantages of the present application will be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings.
[0059] The technical solutions of the present application are described in further detail below with the help of the accompanying drawings and examples. BRIEF DESCRIPTION OF DRAWINGS
[0060] The accompanying drawings are included to provide a further understanding of the present application and are incorporated in and constitute a part of the specification, illustrate embodiments of the present application and are used to explain the present application, but do not limit the present application. In the drawings:
[0061] Figure 1 The flowchart of the intelligent car human-like decision-making method based on rapid optimization of sequential game provided for the embodiment of the present application.
[0062] Figure 2 The schematic diagram of the multi-stage multi-car sequential game interactive decision-making process provided for the embodiment of the present application.
[0063] Figure 3 The schematic diagram of the slice of the expected utility induced functional of the following car and the acceleration strategy of the leading car and the acceleration strategy of the following car provided for the embodiment of the present application.
[0064] Figure 4 The schematic diagram of the relationship between the acceleration strategy of the leading car and the optimal acceleration strategy of the following car provided for the embodiment of the present application.
[0065] Figure 5 The schematic diagram of the change of the expected utility induced functional of the leading car with the acceleration strategy of the leading car provided for the embodiment of the present application.
[0066] Figure 6 The flowchart of the new particle generation of the leading car and the following car fused with intelligent algorithm provided for the embodiment of the present application.
[0067] Figure 7 The flowchart of the hierarchical spatial linear approximation optimization provided for the embodiment of the present application.
[0068] Figure 8 The flowchart of the hierarchical spatial quasi-quadratic programming provided for the embodiment of the present application. DETAILED DESCRIPTION
[0069] Exemplary embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be thoroughly understood, and so that the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0070] Referring to Figure 1 As shown, the embodiment of the present application provides a kind of intelligent car human-like decision-making method based on rapid optimization of sequential game, specifically comprising the following steps:
[0071] S1, the intelligent car in traffic environment and other cars realize dynamic information interaction;
[0072] S2, based on the dynamic information interaction, establish intelligent car multi-stage sequential game decision model;
[0073] S3, each car participating in dynamic information interaction predicts the strategy to be taken by other cars except itself in the future, establishes the expected utility induced functional of leading car and the expected utility induced functional of following car, and optimizes the expected utility induced functional of itself in multi-stage sequential game according to the prediction result;
[0074] S4, based on the expected utility induced functional of leading car and the expected utility induced functional of following car;The intelligent car multi-stage sequential game decision model is solved, and the acceleration of each car at H instantaneous is output;
[0075] S5, the intelligent car executes the acceleration of itself at H instantaneous;
[0076] S6, the motion information, prediction information and cognitive information of the intelligent car and other cars in the traffic environment are updated synchronously at H+h instantaneous;
[0077] S7, repeat steps S1-S6 until the intelligent car autonomously exits the game system and restores the free driving mode.
[0078] The above steps will be described in detail below.
[0079] In the above step S1, the intelligent car and other cars in the traffic environment realize dynamic information interaction through vehicle communication device;Wherein, other cars include other intelligent cars and manually driven cars in the traffic environment;Dynamic information interaction includes driving state information interaction, driving style interaction, expected utility induced index interaction and environment cognition interaction, etc.;When the interactive information is incomplete, the intelligent car estimates the default information according to the belief and cognition of the traffic environment, and predicts the evolution of future traffic situation.
[0080] In step S2 above, a multi-stage sequential game decision-making model for intelligent vehicles is established based on dynamic information interaction, integrating the multi-objective cooperative driving requirements of personalized driving characteristics with the requirements of the transportation system. The multi-objective cooperative driving requirements include the risk characteristics, impact characteristics, and following characteristics of the intelligent vehicle; the transportation system requirements include the ergonomic characteristics and safety characteristics of the transportation system. The established multi-stage sequential game decision-making model for intelligent vehicles allows the driving strategy of the intelligent vehicle to vary within a continuous range that satisfies vehicle dynamics and kinematic constraints.
[0081] The risk characteristics in the aforementioned multi-objective cooperative driving requirements are quantified using the power of the square of the future travel time difference between vehicles. A larger future travel time difference indicates a lower risk level for the vehicle at the current instant, while a smaller future travel time difference indicates a higher risk level. The future travel time of each vehicle is determined by its current driving state, the strategy it adopts from its game strategy set, and the strategies adopted by other vehicles from their respective game strategy sets. The established quantitative indicators for risk characteristics are as follows:
[0082]
[0083]
[0084]
[0085]
[0086]
[0087] Where H_instantaneous represents the absolute time; h_increment represents the sampling interval, which can be formed as H_time, H+h_time, H+2h_time, H+3h_time, etc. This represents the integral variable over a continuous interval of [H, H+h]. and They represent in Instantaneous accelerations of car j and car k; v j (Hh) and v k (Hh) represent the velocities of car j and car k at instant Hh, respectively; Z j (H) and Z k (H) represents the absolute coordinates of car j and car k at instant H, respectively; Z j (H+h) and Z k (H+h) represent the absolute coordinates of vehicle j and vehicle k in the pre-aiming potential conflict zone, respectively; Ξ h This represents a function that solves for the value of h in the formula; and These represent the future travel times of car j and car k through the anticipated potential conflict zone at instant H, respectively. This represents the future travel time difference between car j and car k at instant H; and It is used to replace a certain type of integration process, with the purpose of increasing the readability of the calculation formula;
[0088] The impact characteristics in the aforementioned multi-objective cooperative driving requirements aim to improve the riding experience of passengers in intelligent vehicles, avoid continuous abrupt changes in acceleration and deceleration, and enable intelligent vehicles to drive smoothly and comfortably as much as possible, thereby improving the acceptability and adaptability of intelligent vehicles. The quantitative indicators of the established impact characteristics are as follows:
[0089]
[0090]
[0091] in, and They represent in Instantaneous accelerations of car j and car k; and They represent in Instantaneous accelerations of car j and car k; and Let represent the product of the square roots of the absolute values of the changes in acceleration of car j and car k in every two adjacent stages of the game, from the start of the sequential game to instant H.
[0092] The following feature in the aforementioned multi-objective cooperative driving requirement aims to limit the deviation of the intelligent vehicle's driving path from the reference trajectory output by the planning module caused by multi-stage interactive decision-making in the sequential game process, thereby enabling the intelligent vehicle to maintain optimal trajectory following characteristics and reducing the disturbance of yaw on the coordination of the traffic system. The quantitative indicators of the established following feature are as follows:
[0093]
[0094]
[0095]
[0096] Among them, E j (H) and E k (H) represents the quantitative indicators of the following characteristics of car j and car k at instant H, respectively; and They represent in At any given moment, the absolute coordinates of the reference trajectory points of car j and car k; and respectively, at the instant instantaneous absolute coordinates of the car j and the car k, respectively, at the instant the acceleration results of the instantaneous stage sequential game decision and are functions of and respectively, at the instant instantaneous, the degree of inclination of the car j and the car k to the expected arrival position, which is related to the personalized driving style and driving habit of the intelligent car; θ and μ both represent factors related to the personalized driving style of the intelligent car; ρ and γ both represent factors related to the driving habit of the intelligent car.
[0097] The multi-stage sequential game decision model of the intelligent car includes but is not limited to the risk characteristics, impact characteristics and following characteristics in the above-mentioned multi-objective cooperative driving demand, and other demand characteristics can be integrated into the established model, denoted as
[0098] The ergonomics characteristics in the above-mentioned traffic system demand aim to fully optimize the traffic system speediest dredging mechanism, and improve the dredging efficiency of the traffic system traffic flow under the condition of meeting the traffic system speed limit and other legalities. The quantitative indicators of the established ergonomics characteristics are as follows:
[0099]
[0100]
[0101] wherein, v j (H) and v k (H) represent the speed of the car j and the car k at the instant and respectively, represent the coefficients related to the traffic system speed limit; and respectively, represent the lower limit of speed and the upper limit of speed under the traffic legality regulation.
[0102] The safety characteristics in the above-mentioned traffic system demand aim to control the car-car conflict risk in the whole process of multi-stage sequential game within a safe and reasonable range. The quantitative indicators of the established safety characteristics are as follows:
[0103]
[0104] wherein, and respectively, represent the remaining distance of the car j and the car k to the potential conflict zone of the two at the instant and respectively, represent the remaining distance of the car j and the car k to the potential conflict zone of the two at the instant instantaneous acceleration of car j and car k; and respectively represent the instantaneous speed of car j and car k. instantaneous acceleration of car j and car k;
[0105] The intelligent car multi-stage sequential game decision model includes but is not limited to the ergonomics features and safety features in the above traffic system requirements, and other requirement features can be integrated into the established model, denoted as
[0106] In the above step S3, each car participating in dynamic information interaction optimizes its expected utility induced functional in a multi-stage sequential game by using a fast optimization method based on the prediction of the strategy that other cars may take in the future. The upper space and the lower space of the hierarchical decision correspond to the continuous strategy set of the leading car and the following car in the sequential game, respectively, and also include the expected utility induced functional and the constraint condition of the leading car and the expected utility induced functional and the constraint condition of the following car, respectively. Each constraint condition is related to the configuration parameters of each car; wherein the leading car refers to the car that takes action first in the sequential game; the following car refers to the car that follows to take action in the sequential game; the fast optimization method proposed in the embodiment of the application is used to solve the above intelligent car multi-stage sequential game decision model, and the specific acceleration of each car at H instant is output. The schematic diagram of the multi-stage multi-car sequential game interactive decision process can be seen from Figure 2 .
[0107] The expected utility induced functional and the constraint condition of the leading car in the multi-stage sequential game can be integrated as:
[0108]
[0109] A1(a j ,a k )=a max -a j ≥0
[0110] A2(a j ,a k )=a min -a j ≤0
[0111] A3(a j ,a k )=v max -v j ≥0
[0112] A4(a j ,a k )=v min -v j ≤0
[0113]
[0114]
[0115]
[0116] where, is the expected utility induced functional of the leader vehicle; A1(a j ,a k ), A2(a j ,a k ), A3(a j ,a k ), A4(a j ,a k ), A5(a j ,a k ) are inequality constraints of the leader vehicle; B1(a j ,a k ), B2(a j ,a k ) are equality constraints of the leader vehicle; a j represents the acceleration strategy of the leader vehicle; a k represents the acceleration strategy of the follower vehicle.
[0117] The expected utility induced functional and the constraints of the follower vehicle in the multi-stage sequential game can be integrated as:
[0118]
[0119] C1(a j ,a k ) = a max -a k ≥ 0
[0120] C2(a j ,a k ) = a min -a k ≤ 0
[0121] C3(a j ,a k ) = v max -v k ≥ 0
[0122] C4(a j ,a k ) = v min -v k ≤ 0
[0123]
[0124]
[0125]
[0126] in, To follow the expected utility induced functional of the car; C1(a j ,a k ),C2(a j ,a k ),C3(a j ,a k ),C4(a j ,a k ),C5(a j ,a k All of these are inequality constraints following the car; D1(a) j ,a k ),D2(a j ,a k All of these are equality constraints that follow the car.
[0127] This invention addresses the subgame perfect Nash equilibrium problem in a multi-stage sequential game decision-making model for intelligent vehicles in a continuous action domain. It proposes a method combining classical optimization ideas and intelligent optimization algorithms. The aforementioned sequential game intelligent vehicle decision-making process can be expressed as follows: Figures 3-5 Schematic representation;
[0128] Figure 3 The expected utility-induced functional of the following car and the acceleration strategy of the leading car are represented by a. j And the acceleration strategy of following the car a k The relationship is shown by parameterizing the decision of the lead car with different acceleration strategies to slide slice the expected utility induced functional of the following car. It can be seen that the expected utility induced functional curves of the following car are different in the two slices. Fixing the acceleration strategy of the lead car, the following car adopting different acceleration strategies in the same slice also affects the optimality trend of its expected utility induced functional. In the process of optimizing its expected utility induced functional, the following car is simultaneously affected by the coupling effect of the acceleration strategy of the lead car and its own acceleration strategy. Figure 4 This represents the relationship between the acceleration strategy of the leading car and the optimal acceleration strategy of the following car. Figure 5 This indicates that when the following car has adopted an acceleration strategy that conforms to the steady state of the subgame's refined Nash equilibrium, the expected utility induced functional of the lead car changes with the lead car's acceleration strategy.
[0129] In step S4, the embodiment of the application optimizes the problem layer by layer and compensates the optimal solutions for each other from the perspective of hierarchical space coupling interaction, and the specific method flow is as follows:
[0130] In the continuous action domain of the leading vehicle, initial acceleration strategy particles satisfying the kinematic and dynamic constraints of the vehicle are randomly generated, and an initial particle group is formed, denoted as Each particle in the initial particle group is taken as a fixed parameter of the action of the leading vehicle, and a slice of the expected utility space of the following vehicle corresponding to each parameter is found; the optimal response acceleration strategy of the following vehicle corresponding to the action of the leading vehicle in each space slice is sequentially solved, and the reactive particle group of the following vehicle is denoted as The hierarchical particle group can be regarded as an initial random iterative solution of the subgame perfect Nash equilibrium target.
[0131] The utility degree of the initial particle group and the reactive particle group is measured by the expected utility induced functional of the leading vehicle and the expected utility induced functional of the following vehicle, and the utility degrees of the initial particle group and the reactive particle group are respectively sorted in descending order. In order to select excellent particle groups, the probability disc replication method is used to copy particles from the sorted initial particle group and the reactive particle group , denoted as the leading vehicle core particle group and the following vehicle core particle group The higher the utility degree of a particle, the greater the probability of being selected by the probability disc; the lower the utility degree of a particle, the smaller the probability of being selected by the probability disc, but there is still a possibility of being selected, which is beneficial to increase the diversity of the new particle group. The selection probability of the probability disc method is related to the relative utility degree of the particle, which can be represented as:
[0132]
[0133]
[0134] Where P j and P k represent the probability of the initial particle and the reactive particle being selected by the probability disc, respectively; and represent the utility degree of the initial particle and the reactive particle, respectively; and represent the cumulative utility degree of the initial particle group and the cumulative utility degree of the reactive particle group, respectively.
[0135] The replicated leader car core particle population and the follower car core particle population are encoded, and the acceleration strategy in decimal is encoded into a series of binary strings by the following relationship:
[0136]
[0137] wherein, and represent the interval left and right endpoints of the continuous domain game strategy set of the leader car and the follower car respectively; represents the length of the binary string encoding; represents the binary string form corresponding to a j,k . represents the conversion gate from the binary string to the acceleration strategy in decimal.
[0138] For the leader car core particle population and the follower car core particle population encoded into binary strings, further exchange operation and mutation operation are performed to generate new particles and wherein, the exchange operation exchanges the strings of different particles with each other, and the probability of exchange in the core particle population is P c , and the string position probability of the exchanged particle is represents the string of the new particle generated by cutting, exchanging and reassembling the string fragments from the rth position to the sth position. The mutation operation makes the fragments of the string of the same particle undergo 0 / 1 replacement, if a position in the string of the original particle is 0, it will be mutated to 1 after the mutation operation; if a position in the string of the original particle is 1, it will be mutated to 0 after the mutation operation. The probability of mutation in the core particle population is P m , and the string position probability of the mutated particle is represents the string of the new particle generated by the mutation operation on the string fragments from the kth position to the lth position. The string of the new particle generated by the exchange operation and the mutation operation is decoded into a new acceleration strategy, generating a new particle of the leader car acceleration strategy in the upper space and a new particle of the follower car acceleration strategy in the lower space. The above process of generating new particles is shown in Figure 6 .
[0139] Based on the new particle of the leader car acceleration strategy in the upper space, the optimal acceleration strategy of the follower car in the lower space is calculated. Here, κ is defined as the initial particle population The acceptance degree of the particle swarm, K is a form of proportion, which can be set flexibly according to the demand of the user, and the dimension of the acceptance degree of the particle swarm in the upper space leading car is Next, according to the relationship between the size of the new particle and the dimension of the acceptance degree of the particle swarm , the following several optimization methods are adopted as the judgment basis:
[0140] When the size of the new particle is greater than the dimension of the acceptance degree of the particle swarm , for each particle in the new particle of the leading car, the initial particle closest to it in the initial particle swarm is found by vector norm, and the optimal acceleration strategy of the follower car in the lower space corresponding to the initial particle is excavated in the hierarchical particle swarm , and the particle corresponding to the optimal acceleration strategy of the follower car is taken as the reference to construct a linear function near the particle, and the acceleration strategy of the follower car corresponding to each particle in the new particle of the leading car is approximately solved by the linear function, as shown in Figure 7 . The approximate solving process sacrifices the accuracy of the balanced strategy to some extent, but can effectively improve the calculation efficiency of the method and is beneficial to the deployment in the car controller for real-time operation. The particles corresponding to the acceleration strategy of the follower car obtained by the approximation in the embodiment of the application are screened, and the particles with an approximation error exceeding ε err are removed; and the particles with an approximation error not exceeding ε err are retained.
[0141] When the size of the new particle is less than the dimension of the acceptance degree of the particle swarm , similarly, for each particle in the new particle of the leading car, the initial particle closest to it in the initial particle swarm is found by vector norm, and the optimal acceleration strategy of the follower car in the lower space corresponding to the initial particle is excavated in the hierarchical particle swarm , and the particle corresponding to the optimal acceleration strategy of the follower car is taken as the starting point to construct a quadratic approximation function of the expected utility induced functional of the follower car in the lower space as a quasi-expected utility induced functional, and a linear approximation function of the constraint as a quasi-constraint function, and the quasi-expected utility induced functional and the quasi-constraint function are as follows:
[0142]
[0143]
[0144]
[0145] in, This represents the acceleration strategy of the lower-level space following the car during the iteration process; the quasi-expected utility induced functional and quasi-constraint function transform the problem into a quadratic programming problem, determining the new search direction and multiplier vector, based on the above hierarchical particle swarm. The particle points corresponding to the optimal acceleration strategy for following the car, extracted from the data, are subjected to a one-dimensional search along the search direction to obtain the search step size, and then a new approximate optimal acceleration strategy is calculated. This process is iterated until the convergence accuracy meets the specified requirements. The iteration stops, and the output follows the optimal acceleration strategy of the car. See details below. Figure 8 As shown.
[0146] When the above two methods fail to obtain the optimal acceleration response strategy for the vehicle in the lower space, new particles generating the lower space vehicle-following acceleration strategy using the aforementioned copying, swapping, and mutation operations are used. Each new particle is measured individually using the expected utility induced by the aforementioned established car-following functional. Utility, from the reactive particle swarm following the car and new particles The system comprehensively considers the utility values and selects the single particle with the optimal utility value as the optimal acceleration strategy for following the car in the lower space. To further improve accuracy, users can repeatedly perform copying, swapping, and mutation operations to generate more new particles that more closely resemble the actual subgame refined Nash equilibrium, thus selecting the optimal utility particle until a satisfactory result is achieved.
[0147] The particle corresponding to the optimal acceleration strategy for following the car in the lower space, calculated through the above three processes, is taken as the optimal particle; this optimal particle is then compared with the new particle generated by the leading car. In combination, the new particles of the leading car are measured using the expected utility-induced functional of the leading car established above. The combined utility of the best particle following the car is simply referred to as the combined utility of the leading new particle. From the initial particle swarm... utility and new particles Taking into account the overall utility, the particle with the optimal utility is selected as the optimal acceleration strategy for the lead car in the upper space, thus completing one round of calculation. Then, the satisfaction of the acceleration strategies of the lead car and the following cars is calculated:
[0148] Condition 1:
[0149] Condition 2:
[0150] Condition 3:
[0151] Condition 4:
[0152] Condition 5:
[0153] wherein, and denote the acceleration strategy of the upper space leader vehicle and the lower space follower vehicle in the (p+1)th iteration, respectively; and denote the acceleration strategy of the upper space leader vehicle and the lower space follower vehicle in the pth iteration, respectively; γ1, γ2 and γ3 denote the threshold values of the satisfaction conditions of the three categories, respectively; denotes the norm between the subgame perfect Nash equilibrium solutions obtained by the two rounds of calculation, respectively;
[0154] Condition 1 denotes the satisfaction condition between the subgame perfect Nash equilibrium solutions obtained by the two rounds of calculation, respectively; Condition 2 denotes the satisfaction condition of the increment of the expected utility induced functional of the leader vehicle obtained by the two rounds of calculation, respectively; Condition 3 denotes the satisfaction condition of the increment of the expected utility induced functional of the follower vehicle obtained by the two rounds of calculation, respectively; Condition 4 denotes the satisfaction condition of the rate of change of the expected utility induced functional of the leader vehicle obtained by the two rounds of calculation, respectively; Condition 5 denotes the satisfaction condition of the rate of change of the expected utility induced functional of the follower vehicle obtained by the two rounds of calculation, respectively;
[0155] If the above five conditions are met at the same time, it indicates that the acceleration strategy of the upper space leader vehicle and the acceleration strategy of the lower space follower vehicle obtained by the calculation have gradually converged to the subgame perfect Nash equilibrium state, and the solution task of the intelligent vehicle sequential game decision model of the instantaneous H continuous action domain is completed. If the above five conditions are not met at the same time, new particles and new particles repeat all the above processes.
[0156] In the above step S5, the advanced chassis system of the intelligent vehicle controls the driving force and the braking force, and executes the acceleration of the decision layer output which conforms to the instantaneous H consistent game stable equilibrium.
[0157] In the above step S6, the motion information, the prediction information and the cognition information of the intelligent vehicle and other vehicles in the traffic environment are updated synchronously in H+h;
[0158] In step S7, steps S1-S6 are repeatedly executed to solve the acceleration that meets the game stable equilibrium at each moment through the continuous multi-stage sequential game process and the rolling fast optimization strategy until the intelligent vehicle autonomously exits the game system and restores the free driving mode.
[0159] The embodiment of the application discloses an intelligent vehicle human-like decision-making method based on sequential game fast optimization, which is used for solving the hierarchical challenge, fast optimization problem and human-like decision-making problem of the multi-stage sequential game model in a driving scene involving multi-vehicle strong interaction. In order to make up for the rough solution caused by the discrete action domain and solve the hierarchical decision-making problem of the multi-stage sequential game, the application is extended to the general problem of continuous action domain multi-vehicle interactive decision-making, and a human-like decision-making method based on sequential game fast optimization is proposed by combining the classical optimization idea and the intelligent optimization algorithm for the fast optimization task of the driving decision-making in the hierarchical continuous space.
[0160] Obviously, those skilled in the art can make various modifications and variations to the application without departing from the spirit and scope of the application. Thus, if these modifications and variations of the application belong to the scope of the claims of the application and the equivalent technology thereof, the application also intends to include these modifications and variations.
Claims
1. An intelligent vehicle human-like decision-making method based on rapid optimization of sequential games, characterized in that, The application comprises the following steps: S1, the intelligent car and other cars in the traffic environment realize dynamic information interaction; S2, based on the dynamic information interaction, an intelligent car multi-stage sequential game decision model is established; S3, each car participating in the dynamic information interaction predicts the strategy to be taken by other cars in the future, establishes the expected utility induced functional of the leading car and the expected utility induced functional of the following car, and optimizes the expected utility induced functional of itself in the multi-stage sequential game according to the prediction result; S4, based on the expected utility induced functional of the leading car and the expected utility induced functional of the following car, the intelligent car multi-stage sequential game decision model is solved, and the acceleration of each car at H moment is output; S5, the intelligent car executes the acceleration of itself at H moment; S6, the motion information, prediction information and cognitive information of the intelligent car and other cars in the traffic environment are synchronously updated at H+h moment; S7, steps S1-S6 are repeatedly executed until the intelligent car autonomously exits the game system and restores the free driving mode; In S2, the intelligent car multi-stage sequential game decision model integrates the multi-objective cooperative driving demand of individual driving characteristics and the traffic system demand; The multi-objective cooperative driving demand includes the risk characteristics, impact characteristics and following characteristics of the intelligent car; The traffic system demand includes the ergonomics characteristics and safety characteristics of the traffic system; S4 specifically comprises: S41, generating initial acceleration strategy particles satisfying the kinematic and dynamic constraints of the car in the continuous action domain of the upper layer space leading car, and forming an initial particle group S42, obtaining the reactive particle swarm of each following vehicle based on the initial particle swarm of the upper space leader vehicle obtaining the reactive particle swarm of each following vehicle S43, based on the initial population of particles and the population of reaction particles constructing a hierarchical population of particles S44, measuring the utility degree of the expected utility inducing functional of the leading vehicle and the expected utility inducing functional of the following vehicle through the hierarchical particle population the initial particle population and the reactive particle population , and respectively independently ranking the utility degree of the initial particle population and the reactive particle population in descending order S45, copying out 2 particles from the sorted initial particle population and the reaction particle population respectively by using the probabilistic disc copying method, denoted as the leader car core particle population and the follower car core particle population S46, encoding the leader car core particle population and the follower car core particle population successively performing the encoding operation, the exchange operation, and the mutation operation to generate new particles corresponding to the upper space leader car acceleration strategy and the lower space follower car acceleration strategy S47, new particles according to the upper space leader car acceleration strategy and lower space follower car acceleration strategy calculate the optimal acceleration strategy for the corresponding lower space follower car; S48, the optimal acceleration strategy of the following car in the lower space is used to deduce the optimal acceleration strategy of the leading car in the upper space; S49, the satisfaction degree of the optimal acceleration strategy of the following car in the lower space and the optimal acceleration strategy of the leading car in the upper space is calculated, and the acceleration of each car at H moment is output.
2. The intelligent vehicle human-like decision-making method based on fast optimization of sequential games according to claim 1, wherein, In S1, the dynamic information interaction includes driving state information interaction, driving style interaction, expected utility induced index interaction and environment cognition interaction.
3. The intelligent vehicle human-like decision-making method based on fast optimization of sequential games according to claim 1, characterized in that, In S1, when the interactive information is incomplete, the intelligent car estimates the default information according to the belief and cognition of the traffic environment, and predicts the evolution of the future traffic situation.
4. The intelligent vehicle human-like decision-making method based on fast optimization of sequential games according to claim 1, wherein, S42 specifically comprises: Each particle in the initial particle group is taken as a fixed parameter of the action of the leading car, and the slice of the expected utility space of the following car corresponding to each parameter is found; solving the optimal response acceleration strategy corresponding to the action of the leading vehicle and the following vehicle in each spatial slice in turn, and obtaining the reactive particle swarm of each following vehicle 5. The intelligent vehicle human-like decision-making method based on fast optimization of sequential games according to claim 1, characterized in that, In the S47, a new particle according to the upper space leading vehicle acceleration strategy is calculated The optimal acceleration strategy of the corresponding lower space following vehicle is calculated, specifically including: acquiring the dimension of the swarm acceptance of the upper space leader vehicle judging the new particle and the dimension of the swarm acceptance between the relationship; (1) when the new particle has a size greater than the dimensionality of the particle population acceptance : For each particle in the new population of particles find the initial particle in the initial population of particles that is closest to it by vector norm. And in the hierarchical particle swarm Corresponding to the initial particle mining with the lower space following the optimal acceleration strategy of the car; A linear function near the particle corresponding to the optimal acceleration strategy of the following car is constructed with the particle as the reference; Solving the new particle by linear function approximation The lower space corresponding to each particle in the new particle follows the optimal acceleration strategy of the car; (2) when the new particle has a size smaller than the dimension of the particle population acceptance : For each particle in the new population of particles find the initial particle in the initial population of particles that is closest to it by vector norm. And in the hierarchical particle swarm Corresponding to the initial particle mining with the lower space following the optimal acceleration strategy of the car; A quadratic approximation function of the expected utility induced functional of the following car in the lower space is constructed as a quasi-expected utility induced functional with the particle corresponding to the optimal acceleration strategy of the following car as the starting point, and a linear approximation function of the constraint is constructed as a quasi-constraint function; Based on the quasi-expected utility induced functional and the quasi-constraint function, the problem is converted into a quadratic programming problem, and a new search direction and a multiplier vector are determined; based on the optimal acceleration strategy of the following car mined from the hierarchical particle swarm The particle point corresponding to the optimal acceleration strategy of the following car mined from the hierarchical particle swarm is searched along the search direction in one dimension to obtain a search step, and a new approximate optimal acceleration strategy is calculated. Iterate the above process until the convergence accuracy meets the preset condition to stop iteration, and output the new particles Each particle in the lower space corresponds to the optimal acceleration strategy of the car.
6. The intelligent vehicle human-like decision-making method based on fast optimization of sequential games according to claim 1, wherein, In the S47, the new particles according to the lower space following car acceleration strategy Calculate the optimal acceleration strategy of the corresponding lower space following car, specifically including: The new particles are measured one by one by the expected utility-induced functional of the following car. The utility of reactive particle swarm optimization is considered comprehensively. and new particles The utility of each particle is used to select the single particle with the optimal utility as the optimal acceleration strategy for following the car in the lower space.
7. The intelligent vehicle human-like decision-making method based on fast optimization of sequential games according to claim 1, wherein, S48 specifically comprises: The optimal particle in the lower space is the particle corresponding to the optimal acceleration strategy of the leading vehicle, and the optimal particle and the new particle are combined to measure the comprehensive utility of the optimal particle and the new particle by the expected utility induced functional of the leading vehicle ; and the optimal particle with the optimal utility is selected as the optimal acceleration strategy of the leading vehicle in the upper space by comprehensively considering the comprehensive utility and the utility of the initial particle group .
8. The intelligent vehicle human-like decision-making method based on fast optimization of sequential games according to claim 1, wherein, In S49, the conditions of the satisfaction degree include where, and respectively represent the acceleration strategy of the upper space leader vehicle and the acceleration strategy of the lower space follower vehicle in the (p+1)th iteration; and respectively represent the acceleration strategy of the upper space leader vehicle and the acceleration strategy of the lower space follower vehicle in the pth iteration; γ1, γ2 and γ3 respectively represent the threshold values of several kinds of satisfaction conditions; respectively represent the norm between the subgame perfect Nash equilibrium solutions obtained by two rounds of calculation; if the above five conditions are met at the same time, it indicates that the acceleration strategy of the upper space leader vehicle and the acceleration strategy of the lower space follower vehicle obtained by calculation have gradually converged to the subgame perfect Nash equilibrium state, and the solution task of the intelligent vehicle sequential game decision model in the instantaneous H-continuous action domain is completed; if the above five conditions are not met at the same time, new particles and new particles repeat all the above processes.