Control method and control device of automatic driving vehicle and vehicle

By obtaining and analyzing the current status, road and obstacle information of autonomous vehicles and determining and optimizing driving strategies, the problem of inaccurate vehicle planning methods in complex traffic scenarios is solved, and more accurate and efficient road planning is achieved.

CN120057033APending Publication Date: 2025-05-30SUZHOU ZHITU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311616698.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The road planning methods of the prior art autonomous vehicles in complex traffic scenarios are inaccurate, and it is difficult to output safe, reasonable and efficient decisions in ramp remittance and automatic road change scenarios.

Method used

By obtaining the current status information, road scene information and obstacle scene information of the controlled vehicle, multiple driving strategies are determined, and driving actions and probability distribution models are determined based on these strategies, confidence is calculated, and the state value function is established, and the vehicle is finally controlled to drive according to the driving strategy corresponding to the highest state value function value.

Benefits of technology

In complex road situations, selecting the optimal driving strategy from multiple driving strategies will alleviate the problem of inaccurate vehicle road planning methods in the prior art and improve the accuracy of road planning methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120057033A_ABST
    Figure CN120057033A_ABST
Patent Text Reader

Abstract

The invention provides a control method and device of an automatic driving vehicle and the vehicle. The method comprises the following steps: acquiring current state information, road scene information and obstacle scene information of a controlled vehicle; determining a plurality of driving strategies according to the current state information, the road scene information and the obstacle scene information, determining a driving action according to the plurality of driving strategies, determining a probability distribution model of the driving action, and calculating a confidence coefficient according to the probability distribution model; establishing a state value function according to the confidence coefficient, taking the driving strategy corresponding to the highest state value function value as a target driving strategy, and controlling the controlled vehicle to drive according to the target driving strategy; and establishing a state value function according to the confidence coefficient, taking the driving strategy corresponding to the highest state value function value as a target driving strategy, and controlling the controlled vehicle to drive according to the target driving strategy. According to the invention, the accuracy of automatic driving vehicle road planning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of control of autonomous vehicles, and more particularly, to a control method, a control device, a computer-readable storage medium, and a vehicle for an autonomous vehicle. Background Art

[0002] With the continuous development of autonomous driving technology, many SAE 2 (Society of Automotive Engineers standard) level driving automation systems have been deployed in commercial vehicles, which can provide a more comfortable driving experience. By dynamically perceiving the environment and making decisions in advance, traffic accidents can be reduced. Automobile manufacturers and high-tech companies around the world are working hard on the development of higher-level autonomous driving systems. The control of autonomous vehicles on empty highways is relatively easy, but there are still many imperfections in the current solutions for more congested traffic scenarios with complex traffic flows. One of the problems is that in scenarios such as ramp merging and diverging, and automatic lane change, autonomous vehicles need to output safe, reasonable, and efficient decisions based on the predicted information of traffic participants to complete the trajectory planning of the vehicle in ramp and congestion scenarios. That is, for the planning method in uncertain scenarios, the current planning effect is still not perfect enough.

[0003] In the prior art, the following methods are usually adopted: (1) Rule-based decision-making method. By relying on rules, various state machines are designed. Considering from the extreme perspective, to achieve complete safety and comfort of the system, the system needs to perform a large number of samplings in a high-dimensional continuous state space and a continuous action space to find the optimal strategy. However, the current hardware computing power platform and embedded real-time system cannot achieve the real-time performance of high-density computing, which is a systematic problem of software and hardware. Even if a hierarchical rule-based decision-making method is adopted, the rule parameters are derived from human driving experience and require a large amount of debugging and optimization, and the generalization ability is extremely poor. (2) The POMDP (Partially Observable Markov Decision Process) method is used to make decisions in an uncertain environment. Monte Carlo search or particle filter is used to search the state space of the ego vehicle and obstacles to generate state transition probabilities, and search in continuous state spaces such as position, speed, and acceleration. A large amount of computing power is required for this process. Therefore, the mainstream method performs pruning optimization and makes state transitions in the pruned space. However, currently, these methods still do not meet the real-time requirement of decision-making, and the time consumption for one decision is in seconds. Since the planning system requires real-time decision-making information, the low real-time performance makes the POMDP method unable to be applied in actual production. (3) The reinforcement learning method is adopted, which requires a large amount of data training at the same time and does not have interpretability. To achieve reasonable, efficient, and safe decision-making output, a large amount of data is required to train the network, lacking safety analysis of the output decision, and there is a possibility of dangerous decisions. (4) The vehicle networking method is adopted, which is costly and not universal. The vehicle networking needs to transform the field end and the vehicle end. By transforming, the environmental information around the vehicle is obtained. The transformation of the factory end requires gradual improvement over time. Currently, it is only piloted locally and cannot be promoted to the market. The transformation of the vehicle end requires a unified, reliable, and universal vehicle networking protocol and standard. Currently, the vehicle networking industry is constantly developing and integrating, and the standard is still not mature. (5) The spatio-temporal trajectory planning method is adopted. In the Cartesian coordinate system, the obstacle method is often used to add non-linear constraints to the objective function. For example, in the prior art, the interior point obstacle penalty function, an interior point method, is used. The condition number of the Hessian matrix of the penalty function will tend to infinity, and the solution will become very difficult. This problem of the Hessian matrix becoming ill-conditioned exists for both interior penalty functions and exterior penalty functions. Therefore, there is a possibility of trajectory solution failure when using the penalty function method.

[0004] Therefore, a planning method for autonomous vehicles that overcomes the above problems is needed. Summary of the Invention

[0005] The main purpose of this application is to provide a control method, a control device, a computer-readable storage medium, and a vehicle for autonomous vehicles, so as to at least solve the problem that the road planning method for autonomous vehicles in the prior art is inaccurate.

[0006] To achieve the above object, according to one aspect of the present application, there is provided a control method for an autonomous vehicle, including: obtaining the current state information, road scene information, and obstacle scene information of the controlled vehicle, where the current state information is the position and speed of the controlled vehicle at the current moment, the road scene information represents the road condition information where the controlled vehicle is located, and the obstacle scene information represents the information of vehicles other than the controlled vehicle on the road and at least includes the position and speed; determining a plurality of driving strategies according to the current state information, the road scene information, and the obstacle scene information, determining a driving action according to the plurality of driving strategies, determining a probability distribution model of the driving action, and calculating a confidence level according to the probability distribution model, where each driving strategy includes the driving actions at multiple moments, and the probability distribution model represents the probability that the controlled vehicle drives according to the driving strategy; establishing a state value function according to the confidence level, taking the driving strategy corresponding to the highest state value function value as the target driving strategy, and controlling the controlled vehicle to drive according to the target driving strategy, where the state value function represents the pros and cons of each driving strategy, and the higher the state value function value, the better the corresponding driving strategy.

[0007] Optionally, determining a driving action according to the plurality of driving strategies includes: obtaining the lateral semantic action and longitudinal semantic action of the controlled vehicle, mapping the lateral semantic action and the longitudinal semantic action to a multi-dimensional state space, where the lateral semantic action is a word representing the lateral action made by the controlled vehicle and includes lane keeping, left lane change, and right lane change, and the longitudinal semantic action is a word representing the longitudinal action made by the controlled vehicle and includes acceleration, deceleration, and constant speed; using each driving strategy to search in the multi-dimensional state space to determine the driving action corresponding to each moment in the driving strategy.

[0008] Optionally, determining the probability distribution model of the driving action includes: obtaining the controlled vehicle state variable and the obstacle decision variable, where the controlled vehicle decision variable is a variable representing the action of the controlled vehicle, and the obstacle decision variable is a variable affecting the decision-making of the obstacle; taking the controlled vehicle state variable and the obstacle decision variable as the features of a dynamic Bayesian network, establishing a dynamic Bayesian network prediction model, and updating the dynamic Bayesian network prediction model to obtain the probability distribution model.

[0009] Optionally, calculating the confidence level according to the probability distribution model includes: calculating the action probability of the controlled vehicle according to the probability distribution model, and calculating the action probability of the obstacle; calculating the product of the action probability of the controlled vehicle and the action probability of the obstacle to obtain the confidence level.

[0010] Optionally, controlling the controlled vehicle to travel according to the target driving strategy includes: establishing a kinematic model of the control vehicle in the Cartesian coordinate system and obtaining constraint conditions, and establishing an objective function according to the kinematic model and the constraint conditions; solving the objective function to obtain control variables and optimization variables corresponding to the target driving strategy, and controlling the vehicle to travel according to the control variables and the optimization variables, so that the controlled vehicle travels according to the target driving strategy, where the control variables include the steering wheel angle change rate and the acceleration, and the optimization variables include the position, speed, body heading angle, and steering wheel angle.

[0011] Optionally, solving the objective function includes: establishing an augmented Lagrangian equation according to the objective function and the constraint conditions, and performing iterative solution on the augmented Lagrangian equation; stopping the iteration when the augmented Lagrangian equation satisfies the target convergence condition, and obtaining the control variables and the optimization variables.

[0012] Optionally, establishing a state value function according to the confidence includes: establishing a partially observable Markov decision process model; substituting the confidence into the partially observable Markov decision process model to obtain the state value function.

[0013] According to another aspect of the present application, a control device for an autonomous vehicle is provided, including: an acquisition unit, configured to acquire the current state information, road scene information, and obstacle scene information of the controlled vehicle, where the current state information is the position and speed of the controlled vehicle at the current moment, the road scene information represents the road condition information where the controlled vehicle is located, and the obstacle scene information represents the information of vehicles other than the controlled vehicle on the road and at least includes the position and speed; a calculation unit, configured to determine multiple driving strategies according to the current state information, the road scene information, and the obstacle scene information of the controlled vehicle, determine a driving action according to the multiple driving strategies, determine a probability distribution model of the driving action, and calculate a confidence according to the probability distribution model, where each driving strategy includes the driving actions at multiple moments, and the probability distribution model represents the probability that the controlled vehicle travels according to the driving strategy; a control unit, configured to establish a state value function according to the confidence, use the driving strategy corresponding to the highest state value function value as the target driving strategy, and control the controlled vehicle to travel according to the target driving strategy, where the state value function represents the pros and cons of each driving strategy, and the higher the state value function value, the better the corresponding driving strategy.

[0014] According to another aspect of the present application, there is provided a computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the above control methods.

[0015] According to yet another aspect of the present application, there is provided a vehicle, including: one or more processors, a memory, and one or more programs. The one or more programs are stored in the memory and are configured to be executed by the one or more processors. The one or more programs include those for executing any one of the above control methods.

[0016] Applying the technical solution of the present application, based on the current state information, road scene information, and obstacle scene information of the vehicle, multiple possible driving strategies are determined. Then, the driving action, probability distribution model, and confidence are determined. After that, a state value function is established, and the state value function value is calculated. The vehicle is controlled to drive according to the driving strategy corresponding to the highest state value function value. In this way, in complex road conditions, the optimal driving strategy can be selected from multiple driving strategies to control the driving of the vehicle. Compared with the inaccurate road planning methods of vehicles in complex road scenarios in the prior art, the present application can select the optimal driving strategy from multiple driving strategies determined by the current state information, road scene information, and obstacle scene information of the vehicle. Therefore, it can alleviate the problem of inaccurate vehicle road planning methods in the prior art and improve the accuracy of road planning methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0018] Figure 1 It shows a hardware structure block diagram of a mobile terminal for executing a control method of an autonomous vehicle provided by an embodiment of this application;

[0019] Figure 2 It shows a schematic flowchart of a control method of an autonomous vehicle provided by an embodiment of this application;

[0020] Figure 3 It shows a schematic flowchart of a specific control method of an autonomous vehicle provided by an embodiment of this application;

[0021] Figure 4 It shows a schematic flowchart of generating a key scene in a specific control method of an autonomous vehicle provided by an embodiment of this application;

[0022] Figure 5Shows a schematic flow diagram of action space search in a specific control method for an autonomous driving vehicle provided by an embodiment of the present application;

[0023] Figure 6 Shows a schematic flow diagram of generating an optimal decision in a specific control method for an autonomous driving vehicle provided by an embodiment of the present application;

[0024] Figure 7 Shows a schematic diagram of a vehicle planning model in a specific control method for an autonomous driving vehicle provided by an embodiment of the present application;

[0025] Figure 8 Shows a structural block diagram of a control device for an autonomous driving vehicle provided by an embodiment of the present application.

[0026] Among them, the above-mentioned drawings include the following reference numerals:

[0027] 102, processor; 104, memory; 106, transmission device; 108, input / output device. Detailed implementation manners

[0028] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0029] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of the present application here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] As introduced in the background art, the prior art method for road planning of autonomous vehicles is inaccurate. To solve the problem of inaccurate road planning method for autonomous vehicles, embodiments of the present application provide a control method, a control device, a computer-readable storage medium, and a vehicle for an autonomous vehicle.

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0033] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a control method of an autonomous vehicle according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in Figure 1 a processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown, or have a different configuration from

[0034] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the control method of the autonomous vehicle in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0035] In this embodiment, a control method for an autonomous vehicle running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0036] Figure 2 is a flowchart of the control method for an autonomous vehicle according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0037] Step S201, obtain the current state information, road scene information, and obstacle scene information of the controlled vehicle, where the above-mentioned current state information is the position and speed of the controlled vehicle at the current moment, the above-mentioned road scene information represents the road condition information where the controlled vehicle is located, and the above-mentioned obstacle scene information represents the information of vehicles other than the controlled vehicle on the road and at least includes the position and speed;

[0038] Specifically, first generate a key scenario based on the obtained current state information, road scenario information, and obstacle scenario information: (1) Identify the road scenario through map information or the perception information of road signs. This scenario is an actual geographical scenario information generated based on geographical information, that is, the above-mentioned road scenario information, such as ramp merging and diverging, and the scenario where there are merging vehicles in adjacent lanes; (2) Obtain the obstacle intention information from the vehicle intention module, that is, the above-mentioned obstacle scenario information, such as vehicle lateral lane change, longitudinal acceleration and deceleration information; (3) Obtain the position, speed, etc. of the obstacle from the upstream perception module, and at the same time add the lane attribute of the lane where the obstacle is located; (4) In the geographical scenario, generate a decision environment according to the vehicle decision-making strategy and the intention of the obstacles around the vehicle, and this environment serves as the basis for generating the driving strategy.

[0039] Step S202, determine multiple driving strategies according to the above-mentioned current state information, the above-mentioned road scenario information, and the above-mentioned obstacle scenario information, determine the driving actions according to the multiple above-mentioned driving strategies, determine the probability distribution model of the above-mentioned driving actions, and calculate the confidence level according to the above-mentioned probability distribution model. Each of the above-mentioned driving strategies includes the above-mentioned driving actions at multiple moments, and the above-mentioned probability distribution model represents the probability that the controlled vehicle drives according to the above-mentioned driving strategy.

[0040] Specifically, the driving strategy is, for example: go straight at time T1, turn right at time T2, and turn left at time T3. That is, a series of driving strategies composed of time and the actions at each time. After generating multiple decision environments, it is necessary to select the optimal decision among the multiple decision environments, that is, it is necessary to calculate the confidence level through the probability distribution model. Since the problem form constructed based on the Markov decision method requires searching in the action space.

[0041] Step S203, establish a state value function according to the above-mentioned confidence level, use the above-mentioned driving strategy corresponding to the highest state value function value as the target driving strategy, and control the above-mentioned controlled vehicle to drive according to the above-mentioned target driving strategy. The above-mentioned state value function represents the quality of each of the above-mentioned driving strategies. The higher the above-mentioned state value function value, the better the corresponding above-mentioned driving strategy.

[0042] Specifically, then perform probability distribution calculation, establish a state value function and determine the optimal state value function to obtain the optimal strategy, that is, the target driving strategy, and then control the vehicle to drive according to the optimal driving strategy.

[0043] Through this embodiment, multiple possible driving strategies can be determined based on the current state information, road scene information, and obstacle scene information of the vehicle. Then, the driving action, probability distribution model, and confidence level are determined. After that, the state value function is established, and the state value function value is calculated to control the vehicle to drive according to the driving strategy corresponding to the highest state value function value. In this way, in complex road conditions, the optimal driving strategy can be selected from multiple driving strategies to control the driving of the vehicle. Compared with the inaccurate road planning method of vehicles in complex road scenarios in the prior art, the present application can select the optimal driving strategy from multiple driving strategies determined by the current state information, road scene information, and obstacle scene information of the vehicle. Therefore, it can alleviate the problem of inaccurate vehicle road planning methods in the prior art and improve the accuracy of the road planning method.

[0044] In the specific implementation process, the driving action can be determined according to the multiple driving strategies in step S202 through the following steps: Obtain the lateral semantic action and longitudinal semantic action of the controlled vehicle, and map the lateral semantic action and the longitudinal semantic action to a multi-dimensional state space. Among them, the lateral semantic action is a word representing the lateral action made by the controlled vehicle and includes lane keeping, left lane change, and right lane change. The longitudinal semantic action is a word representing the longitudinal action made by the controlled vehicle and includes acceleration, deceleration, and constant speed; Use each of the driving strategies to search in the multi-dimensional state space to determine the driving action corresponding to each moment in the driving strategy. This method can quickly determine the transition probability between different states by constructing a multi-dimensional state space and performing a search to determine the driving action corresponding to each driving strategy.

[0045] Specifically, the conversion between states needs to be achieved through actions. For example: The state S1 is: the vehicle speed is 20 m / s, and the action a1 is acceleration, then the state S1 is converted to the state S2, and the state S2 can be: the vehicle speed is 40 m / s. Suppose at time T, the host vehicle is in the state S1, and multiple driving strategies are used to search for the action at time T in the N-dimensional state space. When the action a1 is selected, the state at time T+1 is selected according to the state transition probability. There will be a large number of sampling searches in the whole process. In the constructed POMDP (Partially Observable Markov Decision Process), the continuous action space is simplified, semantic actions are used for modeling, and N driving actions are discretized. The longitudinal semantic actions of autonomous driving are defined as three semantic actions: acceleration, deceleration, and constant speed; The lateral semantic action tree is defined as lane keeping, left lane change, and right lane change; Semantic actions can be added or reduced according to the platform computing power; Map the semantic actions to the state space, and at the same time ensure that the semantic actions do not change frequently in the decision sequence.

[0046] In order to accurately determine the action intentions of the controlled vehicle and the obstacle on the road, in some alternative embodiments, the probability distribution model for determining the above-mentioned driving actions in the above-mentioned step S202 can be implemented through the following steps: Obtain the controlled vehicle state variables and the obstacle decision variables, where the above-mentioned controlled vehicle decision variable is a variable representing the actions of the controlled vehicle, and the above-mentioned obstacle decision variable is a variable affecting the decision-making of the above-mentioned obstacle; Use the above-mentioned controlled vehicle state variables and the above-mentioned obstacle decision variables as the features of the dynamic Bayesian network, establish a dynamic Bayesian network prediction model, and update the above-mentioned dynamic Bayesian network prediction model to obtain the above-mentioned probability distribution model. By establishing the probability distribution model through the above steps, the action intentions of the controlled vehicle and the obstacle can be predicted respectively, so as to better control the operation of the controlled vehicle.

[0047] Specifically, a probability graph model (PGM) is used to estimate the probability distribution. A dynamic Bayesian network is built by establishing a directed graph network to calculate the probability distribution of the obstacle state intention. Among the selection of dynamic Bayesian network features, the characteristic variables affecting the lateral and longitudinal decisions of the obstacle include the angle between the heading angle and the lane, the collision time with the obstacle in front, the average speed with the obstacle in front, and the collision time with the obstacles in the adjacent left and right lanes (assuming in the same lane). Each of the four features corresponds to two states, so the combined state space of the features is 16 states. The vehicle state space is divided into lateral and longitudinal, corresponding to the lateral and longitudinal actions in POMDP. Lateral state: going straight and changing lanes; Longitudinal state: accelerating, moving at a constant speed, and decelerating. Thus, a vehicle state equation can be established. The dynamic Bayesian network is divided into two stages: prediction and update. Probability distribution models can be established in the prediction stage and the update stage respectively. Establish the vehicle state equation:

[0048]

[0049]

[0050] Among them, X t represents the vehicle state, and the state includes the position relative to the center line of the lane and the lateral speed; Y t is the measurement output or the so-called observation value; A and B are the system state transition matrices, and C is the output matrix; ε t and η t are Gaussian noises, and the Q and R matrices are the variances of the Gaussian noises; The dynamic Bayesian network is divided into two stages: prediction and update. Prediction stage: Probability distribution of vehicle characteristic states

[0051]

[0052] Among them, Z represents the environmental feature state, i.e., the combination of various features of surrounding obstacles, and the features are represented by E t ; M represents the state space of the vehicle. Horizontally, there are two states: going straight and changing lanes. Vertically, there are three states: accelerating, decelerating, and maintaining a constant speed. It represents that the environmental state changes from Z at time t - 1 t-1 to Z at time t t while the vehicle state changes from M at time t - 1 t-1 to the vehicle state M at time t t The joint probability; therefore is obtained by the product of P(Z t |Z t-1 ), P(M t |Z t , M t-1 ). This distribution is the value calculated in the previous cycle.

[0053] Among them P(Z t |Z t-1 ), P(M t |Z t , M t-1 ), these four probability distributions are respectively: It represents the joint probability that the vehicle state is M at time t - 1 t-1 and the environmental state is Z t-1 . P(Z t |Z t-1 ) represents the conditional probability distribution that the environmental state changes from Z at time t - 1 t-1 to Z at time t t . P(M t |Z t , M t-1 ) represents the conditional probability that when the vehicle state is M at time t - 1 t-1 and the environmental state is Z t , it changes to the environmental state M t ; Vehicle state transition probability distribution:

[0054]

[0055]

[0056] Among them, It represents that when the vehicle state is M at time t - 1 t-1 , and the vehicle state is M at time t t , the position and speed of the vehicle follow a Gaussian distribution with a certain mean and variance; represents the mean of the Gaussian distribution, represents the variance of the Gaussian distribution.

[0057] Update phase:

[0058]

[0059] Among them, P(Y t |M t ,M t-1 ) represents the probability distribution of the observed value when the vehicle state at time t - 1 is M t-1 , and the vehicle state at time t is M t ;

[0060] Vehicle characteristic state probability distribution:

[0061]

[0062] Among them, represents the joint probability that the vehicle state transfers from M t-1 at time t - 1 to M t at time t when the environmental state updates from Z t-1 at time t - 1 to Z t at time t, preparing for the next calculation; P(E t |Z t ) represents the probability of a certain characteristic observed when the environmental state at time t is Z t , where E t represents the angle between the obstacle orientation and the current lane, the collision time between the obstacle and the obstacle in front, the average speed of the obstacle and the obstacle in front, and the collision time between the obstacle and the adjacent obstacles on the left and right.

[0063] Vehicle state transfer calculation:

[0064]

[0065] Among them, represents a Gaussian distribution with a certain mean and variance that the position and speed of the vehicle follow when the vehicle state at time t - 1 is M t-1 , and the vehicle state at time t is M t , which is mainly for the prediction process, i.e., preparing for the next calculation;

[0066] Preparing for the next prediction process, substituting the above calculation results can obtain the prediction result at the next moment:

[0067]

[0068]

[0069]

[0070]

[0071] In some alternative embodiments, calculating the confidence level according to the above probability distribution model in step S202 can be implemented through the following steps: calculating the action probability of the controlled vehicle according to the above probability distribution model, and calculating the action probability of the obstacle; calculating the product of the action probability of the controlled vehicle and the action probability of the obstacle to obtain the confidence level. This method calculates the confidence level through the action probability obtained by the above probability distribution model, so that the confidence level can be accurately calculated.

[0072] In the specific implementation process, the partially observable Markov decision process (POMDP) method is used to construct a POMDP model, generate a decision sequence of the controlled vehicle, set the state set as S, the action set as A, the reward function as R, the discount factor as γ, the state transition probability as P, and the state value function as V(s). According to the Bellman state value function equation, the expectation of the value function on the trajectory taken under the guidance of the policy can be obtained when a given policy π is given, where s’ represents the state at the next time step:

[0073]

[0074] Optimal state value function:

[0075]

[0076] State transition function:

[0077]

[0078] Among them, represents the action of the host vehicle, represents the action of the k-th obstacle, and A k represents the action set of the k-th obstacle; semantic actions are used to map continuous actions, where is the semantic intention of the k-th obstacle,

[0079]

[0080]

[0081] Replace the state transition probability s'~p(.|s,a) from the current state to the next state to obtain

[0082]

[0083] If all environmental states are observable and the model is known, the optimal state value function can be obtained through the above formula. However, since not all environmental states can be observed, the posterior probability of the state, i.e., the confidence, is used to calculate the optimal state value function. Then the above formula is modified to

[0084]

[0085] where B is the confidence space of the state. Using the confidence to model POMDP through the above formula requires calculating the confidence b, and its calculation method is as follows:

[0086]

[0087] where t E is the time period of the decision sequence.

[0088] To accurately control the vehicle's driving, the above step S203 of controlling the controlled vehicle to drive according to the above target driving strategy can be achieved through the following steps: establish the kinematic model of the controlled vehicle in the Cartesian coordinate system and obtain the constraint conditions, and establish the objective function according to the above kinematic model and the above constraint conditions; solve the above objective function to obtain the control variables and optimization variables corresponding to the above target driving strategy, and control the vehicle to drive according to the above control variables and the above optimization variables, so that the controlled vehicle drives according to the above target driving strategy. Among them, the above control variables include the steering wheel angle change rate and acceleration, and the above optimization variables include position, speed, vehicle body heading angle, and steering wheel angle. This method controls the vehicle to drive according to the optimal driving strategy by establishing a kinematic model, so that the vehicle's driving can be accurately controlled.

[0089] Specifically, after generating the optimal strategy, the planning model of the controlled vehicle is established. The entire planning process uses the Cartesian coordinate system, so the kinematic model of the vehicle is a non-linear model. At the same time, the bicycle model is used to model the vehicle model, and the vehicle kinematic equation, where X x represents the Cartesian coordinate position x, X y represents the Cartesian coordinate position y, θ represents the vehicle orientation angle, v represents the vehicle speed, δ represents the vehicle steering wheel angle, and L represents the wheelbase between the front and rear wheels of the vehicle:

[0090]

[0091]

[0092] The equality constraint composed of the state transition equation is as follows: X = [X x X y θ v δ] = f(X,U), and the acceleration constraint is The road boundary constraint uses the Lp distance to describe the distance constraint between the ego vehicle and the road boundary. The Lp distance expression is as follows:

[0093]

[0094] In this method, the 3rd order is used for the numerical expression of the distance as follows: Among them, γ i is the angle between the line connecting the point and the two ends of the line segment. For the obstacle constraint, an ellipse is used to represent the obstacle, and the ellipse equation is constructed:

[0095]

[0096] The objective function is divided into four parts, namely: the cost of the vehicle deviating from the center line of the lane The cost of the vehicle acceleration The cost of the change rate of the steering wheel angle The cost of the vehicle orientation The total objective function is as follows:

[0097]

[0098] In some alternative embodiments, to solve the above objective function, it can be achieved through the following steps: establish an augmented Lagrangian equation according to the above objective function and the above constraint conditions, and perform iterative solution on the above augmented Lagrangian equation; when the above augmented Lagrangian equation satisfies the objective convergence condition, stop the iteration to obtain the above control variables and the above optimization variables. This method uses the augmented Lagrangian method to solve the above non - linear equality constraints, so as to accurately perform the solution.

[0099] Specifically, the kinematic model, that is, the equality constraints, inequality constraints, and objective function composed of the state transition equation are analyzed, and finally the vehicle position, speed, body heading angle, steering wheel angle, as well as the change rate of the steering wheel angle, acceleration, etc. are obtained. Since there are non - linear equality constraints in the process of vehicle kinematic modeling, and both the boundary constraint and the obstacle constraint are non - linear constraints, the augmented Lagrangian differential dynamic programming method is used for solution; the optimal decision is used as the rough trajectory to participate in the solution of the spatio - temporal trajectory generation. By providing a warm start through the rough trajectory, the convergence speed of the augmented Lagrangian differential dynamic programming (AL - DDP) can be accelerated. The AL - DDP method is used to solve the problem, and an augmented Lagrangian equation is constructed for the objective function and constraints in the model in step seven

[0100]

[0101] Among them, in the formula, L A represents the augmented Lagrangian function of the optimization problem, and x kDenote the state, u k Denote the control, λ N Denote the Lagrange multiplier, I u,k Denote the Lagrange penalty factor, c N Denote the terminal constraint, c k Denote the constraints at times 0-(N-1), where the constraints include equality constraints and inequality constraints, l k Denote the objective function from time 0 to N-1. Construct the Cost-to-go iterative equation

[0102]

[0103] Among them, V N (x N ) denotes the state value function at the terminal, V k (x k ) denotes the state value function from time 0 to N-1. Since it is an iterative process, it is necessary to control the perturbation amount of the reference value. Considering the approximation accuracy, perform a second-order Taylor expansion on the constraints and objective function of time step N

[0104]

[0105] Among them

[0106]

[0107]

[0108] Among them, δV k denotes the state value function generated by the perturbation; in this way, the action-state value function δQ can be obtained

[0109]

[0110] Among them,

[0111]

[0112]

[0113]

[0114]

[0115]

[0116] Among them, Q represents the perturbation state-action value function, Qxx represents the second-order derivative of the state-action value function with respect to state x, Quu represents the second-order derivative of the state-action value function with respect to action u, Qux represents the second-order partial derivative of the state-action value function with respect to state x and action u, and Qx and Qu respectively represent the first-order derivatives of the state-action value function with respect to state x and action u; minimize δQ k , according to the necessary condition for the extreme value that the first derivative is 0, that is, for δQ k Taking the first derivative can obtain the control that makes δQ k an extreme value

[0117]

[0118] The update methods of the Lagrange multiplier and the penalty factor are as follows:

[0119]

[0120]

[0121] Among them, λ ki represents the Lagrange multiplier, and μ ki represents the Lagrange penalty factor. The iteration terminates when the following convergence conditions are reached: reaching the maximum number of iterations; the cost function is less than a certain value; the feedback gain K is less than zero; the solution of the entire augmented Lagrangian DDP problem is divided into two parts: the differential dynamic programming process (DDP) and the update of the Lagrange multiplier and the penalty factor. Among them, DDP contains two steps, namely the forward vehicle passing and the backward process; the overall process of the augmented Lagrangian DDP, setting the initial values of λ, μ, and φ; the DDP process; updating the Lagrange multiplier; calculating the cost L A , if it is less than the set value, the loop ends, otherwise it returns to step 2 until the cost converges to less than the set value. Steps of the DDP process: Initialize the state variable X 0 and the control variable U 0 ; Execute the backward process of DDP to generate the feedback gain K t at each time step, execute the forward process of DDP to calculate the cost L A , if it is greater than the steps of the backward process of DDP: Set P N , p N values at time N; calculate δQ; judge Q uu >0, if it is greater than 0, update the feedback gain K, d and end the loop, output K, d, if it is less than 0, increase the ρ value, and jump to step 2 to continue the calculation; Steps of the DDP forward process: Set the initial value x 0 ; According to the feedback gain K and d output by the DDP backward process, calculate the control U at each time step through the control law, and the vehicle state X is calculated by the state transition equation at each time step; Output the sequence of control variables and state variables from time 0 to N

[0122] To establish the state value function, the above step S203 can be implemented through the following steps: establishing a partially observable Markov decision process model; substituting the above confidence into the above partially observable Markov decision process model to obtain the above state value function. This method establishes the state value function through the partially observable Markov decision process model, so that the state value functions of each decision can be accurately obtained.

[0123] Specifically, as mentioned above, the partially observable Markov decision process (POMDP) method is used to generate the ego-vehicle decision sequence, and then the state value function is established. The specific process is as above.

[0124] To enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the control method for an autonomous vehicle of the present application will be described in detail below in conjunction with specific embodiments.

[0125] This embodiment relates to a specific control method for an autonomous vehicle, which provides safe, efficient and stable decision-making information for the autonomous vehicle in a complex dynamic environment in an embedded platform, and completes vehicle trajectory planning based on the iterative LQR method in the Cartesian three-dimensional space. The decision generation method is mainly based on the Markov decision process (MDP), uses the partially observable Markov process (POMDP) as the decision problem modeling method, and simplifies the action space searched by the POMDP at the same time. The overall process is as Figure 3 shown: obtaining the environmental vehicle encoding; generating actions through the action space generation module and the intention generation module; outputting the simulated actions through the action simulation module; and then outputting the lateral and longitudinal semantic actions through the safety inspection module, corresponding to different planners, and converting them into optimization constraints; accessing the planning algorithm interface for path planning. A specific control method for an autonomous vehicle includes the following steps:

[0126] Step S1: Generate key scenarios. As Figure 4 shown, determine the lane information according to the map information, such as large curvature bends or long uphill and downhill sections, etc., and determine the map scenarios, such as cruise roads, the presence of merging lanes in adjacent lanes of the main road, and ramp merging and diverging, road diverging and merging, etc., obtain the perception information, input it into the obstacle intention update module, and determine the key scenarios. For example, the scenario corresponding to a cruise road is cruise, the scenario corresponding to the presence of a merging lane in an adjacent lane of the main road is following / overtaking, and the scenarios corresponding to ramp merging and diverging, road diverging and merging are automatic lane change. All key scenarios generate decision sequences. The decision sequences are, for example, a series of action decisions composed of a time series such as accelerating straight at time T1, decelerating and turning left at time T2, and accelerating and turning right at time T3, etc.

[0127] Step S2: Construct the semantic action trees for the ego vehicle and obstacles. In step S1, multiple decision-making environments are generated. It is necessary to select the optimal decision among multiple decision-making environments. Since the problem form constructed based on the Markov decision-making method requires searching in the action space, as Figure 5 shown, at time T, the ego vehicle is in state S 1 . Use the policy π to search for the action at time T in the N-dimensional state space. When selecting action a 1 , according to the state transition probability, select the state at time T+1 (which may be state S 2 , S 3 , S 4 , …… S t ). When selecting actions a 2 , a 3 , a 4 , ……, a n , the steps are the same. There will be a large number of sampling searches in the whole process. Simplify the continuous action space in the constructed POMDP, use semantic actions for modeling, discretize a total of N states from a 1 to a n . The longitudinal semantic actions of autonomous driving are defined as three semantic actions: acceleration, deceleration, and constant speed; the lateral semantic action tree is defined as lane keeping, left lane change, and right lane change; semantic actions can be added or reduced according to the platform computing power; map the semantic actions to the state space, and at the same time ensure that the semantic actions do not change frequently in the decision sequence;

[0128] Step S3: Probability distribution calculation. Use the probability graph model (PGM) to estimate the probability distribution. Build a dynamic Bayesian network by establishing a directed graph network to calculate the probability distribution of the obstacle state intention, and bring the generated obstacle intention distribution into the confidence update formula in the POMDP to generate the confidence of each time step;

[0129] Step S4: Construct the POMDP model. Use the partially observable Markov decision process (POMDP) method to generate the decision sequence of the ego vehicle. Set the state set as S, the action set as A, the reward function as R, the discount factor as γ, the state transition probability as P, and the state value function as V(s). According to the Bellman state value function equation, the expectation of the value function on the trajectory taken under the guidance of the given policy π can be obtained, and the optimal state value function is established:

[0130] Step S5: Semantic space search, generate the optimal decision π. As Figure 6 shown, update and initialize the ego vehicle state and obstacle intention information, generate key scenarios, create a thread pool parallel computing allocation module. Thread 1: Policy π 1 , policy report R 1 ; Thread 2: Policy π 2 , policy report R2 ; Thread 3: Policy π 3 , Policy report R 3 ;... Thread n: Policy π n , Policy report R n ; Finally, generate a return value queue and produce the optimal policy π; that is, according to the scenarios generated in step S1, traverse the ego-vehicle policies π generated in step S2, search in each scenario, calculate the state value function of each policy to obtain the optimal policy π. To accelerate the decision-making generation efficiency, a parallel computing method is used for calculating the return value of each policy search. The return value is comprehensively calculated by the efficiency and safety evaluation functions;

[0131] Step S6: Establish a planning model. As Figure 7 shown, the entire planning process coordinate system uses a Cartesian coordinate system. Therefore, the kinematic model of the vehicle is a nonlinear model. At the same time, a bicycle model is used to model the vehicle model. Using the Cartesian coordinate system, it is shown as follows:

[0132]

[0133]

[0134] where X x represents the position x in the Cartesian coordinate system, X y represents the position y in the Cartesian coordinate system, θ represents the vehicle orientation angle, v represents the vehicle speed, δ represents the vehicle steering wheel angle, L represents the wheelbase between the front and rear wheels of the vehicle. The objective function is divided into four parts, namely: the cost of the vehicle deviating from the center line of the lane the cost of vehicle acceleration the cost of the change rate of the steering wheel angle the cost of vehicle orientation The total objective function is as follows:

[0135]

[0136] Step S7: Establish an optimal control (OCP) model. The optimization variables X are: position, speed, body heading angle, steering wheel angle, and the control variables U are: the change rate of the steering wheel angle, acceleration;

[0137] Step S8: Solve the augmented Lagrangian DDP problem. Through Step S7, the modeling of the vehicle's longitudinal and lateral trajectory planning problem in the Cartesian coordinate system is completed. Since there are non-linear equality constraints in the process of vehicle kinematic modeling, and both boundary constraints and obstacle constraints are non-linear constraints, the augmented Lagrangian differential dynamic programming method is used for solving; the output of Step S5 and the optimal decision π are used as the rough trajectory to participate in the solution of the spatio-temporal trajectory generation. By providing a warm start through the rough trajectory, the convergence speed of the augmented Lagrangian differential dynamic programming (AL-DDP) can be accelerated. The AL-DDP method is used to solve the problem. An augmented Lagrangian equation is constructed for the objective function and constraints in the model in Step Seven to meet the following convergence conditions and terminate the iteration: reaching the maximum number of iterations; the cost function is less than a certain value; the feedback gain K is less than zero; the solution of the entire augmented Lagrangian DDP problem is divided into two parts: the differential dynamic programming process (DDP) and the update of the Lagrange multiplier and penalty factor. Among them, DDP contains two steps, namely the forward vehicle passing and the backward process; the overall process of augmented Lagrangian DDP, set the initial values of λ, μ, φ; the DDP process; update the Lagrange multiplier; calculate the cost L A , if it is less than the set value, the loop ends, otherwise return to Step 2 until the cost converges to less than the set value. Steps of the DDP process: Initialize the state variable X 0 and the control variable U 0 ; Execute the backward process of DDP to generate the feedback gain K at each time step t , execute the forward process of DDP to calculate the cost L A , if it is greater than the steps of the DDP backward process: Set P N , p N values at time N; Calculate δQ; Judge Q uu >0, if it is greater than 0, update the feedback gain K, d and end the loop, output K, d, if it is less than 0, increase the value of ρ, jump to Step 2 to continue the calculation; Steps of the DDP forward process: Set the initial value x 0 ; According to the feedback gain K, d output by the DDP backward process, calculate the control U at each time step through the control law, and the vehicle state X is calculated at each time step by the state transition equation; Output the sequence of control variables and state variables from time 0 to N

[0138] The embodiments of the present application also provide a control device for an autonomous vehicle. It should be noted that the control device for an autonomous vehicle in the embodiments of the present application can be used to execute the control method for an autonomous vehicle provided by the embodiments of the present application. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0139] The following introduces the control device for an autonomous vehicle provided by the embodiments of the present application.

[0140] Figure 8 is a schematic diagram of the control device for an autonomous vehicle according to the embodiments of the present application. As Figure 8 shown, the device includes:

[0141] An acquisition unit 10, configured to acquire the current state information, road scene information, and obstacle scene information of the vehicle to be controlled. Among them, the above current state information is the position and speed of the vehicle to be controlled at the current moment, the above road scene information represents the road condition information where the vehicle to be controlled is located, and the above obstacle scene information represents the information of vehicles other than the vehicle to be controlled on the road and at least includes the position and speed;

[0142] Specifically, first generate a key scene based on the acquired current state information, road scene information, and obstacle scene information: (1) Identify the road scene through map information or the perceived information of road signs. This scene is an actual geographical scene information generated according to geographical information, that is, the above road scene information, such as ramp merging, exiting, and the scene of a vehicle merging into the adjacent lane; (2) Obtain the obstacle intention information from the vehicle intention module, that is, the above obstacle scene information, such as vehicle lateral lane change, longitudinal acceleration and deceleration information; (3) Obtain the position, speed, etc. of the obstacle from the upstream perception module, and at the same time add the lane attribute of the lane where the obstacle is located; (4) In the geographical scene, generate a decision-making environment according to the vehicle decision-making strategy and the intention of the obstacles around the vehicle, and this environment serves as the basis for generating the driving strategy.

[0143] A calculation unit 20, configured to determine multiple driving strategies according to the current state information, the road scene information, and the obstacle scene information of the vehicle to be controlled, determine a driving action according to the multiple driving strategies, determine a probability distribution model of the driving action, and calculate a confidence level according to the probability distribution model. Among them, each of the above driving strategies includes the above driving actions at multiple moments, and the probability distribution model represents the probability that the vehicle to be controlled travels according to the above driving strategy;

[0144] Specifically, the driving strategies are, for example: go straight at time T1, turn right at time T2, and turn left at time T3. That is, a series of driving strategies composed of time and the actions at each time. After generating multiple decision-making environments, it is necessary to select the optimal decision from the multiple decision-making environments, that is, it is necessary to calculate the confidence level through a probability distribution model. Since the problem form constructed based on the Markov decision-making method requires searching in the action space.

[0145] The control unit 30 is configured to establish a state value function according to the above-mentioned confidence level, use the driving strategy corresponding to the highest state value function value as the target driving strategy, and control the controlled vehicle to drive according to the target driving strategy. Wherein, the state value function represents the advantages and disadvantages of each driving strategy, and the higher the state value function value, the better the corresponding driving strategy.

[0146] Specifically, probability distribution calculation is performed later, a state value function is established and the optimal state value function is determined to obtain the optimal strategy, that is, the target driving strategy. Then, the vehicle is controlled to drive according to the optimal driving strategy.

[0147] Through this embodiment, multiple possible driving strategies can be determined according to the current state information, road scene information, and obstacle scene information of the vehicle. Then, the driving actions, probability distribution model, and confidence level are determined. Then, a state value function is established, and the state value function value is calculated. The vehicle is controlled to drive according to the driving strategy corresponding to the highest state value function value. In this way, in complex road conditions, the optimal driving strategy can be selected from multiple driving strategies to control the driving of the vehicle. Compared with the inaccurate road planning methods of vehicles in complex road scenarios in the prior art, the present application can select the optimal driving strategy from multiple driving strategies determined from the current state information, road scene information, and obstacle scene information of the vehicle. Therefore, the problem of inaccurate vehicle road planning methods in the prior art can be alleviated, and the accuracy of the road planning method can be improved.

[0148] In the specific implementation process, the calculation unit includes a mapping module and a determination module. The mapping module is configured to obtain the lateral semantic actions and longitudinal semantic actions of the controlled vehicle, and map the lateral semantic actions and the longitudinal semantic actions to a multi-dimensional state space. The lateral semantic actions are words representing the lateral actions made by the controlled vehicle and include lane keeping, left lane change, and right lane change. The longitudinal semantic actions are words representing the longitudinal actions made by the controlled vehicle and include acceleration, deceleration, and constant speed. The determination module is configured to search in the multi-dimensional state space using each driving strategy and determine the driving actions corresponding to each time in the driving strategy. This method can quickly determine the transition probability between different states by constructing a multi-dimensional state space and performing a search to determine the driving actions corresponding to each driving strategy.

[0149] Specifically, the transition between states needs to be achieved through actions. For example, state S1 is: the vehicle speed is 20 m / s, and action a1 is acceleration, then state S1 is converted to state S2, and state S2 can be: the vehicle speed is 40 m / s. Suppose at time T, the ego vehicle is in state S1, and multiple driving strategies are used to search for the action at time T in the N-dimensional state space. When action a1 is selected, the state at time T+1 is selected according to the state transition probability. There will be a large number of sampling searches in the whole process. In the constructed POMDP (Partially Observable Markov Decision Process), the continuous action space is simplified, semantic actions are used for modeling, and N driving actions are discretized. The longitudinal semantic actions for autonomous driving are defined as three semantic actions: acceleration, deceleration, and constant speed; the lateral semantic action tree is defined as lane keeping, left lane change, and right lane change; the semantic actions can be added or reduced according to the platform computing power; the semantic actions are mapped to the state space, and at the same time, it is ensured that the semantic actions do not change frequently in the decision sequence.

[0150] In order to accurately determine the action intentions of the controlled vehicle and the obstacle on the road, in some alternative embodiments, the above computing unit further includes an acquisition module and an update module. The acquisition module is used to acquire the state variables of the controlled vehicle and the decision variables of the obstacle. The decision variable of the controlled vehicle is a variable representing the action of the controlled vehicle, and the decision variable of the obstacle is a variable affecting the decision-making of the obstacle; the update module is used to use the state variables of the controlled vehicle and the decision variables of the obstacle as the features of the dynamic Bayesian network, establish a dynamic Bayesian network prediction model, and update the dynamic Bayesian network prediction model to obtain the probability distribution model. The method establishes the probability distribution model through the above steps, so as to be able to predict the action intentions of the controlled vehicle and the obstacle respectively, and thus better control the operation of the controlled vehicle.

[0151] Specifically, a probability graph model (PGM) is used to estimate the probability distribution. A dynamic Bayesian network is built by establishing a directed graph network to calculate the probability distribution of the obstacle state intention. In the selection of the dynamic Bayesian network features, the characteristic variables that affect the lateral and longitudinal decisions of the obstacle include the angle between the heading angle and the lane, the collision time with the obstacle in front, the average speed with the obstacle in front, and the collision time with the obstacles in the adjacent left and right lanes (assuming in the same lane). Each of the four features corresponds to two states, so the combined state space of the features is 16 states. The vehicle state space is divided into lateral and longitudinal, corresponding to the lateral and longitudinal actions in POMDP. Lateral state: going straight and lane changing; Longitudinal state: accelerating, maintaining a constant speed, and decelerating. Thus, a vehicle state equation can be established. The dynamic Bayesian network is divided into two stages: prediction and update. Probability distribution models can be established separately for the prediction stage and the update stage. Establish the vehicle state equation:

[0152]

[0153]

[0154] Among them, X t represents the vehicle state, and the state includes the position relative to the lane centerline and the lateral speed; Y t is the measurement output or the observed value; A and B are the system state transition matrices, and C is the output matrix; ε t , η t are Gaussian noises, and the Q and R matrices are the variances of the Gaussian noises; The dynamic Bayesian network is divided into two stages: prediction and update. Prediction stage: the probability distribution of the vehicle characteristic state

[0155]

[0156] Among them, Z represents the environmental characteristic state, that is, the combination of the characteristics of the surrounding obstacles, and the characteristics are represented by E t ; M represents the vehicle state space, where the lateral is two states: going straight and changing lanes, and the longitudinal is three states: accelerating, decelerating, and maintaining a constant speed. represents the joint probability that the vehicle state transfers from the M t-1 at the t - 1 moment to the M t at the t moment while the environmental state transfers from Z t-1 at the t - 1 moment to Z t at the t moment; So is obtained by the product of P(Z t |Z t-1 ), P(M t |Z t , M t-1 ). This distribution is the value calculated in the previous cycle;

[0157] Among them P(Z t |Z t-1 ), P(M t |Z t , M t-1 ), these four probability distributions are respectively: Indicates the joint probability that the vehicle state at time t - 1 is M t-1 and the environmental state is Z t-1 ; P(Z t |Z t-1 ) indicates the conditional probability distribution that the vehicle state at time t - 1 is Z t-1 transfers to the environmental state Z at time t t ; P(M t |Z t , M t-1 ) indicates the conditional probability that the vehicle state at time t - 1 is M t-1 and the environmental state is Z t when it transfers to the environmental state M t ; Vehicle state transition probability distribution:

[0158]

[0159]

[0160]

[0161] Among them, Indicates that when the vehicle state at time t - 1 is M t-1 , and the vehicle state at time t is M t , the position and speed of the vehicle follow a Gaussian distribution with a certain mean and variance; Indicates the mean of the Gaussian distribution, Indicates the variance of the Gaussian distribution.

[0162] Update phase:

[0163]

[0164] Among them, P(Y t |M t , M t-1 ) indicates the probability distribution of the observed value when the vehicle state at time t - 1 is M t-1 , and the vehicle state at time t is M t ; Vehicle characteristic state probability distribution

[0165]

[0166] ​

[0167] Among them, represents the joint probability that the vehicle state at time t after update is transferred from time t-1 to time t when the environmental state is transferred from time t-1 to time t, which is prepared for the next calculation; represents the probability of a certain feature observed when the environmental state at time t is, where represents the angle between the obstacle orientation and the current lane, the collision time between the obstacle and the obstacle in front, the average speed of the obstacle and the obstacle in front, and the collision time between the obstacle and the adjacent obstacles on the left and right.

[0168] Vehicle state transition calculation

[0169]

[0170] Among them, represents that the vehicle state at time t-1 is M t-1 , and when the vehicle state at time t is M t , the position and speed of the vehicle follow a Gaussian distribution with a certain mean and variance, which is mainly prepared for the prediction process, that is, the next calculation;

[0171] Prepare for the next prediction process. Substituting the above calculation results can obtain the prediction result at the next moment:

[0172]

[0173]

[0174]

[0175]

[0176] In some alternative embodiments, the calculation unit further includes a first calculation module and a second calculation module. Among them, the first calculation module is used to calculate the action probability of the controlled vehicle according to the above probability distribution model and calculate the action probability of the obstacle; the second calculation module is used to calculate the product of the action probability of the controlled vehicle and the action probability of the obstacle to obtain the above confidence level. This method calculates the confidence level through the action probability obtained by the above probability distribution model, so that the confidence level can be accurately calculated.

[0177] In the specific implementation process, the partially observable Markov decision process (POMDP) method is used to construct a POMDP model, generate a decision sequence of the controlled vehicle, set the state set as S, the action set as A, the reward function as R, the discount factor as γ, the state transition probability as P, and the state value function as V(s). According to the Bellman state value function equation, the expectation of the value function on the trajectory taken under the guidance of the policy can be obtained when a given policy π is given, where s’ represents the state at the next time step:

[0178]

[0179] Optimal state value function:

[0180]

[0181] State transition function:

[0182]

[0183] where, represents the ego vehicle's action, represents the action of the k-th obstacle, A k represents the action set of the k-th obstacle; semantic actions are used to map continuous actions, where is the semantic intention of the k-th obstacle,

[0184]

[0185]

[0186] replaces the state transition probability s' ∼ p(.|s,a) from the current state to the next state, obtaining

[0187]

[0188] If all environmental states are observable and the model is known, the optimal state value function can be obtained through the above formula. However, not all environmental states can be observed, so the posterior probability of the state, i.e., the confidence, is used to calculate the optimal state value function, and the above formula is modified to

[0189]

[0190] where, B is the confidence space of the state. Using the confidence to model POMDP by the above formula requires calculating the confidence b, and its calculation method is:

[0191]

[0192] where, t E is the time period of the decision sequence.

[0193] To accurately control the vehicle's driving, the above control unit includes a building module and a control module. Specifically, the building module is used to establish the kinematic model of the above vehicle to be controlled in the Cartesian coordinate system, obtain the constraint conditions, and establish an objective function based on the above kinematic model and the above constraint conditions. The control module is used to solve the above objective function to obtain the control variables and optimization variables corresponding to the above target driving strategy, and control the above vehicle to drive according to the above control variables and the above optimization variables, so that the above vehicle to be controlled drives according to the above target driving strategy. Among them, the above control variables include the steering wheel angle change rate and the acceleration, and the above optimization variables include the position, speed, vehicle body heading angle, and steering wheel angle. This method controls the vehicle to drive according to the optimal driving strategy by establishing a kinematic model, so that the vehicle's driving can be accurately controlled.

[0194] Specifically, after generating the optimal strategy, the planning model of the vehicle to be controlled is established. The entire planning process uses the Cartesian coordinate system, so the kinematic model of the vehicle is a non-linear model. At the same time, the bicycle model is used to model the vehicle model, and the vehicle kinematic equation, where X x represents the position x in the Cartesian coordinate system, X y represents the position y in the Cartesian coordinate system, θ represents the vehicle orientation angle, v represents the vehicle speed, δ represents the vehicle steering wheel angle, and L represents the wheelbase between the front and rear wheels of the vehicle:

[0195]

[0196]

[0197] The equality constraints formed by the state transition equation are as follows: X = [X x X y θ v δ] = f(X,U), and the acceleration constraint is The road boundary constraint uses the Lp distance to describe the distance constraint between the vehicle itself and the road boundary. The Lp distance expression is as follows:

[0198]

[0199] In this method, the numerical expression of the distance using the third order is as follows: Among them, γ i is the angle between the connection lines at both ends of the point and the line segment. For the obstacle constraint, an ellipse is used to represent the obstacle, and the ellipse equation is constructed:

[0200]

[0201] The objective function is divided into four parts, namely: the cost of the vehicle deviating from the center line of the lane The cost of vehicle acceleration Steering wheel angle change rate cost Vehicle orientation cost The total objective function is as follows:

[0202]

[0203] In some alternative embodiments, the control module includes an iterative sub-module and a stopping module. The iterative sub-module is configured to establish an augmented Lagrangian equation according to the above objective function and the above constraints, and iteratively solve the above augmented Lagrangian equation; the stopping module is configured to stop the iteration and obtain the above control variables and the above optimization variables when the above augmented Lagrangian equation satisfies the objective convergence condition. This method uses the augmented Lagrangian method to solve the above non-linear equality constraints, so as to accurately solve the problem.

[0204] Specifically, the equality constraints, inequality constraints, and objective function constituted by the kinematic model, i.e., the state transition equation, are analyzed, and finally the vehicle position, speed, body heading angle, steering wheel angle, steering wheel angle change rate, acceleration, etc. are obtained. Since there are non-linear equality constraints in the process of vehicle kinematic modeling, and both the boundary constraints and obstacle constraints are non-linear constraints, they are solved by the augmented Lagrangian differential dynamic programming method; the optimal decision is used as the rough trajectory to participate in the solution of the spatio-temporal trajectory generation. By providing a warm start through the rough trajectory, the convergence speed of the augmented Lagrangian differential dynamic programming (AL-DDP) can be accelerated. The AL-DDP method is used to solve the problem, and an augmented Lagrangian equation is constructed for the objective function and constraints in the model in step seven.

[0205]

[0206] Among them, in the formula, L A represents the augmented Lagrangian function of the optimization problem, x k represents the state, u k represents the control, λ N represents the Lagrange multiplier, I u,k represents the Lagrange penalty factor, c N represents the terminal constraint, c k represents the constraints at time 0-(N - 1), where the constraints include equality constraints and inequality constraints, l k represents the objective function from time 0 to N - 1. Construct the Cost-to-go iterative equation

[0207]

[0208]

[0209] Among them, V N (x N)The state value function representing the end point, V k (x k )represents the state value function from time 0 to N-1. Since it is an iterative process, it is necessary to control the perturbation amount of the reference value. Considering the approximation accuracy, a second-order Taylor expansion is performed on the constraint of time step N and the objective function.

[0210]

[0211] where

[0212]

[0213]

[0214] where, δV k represents the state value function generated by the perturbation; in this way, the action-state value function δQ can be obtained

[0215]

[0216] where,

[0217]

[0218]

[0219]

[0220]

[0221]

[0222] where, Q represents the perturbation state-action value function, Qxx is the second derivative of the state-action value function with respect to state x, Quu is the second derivative of the state-action value function with respect to action u, Qux is the second partial derivative of the state-action value function with respect to state x and action u, Qx and Qu respectively represent the first derivatives of the state-action value function with respect to state x and action u; minimize δQ k , according to the necessary condition for the extreme value that the first derivative is 0, that is, taking the first derivative of δQ k can obtain the control that makes δQ k an extreme value

[0223]

[0224] The update methods of the Lagrange multiplier and the penalty factor are as follows:

[0225]

[0226]

[0227] Among them, λ ki represents the Lagrange multiplier, and μ ki represents the Lagrange penalty factor. The iteration terminates when the following convergence conditions are met: reaching the maximum number of iterations; the cost function is less than a certain value; the feedback gain K is less than zero; the solution of the entire augmented Lagrangian DDP problem is divided into two parts: the differential dynamic programming process (DDP) and the update of the Lagrange multiplier and penalty factor. Among them, DDP contains two steps, namely the forward vehicle passing and the backward process; the total process of the augmented Lagrangian DDP, setting the initial values of λ, μ, and φ; the DDP process; updating the Lagrange multiplier; calculating the cost L A , if it is less than the set value, the loop ends, otherwise it returns to step 2 until the cost converges to less than the set value. Steps of the DDP process: Initialize the state variable X 0 and the control variable U 0 ; Execute the backward process of DDP to generate the feedback gain K t at each time step, execute the forward process of DDP to calculate the cost L A , if it is greater than the steps of the DDP backward process: Set P N , p N values at time N; Calculate δQ; Judge Q uu >0, if it is greater than 0, update the feedback gain K and d to end the loop, output K and d, if it is less than 0, increase the ρ value, and jump to step 2 to continue the calculation; Steps of the DDP forward process: Set the initial value x 0 ; According to the feedback gain K and d output by the DDP backward process, calculate the control U at each time step through the control law, and the vehicle state X is calculated from the state transition equation at each time step; Output the control variable and state variable sequences from time 0 to N steps.

[0228] To establish the state value function, the above control unit further includes an establishment module and a substitution module. Among them, the establishment module is used to establish a partially observable Markov decision process model; the substitution module is used to substitute the above confidence into the above partially observable Markov decision process model to obtain the above state value function. This method establishes the state value function through the partially observable Markov decision process model, so that the state value functions of each decision can be accurately obtained.

[0229] Specifically, as mentioned above, the partially observable Markov decision process (POMDP) method is used to generate the ego-vehicle decision sequence, and then the state value function is established. The specific process is as above.

[0230] The control device of the above-mentioned autonomous vehicle includes a processor and a memory. The above-mentioned acquisition unit, calculation unit, control unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions. The above-mentioned modules are all located in the same processor; alternatively, the above-mentioned each module is separately located in different processors in any combination form.

[0231] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and the accuracy of the road planning of the autonomous vehicle is improved by adjusting the kernel parameters.

[0232] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory includes at least one storage chip.

[0233] The embodiment of the present invention provides a computer-readable storage medium. The above-mentioned computer-readable storage medium includes a stored program. Wherein, when the above-mentioned program runs, it controls the device where the above-mentioned computer-readable storage medium is located to execute the control method of the above-mentioned autonomous vehicle.

[0234] The embodiment of the present invention provides a device. The device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the control method of the above-mentioned autonomous vehicle.

[0235] The device herein can be a server, a PC, a PAD, a mobile phone, etc.

[0236] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0237] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0238] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0239] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0240] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0241] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0242] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0243] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0244] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0245] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0246] 1) In the control method of the autonomous driving vehicle of the present application, multiple possible driving strategies can be determined based on the current state information, road scene information and obstacle scene information of the vehicle, and then the driving action, probability distribution model and confidence are determined, and then the state value function is established, and the state value function value is calculated, and the vehicle is controlled to drive according to the driving strategy corresponding to the highest state value function value. In this way, the optimal driving strategy can be selected from multiple driving strategies to control the driving of the vehicle under complex road conditions. Compared with the inaccurate road planning method of the vehicle in the prior art under complex road scenes, the present application can select the optimal driving strategy from multiple driving strategies determined by the current state information, road scene information and obstacle scene information of the vehicle, and therefore, it can alleviate the problem of inaccurate vehicle road planning methods in the prior art and improve the accuracy of road planning methods.

[0247] 2) In the control device of the autonomous vehicle of the present application, multiple possible driving strategies can be determined based on the current state information of the vehicle, road scene information, and obstacle scene information. Then, the driving action, probability distribution model, and confidence level are determined. Subsequently, a state value function is established, and the state value function value is calculated. The vehicle is controlled to drive according to the driving strategy corresponding to the highest state value function value. In this way, in complex road conditions, the optimal driving strategy can be selected from multiple driving strategies to control the driving of the vehicle. Compared with the inaccurate road planning methods of vehicles in the prior art in complex road scenarios, the present application can select the optimal driving strategy from multiple driving strategies determined from the current state information of the vehicle, road scene information, and obstacle scene information. Therefore, the problem of inaccurate vehicle road planning methods in the prior art can be alleviated, and the accuracy of the road planning method can be improved.

[0248] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A control method for an autonomous vehicle, characterized in that, it includes: Obtain the current state information, road scene information, and obstacle scene information of the vehicle to be controlled. Among them, the current state information is the position and speed of the vehicle to be controlled at the current moment, the road scene information represents the road condition information where the vehicle to be controlled is located, and the obstacle scene information represents the information of vehicles other than the vehicle to be controlled on the road and at least includes position and speed; Determine multiple driving strategies based on the current state information, the road scene information, and the obstacle scene information, determine driving actions according to the multiple driving strategies, determine the probability distribution model of the driving actions, and calculate the confidence according to the probability distribution model. Among them, each driving strategy includes the driving actions at multiple moments, and the probability distribution model represents the probability that the vehicle to be controlled drives according to the driving strategy; Establish a state value function according to the confidence, use the driving strategy corresponding to the highest state value function value as the target driving strategy, and control the vehicle to be controlled to drive according to the target driving strategy. Among them, the state value function represents the pros and cons of each driving strategy, and the higher the state value function value, the better the corresponding driving strategy.

2. The control method according to claim 1, characterized in that, Determining the driving action according to the multiple driving strategies includes: Obtain the lateral semantic action and longitudinal semantic action of the vehicle to be controlled, and map the lateral semantic action and the longitudinal semantic action to a multi-dimensional state space. Among them, the lateral semantic action is a word representing the lateral action made by the vehicle to be controlled and includes lane keeping, left lane change, and right lane change, and the longitudinal semantic action is a word representing the longitudinal action made by the vehicle to be controlled and includes acceleration, deceleration, and constant speed; Use each driving strategy to search in the multi-dimensional state space to determine the driving action corresponding to each moment in the driving strategy.

3. The control method according to claim 1, characterized in that, Determining the probability distribution model of the driving action includes: Obtain the state variables of the vehicle to be controlled and the obstacle decision variables. Among them, the decision variable of the vehicle to be controlled is a variable representing the action of the vehicle to be controlled, and the obstacle decision variable is a variable affecting the decision-making of the obstacle; Use the state variables of the vehicle to be controlled and the obstacle decision variables as the features of a dynamic Bayesian network, establish a dynamic Bayesian network prediction model, and update the dynamic Bayesian network prediction model to obtain the probability distribution model.

4. The control method according to claim 3, characterized in that, Calculating the confidence according to the probability distribution model includes: Calculate the action probability of the vehicle to be controlled according to the probability distribution model, and calculate the action probability of the obstacle; Calculate the product of the action probability of the vehicle to be controlled and the action probability of the obstacle to obtain the confidence.

5. The control method according to claim 1, characterized in that, Controlling the vehicle to be controlled to drive according to the target driving strategy includes: A kinematic model for controlling a vehicle is established in a Cartesian coordinate system and constraint conditions are obtained, and an objective function is established according to the kinematic model and the constraint conditions; The objective function is solved to obtain control variables and optimization variables corresponding to the target driving strategy, and the vehicle is controlled to drive according to the control variables and the optimization variables, so that the controlled vehicle drives according to the target driving strategy, wherein the control variables include the steering wheel angle change rate and the acceleration, and the optimization variables include the position, the speed, the vehicle body heading angle and the steering wheel angle.

6. The control method according to claim 5, characterized in that solving the objective function includes: establishing an augmented Lagrangian equation according to the objective function and the constraint conditions, and performing iterative solution on the augmented Lagrangian equation; when the augmented Lagrangian equation satisfies the target convergence condition, stopping the iteration to obtain the control variables and the optimization variables.

7. The control method according to claim 1, characterized in that establishing a state value function according to the confidence includes: establishing a partially observable Markov decision process model; substituting the confidence into the partially observable Markov decision process model to obtain the state value function.

8. A control device for an autonomous vehicle, characterized in that comprising: an acquisition unit, configured to acquire the current state information, road scene information, and obstacle scene information of the controlled vehicle, wherein the current state information is the position and speed of the controlled vehicle at the current moment, the road scene information represents the road condition information where the controlled vehicle is located, and the obstacle scene information represents the information of vehicles other than the controlled vehicle on the road and at least includes the position and speed; a calculation unit, configured to determine multiple driving strategies according to the current state information, the road scene information, and the obstacle scene information of the controlled vehicle, determine driving actions according to the multiple driving strategies, determine a probability distribution model of the driving actions, and calculate a confidence according to the probability distribution model, wherein each driving strategy includes the driving actions at multiple moments, and the probability distribution model represents the probability that the controlled vehicle drives according to the driving strategy; a control unit, configured to establish a state value function according to the confidence, use the driving strategy corresponding to the highest state value function value as the target driving strategy, and control the controlled vehicle to drive according to the target driving strategy, wherein the state value function represents the pros and cons of each driving strategy, and the higher the state value function value, the better the corresponding driving strategy.

9. A computer-readable storage medium, characterized in that the computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the control method according to any one of claims 1 to 7.

10. A vehicle, characterized in that comprising: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a control method for performing any one of claims 1 to 7.