Vehicle control method and device, electronic equipment and storage medium

CN121650699BActive Publication Date: 2026-09-25SHENZHEN CONSYS SCI&TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511465050.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-09-25
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

[0004]本申请实施例的主要目的在于提出一种车辆控制方法、装置、电子设备及存储介质,能够解决现有技术中智能体车辆在车辆控制方面的安全性较低的问题,提升智能体在行人行为不确定和环境动态变化场景下的决策能力与控制性能

Benefits of technology

[0015]本申请提出的车辆控制方法、装置、电子设备及存储介质,其通过获取车辆当前状态与环境信息,随后基于贝叶斯法则处理环境中的未知信息,推断出描述环境真实状态概率分布的当前信念。以当前信念为根节点构建信念树,树中分支代表可执行动作,子节点对应动作后不同观测结果更新的新信念。最后,通过蒙特卡洛树搜索算法从信念树中决策出最优目标动作,并控制车辆执行目标动作,实现复杂不确定环境下的鲁棒性实时决策。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121650699B_ABST
    Figure CN121650699B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a vehicle control method and device, electronic equipment and storage medium, belonging to the technical field of intelligent driving. The method comprises: obtaining state information of a vehicle at a current time point and environment information of an environment in which the vehicle is located; obtaining a belief corresponding to the current time point based on Bayes rule according to the state information and the environment information, the belief being a probability distribution obtained by inferring the true state of the environment information according to unknown information in the environment information; constructing a belief tree with the belief corresponding to the current time point as an initial belief, the root node of the belief tree being the initial belief; obtaining a target action corresponding to the current time point from the belief tree based on a Monte Carlo tree search algorithm; and controlling the vehicle to perform the target action. The embodiments of the present application can improve the decision-making ability and control performance of intelligent driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent driving technology, and in particular to a vehicle control method, device, electronic device and storage medium. Background Technology

[0002] With the continuous development of intelligent driving technology, research on control strategies for intelligent vehicles in complex traffic environments has become a hot topic. Traditional vehicle control methods typically rely on deterministic environmental assumptions, which assume that all environmental information (such as the intentions of other traffic participants, road conditions, etc.) is completely known and accurately perceived. However, there are many uncertainties in the actual driving environment, such as sensor noise, occlusion, and the uncertainty of dynamic obstacle behavior. These factors often make environmental information partially observable or incomplete, which can easily lead to decision-making errors due to missing information.

[0003] Traditional vehicle control systems typically rely on preset rules or deterministic model-based algorithms for decision-making, such as rule-based state machines, dynamic window methods, or optimization-based trajectory planning methods. However, these methods have significant limitations in handling environmental uncertainties, resulting in lower vehicle safety in partially observable environments. Summary of the Invention

[0004] The main objective of this application is to propose a vehicle control method, device, electronic device, and storage medium that can solve the problem of low safety in vehicle control of intelligent agent vehicles in the prior art, and improve the decision-making ability and control performance of intelligent agents in scenarios with uncertain pedestrian behavior and dynamic environmental changes.

[0005] To achieve the above objectives, a first aspect of this application provides a vehicle control method, the method comprising: Obtain the vehicle's status information at the current point in time and the environmental information of the vehicle's surroundings; Based on the state information and the environment information, the belief corresponding to the current time point is obtained based on Bayes' theorem. The belief is a probability distribution obtained by inferring the true state of the environment information based on unknown information in the environment information. A belief tree is constructed using the belief corresponding to the current time point as the initial belief. The root node in the belief tree is the initial belief. Each branch extending from the parent node in the belief tree corresponds to an executable action. The action is connected to one or more child nodes. Each child node corresponds to a new belief corresponding to different observation results obtained by executing the action. Based on the Monte Carlo tree search algorithm, the target action corresponding to the current time point is obtained from the belief tree; Control the vehicle to perform the target action.

[0006] In some embodiments, obtaining the belief corresponding to the current time point based on the state information and the environmental information using Bayesian rules includes: Based on the vehicle's current state information and environmental information, a partially observable Markov decision process model is constructed. Based on the partially observable Markov decision process model, the belief corresponding to the current time point is obtained through Bayes' rule.

[0007] In some embodiments, the step of constructing a partially observable Markov decision process model based on the vehicle's current state information and environmental information includes: Based on the state information and environmental information, a state space, action space, observation space, state transition probability, observation probability, reward function, and discount factor are defined. The state space includes the state of the vehicle and pedestrians in the surrounding environment. The action space is the set of actions that the vehicle can perform. The observation space is the set of all observations obtained by the vehicle after performing an action at the current time point. The state transition probability is the conditional probability that the vehicle obtains new state information after performing an action at the current time point. The observation probability is the conditional probability that the vehicle obtains corresponding environmental information after performing an action. The reward function is the reward or penalty obtained by the vehicle after performing an action at the current time point. The discount factor is used to balance the weights of immediate rewards and future rewards. The partially observable Markov decision process model is obtained based on the state space, the action space, the observation space, the state transition probability, the observation probability, the reward function, and the discount factor.

[0008] In some embodiments, obtaining the belief corresponding to the current time point based on the partially observable Markov decision process model using Bayes' theorem includes: Obtain the belief at the previous time point, and the state transition probability and the observation probability at the current time point in the partially observable Markov decision process model; Based on Bayes' theorem, using the belief at the previous time point as a basis, the belief state corresponding to the vehicle's action at the current time point is calculated according to the state transition probability and the observation probability, thus obtaining the belief corresponding to the current time point.

[0009] In some embodiments, constructing a belief tree using the belief corresponding to the current time point as the initial belief includes: Using the belief corresponding to the current time point as the root node, starting from the root node, traverse all executable actions corresponding to the root node; For each executable action of the root node, the following processing is performed: After performing the action under the belief corresponding to the root node and obtaining the corresponding observation results, the belief at the next time point is calculated based on Bayes' theorem. The belief at the next time point is taken as a child node and connected to the root node to obtain the edge in the belief tree; The child node is taken as the root node, and the process jumps to the step of traversing all executable actions corresponding to the root node, starting from the root node, until the depth of the belief tree reaches the preset search boundary.

[0010] In some embodiments, obtaining the target action corresponding to the current time point from the belief tree based on the Monte Carlo tree search algorithm includes: Starting from the root node, at each node of the belief tree, the upper limit confidence interval algorithm is used to select the child node with the largest value of the value function as the node to be expanded, until the maximum depth of the belief tree is reached, thus obtaining the target search path; The target search path is simulated, and the long-term expected reward value of the target search path is calculated; The long-term expected reward value is backpropagated to all child nodes traversed by the target search path, and the value function corresponding to each child node traversed by the target search path is updated to obtain the updated value function of each child node. The target search path is replanned based on the updated value function of each child node until the preset iteration stopping condition is met; The target action is obtained based on the target search path obtained in the last iteration.

[0011] In some embodiments, the target action includes a first acceleration and a first wheel steering angle; Controlling the vehicle to perform the target action includes: Based on the environmental information and the belief corresponding to the current time point, a path planning algorithm is used to search for the path with the minimum comprehensive cost. The comprehensive cost includes at least: static obstacle collision cost, pedestrian collision probability cost, and path smoothness cost. Calculate the initial acceleration and initial wheel steering angle of the path with the minimum overall cost to obtain the second acceleration and the second wheel steering angle; The first acceleration and the second acceleration are vector-fused to obtain the target acceleration; The first wheel steering angle and the second wheel steering angle are vector-fused to obtain the target wheel steering angle; If the target wheel steering angle exceeds the preset maximum reasonable range, then the target wheel steering angle will be adjusted to the boundary value of the maximum reasonable range; Control the vehicle to perform a new target action, the new target action including the target acceleration and the target wheel steering angle.

[0012] To achieve the above objectives, a second aspect of this application provides a vehicle control device, the device comprising: The acquisition module is used to acquire the vehicle's status information at the current time and the environmental information of the environment in which the vehicle is located; The calculation module is used to obtain the belief corresponding to the current time point based on the state information and the environment information and the Bayesian rule. The belief is a probability distribution obtained by inferring the true state of the environment information based on unknown information in the environment information. The construction module is used to construct a belief tree with the belief corresponding to the current time point as the initial belief. The root node in the belief tree is the initial belief. Each branch extending from the parent node in the belief tree corresponds to an executable action. The action is connected to one or more child nodes. Each child node corresponds to a new belief corresponding to different observation results obtained by executing the action. The search module is used to obtain the target action corresponding to the current time point from the belief tree based on the Monte Carlo tree search algorithm. The control module is used to control the vehicle to perform the target action.

[0013] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0014] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0015] The vehicle control method, device, electronic equipment, and storage medium proposed in this application acquire the vehicle's current state and environmental information, then process unknown information in the environment based on Bayesian rules to infer the current belief describing the probability distribution of the true state of the environment. A belief tree is constructed with the current belief as the root node, where branches represent executable actions, and child nodes correspond to new beliefs updated after different observations following the action. Finally, the optimal target action is determined from the belief tree using a Monte Carlo tree search algorithm, and the vehicle is controlled to execute the target action, achieving robust real-time decision-making in complex and uncertain environments. Attached Figure Description

[0016] Figure 1 This is a schematic flowchart of the vehicle control method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the vehicle observation information provided in the embodiments of this application; Figure 3 This is a logical schematic diagram of the vehicle control method provided in the embodiments of this application; Figure 4 This is a schematic diagram of the vehicle control device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0018] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0020] With the continuous evolution of intelligent driving technology, control strategies for intelligent vehicles in complex traffic environments involving dynamic elements such as pedestrians and other vehicles have become a core research direction in this field that urgently needs breakthroughs and has attracted much attention. Traditional vehicle control methods typically rely on deterministic environmental assumptions, that is, assuming that all environmental information (such as the intentions of other traffic participants, road conditions, etc.) is completely known and accurately perceived. However, there are many uncertainties in the actual driving environment, such as sensor noise, occlusion, and the uncertainty of dynamic obstacle behavior. These factors often make environmental information partially observable or incomplete, which can easily lead to decision-making errors due to missing information.

[0021] In the field of intelligent agent decision-making, traditional vehicle control systems typically rely on preset rules or algorithms based on deterministic models for decision-making, such as rule-based state machines, dynamic window methods, or optimization-based trajectory planning methods. However, these methods have significant limitations when dealing with environmental uncertainties. In complex traffic environments, pedestrian motion is highly dynamic and uncertain. Traditional methods struggle to accurately predict pedestrian trajectories and intentions, and cannot quickly respond to or adapt to sudden environmental events such as changes in pedestrian behavior. This leads to delayed or biased decision-making by the intelligent agent vehicle, severely impacting its decision-making performance and consequently affecting vehicle safety in partially observable environments.

[0022] Based on this, embodiments of this application provide a vehicle control method, device, electronic device, and storage medium, aiming to provide an intelligent decision-making method utilizing belief iteration, solving the problem of low efficiency in cooperative vehicle control in the prior art. Embodiments of this application utilize beliefs (probability distributions) to describe the observation information obtained by sensors, and use the POMDP paradigm and Bayesian theory to provide a set of belief iteration calculation methods and a belief tree construction method; to address the large search space of the belief tree, Monte Carlo tree search technology is introduced to achieve efficient search; and to address the agent control problems that may arise from output actions, a path correction method is provided to make the agent control smoother and more reasonable.

[0023] The vehicle control method, device, electronic equipment, and storage medium provided in this application are specifically described through the following embodiments. First, the vehicle control method in the embodiments of this application is described.

[0024] The vehicle control method provided in this application relates to the field of intelligent driving technology. The vehicle control method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the vehicle control method, but is not limited to the above forms.

[0025] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0026] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0027] Figure 1 This is an optional flowchart of a vehicle control method provided in an embodiment of this application. This embodiment applies to radar. Figure 1 The method may include, but is not limited to, steps S100 to S500.

[0028] Step S100: Obtain the vehicle's status information at the current time and the environmental information of the environment in which the vehicle is located.

[0029] In this embodiment, the system collects vehicle status information and environmental information in real time using onboard sensors (such as cameras, radar, lidar, etc.). Vehicle status information refers to physical quantities describing the instantaneous kinematics and dynamics of the vehicle itself. Environmental information refers to the state of static and dynamic objects in the vehicle's external environment that are relevant to decision-making.

[0030] Specifically, the environment is first modeled as a Partial Observable Markov Decision Process (POMDP), i.e., a seven-tuple. At each time step, the vehicle agent performs an action 'a', determined by the state. Transition to state Then obtain observation information. State transition probability Describes the state transition of a vehicle taking action a in state s; observation probability. Describes how the vehicle reaches a state through action a. Subsequent observations: After performing action a in state s, the vehicle receives an immediate reward R(s,a).

[0031] S is the state space, which is the set of all possible and related situations in the vehicle's environment, including the vehicle's state space and the pedestrian's state space; A is the action space, which is the set of all possible actions that the vehicle can perform in all states; Z is the observation space, which is the set of all possible incomplete information that the vehicle can obtain through sensors after performing an action. Observation is the direct input of the vehicle's perception of the environment and usually cannot completely determine the true state; T is the state transition probability, which is a probability function describing the dynamic changes of the environment and represents the probability that the environment will transition to a new state s' after performing action a in state s; O is a probability function simulating sensor noise and uncertainty, representing the probability that the vehicle will observe z after performing action a and reaching a new state s'; R is the reward function, which provides a criterion for the vehicle's decision-making and represents the immediate reward (or penalty) obtained by the vehicle after performing action a in state s. The ultimate goal is to maximize the long-term cumulative reward. γ is the discount factor, a constant between 0 and 1, used to weigh the importance of immediate rewards against future rewards. The closer γ is to 1, the more forward-looking the vehicle is; the closer it is to 0, the more it focuses on short-term gains.

[0032] The vehicle's status information includes instantaneous speed, acceleration, steering angle, position coordinates, heading angle, etc., which are used to comprehensively describe the vehicle's physical state at the current moment.

[0033] The state space of a vehicle can be represented as: Where (x,y) are the real-time position coordinates of the vehicle in the two-dimensional coordinate system, θ is the attitude angle (i.e., the angle between the vehicle's direction of travel and the reference coordinate system), and v is the instantaneous velocity of the vehicle.

[0034] The vehicle's motion space can be represented as: .in, The rotation angle of the wheel, where 'a' is the acceleration of the vehicle.

[0035] Environmental information focuses on key factors surrounding the vehicle that influence driving decisions, including both dynamic targets and static environmental information. Dynamic targets are centered on pedestrians, whose state space can be represented as: Where (x, y) are the real-time position coordinates of the pedestrian in the two-dimensional coordinate system, g is the pedestrian's target position, and v is the pedestrian's instantaneous velocity. In this embodiment, for a specific pedestrian, the coordinates of the target position are fixed, and the pedestrian will move towards the target position g at each time step. The distance. Static environmental information includes static elements around the vehicle, such as road structure information like road boundaries and lane lines; information on stationary obstacles such as roadside guardrails, stopped vehicles, and traffic signs.

[0036] The reward value R(s,a) is calculated as follows: .in, It controls risk avoidance; if other pedestrians or vehicles are less than the safe threshold, there is a penalty value. Controlling the vehicle's movement forward, a bonus is awarded if the distance between the vehicle and the target decreases; Controlling the speed and encouraging the vehicle to travel as close to its maximum speed as possible. To travel at the speed of; Comfort is controlled, and there is a penalty value if the vehicle has large acceleration or deceleration actions.

[0037] It should be noted that in this embodiment, environmental information includes both directly observable information (such as the current location of a pedestrian) and information that is not directly observable or unknown (such as the pedestrian's intention or target location).

[0038] Step S200: Based on the state information and the environment information, obtain the belief corresponding to the current time point based on Bayes' theorem. The belief is a probability distribution obtained by inferring the true state of the environment information based on unknown information in the environment information.

[0039] In this embodiment, within the intelligent driving system, the vehicle needs to make decisions based on its own state information (such as position, speed, and direction) and environmental information (such as surrounding vehicles, pedestrians, and traffic signals). However, environmental information often contains uncertainties, such as the driving intentions of other vehicles and the next actions of pedestrians. To address this uncertainty, the system needs to construct beliefs using Bayes' theorem, which is an inference of the probability distribution of the true state of the environment. A belief is a probability distribution used to represent the vehicle's uncertainty estimate of unknown information (i.e., state variables that cannot be directly observed) in the current environmental information. The vehicle's observation information is a vector composed of the vehicle's state and the positions of all pedestrians. Since the target coordinates of pedestrians are unknown, inference is required. The belief is this inference.

[0040] For example, such as Figure 2As shown, if a pedestrian continues to walk straight on the crosswalk (solid line direction), the optimal action of the vehicle might be to accelerate past the pedestrian; if the pedestrian intends to cross the street, the optimal action of the vehicle might be to slow down and yield. The belief vector for these two scenarios is (p, 1-p). At the beginning of driving, due to a lack of information, the belief vector is initialized to (0.5, 0.5). As driving progresses, if the vehicle's sensor data shows that the pedestrian has a turning trajectory (dashed line direction), then the pedestrian may intend to cross the street, and the belief vector value at this time might be (0.3, 0.7), eventually converging to (0, 1).

[0041] Each time step t corresponds to a belief, which contains the vehicle agent's predictions about all possibilities. The essence of a belief is a probability distribution, denoted as . At time step t, record the action performed by the vehicle. Then obtain observation information This embodiment utilizes Bayes' theorem to construct beliefs. Update formula: .in, For the beliefs of the previous time step, It is a standardized constant. It can be viewed as a function, derived from a series of previous observations. , It is the initial belief.

[0042] Step S300: Construct a belief tree with the belief corresponding to the current time point as the initial belief. The root node in the belief tree is the initial belief. Each branch extending from the parent node in the belief tree corresponds to an executable action. The action is connected to one or more child nodes. Each child node corresponds to a new belief corresponding to different observation results obtained by executing the action.

[0043] In this embodiment, the belief tree is a tree-like data structure used to represent all possible future belief evolution paths starting from the current belief. Its nodes represent the agent's (vehicle's) cognitive state (belief) regarding the environment, and branches represent the agent's action choices. Hierarchical expansion enables a complete enumeration of future decision paths. Each node represents a belief state b, i.e., the probability distribution of the environmental state. The root node is the initial belief b0 at the current time point. Each edge emanating from the parent node represents an executable action a∈A, where A is the vehicle's action space. Each action edge a connects to one or more child nodes, and each child node represents the updated belief state obtained after executing action a and receiving a certain observation z∈Z.

[0044] Specifically, in a POMDP environment, the agent's action policy is essentially a probability distribution mapping derived from its current belief, which can be written as: For belief b, the action strategy is... The value function is: .in, It is a discount factor.

[0045] In this embodiment, a belief tree is defined, where each node represents a belief, each edge represents an action-belief pair, and the root node is the initial belief. For parent node b, if there is an edge (a, z) pointing to one of its child nodes. Then the relation is established. .

[0046] Step S400: Based on the Monte Carlo tree search algorithm, obtain the target action corresponding to the current time point from the belief tree.

[0047] In this embodiment, Monte Carlo tree search is an iterative search algorithm that performs a directional search on the belief tree through four stages: selection, expansion, simulation, and backtracking, and finally selects the optimal one from all possible actions of the root node.

[0048] Specifically, after defining the belief tree, in the current state, the belief tree is searched with the current belief as the root node to solve for the vehicle's optimal action. At each node, this embodiment, based on the Johnson-Bellman rule, sets the value function of that node as the case under a greedy strategy: The formula for calculating the value function of edge (a,z) is: .

[0049] Belief trees have numerous branches, resulting in a large search space. This patent introduces the concept of Monte Carlo tree search into the belief tree search process. The search formula for the tree strategy is as follows: The constant c takes the value of N and n represent the total number of tree search iterations and the number of times the current leaf node has been searched, respectively. In each round of the search, the node with the largest UCB value is selected for expansion, Monte Carlo simulation is run to the maximum depth, and then the value of the value function is updated.

[0050] After the search is complete, use a greedy strategy of Q(b,a) to select actions. .

[0051] Step S500: Control the vehicle to perform the target action.

[0052] In this embodiment, decision-making algorithms such as Monte Carlo Tree Search (MCTS) search based on discrete models. The output action is theoretically optimal, but may exceed the execution capability (e.g., the steering angle may exceed the physical limits of the vehicle's steering mechanism) or fail to consider global path smoothness (e.g., vehicle instability caused by sharp turns). Therefore, this embodiment corrects the target action through a path correction algorithm and controls the vehicle's execution.

[0053] Specifically, the vehicle's travel path is first abstracted into computable structured data, and a path is defined. , consisting of a series of coordinate points To represent. These coordinate points are evenly distributed in In the above, the distance between each pair of adjacent coordinate points is l (fixed step size).

[0054] The cost-efficiency function C(ρ) quantifies the quality of each path, with lower costs indicating better paths. This function comprehensively considers collision risk and driving smoothness, and assigns higher weight to nearest points through a discount factor λ, prioritizing the safety of the shortest paths.

[0055] path The cost-efficiency function is: The first item This indicates the penalty value for a vehicle hitting a stationary obstacle; the second item. This indicates the penalty value for a vehicle colliding with a moving pedestrian; the third item. This represents the penalty value for an insufficiently smooth vehicle path. (Constant) This represents the discount factor; closer coordinates should have a higher weight than farther coordinates. It also represents the penalty value for hitting a pedestrian. The calculation process takes belief into account. For a specific coordinate point (x, y), the calculation method is as follows: Where C is a constant representing the penalty value.

[0056] This embodiment uses the Hybrid A* algorithm to search for the path with the lowest cost-efficiency value. The nodes in the search process are... , representing the vehicle's position coordinates (x, y) and attitude angles, respectively. The search step size is defined as follows: Each step of the search is performed, and the node... It is possible to reach the next node. After the search is complete, consider starting from the point. arrive For the first road segment, calculate the expected acceleration. and wheel steering angle Let the maximum rotation angle of the wheel be . The acceleration and wheel rotation angle output by the Monte Carlo tree are respectively and ,remember and The corresponding unit vectors are respectively and The following correction algorithm is constructed in this embodiment: ; If the above... Falling in the range If the inside is inside, then the final rotation angle is... ;if Exceeding the range Then take and Closer to China The direction, as the final After the correction, the vehicle... As the final action selected, it is passed to the action control module for intelligent driving control.

[0057] The logic diagram of this embodiment is as follows: Figure 3As shown, through Bayesian belief iteration, the system can transform raw, noisy sensor observation information, fused with historical data, into probabilistic and quantitative inferences (beliefs) about hidden states in the environment (such as pedestrian intentions and vehicle targets). This makes decision-making no longer rely on a single, potentially erroneous "guess," but rather on "expectations" based on all possibilities, fundamentally solving the technical problem of inaccurate prediction of dynamic obstacle (such as pedestrian) behavior and providing a solid and reliable information foundation for subsequent decisions. By constructing a belief tree, the continuous, high-dimensional belief space is transformed into a structured discrete search problem, providing a framework for algorithm optimization. Furthermore, the Monte Carlo tree search algorithm is introduced, which can efficiently approximate the optimal solution without traversing all nodes of the belief tree. This mechanism greatly reduces computational complexity, making it possible to evaluate complex future scenarios and search for optimal actions within limited hardware resources and time constraints, overcoming the technical bottlenecks of computational explosion and inability to be applied online in traditional POMDP exact algorithms. The path correction algorithm, acting as a "safety filter" and "smoothener" between decision-making and execution, performs secondary optimization on the theoretically optimal action output by the Monte Carlo tree search algorithm. By considering vehicle dynamics constraints, global path smoothness, and fusing belief-based collision probability costs, this method ensures that the final executed action commands conform to the vehicle's physical limits while avoiding abrupt acceleration, deceleration, and steering, thus significantly improving vehicle handling stability and passenger comfort. This method forms a perfect technological closed loop from "perception-inference-decision-execution-re-perception." The system can continuously learn from the environment (updating beliefs) and adjust future decisions based on the learning results (searching for new belief trees), ultimately achieving a balance between safety, reliability, comfort, and real-time performance in vehicle control.

[0058] In some embodiments, step S200 may include, but is not limited to, steps S210 to S220: Step S210: Based on the vehicle's current state information and environmental information, construct a partially observable Markov decision process model; Step S220: Based on the partially observable Markov decision process model, obtain the belief corresponding to the current time point using Bayes' rule.

[0059] In some embodiments, step S210 may include, but is not limited to, steps S211 to S212: Step S211: Define a state space, action space, observation space, state transition probability, observation probability, reward function, and discount factor based on the state information and environment information. The state space includes the state of the vehicle and pedestrians in the surrounding environment. The action space is the set of actions that the vehicle can perform. The observation space is the set of all observation results obtained by the vehicle after performing an action at the current time point. The state transition probability is the conditional probability that the vehicle obtains new state information after performing an action at the current time point. The observation probability is the conditional probability that the vehicle obtains corresponding environmental information after performing an action. The reward function is the reward or penalty obtained by the vehicle after performing an action at the current time point. The discount factor is used to balance the weights of immediate rewards and future rewards. Step S212: Based on the state space, the action space, the observation space, the state transition probability, the observation probability, the reward function, and the discount factor, the partially observable Markov decision process model is obtained.

[0060] In this embodiment, a partially observable Markov Decision Process (POMDP) ​​model adapted to the intelligent driving scenario is constructed by combining the vehicle's current state information and environmental information. The POMDP model uses a seven-tuple... Using this as the core framework, this embodiment combines the physical characteristics of intelligent driving scenarios to customize the definition of each element of the seven-tuple, thereby achieving a precise mapping between the model and the actual environment.

[0061] Specifically, the state space S is a set containing the vehicle state and the pedestrian state in the surrounding environment, covering all key state parameters that affect vehicle decisions. Among them, the vehicle state includes the current vehicle speed, acceleration, steering angle, position coordinates, and attitude angle, etc.; the pedestrian state includes the pedestrian's position, moving speed, direction of movement, and crossing intention, etc., including both explicit states that can be directly observed by sensors and implicit states that are occluded or undetected.

[0062] Action space A is the set of all actions that a vehicle can perform under current physical constraints and traffic rules, including but not limited to acceleration, deceleration, constant speed driving, left turn, right turn, maintaining steering, and emergency braking. The settings of actions must match the vehicle's dynamic performance (such as maximum acceleration / deceleration) and road traffic rules (such as prohibition of going straight at red lights). For example, in road sections where left turns are prohibited, left turn actions are excluded from the action space; and the deceleration of emergency braking is limited to a reasonable range based on the performance of the vehicle's braking system.

[0063] The observation space O is the set of all observations acquired by the perception module after the vehicle performs an action at the current time point, encompassing observations of both environmental and vehicle status information. Environmental information observations include obstacle distances detected by lidar, pedestrian positions and attitudes identified by cameras, and target speeds measured by millimeter-wave radar. Vehicle status observations include the actual vehicle speed fed back by the speed sensor and the steering angle collected by the steering angle sensor. Observations are incomplete and noisy reflections of the true state. For example, observations of a pedestrian might include {pedestrian located at coordinates (x, y), pedestrian speed v, pedestrian orientation θ}, but due to sensor errors, these observations may deviate from the true state, and the pedestrian's intentions may not be observable.

[0064] The state transition probability T(s'|s,a) is a conditional probability representing the conditional probability that the system state will transition to a new state s' after the vehicle performs action a in the current state s. It describes the impact of the action on state evolution. This probability is constructed based on the vehicle dynamics model, pedestrian movement patterns, and historical traffic data. For example, if the current state s is "vehicle speed 30km / h, no pedestrians at the intersection", after performing action a "maintain speed", the probability of transitioning to the new state s' "vehicle speed 30km / h, still no pedestrians at the intersection" is 0.9, and the probability of transitioning to s'' "vehicle speed 30km / h, pedestrians appear at the intersection" is 0.1.

[0065] The observation probability O(o|s',a) is a conditional probability, representing the conditional probability that the perception module will obtain the observation result o when the vehicle transitions to state s' after performing action a. It reflects the uncertainty of sensor observation. This probability is directly related to the sensor's measurement error characteristics. For example, when the state s' after performing the "deceleration" action is "pedestrians at the intersection," due to low visibility in rainy weather, the probability of the camera observing "pedestrians at the intersection" is 0.8, and the probability of observing "no pedestrians at the intersection" is 0.2; when the state s' is "no pedestrians at the intersection," the probability of observing "no pedestrians at the intersection" is 1.0.

[0066] The reward function R(s,a) quantifies the reward (or penalty) a vehicle receives immediately after performing action a in its current state s. Its core objective is to guide decisions towards safety, efficiency, and comfort. The reward / penalty setting incorporates multiple dimensions, including at least collision penalties, target approach rewards, speed maintenance rewards, and comfort penalties.

[0067] The discount factor γ is used to balance the weights of immediate and future rewards, reflecting the long-term nature of the decision. Its value ranges from 0 to 1. When γ is close to 1, the decision prioritizes long-term future rewards; when γ is close to 0, the decision favors immediate rewards. The specific value of the discount factor can be determined through training and optimization using historical data for different driving scenarios.

[0068] By integrating the state space S, action space A, observation space Z, state transition probability T, observation probability O, reward function R, and discount factor γ defined above, a complete partially observable Markov decision process model is obtained. This model fully maps the physical process of "vehicle performing actions → environmental state transitions → sensor acquiring observations → evaluating action rewards," providing a structured mathematical framework for subsequent belief inference.

[0069] This embodiment defines a seven-tuple model that includes observable and unobservable states, and combines state transition probabilities and observation probabilities to achieve probabilistic modeling of unknown environmental states.

[0070] In some embodiments, step S220 may include, but is not limited to, steps S221 to S222: Step S221: Obtain the belief at the previous time point, and the state transition probability and the observation probability at the current time point in the partially observable Markov decision process model; Step S222: Based on Bayes' theorem, using the belief at the previous time point as a basis, calculate the belief state corresponding to the vehicle's action at the current time point according to the state transition probability and the observation probability, and obtain the belief corresponding to the current time point.

[0071] In this embodiment, after constructing the POMDP model, beliefs are updated based on Bayes' theorem. A belief is a posterior probability distribution of the true state of the environment. Based on the belief at the previous time point, and combined with the state transition probabilities and observation probabilities of the POMDP model, it is iteratively updated using Bayes' theorem.

[0072] Specifically, firstly, the prior probabilities of the environment in each possible new state after taking an action are calculated by combining the belief from the previous time point with the state transition probabilities. The belief from the previous time point is the agent's (vehicle's) inference about the unknown state of the environment in the previous decision cycle, essentially a probability distribution. For example, if the previous time point inferred that "the probability of a pedestrian crossing the street is 50%, and the probability of going straight is 50%", this probability combination (0.5, 0.5) is the belief from the previous time point, representing the agent's existing knowledge of the environment before acquiring the current observation and taking the current action. Then, when the vehicle actually receives the new sensor observations, it uses the observation probabilities to correct the previous predictions, amplifying the probability of states consistent with the current observation and reducing the probability of inconsistent states, making the belief closer to the actual observation evidence. The sum of the corrected probabilities may not be 1 (not meeting the basic requirements of a probability distribution), so it needs to be normalized using a standardization constant η. After the above calculations and normalization, the belief corresponding to the current time point is obtained. It is the optimal probability estimate of the current environmental state, derived after integrating all historical information, considering predictions of environmental dynamics, and incorporating the latest sensor evidence.

[0073] This embodiment uses Bayesian updates to fuse multi-source observation data in real time and dynamically update the distribution of beliefs about the environmental state, significantly improving the decision robustness of the autonomous driving system in complex scenarios such as occlusion and noise.

[0074] In some embodiments, step S300 may include, but is not limited to, steps S310 to S340: Step S310: Taking the belief corresponding to the current time point as the root node, starting from the root node, traverse all executable actions corresponding to the root node; For each executable action of the root node, the following processing is performed: Step S320: After performing the action under the belief corresponding to the root node and obtaining the corresponding observation results, calculate the belief at the next time point based on Bayes' theorem. Step S330: Take the belief at the next time point as a child node and connect it to the root node to obtain the edge in the belief tree; Step S340: Take the child node as the root node and jump to the step of traversing all executable actions corresponding to the root node starting from the root node, until the depth of the belief tree reaches the preset search boundary.

[0075] In this embodiment, the belief tree is a tree-like data structure used to enumerate and deduce all possible action sequences and observation sequences within a finite number of future time steps, starting from the current moment, and to evaluate their long-term benefits. Its construction is a recursive process. The belief tree is created with the current belief state as the root node; for each node in the current layer, all possible actions are traversed; for each action, all possible observation results are considered, and corresponding child nodes are generated; the above process is repeated with the newly generated child node as the current node; expansion stops when the tree depth reaches a preset search boundary. The belief tree is completed when all nodes have expanded to the preset search boundary (or no new nodes can be expanded). The final belief tree is rooted with the initial belief, uses action-observation combinations as edges, and the updated belief as child nodes, forming a hierarchical structure of "root node → first-level child node → second-level child node → ... → leaf node," completely enumerating all possible "belief-action-observation" evolution paths within the preset depth.

[0076] This embodiment uses the current belief as the root node, traverses the action space, simulates the observation results after the action is executed, generates child node beliefs using Bayes' theorem, and recursively constructs a probabilistic decision tree. This method achieves efficient search under resource constraints through deep boundary control, provides an uncertainty propagation mechanism for generating action sequences of vehicles in partially observable environments, and significantly improves decision robustness in complex scenarios.

[0077] In some embodiments, step S400 may include, but is not limited to, steps S410 to S450: Step S410: Starting from the root node, at each node of the belief tree, the upper limit confidence interval algorithm is used to select the child node with the largest value of the value function as the node to be expanded, until the maximum depth of the belief tree is reached, thus obtaining the target search path. Step S420: Simulate the target search path and calculate the long-term expected reward value of the target search path; Step S430: Backpropagate the long-term expected reward value to all child nodes traversed by the target search path, and update the value function corresponding to each child node traversed by the target search path to obtain the updated value function of each child node; Step S440: Replan the target search path based on the updated value function of each child node until the preset iteration stopping condition is met; Step S450: Based on the target search path obtained in the last iteration, the target action is obtained.

[0078] In this embodiment, the Upper Confidence Bound (UCB) algorithm is a quantitative criterion for balancing the "exploration" and "exploitation" of a search. By integrating action value and exploration level, it selects nodes to ensure that high-value paths are prioritized while also exploring potentially better unknown paths to a certain extent. The value function is a quantitative indicator used to evaluate the quality of belief tree nodes; a higher value indicates a greater decision benefit for the corresponding belief or action. The long-term expected reward is the sum of discounted rewards accumulated after future simulations of the target search path, reflecting the long-term decision benefit of the path and integrating multiple dimensions such as safety, efficiency, and comfort.

[0079] Specifically, starting from the root node of the belief tree (the initial belief at the current time point), the upper bound confidence interval algorithm is used to recursively select child nodes. If a node has untried actions (unexpanded child nodes), the child node corresponding to that action is selected for expansion. If all actions of a node have been tried, the UCB value of each child node is calculated, and the child node with the largest UCB value is selected to proceed to the next level. This process is repeated until the leaf node of the belief tree (the preset tree depth) is reached, forming a target search path from the root node to the leaf node. At each node, the UCB value of all tried actions is calculated, and the child node corresponding to the action with the largest UCB value is selected, thus prioritizing high-value paths while also considering insufficiently explored paths. When the selection phase reaches an incompletely expanded node, an action is selected from the set of untried actions of that node, all possible observations of that action are predicted, and a new belief corresponding to each observation is calculated based on Bayes' theorem. A child node is created for each new belief in the belief tree, and it is connected to the branch of that action.

[0080] The generated target search path is simulated, and the environmental feedback after the action is executed is simulated. Based on the reward function and discount factor of the POMDP model, the long-term expected reward value of the target search path is calculated to quantify the long-term decision benefit of the path. The calculated long-term expected reward value is backpropagated from the leaf nodes of the target search path to the root node, updating the value function and search count for all nodes on the path, achieving dynamic optimization of value assessment. Based on the updated value function, a new target search path is selected again using the UCB algorithm, and the simulation, backpropagation, and value update process is repeated until the preset iteration stopping condition is reached. The preset iteration stopping condition can be: the iteration reaches the preset number of iterations / the search time reaches the upper limit / the value of the value function converges.

[0081] When the iteration meets the stopping condition, the target search path obtained in the last iteration is extracted. The first action of this path (the action corresponding to the edge from the root node to the child node) is the target action corresponding to the current time point.

[0082] This embodiment selects the optimal path in the probability belief space using the UCT algorithm, calculates the long-term expected reward using Monte Carlo simulation, and updates the node value function through backpropagation. It balances exploration and utilization in iterative optimization and finally outputs the target action that maximizes long-term benefits, significantly improving the vehicle's decision robustness and real-time performance in partially observable environments.

[0083] In some embodiments, step S500 may include, but is not limited to, steps S510 to S560: Step S510: Based on the environmental information and the belief corresponding to the current time point, a path planning algorithm is used to search for the path with the minimum comprehensive cost. The comprehensive cost includes at least: static obstacle collision cost, pedestrian collision probability cost, and path smoothness cost; the target action includes a first acceleration and a first wheel steering angle. Step S520: Calculate the initial acceleration and initial wheel steering angle of the path with the minimum overall cost to obtain the second acceleration and the second wheel steering angle; Step S530: Perform vector fusion of the first acceleration and the second acceleration to obtain the target acceleration; Step S540: Perform vector fusion on the first wheel steering angle and the second wheel steering angle to obtain the target wheel steering angle; Step S550: If the target wheel steering angle exceeds the preset maximum reasonable range, then adjust the target wheel steering angle to the boundary value of the maximum reasonable range. Step S560: Control the vehicle to perform a new target action, the new target action including the target acceleration and the target wheel steering angle.

[0084] In this embodiment, the target action obtained by Monte Carlo tree search includes the first acceleration and the first wheel steering angle. To further improve the safety and smoothness of action execution, parameter optimization and constraint verification need to be performed in conjunction with path planning. Based on the environmental information and beliefs at the current time point, a path planning algorithm is used to search for the driving path with the minimum overall cost.

[0085] Specifically, the Hybrid A* algorithm can be used for path search. This algorithm combines the heuristic optimization of the A* algorithm with vehicle kinematic constraints. The search node is defined as (x, y, θ) (position + attitude angle), ensuring that the generated path conforms to the vehicle's steering characteristics (e.g., the attitude angle θ determines the direction of the next path point). Based on environmental information (position coordinates of static obstacles, road boundary geometric parameters; real-time position and velocity of dynamic targets), current-time beliefs (probability distribution of pedestrian target position and behavioral intentions), and vehicle state (current position, attitude angle, instantaneous velocity, etc.), the comprehensive cost of the path is calculated. The Hybrid A* algorithm is then used to traverse all feasible nodes, selecting the path with the minimum comprehensive cost.

[0086] Then, based on the position-attitude change pattern of the optimal path, physically feasible initial action parameters are derived, including the second acceleration and the second wheel steering angle. A weighted vector fusion method is used to fuse the first and second accelerations, and the first and second wheel steering angles, respectively, to obtain the target acceleration and the target wheel steering angle. For example, the weights can be dynamically adjusted based on the uncertainty of the belief; the higher the belief entropy, the higher the weight of the path planning result. During the fusion process, the vector direction of the parameters must be kept consistent (e.g., the sign of acceleration corresponds to acceleration / deceleration, and the sign of steering angle corresponds to left / right turn).

[0087] The reasonableness of the target wheel steering angle obtained by fusion is checked. If it exceeds the preset maximum reasonable range, boundary adjustment is performed. The maximum reasonable range can be determined based on the physical constraints of the vehicle steering system (such as the maximum steering motor angle) and road driving safety requirements. If the target wheel steering angle exceeds the maximum reasonable range, the target wheel steering angle is adjusted to the boundary value of the maximum reasonable range.

[0088] The adjusted target acceleration and target wheel steering angle are used as new target actions, which are converted into control signals that can be recognized by the vehicle actuators. The control signals are sent to the corresponding execution units to control the vehicle to drive according to the optimized target actions. At the same time, execution feedback data is collected in real time to provide input for the next round of decision-making.

[0089] This embodiment generates control commands that balance strategy consistency and trajectory feasibility by fusing the target actions of the decision-making layer and the constraint actions of the path planning layer. The comprehensive cost function quantifies the collision probabilities of static obstacles and pedestrians, as well as path smoothness, and employs vector fusion and steering angle safety constraint mechanisms to ensure safe and smooth motion control of the vehicle in partially observable environments.

[0090] This application's embodiments utilize a Bayesian-based belief iteration method to achieve accurate probabilistic inference of unknown environmental information in pedestrian trajectory prediction, significantly improving prediction accuracy. The constructed belief tree structure provides an efficient decision-making framework for the Monte Carlo tree search algorithm, supporting rapid search for optimal actions in complex dynamic environments. Combined with a path correction algorithm, the decision actions are subject to safety limits and dynamic fusion, ensuring that vehicle acceleration and steering angle remain within reasonable ranges, avoiding exceeding dynamic constraints. The overall solution adopts a modular design, significantly reducing computational complexity through sparse belief updates, lightweight tree search, and linear fusion strategies. While ensuring real-time performance, it achieves highly reliable autonomous driving decision control without requiring high-performance hardware.

[0091] Please see Figure 4 This application also provides a vehicle control device 600 that can implement the above-described vehicle control method. The device includes: The acquisition module 10 is used to acquire the vehicle's status information at the current time and the environmental information of the environment in which the vehicle is located; The calculation module 20 is used to obtain the belief corresponding to the current time point based on the state information and the environment information and the Bayesian rule. The belief is a probability distribution obtained by inferring the true state of the environment information based on the unknown information in the environment information. Construction module 30 is used to construct a belief tree with the belief corresponding to the current time point as the initial belief. The root node in the belief tree is the initial belief. Each branch extending from the parent node in the belief tree corresponds to an executable action. The action is connected to one or more child nodes. Each child node corresponds to a new belief corresponding to different observation results obtained by executing the action. Search module 40 is used to obtain the target action corresponding to the current time point from the belief tree based on the Monte Carlo tree search algorithm; The control module 50 is used to control the vehicle to perform the target action.

[0092] In some implementations, the computing module 20 may include: A submodule is constructed to build a partially observable Markov decision process model based on the vehicle's current state information and environmental information. The first calculation submodule is used to obtain the belief corresponding to the current time point based on the partially observable Markov decision process model using Bayes' rule.

[0093] In some implementations, building a submodule may include: A definition unit is used to define a state space, action space, observation space, state transition probability, observation probability, reward function, and discount factor based on the state information and environmental information. The state space includes the state of the vehicle and pedestrians in the surrounding environment. The action space is the set of actions that the vehicle can perform. The observation space is the set of all observation results obtained by the vehicle after performing an action at the current time point. The state transition probability is the conditional probability that the vehicle obtains new state information after performing an action at the current time point. The observation probability is the conditional probability that the vehicle obtains corresponding environmental information after performing an action. The reward function is the reward or penalty obtained by the vehicle after performing an action at the current time point. The discount factor is used to balance the weights of immediate rewards and future rewards. The construction unit is used to obtain the partially observable Markov decision process model based on the state space, the action space, the observation space, the state transition probability, the observation probability, the reward function, and the discount factor.

[0094] In some implementations, the first computing submodule may include: The acquisition unit is used to acquire the belief at the previous time point, as well as the state transition probability and the observation probability at the current time point in the partially observable Markov decision process model; The calculation unit is used to calculate the belief state corresponding to the control of the vehicle to perform an action at the current time point based on Bayes' theorem and the belief at the previous time point, according to the state transition probability and the observation probability, so as to obtain the belief corresponding to the current time point.

[0095] In some implementations, the building module 30 may include: The traversal submodule is used to traverse all executable actions corresponding to the current time point, starting from the belief corresponding to the current time point. The second calculation submodule is used to perform the following processing for each executable action of the root node: after executing the action under the belief corresponding to the root node and obtaining the corresponding observation results, calculate the belief at the next time point based on Bayes' rule; The connection submodule is used to connect the belief at the next time point as a child node to the root node to obtain the edge in the belief tree; The jump-rotor module is used to take the child node as the root node and jump to the step of executing all executable actions corresponding to the root node, starting from the root node, until the depth of the belief tree reaches the preset search boundary.

[0096] In some implementations, the search module 40 may include: The selection submodule is used to select the child node with the largest value of the value function as the expansion node at each node of the belief tree, starting from the root node, and using the upper limit confidence interval algorithm, until the maximum depth of the belief tree, to obtain the target search path; The simulation submodule is used to simulate the target search path and calculate the long-term expected reward value of the target search path. The update submodule is used to backpropagate the long-term expected reward value to all child nodes traversed by the target search path, and update the value function corresponding to each child node traversed by the target search path to obtain the updated value function of each child node. The planning submodule is used to replan the target search path based on the updated value function of each child node until the preset iteration stopping condition is met. The determination submodule is used to obtain the target action based on the target search path obtained in the last iteration.

[0097] In some implementations, the control module 50 may include: The search submodule is used to search for the path with the minimum comprehensive cost based on the environmental information and the belief corresponding to the current time point using a path planning algorithm. The comprehensive cost includes at least: static obstacle collision cost, pedestrian collision probability cost and path smoothness cost. The target action includes a first acceleration and a first wheel steering angle. The third calculation submodule is used to calculate the initial acceleration and initial wheel steering angle of the path with the minimum comprehensive cost, and to obtain the second acceleration and the second wheel steering angle. The first fusion submodule is used to perform vector fusion of the first acceleration and the second acceleration to obtain the target acceleration.

[0098] The second fusion submodule is used to perform vector fusion of the first wheel steering angle and the second wheel steering angle to obtain the target wheel steering angle. The judgment submodule is used to adjust the target wheel steering angle to the boundary value of the maximum reasonable range if the target wheel steering angle exceeds the preset maximum reasonable range. The control submodule is used to control the vehicle to perform a new target action, which includes the target acceleration and the target wheel steering angle.

[0099] The specific implementation of this vehicle control device is basically the same as the specific implementation of the vehicle control method described above, and will not be repeated here.

[0100] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described vehicle control method. This electronic device can be any smart terminal, including a tablet computer, an in-vehicle computer, or similar device.

[0101] Please see Figure 5 , Figure 5 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 801 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 802 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 using the vehicle control method of the embodiments of this application. The 803 input / output interface is used to implement information input and output. The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804); The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.

[0102] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described vehicle control method.

[0103] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0104] The vehicle control method, vehicle control device, electronic device, and storage medium provided in this application acquire the vehicle's current state and environmental information, then process unknown information in the environment based on Bayes' theorem to infer the current belief describing the probability distribution of the true state of the environment. A belief tree is constructed with the current belief as the root node, where branches represent executable actions, and child nodes correspond to new beliefs updated after different observations following the action. Finally, the optimal target action is determined from the belief tree using a Monte Carlo tree search algorithm, and the vehicle is controlled to execute the target action, achieving robust real-time decision-making in complex and uncertain environments.

[0105] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0106] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0107] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0108] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0109] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0110] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0111] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0112] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0113] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0114] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0115] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A vehicle control method, characterized in that, The method includes: Obtain the vehicle's status information at the current point in time and the environmental information of the vehicle's surroundings; Based on the state information and the environment information, the belief corresponding to the current time point is obtained based on Bayes' theorem. The belief is a probability distribution obtained by inferring the true state of the environment information based on unknown information in the environment information. A belief tree is constructed using the belief corresponding to the current time point as the initial belief. The root node in the belief tree is the initial belief. Each branch extending from the parent node in the belief tree corresponds to an executable action. The action is connected to one or more child nodes. Each child node corresponds to a new belief corresponding to different observation results obtained by executing the action. Based on the Monte Carlo tree search algorithm, the target action corresponding to the current time point is obtained from the belief tree; Control the vehicle to perform the target action; The target motion includes a first acceleration and a first wheel steering angle; Controlling the vehicle to perform the target action includes: Based on the environmental information and the belief corresponding to the current time point, a path planning algorithm is used to search for the path with the minimum comprehensive cost. The comprehensive cost includes at least: static obstacle collision cost, pedestrian collision probability cost, and path smoothness cost. Calculate the initial acceleration and initial wheel steering angle of the path with the minimum overall cost to obtain the second acceleration and the second wheel steering angle; The first acceleration and the second acceleration are vector-fused to obtain the target acceleration; The first wheel steering angle and the second wheel steering angle are vector-fused to obtain the target wheel steering angle; If the target wheel steering angle exceeds the preset maximum reasonable range, then the target wheel steering angle will be adjusted to the boundary value of the maximum reasonable range; Control the vehicle to perform a new target action, the new target action including the target acceleration and the target wheel steering angle.

2. The method according to claim 1, characterized in that, The step of obtaining the belief corresponding to the current time point based on the state information and the environment information and using Bayes' theorem includes: Based on the vehicle's current state information and environmental information, a partially observable Markov decision process model is constructed. Based on the partially observable Markov decision process model, the belief corresponding to the current time point is obtained through Bayes' rule.

3. The method according to claim 2, characterized in that, The construction of a partially observable Markov decision process model based on the vehicle's current state information and environmental information includes: Based on the state information and environmental information, a state space, action space, observation space, state transition probability, observation probability, reward function, and discount factor are defined. The state space includes the state of the vehicle and pedestrians in the surrounding environment. The action space is the set of actions that the vehicle can perform. The observation space is the set of all observations obtained by the vehicle after performing an action at the current time point. The state transition probability is the conditional probability that the vehicle obtains new state information after performing an action at the current time point. The observation probability is the conditional probability that the vehicle obtains corresponding environmental information after performing an action. The reward function is the reward or penalty obtained by the vehicle after performing an action at the current time point. The discount factor is used to balance the weights of immediate rewards and future rewards. The partially observable Markov decision process model is obtained based on the state space, the action space, the observation space, the state transition probability, the observation probability, the reward function, and the discount factor.

4. The method according to claim 3, characterized in that, The belief corresponding to the current time point, obtained through Bayes' theorem based on the partially observable Markov decision process model, includes: Obtain the belief at the previous time point, and the state transition probability and the observation probability at the current time point in the partially observable Markov decision process model; Based on Bayes' theorem, using the belief at the previous time point as a basis, the belief state corresponding to the vehicle's action at the current time point is calculated according to the state transition probability and the observation probability, thus obtaining the belief corresponding to the current time point.

5. The method according to claim 1, characterized in that, The construction of the belief tree using the belief corresponding to the current time point as the initial belief includes: Using the belief corresponding to the current time point as the root node, starting from the root node, traverse all executable actions corresponding to the root node; For each executable action of the root node, the following processing is performed: After performing the action under the belief corresponding to the root node and obtaining the corresponding observation results, the belief at the next time point is calculated based on Bayes' theorem. The belief at the next time point is taken as a child node and connected to the root node to obtain the edge in the belief tree; The child node is taken as the root node, and the process jumps to the step of traversing all executable actions corresponding to the root node, starting from the root node, until the depth of the belief tree reaches the preset search boundary.

6. The method according to claim 1, characterized in that, The Monte Carlo tree search algorithm, which obtains the target action corresponding to the current time point from the belief tree, includes: Starting from the root node, at each node of the belief tree, the upper limit confidence interval algorithm is used to select the child node with the largest value of the value function as the node to be expanded, until the maximum depth of the belief tree is reached, thus obtaining the target search path; The target search path is simulated, and the long-term expected reward value of the target search path is calculated; The long-term expected reward value is backpropagated to all child nodes traversed by the target search path, and the value function corresponding to each child node traversed by the target search path is updated to obtain the updated value function of each child node. The target search path is replanned based on the updated value function of each child node until the preset iteration stopping condition is met; The target action is obtained based on the target search path obtained in the last iteration.

7. A vehicle control device, characterized in that, The device includes: The acquisition module is used to acquire the vehicle's status information at the current time and the environmental information of the environment in which the vehicle is located; The calculation module is used to obtain the belief corresponding to the current time point based on the state information and the environment information and the Bayesian rule. The belief is a probability distribution obtained by inferring the true state of the environment information based on unknown information in the environment information. The construction module is used to construct a belief tree with the belief corresponding to the current time point as the initial belief. The root node in the belief tree is the initial belief. Each branch extending from the parent node in the belief tree corresponds to an executable action. The action is connected to one or more child nodes. Each child node corresponds to a new belief corresponding to different observation results obtained by executing the action. The search module is used to obtain the target action corresponding to the current time point from the belief tree based on the Monte Carlo tree search algorithm. A control module is used to control the vehicle to perform the target action; the target action includes a first acceleration and a first wheel steering angle; based on the environmental information and the belief corresponding to the current time point, a path planning algorithm is used to search for the path with the minimum comprehensive cost, the comprehensive cost including at least: static obstacle collision cost, pedestrian collision probability cost, and path smoothness cost; the initial acceleration and initial wheel steering angle of the path with the minimum comprehensive cost are calculated to obtain a second acceleration and a second wheel steering angle; the first acceleration and the second acceleration are vector-fused to obtain a target acceleration; the first wheel steering angle and the second wheel steering angle are vector-fused to obtain a target wheel steering angle; if the target wheel steering angle exceeds a preset maximum reasonable range, the target wheel steering angle is adjusted to the boundary value of the maximum reasonable range; the vehicle is controlled to perform a new target action, the new target action including the target acceleration and the target wheel steering angle.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the vehicle control method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the vehicle control method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Vehicle control method and device based on decision tree and Bayesian network

    CN113232674A

  • Automatic driving behavior decision method, driving decision device and readable storage medium

    CN119773807A