Intelligent agent assisted driving method and device, electronic device, and storage medium

CN121404312BActive Publication Date: 2026-08-11SHENZHEN CONSYS SCI&TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]然而,相关技术中,智能体辅助行驶的效率和准确性依然较低,难以满足智能驾驶对于快速响应和精准决策的需求

Benefits of technology

[0016] The intelligent agent-assisted driving method, device, electronic device, and storage medium proposed in this application acquire assisted driving commands and, based on these commands, utilize an intelligent agent to obtain the driving state space information and driving action space information of the vehicle to be assisted at the current time. The driving action space information includes multiple preset driving commands. Then, based on the driving state space information, the intelligent agent performs multiple value evaluation operations on the multiple preset driving commands to obtain action reward quantification values ​​corresponding to the preset driving commands. Each value evaluation operation includes: using an intelligent agent to determine one preset driving command as a candidate driving command from the multiple preset driving commands based on an initial reward quantification value; using the intelligent agent to simulate the action trajectory based on the candidate driving command and the driving state space information to obtain the driving trajectory information of the vehicle to be assisted; then, using the intelligent agent to evaluate the reward of the driving trajectory information based on the initial reward quantification value and obtain the action reward quantification value of the candidate driving command; and using the intelligent agent to use the action reward quantification value of the candidate driving command as the initial reward quantification value for the next value evaluation operation. Finally, based on the action reward quantification values ​​corresponding to the multiple preset driving commands, the intelligent agent determines the target driving command of the vehicle to be assisted, and uses the intelligent agent to assist in controlling the driving action of the vehicle to be assisted based on the target driving command.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121404312B_ABST
    Figure CN121404312B_ABST
Patent Text Reader

Abstract

This application provides an intelligent agent-assisted driving method, device, electronic device, and storage medium, belonging to the field of intelligent driving technology. The method includes: acquiring assisted driving instructions; using an intelligent agent to acquire the driving state space information and driving action space information of the vehicle to be assisted at the current time, the driving action space information including multiple preset driving instructions; using the intelligent agent to perform multiple value evaluation operations on the multiple preset driving instructions based on the driving state space information to obtain action reward quantification values ​​corresponding to the multiple preset driving instructions; using the intelligent agent to determine the target driving instruction of the vehicle to be assisted based on the action reward quantification values ​​corresponding to the multiple preset driving instructions; and using the intelligent agent to assist in controlling the driving action of the vehicle to be assisted based on the target driving instruction. This application embodiment can improve the efficiency and accuracy of intelligent agent-assisted driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent driving technology, and in particular to an intelligent agent-assisted driving method and device, electronic device and storage medium. Background Technology

[0002] With the rapid development of technology, intelligent agents (AI agents), as entities based on artificial intelligence technology, are capable of autonomously perceiving their environment and making decisions, and have been widely used in many fields. For example, in the field of intelligent driving, intelligent agents can assist drivers in driving vehicles, such as automatically adjusting driving speed, thereby improving driving efficiency.

[0003] In the field of intelligent driving, due to the increasing complexity of traffic environments, drivers may be unable to react in time to sudden situations or complex road conditions. Therefore, intelligent agents need to quickly perceive the vehicle's surroundings and make decisions to assist the driver in controlling the vehicle's actions, such as accelerating to change lanes or accelerating straight ahead, thereby avoiding potential traffic risks and ensuring driver safety.

[0004] However, among related technologies, the efficiency and accuracy of agent-assisted driving are still relatively low, making it difficult to meet the needs of intelligent driving for rapid response and accurate decision-making. Summary of the Invention

[0005] The main objective of this application is to provide an intelligent agent-assisted driving method, device, electronic device, and storage medium, aiming to improve the efficiency and accuracy of intelligent agent-assisted driving.

[0006] To achieve the above objectives, a first aspect of this application proposes an intelligent agent-assisted driving method, the method comprising: Obtain assisted driving instructions, and use an intelligent agent to obtain the driving state space information and driving action space information of the vehicle to be assisted at the current time based on the assisted driving instructions. The driving action space information includes multiple preset driving instructions. Based on the driving state space information, the intelligent agent performs multiple value evaluation operations on the multiple preset driving instructions to obtain the action reward quantification value corresponding to the multiple preset driving instructions; Based on the action reward quantification value corresponding to the multiple preset driving instructions, the intelligent agent determines the target driving instruction of the vehicle to be assisted, and based on the target driving instruction, the intelligent agent assists in controlling the driving action of the vehicle to be assisted. Each of the aforementioned value assessment operations includes: Based on the initial reward quantification value, the agent determines one preset driving instruction as a candidate driving instruction from the plurality of preset driving instructions. Based on the candidate driving instructions, the intelligent agent simulates the motion trajectory of the driving state space information to obtain the driving trajectory information of the vehicle to be assisted. Based on the initial reward quantification value, the agent performs a reward evaluation on the driving trajectory information to obtain the action reward quantification value of the candidate driving instruction; The agent uses the action reward quantification value of the candidate driving instruction as the initial reward quantification value for the next value assessment operation.

[0007] In some embodiments, the step of using the agent to evaluate the driving trajectory information based on the initial reward quantification value to obtain the action reward quantification value of the candidate driving instruction includes: The intelligent agent is used to evaluate the trajectory utility of the driving trajectory information to obtain trajectory utility parameters; The intelligent agent is used to obtain the first number of times the candidate driving command is selected. Based on the number of times the first instruction is selected and the initial reward quantification value, the agent adjusts the trajectory utility parameters to obtain the action reward quantification value of the candidate driving instruction.

[0008] In some embodiments, the step of adjusting the trajectory utility parameters using the agent based on the number of times the first instruction is selected and the initial reward quantification value to obtain the action reward quantification value of the candidate driving instruction includes: The agent is used to obtain the initial reward quantification value corresponding to all executed value assessment operations, and the agent is used to determine the maximum initial reward quantification value based on the initial reward quantification value corresponding to all executed value assessment operations. The importance weight of the driving trajectory information is determined by the agent based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and the preset greedy strategy parameters. The trajectory utility parameters are weighted according to the importance weights by the intelligent agent to obtain weighted trajectory utility parameters; Based on the number of times the first instruction is selected and the initial reward quantification value, the agent adjusts the weighted trajectory utility parameters to obtain the action reward quantification value of the candidate driving instruction.

[0009] In some embodiments, determining the importance weight of the driving trajectory information using the agent based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and preset greedy strategy parameters includes: Based on the maximum initial reward quantification value, the agent generates a behavior probability distribution for the driving trajectory information to obtain the target policy probability distribution. Based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and the preset greedy strategy parameters, the agent generates a target probability distribution for the driving trajectory information to obtain a behavior strategy probability distribution. The ratio of the probability distribution of the target policy to the probability distribution of the behavior policy is used by the intelligent agent to determine the importance weight of the driving trajectory information.

[0010] In some embodiments, the step of using the intelligent agent to evaluate the trajectory utility of the driving trajectory information and obtain trajectory utility parameters includes: The intelligent agent is used to obtain the preset maximum speed of the vehicle to be assisted; Based on the preset maximum vehicle speed, the intelligent agent evaluates the driving efficiency of the driving trajectory information to obtain a quantitative value of driving efficiency. The intelligent agent is used to assess driving safety based on the driving trajectory information to obtain a quantitative value of driving safety. The trajectory utility parameters are determined by the intelligent agent based on preset weighting coefficients, the quantified value of driving efficiency, and the quantified value of driving safety.

[0011] In some embodiments, determining a preset driving instruction as a candidate driving instruction from the plurality of preset driving instructions using the agent based on an initial reward quantification value includes: The intelligent agent is used to obtain the number of times the second instruction is selected for each preset driving instruction, and the intelligent agent is used to obtain the number of times the evaluation operation is performed to perform the value evaluation operation. Based on the initial reward quantification value, the number of times the second instruction was selected, and the number of evaluation operations, the agent performs a confidence score on each preset driving instruction to obtain the instruction confidence score corresponding to each preset driving instruction; Based on the instruction confidence score corresponding to each preset driving instruction, the agent determines one preset driving instruction as a candidate driving instruction from the plurality of preset driving instructions.

[0012] In some embodiments, the step of using the intelligent agent to perform multiple value evaluation operations on the plurality of preset driving instructions based on the driving state space information to obtain the action reward quantification value corresponding to the plurality of preset driving instructions includes: Based on the driving state space information, the intelligent agent performs multiple value assessment operations on the multiple preset driving commands, and obtains the cumulative operation time of all executed value assessment operations using the intelligent agent. When the cumulative operation time is less than or equal to a preset time threshold, the agent continues to perform value evaluation operations on the multiple preset driving instructions based on the driving state space information, so as to obtain the action reward quantification value corresponding to the multiple preset driving instructions. When the cumulative operation time exceeds a preset time threshold, the value assessment operation for multiple preset driving commands is stopped.

[0013] To achieve the above objectives, a second aspect of this application provides an intelligent agent-assisted driving device, the device comprising: An environmental perception unit is used to acquire assisted driving instructions and to acquire, based on the assisted driving instructions, the driving state spatial information and driving action spatial information of the vehicle to be assisted at the current time using the intelligent agent. The driving action spatial information includes multiple preset driving instructions. The value assessment unit is used to perform multiple value assessment operations on the multiple preset driving instructions using the intelligent agent based on the driving state space information, so as to obtain the action reward quantification value corresponding to the multiple preset driving instructions. The assisted driving unit is used to determine the target driving instruction of the vehicle to be assisted based on the action reward quantification value corresponding to the plurality of preset driving instructions using the intelligent agent, and to assist in controlling the driving action of the vehicle to be assisted based on the target driving instruction using the intelligent agent. Each of the aforementioned value assessment operations includes: Based on the initial reward quantification value, the agent determines one preset driving instruction as a candidate driving instruction from the plurality of preset driving instructions. Based on the candidate driving instructions, the intelligent agent simulates the motion trajectory of the driving state space information to obtain the driving trajectory information of the vehicle to be assisted. Based on the initial reward quantification value, the agent performs a reward evaluation on the driving trajectory information to obtain the action reward quantification value of the candidate driving instruction; The agent uses the action reward quantification value of the candidate driving instruction as the initial reward quantification value for the next value assessment operation.

[0014] To achieve the above objectives, a third aspect of the present application provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described in the first aspect.

[0015] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0016] The intelligent agent-assisted driving method, device, electronic device, and storage medium proposed in this application acquire assisted driving commands and, based on these commands, utilize an intelligent agent to obtain the driving state space information and driving action space information of the vehicle to be assisted at the current time. The driving action space information includes multiple preset driving commands. Then, based on the driving state space information, the intelligent agent performs multiple value evaluation operations on the multiple preset driving commands to obtain action reward quantification values ​​corresponding to the preset driving commands. Each value evaluation operation includes: using an intelligent agent to determine one preset driving command as a candidate driving command from the multiple preset driving commands based on an initial reward quantification value; using the intelligent agent to simulate the action trajectory based on the candidate driving command and the driving state space information to obtain the driving trajectory information of the vehicle to be assisted; then, using the intelligent agent to evaluate the reward of the driving trajectory information based on the initial reward quantification value and obtain the action reward quantification value of the candidate driving command; and using the intelligent agent to use the action reward quantification value of the candidate driving command as the initial reward quantification value for the next value evaluation operation. Finally, based on the action reward quantification values ​​corresponding to the multiple preset driving commands, the intelligent agent determines the target driving command of the vehicle to be assisted, and uses the intelligent agent to assist in controlling the driving action of the vehicle to be assisted based on the target driving command.

[0017] This application utilizes the current driving state space information of the vehicle to be assisted and multiple preset driving commands. An intelligent agent performs multiple value evaluation operations on these preset driving commands. Each value evaluation operation involves the intelligent agent selecting a preset driving command to simulate a motion trajectory based on the driving state space information. The simulated driving trajectory information is then evaluated for rewards until a quantitative value of the action reward corresponding to each preset driving command is obtained. Based on the quantitative value of the action reward for each preset driving command, a target driving command is determined. Finally, the intelligent agent assists in controlling the driving action of the vehicle to be assisted according to the target driving command. In this way, the simulated trajectory can reach more random events based on the preset driving commands, reducing the risk of the simulated trajectory repeatedly occurring between a few local optima. Therefore, this application can improve the efficiency and accuracy of agent-assisted driving. Attached Figure Description

[0018] Figure 1 This is a flowchart of the intelligent agent-assisted driving method provided in the embodiments of this application; Figure 2 This is a flowchart of the value assessment operation provided in the embodiments of this application; Figure 3 This is a flowchart illustrating a specific implementation of the intelligent agent-assisted driving method provided in this application embodiment; Figures 4A to 4C This is an experimental result diagram of the first application of the intelligent agent-assisted driving method provided in the embodiments of this application; Figures 5A to 5C This is an experimental result diagram of the second application of the intelligent agent-assisted driving method provided in the embodiments of this application; Figures 6A to 6C This is an experimental result diagram of the third application of intelligent agent-assisted driving method provided in the embodiments of this application; Figures 7A to 7C This is an experimental result diagram of the fourth application of intelligent agent-assisted driving method provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of the intelligent agent-assisted driving device provided in the embodiments of this application; Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0021] The widespread application of agent-assisted driving methods has significantly improved the real-time decision-making and trajectory smoothness in complex urban roads, highway ramps, and dynamic obstacle scenarios, providing strong support for intelligent driving and effectively improving overall vehicle traffic efficiency, energy economy, and passenger comfort. However, existing agent-assisted driving methods still have significant shortcomings in practical applications. For example, existing technologies typically use rule-based policy networks for end-to-end training to output the vehicle's instantaneous optimal driving commands. However, this approach relies on a large number of prior rules and offline reward calibration, which can easily weaken the policy robustness due to mismatch between the environment and vehicle state, thereby reducing the reliability of assisted driving in complex scenarios. Alternatively, related technologies also use Monte Carlo Tree Search (MCTS) algorithms to obtain the vehicle's future driving trajectory to control the vehicle's acceleration, steering, and lane-changing actions in subsequent moments. However, this approach relies solely on a single probability distribution for sampling during the trajectory selection phase. If a greedy strategy is directly used to expand nodes, the simulated trajectory is easily trapped in a few local optima, leading to biased leaf node value estimation. Simultaneously, MCTS requires unfolding a massive search tree layer by layer, with computational complexity increasing exponentially with the planning time domain, making it difficult to complete real-time decisions within milliseconds, thus weakening the reliability and timeliness of assisted driving in complex dynamic scenarios. Therefore, this application provides an intelligent agent-assisted driving method and device, electronic device, and storage medium, aiming to improve the efficiency and accuracy of intelligent agent-assisted driving.

[0022] The intelligent agent-assisted driving method provided in this application relates to the field of intelligent driving technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the intelligent agent-assisted driving method, but is not limited to the above forms.

[0023] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0024] Figure 1 This is an optional flowchart of the intelligent agent-assisted driving method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S103: Step S101: Obtain assisted driving instructions, and use the intelligent agent to obtain the driving state space information and driving action space information of the vehicle to be assisted at the current time according to the assisted driving instructions. The driving action space information includes multiple preset driving instructions. Step S102: Based on the driving state space information, the intelligent agent performs multiple value evaluation operations on multiple preset driving instructions to obtain the action reward quantification value corresponding to the multiple preset driving instructions. Step S103: Based on the action reward quantification values ​​corresponding to multiple preset driving instructions, the intelligent agent determines the target driving instruction of the vehicle to be assisted, and based on the target driving instruction, the intelligent agent assists in controlling the driving action of the vehicle to be assisted. Each valuation operation includes: Based on the initial reward quantification value, the agent determines one preset driving instruction as a candidate driving instruction from multiple preset driving instructions. Based on the candidate driving instructions, the intelligent agent simulates the motion trajectory of the driving state space information to obtain the driving trajectory information of the vehicle to be assisted. Based on the initial reward quantification value, the agent evaluates the driving trajectory information to obtain the action reward quantification value of the candidate driving command. The agent uses the action reward quantification value of the candidate driving command as the initial reward quantification value for the next value evaluation operation.

[0025] Steps S101 to S103 of this embodiment involve using the driving state space information of the vehicle to be assisted at the current time and multiple preset driving commands. An intelligent agent performs multiple value evaluation operations on these preset driving commands. Each value evaluation operation uses the intelligent agent to select a preset driving command to simulate the motion trajectory of the driving state space information. The simulated driving trajectory information is then evaluated for rewards until a quantitative value of the action reward corresponding to each preset driving command is obtained. A target driving command is then determined based on the quantitative value of the action reward corresponding to each preset driving command. Finally, the intelligent agent assists in controlling the driving action of the vehicle to be assisted according to the target driving command. In this way, the simulated trajectory can reach more random events based on the preset driving commands, reducing the risk of the simulated trajectory repeatedly occurring among a few local optima. Therefore, this application can improve the efficiency and accuracy of intelligent agent-assisted driving.

[0026] In step S101 of some embodiments, the assisted driving command may refer to a set of commands generated by an onboard computing platform or cloud-based decision-making system, used to activate the intelligent agent to enter the assisted decision-making mode and obtain the driving state space information and driving action space information of the vehicle to be assisted at the current time. The vehicle to be assisted may refer to a vehicle requiring assisted driving decision-making and action control. The intelligent agent may refer to an artificial intelligence entity installed on the vehicle to be assisted, capable of autonomously perceiving the environment and making decisions. The driving state space information may refer to a quantifiable set of states of the vehicle to be assisted at the current time, including road information, position, driving direction, speed, acceleration, and surrounding vehicle information. For example, the driving state space information can be defined as follows:

[0027]

[0028] in, Indicates the current time step. This indicates road information at the current time (such as lane type, curvature, slope, drivable area boundaries, and road obstacle conditions). This represents the set of quantifiable states of the vehicle to be assisted, including its position, velocity, acceleration, and direction of travel at the current time. This indicates the current position of the vehicle to be assisted (which can be represented by three-dimensional coordinates). (The unit can be expressed as meters per second, " "" indicates the speed of the vehicle to be assisted at the current time. (The unit can be expressed as meters per second squared.) "" indicates the acceleration of the vehicle to be assisted at the current time. This indicates the direction of travel of the vehicle to be assisted at the current time (which can be represented by a three-dimensional vector). This indicates the information of surrounding vehicles for the vehicle requiring assistance at the current time. This also includes a quantifiable set of states related to the position, speed, acceleration, and direction of travel of surrounding vehicles. It's understandable that if there are multiple surrounding vehicles, then... It can be represented as a list, where each element in the list has a and Same data structure. This is represented as spatial information about the driving status.

[0029] Driving motion space information can refer to a set of selectable actions containing multiple preset driving commands. Preset driving commands are a collection of commands generated by the agent based on preset acceleration and preset lane change commands, used to directly control the driving actions of the vehicle to be assisted. For example, preset driving commands can be defined as follows:

[0030] in, Indicates the preset acceleration, for example, It can be , , , or etc., without specifying the exact type. This indicates a preset lane change command. For example, the preset lane change command could be left lane change, lane keeping, right lane change, etc., without any specific limitation. This refers to the preset driving instructions generated by the intelligent agent based on preset acceleration and preset lane change commands.

[0031] In step S102 of some embodiments, the value assessment operation can refer to the process of using an intelligent agent to simulate trajectories for multiple preset driving instructions one by one based on driving state space information, so as to quantify the reward value of the simulated trajectory corresponding to each preset driving instruction. The action reward quantification value can refer to the quantification value used to indicate the quality of the simulated trajectory after performing multiple value assessment operations on multiple preset driving instructions based on driving state space information. It is understood that the value assessment operation can be executed repeatedly until a corresponding action reward quantification value is obtained for each preset driving instruction; or, when computing resources are scarce or real-time requirements are strict, it can be terminated early after only completing the assessment of some instructions; or, after obtaining the action reward quantification value corresponding to each preset driving instruction, the value assessment operation can continue to be executed, and the original reward value can be iteratively optimized using new simulation results to further improve the accuracy of instruction selection.

[0032] It should be noted that you should refer to [link / reference]. Figure 2 , Figure 2This is a flowchart of a value assessment operation provided in an embodiment of this application. Specifically, it includes: firstly, using an intelligent agent to determine a candidate driving instruction from multiple preset driving instructions based on an initial reward quantification value. When this is the first value assessment operation, the initial reward quantification value can refer to a pre-set quantification value used to determine the candidate driving instruction. For example, the initial reward quantification value can be zero or other set values, without specific limitations. Alternatively, when this is not the first value assessment operation, the initial reward quantification value can refer to the action reward quantification value obtained from the previous value assessment operation. The candidate driving instruction can refer to a preset driving instruction determined by the intelligent agent from multiple preset driving instructions based on the initial reward quantification value. For example, if there are five preset driving instructions, they are represented as follows: , , , , The initial reward quantification value is 0. The agent can calculate the confidence score corresponding to each preset driving command under this initial reward quantification value. Since the initial reward quantification value is 0, the confidence scores corresponding to the five preset driving commands are all the same. Therefore, the first value assessment operation randomly selects one of the five preset driving commands. As candidate driving instructions, it is understandable that in subsequent value assessment operations, since the initial reward quantification value will be updated based on the action reward quantification value obtained from the previous value assessment operation, the confidence score corresponding to each preset driving instruction will be different under the initial reward quantification value calculated by the agent in subsequent value assessment operations. At this time, the preset driving instruction with the highest confidence score can be selected.

[0033] Then, based on the candidate driving commands, an intelligent agent simulates the motion trajectory of the driving state space information to obtain the driving trajectory information of the vehicle to be assisted. Motion trajectory simulation refers to the process of using an intelligent agent to extrapolate several steps forward from the candidate driving commands as input to obtain the driving state space information of the vehicle to be assisted at consecutive time points. Driving trajectory information refers to the path information obtained after simulating the motion trajectory of the driving state space information using an intelligent agent based on the candidate driving commands, indicating the changes in the driving state space information of the vehicle to be assisted over time in future periods. For example, driving trajectory information can be defined as follows: S

[0034] in, Indicates the current time step. Indicates the termination time step. Indicates a candidate driving instruction. This represents the spatial information of the driving state at the current time step. to Indicates in candidate driving instructions In the future, the agent will randomly or strategically select subsequent driving commands from multiple preset driving commands at each time step. to This indicates the spatial information of the driving status of the vehicle to be assisted in a future time period under the corresponding driving command. Indicates in candidate driving instructions Down, The corresponding action reward quantification value. It is understandable that the candidate driving instructions are determined in each value assessment operation, but the preset driving instructions used in each future time step are randomly selected or selected according to a strategy. Therefore, multiple driving trajectory information can be generated, and the one with the best action reward quantification value is selected as the representative of this value assessment operation.

[0035] Furthermore, based on the initial reward quantification value, the agent performs reward evaluation on the driving trajectory information to obtain the action reward quantification value of the candidate driving command. Here, reward evaluation refers to the process of quantifying the reward value of the simulated driving trajectory information using the agent based on the initial reward quantification value. The action reward quantification value refers to the quantification value obtained after the agent performs reward evaluation on the driving trajectory information based on the initial reward quantification value, used to indicate the quality of the simulated trajectory. Finally, the agent uses the action reward quantification value of the candidate driving command as the initial reward quantification value for the next value evaluation operation.

[0036] In step S103 of some embodiments, the target driving instruction may refer to the preset driving instruction with the largest action reward quantification value selected by comparing the action reward quantification values ​​corresponding to multiple preset driving instructions using an intelligent agent. The intelligent agent assists in controlling the driving actions of the vehicle to be assisted according to the target driving instruction by outputting corresponding acceleration, deceleration, and steering commands in real time through the underlying execution interface of the vehicle to be assisted. For example, if the target driving instruction is " If the signal is "accelerate longitudinally and change lanes to the left," the intelligent agent immediately sends acceleration and left-turn signals to the engine control module and electric steering module of the vehicle to be assisted, assisting the driver in changing lanes and overtaking, and avoiding potential traffic risks in time.

[0037] In some embodiments, the agent evaluates the driving trajectory information based on the initial reward quantification value to obtain the action reward quantification value of the candidate driving command, including: The trajectory utility parameters are obtained by using an intelligent agent to evaluate the trajectory information of the driving trajectory. The first instruction selection count of candidate driving instructions is obtained using an intelligent agent; Based on the number of times the first instruction is selected and the initial reward quantification value, the agent adjusts the trajectory utility parameters to obtain the action reward quantification value of the candidate driving instruction.

[0038] In this embodiment, trajectory utility evaluation can refer to the process of using an intelligent agent to comprehensively evaluate the safety, comfort, traffic efficiency, and traffic rule compliance of driving trajectory information. The trajectory utility parameter can refer to the numerical value obtained after using an intelligent agent to evaluate the trajectory utility of the driving trajectory information, used to quantify the quality of the driving trajectory information. The first instruction selection count can refer to the cumulative number of times the candidate driving instruction has been selected by the intelligent agent and the trajectory simulation has been executed in the current value evaluation operation. For example, if the current value evaluation operation is the first time, and the candidate driving instruction is... ,but The first instruction is selected once; or, if the current valuation operation is the tenth time, and three out of the previous nine operations were selected. ,but The first instruction is selected 3 times. Parameter adjustment refers to the process of using an agent to correct the trajectory utility parameters based on the first instruction selection count and the initial reward quantification value, in order to obtain the corrected trajectory value that integrates exploration gains and experience reliability. The action reward quantification value refers to the quantification value that better reflects the comprehensive expected gains of candidate driving instructions after adjusting the trajectory utility parameters using an agent based on the first instruction selection count and the initial reward quantification value. For example, the action reward quantification value can be defined as follows:

[0039] in, This represents the initial reward quantification value. Represents the trajectory utility parameter. Indicates the number of times the first instruction is selected. This represents the quantified value of the action reward.

[0040] It is understood that the embodiments of this application first utilize an intelligent agent to evaluate the trajectory utility of the driving trajectory information, and then, based on the number of times the first instruction of the candidate driving instruction is selected and the initial reward quantification value, the intelligent agent adjusts the trajectory utility parameters obtained from the evaluation to obtain the action reward quantification value of the candidate driving instruction. In this way, adaptive rewards and penalties can be achieved for driving instructions that have been fully simulated and driving instructions that have been explored less, reducing the value estimation bias caused by insufficient sampling, thereby improving the accuracy of intelligent agent-assisted driving.

[0041] In some embodiments, the agent adjusts the trajectory utility parameters based on the number of times the first instruction is selected and the initial reward quantization value to obtain the action reward quantization value of the candidate driving instruction, including: The agent is used to obtain the initial reward quantification value corresponding to all executed value assessment operations, and the agent is used to determine the maximum initial reward quantification value based on the initial reward quantification value corresponding to all executed value assessment operations. The importance weight of the driving trajectory information is determined by the agent based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and the preset greedy strategy parameters. The trajectory utility parameters are weighted by an intelligent agent based on importance weights to obtain weighted trajectory utility parameters; Based on the number of times the first instruction is selected and the initial reward quantification value, the agent adjusts the weighted trajectory utility parameters to obtain the action reward quantification value of the candidate driving instruction.

[0042] In this embodiment, the maximum initial reward quantification value can refer to the maximum reward quantification value determined among all the initial reward quantification values ​​corresponding to the executed value assessment operations obtained by the agent. The importance weight can refer to a quantification coefficient determined by the agent based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and preset greedy strategy parameters, used to indicate the relative importance of the driving trajectory information relative to the optimal trajectory. The preset greedy strategy parameters can refer to a pre-set coefficient used to adjust the bias of the importance weight between the maximum initial reward quantification value and the current initial reward quantification value. The weighted trajectory utility parameter can refer to an intermediate quantification value obtained by weighting the trajectory utility parameter according to the importance weight using the agent, used to reflect the contribution of the current trajectory relative to the optimal trajectory. The action reward quantification value can refer to a quantification value further reflecting the comprehensive expected return of the candidate driving instructions, obtained by adjusting the weighted trajectory utility parameter according to the number of times the first instruction was selected and the initial reward quantification value using the agent. For example, the action reward quantification value can be defined as follows:

[0043] in, This represents the initial reward quantification value. Represents the trajectory utility parameter. Indicates importance weight, This represents the weighted trajectory utility parameter. Indicates the number of times the first instruction is selected. This represents the quantified value of the action reward.

[0044] Understandably, in this embodiment, the importance weights corresponding to the currently simulated driving trajectory information are first determined by the intelligent agent based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and preset greedy strategy parameters. Then, the trajectory utility parameters are weighted using these importance weights. Finally, the intelligent agent adjusts the weighted trajectory utility parameters using the number of times the first instruction is selected and the initial reward quantification value to obtain the action reward quantification value for the candidate driving instruction. This allows for adaptive weighted assessment, assigning higher weights to high-value trajectories generated in each value assessment operation and suppressing the influence of low-value trajectories. This reduces value estimation bias caused by random errors in single simulations and insufficient early sampling, thereby improving the accuracy of agent-assisted driving.

[0045] In some embodiments, this application can utilize an agent to perform multiple rounds of trajectory simulation starting from candidate driving instructions in each value assessment operation, generating multiple driving trajectory information at once. Then, the trajectory utility parameters of each driving trajectory information are calculated, and weighted summation is performed according to their respective importance weights. The final weighted trajectory utility parameters for each value assessment operation are determined based on the weighted summation results. It is understood that the action reward quantification value at this time can be obtained by adjusting the final weighted trajectory utility parameters using the agent based on the number of times the first instruction is selected and the initial reward quantification value. For example, the action reward quantification value at this time can be defined as follows:

[0046]

[0047] in, Indicates the number of times the trajectory simulation was performed. The value ranges from 1 to , Indicates the first The importance weights corresponding to the driving trajectory information generated by the secondary trajectory simulation. Indicates the first The trajectory utility parameters corresponding to the driving trajectory information generated by the secondary trajectory simulation. This represents the final weighted trajectory utility parameter after n rounds of trajectory simulation in this value assessment operation. It should be noted that if n is 1, meaning only one round of trajectory simulation is performed, the final weighted trajectory utility parameter... That is, the weighted trajectory utility parameter corresponding to this round. .

[0048] To better understand the calculation principle of the final weighted trajectory utility parameter, the following will explain it in detail: First, define .in, Let $\sum$ be the sum of the importance weights corresponding to the first $n$ rounds of trajectory simulation. Then, the final weighted trajectory utility parameter for the $n$ round can also be expressed as: .

[0049] For the (n+1)th simulation, there exists .in, Indicates the first The importance weights corresponding to the driving trajectory information generated by the secondary trajectory simulation.

[0050] but and The following recursive relationship holds between them:

[0051] in, That is, the trajectory utility parameter corresponding to the driving trajectory information generated in the (n+1)th trajectory simulation. This represents the final weighted trajectory utility parameter after n+1 rounds of trajectory simulation in this value assessment operation. This is the summation of the importance weights corresponding to the simulation of the first n+1 trajectories.

[0052] In some embodiments, the importance weight of the driving trajectory information is determined by the agent based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and preset greedy policy parameters, including: Based on the maximum initial reward quantification value, the agent generates a behavior probability distribution of the driving trajectory information to obtain the target policy probability distribution. Based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and the preset greedy strategy parameters, the agent generates the target probability distribution of the driving trajectory information to obtain the behavior strategy probability distribution. The ratio of the probability distribution of the target policy to the probability distribution of the behavior policy is used to determine the importance weight of the driving trajectory information by the intelligent agent.

[0053] In this embodiment, behavior probability distribution generation can refer to the process of mapping driving trajectory information to a corresponding probability distribution using an agent based on the maximum initial reward quantification value. Target policy probability distribution can refer to the probability distribution obtained after generating behavior probability distributions for driving trajectory information using an agent based on the maximum initial reward quantification value, representing the probability that the current trajectory should be adopted under an ideal high-reward policy. Target probability distribution generation involves mapping driving trajectory information to a corresponding probability distribution using an agent based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and preset greedy policy parameters. Behavior policy probability distribution can refer to the probability distribution obtained after generating target probability distributions for driving trajectory information using an agent based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and preset greedy policy parameters, representing the probability that the current trajectory will be adopted under the actual exploration policy. Importance weight can refer to a weight coefficient determined by using an agent to calculate the ratio of the target policy probability distribution to the behavior policy probability distribution, used to quantify the degree of difference between the current driving trajectory information under the ideal policy and the actual exploration policy. For example, importance weight can be defined as follows:

[0054] in, Indicates importance weight, Represents the probability distribution of the target policy. This represents the probability distribution of behavioral strategies.

[0055] It should be noted that the target strategy probability distribution in the embodiments of this application can also be understood as the probability distribution corresponding to the driving trajectory information under the target strategy. For example, the target strategy (labeled as " ) can be represented as:

[0056] in, This represents the maximum initial reward quantification value.

[0057] The probability distribution of the behavior strategy in this application embodiment can also be understood as the probability distribution corresponding to the driving trajectory information under the behavior strategy. For example, the behavior strategy (labeled as " ) can be represented as:

[0058] in, This indicates the preset greedy strategy parameters. This represents the initial reward quantification value. This represents the maximum initial reward quantification value.

[0059] It is understood that this embodiment maps the target policy and behavior policy corresponding to the driving trajectory information to corresponding target policy probability distributions and behavior policy probability distributions, respectively, and uses an agent to determine the ratio of the target policy probability distribution to the behavior policy probability distribution as the importance weight of the driving trajectory information. In this way, the reliability of the trajectory can be directly quantified using the ratio of the two policies, reducing the overhead of complex modeling and quickly highlighting the contribution of high-reward trajectories, thereby improving the efficiency and accuracy of agent-assisted driving.

[0060] In some embodiments, an intelligent agent is used to evaluate the trajectory utility of driving trajectory information to obtain trajectory utility parameters, including: The preset maximum speed of the vehicle to be assisted is obtained using an intelligent agent; Based on the preset maximum vehicle speed, the intelligent agent evaluates the driving efficiency of the driving trajectory information to obtain a quantitative value of driving efficiency. Using intelligent agents to assess driving safety based on driving trajectory information, a quantitative value of driving safety is obtained; The trajectory utility parameters are determined by an intelligent agent based on preset weighting coefficients, driving efficiency quantification values, and driving safety quantification values.

[0061] In this embodiment, the preset limit speed is used to indicate the maximum speed that the vehicle to be assisted is allowed to reach under the current road and traffic conditions. Driving efficiency assessment can refer to the process of comprehensively measuring the average speed covered by the driving trajectory information using an intelligent agent based on the preset limit speed. The driving efficiency quantification value can refer to the numerical value obtained after assessing the driving efficiency of the driving trajectory information using an intelligent agent based on the preset limit speed, used to characterize the speed of trajectory movement and the efficiency of traffic flow. Driving safety assessment can refer to the process of using an intelligent agent to determine the safety level of following distance, collision risk, and driving status corresponding to the driving trajectory information. The driving safety quantification value can refer to the numerical value obtained after assessing the driving safety of the driving trajectory information using an intelligent agent, used to quantify the trajectory safety risk. The preset weighting coefficient can refer to the weighting coefficients preset for the driving efficiency quantification value and the driving safety quantification value, used to reflect the relative importance of the two when comprehensively calculating the trajectory utility parameters; for example, the weighting coefficient corresponding to the driving efficiency quantification value can be set to 0.6, and the weighting coefficient corresponding to the driving safety quantification value can be set to 0.4, and the specific values ​​are not limited. Trajectory utility parameters can refer to a comprehensive value determined by an intelligent agent based on preset weighting coefficients, driving efficiency quantification values, and driving safety quantification values, used to quantify the overall quality of the current driving trajectory. For example, trajectory utility parameters can be defined as follows:

[0062]

[0063]

[0064] in, This represents the quantified value of driving efficiency. Indicates the preset maximum speed. This indicates the average vehicle speed covered by the driving trajectory information. Vehicles requiring assistance are encouraged to travel at the highest possible speed. This represents the quantitative value of driving safety. Indicates the time when a collision is expected to occur. This indicates whether the vehicle to be assisted is in a legal lane; if it is, a value is assigned. Otherwise ; This indicates whether the vehicle to be assisted obeys traffic signals; if it does, a value is assigned. Otherwise ; This indicates whether the vehicle to be assisted has been involved in a collision; if no collision occurred, a value is assigned. Otherwise .constant , This indicates the preset weighting coefficient. This refers to the weighting coefficient corresponding to the quantified value of driving efficiency. This refers to the weighting coefficient corresponding to the quantitative value of driving safety.

[0065] It is understood that the embodiments of this application utilize an intelligent agent to evaluate the driving efficiency of the driving trajectory information by setting a preset maximum speed, thereby obtaining a quantified value of driving efficiency. Then, the intelligent agent is used to evaluate the driving safety of the driving trajectory information, obtaining a quantified value of driving safety. Finally, based on preset weighting coefficients, the quantified values ​​of driving efficiency and driving safety, the intelligent agent determines the trajectory utility parameters. In this way, a balanced quantification of driving trajectory information from both efficiency and safety dimensions can be achieved, reducing misjudgments of trajectory quality caused by a single indicator, thereby improving the accuracy of intelligent agent-assisted driving.

[0066] In some embodiments, the agent determines a preset driving instruction as a candidate driving instruction from a plurality of preset driving instructions based on an initial reward quantification value, including: The intelligent agent is used to obtain the number of times the second instruction is selected for each preset driving instruction, and the intelligent agent is used to obtain the number of evaluation operations to perform the value evaluation operation. Based on the initial reward quantification value, the number of times the second instruction is selected, and the number of evaluation operations, the agent performs a confidence score on each preset driving instruction to obtain the instruction confidence score corresponding to each preset driving instruction. Based on the instruction confidence score corresponding to each preset driving instruction, the agent determines one preset driving instruction as a candidate driving instruction from multiple preset driving instructions.

[0067] In this embodiment, the second instruction selection count can refer to the cumulative number of times each preset driving instruction has been selected and trajectory simulation executed by the agent in the current value assessment sequence. The number of assessment operations can refer to the total number of all value assessment operations completed so far. The confidence score can refer to the process by which the agent calculates the confidence level of the second instruction selection count for each preset driving instruction by combining the initial reward quantification value and the number of assessment operations. The instruction confidence score can refer to the value used to quantify the credibility priority of the corresponding preset driving instruction after the agent performs a confidence score on each preset driving instruction based on the initial reward quantification value, the second instruction selection count, and the number of assessment operations. For example, the instruction confidence score can be defined as follows:

[0068] in, This represents the initial reward quantification value. Indicates the number of times the second instruction is selected. Indicates the number of evaluation operations. This indicates the confidence score of the instruction.

[0069] Candidate driving instructions refer to the preset driving instructions selected by the agent based on the instruction confidence scores of each preset driving instruction, with the highest confidence score being the one corresponding to the maximum confidence score. It should be noted that when there are multiple preset driving instructions corresponding to the highest confidence score, one of them will be randomly selected as the candidate driving instruction.

[0070] It is understood that, in this embodiment, a confidence score is calculated for each preset driving instruction by the number of times the second instruction is selected, the number of evaluation operations performed, and the initial reward quantification value. This yields a confidence score for each preset driving instruction. Then, based on the confidence score, an agent selects one preset driving instruction from among multiple preset driving instructions as a candidate driving instruction. This allows for the performance of value evaluation operations on unexplored preset driving instructions each time, and even after all preset driving instructions have been explored, continued value evaluation operations can be performed on preset driving instructions with higher confidence scores. This reduces computational waste caused by uneven sampling or repeated evaluation of low-potential instructions, thereby improving the efficiency of agent-assisted driving.

[0071] In some embodiments, based on driving state space information, an intelligent agent performs multiple value evaluation operations on multiple preset driving commands to obtain quantified action reward values ​​corresponding to the multiple preset driving commands, including: Based on the driving state spatial information, the intelligent agent performs multiple value assessment operations on multiple preset driving commands, and uses the intelligent agent to obtain the cumulative operation time of all executed value assessment operations. When the cumulative operation time is less than or equal to the preset time threshold, the agent continues to perform value evaluation operations on multiple preset driving instructions based on the driving state space information, so as to obtain the action reward quantification value corresponding to the multiple preset driving instructions. When the cumulative operation time exceeds the preset time threshold, stop performing value assessment operations on multiple preset driving commands.

[0072] In this embodiment, the cumulative operation time can refer to the total actual running time consumed by the agent from the start of the first round of value assessment operations to the end of the last round of value assessment operations. The preset time threshold can refer to a pre-set maximum allowable time limit for limiting the cumulative time consumed by all value assessment operations in a single assisted decision-making process.

[0073] It should be noted that when the cumulative operation time is less than or equal to the preset time threshold, it means that the agent can continue to perform the value assessment operation within the allowed time budget, so as to complete as much simulation and reward update as possible before the deadline, thereby improving the accuracy of the action reward quantification value corresponding to multiple preset driving instructions; when the cumulative operation time is greater than the preset time threshold, it means that the allowed time budget has been exceeded, and continuing to run will delay the decision output. Therefore, the value assessment operation is terminated immediately, and the currently accumulated action reward quantification value is used as the final decision basis to ensure the real-time performance of the assisted driving decision.

[0074] It is understood that, in this embodiment of the application, after performing multiple value assessment operations, the agent obtains the cumulative operation time consumed by all executed value assessment operations, and continues to perform value assessment operations on multiple preset driving commands until the time budget is exhausted, provided that the cumulative operation time does not exceed a preset time threshold. In this way, more rounds of value assessment operations can be completed preferentially within a limited time, and the process can be terminated immediately when the time limit is reached, reducing decision-making delays caused by timeouts and thus improving the real-time performance and accuracy of agent-assisted driving.

[0075] Please see Figure 3 , Figure 3This is a flowchart illustrating a specific implementation of the intelligent agent-assisted driving method provided in this application. Specifically, it includes: first, reading the state, i.e., acquiring the driving state space information of the vehicle to be assisted at the current moment and multiple preset driving commands that can be invoked. Then, it enters the search parameter update stage, where the intelligent agent updates its internal search parameters (i.e., initial reward quantification, command selection count, and cumulative operation time) based on the accumulated number of evaluations, operation time, and action reward quantification value. Finally, it determines whether the conditions for continuing the search are met. If the cumulative operation time does not exceed the preset threshold and there are still preset driving instructions to be evaluated, the value assessment operation continues. This involves the agent selecting candidate driving instructions from multiple preset instructions, simulating the driving state space information based on these instructions, quantifying the trajectory utility parameters corresponding to the simulated trajectory, determining the importance weights of these parameters based on the behavior and target strategies, and calculating the action reward quantification value for each candidate driving instruction based on the importance weights and trajectory utility parameters. The calculated action reward quantification value is then used to update the initial reward quantification value and the number of times the corresponding preset driving instruction is selected. The time consumed in this value assessment operation is added to the cumulative operation time. If the search conditions are not met, the process stops searching. The agent immediately terminates the subsequent simulation, compares the action reward quantification values ​​of all preset driving instructions, and selects the one with the highest score as the final target driving instruction. This target driving instruction is then sent to the vehicle control unit, completing this assisted driving decision. Understandably, compared to the traditional Monte Carlo method, the agent-assisted driving method of this application employs behavioral and target strategies, enabling the simulated trajectory to reach more random events and ensuring the accuracy of agent-assisted driving. Furthermore, by calculating the quantified action reward value for each preset driving instruction through the importance weights determined by the dual strategies, high-value trajectories can receive greater weight, while the contribution of low-value trajectories is suppressed, thereby reducing random errors in value estimation and improving the stability and reliability of decision-making. In addition, the entire calculation process of this application embodiment is logically simple, computationally inexpensive, and requires no high-end computer configuration.

[0076] Please see Figures 4A to 4C , Figures 4A to 4C This is an experimental result diagram of the first application of intelligent agent-assisted driving method provided in the embodiments of this application. Figures 4A to 4C The experiments were conducted on the well-known open-source simulation platform CARLA. CARLA has a benchmark test case, Leaderboard, which includes various challenging scenarios in autonomous driving to test the performance of the intelligent vehicle under complex conditions. It should be noted that all experiments in this application's embodiments were conducted on this platform. Figure 4AThe illustration shows the experimental environment of an embodiment of this application. The vehicle (i.e., the vehicle to be assisted) is traveling straight on a two-lane road, and then an animal appears in front of it and crosses the road from right to left. Figure 4B This displays the spatial information of the friendly vehicle's driving state, including the quantifiable set of its position, velocity, acceleration, and direction of travel at the current time. A quantifiable set of states of the surrounding animals, including their position, velocity, acceleration, and direction of travel. . Figure 4C The comparison between the strategy of this application and existing completely random strategies, greedy strategies, and temperature parameter strategies demonstrates that... Figure 4C As can be seen, when an animal suddenly crosses the lane, after a certain number of iterations, the strategy of this application is significantly superior to other strategies. Furthermore, at the end of the iteration, the strategy of this application can obtain the most accurate trajectory utility parameters with the fewest iterations, enabling the vehicle to be assisted to brake or change lanes in advance to avoid a collision.

[0077] Please see Figures 5A to 5C , Figures 5A to 5C This is an experimental result diagram of the second application of the intelligent agent-assisted driving method provided in the embodiments of this application. Among them, Figure 5A The illustration shows the experimental environment of an embodiment of this application. The vehicle (i.e., the vehicle to be assisted) is traveling straight on a two-lane road, and then other vehicles are coming from the left at the intersection ahead. Figure 5B This displays the spatial information of the friendly vehicle's driving state, including the quantifiable set of its position, velocity, acceleration, and direction of travel at the current time. The set of quantifiable states of the surrounding vehicles, including their position, speed, acceleration, and direction of travel. . Figure 5C The comparison between the strategy of this application and existing completely random strategies, greedy strategies, and temperature parameter strategies demonstrates that... Figure 5C As can be seen, in the conflict scenario where a vehicle suddenly exits from the left at the intersection, the trajectory utility parameter of the proposed strategy rises the fastest and has the most stable trend. Under the same number of iterations, its trajectory utility parameter is significantly higher than that of the completely random, greedy, and temperature-parameter strategies. Furthermore, at the end of the iteration, the proposed strategy achieves the highest trajectory utility parameter with the fewest evaluation rounds, indicating that it can quickly assess conflict risk and make accurate decisions, enabling friendly vehicles to slow down or yield in time, ensuring safe and smooth passage through the intersection.

[0078] Please see Figures 6A to 6C , Figures 6A to 6C This is an experimental result diagram of the third application of the intelligent agent-assisted driving method provided in the embodiments of this application. Among them, Figure 6AThe illustration shows the experimental environment of an embodiment of this application. The vehicle (i.e., the vehicle to be assisted) is traveling straight in a two-lane road. Then, other vehicles are traveling forward in the same lane ahead. There is an obstacle blocking both lanes in front of the other vehicles. Figure 6B This displays the spatial information of the friendly vehicle's driving state, including the quantifiable set of its position, velocity, acceleration, and direction of travel at the current time. The set of quantifiable states of the surrounding vehicles, including their position, speed, acceleration, and direction of travel. . Figure 6C The comparison between the strategy of this application and existing completely random strategies, greedy strategies, and temperature parameter strategies demonstrates that... Figure 6C As can be seen, in scenarios where the vehicle in front in the same lane suddenly brakes due to a two-lane obstacle, the curve of the proposed strategy maintains the fastest rate of ascent from the first iteration, and its trajectory utility parameter at the same number of iterations is significantly higher than that of the completely random, greedy, and temperature-parameter strategies. Furthermore, at the end of the iteration, the proposed strategy still converges to the optimal trajectory with the fewest evaluation rounds, indicating that the proposed strategy can quickly identify obstructions and lane-changing opportunities ahead, make accurate decisions, enable the vehicle to decelerate in time and change lanes smoothly, avoid rear-end collisions and sudden braking, and ensure driving safety and comfort.

[0079] Please see Figures 7A to 7C , Figures 7A to 7C This is an experimental result diagram of the fourth application of the intelligent agent-assisted driving method provided in the embodiments of this application. Among them, Figure 7A The illustration shows the experimental environment of an embodiment of this application. The vehicle (i.e., the vehicle to be assisted) is moving straight forward in the left lane. There are other vehicles in front of it moving forward in the same lane, and there are other vehicles in the right lane moving towards the vehicle in the opposite direction. The vehicle needs to change lanes to overtake the vehicles in front. Figure 7B This displays the spatial information of the friendly vehicle's driving state, including the quantifiable set of its position, velocity, acceleration, and direction of travel at the current time. The set of quantifiable states of the surrounding vehicles, including their position, speed, acceleration, and direction of travel. , . Figure 7C The comparison between the strategy of this application and existing completely random strategies, greedy strategies, and temperature parameter strategies demonstrates that... Figure 7CAs can be seen, in complex oncoming traffic scenarios where the vehicle needs to use the oncoming right lane to overtake, the curve of the proposed strategy rises the fastest and remains stable. Within the same number of iterations, the trajectory utility parameter is significantly higher than that of the completely random, greedy, and temperature-parameter strategies. Furthermore, at the end of the iteration, the proposed strategy converges to the optimal effect with the fewest evaluation rounds, indicating that it can quickly weigh the overtaking opportunity against the oncoming traffic risk, making accurate decisions, enabling the vehicle to accelerate in time and smoothly merge into the right lane, successfully completing the overtaking maneuver, avoiding collisions and sudden braking, and ensuring the safety and smoothness of the lane-changing process.

[0080] Please see Figure 8 This application also provides an intelligent agent-assisted driving device that can implement the above-described intelligent agent-assisted driving method. The device includes: The environmental perception unit 801 is used to acquire assisted driving instructions and to acquire, according to the assisted driving instructions, the driving state space information and driving action space information of the vehicle to be assisted at the current time using an intelligent agent. The driving action space information includes multiple preset driving instructions. The value assessment unit 802 is used to perform multiple value assessment operations on multiple preset driving instructions using an intelligent agent based on driving state space information, so as to obtain the action reward quantification value corresponding to the multiple preset driving instructions. The assisted driving unit 803 is used to determine the target driving command of the vehicle to be assisted by an intelligent agent based on the action reward quantification value corresponding to multiple preset driving commands, and to use the intelligent agent to assist in controlling the driving action of the vehicle to be assisted according to the target driving command. Each valuation operation includes: Based on the initial reward quantification value, the agent determines one preset driving instruction as a candidate driving instruction from multiple preset driving instructions. Based on the candidate driving instructions, the intelligent agent simulates the motion trajectory of the driving state space information to obtain the driving trajectory information of the vehicle to be assisted. Based on the initial reward quantification value, the agent evaluates the driving trajectory information to obtain the action reward quantification value of the candidate driving command. The agent uses the action reward quantification value of the candidate driving command as the initial reward quantification value for the next value evaluation operation.

[0081] The specific implementation of this intelligent agent-assisted driving device is basically the same as the specific implementation of the intelligent agent-assisted driving method described above, and will not be repeated here.

[0082] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described intelligent agent-assisted driving method. This electronic device can be any intelligent terminal, including tablet computers, in-vehicle computers, etc.

[0083] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 901 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the intelligent agent-assisted driving method of the embodiments of this application. The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904); The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0084] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described intelligent agent-assisted driving method.

[0085] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0086] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0087] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0088] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0089] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0090] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0091] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0092] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. An agent-assisted driving method, characterized by, The method includes: Obtain assisted driving instructions, and use an intelligent agent to obtain the driving state space information and driving action space information of the vehicle to be assisted at the current time based on the assisted driving instructions. The driving action space information includes multiple preset driving instructions. Based on the driving state space information, the intelligent agent performs multiple value evaluation operations on the multiple preset driving instructions to obtain the action reward quantification value corresponding to the multiple preset driving instructions; Based on the action reward quantification value corresponding to the multiple preset driving instructions, the intelligent agent determines the target driving instruction of the vehicle to be assisted, and based on the target driving instruction, the intelligent agent assists in controlling the driving action of the vehicle to be assisted. Each of the aforementioned value assessment operations includes: Based on the initial reward quantification value, the agent determines one preset driving instruction as a candidate driving instruction from the plurality of preset driving instructions. Based on the candidate driving instructions, the intelligent agent simulates the motion trajectory of the driving state space information to obtain the driving trajectory information of the vehicle to be assisted. Based on the initial reward quantification value, the agent performs a reward evaluation on the driving trajectory information to obtain the action reward quantification value of the candidate driving instruction; The agent uses the action reward quantification value of the candidate driving instruction as the initial reward quantification value for the next value assessment operation.

2. The method of claim 1, wherein, The step of evaluating the driving trajectory information using the agent based on the initial reward quantification value to obtain the action reward quantification value of the candidate driving instruction includes: The intelligent agent is used to evaluate the trajectory utility of the driving trajectory information to obtain trajectory utility parameters; The intelligent agent is used to obtain the first number of times the candidate driving command is selected. Based on the number of times the first instruction is selected and the initial reward quantification value, the agent adjusts the trajectory utility parameters to obtain the action reward quantification value of the candidate driving instruction.

3. The method of claim 2, wherein, The step of adjusting the trajectory utility parameters using the agent based on the number of times the first instruction is selected and the initial reward quantification value to obtain the action reward quantification value of the candidate driving instruction includes: The agent is used to obtain the initial reward quantification value corresponding to all executed value assessment operations, and the agent is used to determine the maximum initial reward quantification value based on the initial reward quantification value corresponding to all executed value assessment operations. The importance weight of the driving trajectory information is determined by the agent based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and the preset greedy strategy parameters. The trajectory utility parameters are weighted according to the importance weights by the intelligent agent to obtain weighted trajectory utility parameters; Based on the first instruction selection count and the initial reward quantification value, the agent adjusts the weighted trajectory utility parameters to obtain the action reward quantification value of the candidate driving instruction.

4. The method of claim 3, wherein, The step of determining the importance weight of the driving trajectory information using the agent based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and preset greedy strategy parameters includes: Based on the maximum initial reward quantification value, the agent generates a behavior probability distribution for the driving trajectory information to obtain the target policy probability distribution. Based on the maximum initial reward quantification value, the initial reward quantification value corresponding to the currently executed value assessment operation, and the preset greedy strategy parameters, the agent generates a target probability distribution for the driving trajectory information to obtain a behavior strategy probability distribution. The ratio of the probability distribution of the target policy to the probability distribution of the behavior policy is used by the intelligent agent to determine the importance weight of the driving trajectory information.

5. The method of claim 2, wherein, The process of using the intelligent agent to evaluate the trajectory utility of the driving trajectory information and obtain trajectory utility parameters includes: The intelligent agent is used to obtain the preset maximum speed of the vehicle to be assisted; Based on the preset maximum vehicle speed, the intelligent agent evaluates the driving efficiency of the driving trajectory information to obtain a quantitative value of driving efficiency. The intelligent agent is used to assess driving safety based on the driving trajectory information to obtain a quantitative value of driving safety. The trajectory utility parameters are determined by the intelligent agent based on preset weighting coefficients, the quantified value of driving efficiency, and the quantified value of driving safety.

6. The method of claim 1, wherein, The step of using the agent to determine a preset driving instruction as a candidate driving instruction from the plurality of preset driving instructions based on the initial reward quantification value includes: The intelligent agent is used to obtain the number of times the second instruction is selected for each preset driving instruction, and the intelligent agent is used to obtain the number of times the evaluation operation is performed to perform the value evaluation operation. Based on the initial reward quantification value, the number of times the second instruction was selected, and the number of evaluation operations, the agent performs a confidence score on each preset driving instruction to obtain the instruction confidence score corresponding to each preset driving instruction; Based on the instruction confidence score corresponding to each preset driving instruction, the agent determines one preset driving instruction as a candidate driving instruction from the plurality of preset driving instructions.

7. The method of claim 1, wherein, The step of using the intelligent agent to perform multiple value evaluation operations on the multiple preset driving commands based on the driving state space information to obtain the action reward quantification value corresponding to the multiple preset driving commands includes: Based on the driving state space information, the intelligent agent performs multiple value assessment operations on the multiple preset driving commands, and obtains the cumulative operation time of all executed value assessment operations using the intelligent agent. When the cumulative operation time is less than or equal to a preset time threshold, the agent continues to perform value evaluation operations on the multiple preset driving instructions based on the driving state space information, so as to obtain the action reward quantification value corresponding to the multiple preset driving instructions. When the cumulative operation time exceeds a preset time threshold, the value assessment operation for multiple preset driving commands is stopped.

8. An agent-assisted driving apparatus, characterized by comprising: The device includes: An environmental perception unit is used to acquire assisted driving instructions and to acquire, based on the assisted driving instructions, the driving state spatial information and driving action spatial information of the vehicle to be assisted at the current time using the intelligent agent. The driving action spatial information includes multiple preset driving instructions. The value assessment unit is used to perform multiple value assessment operations on the multiple preset driving instructions using the intelligent agent based on the driving state space information, so as to obtain the action reward quantification value corresponding to the multiple preset driving instructions. The assisted driving unit is used to determine the target driving instruction of the vehicle to be assisted based on the action reward quantification value corresponding to the plurality of preset driving instructions using the intelligent agent, and to assist in controlling the driving action of the vehicle to be assisted based on the target driving instruction using the intelligent agent. Each of the aforementioned value assessment operations includes: Based on the initial reward quantification value, the agent determines one preset driving instruction as a candidate driving instruction from the plurality of preset driving instructions. Based on the candidate driving instructions, the intelligent agent simulates the motion trajectory of the driving state space information to obtain the driving trajectory information of the vehicle to be assisted. Based on the initial reward quantification value, the agent performs a reward evaluation on the driving trajectory information to obtain the action reward quantification value of the candidate driving instruction; The agent uses the action reward quantification value of the candidate driving instruction as the initial reward quantification value for the next value assessment operation.

9. An electronic device, comprising: The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the intelligent agent-assisted driving method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. When the computer program is executed by the processor, it implements the intelligent agent-assisted driving method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Vehicle risk avoidance method and system, vehicle, and storage medium

    EP4497644A1

  • Behavior planning for autonomous vehicles

    US20210253128A1