A Human-Machine Co-Integration Autonomous Driving Decision-Making Method Based on Fast and Slow Systems

The hybrid DRL-LLM system integrates human driving intentions into automatic driving systems, enhancing adaptability and safety by allowing real-time response to dynamic environments and providing explainable decisions.

CN120066281BActive Publication Date: 2025-07-15TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510536060.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-15
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing autonomous driving technology lacks the ability to integrate human-machine machines in complex traffic environments, making it difficult to take into account the personalized needs and safety of users, the rule-driven system is not flexible enough, and the data-driven system lacks controllability and interpretability.

Method used

The fast system based on deep reinforcement learning is used to combine it with a slow system of large language models. The fast system is responsible for real-time driving decisions and control. The slow system analyzes human instructions and generates high-level strategies. Human-computer integration is achieved through design and extended observation space and reward functions.

Benefits of technology

It realizes flexible response to user instructions while ensuring safe driving, improves the controllability and interpretability of the system, and enhances the adaptability and user experience to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066281B_ABST
    Figure CN120066281B_ABST
Patent Text Reader

Abstract

The present invention discloses a human-machine co-integration autonomous driving decision-making method based on a fast-slow system, aiming to balance safety, flexibility, and controllability in autonomous driving scenarios. The model is jointly composed of a fast system based on deep reinforcement learning and a slow system based on a large language model: among them, the fast system is responsible for real-time driving decision-making and control, and can quickly respond in a short-term and high-frequency dynamic traffic environment; the slow system makes high-level decisions and target lane selections by understanding and parsing human user instructions, combined with environmental perception information, and transmits this information to the fast system for execution. By introducing target lane and human instruction information into the observation space of the fast system and designing corresponding network structures and reward functions, the system can listen to human instructions while ensuring safe driving, realizing "human-machine co-integration" autonomous driving. The present invention has good generality, scalability, and interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of autonomous driving and artificial intelligence, and particularly to a driving decision-making method that collaborates based on a deep reinforcement learning algorithm and a large language model. Background Art

[0002] With the continuous evolution of artificial intelligence and autonomous driving technologies, there are increasingly high requirements for the autonomous decision-making capabilities of vehicles in complex traffic environments. Current autonomous driving decision-making mainly focuses on two major categories of methods: rule-driven expert systems and data-driven expert systems (such as end-to-end systems based on deep learning or deep reinforcement learning). However, both of these methods have their respective limitations in practical applications:

[0003] Rule-driven expert systems have strong controllability but insufficient flexibility: Traditional autonomous driving systems usually rely on a large number of pre-written rules or state machine-based decision-making logics. In scenarios such as urban roads or highways, the system will execute operations according to fixed triggering conditions. Although such methods have high interpretability and controllability, ensuring safe and legal driving of the system in common scenarios, their response to emergencies or long-tail scenarios often lacks sufficient adaptability. When the external environment exceeds the designed scope, it is difficult to make reasonable decisions in real time solely relying on pre-written rules, and the robustness and generality of the system will decrease significantly.

[0004] Data-driven expert systems have high flexibility but lack controllability and interpretability: In recent years, emerging end-to-end autonomous driving decision-making systems based on deep reinforcement learning, with the powerful representation ability of neural networks, can learn relatively excellent driving strategies in changing traffic environments and have high adaptability to environmental changes. However, such methods often have the "black box" problem, making it difficult to visualize or explain their internal decision-making logics. Once there are safety threats or decisions that do not meet the driver's expectations on real roads, there are no clear means to intervene or quickly correct them. In addition, data-driven algorithms usually cannot actively "understand" the high-level preferences or intentions of human occupants, such as subjective needs like passengers wanting to "arrive as soon as possible" or "enjoy the scenery along the way", making it difficult to achieve a "human-machine co-integration" driving mode.

[0005] The increasing prominence of human-machine collaboration and personalized needs: As autonomous driving technology gradually moves from testing to practical applications, people not only focus on the safety and efficiency of the system in a single scenario but also increasingly pay attention to the personalized experience of passengers and the sense of control over the vehicle's driving strategy. Different users may have different needs on the same road section, and pure algorithm-driven or pure rule-driven models often have difficulty well accommodating these subjective needs. Summary of the Invention

[0006] The object of the present invention is to overcome the problems of the lack of integration of human driving goals and intentions, as well as insufficient interpretability and security in the prior art, and provide a human-machine co-integrated autonomous driving decision-making method based on a fast-slow system. The autonomous driving decision-making is jointly constituted by a fast system based on Deep Reinforcement Learning (DRL) and a slow system based on a Large Language Model (LLM). Among them, the fast system is responsible for real-time driving decision-making and control, and can quickly respond in a short-term and high-frequency dynamic traffic environment; the slow system makes high-level decisions and target lane selections by understanding and parsing human user instructions, combining environmental perception information, and transmits this information to the fast system for execution. By introducing target lane and human instruction information into the observation space of the fast system and designing the corresponding network structure and reward function, the system can flexibly follow human instructions while ensuring safe driving, realizing "human-machine co-integration" autonomous driving.

[0007] The object of the present invention is achieved by the following technical solutions:

[0008] A human-machine co-integrated autonomous driving decision-making method based on a fast-slow system, comprising the following steps:

[0009] S1 Data collection and environmental perception, obtaining real-time vehicle and environmental state information:

[0010] Deploy multi-modal sensors on the vehicle, and obtain the following key information through real-time perception of the external environment and the vehicle's own state: ① The vehicle's own state: including position 、 speed and acceleration ; ② The surrounding environmental state: the positions and speeds of adjacent vehicles, lane line information, traffic signals and obstacle positions.

[0011] Specifically, the multi-modal sensors include: cameras, millimeter-wave radars, lidars and GPS / IMU.

[0012] Then fuse and synchronize the recognized information into the vehicle coordinate system or the global coordinate system to obtain the following key elements: vehicle position 、 vehicle speed 、 the relative positions of surrounding vehicles or obstacles and relative speeds . At time , organize the preprocessed information into a state vector , and this vector contains the following elements:

[0013]

[0014] Where is the number of other vehicles within the vehicle's perception range. The state vector after the timing synchronization will be one of the important inputs for the subsequent system to perform high-level parsing and low-level control decisions.

[0015] S2 Slow System (LLM) Parsing and High-Level Instruction Generation:

[0016] The large language model (LLM) serves as the slow system. The information input into the LLM includes: the identity positioning information of the LLM, the real-time vehicle and environmental state information, and the human instruction information.

[0017] The identity positioning information of the said LLM is used to enable the LLM to confirm its own identity positioning;

[0018] For example:

[0019] "You are a large language model. Now please act as a mature driving assistant who can provide accurate and correct advice and guidance to human drivers in complex urban driving scenarios. You will be provided with a detailed description of the driving scenario and the intention indication of the human in the current scenario. You need to fully understand the human intention and give appropriate expected lanes and driving styles in combination with the current scenario."

[0020] The said human instruction information is that the user (driver or passenger) inputs abstract or specific driving instructions through voice or text, such as "I'm in a hurry to go to work", "want to enjoy the scenery along the way", "need to bypass the construction section", etc.; if it is voice input, the voice recognition (ASR) module is used to transcribe the voice into text; if it is text input, it is directly obtained through the in-vehicle human-machine interface.

[0021] The real-time vehicle and environmental state information is obtained from step S1 and needs to be transformed into a standard expression that conforms to natural language rules.

[0022] In the real-time vehicle and environmental state information, first, the LLM can distinguish the surrounding scene states and make classification judgments. Then, the position of the autonomous driving vehicle itself is informed to the LLM, including vehicle position information, vehicle speed information, and acceleration information. Then, based on the road topology information, the environmental vehicles that may conflict with the autonomous driving vehicle and the vehicle closest to it in the surrounding area are obtained, and the position, speed, and acceleration information of the corresponding vehicle are informed to the LLM. If there are no other vehicles around, it is also informed to the LLM.

[0023] The above information is input into the large language model (LLM), and using its natural language understanding and reasoning capabilities, a comprehensive analysis of human intentions and external environmental constraints is carried out; inside the large language model, through semantic vector representation, the abstract instructions (such as "safety first", "change lanes as few as possible") are structurally mapped to generate high-level policy information.

[0024] To constrain the output format of the LLM and require it to enhance decision-making quality through reasoning, the LLM is required to output its decision content in a fixed output format. For example, in the system information, add a requirement for the LLM to output according to the format of "reason - reason - repeat the reason until a decision is obtained. After obtaining the decision, it should be output according to #<desired lane>, #<desired driving style>".

[0025] To effectively guide the fast system (DRL, Deep Reinforcement Learning), the slow system needs to output clear driving strategy elements, including the target lane ( ), and the driving mode ( ). The high-level instructions output by the slow system are as follows:

[0026]

[0027] Among them,

[0028] Target lane ( ): According to semantic parsing and road information, specify the lane that the vehicle should prefer or maintain;

[0029] Driving mode ( ): Such as "FAST", "COMFORT", "ECO";

[0030] If the LLM does not normally return the corresponding driving strategy, the LLM is required to rethink and output the corresponding decision, emphasizing that it should be output according to the format of #<desired lane>, #<desired driving style>".

[0031] The slow system will recalculate and update this high-level policy information regularly (or when detecting user instructions or environmental changes) to ensure that it can continuously meet user needs in dynamic scenarios.

[0032] S3 Fast system (DRL) real-time decision-making and control execution:

[0033] The fast system is built based on the deep reinforcement learning DRL method.

[0034] Combine the high-level instructions output by the slow system with the vehicle / environment state to form the extended observation space of the fast system (DRL) :

[0035]

[0036] Including the comprehensive perception information of the vehicle at time extracted in S1 ), and instruction elements such as the target lane and driving mode given by the slow system at the current moment.

[0037] Finally, the observation information obtained by the vehicle will jointly form an observation space matrix with the high-level instructions. Each row of the matrix represents the information of a vehicle, including the corresponding position, speed, acceleration of the vehicle, and high-level instruction information. In addition, in order to enable the DRL agent to clarify the information of its own vehicle, the first row of the matrix is the information of its own vehicle, and the remaining rows are arranged according to the Euclidean distance between the environmental vehicle and the autonomous driving vehicle curtain. The last two columns of the matrix are high-level instruction information, which are the ordinate value of the expected lane center line and the number corresponding to different driving styles respectively. For the surrounding environmental vehicles, there is no corresponding high-level instruction, and the ordinate value of the vehicle and the default driving style in the current state are directly adopted.

[0038] The present invention uses a deep reinforcement learning (DRL) network to make real-time decisions to obtain the underlying control actions :

[0039]

[0040] where represents the steering wheel angle, represents the acceleration (deceleration) or throttle-brake control amount.

[0041] The training of the deep reinforcement learning (DRL) network uses a policy gradient algorithm. Define the policy and optimize the following expected return function:

[0042]

[0043] where, represents the parameters of the policy network, represents the reward obtained by taking the action in the state ; is the state visit distribution induced by the policy. For a given agent policy , when the agent starts from the initial distribution and runs infinitely many steps according to the discount factor , then its discount-dominated state distribution is defined as:

[0044]

[0045] represents randomly sampling a state from the state distribution induced by the policy; represents the policy distribution in the state Random extraction action ; Indicates the expectation under the above conditions;

[0046] Among them, the reward is designed as:

[0047]

[0048] Among them,

[0049] Related to safe driving, such as giving positive rewards when maintaining a safe distance, having no collisions or violations, and punishing in case of danger;

[0050] Rewards for the degree of compliance with slow system instructions (such as target lane, driving mode); if the current actual lane or speed of the vehicle deviates significantly from the target requirements, negative rewards will be given;

[0051] Related to factors such as driving efficiency and comfort, such as reducing unnecessary lane changes and avoiding frequent acceleration and deceleration. Is a weighting coefficient, which is adjusted according to different scenarios or requirements.

[0052] Through training, the fast system network can gradually learn how to make optimal driving decisions on the premise of ensuring safety and obeying instructions, and output control signals at a high frequency (100Hz) during actual operation.

[0053] In addition, during the execution process, the underlying actions output by the fast system network Are transmitted to the vehicle execution unit (steering, throttle, brake, etc.) to complete the real-time control of the vehicle; if the system detects potential dangers or the output of instructions violating traffic rules, a safety filtering module can be introduced to clip or alarm the actions to ensure the safe driving of the vehicle.

[0054] S4 Human-machine co-integration and dynamic feedback:

[0055] Dynamic parsing of the slow system: When the external environment or user instructions change, the slow system can re-perform semantic parsing and high-level planning; using the current state of the vehicle And new user requirements, generate a new round , such as switching the target lane, adjusting the driving mode, etc.

[0056] Immediate response of the fast system: After the fast system obtains new , it will incorporate it into the extended observation space , the DRL network quickly adjusts the underlying actions. For example, when construction on a certain section causes congestion, if the slow system gives an instruction of "reduce the vehicle speed and try to change lanes to avoid congestion", the fast system will update the lane-changing and speed control actions to enter the appropriate lane and maintain a safe distance in the shortest time.

[0057] Multiple rounds of interaction and user experience improvement: The system can interact with the user multiple times during the whole process. If the user is not satisfied with the selected plan, they can input adjustment instructions again, and the slow system will re-plan the route or speed; the fast system will immediately execute the new high-level instruction.

[0058] Beneficial effects

[0059] Compared with the prior art, the present invention has the following advantages:

[0060] (1) Combining controllability and flexibility: By incorporating user instructions and information such as the target lane into the observation space of DRL, the controllability of the algorithm behavior is achieved; at the same time, the efficient learning and rapid response capabilities of DRL for unknown complex scenarios are retained, overcoming the limitation of the lack of adaptability of the pure rule method.

[0061] (2) Enhancing interpretability and human-machine interaction: The slow system is based on a large language model and can generate explanatory descriptions according to user needs or usage scenarios, such as why a certain lane is selected and how to balance safety and time; users can adjust or inquire about high-level instructions through multiple interactions, thus achieving true human-machine integration.

[0062] (3) Safety and robustness: A safety-first reward function is designed in the training stage; during the execution stage, through a safety filtering mechanism and rule constraints, potential dangerous actions can be corrected in a timely manner; the system can adaptively adjust to long-tail scenarios or emergencies, improving the overall robustness of the system.

[0063] (4) Expandability and applicability: This framework is not only applicable to high-level unmanned driving systems, but also can be applied to driving assistance scenarios that require human-machine collaboration; the large language model part can be replaced or upgraded according to user needs or business scenarios, and the DRL model of the fast system can also be coupled with other AI algorithms, having good scalability. Brief description of the drawings

[0064] Figure 1 It is the processing flow chart of the method of the present invention;

[0065] Figure 2 It is the overall architecture schematic diagram of the method of the present invention;

[0066] Figure 3 It is the schematic diagram of the slow system processing flow of the method of the present invention;

[0067] Figure 4Schematic diagram of the observation space and training process in the fast system of the method of the present invention;

[0068] Figure 5 Schematic diagram of the dual-lane road environment simulated in the embodiment of the present invention;

[0069] Figure 6 Test comparison chart of the effects of the method of the present invention and the comparative method under multiple scenarios. Detailed implementation manners

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] Embodiment

[0072] A human-machine co-integrated autonomous driving decision-making method based on a fast-slow system, the overall process of which is as Figure 1 shown, including the following steps:

[0073] S1: Data collection and environmental perception, obtaining real-time vehicle and environmental state information

[0074] S11: Test environment construction: Before the test starts, establish a virtual simulation environment or a closed test field with various typical road conditions (high-speed multi-lanes, merging areas, oncoming dual-lanes, etc.); configure traffic participants, such as surrounding vehicles, pedestrians, and traffic lights, to simulate a real road environment. Next, take Figure 5 the oncoming dual-lane shown as a specific case for further illustration.

[0075] S12: Vehicle and sensor initialization: The vehicle described in the present invention (hereinafter referred to as the "tested autonomous driving vehicle") is equipped with in-vehicle sensing devices such as cameras, millimeter-wave radars, lidars, GPS / IMU, etc.; the vehicle is also equipped with in-vehicle communication devices and computing units, which can share data with the cloud controller or the edge server; start the vehicle and calibrate the sensors to obtain initial state information (vehicle position, speed, lane information, etc.), and at the same time confirm whether the network communication is normal.

[0076] S2: Slow system (LLM) parsing and high-level instruction generation

[0077] S21: User Instruction Acquisition: The autonomous vehicle under test receives high-level instructions from the driver or passengers, which can be in the form of voice or text, such as "I want to get to the company as soon as possible"; if it is voice input, it is first converted into text that can be parsed by the large language model through the automatic speech recognition (ASR) module; the text mode can be directly input into the in-vehicle human-machine interface.

[0078] S22: Slow System Parsing and Understanding: The user instructions and the current vehicle / environment state are passed to the slow system (large language model, LLM) together, as Figure 3 shown; the slow system generates corresponding high-level decision-making information through semantic analysis and reasoning, combined with pre-set or real-time obtained road information (speed limit, construction information, traffic flow, etc.).

[0079] S23: Data Packing and Visualization: The slow system encapsulates the parsing results into data packets in a specified format, recording the high-level intentions, for example:

[0080]

[0081] S3: Fast System (DRL) Real-time Decision-making and Control Execution

[0082] S31: Observation Space Construction: The high-level decision-making information (such as the target lane) output by the slow system is incorporated into the observation space of the fast system, as Figure 4 shown; in addition, the observation space also includes the vehicle's own state (speed, position, acceleration) and surrounding traffic elements, etc.

[0083] Denote the state vector of the vehicle at time t as , which can be expressed as:

[0084]

[0085] S32: Action Space Definition: The fast system outputs low-level driving control instructions through a deep reinforcement learning network combined with a low-level PID controller:

[0086]

[0087] where represents the acceleration (or deceleration), and represents the steering wheel angle.

[0088] S33: Training and Deployment: The fast system can be trained using DRL algorithms such as Policy Gradient or Value-based. During the training phase, through a large number of simulation interactions (or combined with closed-field tests), the policy parameters are continuously optimized to ensure good safety and compliance with the slow system instructions in complex traffic environments.

[0089] S34: Reward function design and collaborative decision-making: The reward function is a crucial component in deep reinforcement learning (DRL). Through the reward function, the DRL model can learn how to make appropriate decisions to achieve the desired behavior. In the present invention, the design goal of the reward function is to take into account the safety, driving efficiency, and obedience to the slow system instructions of the vehicle.

[0090] Among them, the reward function is composed of: In order to take into account safety, efficiency and obedience to human instructions, the reward function in the present invention is It can be divided into the following parts:

[0091]

[0092] in, Safety reward: If you keep a safe distance from the vehicle in front and avoid collision, you will receive a positive reward; if there is a collision or serious risk, you will receive a negative reward; Efficiency bonus, which is higher if the vehicle travels at a steady speed and meets the requirements of "fast" or "economy" mode; The higher the degree of match between the vehicle and the target lane or speed range issued by the slow system, the greater the reward. It can be set or dynamically adjusted according to actual needs.

[0093] The weight of each part , and It is used to balance the importance of different reward items. The weight coefficient setting can be adjusted according to the actual application scenario to achieve a suitable balance between different goals (such as safety, efficiency, and command compliance).

[0094] Taking the opposite double lane as an example, the specific structure of the reward function is as follows:

[0095] Efficiency bonus items ( ):

[0096] The vehicle's high-speed reward is designed to encourage the vehicle to maintain a higher driving speed, in line with the "fast" driving mode instruction requirements. Based on the vehicle's current speed and the desired speed range, the reward function first calculates the "normalized" value of the current speed:

[0097]

[0098] Among them, forward_speed is the actual forward speed of the vehicle (calculated by the cosine value of the vehicle's velocity vector and the direction of the vehicle's head), and is mapped to the range of [-1, 1];

[0099] The reward_speed_range is the defined ideal speed range, including the corresponding expected minimum speed min_target_speed and the expected maximum speed max_target_speed;

[0100] The calculation method of the lmap function is as follows:

[0101]

[0102] If the vehicle speed is high and within the expected range, the reward value will increase. The specific reward formula is:

[0103]

[0104] Among them, the scaled speed is the calculated "normalized speed value", used to limit the elements in the array to a specified range (NumPy library function). If an element exceeds this range, np.clip will truncate it to the boundary value of this range. Specifically, the function of np.clip can be summarized as:

[0105]

[0106] Among them, x is the input value to be processed; min is the lower bound of the specified value. If the input value is less than this lower bound, it will be truncated to min; max is the upper bound of the specified value. If the input value is greater than this upper bound, it will be truncated to max. In the present invention, np.clip is used to limit the range of the "normalized speed value" (scaled_speed) to keep it between [-1, 1]. The normalized speed value (scaled_speed) is obtained by linearly mapping the actual driving speed of the vehicle, and its function is to convert the vehicle speed value to a unified scale (-1 to 1) for convenient subsequent reward calculation.

[0107] Safety reward item ( )

[0108] The collision reward is used to punish the situation where the vehicle collides and is a guarantee for the safety behavior of autonomous driving. If the vehicle collides, self.vehicle.crashed is 1, and the reward is negative at this time; if there is no collision, the reward is zero.

[0109]

[0110] The weight coefficient of this reward item is very large (-10) to ensure that the vehicle can avoid collisions as much as possible.

[0111] Preference reward item ( )

[0112] The preference reward is used to reward the vehicle for choosing behaviors that match the user's preferences. For example, when the vehicle approaches the desired distance from the target lane, a positive reward is given; if the vehicle deviates from the target lane, a negative reward is given. This reward ensures that the vehicle selects the best path possible according to the user's instructions.

[0113]

[0114] In this way, the system ensures that the vehicle can follow the driving goals set by the user and encourages the vehicle to execute the preference strategy.

[0115] Composite reward calculation:

[0116] Finally, considering the weights of all reward items comprehensively, the system calculates the final reward value. The reward value is the weighted sum of multiple sub-rewards:

[0117]

[0118] where 、 and represent the reward items related to safety, efficiency, and compliance with instructions respectively.

[0119] To ensure the comparability of reward values in different scenarios, we normalize the rewards. The specific method is to map the reward value to the interval [0, 1] to ensure that all reward items are within the same scale range:

[0120]

[0121] Perform a linear mapping to map the range of the reward value from the configured minimum value to the maximum value to the range [0, 1], so as to ensure that the reward value output by the system works under a unified scale. S32: Collaboration between the fast system and the slow system: The slow system continuously outputs high-level information based on the user's instructions and traffic situation, and can instantaneously update the target lane or driving mode; the fast system makes rapid inferences based on the new observation state within each decision cycle (such as 0.1 s or shorter) and outputs decision-making information.

[0122] S4: Human-machine integration and dynamic feedback

[0123] S41: Adversarial or conflict situations: If the user changes the instructions (such as switching from "drive fast" to "safety first"), the slow system re-plans the high-level information and updates the target lane and driving mode; during this human-machine interaction process, the system can dynamically display the reasons or expected effects of the decisions made by the vehicle.

[0124] S42: Multi-round Interaction and User Feedback: During the operation of the vehicle, if it is found that the user issues an instruction conflicting with the road traffic regulations, the slow system can detect it in time and prompt rejection or suggest modification through text or interface in a timely manner to avoid the occurrence of dangerous behaviors; the user can also provide new instructions to the vehicle during driving.

[0125] The entire process of the present invention is carried out in a cyclic iteration during actual operation:

[0126] 1) In step S1, the slow system continuously monitors user instructions and environmental changes;

[0127] 2) In steps S2 and S3, the fast system makes decisions based on the updated observation information and reward function;

[0128] 3) S4 dynamically adjusts the system strategy and reward distribution at the human-machine interaction level;

[0129] 4) Continuously repeat this process until the vehicle journey ends or the user instruction is completed.

[0130] The simulation environment used for training is built with Highway-Env and Gymnasium, where the vehicle position and orientation are controlled through a closed-loop PID:

[0131]

[0132]

[0133] where is the relative lateral distance of the vehicle with respect to the center line of the corresponding target lane, is the lateral speed control instruction, is the control instruction for controlling the steering angle of the vehicle.

[0134]

[0135] where is the lane heading, is the target orientation of the heading and position of the desired lane, is the lateral control rate instruction, is the control amount of the steering angle of the front wheel, and are the control gains of the position and heading angle respectively.

[0136]

[0137] where the motion control of the vehicle is implemented according to the above formula, where is the position of the vehicle, is the forward speed of the vehicle, is the acceleration instruction of the vehicle, is the slip angle at the center of gravity. The longitudinal motion of the surrounding vehicles adopts the IDM algorithm, and the lateral control adopts the MOBIL lane-changing strategy.

[0138] Multiple scenarios are selected for testing, including highway car-following scenarios, ramp lane-changing scenarios, and passing-on-the-right scenarios. The performance of the SOTA algorithm (Dilu) based on LLM, the value-based DRL algorithm DQN, and the policy-based DRL algorithm PPO are tested and compared. The results are as Figure 6 shown. The proposed model achieves the highest success rate in various scenarios.

[0139] The aggregated results are analyzed, and the data results are shown in the following table. It is found that the proposed model can achieve the best balance among safety, efficiency, and compliance with human guidance in multiple scenarios:

[0140]

[0141] The above description is only for the description of the preferred embodiments of the present application and is not a limitation on the scope of the present application. Any modification or variation made by any person skilled in the art based on the disclosed technical content shall be regarded as an equivalent effective embodiment and fall within the scope of the technical solutions of the present application.

Claims

1. A human-machine co-integrated autonomous driving decision-making method based on a fast-slow system, characterized in that, It includes the following steps: S1 Data collection and environment perception to obtain real-time vehicle and environment status information; S2 Slow system parsing and high-level instruction generation; S3 Fast system real-time decision-making and control execution; S4 Human-machine co-integration and dynamic feedback; Specifically, step S2 is as follows: The large language model LLM serves as the slow system. The information input into the LLM includes: the identity positioning information of the LLM, real-time vehicle and environment status information, and human instruction information; The identity positioning information of the LLM is used to let the LLM confirm its own identity positioning; The human instruction information is that the user inputs abstract or specific driving instructions through voice or text; The real-time vehicle and environment status information is obtained from step S1 and needs to be transformed into a standard expression that conforms to natural language rules; The above information is input into the large language model. Using its natural language understanding and reasoning capabilities, it comprehensively analyzes human intentions and external environment constraints; inside the large language model, through semantic vector representation, it structurally maps abstract instructions to generate high-level policy information; To effectively guide the fast system DRL, the slow system needs to output clear driving strategy elements, including the target lane and driving mode , and the high-level instructions output by the slow system are as follows: Among them, Target lane : The lane that a vehicle should preferentially select or maintain according to semantic parsing and road information; Driving mode : including "fast", "comfortable" and "energy-saving"; The slow system recalculates and updates this high-level policy information regularly or when detecting user instructions or environmental changes to ensure that it can continuously meet user needs in dynamic scenarios; In step S3, the fast system is constructed based on the deep reinforcement learning DRL method, specifically as follows: Combine the high-level instructions output by the slow system with the vehicle / environment state to form an extended observation space for the fast system : Including the comprehensive perception information of the vehicle at the moment and the target lane and driving mode instruction elements given by the slow system at the current moment; Use a deep reinforcement learning network to make real-time decisions to obtain underlying control actions : wherein represents the steering angle of the steering wheel, represents the acceleration (deceleration) or the control amount of the accelerator and brake; The training of the deep reinforcement learning network uses policy gradient algorithms to define the policy and optimize the following expected return function: where, represents the parameters of the policy network, denotes the reward obtained by taking action in state ; is the state visitation distribution induced by the policy. For a given agent policy , when the agent starts from the initial distribution and runs for an infinite number of steps according to the discount factor , its discounted dominant state distribution is defined as: Denote the state distribution induced by the policy Randomly sample a state ; Denote sampling an action according to the policy distribution in the state Randomly sample an action ; Denote the expectation under the above conditions; Among them, the reward is designed as: Among them, Related to safe driving, such as giving positive rewards when maintaining a safe following distance, without collision or violation, and imposing penalties in case of danger; Reward for compliance with slow system instructions; if the current actual lane or speed of the vehicle deviates significantly from the target requirements, a negative reward is given; It is related to factors such as driving efficiency and comfort, such as reducing unnecessary lane changes and avoiding frequent acceleration and deceleration; is a weighting coefficient, which is adjusted according to different scenarios or requirements; Through training, the fast system network can gradually learn how to make optimal driving decisions while ensuring safety and obeying instructions, and output control signals at a high frequency during actual operation; In addition, during the execution process, the underlying actions output by the fast system network are transmitted to the vehicle execution unit to complete the real-time control of the vehicle. If the system detects the output of instructions for potential danger or traffic rule violations, a safety filtering module can be introduced to clip or alarm the actions to ensure the safe driving of the vehicle.

2. The method for making an autonomous driving decision for human-machine coexistence based on a fast-slow system according to claim 1, wherein Specifically, step S1 is as follows: Deploy multi-modal sensors on the vehicle to obtain the following key information through real-time perception of the external environment and the vehicle's own state: ① Vehicle's own state: including position , speed and acceleration ; ② Surrounding environment state: positions and speeds of adjacent vehicles, lane line information, traffic signals and obstacle positions; Then, fuse and synchronize the recognized information to the vehicle coordinate system or the global coordinate system to obtain the following elements: vehicle position , vehicle speed , relative positions of surrounding vehicles or obstacles and relative speeds ; At time , organize the preprocessed information into a state vector , which contains the following elements: wherein is the number of other vehicles within the vehicle's sensing range.

3. The human-machine co-integrated autonomous driving decision-making method based on a fast-slow system according to claim 1, wherein The human instruction information is that the user inputs abstract or specific driving instructions through voice or text: if it is voice input, the voice recognition module is used to transcribe the voice into text; if it is text input, it is directly obtained through the in-vehicle human-machine interface.

4. A human-machine co-integrated automatic driving decision-making method based on a fast-slow system according to claim 1, characterized in that, The observation information obtained by the vehicle and the high-level instructions together form an observation space matrix. Each row of the matrix represents the information of a vehicle, including the corresponding position, speed, acceleration of the vehicle, and high-level instruction information; in addition, in order to enable the DRL agent to clarify its own vehicle information, the first row of the matrix is the information of its own vehicle, and the remaining rows are arranged according to the Euclidean distance between the environmental vehicle and the autonomous driving vehicle; the last two columns of the matrix are high-level instruction information, which are the ordinate value of the expected lane center line and the number corresponding to different driving styles respectively; for surrounding environmental vehicles, there is no corresponding high-level instruction, and the ordinate value of the vehicle in the current state and the default driving style are directly adopted.

5. The human-machine co-integrated autonomous driving decision-making method based on a fast-slow system according to claim 1, wherein In step S4, human-machine co-integration and dynamic feedback include: Dynamic parsing of the slow system: When the external environment or user instructions change, the slow system re-performs semantic parsing and high-level planning; using the current vehicle state and the new user requirements, generate a new round of high-level instructions ; Immediate response of the fast system: After the fast system obtains new , it incorporates it into the extended observation space , and quickly adjusts the underlying actions through the DRL network; Multi-round interaction and user experience improvement: The system can interact with the user multiple times throughout the process; if the user is not satisfied with the selected solution, they can input adjustment instructions again, and the slow system replans the route or vehicle speed; the fast system immediately executes the new high-level instructions.

Citation Information

Patent Citations

  • High-consistency man-machine hybrid decision-making method based on hybrid enhanced intelligence

    CN115564029A

  • Automatic driving automobile external man-machine interaction system and method based on vehicle-road cloud cooperation

    CN116353585A