Man-machine co-fusion automatic driving decision-making method based on fast and slow system
By combining deep reinforcement learning and large language models, the fast and slow systems used for autonomous driving systems work together, solving the shortcomings of existing systems in complex environments and human needs processing, and achieving safe, controllable and interpretable autonomous driving decisions.
Patent Information
- Application Number
- CN202510536060.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Existing autonomous driving decision-making systems lack interpretability and safety when dealing with complex traffic environments and human driving goals, making it difficult to achieve human-machine integration and personalized needs.
The fast system based on deep reinforcement learning and the slow system working together with the large language model. The fast system is responsible for real-time driving decisions. The slow system generates high-level decisions by understanding human instructions and environmental information. The two combine to achieve autonomous driving with human-machine integration.
It realizes the flexibility of obeying human instructions while ensuring safe driving, enhances the controllability and interpretability of the system, and can better take into account human personalized needs.
Smart Images

Figure CN120066281A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of autonomous driving and artificial intelligence, and particularly to a driving decision-making method based on the collaborative work of a deep reinforcement learning algorithm and a large language model. Background Art
[0002] With the continuous evolution of artificial intelligence and autonomous driving technologies, people have put forward increasingly high requirements for the autonomous decision-making ability of vehicles in complex traffic environments. The current autonomous driving decisions mainly focus on two major types of methods: rule-driven expert systems and data-driven expert systems (such as end-to-end systems based on deep learning or deep reinforcement learning, etc.). However, both of these two types of methods have their own limitations in practical applications: Rule-driven expert systems have strong controllability but insufficient flexibility: Traditional autonomous driving systems usually rely on a large number of pre-written rules or state machine-based decision-making logics. In the scenarios of urban roads or highways, the system will perform operations according to fixed trigger conditions. Although such methods have high interpretability and controllability, and can ensure the safe and legal driving of the system in common scenarios, their response to unexpected events or long-tail scenarios often lacks sufficient adaptability. When the external environment exceeds the design scope, it is difficult to make reasonable decisions in real time simply relying on pre-written rules, and the robustness and generality of the system will significantly decline.
[0003] Data-driven expert systems have high flexibility but lack controllability and interpretability: In recent years, the emerging end-to-end autonomous driving decision-making systems based on deep reinforcement learning can learn relatively excellent driving strategies in changing traffic environments with the help of the powerful representation ability of neural networks, and have high adaptability to environmental changes. However, such methods often have the "black box" problem, and it is difficult to visualize or explain their internal decision-making logics. Once there are safety threats or decisions that do not meet the driver's expectations on real roads, there is no clear means to intervene or quickly correct them. In addition, data-driven algorithms usually cannot actively "understand" the high-level preferences or intentions of human occupants, such as subjective needs like passengers hoping to "arrive as soon as possible" or "enjoy the scenery along the way", and it is difficult to achieve a "human-machine co-integration" driving mode.
[0004] The increasing prominence of human-machine collaboration and personalized needs: As autonomous driving technology gradually moves from testing to practical applications, people not only focus on the safety and efficiency of the system in a single scenario, but also increasingly pay attention to the personalized experience of passengers and the sense of control over the vehicle's driving strategy. Different users may have different needs on the same road section, and pure algorithm-driven or pure rule-driven modes often have difficulty in well considering these subjective needs. Summary of the Invention
[0005] The object of the present invention is to overcome the problems of the lack of integration of human driving goals and intentions, as well as insufficient interpretability and safety in the prior art, and provide a human-machine co-integrated autonomous driving decision-making method based on a fast-slow system. The autonomous driving decision-making is jointly composed of a fast system based on Deep Reinforcement Learning (DRL) and a slow system based on a Large Language Model (LLM): among them, the fast system is responsible for real-time driving decision-making and control, and can quickly respond in a short-term and high-frequency dynamic traffic environment; the slow system makes high-level decisions and target lane selections by understanding and parsing human user instructions, combined with environmental perception information, and transmits this information to the fast system for execution. By introducing target lane and human instruction information into the observation space of the fast system and designing the corresponding network structure and reward function, the system can flexibly follow human instructions while ensuring safe driving, realizing "human-machine co-integrated" autonomous driving.
[0006] The object of the present invention is achieved by the following technical solutions: A human-machine co-integrated autonomous driving decision-making method based on a fast-slow system, comprising the following steps: S1 Data collection and environmental perception, obtaining real-time vehicle and environmental state information: Deploy multi-modal sensors on the vehicle, and obtain the following key information through real-time perception of the external environment and the vehicle's own state: ① The vehicle's own state: including position , speed and acceleration ; ② The surrounding environment state: the positions and speeds of adjacent vehicles, lane line information, traffic signals and obstacle positions.
[0007] Specifically, the multi-modal sensors include: cameras, millimeter-wave radars, lidars and GPS / IMU.
[0008] Then fuse and synchronize the recognized information to the vehicle coordinate system or the global coordinate system to obtain the following key elements: vehicle position , vehicle speed , the relative positions of surrounding vehicles or obstacles and relative speeds . At time , organize the preprocessed information into a state vector , and this vector contains the following elements: where is the number of other vehicles within the vehicle's perception range. This time-series synchronized state vector It will serve as one of the important inputs for high-level parsing and low-level control decision-making in subsequent systems.
[0009] S2 Slow System (LLM) Parsing and High-Level Instruction Generation: As a slow system, the information input into the large language model (LLM) includes: the identity positioning information of the LLM, real-time vehicle and environmental status information, and human instruction information.
[0010] The identity positioning information of the LLM is used to let the LLM confirm its own identity positioning; For example: "You are a large language model. Now please act as a mature driving assistant who can provide accurate and correct advice and guidance to human drivers in complex urban driving scenarios. You will receive a detailed description of the current driving scenario and the intention indication of humans. You need to fully understand the human intention and give appropriate expected lanes and driving styles based on the current scenario."
[0011] The human instruction information is that the user (driver or passenger) inputs abstract or specific driving instructions through voice or text, such as "I'm in a hurry to go to work", "want to enjoy the scenery along the way", "need to bypass the construction section", etc.; if it is voice input, the voice recognition (ASR) module is used to transcribe the voice into text; if it is text input, it is directly obtained through the in-vehicle human-machine interaction interface.
[0012] The real-time vehicle and environmental status information is obtained in step S1 and needs to be transformed into a standard expression that conforms to natural language rules.
[0013] In the real-time vehicle and environmental status information, first, the LLM can identify the surrounding scene status and make a classification judgment. Then, the position of the autonomous driving vehicle itself is informed to the LLM, including vehicle position information, vehicle speed information, and acceleration information. Then, based on the road topology information, the environmental vehicles that may conflict with the autonomous driving vehicle and the vehicle closest to it around are obtained, and the position, speed, and acceleration information of the corresponding vehicle are informed to the LLM. If there are no other vehicles around, the LLM is also informed.
[0014] The above information is input into the large language model (LLM). Using its natural language understanding and reasoning capabilities, a comprehensive analysis of human intentions and external environmental constraints is carried out; inside the large language model, through semantic vector representation, abstract instructions (such as "safety first", "change lanes as few times as possible") are structurally mapped to generate high-level policy information.
[0015] To constrain the output format of the LLM and require it to enhance the decision-making quality through reasoning, the LLM is required to output its decision content in a fixed output format. For example, in the system information, the LLM is required to output in the format of "reason - reason - repeat the reason until a decision is obtained. After obtaining the decision, it should be in the format of #<expected lane>, #<expected driving style>".
[0016] To effectively guide the fast system (DRL, Deep Reinforcement Learning), the slow system needs to output clear driving strategy elements, including the target lane ( ), and the driving mode ( ). The high-level instructions output by the slow system are as follows: Among them, Target lane ( ): According to semantic parsing and road information, specify the lane that the vehicle should preferentially choose or maintain; Driving mode ( ): Such as "Fast (FAST)", "Comfort (COMFORT)", "Energy-saving (ECO)"; If the LLM does not normally return the corresponding driving strategy, the LLM is required to rethink and output the corresponding decision, emphasizing that it should be in the format of #<expected lane>, #<expected driving style>".
[0017] The slow system will recalculate and update this high-level policy information regularly (or when detecting user instructions or environmental changes) to ensure that it can continuously meet user needs in dynamic scenarios.
[0018] S3 Fast system (DRL) real-time decision-making and control execution: The fast system is built based on the deep reinforcement learning DRL method.
[0019] Combine the high-level instructions output by the slow system with the vehicle / environment state to form the extended observation space of the fast system (DRL) : This includes the comprehensive perception information of the vehicle at time (extracted in S1 ), as well as instruction elements such as the target lane and driving mode given by the slow system at the current moment.
[0020] The observed information obtained by the final vehicle will jointly form an observation space matrix with the high-level instructions. Each row of the matrix represents the information of a vehicle, including the corresponding position, speed, acceleration of the vehicle, and high-level instruction information. In addition, to enable the DRL agent to identify its own vehicle information, the first row of the matrix is the information of its own vehicle, and the remaining rows are arranged according to the Euclidean distance between the surrounding vehicles and the autonomous driving vehicle curtain. The last two columns of the matrix are the high-level instruction information, which are the ordinate value of the expected lane centerline and the number corresponding to different driving styles. For the surrounding environmental vehicles, there is no corresponding high-level instruction, and the ordinate value of the vehicle in the current state and the default driving style are directly adopted.
[0021] The present invention uses a deep reinforcement learning (DRL) network to make real-time decisions to obtain the underlying control actions : where represents the steering wheel angle, represents the acceleration (deceleration) or throttle-brake control amount.
[0022] The training of the deep reinforcement learning (DRL) network adopts a policy gradient algorithm, defining the policy and optimizing the following expected return function: where, represents the parameters of the policy network, represents the reward obtained by taking the action in the state ; is the state visitation distribution induced by the policy. For a given agent policy , when the agent starts from the initial distribution and runs infinitely many steps according to the discount factor , then its discount-dominated state distribution is defined as: represents randomly sampling a state from the state distribution induced by the policy ; represents randomly sampling an action in the state according to the policy distribution ; represents the expectation under the above conditions; where the reward is designed as: where, Related to safe driving, such as giving positive rewards when maintaining a safe distance, having no collisions or violations, and punishing when danger occurs; Rewards for compliance with slow system instructions (such as target lane, driving mode); if the vehicle's current actual lane or speed deviates significantly from the target requirements, negative rewards will be given; Related to factors such as driving efficiency and comfort, such as reducing unnecessary lane changes and avoiding frequent acceleration and deceleration. is a weighting coefficient, which is adjusted according to different scenarios or requirements.
[0023] Through training, the fast system network can gradually learn how to make optimal driving decisions while ensuring safety and compliance with instructions, and output control signals at a high frequency (100Hz) during actual operation.
[0024] In addition, during the execution process, the underlying actions output by the fast system network are transmitted to the vehicle execution unit (steering, throttle, brakes, etc.) to complete real-time control of the vehicle; if the system detects potential dangers or the output of instructions that violate traffic rules, a safety filtering module can be introduced to clip or alarm the actions to ensure the safety of vehicle driving.
[0025] S4 Human-Machine Integration and Dynamic Feedback: Dynamic parsing of the slow system: When the external environment or user instructions change, the slow system can re-perform semantic parsing and high-level planning; using the current state of the vehicle and new user requirements, generate a new round , such as switching the target lane, adjusting the driving mode, etc.
[0026] Immediate response of the fast system: After the fast system obtains new , it will incorporate it into the extended observation space , and quickly adjust the underlying actions through the DRL network; for example, when construction causes congestion on a certain section of the road, if the slow system gives an instruction of "reduce speed and try to change lanes to avoid congestion", the fast system will update the lane change and speed control actions to enter the appropriate lane and maintain a safe distance in the shortest time.
[0027] Multiple rounds of interaction and improvement of user experience: The system can interact with the user multiple times during the whole process. If the user is not satisfied with the selected plan, they can input adjustment instructions again, and the slow system will re-plan the route or speed; the fast system will immediately execute the new high-level instructions.
[0028] Beneficial Effects Compared with the prior art, the present invention has the following advantages: (1)Both controllability and flexibility: By incorporating user instructions and information such as the target lane into the observation space of the DRL, the controllability of the algorithm's behavior is achieved; at the same time, the DRL's efficient learning and rapid response capabilities for unknown complex scenarios are retained, overcoming the limitation of the lack of adaptability of the pure rule method.
[0029] (2)Enhanced interpretability and human-machine interaction: The slow system is based on a large language model and can generate explanatory descriptions according to user needs or usage scenarios, such as why a certain lane is selected and how to balance safety and time; users can adjust or query high-level instructions through multiple interactions, thus achieving true human-machine integration.
[0030] (3)Safety and robustness: A safety-first reward function is designed in the training phase; during the execution phase, potential dangerous actions can be corrected in a timely manner through a safety filtering mechanism and rule constraints; adaptive adjustments can be made to long-tail scenarios or emergencies, improving the overall robustness of the system.
[0031] (4)Scalability and applicability: This framework is not only applicable to high-level unmanned driving systems, but also can be applied to driving assistance scenarios that require human-machine collaboration; the large language model part can be replaced or upgraded according to user needs or business scenarios, and the DRL model of the fast system can also be coupled with other AI algorithms, having good scalability. Description of the Drawings
[0032] Figure 1 It is a processing flow chart of the method of the present invention; Figure 2 It is a schematic diagram of the overall architecture of the method of the present invention; Figure 3 It is a schematic diagram of the slow system processing flow of the method of the present invention; Figure 4 It is a schematic diagram of the observation space and training process in the fast system of the method of the present invention; Figure 5 It is a schematic diagram of the simulated two-lane road environment in the embodiment of the present invention; Figure 6 It is a comparison chart of the effect tests of the method of the present invention and the comparative method under various scenarios. Detailed Embodiments
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] Embodiment A human-machine co-integrated autonomous driving decision-making method based on a fast-slow system, and its overall process is as Figure 1 shown, including the following steps: S1: Data collection and environment perception, obtaining real-time vehicle and environment status information S11: Test environment setup: Before the test starts, establish a virtual simulation environment or a closed test field with various typical road conditions (such as multi-lane highways, merging areas, oncoming two-lane roads, etc.); configure traffic participating elements, such as surrounding vehicles, pedestrians, and traffic lights, to simulate a real road environment. Below, take Figure 5 the oncoming two-lane road shown as a specific case for further illustration.
[0035] S12: Vehicle and sensor initialization: The vehicle described in the present invention (hereinafter referred to as the "tested autonomous driving vehicle") is equipped with on-vehicle sensing devices such as cameras, millimeter-wave radars, lidars, GPS / IMUs, etc.; the vehicle is also equipped with on-vehicle communication devices and computing units, which can share data with a cloud controller or an edge server; start the vehicle and calibrate the sensors to obtain initial state information (such as vehicle position, speed, lane information, etc.), and at the same time confirm whether the network communication is normal.
[0036] S2: Slow system (LLM) parsing and high-level instruction generation S21: User instruction acquisition: The tested autonomous driving vehicle receives high-level instructions from the driver or passenger, which can be in the form of voice or text, such as "I want to get to the company as soon as possible", etc.; if it is a voice input, it is first converted into text that can be parsed by the large language model through an automatic speech recognition (ASR) module; the text mode can be directly input into the on-vehicle human-machine interface.
[0037] S22: Slow system parsing and understanding: Transmit the user instruction and the current vehicle / environment status to the slow system (large language model, LLM) together, as Figure 3 shown; the slow system generates corresponding high-level decision-making information through semantic analysis and reasoning, combined with pre-set or real-time obtained road information (such as speed limits, construction information, traffic flow, etc.).
[0038] S23: Data packaging and visualization: The slow system encapsulates the parsing result into a data packet in a specified format, recording the high-level intention, for example: S3: Fast system (DRL) real-time decision-making and control execution S31: Observation space construction: The high-level decision-making information (such as the target lane) output by the slow system is incorporated into the observation space of the fast system, as Figure 4 shown; in addition, the observation space also includes the vehicle's own state (speed, position, acceleration) and surrounding traffic elements, etc.
[0039] The state vector of the vehicle at time t is , which can be expressed as: S32: Action space definition: The fast system outputs the underlying driving control instructions through the deep reinforcement learning network and combined with the underlying PID controller: in represents acceleration (or deceleration), Indicates the steering wheel angle.
[0040] S33: Training and deployment: The fast system can be trained using DRL algorithms such as Policy Gradient or Value-based. During the training phase, through a large number of simulation interactions (or combined with closed field tests), the policy parameters are continuously optimized to ensure good safety and obedience to the slow system's instructions in complex traffic environments.
[0041] S34: Reward function design and collaborative decision-making: The reward function is a crucial component in deep reinforcement learning (DRL). Through the reward function, the DRL model can learn how to make appropriate decisions to achieve the desired behavior. In the present invention, the design goal of the reward function is to take into account the safety, driving efficiency, and obedience to the slow system instructions of the vehicle.
[0042] Among them, the reward function is composed of: In order to take into account safety, efficiency and obedience to human instructions, the reward function in the present invention is It can be divided into the following parts: in, Safety reward: If you keep a safe distance from the vehicle in front and avoid collision, you will receive a positive reward; if there is a collision or serious risk, you will receive a negative reward; Efficiency bonus, which is higher if the vehicle travels at a steady speed and meets the requirements of "fast" or "economy" mode; The higher the degree of match between the vehicle and the target lane or speed range issued by the slow system, the greater the reward. It can be set or dynamically adjusted according to actual needs.
[0043] The weight of each part , and It is used to balance the importance of different reward items. The weight coefficient setting can be adjusted according to the actual application scenario to achieve a suitable balance between different goals (such as safety, efficiency, and command compliance).
[0044] Taking a two-way two-lane as an example, the specific structure of its reward function is as follows: Efficiency reward term ( ): The high-speed reward for vehicles aims to encourage vehicles to maintain a relatively high driving speed, meeting the requirements of the "fast" driving mode instruction. According to the current speed of the vehicle and the desired speed range, the reward function first calculates the "normalized" value of the current vehicle speed: where forward_speed is the actual forward speed of the vehicle (calculated by the cosine value of the vehicle's speed vector and the vehicle head direction), and is mapped to the range of [-1, 1]; reward_speed_range is the defined ideal speed range, including the corresponding desired minimum speed min_target_speed and desired maximum speed max_target_speed; The calculation method of the lmap function is as follows: If the vehicle speed is relatively high and within the desired range, the reward value will increase. The specific reward formula is: where scaled speed is the calculated "normalized speed value", used to limit the elements in the array to a specified range (NumPy library function). If an element exceeds this range, np.clip will truncate it to the boundary value of this range. Specifically, the function of np.clip can be summarized as: where x is the input value to be processed; min is the lower bound of the specified value, if the input value is less than this lower bound, it will be truncated to min; max is the upper bound of the specified value, if the input value is greater than this upper bound, it will be truncated to max. In the present invention, np.clip is used to limit the range of the "normalized speed value" (scaled_speed) to keep it between [-1, 1]. The normalized speed value (scaled_speed) is obtained by linearly mapping the actual driving speed of the vehicle, and its function is to convert the vehicle speed value to a unified scale (-1 to 1) for convenient subsequent reward calculation.
[0045] Safety reward term ( ): The collision reward is used to penalize the situation where the vehicle collides, which is a guarantee for the safety behavior of autonomous driving. If the vehicle collides, self.vehicle.crashed is 1, and the reward is negative at this time; if there is no collision, the reward is zero.
[0046] The weight coefficient of this reward item is very large (-10) to ensure that the vehicle can avoid collisions as much as possible.
[0047] Preference reward item ( ) The preference reward is used to reward the vehicle for choosing behaviors that conform to the user's preferences. For example, when approaching the desired distance from the target lane, a positive reward is given; if the vehicle deviates from the target lane, a negative reward is given. This reward ensures that the vehicle selects the best path as much as possible according to the user's instructions.
[0048] In this way, the system ensures that the vehicle can follow the driving goals set by the user and encourages the vehicle to execute the preference strategy.
[0049] Comprehensive reward calculation: Finally, considering the weights of all reward items, the system calculates the final reward value. The reward value is the weighted sum of multiple sub-rewards: Among them, , and represent the reward items related to safety, efficiency, and instruction compliance respectively.
[0050] To ensure the comparability of the reward values in different scenarios, we normalize the rewards. The specific method is to map the reward values to the interval [0, 1] to ensure that all reward items are within the same scale range: Perform a linear mapping to map the range of the reward values from the configured minimum value to the maximum value to the range [0, 1], so as to ensure that the reward values output by the system work under a unified scale. S32: Collaboration between the fast system and the slow system: The slow system continuously outputs high-level information according to the user's instructions and traffic situation, and can instantly update the target lane or driving mode; the fast system makes rapid inferences based on the new observation states within each decision cycle (such as 0.1s or shorter) and outputs decision-making information.
[0051] S4: Human-machine co-integration and dynamic feedback S41: Confrontation or conflict situation: If the user changes the instruction (such as switching from "drive fast" to "safety first"), the slow system re-plans the high-level information, updates the target lane and driving mode; during this human-machine interaction process, the system can dynamically display the reasons or expected effects of the decisions made by the vehicle.
[0052] S42: Multi-round interaction and user feedback: During the operation of the vehicle, if it is found that the user issues an instruction that conflicts with the road traffic regulations, the slow system can detect it in time and prompt rejection or suggest modification through text or interface in time to avoid the occurrence of dangerous behaviors; the user can also provide new instructions to the vehicle during driving.
[0053] The entire process of the present invention is carried out in a cyclic iteration during actual operation: 1) In step S1, the slow system continuously monitors the user's instructions and environmental changes; 2) In steps S2 and S3, the fast system makes decisions based on the updated observation information and reward function; 3) S4 dynamically adjusts the system strategy and reward distribution at the human-machine interaction level; 4) Continuously repeat this process until the vehicle journey ends or the user's instruction is completed.
[0054] The simulation environment used for training is built with Highway-Env and Gymnasium, where the vehicle position and orientation are controlled by a closed-loop PID: where is the relative lateral distance of the vehicle with respect to the center line of the corresponding target lane, is the lateral speed control instruction, is the control instruction for controlling the steering angle of the vehicle.
[0055] where is the lane heading, is the target orientation of the heading and position of the desired lane, is the lateral control rate instruction, is the control amount of the steering angle of the front wheel, and are the control gains of the position and heading angle respectively.
[0056] where the motion control of the vehicle is implemented according to the above formula, where is the position of the vehicle, is the forward speed of the vehicle, is the acceleration command of the vehicle, is the slip angle at the center of gravity. The longitudinal direction of the surrounding vehicle adopts the IDM algorithm, and the lateral control adopts the MOBIL lane-changing strategy.
[0057] Multiple scenarios are selected, including high-speed following scenarios, ramp lane-changing scenarios, and passing-by overtaking scenarios for testing. The SOTA algorithm (Dilu) based on LLM, the value-based DRL algorithm DQN, and the policy-based DRL algorithm PPO are selected for performance test comparison. The results are as Figure 6 shown. The proposed model achieves the highest success rate in multiple scenarios.
[0058] Aggregate result analysis is carried out. The data results are shown in the following table. It is found that the proposed model can achieve the best balance among safety, efficiency, and compliance with human guidance in multiple scenarios: The above description is only a description of the preferred embodiments of the present application, and is not any limitation on the scope of the present application. Any change or modification made by any person skilled in the art based on the technical content disclosed above shall be regarded as an equivalent effective embodiment, and all fall within the scope of protection of the technical solution of the present application.
Claims
1. A human-machine collaborative automatic driving decision-making method based on a fast and slow system, characterized in that: The following steps are involved: S1 data collection and environmental perception, obtaining real-time vehicle and environmental status information; S2 slow system parsing and high-level instruction generation; S3 fast system real-time decision making and control execution; S4 Human-machine integration and dynamic feedback; Step S2 specifically includes: The large language model LLM is a slow system. The information input to LLM includes: LLM's identity positioning information, real-time vehicle and environment status information, and human command information; The LLM's identity location information is used to allow the LLM to confirm its own identity location; The human instruction information is an abstract or specific driving instruction input by the user through voice or text; The real-time vehicle and environment status information is obtained in step S1 and needs to be converted into a standard expression that complies with natural language rules; The above information is input into the large language model, which uses its natural language understanding and reasoning capabilities to conduct a comprehensive analysis of human intentions and external environmental constraints. The large language model uses semantic vector representation to perform structured mapping of abstract instructions and generate high-level strategy information. To effectively guide the fast system DRL, the slow system needs to output clear driving strategy elements, including the target lane and driving mode , high-level instructions output by the slow system as follows: in, Target lane : According to semantic analysis and road information, specify the lane that the vehicle should prioritize or maintain; Driving Mode : Including "fast", "comfortable" and "energy saving"; The slow system recalculates and updates the high-level policy information periodically or when it detects user instructions or environmental changes to ensure that user needs can continue to be met in dynamic scenarios.
2. According to claim 1, a human-machine collaborative automatic driving decision-making method based on a fast and slow system is characterized in that: Step S1 specifically comprises: Deploy multimodal sensors on the vehicle to obtain the following key information through real-time perception of the external environment and the vehicle's own status: ① Vehicle's own status: including location ,speed and acceleration ; ②Surrounding environment status: adjacent vehicle position and speed, lane line information, traffic signals and obstacle positions; The identified information is then integrated and synchronized to the vehicle coordinate system or the global coordinate system to obtain the following elements: vehicle position , vehicle speed , the relative positions of surrounding vehicles or obstacles and relative speed ; at the time , organize the preprocessed information into a state vector , which contains the following elements: in is the number of other vehicles within the vehicle's perception range.
3. According to claim 1, a human-machine collaborative automatic driving decision-making method based on a fast and slow system is characterized in that: The human command information is abstract or specific driving instructions input by the user through voice or text: if it is voice input, the voice recognition module is used to transcribe the voice into text; if it is text input, it is directly obtained through the on-board human-computer interaction interface.
4. According to claim 1, a human-machine collaborative automatic driving decision-making method based on a fast and slow system is characterized in that: In step S3, the fast system is constructed based on the deep reinforcement learning (DRL) method, as follows: High-level instructions that output slow systems and vehicle / environment status Combined to form an extended observation space for fast systems : in Including vehicles at time The comprehensive perception information and the target lane and driving mode instruction elements given by the slow system at the current moment; Using deep reinforcement learning network Make real-time decisions to get low-level control actions : in Indicates the steering wheel angle, Indicates the acceleration (deceleration) or throttle and brake control amount; The training of deep reinforcement learning network adopts policy gradient algorithm to define the strategy And optimize the following expected reward function: in, represents the parameters of the policy network, Indicates in status Take action Rewards received; is the policy-induced state visit distribution. For a given agent strategy , when the agent is from the initial distribution Departure, according to the discount factor If the operation is infinite, the discounted state distribution is defined as: represents the state distribution induced by the policy Randomly select a state ; Indicates distribution by strategy in state Randomly extract actions ; Express expectations under the above conditions; The rewards are designed as follows: in, Related to safe driving, such as giving positive rewards for maintaining a safe distance, avoiding collisions or violations, and penalties for dangerous situations; Rewards for slow system compliance; negative rewards if the vehicle's current lane or speed differs significantly from the target requirement; Related to factors such as driving efficiency and comfort, such as reducing unnecessary lane changes and avoiding frequent acceleration and deceleration; It is a weighted coefficient, which is adjusted according to different scenarios or needs; Through training, the fast system network can gradually learn how to make optimal driving decisions while ensuring safety and obeying instructions, and output control signals at a high frequency during actual operation; In addition, during execution, the underlying actions of the fast system network output are The information is transmitted to the vehicle execution unit to complete real-time control of the vehicle. If the system detects a potential danger or a command output that violates traffic rules, a safety filtering module can be introduced to cut the action or issue an alarm to ensure vehicle driving safety.
5. According to claim 4, a human-machine collaborative automatic driving decision-making method based on a fast and slow system is characterized in that: The observation information obtained by the vehicle and the high-level instructions together constitute the observation space matrix. Each row of the matrix represents the information of a vehicle, including the corresponding position, speed, acceleration and high-level instruction information of the vehicle. In addition, in order to enable the DRL intelligent agent to clearly understand the information of its own vehicle, the first row of the matrix is the information of its own vehicle, and the remaining rows are arranged according to the Euclidean distance between the environment vehicle and the autonomous driving vehicle. The last two columns of the matrix are high-level instruction information, which are the ordinate value of the center line of the desired lane and the corresponding number of different driving styles. For the surrounding environment vehicles, there is no corresponding high-level instruction, and the ordinate value of the vehicle in the current state and the default driving style are directly used.
6. According to claim 1, a human-machine collaborative automatic driving decision-making method based on a fast and slow system is characterized in that: In step S4, human-machine integration and dynamic feedback include: Dynamic analysis of the slow system: When the external environment or user instructions change, the slow system re-performs semantic analysis and high-level planning; using the current state of the vehicle and new user needs, generating a new round of high-level instructions ; Fast system instant response: Fast system in obtaining new Then, it is included in the extended observation space , quickly adjust the underlying actions through the DRL network; Multiple rounds of interaction and improved user experience: The system can interact with the user multiple times during the entire process; if the user is not satisfied with the selected solution, he or she can input adjustment instructions again, and the slow system will re-plan the route or speed; the fast system will immediately execute the new high-level instructions.
Citation Information
Patent Citations
High-consistency man-machine hybrid decision-making method based on hybrid enhanced intelligence
CN115564029A
Automatic driving automobile external man-machine interaction system and method based on vehicle-road cloud cooperation
CN116353585A
Human feedback-based interactive adaptive decision control method for autonomous vehicle
CN117719535A
Cited By
Highway ramp traffic control method based on UAV-LLM-FCS
CN120894915A
Self-adaptive automatic driving decision-making method fusing user preferences and traffic rules
CN120942370A
An adaptive autonomous driving decision-making method fusing user preferences and traffic rules
CN120942370B
Intelligent driving cooperative control terminal based on visual language action model
CN121133732A
Intelligent driving cooperative control terminal based on visual language action model
CN121133732B