Vehicle control method and device, vehicle and storage medium
By combining empirical memory and large language models for logical reasoning in the vehicle autonomous driving system, the problems of insufficient generalization capabilities of deep learning models and unexplainable decision-making are solved, and high adaptability and transparent decision-making are achieved for complex environments.
Patent Information
- Application Number
- CN202510456334.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
AI Technical Summary
In existing vehicle autonomous driving systems, the generalization ability of deep learning models is limited and the decision-making process lacks interpretability, resulting in poor performance in complex and uncovered scenarios, making it difficult for users to understand the basis for model decision-making.
By obtaining the surrounding environment data of the vehicle, generating driving decision problems, combining the vehicle's experience memory and large language model for reasoning, generating adaptive autonomous driving strategies, using thinking chain technology to decompose complex scenarios for logical reasoning for sub-problems, and introducing reflection mechanisms and common sense memory optimization decisions.
It significantly improves the system's ability to adapt to complex driving environments, the decision-making process is transparent and logical, and enhances user trust and system interpretability.
Smart Images

Figure CN120288069A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of vehicle control, and particularly to a vehicle control method, device, vehicle, and storage medium. Background Art
[0002] The current vehicle's autopilot system generally uses a deep learning model as the core decision-making architecture. Although this technical solution can achieve basic autopilot functions in structured driving scenarios, it has the following defects in practical applications: First, in terms of the model generalization ability, this single-model architecture has serious limitations. Since the model training completely depends on the pre-collected fixed data set, and the real driving environment is highly dynamic and unpredictable, the system performs poorly when facing scenarios not covered by the training data.
[0003] Secondly, the deep learning model calculates the probability distribution as a "black box", and this characteristic leads to the lack of interpretability in the decision-making process. Users cannot understand the specific decision-making basis of the model, thus reducing the credibility of the system. Summary of the Invention
[0004] In view of the above problems, this application provides a vehicle control method, device, vehicle, and storage medium that overcome or at least partially solve the above problems. The technical solutions are as follows: A vehicle control method includes: Obtain the first driving environment data sensed by the first vehicle; Generate a driving decision problem corresponding to the first driving environment data, and reason about the driving decision problem based on the experience memory corresponding to the first vehicle to obtain a first autopilot strategy adapted to the driving decision problem; wherein, the experience memory includes the autopilot strategies obtained through historical reasoning and their corresponding driving environment data; Perform autopilot control on the first vehicle according to the first autopilot strategy.
[0005] Optionally, the method further includes: obtaining the driver's operation of the first vehicle under the second driving environment data; if the second autopilot strategy corresponding to the second driving environment data is already included in the experience memory and the driver's operation does not match the second autopilot strategy, then reflect on the second autopilot strategy based on the driver's operation to generate a third autopilot strategy; associate and add the third autopilot strategy and the second driving environment data to the reflection memory corresponding to the first vehicle; the reasoning about the driving decision problem based on the experience memory corresponding to the first vehicle includes: reasoning about the driving decision problem based on the experience memory and the reflection memory.
[0006] Optionally, the method further includes: if the second autonomous driving strategy corresponding to the second driving environment data is not included in the experience memory, generating a fourth autonomous driving strategy based on the driver's operation, and associating and adding the fourth autonomous driving strategy and the second driving environment data to the experience memory.
[0007] Optionally, the method further includes: obtaining a fifth autonomous driving strategy provided by a second vehicle and its corresponding third driving environment data; if the sixth autonomous driving strategy corresponding to the third driving environment data is already included in the experience memory and the sixth autonomous driving strategy does not match the fifth autonomous driving strategy, reflecting on the sixth autonomous driving strategy based on the fifth autonomous driving strategy to obtain a seventh autonomous driving strategy; associating and adding the seventh autonomous driving strategy and the third driving environment data to the reflection memory corresponding to the first vehicle; the reasoning about the driving decision problem based on the experience memory corresponding to the first vehicle to obtain a first autonomous driving strategy suitable for the driving decision problem includes: reasoning about the driving decision problem based on the experience memory and the reflection memory to obtain a first autonomous driving strategy suitable for the driving decision problem.
[0008] Optionally, the method further includes: if the sixth autonomous driving strategy corresponding to the third driving environment data is not included in the experience memory, associating and adding the fifth autonomous driving strategy and the third driving environment data to the experience memory.
[0009] Optionally, the reasoning about the driving decision problem based on the experience memory corresponding to the first vehicle includes: reasoning about the driving decision problem based on the experience memory corresponding to the first vehicle and the common sense memory; wherein, the common sense memory includes autonomous driving strategies that conform to traffic rules and / or vehicle dynamics constraints.
[0010] Optionally, the generating the driving decision problem corresponding to the first driving environment data and reasoning about the driving decision problem based on the experience memory corresponding to the first vehicle includes: constructing the driving decision problem corresponding to the first driving environment data based on a large language model using the chain of thought technique, and reasoning about the driving decision problem in combination with the experience memory corresponding to the first vehicle.
[0011] A vehicle control device includes: An acquisition module, configured to acquire first driving environment data sensed by a first vehicle; An inference module, configured to generate a driving decision problem corresponding to the first driving environment data, and infer the driving decision problem based on the experience memory of the first vehicle to obtain a first autonomous driving strategy adapted to the driving decision problem; wherein, the experience memory includes the autonomous driving strategies obtained through historical inferences and their corresponding driving environment data; A control module, configured to perform autonomous driving control on the first vehicle according to the first autonomous driving strategy.
[0012] A vehicle, comprising: a processor; and a memory arranged to store computer-executable instructions that, when executed, cause the processor to execute the above torque control method.
[0013] A computer-readable storage medium storing a computer program that, when executed, implements the above torque control method.
[0014] In the embodiment of the present application, by acquiring the driving environment data around the first vehicle and generating corresponding driving decision problems; then calling the association relationship between the historical autonomous driving strategies and driving environment data stored in the experience memory of the first vehicle, using a large language model to infer the driving decision problems, and finally generating an adapted autonomous driving strategy and converting it into a vehicle control instruction for execution. This solution significantly improves the system's adaptability to complex driving environments by dynamically associating historical experience with real-time scenarios. Compared with the classification prediction method of traditional deep learning models (outputting results in the form of probability distributions), this solution not only has a richer applicable scenario, but also directly generates an autonomous driving strategy through inference instead of a probability distribution result. This inference process is transparent and logically clear, facilitating user understanding and trust.
[0015] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically gives the specific embodiments of the present application. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is the first flow schematic diagram of the vehicle control method in the embodiment of the present application.
[0018] Figure 2 Schematic diagram of the application architecture of the vehicle control method according to an embodiment of the present application.
[0019] Figure 3 Schematic diagram of the second process of the vehicle control method according to an embodiment of the present application.
[0020] Figure 4 Schematic diagram of the third process of the vehicle control method according to an embodiment of the present application.
[0021] Figure 5 Schematic diagram of the structure of the vehicle control device according to an embodiment of the present application.
[0022] Figure 6 Schematic diagram of the structure of the vehicle according to an embodiment of the present application. Detailed implementation manners
[0023] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0024] The current vehicle's autonomous driving system generally uses a deep learning model as the core decision-making architecture. Although this technical solution can achieve basic autonomous driving functions in structured driving scenarios, it has the following defects in actual applications: The generalization ability of the deep learning model is relatively limited. Since the training of the deep learning model depends on a pre-collected fixed data set, and the real driving environment is highly dynamic and unpredictable, this causes the system to perform poorly in scenarios not covered by the training data.
[0025] The "black box" characteristic of the deep learning model results in the lack of interpretability of its decision-making process. This characteristic makes it difficult for users to understand the specific decision-making basis of the model, thereby reducing the credibility of the system. For example, from the input sensor data to the output vehicle control instructions, the deep learning model undergoes complex neural network calculations, and the intermediate process is difficult to intuitively explain.
[0026] To solve the above problems, the present application proposes a vehicle control solution, which generates corresponding driving decision problems by obtaining data on the vehicle's surrounding environment. This process combines the autonomous driving strategies stored in the vehicle's historical experience memory and the correlation relationships between the driving environment data, and reasons about the driving decision problems, thereby generating an adaptable autonomous driving strategy and converting it into vehicle control instructions for execution. This solution breaks through the limitations of traditional deep learning models: first, it dynamically combines historical experience and real-time scenario data, effectively improving the adaptability to complex driving environments; second, through the reasoning process, the decision-making basis is clearly traceable, significantly enhancing interpretability and credibility.
[0027] Specifically, the vehicle control solution of the present application includes a vehicle control method, device, vehicle, and storage medium. The following will be introduced in detail in combination with their respective embodiments.
[0028] An embodiment of the present application provides a vehicle control method. Among them, Figure 1 is a schematic flowchart of the vehicle control method, including the following steps: S101, obtain the first driving environment data sensed by the first vehicle.
[0029] In this embodiment, the driving environment data is not limited to road condition information (such as lane lines, traffic signs, road surface quality), obstacle information (such as the position and movement state of vehicles, pedestrians, and static objects), traffic participant behavior information (such as the turn signals and speed changes of surrounding vehicles), and environmental factor (such as weather, lighting conditions) information, etc.
[0030] Specifically, the first vehicle obtains the first driving environment data in the following ways: 1) Visual sensor (camera): The camera collects image and video data for identifying static and dynamic information such as lane lines, traffic signs, and obstacles. The camera can also capture the driving trajectories and behavior characteristics of surrounding vehicles, such as turn signal signals and speed changes; 2) Radar system (millimeter-wave radar, lidar): Millimeter-wave radar and lidar can detect the distance and speed of surrounding objects. Millimeter-wave radar is suitable for long-distance detection, while lidar can generate high-precision three-dimensional point cloud data, providing a detailed representation of the environment, especially performing well under low-light or nighttime conditions.
[0031] 3) Positioning system (GPS global positioning system, IMU inertial measurement unit): The GPS global positioning system provides global positioning services. Combining the IMU inertial measurement unit can calculate the vehicle's position, attitude, and speed information in real time. The IMU can still provide continuous positioning support when the GPS signal is weak or lost, ensuring the precise positioning of the vehicle in complex environments.
[0032] 4) Vehicle communication equipment (V2X vehicle networking communication technology): V2X technology allows vehicles to receive data from other traffic participants and infrastructure, such as traffic signal status, road construction information, etc. This information helps vehicles better understand the surrounding environment and make decisions.
[0033] Through the collaborative work of the above-mentioned multiple sensors and devices, the first vehicle can comprehensively perceive its driving environment and provide high-quality data support for subsequent autonomous driving strategy reasoning.
[0034] S102, generate a driving decision problem corresponding to the first driving environment data, and reason about the driving decision problem based on the experience memory corresponding to the first vehicle to obtain a first autonomous driving strategy suitable for the driving decision problem; among them, the experience memory includes the autonomous driving strategies obtained through historical reasoning and their corresponding driving environment data.
[0035] In this embodiment, a large language model is used as the reasoning engine, combined with the Chain-of-Thought technology, to decompose complex driving scenarios (the first driving environment data) into multiple logically related sub-problems, and optimize the decision-making based on the historical autonomous driving strategies in the experience memory module.
[0036] Among them, the driving decision problem can be presented in the form of a specific scenario description, such as "how to safely change lanes in the front construction area" or "what braking strategy should be taken when a pedestrian suddenly crosses the road". For these problems, the system performs step-by-step reasoning through the large language model, including: evaluating the current lane state, analyzing the feasibility of adjacent lanes, calculating the safe operation distance, and determining the best execution timing, etc. The driving decision problem can be transformed into a description that can be intuitively understood by users through the natural language processing ability of the large language model.
[0037] The experience memory stores the autonomous driving strategies obtained through historical reasoning and their corresponding driving environment data. These data record the autonomous driving decision-making schemes and their execution effects in past similar driving environments in a structured form. In this embodiment, relevant autonomous driving strategies can be queried from the experience memory module through the first driving environment data as the reasoning basis.
[0038] The large language model gradually reasons out the optimal decision through the Chain-of-Thought technology. For example, in the lane-changing decision, the model will sequentially evaluate the current lane state, analyze the feasibility of adjacent lanes, calculate the lane-changing safety distance, and determine the best execution timing. During this process, the large language model will real-time call the first driving environment data and combine the autonomous driving strategies in the experience memory module related to it to reason about the driving decision problem. This reasoning method not only ensures the transparency and interpretability of the decision-making process, but also significantly improves the system's adaptability to complex scenarios through the reuse of experience memory.
[0039] In addition, through the driver behavior feedback mechanism in this embodiment, the reasoning ability of the autonomous driving system is further enhanced, and the specific implementation process is as follows: Continuously monitor the driver's operation behavior of the first vehicle in a specific driving environment (second driving environment data).
[0040] If the second autonomous driving strategy corresponding to the second driving environment data is already included in the experience memory and the driver's operation does not match the second autonomous driving strategy (for example, the vehicle control parameters are inconsistent or there are significant differences), then reflect on the second autonomous driving strategy based on the driver's operation, generate an optimized new strategy (third autonomous driving strategy), and add the third autonomous driving strategy and the second driving environment data to the reflection memory corresponding to the first vehicle in an associated manner to form an autonomous driving strategy set different from the experience memory.
[0041] If the second autonomous driving strategy corresponding to the second driving environment data is not included in the experience memory, then the fourth autonomous driving strategy can be generated based on the driver's operation, and the fourth autonomous driving strategy and the second driving environment data are added to the experience memory in an associated manner.
[0042] In the subsequent decision-making process, in this embodiment, the large language model is used to parallelly retrieve the experience memory and the reflection memory to reason about the driving decision problem. For example, after generating the driving decision problem corresponding to the first driving environment data, reason about the driving decision problem based on the experience memory and the reflection memory. This design breaks through the limitation that traditional deep learning models need to stop training for updating, and realizes the continuous online optimization of the autonomous driving strategy. This mechanism realizes the continuous online update of the experience memory and the reflection memory, enabling the reasoning of the large language model to be iterated without the need to stop and fine-tune the training like a deep learning model.
[0043] In addition, through the cross-vehicle knowledge sharing mechanism in this embodiment, the reasoning process of the autonomous driving strategy is optimized. The specific implementation process is as follows: Real-time receive the fifth autonomous driving strategy provided by the second vehicle and its corresponding third driving environment data. Among them, the fifth autonomous driving strategy can be obtained by the second vehicle through the reasoning or reflection of the large language model, that is, it comes from the experience memory or reflection memory of the second vehicle.
[0044] If the sixth autonomous driving strategy corresponding to the third driving environment data is already included in the experience memory of the first vehicle and the sixth autonomous driving strategy does not match the fifth autonomous driving strategy, then reflect on the sixth autonomous driving strategy based on the fifth autonomous driving strategy, generate the seventh autonomous driving strategy, and add the seventh autonomous driving strategy and the third driving environment data to the reflection memory of the first vehicle in an associated manner.
[0045] If the sixth autonomous driving strategy corresponding to the third driving environment data is not included in the experience memory of the first vehicle, the fifth autonomous driving strategy is associated with the third driving environment data and added to the experience memory.
[0046] Through the above cross-vehicle knowledge sharing mechanism, the first vehicle can integrate the autonomous driving strategies and experience data accumulated by other vehicles (such as the second vehicle) in special driving scenarios (such as extreme weather, complex intersection games, etc., which the vehicle itself has not actually encountered). When receiving the fifth strategy and the corresponding third driving environment data provided by the second vehicle, the ability to handle special scenarios can be improved: 1) For unfamiliar scenarios not covered by the vehicle's own experience memory (such as the sixth autonomous driving strategy for which the third driving environment data is not recorded), it is directly stored as prior knowledge, establishing a cross-scenario strategy mapping relationship to achieve zero-shot learning ability.
[0047] 2) For scenarios with existing conflicting strategies (the sixth and fifth autonomous driving strategies do not match), through the reflection mechanism, the experiences of multiple vehicles are integrated to generate a more robust seventh autonomous driving strategy, effectively solving the decision-making bias caused by the limited scenario coverage of a single vehicle.
[0048] As an exemplary introduction, autonomous driving strategies that comply with traffic regulations may include: speed limit rules (such as "the minimum speed limit on the highway is 60 km / h"); lane-keeping principles (such as "no lane change across solid lines"); priority avoidance rules (such as "pedestrians have the right of way at zebra crossings"), etc.
[0049] Autonomous driving strategies that comply with vehicle dynamics constraints may include: lateral acceleration limit (≤0.8g); braking safety redundancy (the following distance ≥ 1.5 seconds); steering angle rate threshold (≤90° / s), etc.
[0050] In the subsequent decision-making process, in this embodiment, the large language model is used to parallelly retrieve the experience memory, reflection memory, and common sense memory to reason about the driving decision problem. For example, after generating the driving decision problem corresponding to the first driving environment data, the driving decision problem is reasoned based on the experience memory, reflection memory, and common sense memory. It should be understood that introducing the common sense memory in the reasoning process can ensure that the finally generated first autonomous driving strategy complies with traffic rules and does not exceed the vehicle performance limit.
[0051] It should be noted that the autonomous driving strategies adopted from the experience memory, reflection memory, and common sense memory in this embodiment can be presented to the user in natural language through the large language model, for example, displayed on the screen of the in-vehicle system, and these can be understood by the user as the basis for reasoning.
[0052] S103, perform autonomous driving control on the first vehicle according to the first autonomous driving strategy.
[0053] Specifically, in this embodiment, the first autonomous driving strategy is parsed into control instructions for the vehicle and then initiated for execution. As an exemplary introduction, the control instructions include operations such as speed adjustment, steering control, and lighting use, all of which are configured according to the interface specifications of the vehicle's actuators.
[0054] During the execution of the control instructions, real-time data in two dimensions is continuously monitored: one is the vehicle state feedback data, which is used to verify whether the actual driving trajectory conforms to the strategy expectation; the other is the driving environment data perceived by the first vehicle. When a significant change in the driving environment is detected, the system will automatically re-infer a new autonomous driving strategy. In addition, vehicles can also share in real time the autonomous driving strategies they are currently adopting. For this purpose, in this embodiment, by introducing the non-cooperative game theory, an interaction model between vehicles is established. The autonomous driving strategy of each vehicle is regarded as a game participant. Through the Nash equilibrium solution algorithm, on the premise of ensuring the optimal individual interests, the currently adopted autonomous driving strategy is adjusted, so as to maximize the overall traffic efficiency. This embodiment can add a "common sense memory" to the first vehicle to store autonomous driving strategies that comply with traffic regulations and vehicle dynamics constraints.
[0055] To ensure the safety of the control process, before executing each control instruction, the system will call the common sense memory module for compliance verification. For example, in a lane change operation, the system can continuously monitor the lateral acceleration data to ensure that it never exceeds the preset safety threshold of 0.8g.
[0056] In addition, this embodiment adopts a continuous learning mechanism, storing the successfully executed autonomous driving strategy cases in the experience memory module; for the autonomous driving strategy cases that fail to be successfully executed (such as driver takeover operations), the system will compare and reflect with the driver's operations. If a new autonomous driving strategy is obtained after reflection, the new autonomous driving strategy and its corresponding driving environment data will be recorded in the reflection memory.
[0057] In practical applications, this embodiment can also call a large language model to implement the following functional systems: 1) Traffic signal optimization system: By analyzing multi-dimensional data such as real-time traffic flow and weather through a large language model, a spatio-temporal correlation model is established to predict the future traffic situation. The system generates an optimal signal control strategy based on a reinforcement learning framework to achieve regional collaborative optimization, and controls the average waiting time and intersection passing rate above the normal level.
[0058] 2) Intelligent parking guidance system: The large language model analyzes the user's parking needs (such as "need a charging parking space") through semantic understanding technology, combines the real-time sensor data of the parking lot, establishes a demand-resource matching model, recommends the optimal parking plan for the user, and provides dynamic navigation guidance.
[0059] 3) Road safety monitoring system: The large language model uses computer vision technology to parse surveillance videos in real time, constructs an event understanding framework, and accurately identifies traffic accidents and violations. Based on a knowledge graph, the system automatically associates response plans to achieve intelligent early warning and emergency response.
[0060] 4) Driver behavior analysis system: The large language model uses multi-modal fusion technology to comprehensively analyze driving characteristics from multiple dimensions such as OBD data and visual information, establishes a personalized scoring model (60% for safety, 20% for economy, 20% for comfort), and generates an interpretable driving behavior analysis report.
[0061] 5) Intelligent fleet management system: The large language model uses operations research and optimization algorithms to process structured data such as vehicle positions and cargo types, and at the same time analyzes unstructured information such as dispatching instructions, and establishes a multi-objective optimization model (10% reduction in fuel costs, empty load rate ≤ 5%). The system supports natural language interaction for dispatching instruction input and optimization plan explanation.
[0062] 6) Vehicle networking security system: The large language model analyzes network traffic patterns through anomaly detection algorithms, establishes a feature library of attack behaviors to identify potential network attacks and take corresponding defense measures. The system uses distributed key management and dynamic encryption algorithms to prevent information leakage and tampering, and ensure collaborative cooperation and data security between vehicles.
[0063] In summary, the vehicle control method of this embodiment significantly improves the intelligent levels of traffic management, parking services, road safety, driver behavior analysis, and fleet management through multi-scenario applications of the large language model, providing comprehensive technical support for the development of future autonomous driving technologies. Among them, Figure 2 illustrates the application framework of the vehicle control method of this embodiment, and this application framework includes: I. Observation module The observation module uses multi-source sensor fusion technology to collect and process driving environment data in real time, including key information such as lane topology and surrounding vehicle dynamics. Through the large language model, semantic encoding and intention understanding are performed on the original data, and heterogeneous sensor data is uniformly mapped to the driving intention semantic space, realizing deep fusion and standardized representation of cross-modal data, and providing structured environment perception input for the downstream inference module. This data fusion method based on semantic encoding not only improves the accuracy of environment perception, but also ensures the semantic consistency of sensor data from different sources and in different formats, thus significantly enhancing the autonomous driving system's understanding ability of complex traffic scenarios.
[0064] II. Memory module The memory module is set with common sense memory, experience memory, and reflection memory.
[0065] Common sense memory stores autonomous driving strategies that comply with traffic rules and vehicle dynamics constraints. These strategies are saved in structured text form for quick retrieval and application, ensuring that the autonomous driving system can follow basic safety norms in complex traffic environments.
[0066] Experience memory stores past driving environment data and their corresponding autonomous driving strategies, providing a reference for autonomous driving decision-making inference in the current driving environment. This storage method is similar to the experience accumulated by human drivers during the learning process. By retrieving similar scenarios and referring to historical decisions, the system's adaptability and robustness are enhanced.
[0067] Reflective memory stores autonomous driving strategies generated by reflecting on experience memory. These autonomous driving strategies can provide lessons from past possible mistakes for autonomous driving decision-making inference in the current driving environment.
[0068] III. Inference Module As the intelligent center of the autonomous driving system, the inference module constructs a complete cognition-decision closed-loop: First, it converts the real-time environment data collected by the observation module into semantic embedding vectors, deeply correlates them with three types of knowledge in the memory module (common sense memory of traffic rules, experience memory of historical driving scenarios, and reflective memory of optimized strategies), and forms a multi-dimensional decision context. Based on the thought chain technology of large language models, the module adopts a "problem decomposition-step-by-step reasoning" working mechanism: decomposes complex road conditions into sub-problems such as lane selection and speed adjustment, generates traceable intermediate reasoning steps by simulating the human thinking process, and finally outputs autonomous driving decisions that meet safety requirements and driving goals. In addition, the inference module has the ability of dynamic optimization and can adjust and infer new autonomous driving decisions based on real-time driving environment changes.
[0069] Among them, in this embodiment, the large language model can be guided to complete reasoning through structured prompts. The prompts can include the following content: Task instructions: Define decision-making goals and constraints; Scenario description: Describe the current road environment and vehicle state; Historical experience: Autonomous driving strategies provided by experience memory; Optimization goals: Clarify key indicators such as safety and efficiency; Feasible solutions: List pre-screened candidate strategies; Collaborative information: Autonomous driving strategies provided by reflective memory.
[0070] IV. Reinforcement and Reflection Module The reinforcement and reflection module continuously improves the autonomous driving strategies in the experience memory through a dual feedback mechanism.
[0071] On the one hand, it monitors the differences between the driver's operations and the autopilot strategies in the experience memory in real time, analyzes the advantages of human driving, and reflects on the strategies in the experience memory based on this to generate new autopilot decision-making schemes. On the other hand, it receives the autopilot strategies shared by other surrounding intelligent agents (such as the second vehicle) and reflects on them in combination with the autopilot strategies in its own experience memory to generate new autopilot decision-making schemes. In addition, the reinforcement reflection module adopts an incremental learning method, stores the newly generated strategies in the reflection memory module, and retains the original strategies at the same time to ensure that the system does not lose existing experience while absorbing new knowledge.
[0072] Specifically, the reinforcement reflection module adopts an "evaluation-reflection" double-loop mechanism. The evaluator generates multi-dimensional quantitative feedback (such as safety scores, efficiency metrics, etc.) to provide a clear direction for strategy optimization; while the reflector uses the COT thinking chain technology to systematically identify reasoning biases and strategy defects through a three-step method of "decision motivation analysis - process deduction - result verification". The two components work together. The evaluator focuses on "what to improve", while the reflector solves the problems of "why to improve" and "how to improve". Finally, all reflection results are transformed into new autopilot decisions and stored in the reflection memory module.
[0073] V. Communication Module The communication module is mainly responsible for three core functions: First, it dynamically judges the optimal communication timing through an intelligent trigger mechanism (such as in scenarios like lane-changing negotiation and intersection coordination), and precisely controls the content of information interaction; Second, it uses V2X wireless communication technology combined with an edge computing architecture to build an inter-vehicle data channel with low latency (<50ms) and high reliability, supporting multiple communication modes including handshake protocols and status synchronization; Finally, it realizes end-to-end encryption through the national cryptography SM9 encryption algorithm, and cooperates with a blockchain-based anti-tampering verification mechanism to ensure the integrity and authenticity of communication content.
[0074] The following explains the actual application scenarios and introduces the vehicle control method of this embodiment.
[0075] Scenario Description: Refer to Figure 3As shown, after the second vehicle (the leading vehicle) senses the driving environment data of the accident area, it infers the fifth autonomous driving strategy for the driving environment of the accident area based on the large language model, and uses the large language model to send the driving environment data of the accident area and its corresponding fifth autonomous driving strategy (detour instruction) to the first vehicle (this vehicle) behind in the form of natural language description, so that the driver of the first vehicle can intuitively see the accident area discovered by the second vehicle and the inferred autonomous driving strategy. At the same time, the first vehicle reflects on the sixth autonomous driving strategy (going straight) of the accident area stored in its own experience memory according to the fifth autonomous driving strategy (detour instruction), generates the seventh autonomous driving strategy (detour), and associates and stores the driving environment data of the accident area and the seventh autonomous driving strategy in the local reflection memory. In subsequent autonomous driving involving the accident area ahead, the first vehicle uses the thought chain technology of the large language model to construct a driving decision-making problem corresponding to the currently sensed driving environment data of the accident area, and comprehensively infers the driving decision-making problem based on the basic autonomous driving strategy in the common sense memory, the sixth autonomous driving strategy in the experience memory, and the seventh autonomous driving strategy in the reflection memory to obtain the first autonomous driving strategy suitable for the accident area, and finally executes the first autonomous driving strategy to complete a safe detour.
[0076] The specific process implementation includes: 1) Strategy sharing stage The second vehicle sends the driving environment data of the accident area and the fifth autonomous driving strategy formulated for the accident area to the first vehicle through V2X technology.
[0077] After receiving it, the first vehicle converts the driving environment data of the accident area and the fifth autonomous driving strategy into a natural language description text through the large language model: "There is construction on the right lane 200 meters ahead. It is recommended to change lanes to the left and pass at a speed of 30 km / h", and displays the description text through the in-vehicle system for the user's reference.
[0078] 2) Strategy reflection and reasoning stage Reference Figure 4 As shown, after the enhanced reflection module of the first vehicle obtains the fifth autonomous driving strategy provided by the second vehicle and the driving environment data of the accident area, it reflects in combination with the relevant sixth autonomous driving strategy in the experience memory of the first vehicle to obtain the seventh autonomous driving strategy.
[0079] Specifically, the enhanced reflection module starts the evaluator-reflector during the reflection process.
[0080] The evaluator conducts a collision probability assessment in the accident scenario: the collision probability of the sixth autonomous driving strategy (going straight) is 63%, and the collision probability of the fifth autonomous driving strategy (detour) is 0%.
[0081] Based on the collision probability assessment results of the fifth and sixth autonomous driving strategies, the reflector reflects on the original sixth autonomous driving strategy locally, confirms that there is a high probability of collision directly in the accident area, and obtains the seventh autonomous driving strategy for the accident area again (start changing lanes 100 meters in advance + decelerate to 25 km / h), and associates and stores the seventh autonomous driving strategy with the driving environment data of the accident area in the reflection memory of the first vehicle.
[0082] When the first vehicle approaches the accident area (such as entering within a radius of 300 meters of the accident area), the inference module matches the autonomous driving strategies related to common sense memory, experience memory, and reflection memory according to the driving environment data of the accident area to infer the first autonomous driving strategy.
[0083] For example: Common sense memory: The minimum speed limit clause in the accident area; Experience memory: The sixth autonomous driving strategy (going straight); Reflection memory: The seventh autonomous driving strategy (detouring).
[0084] Finally, the first autonomous driving strategy is generated: Start changing lanes at 150 meters (adding a 50-meter buffer compared to the seventh autonomous driving strategy), and maintain a passing speed of 28 km / h.
[0085] Among them, the minimum speed limit clause, the sixth autonomous driving strategy (going straight) 3), and the seventh autonomous driving strategy (detouring) used to infer the first autonomous driving strategy can be displayed on the in-vehicle screen of the first vehicle for the user to refer to as the inference basis.
[0086] Execution phase The first vehicle selects a detour according to the first autonomous driving strategy to avoid the accident area.
[0087] To sum up, the method of this embodiment realizes the inference and reflection of autonomous driving strategies through soft architecture design (refer to Figure 2 ) and the application of large language models. Whether it is the first vehicle (this vehicle) or the second vehicle (other vehicles), deployment can be achieved. At the same time, these vehicles do not operate in isolation, but cooperate to optimize the large language model through federated learning technology. Specifically, after each vehicle trains its own large language model locally, only the trained model parameters (non-sensitive data) are uploaded to the centralized computing platform for aggregation, and then the centralized computing platform distributes the aggregated model parameters back to each vehicle, and then each vehicle iterates the local large language model. This design makes full use of the advantages of federated learning, including reducing communication costs, protecting the privacy of training data, and improving the generalization ability of large language models.
[0088] In addition, corresponding to Figure 1For the method shown above, another embodiment of the present application further provides a vehicle control device. Figure 5 It is a schematic structural diagram of the vehicle control device 500, including: An acquisition module 510, configured to acquire first driving environment data sensed by a first vehicle.
[0089] An inference module 520, configured to generate a driving decision problem corresponding to the first driving environment data, and perform inference on the driving decision problem based on the experience memory corresponding to the first vehicle to obtain a first autonomous driving strategy adapted to the driving decision problem; wherein, the experience memory includes autonomous driving strategies obtained through historical inference and their corresponding driving environment data.
[0090] A control module 530, configured to perform autonomous driving control on the first vehicle according to the first autonomous driving strategy.
[0091] The device in this embodiment acquires the driving environment data around the first vehicle and generates corresponding driving decision problems; subsequently, it calls the association relationship between the historical autonomous driving strategies and the driving environment data stored in the experience memory of the first vehicle, uses a large language model to perform inference on the driving decision problems, and finally generates an adapted autonomous driving strategy and converts it into a vehicle control instruction for execution. This solution significantly improves the system's adaptability to complex driving environments by dynamically associating historical experience with real-time scenarios. Compared with the classification prediction method of traditional deep learning models (outputting results in a probability distribution), this solution not only has a richer applicable scenario, but also directly generates an autonomous driving strategy through inference instead of a probability distribution result. This inference process is transparent and logically clear, facilitating user understanding and trust.
[0092] Optionally, the vehicle control device in this embodiment further includes: A first strategy adjustment module, configured to acquire the driver's operations of the first vehicle under second driving environment data; if the second driving environment data corresponding to the second autonomous driving strategy is already included in the experience memory and the driver's operations do not match the second autonomous driving strategy, then reflect on the second autonomous driving strategy based on the driver's operations to generate a third autonomous driving strategy; and add the third autonomous driving strategy and the second driving environment data in an associated manner to the reflection memory corresponding to the first vehicle. Correspondingly, when the inference module 520 performs inference on the driving decision problem based on the experience memory corresponding to the first vehicle, it includes: performing inference on the driving decision problem based on the experience memory and the reflection memory.
[0093] Optionally, the first policy adjustment module is further configured to: if the second autonomous driving policy corresponding to the second driving environment data is not included in the experience memory, generate a fourth autonomous driving policy based on the driver's operation, and add the fourth autonomous driving policy and the second driving environment data to the experience memory in an associated manner.
[0094] Optionally, the vehicle control device of this embodiment further includes: A second policy adjustment module, configured to obtain a fifth autonomous driving policy provided by a second vehicle and its corresponding third driving environment data; if the sixth autonomous driving policy corresponding to the third driving environment data is already included in the experience memory and the sixth autonomous driving policy does not match the fifth autonomous driving policy, reflect on the sixth autonomous driving policy based on the fifth autonomous driving policy to obtain a seventh autonomous driving policy; add the seventh autonomous driving policy and the third driving environment data to the reflection memory corresponding to the first vehicle in an associated manner. Correspondingly, the inference module 520 performs inference on the driving decision problem based on the experience memory corresponding to the first vehicle, including: performing inference on the driving decision problem based on the experience memory and the reflection memory.
[0095] Optionally, the second policy adjustment module is further configured to: if the sixth autonomous driving policy corresponding to the third driving environment data is not included in the experience memory, add the fifth autonomous driving policy and the third driving environment data to the experience memory in an associated manner.
[0096] Optionally, the inference module 520 performs inference on the driving decision problem based on the experience memory corresponding to the first vehicle, including: performing inference on the driving decision problem based on the experience memory and the common sense memory corresponding to the first vehicle; where the experience memory includes the autonomous driving policies obtained through historical inferences and their corresponding driving environment data. Optionally, the inference module 520 generates a driving decision problem corresponding to the first driving environment data and performs inference on the driving decision problem based on the experience memory corresponding to the first vehicle, including: constructing the driving decision problem corresponding to the first driving environment data based on a large language model using the chain of thought technique, and performing inference on the driving decision problem in combination with the experience memory corresponding to the first vehicle.
[0097] It should be noted that regarding the vehicle control device in the above embodiments, the specific manners in which each model performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0098] In addition, another embodiment of the present application further provides a vehicle. Figure 5It is a schematic structural diagram of the vehicle, including a memory 601 and a processor 602. Among them, an executable program code 6011 is stored in the memory 601, and the processor 602 is used to call and execute the executable program code 5011 to execute the vehicle control method provided in the above embodiment.
[0099] In this embodiment, the vehicle can be divided into functional modules according to the above method example. For example, it can correspond to each functional module, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0100] In the case of dividing each functional module according to each function, the vehicle can include: an acquisition module, an inference module, a control module, a first policy adjustment module, and a second policy adjustment module. It should be noted that all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be repeated here.
[0101] It should be understood that the vehicle provided in this embodiment is used to execute the above vehicle control method, so the same effect as the above implementation method can be achieved. That is, the vehicle in this embodiment obtains the driving environment data around the first vehicle and generates corresponding driving decision-making problems; then calls the association relationship between the historical autonomous driving strategies stored in the first vehicle experience memory and the driving environment data, uses the large language model to reason about the driving decision-making problems, and finally generates an adapted autonomous driving strategy and converts it into a vehicle control instruction for execution. This solution significantly improves the system's adaptability to complex driving environments by dynamically associating historical experience with real-time scenarios. Compared with the classification prediction method of traditional deep learning models (outputting results in the form of probability distributions), this solution not only has a richer applicable scenario, but also directly generates an autonomous driving strategy through reasoning instead of a probability distribution result. This reasoning process is transparent and logically clear, facilitating user understanding and trust.
[0102] In the case of adopting an integrated unit, the vehicle can include a processing module and a storage module. Among them, the processing module can be used to control and manage the actions of the vehicle. The storage module can be used to support the vehicle to execute mutual program codes and data, etc.
[0103] Among them, the processing module can be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc. The storage module can be a memory.
[0104] In addition, another embodiment of the present application further provides a computer-readable storage medium, in which computer program code is stored. When the computer program code runs on a computer, the computer is caused to execute the above-mentioned related method steps to implement a vehicle control method provided in the above embodiment.
[0105] Among them, the beneficial effects of the above embodiment can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0106] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0107] In the embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0108] In the description of the present application, it should be understood that if terms such as "upper", "lower", "front", "rear", "left", and "right" are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the indicated position or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present application.
[0109] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.
[0110] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A vehicle control method, characterized in that, Including: Obtain the first driving environment data sensed by the first vehicle; Generate a driving decision-making problem corresponding to the first driving environment data, and reason about the driving decision-making problem based on the experience memory corresponding to the first vehicle to obtain a first autonomous driving strategy adapted to the driving decision-making problem; wherein, the experience memory includes the autonomous driving strategies obtained by historical reasoning and their corresponding driving environment data; Perform autonomous driving control on the first vehicle according to the first autonomous driving strategy.
2. The method according to claim 1, characterized in that The method further includes: Obtain the driver's operation of the first vehicle under the second driving environment data; If the second autonomous driving strategy corresponding to the second driving environment data is already included in the experience memory and the driver's operation does not match the second autonomous driving strategy, then reflect on the second autonomous driving strategy based on the driver's operation to generate a third autonomous driving strategy; Associate and add the third autonomous driving strategy and the second driving environment data to the reflection memory corresponding to the first vehicle; The reasoning about the driving decision-making problem based on the experience memory corresponding to the first vehicle includes: Reason about the driving decision-making problem based on the experience memory and the reflection memory.
3. The method according to claim 2, wherein The method further includes: If the second autonomous driving strategy corresponding to the second driving environment data is not included in the experience memory, then generate a fourth autonomous driving strategy based on the driver's operation, and associate and add the fourth autonomous driving strategy and the second driving environment data to the experience memory.
4. The method according to claim 1, wherein It further includes: Obtain the fifth autonomous driving strategy provided by the second vehicle and its corresponding third driving environment data; If the sixth autonomous driving strategy corresponding to the third driving environment data is already included in the experience memory and the sixth autonomous driving strategy does not match the fifth autonomous driving strategy, then reflect on the sixth autonomous driving strategy based on the fifth autonomous driving strategy to obtain a seventh autonomous driving strategy; Associate and add the seventh autonomous driving strategy and the third driving environment data to the reflection memory corresponding to the first vehicle; The reasoning about the driving decision-making problem based on the experience memory corresponding to the first vehicle to obtain a first autonomous driving strategy adapted to the driving decision-making problem includes: Reason about the driving decision-making problem based on the experience memory and the reflection memory to obtain a first autonomous driving strategy adapted to the driving decision-making problem.
5. The method according to claim 4, wherein The method further includes: If the sixth autonomous driving strategy corresponding to the third driving environment data is not included in the experience memory, then associate and add the fifth autonomous driving strategy and the third driving environment data to the experience memory.
6. The method according to claim 1, wherein The reasoning about the driving decision-making problem based on the experience memory corresponding to the first vehicle includes: Reason about the driving decision-making problem based on the experience memory corresponding to the first vehicle and the common sense memory; wherein, the common sense memory includes autonomous driving strategies that conform to traffic rules and / or vehicle dynamics constraints.
7. The method according to claim 1, wherein Generating a driving decision-making problem corresponding to the first driving environment data and reasoning about the driving decision-making problem based on the experience memory corresponding to the first vehicle, including: Based on a large language model, using the chain of thought technique to construct a driving decision-making problem corresponding to the first driving environment data, and reasoning about the driving decision-making problem in combination with the experience memory corresponding to the first vehicle.
8. A vehicle control device, characterized in that, Including: An acquisition module for acquiring first driving environment data sensed by a first vehicle; An inference module for generating a driving decision-making problem corresponding to the first driving environment data and reasoning about the driving decision-making problem based on the experience memory corresponding to the first vehicle to obtain a first autonomous driving strategy adapted to the driving decision-making problem; wherein the experience memory includes autonomous driving strategies obtained through historical reasoning and their corresponding driving environment data; A control module for performing autonomous driving control on the first vehicle according to the first autonomous driving strategy.
9. A vehicle, comprising: A processor; And a memory arranged to store computer-executable instructions, wherein the executable instructions, when executed, cause the processor to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Intelligent agent question answering system and method based on engineering project
CN121350319A