Automatic driving planning method based on visual language large model and related equipment

By employing a multi-stage reasoning method based on a large visual language model, and utilizing sample driving scenario information and fine-tuning of the thought chain, the problem of visual illusion in complex scenarios of autonomous driving systems is solved, thereby improving safety and output reliability and achieving better understanding and planning of the driving environment.

CN119514716BActive Publication Date: 2025-11-18INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411318441.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-11-18
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

Existing autonomous driving planning methods suffer from visual illusions when dealing with complex and dynamically changing driving scenarios, leading to incorrect judgments and affecting driving safety.

Method used

A multi-stage reasoning method based on a large visual language model is adopted. By fine-tuning sample driving scenario information and thought chain, autonomous driving planning parameters suitable for the current driving scenario are obtained, which enhances the system's semantic understanding and contextual reasoning ability and alleviates the visual illusion problem.

Benefits of technology

It improves the safety and output reliability of autonomous driving systems, enabling a more comprehensive and in-depth understanding of complex and ever-changing driving environments, generating planning parameters that are more suitable for the current scenario, and enhancing system safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119514716B_ABST
    Figure CN119514716B_ABST
Patent Text Reader

Abstract

The application discloses an automatic driving planning method based on a visual language large model and related equipment, and relates to the technical field of artificial intelligence. The method comprises the following steps: acquiring driving scene information; inputting the driving scene information into a visual language large model to acquire automatic driving planning parameters output by the visual language large model, wherein the visual language large model is obtained by sample driving scene information and a thought chain, and the thought chain is used for assisting the visual language large model in performing multi-stage inference analysis on the sample driving scene information. The technical scheme provided by the application improves the reliability of model output through multi-stage inference, and enhances the safety of an automatic driving system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an autonomous driving planning method and related equipment based on a large visual language model. Background Technology

[0002] With the rapid development of autonomous driving technology, how to achieve efficient and safe autonomous driving planning has become a hot research topic. In complex and ever-changing road environments, autonomous driving systems need to quickly understand the driving scenario and make appropriate driving decisions, which places high demands on autonomous driving planning algorithms.

[0003] Currently, most autonomous driving planning methods rely primarily on manual rules and traditional machine learning algorithms. These methods typically use predefined rules or models to analyze driving scenarios and then generate corresponding driving plans based on the analysis results.

[0004] However, these methods still have significant limitations when dealing with complex and dynamically changing driving scenarios. In particular, the introduction of large visual language models can lead to visual illusions—that is, misinterpreting elements or situations that do not actually exist when analyzing driving scenarios. These visual illusions can cause autonomous driving systems to make incorrect judgments, seriously affecting driving safety. Summary of the Invention

[0005] This invention provides an autonomous driving planning method and related equipment based on a large visual language model. Through multi-stage reasoning, the reliability of the model output is improved, thereby enhancing the safety of the autonomous driving system.

[0006] In a first aspect of the present invention, an autonomous driving planning method based on a large visual language model is provided, comprising:

[0007] Obtain driving scenario information;

[0008] The driving scenario information is input into the visual language big model to obtain the autonomous driving planning parameters output by the visual language big model. The visual language big model is obtained by fine-tuning the sample driving scenario information and the thought chain. The thought chain is used to assist the visual language big model in performing multi-stage reasoning analysis on the sample driving scenario information.

[0009] In a second aspect of the invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the autonomous driving planning method based on a large visual language model as described above.

[0010] In a third aspect of the invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the autonomous driving planning method based on a large visual language model as described above.

[0011] In a fourth aspect of the invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the autonomous driving planning method based on a large visual language model as described above.

[0012] In summary, one or more technical solutions provided in this invention have at least the following technical effects or advantages:

[0013] By inputting driving scenario information into a visual language model fine-tuned using thought chains, autonomous driving planning parameters suitable for the current driving scenario can be obtained. This method has significant advantages over traditional methods that rely on predefined rules or models. First, the visual language model fine-tuned using thought chains can perform multi-stage reasoning analysis of the driving scenario, enabling the model to more comprehensively and deeply understand the complex and ever-changing driving environment, thereby generating planning parameters more suitable for the current scenario. Second, the introduction of the visual language model gives the system stronger semantic understanding and contextual reasoning capabilities, enabling it to better handle dynamically changing driving scenarios and overcome the shortcomings of traditional methods in complex scenarios. Furthermore, the introduction of thought chains helps alleviate the visual illusion problem that may arise in the visual language model, and improves the reliability of the model output through multi-stage reasoning, thereby enhancing the safety of the autonomous driving system. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0015] Figure 1 This is a flowchart illustrating an autonomous driving planning method based on a large visual language model provided in an embodiment of the present invention.

[0016] Figure 2 This is a bird's-eye view diagram provided in an embodiment of the present invention.

[0017] Figure 3 This is a schematic diagram of a lane map provided in an embodiment of the present invention.

[0018] Figure 4 This is a schematic diagram of the fine-tuning process of a large visual language model provided in an embodiment of the present invention.

[0019] Figure 5 This is a schematic diagram illustrating the fine-tuning process of a large visual language model provided in an embodiment of the present invention.

[0020] Figure 6 This is a simulation diagram provided in an embodiment of the present invention.

[0021] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0023] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an autonomous driving planning method based on a large visual language model, provided by an embodiment of the present invention. This method can be implemented using a computer program, a microcontroller, or run on an autonomous driving planning system based on a large visual language model using the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. Specifically, the method may include the following steps:

[0024] S101. Obtain driving scenario information.

[0025] Driving scenario information refers to a comprehensive and structured description of the environment in which the autonomous vehicle is currently located, including but not limited to road layout, the location and status of traffic participants, traffic rules, and other related information. In this embodiment, it can be understood as a highly abstract and structured representation of the scene obtained through the collection and processing of data from multiple sensors, mainly consisting of two parts: a bird's-eye view of the scene and a corresponding text description.

[0026] Furthermore, driving scene information provides comprehensive and easily processed input to the visual language large model regarding the current driving environment. Complex 3D scene information can be transformed into 2D plane and text forms, significantly reducing data complexity and facilitating processing and understanding by the visual language large model. Secondly, by combining image and text modalities, it can more comprehensively capture all aspects of the scene, including both intuitive spatial relationships and semantic information that is difficult to visualize. Thirdly, this representation retains key information in the scene while filtering out unnecessary details, allowing the model to focus on elements crucial for safe driving. Finally, this standardized scene representation helps improve the model's generalization ability, enabling it to better cope with various complex driving environments.

[0027] Based on the above embodiments, as an optional embodiment, in S101, the step of obtaining driving scenario information may further include the following steps:

[0028] S201, Obtain the driving scenario.

[0029] The driving scenario refers to the overall environment in which the autonomous vehicle operates, including external factors that may influence driving decisions, such as road structure, traffic participants, traffic facilities, and weather conditions. In this embodiment, it can be understood as the set of raw data collected in real time by various sensor systems on the autonomous vehicle, which together constitute a comprehensive description of the current driving environment. Specifically, the driving scenario encompasses the road environment (such as lane lines, road surface conditions, road geometry, etc.), traffic participants (such as other vehicles, pedestrians, cyclists, and their dynamic information), traffic facilities (such as traffic lights, traffic signs, roadblocks, etc.), weather and lighting conditions, and special situations (such as traffic accidents, road construction, emergency vehicles, etc.). This diversified scenario definition ensures that the autonomous driving system can comprehensively perceive and understand complex driving environments.

[0030] S202, Obtain road elements in the driving scene.

[0031] In this context, road elements refer to the basic components that constitute a driving scenario, representing a structured and semantic representation of a complex driving environment. These elements include, but are not limited to, lanes, traffic participants, and traffic facilities—key factors that can influence autonomous driving decisions. In this embodiment, they can be understood as entities with specific attributes and semantics extracted and identified from raw driving scenario data, mainly categorized into three types: lanes, dynamic objects, and static objects.

[0032] Lane elements include information such as lane lines, lane type, and lane width. Dynamic object elements mainly refer to other vehicles, pedestrians, cyclists, etc., and include their dynamic characteristics such as position, speed, and direction of movement. Static object elements include traffic signs, traffic lights, road barriers, and other objects that do not move or move very little, and record their static characteristics such as type and position. Each road element is given clear semantic information and attributes, enabling the system to accurately understand its role and impact in the driving scenario.

[0033] S203. Obtain a bird's-eye view composed of road elements.

[0034] Please refer to Figure 2 , Figure 2 This invention provides a bird's-eye view diagram. Specifically, a bird's-eye view refers to a two-dimensional image representation method that observes and represents a driving scene from a high-altitude, top-down perspective. This representation method projects a three-dimensional real-world scene onto a plane, centering on the autonomous vehicle, and displays the layout of the surrounding environment and the spatial relationships of various road elements. In the embodiments of this application, it can be understood as a graphical representation of various road elements, such as lanes, dynamic objects, and static objects, within a two-dimensional coordinate system with the vehicle as the origin.

[0035] Based on the above embodiments, as an optional embodiment, step S203, which involves obtaining a bird's-eye view composed of road elements, may further include the following steps:

[0036] S301. Convert dynamic objects into arrows, where arrows are used to represent the relative motion trend of dynamic objects.

[0037] In practice, the system first identifies and locates dynamic objects in the scene based on sensor data, including other vehicles, pedestrians, and cyclists. Then, using positional changes between consecutive frames or velocity information directly obtained from the sensors, it calculates the motion vector for each dynamic object. This motion vector contains both velocity and direction information. Next, the position of each dynamic object in the bird's-eye view is represented by a specific icon, with an arrow superimposed on the icon. The arrow's direction is consistent with the object's direction of motion, while the arrow's length is proportional to the object's velocity.

[0038] Please refer to Figure 3 , Figure 3 This invention provides a lane map diagram. To enhance readability, different colored arrows can be used to distinguish different types of dynamic objects; for example, red arrows represent other vehicles, and blue arrows represent pedestrians. Furthermore, the thickness or transparency of the arrows can indicate speed uncertainty or prediction reliability. For stationary dynamic objects (such as parked vehicles), special markers (such as dots or small circles) can be used to distinguish them.

[0039] S302. Convert the lanes into a lane graph, where the lane graph includes nodes and edges. Nodes represent the center lines of the lanes, and edges include successor edges and adjacent edges. Successor edges represent nodes of the same lane, and adjacent edges represent nodes of adjacent lanes.

[0040] Specifically, in the process of obtaining a bird's-eye view of road elements, converting lanes into a lane map is a crucial step. A lane map is a structured representation of lane networks that describes the topological relationships between lanes through nodes and edges.

[0041] Furthermore, in the lane diagram, nodes represent the centerlines of lanes, while edges are divided into two types: successor edges and adjacent edges. Successor edges connect adjacent nodes in the same lane, indicating the direction of travel for vehicles in that lane; adjacent edges connect adjacent nodes in different lanes, indicating possible lane-changing paths for vehicles. This representation is important because it simplifies complex road networks into a graph structure that is easily processed by computers, thus facilitating subsequent route planning and decision-making.

[0042] In practice, the system first extracts lane line information from the raw sensor data, including the type, location, and shape of the lane lines. Then, it samples each lane line at equal intervals, generating a series of discrete points that will serve as nodes in the lane map. The sampling interval is typically set to a fixed value, such as 10 meters or 20 meters, to control computational complexity while ensuring accuracy. Next, the system analyzes the relationships between adjacent nodes. If two nodes belong to the same lane, a successor edge is added between them; if two nodes belong to adjacent lanes, an adjacent edge is added. During this process, special cases such as lane branching and merging also need to be considered, and the connection relationships between nodes and edges are adjusted accordingly.

[0043] S303, Construct the bird's-eye view corresponding to the arrows, static objects, and lane map.

[0044] S204. Convert the bird's-eye view into a text description to obtain driving scene information.

[0045] Converting bird's-eye view images into text descriptions allows visual information to be transformed into a form that language models can directly process, thus providing more comprehensive and structured input for subsequent large-scale visual-language models. Text descriptions not only contain spatial relationship information from the bird's-eye view but also express semantic information that is difficult to present directly in the image, such as vehicle speed and pedestrian intentions, making the driving scene information richer and more complete.

[0046] In practice, the system identifies and classifies various elements in the bird's-eye view, including lanes, dynamic objects, and static objects. Then, based on the constructed lane map, the system describes the overall road layout, including the number of lanes, lane types, and relationships between lanes. For dynamic objects, the system records their position, speed, and direction of motion, especially their position and motion relative to the autonomous vehicle. Static objects are described as relatively fixed reference points, including their type and location. Furthermore, the system analyzes key information in the scene, such as potential hazards and traffic rules, and incorporates this information into the text description.

[0047] When generating text descriptions, the system uses predefined templates and natural language generation techniques to ensure accuracy and readability. For example, in a typical urban road scenario, the text description might include the following: "The vehicle is located in the right lane of a two-way four-lane road, currently traveling at 40 km / h. There is a red sedan traveling at 35 km / h 50 meters ahead. There is a blue SUV traveling at 45 km / h in the left lane ahead. There is a traffic light 10 meters to the right, currently showing green. There is a pedestrian waiting to cross the road at the intersection 200 meters ahead." This description not only includes basic information about each element in the scene but also reflects their spatial relationships and dynamic characteristics.

[0048] S102. Input the driving scenario information into the visual language big model and obtain the autonomous driving planning parameters output by the visual language big model. The visual language big model is obtained by fine-tuning the sample driving scenario information and the thought chain. The thought chain is used to assist the visual language big model in performing multi-stage reasoning analysis on the sample driving scenario information.

[0049] The "visual-language large-scale model" refers to a large-scale neural network model capable of simultaneously processing and understanding visual and linguistic information. These models are typically based on the Transformer architecture, pre-trained on massive amounts of multimodal data, and exhibit powerful performance on various visual and linguistic tasks. In this embodiment, it can be understood as a large-scale pre-trained model specifically optimized and fine-tuned for autonomous driving scenarios. It can receive driving scene information, including bird's-eye views and textual descriptions, as input and output autonomous driving planning parameters suitable for the current scenario through an internal multi-stage inference process.

[0050] Correspondingly, a thought chain refers to a structured reasoning framework that breaks down complex decision-making processes into a series of logically interconnected steps or stages. This method mimics the human problem-solving process by breaking down large problems into a series of smaller problems, solving them step by step, and finally arriving at a conclusion. In the embodiments of this application, it can be understood as a predefined, multi-stage reasoning structure integrated into a large visual language model to guide the model in systematically analyzing and making decisions about driving scenarios.

[0051] Specifically, this thought chain is used to organize and structure the reasoning process of large visual language models, ensuring that the model follows a certain logical sequence when dealing with complex driving scenarios. For example, the thought chain may include multiple stages such as scenario understanding, risk assessment, behavior prediction, and decision generation, each with clearly defined inputs and outputs. Secondly, the thought chain improves the interpretability and traceability of the reasoning process. By recording the intermediate results of each reasoning stage, the system can clearly demonstrate the entire reasoning path from input to output, which is crucial for understanding and validating the model's decisions. Furthermore, the thought chain enhances the model's reasoning ability, especially when dealing with complex or rare scenarios. By forcing the model to follow predefined reasoning steps, it can prevent the model from ignoring important information or making unreasonable leaps in reasoning.

[0052] In practice, the pre-processed driving scene information, including a bird's-eye view and text description, is first input into the visual language model according to a predefined format. This input format typically includes scene description, current vehicle status, and surrounding environment information. Next, the visual language model performs multi-stage reasoning analysis based on its built-in thought chain. The thought chain is a structured reasoning framework that breaks down complex decision-making processes into multiple logical steps, such as scene understanding, risk assessment, behavior prediction, and decision generation. In each step, the model generates intermediate reasoning results based on the current reasoning outcome and previously accumulated knowledge. These intermediate results are used as input for the next reasoning step, thus forming a complete reasoning chain.

[0053] During the reasoning process, the visual language big data model considers various factors, such as traffic rules, safe distance, comfort, and efficiency, to balance various goals and constraints. Ultimately, the model outputs a set of autonomous driving planning parameters. These parameters refer to a set of instructions and constraints used to guide the behavior of autonomous vehicles. These parameters typically include, but are not limited to, target lane, target speed, acceleration, steering angle, and safe distance, which together define the vehicle's expected motion state and trajectory in the short term. In this embodiment, it can be understood as a set of specific values ​​or instructions output by the visual language big data model based on its understanding and analysis of the current driving scenario. These values ​​or instructions can be directly interpreted and executed by the underlying control system.

[0054] Please refer to Figure 4 and Figure 5 , Figure 4 This is a schematic diagram of the fine-tuning process of a large visual language model provided by the present invention. Figure 5 This diagram illustrates the fine-tuning process of a large visual language model provided by the present invention. The above embodiments illustrate the application process of the large visual language model. Based on the above embodiments, the following will combine... Figure 4 and Figure 5 The fine-tuning process of the large model of visual language is explained.

[0055] S401. Obtain sample driving scenario information.

[0056] The sample driving scenario information refers to a set of structured datasets used for fine-tuning and evaluating large visual language models. These datasets contain detailed descriptions of various real or simulated driving environments, traffic conditions, and vehicle behaviors. In this embodiment, it can be understood as a series of standardized driving scenario representations, each containing a bird's-eye view, text description, and corresponding standard driving decision or behavior annotations. This sample driving scenario information covers a wide range of driving situations, from common daily commuting scenarios to rare emergency situations, to ensure that the model can learn comprehensive driving knowledge.

[0057] S402. Determine the sample thinking chain based on sample driving scenario information. The sample thinking chain includes a first sample thinking sub-chain and a second sample thinking sub-chain. The logical order of the first sample thinking sub-chain is prior to that of the second sample thinking sub-chain. The first sample thinking sub-chain is used to assist the visual language large model in making mid-level decisions on the sample driving scenario. The mid-level decisions include lateral decisions and longitudinal decisions. The lateral decisions are used to represent lane-changing operations, and the longitudinal decisions are used to represent speed control operations. The second sample thinking sub-chain is used to assist the visual language large model in determining the autonomous driving planning parameters for different stages based on the mid-level decisions.

[0058] The "sample thinking chain" refers to a structured reasoning framework template that simulates the thinking and decision-making process of human drivers in various driving scenarios. This framework decomposes complex driving decision-making tasks into a series of logically interconnected steps to guide the visual language large model in systematic analysis and reasoning. In the embodiments of this application, it can be understood as a hierarchical decision structure consisting of two interconnected sub-chains, including a first sample thinking sub-chain and a second sample thinking sub-chain. The first sample thinking sub-chain is used to generate intermediate-level decisions, while the second sample thinking sub-chain generates specific autonomous driving planning parameters based on the intermediate-level decisions.

[0059] Furthermore, the first sample thinking subchain and the second sample thinking subchain are two interconnected but functionally distinct reasoning framework components constituting the sample thinking chain. Together, they simulate the human driving decision-making process from high-level strategy to specific operations. In the embodiments of this application, the first sample thinking subchain can be understood as a reasoning framework focused on generating mid-level decisions. It mainly handles the overall understanding of the driving scenario and macro-level strategy formulation, including lateral decisions (such as whether to change lanes) and longitudinal decisions (such as approximate speed and direction adjustments). This subchain typically includes steps such as scenario perception, traffic condition assessment, safety analysis, and strategy generation. It aims to guide the model in overall scenario understanding and risk assessment, generate high-level driving strategies, and ensure that the model considers macro-level safety and efficiency when making decisions.

[0060] In contrast, the second sample thinking subchain can be understood as a reasoning framework based on mid-level decisions, generating specific autonomous driving planning parameters. It is responsible for transforming the macro-level decisions of the first subchain into precise, executable driving instructions, typically including steps such as path planning, speed planning, execution timing, and parameter optimization. The main function of this subchain is to concretize high-level strategies, generate precise path planning and speed control instructions, optimize driving parameters to ensure comfort and energy efficiency, and formulate detailed execution plans.

[0061] In this context, mid-level decision-making refers to an intermediate-level decision-making process in an autonomous driving system, situated between high-level driving strategies and low-level control commands. This process involves an overall assessment of the current driving scenario and corresponding behavioral choices, but it is not yet refined to specific execution parameters. In the embodiments of this application, mid-level decision-making can be understood as a set of high-level behavioral commands generated by the first sample thought subchain, mainly including two aspects: lateral decision-making and vertical decision-making.

[0062] Lateral decision-making is primarily used to characterize lane-changing operations, involving the selection and transition of vehicles between different lanes. It may include instructions such as "stay in the current lane," "change lanes to the left," or "change lanes to the right." These decisions are based on a comprehensive analysis of factors such as the positions of surrounding vehicles, lane markings, and road structure, aiming to optimize the vehicle's driving path and position.

[0063] Longitudinal decisions are primarily used for standard speed control operations, involving adjustments to the vehicle's speed in the current direction of travel. These may include commands such as "accelerate," "decelerate," "maintain current speed," or "emergency braking." These decisions are based on assessments of factors such as distance to vehicles ahead, road speed limits, and traffic signals, aiming to ensure the vehicle travels safely, legally, and efficiently.

[0064] For example, suppose an autonomous vehicle is traveling on a three-lane urban main road, currently in the middle lane, with a slow-moving large truck 100 meters ahead, several vehicles traveling normally in the left lane, and the right lane relatively empty, with 500 meters to the next intersection. Based on this scenario, we can systematically determine the structure and content of the sample thought chain.

[0065] First, the system constructs a first-sample thinking sub-chain, primarily responsible for scene perception and mid-level decision-making. In this sub-chain, the system first performs scene analysis, identifying the current road type, its own position, and surrounding key elements. Next, it assesses the potential risks of continuing in the current lane and the feasibility and risks of changing lanes. Simultaneously, traffic rules are considered, confirming whether the current speed complies with the road speed limit and evaluating the legality of lane changing. Furthermore, the impact of different decisions on driving efficiency is analyzed. Based on these comprehensive analyses, mid-level decisions are generated, including lateral decisions (such as "change lanes to the right") and longitudinal decisions (such as "maintain current speed"). This process simulates the initial judgment and decision-making process of a human driver facing complex situations.

[0066] After the first sample thinking subchain completes the intermediate-level decision, it enters the second sample thinking subchain. This subchain is responsible for translating the intermediate-level decision into specific planning parameters. In this subchain, path planning is first performed based on the "change lanes to the right" decision, determining the specific trajectory for the lane change. Simultaneously, based on the "maintain current speed" decision and in conjunction with the lane change operation, a detailed speed curve is developed to ensure the smoothness of the entire process. Next, the execution timing needs to be arranged to determine the optimal time for the lane change operation. Then, we optimize and determine specific operational parameters, such as steering angle, acceleration, and target lane position. Finally, a safety check is performed on the generated parameters to ensure that there will be no conflict with surrounding vehicles. This process is similar to how a human driver plans how to execute a decision after making an initial one.

[0067] The sample thought chain determined through the above method can effectively guide the visual language large model to analyze complex driving scenarios step by step and make reasonable decisions. This hierarchical structure not only improves the interpretability and traceability of the decision-making process, but also enhances the model's ability to handle complex scenarios.

[0068] S403. Input the sample driving scenario information and sample thought chain into the visual language big model to obtain the sample autonomous driving planning parameters output by the visual language big model.

[0069] In practice, the system first needs to transform the sample driving scene information into a format that the visual language model can process. This typically involves encoding the bird's-eye view of the scene into image feature vectors and converting the text description into a sequence of word embeddings. Simultaneously, the sample thought chain is encoded into a series of cues or instructions to guide the model's reasoning process. This encoded information is then combined into a structured input and fed into the visual language model.

[0070] Internally, the processing follows the structure defined by the sample thought chain. First, the model understands and analyzes the input scene information based on the first sample thought chain. This includes steps such as identifying key elements, assessing risks, and considering traffic rules, ultimately generating a mid-level decision. This process fully leverages the multimodal understanding capabilities of the visual language big data model, enabling it to process both image and text information simultaneously, thus achieving a comprehensive understanding of the scene.

[0071] Next, based on the second sample thought chain and the generated mid-level decisions, the model further infers specific autonomous driving planning parameters. This process involves steps such as path planning, speed planning, execution timing, and parameter optimization. At this stage, the model needs to transform abstract decisions into concrete numerical parameters, which fully leverages the powerful reasoning and generation capabilities of the visual language big data model.

[0072] Ultimately, the model outputs sample autonomous driving planning parameters, which may include specific values ​​such as target lane, target speed, acceleration, and steering angle. These output parameters not only reflect the model's understanding of the current driving scenario but also demonstrate its ability to reason according to predefined thought processes.

[0073] Based on the above embodiments, as an optional embodiment, S403 may further include the following steps:

[0074] S501. Input the sample driving scenario information and the first sample thinking subchain into the visual language big model to obtain the intermediate decision.

[0075] S502. Input the second sample thinking sub-chain into the visual language big model. Through the second sample thinking sub-chain and the intermediate decision, determine the output target of the visual language big model, and restrict the input and output paradigms of the visual language big model to obtain the sample autonomous driving planning parameters of the autonomous driving target in the current stage and the next stage.

[0076] The output targets refer to the specific information and parameters that the visual language large model needs to generate in the autonomous driving planning task. In the embodiments of this application, they can be understood as a set of well-defined, structured output requirements that specify the type, format, and range of planner parameters that the model needs to generate. Specifically, the output targets include lateral control parameters (such as the lane centerline to be tracked) and longitudinal control parameters (such as speed limits, maximum acceleration, and maximum deceleration). These output targets guide the model to generate parameters that can be directly used in the autonomous driving control system, ensuring the practicality and executability of the model output. At the same time, clear output targets help standardize the model's behavior, improve the consistency and reliability of the output, and provide clear standards for evaluating model performance.

[0077] Correspondingly, the input paradigm refers to a standardized data structure and format template used to standardize the information received by the large visual language model in autonomous driving tasks. In the embodiments of this application, it can be understood as a structured scene description format that specifies in detail how to transform complex driving environment information into data that the model can process. Specifically, the input paradigm includes the state information of the autonomous vehicle (such as speed, acceleration, and current lane lines), surrounding vehicle information (such as relative distance, speed, and orientation), and the selection range of planner parameters. The input paradigm is designed to ensure that the model receives comprehensive and accurate scene information, providing sufficient data support for making accurate planning decisions. At the same time, by specifying the selection range of parameters, the input paradigm also plays a role in limiting the uncertainty of the model's output.

[0078] An output paradigm refers to a predefined, standardized output format and structure used to regulate the output of a large visual language model in autonomous driving planning tasks. In this embodiment, it can be understood as a structured output template that includes the reasoning process and final parameters. Specifically, the output paradigm requires the model to not only provide the final planner parameters but also demonstrate the complete reasoning process, including analysis of vehicle interactions over a future period, lane-changing analysis based on lateral decisions, analysis of longitudinal parameter generation, and a final parameter summary. The output paradigm is designed for several important purposes: First, it ensures that the model's output information is comprehensive and structured, facilitating subsequent system processing and execution. Second, by requiring the demonstration of the reasoning process, it improves the interpretability and traceability of model decisions, contributing to the safety verification and optimization of the system. Third, this structured output helps evaluate the quality of the model's decisions, providing a basis for continuous model improvement. Finally, by standardizing the output format, it simplifies the interface between the model and other autonomous driving system components, improving the integration efficiency and reliability of the entire system.

[0079] Specifically, the system first inputs the second-sample thinking sub-chain into the visual language large model. As a high-level reasoning framework, the second-sample thinking sub-chain, combined with the previously generated mid-level decisions, guides the model to perform deeper and more specific parameter reasoning. This combined approach has unique advantages: the mid-level decisions provide a macro-level understanding and initial strategies for the current driving scenario, while the second-sample thinking sub-chain provides the reasoning path to transform these macro-level strategies into specific parameters. This combination ensures that the model does not deviate from the overall driving strategy when generating specific parameters, thereby improving the consistency and rationality of the decisions.

[0080] Based on the second-sample thinking subchain and mid-level decision-making, the system further determines the output objectives of the large visual language model. These output objectives explicitly define the specific parameter types that the model needs to generate, including lane centerline tracking information required for lateral control, and speed limits, maximum acceleration, and maximum deceleration required for longitudinal control. By explicitly defining the output objectives, the system ensures that the model's output directly corresponds to the needs of the autonomous driving control system, improving the practicality and executability of the output.

[0081] Meanwhile, the system also constrains the behavior of the large visual language model by defining input and output paradigms. Input paradigms specify the standard format of scene information received by the model, including the vehicle's status, surrounding vehicle information, and parameter selection range. This standardized input format ensures the model receives complete and accurate scene information, providing sufficient data support for making accurate decisions. Output paradigms define the standard structure of the model's output, requiring the model to not only provide the final planning parameters but also demonstrate the complete reasoning process. This design significantly improves the interpretability and traceability of the model's decisions, contributing to the system's secure verification and continuous optimization.

[0082] Through this structured and constrained approach, the visual language large model can ultimately generate sample autonomous driving planning parameters for the current and next stages of the autonomous driving target. These parameters include not only the control commands that need to be executed immediately, but also predictions and plans for recent driving behaviors. For example, for a lane change operation, the model will not only provide the steering angle and acceleration that need to be executed at the moment, but also predict the target lane position and speed after the lane change is completed.

[0083] S404. Obtain the simulation results of the sample autonomous driving planning parameters, and adjust the sample thinking chain based on the simulation results.

[0084] Specifically, the system first conducts simulation tests on the sample autonomous driving planning parameters generated by the large visual language model. This simulation test simulates real road environments and traffic conditions, including various common and extreme scenarios, to comprehensively evaluate the feasibility and safety of the planning parameters. The simulation process considers not only the behavior of the autonomous vehicle but also the reactions of other road users and various possible environmental changes. This comprehensive simulation test is crucial for the safety verification of the autonomous driving system because it can fully expose potential problems and risks without jeopardizing actual road safety.

[0085] After the simulation test is completed, the system will perform a detailed analysis of the simulation results. Based on the analysis of the simulation results, the system will make corresponding adjustments to the sample's thought process. This adjustment is a delicate and complex process that may involve multiple levels:

[0086] For example, if simulation results show that the planning parameters perform poorly in certain types of scenarios, the system may adjust the scenario understanding and mid-level decision generation logic in the first-sample thinking subchain to improve the model's ability to identify and process such scenarios. Secondly, if it is found that the selection of certain specific parameters (such as acceleration, steering angle, etc.) is inappropriate, the system may adjust the parameter generation logic in the second-sample thinking subchain, for example, by modifying the parameter selection range or adjusting the priority of parameter generation. Furthermore, if simulation results show that the model's performance is unbalanced in certain aspects (such as safety, efficiency, etc.), the system may adjust the weight allocation in the thinking chain to better balance various performance indicators.

[0087] Based on the above embodiments, as an optional embodiment, step S404, obtaining the simulation results of sample autonomous driving planning parameters, may further include the following steps:

[0088] S601. Input the sample autonomous driving planning parameters into the simulator to obtain the trajectory prediction within the preset time period.

[0089] Specifically, the system first inputs the sample autonomous driving planning parameters generated by the large visual language model into a high-precision simulator. These parameters typically include lateral control information and longitudinal control information. As a complex software system, the simulator is capable of simulating real-world physical characteristics, traffic rules, and the behavior of other road users. It not only considers the dynamic characteristics of autonomous vehicles but also simulates factors such as tire-road friction, air resistance, and vehicle weight distribution to ensure the realism and reliability of the simulation results.

[0090] The simulator simulates the trajectory of an autonomous vehicle within a preset time period based on the input planning parameters. During the simulation, the system not only simulates the behavior of the autonomous vehicle but also the dynamic changes in the surrounding environment, including the movement of other vehicles, pedestrian activity, and changes in traffic signals. This comprehensive simulation ensures that the evaluation results reflect the actual effectiveness of the planning parameters in complex and dynamic environments. The simulator records the vehicle's detailed trajectory within the preset time period, including its position, speed, acceleration, direction, and relative position to other traffic participants, as well as potential conflict points, at each time point.

[0091] Please refer to Figure 6 , Figure 6 This is a simulation diagram provided by the present invention. Figure 6 The paper presents autonomous driving planning parameters based on the output of a large visual language model. A short-term simulation of 2 seconds is performed on the autonomous vehicle using an intelligent driver model. At the same time, a 2-second trajectory prediction map of moving obstacles is generated using a uniform speed model.

[0092] S602. Score the predicted trajectory to obtain the simulation results.

[0093] Specifically, the system can define a comprehensive and reasonable scoring standard. This standard typically includes multiple key dimensions to comprehensively evaluate autonomous driving performance. The main scoring dimensions may include: safety (such as minimum safe distance from other vehicles, number of emergency braking events, etc.), efficiency (such as average speed, travel time, energy consumption, etc.), comfort (such as rate of change of acceleration, steering smoothness, etc.), rule compliance (such as response to traffic signals, lane keeping, etc.), and task completion (such as whether the destination is reached as expected). Each dimension will be assigned a corresponding weight to reflect its importance in the overall evaluation.

[0094] Next, the system will perform a detailed analysis and scoring of the generated predicted trajectory based on this scoring standard. This process typically involves complex data processing and calculations. For example, when assessing safety, the system will calculate the minimum distance to other vehicles along the entire trajectory and count the number of potential collision risks; when assessing efficiency, it will calculate the average speed and total travel time; and when assessing comfort, it will analyze the rate of change of acceleration and steering angle. These calculations not only consider the static characteristics of the trajectory but also analyze its dynamic changes to comprehensively reflect the quality of autonomous driving decisions.

[0095] S405. Integrate the adjusted sample thinking chain into the visual language large model.

[0096] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions from the memory 730 to execute an autonomous driving planning method based on a large visual language model.

[0097] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0098] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the autonomous driving planning method based on the visual language large model provided by the above methods.

[0099] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the autonomous driving planning method based on a large visual language model provided by the above methods.

[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An autonomous driving planning method based on a large visual language model, characterized in that, include: Obtain driving scenario information; The driving scenario information is input into the visual language big model to obtain the autonomous driving planning parameters output by the visual language big model. The visual language big model is obtained by fine-tuning the sample driving scenario information and the thought chain. The thought chain is used to assist the visual language big model in performing multi-stage reasoning analysis on the sample driving scenario information. The fine-tuning process of the large visual language model includes: Obtain sample driving scenario information; The sample thought chain is determined based on the sample driving scenario information; The sample driving scenario information and the sample thought chain are input into the visual language big model to obtain the sample autonomous driving planning parameters output by the visual language big model. Obtain the simulation results of the autonomous driving planning parameters of the sample, and adjust the thought process of the sample based on the simulation results; The adjusted sample thought chain is integrated into the aforementioned visual language large model; The sample thinking chain includes a first sample thinking sub-chain and a second sample thinking sub-chain, with the logical order of the first sample thinking sub-chain preceding that of the second sample thinking sub-chain. The first sample thinking subchain is used to assist the visual language big model in making mid-level decisions for the sample driving scenario. The mid-level decisions include lateral decisions and longitudinal decisions. The lateral decisions are used to represent lane-changing operations, and the longitudinal decisions are used to represent speed control operations. The second sample thinking subchain is used to assist the visual language big model in determining the autonomous driving planning parameters for different stages based on the mid-level decision.

2. The autonomous driving planning method based on a large visual language model according to claim 1, characterized in that, The acquisition of driving scenario information includes: Obtain the driving scenario; Obtain road elements in the driving scenario; Obtain a bird's-eye view composed of the road elements; The bird's-eye view is converted into a text description to obtain driving scene information.

3. The autonomous driving planning method based on a large visual language model according to claim 2, characterized in that, The road elements include lanes, dynamic objects, and static objects. Obtaining the bird's-eye view composed of these road elements includes: The dynamic object is converted into an arrow, wherein the arrow is used to represent the relative motion trend of the dynamic object; The lane is converted into a lane graph, wherein the lane graph includes nodes and edges, the nodes represent the centerline of the lane, the edges include successor edges and adjacent edges, the successor edges represent nodes of the same lane, and the adjacent edges represent nodes of adjacent lanes; Construct a bird's-eye view corresponding to the arrow, the static object, and the lane map.

4. The autonomous driving planning method based on a large visual language model according to claim 1, characterized in that, The step of inputting the sample driving scenario information and the sample thought chain into the visual language big model to obtain the sample autonomous driving planning parameters output by the visual language big model includes: The sample driving scenario information and the first sample thought subchain are input into the visual language big model to obtain the mid-level decision; The second sample thinking subchain is input into the visual language big model. Through the second sample thinking subchain and the middle-level decision, the output target of the visual language big model is determined, and the input and output paradigms of the visual language big model are restricted to obtain the sample autonomous driving planning parameters of the autonomous driving target in the current stage and the next stage.

5. The autonomous driving planning method based on a large visual language model according to claim 1, characterized in that, The simulation results for obtaining the sample autonomous driving planning parameters include: The sample autonomous driving planning parameters are input into the simulator to obtain the trajectory prediction within a preset time period; The predicted trajectory is scored to obtain the simulation results.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the autonomous driving planning method based on a large visual language model as described in any one of claims 1 to 5.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the autonomous driving planning method based on a large visual language model as described in any one of claims 1 to 5.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the autonomous driving planning method based on a large visual language model as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Intelligent agent decision-making method, control method, electronic equipment and storage medium

    CN117151246A

  • Model training method and device based on thinking chain

    CN118114770A