Dynamic style behavior planning method and device based on multi-modal large language model and storage medium

Through the multimodal large language model integrating natural language, image and three-dimensional spatial information, the problems of insufficient generalization performance and poor interpretability in autonomous driving behavior planning are solved, flexible driving style switching and efficient scenario understanding are achieved, and the safety and reliability of the system are improved.

CN120339755APending Publication Date: 2025-07-18COWA TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510500061.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing autonomous driving behavior planning methods lack generalization performance when facing complex scenarios, and the rule-based methods have error accumulation problems, while the end-to-end model is difficult to interpret and debug, affecting safety and reliability.

Method used

The multimodal large language model is adopted, and through driving data acquisition, image style transfer and multimodal language model training, combined with natural language, image and three-dimensional spatial information, cross-modal information fusion and decision-making logic output are realized. The real-world scenario data is obtained using simulator and image style transfer technology, and flexible driving style switching is achieved through diversified language prompts.

Benefits of technology

It improves the generalization performance of the autonomous driving system in complex scenarios, enhances the interpretability of planning results and model debugging capabilities, realizes diversified driving style switching, and improves the safety and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339755A_ABST
    Figure CN120339755A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic style behavior planning method based on a multi-modal large language model and a storage medium, and the method comprises the steps: collecting driving data, and carrying out the collection of various styles of driving data; image style migration: aligning the style of the simulation image to the visual style of the real world image; and multi-modal language model training: mapping data of multiple modalities to a unified feature space, and performing interaction and reasoning on information of different modalities in the common space. According to the method, a simulator and an image style migration technology are utilized, a large amount of driving data close to a real world scene are efficiently obtained at low cost, and the purpose of flexibly converting the automatic driving style of the vehicle can be achieved by means of the characteristic that diversified responses can be generated according to different prompts by means of a multi-mode language model; and meanwhile, the model has the capability of outputting related information such as scene understanding and decision logic in a text form, so that the interpretability of the system is greatly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, belonging to the cross-modal information fusion and text generation technology fields in the fields of behavior planning and multimodal large language models, and in particular to a dynamic style behavior planning method and storage medium based on a multimodal large language model. Background Art

[0002] With the progress of artificial intelligence, sensor technology, and computing power, autonomous driving technology has achieved rapid development in the past decade. Behavior planning is one of the core components of autonomous driving technology, mainly responsible for providing high-level behavior decisions for vehicles in a given environment and traffic conditions. Simply put, behavior planning is "how to do it", rather than just "which road to take". Its core goal is to make appropriate driving decisions by understanding the surrounding environment, including judging when to change lanes, accelerate, decelerate, stop, etc., to ensure that autonomous vehicles can safely and efficiently complete the driving task from the starting point to the end point, while responding to the dynamic changes of various traffic environments during the process.

[0003] A multimodal large language model is a language model that can achieve cross-modal information fusion, capable of deeper understanding and generation by processing and fusing text and other modal data. The advantage of a multimodal large language model lies in its flexible and powerful cross-modal understanding ability. It can not only "see" and "understand" various data such as images, audio, and video, but also selectively extract and output the content in the data in the form of language according to different text prompts. In the field of autonomous driving, a multimodal large language model can assist the autonomous driving system in understanding the current traffic environment. Especially for some complex and rare long-tail scenarios, compared with traditional deep learning models, a large-scale language model can often show stronger generalization performance, thereby improving the safety of the entire autonomous driving system.

[0004] The mainstream autonomous driving behavior planning solutions in the current industrial circle can be divided into rule-based planning and neural network-based planning according to the implementation methods. The rule-based method mainly constructs the process in a cascaded form, combines the environmental information provided by upstream such as perception, map, and positioning, and judges the vehicle's behavior through a set of artificially predetermined rules and decision trees. The defects of this solution lie in the inevitable cumulative errors and the insufficient coping ability for some complex scenarios under fixed rules. Among the neural network-based planning algorithms, the most concerned one currently is the end-to-end model. The entire decision-making process directly maps from the perception data to the vehicle's control instructions, without relying on traditional hierarchical modules, making the entire autonomous driving process more concise and efficient. However, the end-to-end algorithm is usually a "black box" model, which means that its decision-making process is difficult for developers to understand and debug, which is a potential problem for the safety and reliability of autonomous driving, especially when abnormal behaviors occur, it is difficult to trace and explain the reasons. Summary of the Invention

[0005] To solve the above problems of the existing technical solutions, the present invention proposes a dynamic style behavior planning method, device and storage medium based on a multimodal large language model to solve the above technical problems.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A dynamic style behavior planning method based on a multimodal large language model, including the following steps:

[0007] S1, Driving data collection, collecting driving data of various styles with the help of an autonomous driving simulator or an actually deployed vehicle;

[0008] S2, Image style transfer, aligning the style of the simulation image Img vir to the visual style of the real-world image Img real to reduce or even eliminate the differences;

[0009] S3, Multimodal language model training, mapping data of various modalities into a unified feature space, so that the model can interact and reason about different modality information in this common space, thereby better handling various multimodal tasks.

[0010] Further, in S3, for text information, the input text I text , first passes through a tokenizer implemented based on Byte-Pair Encoding (BPE) to tokenize the input string, represents the tokenized text input in the form of indexes, and then uses an embedding layer to map the indexes to the feature form that can be understood by the language model where N is the input text I textThe length after word segmentation, C l is the dimension of the language model feature space.

[0011] Furthermore, in S3, for the image information, the input image I img passes through two independent branches, two-dimensional and three-dimensional, to introduce scene information of other modalities into the multi-modal language model. In the two-dimensional branch, first, a two-dimensional feature extraction network of the ViT (Vision Transformer) architecture is used to extract image features where L is the sequence length after the image features are unfolded, and C i is the image feature dimension. Then, an adapter is used to project the image features into the language feature space The adapter can use a multi-layer perceptron (MLP) or a cross-attention module (Cross Attention); in the three-dimensional branch, a bird's-eye view encoder (BEV Encoder) is adopted to extract three-dimensional space BEV features from the input two-dimensional image where X and Y respectively correspond to the perception ranges in each direction in the BEV space, and C b is the feature dimension. Then, an adapter is also used to project the BEV features into the language feature space

[0012] Furthermore, in S3, after the feature projection is completed, the features of the three modalities of natural language, two-dimensional image, and three-dimensional space are concatenated in a specific order, and the specific order depends on the design method of the input prompt. Finally, the concatenated features are fed into the language model (LLM), which is usually composed of a large number of stacked attention modules paired with a linear layer network.

[0013] Furthermore, the training of the multi-modal language model includes three stages:

[0014] The first stage: learning the basic knowledge in the field of autonomous driving, only activating the language branch, and instilling basic traffic rules and driving common sense into the model in the form of pure text;

[0015] Phase II: Strengthen scene understanding and decision analysis capabilities, introduce image input, activate image features and spatial feature branches based on the first phase, and conduct cross-modal training; use part of the information in the collected data (target detection results, traffic signals, lane information, etc. corresponding to the scene) and public autonomous driving field data sets (wayveai, drivelm, etc.) to build a graphic training set in the form of command questions and answers, input images and unique language prompts set for different tasks, and output various key information about driving in the scene, including the perception of objective elements, the behavior prediction of dynamic traffic participants, and the driving decisions that the vehicle can currently take;

[0016] The third stage: the decision-making ability is embodied as the driving trajectory. During the data collection process in step S1, the simulator can provide the accurate real-time position of the vehicle. According to the difference between consecutive frames, the motion trajectory label W = [(x1, y1), (x2, y2), ..., (x n ,y n )], where x i and i are the lateral and longitudinal offsets of the path point at the future moment i based on the current coordinate system of the vehicle. Here, the time interval between adjacent path points is set to 0.5 seconds, the number of path points n = 10, and the input is the surround view image Img after style transfer in step S2 real , navigation instructions and language prompts corresponding to driving style; the output is a continuous motion trajectory with special semantic tokens (Special Token) spliced before and after stringification, specifically ``<waypoints_start> (x1,y1),(x2,y2),...,(x n ,y n )<waypoints_end> ",in<waypoints_start> and<waypoints_end> They mark the beginning and end of the predicted path point sequence respectively. In the generation task, such start and end symbols can clarify the boundaries of the sequence, which helps the model to handle specific tasks and generate coherent content.

[0017] Furthermore, in S2, there is also a step of style data screening, in which the driving data is screened again based on the rules. In addition to the original state information such as vehicle speed, acceleration, steering angular velocity, etc. directly captured by each sensor in step S1, the lane offset, headway time (TH), and collision time (TTC) are calculated in combination with the real-time scene information provided by the simulator to reconfirm the driving style of the collected data. TH is used to measure the relationship between the distance and speed of the vehicle when following a vehicle. The calculation formula is: D and Vego The following are the following distance and the speed of the host vehicle respectively; TTC is used to measure whether there is a potential risk of an impending collision between the vehicle and the vehicle in front, and the calculation formula is V ahead is the speed of the following vehicle in front; according to the relevant information during the vehicle driving process, the driving style label is judged. If the result is consistent with the expected driving style of the data, it is retained as valid data; if not, it is determined that there may be driving behaviors that do not conform to the expected style in the data, and the data is discarded.

[0018] The present invention also discloses a dynamic style behavior planning device based on a multimodal large language model, including:

[0019] A driving data acquisition device, which is set in an autonomous driving simulator or an actually deployed vehicle, and is used for collecting various styles of driving data;

[0020] An image style transfer device, which is used to align the style of the simulation image Img vir to the visual style of the real-world image Img real , so as to reduce or even eliminate the difference;

[0021] A multimodal language model training device, which is used to map data of multiple modalities into a unified feature space, so that the model can interact and reason about different modality information in this common space, so as to better process various multimodal tasks.

[0022] Furthermore, the above-mentioned multiple modalities include three modalities: natural language, two-dimensional image, and three-dimensional space;

[0023] Among them, the natural language is converted into text information. For the text information, the input text I text , first passes through a tokenizer implemented based on byte-pair encoding (BPE, Byte-Pair Encoding) to tokenize the input string, and represents the tokenized text input in the form of an index. Then, the encoding layer (Embedding) is used to map the index to the feature form that the language model can understand where N is the length of the input text I text after tokenization, and C l is the dimension of the language model feature space;

[0024] For the image information, the input image I img passes through two independent branches of two dimensions and three dimensions to introduce scene information of other modalities into the multimodal language model. In the two-dimensional branch, first, a two-dimensional feature extraction network based on the ViT (Vision Transformer) architecture is used to extract image features Where L is the sequence length after expanding the image features, and C i is the dimension of the image features. Then, an adapter is used to project the image features into the language feature space The adapter can use a multi-layer perceptron (MLP) or a cross-attention module (Cross Attention); in the 3D branch, a bird's-eye view encoder (BEV Encoder) is used to extract 3D space BEV features from the input 2D images where X and Y respectively correspond to the perception ranges in each direction in the BEV space, and C b is the feature dimension. Then, an adapter is also used to project the BEV features into the language feature space

[0025] After completing the feature projection, the features of the three modalities of natural language, 2D images, and 3D space are concatenated in a specific order, and the specific order depends on the design method of the input prompt. Finally, the concatenated features are fed into a language model (LLM), which is usually composed of a large number of stacked attention modules and a linear layer network

[0026] Furthermore, when the image style transfer device performs style data screening, it re-screens the driving data based on rules. In addition to the original state information such as vehicle speed, acceleration, and steering angular velocity directly captured by each sensor in step S1, other key data such as lane offset, time headway (TH), and time to collision (TTC) are calculated by combining the real-time scene information provided by the simulator to re-confirm the driving style of the collected data. Among them, TH is used to measure the relationship between the distance and speed when following a vehicle, and the calculation formula is D and V ego are the following distance and the speed of the host vehicle respectively; TTC is used to measure whether there is a potential risk of an upcoming collision between the vehicle and the preceding vehicle, and the calculation formula is V ahead is the speed of the preceding following vehicle; according to the relevant information during the vehicle driving process, a decision tree method is used to determine the driving style label. If the result is consistent with the expected driving style of the data, it is retained as valid data; if not, it is determined that there may be driving behaviors inconsistent with the expected style in the data, and the data is discarded

[0027] The present invention also provides a computer-readable storage medium containing a computer program, which, when executed by one or more processors, implements the dynamic style behavior planning method based on a multi-modal large language model described in any one of the above

[0028] The advantages of implementing autonomous driving behavior planning based on multimodal large language models are as follows:

[0029] 1. Compared with rule-based methods, it has stronger generalization performance in the face of some complex long-tail scenarios and avoids the influence brought by error accumulation;

[0030] 2. Compared with end-to-end methods, the language model can directly output scene information and decision logic in text form, greatly improving the interpretability of the planning results and providing references for developers to analyze and debug the model;

[0031] 3. With the help of diverse language prompts, the multimodal large language model can generate diversified language responses under the same image input. Based on this feature, it is convenient to integrate various styles of driving strategies in one model and flexibly switch driving styles according to different text prompt information. Brief Description of the Drawings

[0032] Figure 1 It is a schematic flow diagram of a dynamic style behavior planning method based on a multimodal large language model of the present invention;

[0033] Figure 2 It is a schematic diagram of an image style transfer example of the present invention;

[0034] Figure 3 It is a schematic diagram of the multimodal language model structure of the present invention;

[0035] Figure 4 It is a schematic diagram of an instruction Q&A training sample example of the present invention;

[0036] Figure 5 It is a schematic diagram of a multi-style behavior planning example of the present invention. Detailed Embodiments

[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] As Figures 1 - 3 shown, a dynamic style behavior planning method based on a multimodal large language model, as Figure 1 shown, this solution includes driving data collection, image style transfer, and multimodal language model training. Specifically, it includes the following steps:

[0039] S1, driving data collection, using autonomous driving simulators or actual deployed vehicles to collect driving data of various styles;

[0040] Take the collection of driving data of various styles with the help of an autonomous driving simulator as an example. Before starting data collection, it is necessary to configure the on-board sensors in the simulator to obtain and record relevant information during driving. In this solution, the simulated vehicle is equipped with a global navigation satellite system (GNSS), an inertial measurement unit (IMU), a speedometer, and a surround-car high-definition camera. The GNSS provides the real-time position of the vehicle; the IMU provides the vehicle's acceleration, angular velocity and other information, and calculates the vehicle's motion state in real time for driving style identification; the speedometer outputs the vehicle's current speed; the surround-car camera collects scene images Img during the vehicle's driving process. vir , 1 to 8 camera sensors can be configured according to actual needs. For example, this solution configures 6 cameras in total to collect images in front of the vehicle, in front of the left and right sides, behind the left and right sides, and directly behind the vehicle. The specific parameters of the camera, such as installation position, pitch angle, resolution, and acquisition frequency, should be as consistent as possible with the actual deployed vehicles in reality.

[0041] The high-precision maps used for data collection are based on the open source high-precision geographic data project OpenStreetMap (OSM). OSM is an open geographic information platform that provides street and infrastructure data worldwide. These data can be used to generate road networks, lane markings, intersections and other elements, thereby providing diverse and realistic driving scenarios. Different maps will contain distinctive road environments, such as urban blocks, suburbs, rural areas, highways and special test areas (such as sharp turns, roundabouts, construction areas, etc.).

[0042] The simulator can enable dynamic time, weather, pedestrian density, and traffic density configurations to ensure that the simulated driving scene is more realistic and complex. After the autonomous driving agent is enabled, data can be collected through the simulator. In addition to the raw data recorded by each sensor, high-level information provided by the simulator can also be selectively retained, including navigation instructions (straight, turn, etc.), the category and location of other traffic participants, traffic light status, traffic signs, current lane and road type, etc.

[0043] After that, data collection is carried out through manual takeover. The driver needs to control the vehicle according to the navigation instructions provided by the simulator, and mark the collectors as drivers of different styles according to information such as the average vehicle speed, obstacle avoidance distance, and lane change strategy when controlling the vehicle. The driving tasks on the same route are collected by drivers of different styles to obtain driving data of multiple styles under the same scenario.

[0044] S2. Image style transfer, align the style of the simulation image Img vir to the visual style of the real-world image Img real to reduce or even eliminate the differences;

[0045] As Figure 2 shown, the virtual-style image Img vir collected through the simulator in has certain differences from the real-world scenario. Direct use will affect the effectiveness of the final model. It is necessary to perform style transfer processing on the collected images. Use the model CycleGAN built based on convolutional neural network to align the style of the simulation image Img vir to the visual style of the real-world image Img real to reduce or even eliminate the differences. As Figure 2 shown, after passing through the style transfer model, the virtual image is converted into a real image with the style of the Cityscapes dataset. Different style model weights can be selected according to the actual situation to convert the image into the image style of the actual application scenario of the vehicle.

[0046] In some embodiments, in S2, it further includes the step of filtering style data. Driving styles are usually reflected in multiple aspects of actual driving behavior. For example, a sporty style may be more inclined to overtaking, frequent lane changes, and a short following distance, etc.; while a comfortable style will try to avoid drastic acceleration and braking to provide a smoother riding experience. Because manual control has strong subjectivity and uncertainty, it is necessary to screen the driving data again based on rules. In addition to the original state information such as vehicle speed, acceleration, and steering angular velocity directly captured by each sensor in step S1, combined with the real-time scene information provided by the simulator, other key data such as lane offset, time headway (TH), and time to collision (TTC) are calculated to reconfirm the driving style of the collected data. Among them, TH is used to measure the relationship between distance and speed when following a vehicle, and the calculation formula is D and V ego are the following distance and the vehicle speed of the host vehicle respectively; TTC is used to measure whether there is a potential risk of an upcoming collision between the vehicle and the vehicle in front, and the calculation formula is V aheadis the speed of the vehicle in front; in practical applications, due to differences in different vehicle models and different road environments, the above key data also includes other judgment conditions that can be added according to the actual situation. For example, including but not limited to: acceleration, steering angular velocity, whether speeding, whether going against the traffic, obstacle avoidance distance, and so on.

[0047] In some embodiments, according to the relevant information during vehicle driving, the driving style label is judged in the way of a decision tree. The internal node of the decision tree represents the judgment of a certain attribute. Here, it mainly judges whether TH and TTC exceed the threshold, and other key data mentioned above, such as whether going against the traffic, whether speeding, whether the speed change rate exceeds the threshold, and so on. The judgment continues in the branches of the node or belongs to a certain type of label result. If the result is consistent with the expected driving style of the data, it is retained as valid data; if not, it is determined that there may be driving behaviors inconsistent with the expected style in the data, and the data is discarded.

[0048] S3. Training of the multimodal language model, mapping data of multiple modalities into a unified feature space, so that the model can interact and reason about different modality information in this common space, thereby better processing various multimodal tasks.

[0049] As Figure 3 shown, the multimodal language model structure adopted in this solution uses input information of three modalities: natural language, two-dimensional image, and three-dimensional space. For text information, the input text I text , first passes through a tokenizer implemented based on Byte-Pair Encoding (BPE) to tokenize the input string, representing the tokenized text input in the form of an index, and then uses the Embedding layer to map the index to a feature form understandable by the language model where N is the length of the input text I text after tokenization, and C l is the dimension of the language model feature space.

[0050] For image information, the input image I img passes through two independent branches of two dimensions and three dimensions, introducing scene information of other modalities for the multimodal language model. In the two-dimensional branch, first use a two-dimensional feature extraction network based on the ViT (Vision Transformer) architecture to extract image features where L is the length of the sequence after the image features are unfolded, and C i is the dimension of the image features. Then use an adapter to project the image features into the language feature space The adapter can use a multi-layer perceptron (MLP) or a cross-attention module. In the 3D branch, a bird's-eye view encoder (BEV Encoder) is adopted to extract 3D space BEV features from the input 2D images. where X and Y respectively correspond to the perception ranges in each direction in the BEV space, and C b is the feature dimension. After that, an adapter is also used to project the BEV features into the language feature space. In the pure vision solution, 3D data such as point clouds are not adopted. Therefore, the "input information of three modalities in 3D space" can be understood as "pseudo-3D" for the sake of rigor. The BEV features are 3D information with the height compressed, which can be regarded as 2.5D or pseudo-3D features. There are many mature solutions for mapping images to BEV features, such as LSS, deformable attention, and so on.

[0051] After completing the feature projection, the features of natural language, 2D images, and 3D space in three modalities are concatenated in a specific order, and the specific order depends on the design method of the input prompt. Finally, the concatenated features are fed into a language model (LLM), which is usually composed of a large number of stacked attention modules and a linear layer network.

[0052] The specific order can be determined as follows: placeholders are inserted at specific positions in the input text, and after feature encoding, the features are concatenated according to the input order. For example, the following input:

[0053] "The circumferential vehicle image of the current scene is: , and the 3D space information is:

[0054] <3d_feature>. According to the above information......".

[0055] In some embodiments, the training of the multi-modal language model includes three stages:

[0056] The first stage: learning the basic knowledge in the field of autonomous driving, only activating the language branch, and instilling basic traffic rules and driving common sense into the model in the form of pure text. After the training in the first stage, the model has mastered the theoretical knowledge about driving, such as stopping at red lights, going at green lights, avoiding collisions, and bypassing dangerous areas.

[0057] The second stage: Strengthen scene understanding and decision analysis capabilities, introduce image input, activate image features and spatial feature branches based on the first stage, and conduct cross-modal training; use part of the information in the collected data (target detection results, traffic signals, lane information, etc. corresponding to the scene) and public autonomous driving data sets (wayveai, drivelm, etc.) to build a graphic training set in the form of command questions and answers, input images and unique language prompts set for different tasks, and output various key information about driving in the scene, including the perception of objective elements, the behavior prediction of dynamic traffic participants, and the driving decisions that the vehicle can currently take; after the second stage of training, the model can conduct a comprehensive analysis of the current scene and provide reference opinions for driving behavior. Examples of training samples in the second stage are as follows: Figure 4 shown.

[0058] The third stage: the decision-making ability is embodied as the driving trajectory. During the data collection process in step S1, the simulator can provide the accurate real-time position of the vehicle. According to the difference between consecutive frames, the motion trajectory label W = [(x1, y1), (x2, y2), ..., (x n ,y n )], where x i and i are the lateral and longitudinal offsets of the path point at the future moment i based on the current coordinate system of the vehicle. Here, the time interval between adjacent path points is set to 0.5 seconds, the number of path points n = 10, and the input is the surround view image Img after style transfer in step S2 real , navigation instructions and language prompts corresponding to driving style; the output is a continuous motion trajectory with special semantic tokens (Special Token) spliced before and after stringification, specifically ``<waypoints_start> (x1,y1),(x2,y2),...,(x n ,y n )<waypoints_end> ",in<waypoints_start> and<waypoints_end> They mark the beginning and end of the predicted path point sequence respectively. In the generation task, such start and end symbols can clarify the boundaries of the sequence, which helps the model to handle specific tasks and generate coherent content. After the third stage of training, the model can output a matching predicted trajectory based on the surround image and the instructions containing navigation information and style labels. Specific examples of use are as follows: Figure 5 shown.

[0059] The present invention also discloses a dynamic style behavior planning device based on a multimodal large language model, comprising:

[0060] A driving data acquisition device, which is set in an autonomous driving simulator or an actually deployed vehicle and is used for acquiring driving data of various styles;

[0061] An image style transfer device, which is used to align the style of the simulation image Img vir to the visual style of the real-world image Img real to reduce or even eliminate the difference;

[0062] A multi-modal language model training device, which is used to map data of multiple modalities into a unified feature space, so that the model can interact and reason about different modality information in this common space, thereby better handling various multi-modal tasks.

[0063] The aforementioned multiple modalities include three modalities: natural language, two-dimensional image, and three-dimensional space;

[0064] Among them, to convert natural language into text information, for the text information, the input text I text , first passes through a tokenizer implemented based on Byte-Pair Encoding (BPE) to tokenize the input string, and represents the tokenized text input in the form of indexes. Then, an embedding layer is used to map the indexes to a feature form that can be understood by the language model where N is the length of the input text I after tokenization, and C text is the dimension of the language model feature space; l For image information, the input image I

[0065] passes through two independent branches of two-dimensional and three-dimensional to introduce scene information of other modalities into the multi-modal language model. In the two-dimensional branch, first, a two-dimensional feature extraction network with a Vision Transformer (ViT) architecture is used to extract image features img where L is the sequence length after the image features are unfolded, and C is the dimension of the image features. Then, an adapter is used to project the image features i to the language feature space The adapter can use a multi-layer perceptron (MLP) or a cross-attention module (Cross Attention); in the three-dimensional branch, a Bird's Eye View Encoder (BEV Encoder) is used to extract three-dimensional space BEV features from the input two-dimensional image where X and Y respectively correspond to the perception ranges in each direction in the BEV space, and C is the feature dimension. Then, an adapter is also used to project the BEV features b to the language feature space Projected to the language feature space

[0066] After completing the feature projection, the features of the three modalities of natural language, 2D images, and 3D space are concatenated in a specific order, which specifically depends on the design of the input prompt. Finally, the concatenated features are fed into a language model (LLM), which is usually composed of a large number of stacked attention modules paired with a linear layer network.

[0067] When the image style transfer device performs style data screening, driving styles are usually reflected in multiple aspects of actual driving behavior. For example, a sporty style may be more inclined to overtaking, frequent lane changes, and a short following distance, etc.; while a comfortable style will try to avoid drastic acceleration and braking to provide a smoother riding experience. Because manual control has strong subjectivity and uncertainty, it is necessary to screen the driving data again based on rules. In addition to the original state information such as vehicle speed, acceleration, and steering angular velocity directly captured by each sensor in step S1, other key data such as lane offset, time headway (TH), and time to collision (TTC) are calculated by combining the real-time scene information provided by the simulator to reconfirm the driving style of the collected data. Among them, TH is used to measure the relationship between the distance and speed when following a vehicle, and the calculation formula is D and V ego are the following distance and the speed of the host vehicle respectively; TTC is used to measure whether there is a potential risk of an impending collision between the vehicle and the vehicle in front, and the calculation formula is V ahead is the speed of the vehicle following in front; according to the relevant information during the vehicle driving process, the driving style label is judged by using the decision tree method. If the result is consistent with the expected driving style of the data, it is retained as valid data; if not, it is determined that there may be driving behaviors inconsistent with the expected style in the data, and the data is discarded.

[0068] The present invention also provides a computer-readable storage medium containing a computer program, which, when executed by one or more processors, implements the dynamic style behavior planning method based on a multi-modal large language model described in any one of the above.

[0069] The advantages of implementing autonomous driving behavior planning based on a multi-modal large language model are as follows:

[0070] 1. Compared with rule-based methods, it has stronger generalization performance in the face of some complex long-tail scenarios and avoids the influence brought by error accumulation;

[0071] 2. Compared with the end-to-end method, the language model can directly output the scenario information and decision logic in text form, greatly improving the interpretability of the planning results and providing a reference for developers to analyze and debug the model;

[0072] 3. With the help of diverse language prompts, the multimodal large language model can generate diversified language responses for the same image input. Based on this feature, it is convenient to integrate various driving strategies in one model and flexibly switch the driving style according to different text prompt information.

[0073] The present invention provides a technical solution for realizing dynamic style behavior planning of autonomous driving through a multimodal large language model. By using a simulator and image style transfer technology, a large amount of driving data close to real-world scenarios can be obtained efficiently and at low cost. By virtue of the feature that the multimodal language model can generate diversified responses according to different prompts, the purpose of flexibly switching the autonomous driving style of the vehicle is achieved. At the same time, since the model has the ability to output relevant information such as scenario understanding and decision logic in text form, the interpretability of the system is greatly enhanced.

[0074] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.

Claims

1. A dynamic style behavior planning method based on a multimodal large language model, characterized in that Including the following steps: S1. Driving data collection: Collect driving data of various styles with the help of an autonomous driving simulator or an actually deployed vehicle. S2, Image style transfer, aligning the style of the simulation image Img vir to the visual style of the real-world image Img real to reduce or even eliminate the differences; S3. Multi-modal language model training: Map data of multiple modalities into a unified feature space, enabling the model to interact and reason about different modality information in this common space, so as to better handle various multi-modal tasks.

2. The dynamic style behavior planning method based on a multimodal large language model according to claim 1, wherein In S3, for text information, the input text I text , first passes through a tokenizer Tokenizer implemented based on byte pair encoding (BPE) to tokenize the input string, and represents the tokenized text input in the form of indices. Then, the encoding layer Embedding is used to map the indices to a feature form that can be understood by the language model where N is the length of the input text I text after tokenization, and C l is the dimension of the language model feature space.

3. The dynamic style behavior planning method based on a multimodal large language model according to claim 1 or 2, characterized in that, In S3, for the image information, the input image I img passes through two independent branches of two - dimensional and three - dimensional to introduce scene information of other modalities for the multi - modal language model. In the two - dimensional branch, first, a two - dimensional feature extraction network of the ViT architecture is used to extract image features where L is the sequence length after the image features are unfolded, and C i is the dimension of the image features. After that, an adapter is used to project the image features into the language feature space The adapter can use a multi - layer perceptron MLP or a cross - attention module Cross Attention; In the three-dimensional branch, an aerial view encoder BEVEncoder is used to extract three-dimensional space BEV features from the input two-dimensional images. Where X and Y respectively correspond to the perception ranges in each direction in the BEV space, and C b is the feature dimension. Then, an adapter Adapter is also used to project the BEV features into the language feature space.

4. The dynamic style behavior planning method based on a multimodal large language model according to claim 3, wherein, In S3, after completing the feature projection, the features of the three modalities of natural language, two-dimensional images, and three-dimensional space are concatenated in a specific order, which specifically depends on the design of the input prompt. Finally, the concatenated features are fed into the language model LLM, which is typically composed of a large number of stacked attention modules paired with a linear layer network.

5. The dynamic style behavior planning method based on the multi-modal large language model according to claim 4, characterized in that The training of the multi-modal language model includes three stages: The first stage: Learning basic knowledge in the field of autonomous driving. Only activate the language branch and instill basic traffic rules and driving common sense into the model in the form of pure text. The second stage: Strengthening the ability of scene understanding and decision-making analysis. Introduce image input, activate the image feature and spatial feature branches on the basis of the first stage, and conduct cross-modal training. Use part of the information in the collected data and the publicly available data set in the field of autonomous driving to construct a graphic-text training set in the form of instruction answering. Input images and unique language prompts for different tasks, and output various key driving-related information within the scene, including the perception of objective elements, the behavior prediction of dynamic traffic participants, and the driving decisions that the vehicle itself can take currently. The third stage: Visualize the decision-making ability as a driving trajectory. During the data collection in step S1, the simulator can provide the accurate real-time position of the vehicle. Based on the difference between consecutive frames, the motion trajectory label W = [(x1, y1), (x2, y2),..., (x n , y n )] of the vehicle in the current state within a certain period in the future can be calculated, where x i and y i are the lateral and longitudinal offsets of the path point at the future ith moment based on the current coordinate system of the host vehicle respectively. Here, the time interval between adjacent path points is set to 0.5 seconds, and the number of path points n = 10. The input is the panoramic image Img real after style transfer in step S2, the navigation instruction, and the language prompt corresponding to the driving style; The output is a continuous motion trajectory with special semantic tokens "Special Token" concatenated before and after being stringified, specifically as ``<waypoints_start>(x1, y1),(x2, y2),...,(x n , y n ))<waypoints_end>", where <waypoints_start> and <waypoints_end> mark the start and end of the predicted path point sequence respectively. In the generation task, such start and end symbols can clarify the boundaries of the sequence, which helps the model process specific tasks and generate coherent content.

6. The dynamic style behavior planning method based on a multimodal large language model according to claim 5, characterized in that In S2, it also includes the step of filtering style data, and the driving data is filtered again based on rules. In addition to the original state information of vehicle speed, acceleration, and steering angular velocity directly captured by each sensor in step S1, the lane offset, time headway TH, and time to collision TTC calculated by combining the real-time scene information provided by the simulator are used to reconfirm the driving style of the collected data. Among them, TH is used to measure the relationship between the distance and speed when following a vehicle, and the calculation formula is D and V ego are the following distance and the speed of the host vehicle respectively; TTC is used to measure whether there is a potential risk of an impending collision between the vehicle and the vehicle in front, and the calculation formula is V ahead is the speed of the following vehicle in front; according to the relevant information during the vehicle driving process, the driving style label is judged. If the result is consistent with the expected driving style of the data, it is retained as valid data; if not, it is determined that there may be driving behaviors inconsistent with the expected style in the data, and the data is discarded.

7. A dynamic style behavior planning device based on a multimodal large language model, characterized in that Including: A driving data collection device, which is set in an autonomous driving simulator or an actually deployed vehicle and is used for collecting driving data of various styles. An image style transfer device for aligning the style of a simulation image Img vir to the visual style of a real-world image to reduce or even eliminate the differences; A multi-modal language model training device, which is used to map data of multiple modalities into a unified feature space, enabling the model to interact and reason about different modality information in this common space, so as to better handle various multi-modal tasks.

8. The dynamic style behavior planning device based on a multimodal large language model according to claim 7, characterized in that, The multiple modalities mentioned above include three modalities: natural language, two-dimensional image, and three-dimensional space. Among them, converting natural language into text information, for the text information, the input text I text , first passes through a tokenizer Tokenizer implemented based on byte pair encoding (BPE) to tokenize the input string, and represents the tokenized text input in the form of indices. Then, the encoding layer Embedding is used to map the indices to the feature form that can be understood by the language model where N is the input text I text is the length after tokenization, and C l is the dimension of the language model feature space; For image information, the input image I img passes through two independent branches of two-dimensional and three-dimensional to introduce scene information of other modalities into the multi-modal language model. In the two-dimensional branch, first, a two-dimensional feature extraction network based on the ViT (Vision Transformer) architecture is used to extract image features where L is the sequence length after the image features are unfolded, and C i is the dimension of the image features. Then, the adapter Adapter is used to project the image features into the language feature space The adapter can use a multi-layer perceptron MLP or a cross-attention module Cross Attention; in the three-dimensional branch, a bird's-eye view encoder BEV Encoder is adopted to extract three-dimensional space BEV features from the input two-dimensional image where X and Y respectively correspond to the perception ranges in each direction in the BEV space, and C b is the feature dimension. Then, the adapter Adapter is also used to project the BEV features into the language feature space After the feature projection is completed, the features of the three modalities of natural language, two-dimensional images, and three-dimensional space are concatenated in a specific order, which specifically depends on the design method of the input prompt. Finally, the concatenated features are fed into the language model LLM, which is usually composed of a large number of stacked attention modules paired with a linear layer network.

9. The dynamic style behavior planning device based on a multimodal large language model according to claim 7 or 8, characterized in that When the image style transfer device performs style data screening, it re - screens the driving data based on rules. In addition to the original state information of vehicle speed, acceleration, and steering angular velocity directly captured by each sensor in step S1, it combines the real - time scene information provided by the simulator to calculate the lane offset, time headway TH, and time - to - collision TTC to re - confirm the driving style of the collected data. Among them, TH is used to measure the relationship between the distance and speed when following a vehicle, and the calculation formula is D and V ego are the following distance and the speed of the host vehicle respectively; TTC is used to measure whether there is a potential risk of an upcoming collision between the vehicle and the vehicle in front, and the calculation formula is V ahead is the speed of the following vehicle in front; According to the relevant information during the vehicle driving process, the driving style label is judged. If the result is consistent with the expected driving style of the data, it is retained as valid data; if not, it is determined that there may be driving behaviors inconsistent with the expected style in the data, and the data is discarded.

10. A computer-readable storage medium comprising a computer program, characterized in that, When the computer program is executed by one or more processors, it implements the dynamic style behavior planning method based on the multi-modal large language model according to any one of claims 1-6.

Citation Information

Cited By

  • Vehicle automatic driving test method, device, equipment and medium

    CN120846693A