Cut-in scenario vehicle trajectory prediction method based on traffic context map and vlm

CN122561034BActive Publication Date: 2026-09-25JILIN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611066494.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-25
Estimated Expiration
2046-07-17

AI Technical Summary

Technical Problem

然而,若直接采用视觉语言模型生成的自然语言结果或离散意图标签用于轨迹预测,容易在文本解码过程中造成信息压缩和信息损失,且难以与连续轨迹预测模型充分融合,例如:中国专利CN121210643A公开了“一种基于视觉-语言模型、近端策略优化和扩散模型的自动驾驶轨迹预测方法”,虽然利用了视觉语言模型对场景理解,但是将理解结果输出,然后在进行二次映射,这种方式使得特征的表达受到了严重的信息压缩,影响预测结果准确性

Benefits of technology

(1)本发明面向高速公路Cut-in高风险交互场景,将目标车辆、自车、CIPV及周围交通参与者的历史运动状态、车道几何结构、候选切入间隙(候选目标车道间隙区域)、道路前进方向和关键交互约束统一构建为交通上下文图,使原本分散的轨迹信息、道路信息和车辆交互信息能够以直观、结构化的方式表达,有利于增强模型对Cut-in场景整体空间关系和交通语义关系的理解能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122561034B_ABST
    Figure CN122561034B_ABST
Patent Text Reader

Abstract

The application provides a Cut-in scene vehicle trajectory prediction method based on a traffic context graph and a VLM, first, historical motion states of vehicles in a highway Cut-in scene are acquired, and key interactive vehicles are determined in combination with high-precision map information; then, a traffic context graph is constructed according to vehicle historical trajectories, current postures, lane geometric structures and the like, and a structured scene description feature file corresponding to the traffic context graph is generated; subsequently, the graph and text data are jointly input into a visual language model, hidden layer dense features are extracted in a forward inference process of the visual language model, and dense features related to a task are obtained through mean pooling, maximum pooling and global semantic feature extraction; subsequently, the dense features are fused with historical motion features, road topological features and vehicle interaction features, and a target vehicle future single-mode or multi-mode prediction trajectory and a probability thereof are output; the method improves the accuracy, robustness and explainability of vehicle trajectory prediction in the highway Cut-in scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent transportation and autonomous driving technology, and relates to environmental perception and behavior prediction of intelligent connected vehicles. Specifically, it relates to a method for predicting vehicle trajectory in cut-in scenarios based on traffic context graph and visual language model (VLM). Background Technology

[0002] With the development of autonomous driving technology, vehicle trajectory prediction has become a crucial component of perception, decision-making, and planning in autonomous driving systems. It primarily predicts the target vehicle's future movement trend over a period of time based on the historical motion states of the target vehicle, the resident vehicle, and surrounding traffic participants, as well as road topology and traffic environment information. This provides support for collision risk assessment, behavioral decisions, and path planning. Highway cut-in scenarios are typical high-risk interaction scenarios in autonomous driving. In this scenario, the target vehicle may cut into the resident vehicle's lane from an adjacent lane. Its future trajectory is influenced by factors such as its own motion state, the gap between vehicles in front and behind in the target lane, the resident vehicle's relative position, the nearest on-path vehicle (CIPV), relative speed, collision time, and road structure. Therefore, accurately predicting the future trajectory of the target vehicle in cut-in scenarios is of great significance for improving the driving safety of autonomous vehicles.

[0003] Existing vehicle trajectory prediction methods mainly include rule-based models, traditional machine learning methods, and deep learning methods. Deep learning methods can use structures such as recurrent neural networks, graph neural networks, and Transformers to model historical trajectories, road maps, and vehicle interaction relationships. However, they typically rely primarily on historical trajectories and numerical interaction features, and do not adequately utilize target vehicle cutting intentions, candidate cutting gaps, CIPV constraints, and traffic context semantic information. This makes them prone to problems such as prediction lag, pattern misjudgment, and trajectory instability when cutting intentions are unclear, target lane gaps change dynamically, or vehicle interactions are complex.

[0004] Vision-Language Models (VLMs) possess strong image understanding and text semantic understanding capabilities, offering novel technical approaches for traffic scene understanding. However, directly using natural language results or discrete intent labels generated by VLMs for trajectory prediction can easily lead to information compression and loss during text decoding, and is difficult to fully integrate with continuous trajectory prediction models. For example, Chinese patent CN121210643A discloses "An Autonomous Driving Trajectory Prediction Method Based on a Vision-Language Model, Proximal Policy Optimization, and Diffusion Model." While utilizing a VLM for scene understanding, the output of the understanding results followed by secondary mapping severely compresses the feature representation, affecting the accuracy of the prediction results. Therefore, there is an urgent need to improve existing vehicle trajectory prediction methods to enhance the accuracy, robustness, and scene understanding capabilities of trajectory prediction in highway cut-in scenarios. Summary of the Invention

[0005] In view of the shortcomings and deficiencies of existing technologies, the purpose of this invention is to propose a vehicle trajectory prediction method for cut-in scenes based on traffic context graphs and VLMs. This method first acquires historical motion state data of the vehicle, target vehicle, and surrounding traffic participants in a highway cut-in scene, and combines this data with high-precision map information to determine key interacting vehicles and interaction relationship features. Then, it constructs a traffic context graph based on vehicle historical trajectories, lane geometry, current vehicle posture, candidate cut-in gaps, and road direction of travel, while simultaneously generating a corresponding structured scene description feature file. Next, the traffic context graph and the structured scene description feature file are jointly input into a visual language model to jointly encode traffic image information and scene semantic information. Based on this, the method does not directly use the data generated by the visual language model. Instead of relying on natural language results or discrete intent labels, this method directly extracts dense features from the hidden layer during the forward inference process. These features are then obtained through mean pooling, max pooling, and global semantic feature extraction to obtain task-related visual language dense features. Subsequently, these dense features are fused with historical motion features, road topology features, and vehicle interaction features. A trajectory prediction decoder is then used to output the target vehicle's future single-modal or multi-modal predicted trajectory and its probability. This method integrates spatial geometric relationships, vehicle interaction relationships, and high-level traffic semantic information into the trajectory prediction process, reducing information loss caused by natural language decoding in visual language models. It improves the accuracy, robustness, and interpretability of vehicle trajectory prediction in highway cut-in scenarios, providing a reliable basis for risk assessment, behavioral decisions, and path planning for autonomous vehicles.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A vehicle trajectory prediction method for cut-in scenarios based on traffic context graphs and VLMs includes the following steps: Step 1. Obtain historical motion status data of each vehicle in the highway cut-in scenario, as well as surrounding road environment information, and combine with high-precision map information to determine key interactive vehicles; Step 2. Map and draw the vehicle's historical trajectory, current vehicle posture, lane geometry, candidate target lane gap area, and road direction of travel onto the traffic context graph, and set different visual styles for different objects. Step 3. Based on the data obtained in Step 1, generate a structured scene description feature file corresponding to the traffic context graph; Step 4. Construct Cut-in scene task prompts for visual language models; Step 5. Construct a combined image and text input based on the data from Steps 2 to 4. The image processing unit corresponding to the visual language model is used for preprocessing, followed by forward propagation. During forward propagation, the intermediate hidden states of the visual language model are output. Then, a preset set of layers is selected from the intermediate hidden states. The corresponding hidden state is determined, and the hidden representation of the image token is extracted. The last valid text token hidden representation corresponding to the current traffic scenario ; after that Max pooling and average pooling are performed separately; finally, the image tokens are average pooled for feature extraction. Image token max pooling feature and The features are then spliced ​​together to obtain task-relevant dense features. Step 6. Encode the historical motion state and high-precision map information to obtain the interaction features; stitch the interaction features with the task-related dense features, and then perform mapping, fusion and decoding to obtain the future trajectory of the target vehicle.

[0007] As a preferred embodiment of the present invention, the historical motion state data in step 1 includes the vehicle's position coordinates, velocity, acceleration, heading angle, vehicle length, and width; the key interactive vehicle includes the target vehicle. bicycle The vehicle in front of the bicycle lane The target vehicle is a vehicle intended for cut-in.

[0008] As a preferred embodiment of the present invention, after determining the key interactive vehicle in step 1, a set of Cut-in scene interaction relationship features is constructed based on the key interactive vehicle; wherein, the Cut-in scene interaction relationship features include: and The longitudinal distance and collision time, and The longitudinal distance and collision time, and and The longitudinal distance and collision time.

[0009] As a preferred embodiment of the present invention, the historical observation time window is used in step 2. The historical trajectories of each vehicle are mapped and plotted onto the bird's-eye view traffic context map; based on the vehicle's current position, heading angle, vehicle length, and vehicle width obtained in step 1, the target vehicle is... bicycle The surrounding vehicles are drawn as bird's-eye view vehicle bounding boxes; the lane centerlines and lane boundaries around the current cut-in scene are mapped and drawn onto the bird's-eye view traffic context map, and the road forward direction indicator is determined based on the direction of the drawn lane centerlines; at the same time, the longitudinal area between the vehicle's EGO and the nearest on-path vehicle in front is determined as the candidate target lane gap area. .

[0010] As a preferred embodiment of the present invention, in step 2, differentiated visual style parameters are set for the target vehicle, the vehicle itself, surrounding vehicles, historical trajectory, candidate target lane gap area, and road forward direction in the bird's-eye view traffic context map. For the historical trajectory of the vehicle, trajectory points or trajectory lines are drawn from light to dark by display intensity weight; the candidate target lane gap area is colored differently from the lane; and the road forward direction is displayed using directional arrows or text labels.

[0011] As a preferred embodiment of the present invention, in step 3, based on the road environment information and high-precision map matching results obtained in step 1, and the road forward direction identifier obtained in step 2, a scene basic information description is generated. The scene basic information includes road type, lane type, speed limit information, weather information, time information, lane center line, lane boundary, and road forward direction. Based on the vehicle motion state data obtained in step 1, a state description of TV, EGO, and CIPV at the current moment, as well as the lateral movement trend and speed trend of the target vehicle, are generated. Based on all the generated data and the obtained Cut-in scene interaction relationship feature set, a structured scene description feature file is generated.

[0012] As a preferred embodiment of the present invention, step 4 involves constructing the task background and prediction target prompts. It is used to prompt visual language models for task understanding; it constructs instructions that use constraints and require no truth leakage. It is used to tell the visual language model what information it can and cannot see; to construct visual legends and image element interpretation instructions. It assists the VLM model in scene understanding.

[0013] As a preferred embodiment of the present invention, step 5 involves hiding the state from the intermediate layer. Extract the preset layer set from the middle The features of the two layers are then used to extract the hidden representations of the image tokens in layers 48 and 96 based on the image token location set. , spliced ​​together Based on the valid length of the text token sequence, extract the hidden representation of the last valid text token at layers 48 and 96 corresponding to the current traffic scenario. , spliced ​​together .

[0014] As a preferred embodiment of the present invention, in step 6, the interaction features are concatenated with the dense features related to the current traffic scene task, and then mapped using a multi-layer perceptron. Then, a multi-head self-attention mechanism is used for fusion to obtain the decoded features. After that, the trajectory prediction decoder is used to decode and obtain the future trajectory of the target vehicle.

[0015] As a further preferred embodiment of the present invention, step 2 displays the intensity weight. Represented as: ; In the formula, Representing historical moments The corresponding trajectory displays the intensity; This indicates the display intensity corresponding to the earliest historical moment; This indicates the display intensity in the vicinity of the current moment; Indicates the length of the historical observation time window.

[0016] Advantages and beneficial effects of the present invention: (1) This invention is aimed at high-risk interaction scenarios of highway cut-in. It unifies the historical motion state of the target vehicle, the vehicle itself, CIPV and surrounding traffic participants, lane geometry, candidate cut-in gap (candidate target lane gap area), road direction of travel and key interaction constraints into a traffic context graph. This enables the originally scattered trajectory information, road information and vehicle interaction information to be expressed in an intuitive and structured way, which is conducive to enhancing the model's ability to understand the overall spatial relationship and traffic semantic relationship of the cut-in scenario.

[0017] (2) This invention explicitly describes the current state of the target vehicle, the vehicle, and the CIPV, the historical lateral movement trend and speed change trend of the target vehicle, as well as the longitudinal distance and collision time between the TV, EGO, and CIPV through a structured scene description feature file. This can supplement the high-level traffic semantic information that is difficult to fully express by simple image input or simple numerical trajectory input, thereby improving the model's ability to judge the target vehicle's cutting intention, the feasibility of candidate cutting gaps, and potential collision risks.

[0018] (3) This invention does not directly use the natural language answers, discrete intent labels, or pseudo-labels generated by the visual language model. Instead, it directly extracts dense features from the intermediate hidden layers during the forward propagation of the visual language model and combines them with image token average pooling, max pooling, and effective text token representations to form task-related dense features. This approach can reduce information compression and semantic loss caused by the natural language decoding process, enabling the multimodal understanding capabilities formed within the visual language model to participate more fully in subsequent continuous trajectory prediction tasks.

[0019] (4) The present invention maps the task-related dense features extracted by the visual language model to the latent space of the trajectory prediction model and integrates them with historical motion features, road topology features and vehicle interaction features. This enables the trajectory prediction model to not only utilize the kinematic and interaction features obtained by the traditional trajectory encoder, but also to introduce the understanding results of the traffic context graph and structured scene semantics of the visual language model, thereby improving the completeness and discriminability of the predicted feature expression.

[0020] (5) The present invention sets the task background, prediction target, information usage constraints and no future truth leakage requirement in the prompt words, so that the visual language model can understand the scene only based on historical trajectory, current frame state, lane geometry information and known interaction indicators, avoiding the introduction of future real trajectory, future action label or future lane label, which is conducive to ensuring the rationality, fairness and engineering usability of the training and inference process.

[0021] (6) The present invention can integrate spatial geometric relationships, vehicle interaction relationships and high-level traffic semantic information into the trajectory prediction process, thereby improving the accuracy, robustness and interpretability of vehicle trajectory prediction in highway cut-in scenarios.

[0022] (7) The visual language dense features extracted by the present invention can be used in combination with existing mature trajectory prediction encoders and decoders, and are compatible with various trajectory prediction model structures such as Transformer, LSTM, GRU, and MLP. They do not depend on a specific single prediction network and have good model compatibility and engineering scalability.

[0023] (8) The present invention outputs the future single-modal or multi-modal predicted trajectory of the target vehicle and its probability, which can provide a more reliable prediction basis for the collision risk assessment, behavior decision and path planning of autonomous vehicles in the Cut-in scenario, and help improve the driving safety, decision stability and risk prediction ability of autonomous vehicles in the highway lane change cut-in scenario. Attached Figure Description

[0024] Other objects and results of the invention will become more apparent and readily understood with reference to the following description taken in conjunction with the accompanying drawings. In the drawings: Figure 1 The present invention provides a traffic semantic scene diagram for representing the spatial relationships between the target vehicle (TV), the autonomous vehicle (EGO), surrounding traffic participants, lane centerline, lane boundary, historical movement trajectory, and road direction of travel in a highway cut-in scenario. Figure 2 This invention provides a typical high-speed cut-in scenario diagram to illustrate the positional and interactive relationships between the target vehicle (TV), the vehicle (EGO), the nearest vehicle ahead in the vehicle's lane (CIPV), and the road speed limit information when the target vehicle (TV) cuts into the lane where the vehicle (EGO) is located from an adjacent lane. Figure 3 The flowchart shows the vehicle trajectory prediction method for Cut-in scenarios based on traffic context graph and VLM provided by this invention. Detailed Implementation

[0025] To enable those skilled in the art to better understand the technical solutions and advantages of the present invention, the present application will be described in detail below with reference to the accompanying drawings, but this is not intended to limit the scope of protection of the present invention.

[0026] like Figures 1 to 3 As shown, this embodiment provides a method for predicting vehicle trajectories in a cut-in scenario based on a traffic context graph and a vehicle-to-everything (VLM). The method includes the following steps: Step 1. Construction of basic data for highway cut-in scenarios, including traffic data acquisition, high-precision map matching, and determination of interactive vehicle roles; Specifically, step 1 includes the following steps: Step 1.1. Acquisition of traffic participant motion status data: Obtain the target vehicle in the highway cut-in scenario ( ), bicycle ( and surrounding traffic participants during historical observation time windows The motion state data within the data form a historical motion state set. .

[0027] Furthermore, in this embodiment, the target vehicle in the highway Cut-in scenario is obtained through the perception system, positioning system, vehicle-to-infrastructure (V2I) system, vehicle control system, and vehicle control system mounted on the autonomous vehicle. ), bicycle ( (and the motion data of surrounding traffic participants within the historical observation time window).

[0028] Suppose there are a total of For each traffic participant, the historical observation time window length is... Then the set of historical motion states of traffic participants can be represented as:

[0029] Among them, the Traffic participants The historical state of motion can be represented as:

[0030] In the A historic moment, vehicles motion state vector Represented as:

[0031] In the formula, Indicates the vehicle at time t Position coordinates in the current vehicle coordinate system or unified map coordinate system; Indicates the vehicle at time t Velocity components along the x and y axes; Indicates the vehicle at time t Acceleration components along the x and y axes; Indicates the vehicle's heading angle; These represent the vehicle's length and width, respectively.

[0032] Step 1.2. Obtaining Road Environment Information and Matching it with High-Precision Maps: By using the positioning system and high-precision map system of autonomous vehicles, the road environment information around the current Cut-in scene is obtained, and the positions of traffic participants obtained in step 1.1 are matched with the high-precision map to obtain the lane, adjacent lanes, lane center line, lane boundary and road topology relationship of each vehicle.

[0033] High-precision map information can be represented as:

[0034] In the formula, Represents the set of lane centerlines; Represents the set of lane boundaries; Indicates the road topology connections; The road semantic information includes road type, lane type, speed limit information, weather, time, intersection information, and special road target information.

[0035] Step 1.3. Cut-in prediction object and candidate target lane determination: The target vehicle in the highway cut-in scenario is denoted as... Autonomous vehicles are referred to as self-driving cars. The vehicle in front of the vehicle in the lane is designated as CIPV, and the lane where EGO is located is designated as the candidate target lane for TV.

[0036] In this embodiment, the target vehicle is a vehicle with cut-in intent. During actual operation, the upstream perception module determines cut-in candidate vehicles based on the positional relationship between the vehicle and surrounding vehicles. Specifically, vehicles in lanes adjacent to the vehicle's lane are used as cut-in candidate vehicles. Then, by combining the longitudinal distance, lateral positional relationship, and relative motion state between the cut-in candidate vehicle and the vehicle and CIPV, it is determined whether the cut-in candidate vehicle has cut-in intent. If it has cut-in intent, it is used as the target vehicle, and the future driving trajectory is predicted based on the trajectory prediction method provided by this invention.

[0037] It should be noted that those skilled in the art can refer to existing technologies to determine the cut-in intention, which is not an improvement of this application. The key point of this application is to directly output the future driving trajectory of the target vehicle with the cut-in intention output by the upstream sensing module.

[0038] Step 1.4. Calculation of Cut-in Interaction Relationship Features: Based on the vehicle motion state, high-precision map matching results, and interactive vehicle roles obtained in step 1.1, calculate... , , The longitudinal distances and collision times between the three types of key vehicles are used to obtain a set of interaction relationship features for the Cut-in scene. .

[0039] Specifically, the Cut-in scene interaction relationship features include: and longitudinal distance and collision time , and longitudinal distance and collision time ,as well as and Longitudinal distance and collision time .

[0040] In a single combination , Taking the calculation as an example, , The longitudinal distance at time t can be expressed as:

[0041] , The collision time (TTC) at time t can be expressed as:

[0042] In the formula, , Represents the longitudinal position coordinates of the vehicle and the target vehicle at time t. This represents the longitudinal velocity of the vehicle and the target vehicle at time t.

[0043] Based on the above calculation method, the calculation is performed. , , , Thus, the time is obtained. The following is a set of interactive relationship features for the Cut-in scene:

[0044] Step 1.5. Based on the data obtained in steps 1.1 to 1.4, the final basic data set for the Cut-in scenario is obtained:

[0045] in, This represents the set of interactive relationship features in the Cut-in scene.

[0046] Step 2. Construct a traffic context graph for visual language model understanding: Step 2.1. Extracting elements for constructing the traffic context graph: Based on the basic data set of highway cut-in scenarios obtained in step 1.5 Extracting historical trajectories of traffic participants for constructing a traffic context graph. Current vehicle attitude Lane geometry Elements for constructing traffic context graphs.

[0047] Suppose there are a total of For each traffic participant, the historical observation time window length is... , No. The trajectory of a traffic participant within a historical observation window can be represented as:

[0048] In the formula, Indicates the first Traffic participants at a historical moment Location coordinates, .

[0049] No. The vehicle orientation of a traffic participant at the current moment can be represented as:

[0050] In the formula, Indicates vehicle At the current heading angle, and Representing vehicles Length and width.

[0051] The lane geometry can be represented as:

[0052] In the formula, Represents the set of lane centerlines. This represents the set of lane boundaries.

[0053] Step 2.2. Bird's-eye view coordinate mapping: The vehicle trajectory, current vehicle posture, and lane geometry obtained in step 2.1 are uniformly mapped onto the bird's-eye view plane to obtain traffic context graph drawing coordinates suitable for visual language model understanding.

[0054] Let any point in the vehicle coordinate system or road coordinate system be... Its image coordinates in the traffic context map are Then the mapping relationship can be expressed as:

[0055] In the formula, This represents a mapping function from the vehicle coordinate system or road coordinate system to the bird's-eye view image coordinate system. This mapping function is mainly determined based on the coordinate axis direction and scale ratio between the vehicle / road coordinate system and the bird's-eye view image coordinate system.

[0056] Furthermore, to ensure that the road's direction of travel remains upward in the traffic context map, the following coordinate transformation method can be used:

[0057]

[0058] In the formula, Indicates the vehicle's forward coordinates. Indicates the vehicle's left-hand coordinates. and These represent the horizontal and vertical drawing coordinates in the bird's-eye view traffic context map, respectively. This means that the road's forward direction corresponds to the vertical direction in the image, and the vehicle's lateral position corresponds to the horizontal direction in the image. Therefore, subsequent vehicle trajectories, attitudes, and lane structures will be uniformly converted to the traffic context map according to this relationship.

[0059] Step 2.3. Historical trajectory mapping of traffic participants: Based on the historical movement data of traffic participants obtained in step 1.1 and the bird's-eye view coordinate mapping results obtained in step 2.2, the historical observation time window is... The historical trajectories of each traffic participant are plotted in the traffic context map.

[0060] No. The historical trajectory mapping result of a traffic participant can be represented as:

[0061] In the formula, Indicates vehicle The set of historical trajectory points in the traffic context map.

[0062] Step 2.4. Draw the current vehicle attitude: Based on the location, heading angle, vehicle length, and vehicle width of the traffic participants obtained in step 1.1, the target vehicle... bicycle And surrounding traffic participants drawn as bird's-eye view vehicle frames .

[0063] For the A traffic participant, whose current bird's-eye view vehicle frame can be represented as:

[0064] In the formula, Indicates vehicle Bird's-eye view of vehicles in the traffic context map.

[0065] In this embodiment, the vehicle The bird's-eye view of the vehicle frame is from the current location. Heading angle Vehicle length and vehicle width Together, they determine the spatial position and orientation of the vehicle in the current Cut-in scene.

[0066] Step 2.5. Lane geometry drawing: Based on the high-precision map information obtained and matched in step 1.2 The center lines of the lanes surrounding the current Cut-in scene. and lane boundary Map and plot to the bird's-eye view traffic context graph.

[0067] The lane centerline mapping result can be represented as:

[0068] The lane boundary mapping result can be represented as:

[0069] In the formula, This represents the lane centerline mapped to the traffic context graph. This represents the lane boundaries mapped to the traffic context graph.

[0070] By drawing lane centerlines and lane boundaries, the traffic context map can explicitly represent the target vehicle. bicycle This includes the road structure and lane constraints of surrounding traffic participants. Furthermore, the direction of travel on the road can be determined based on the drawn lane centerline.

[0071] Step 2.6. Setting the visual style for key areas and traffic context map: The future trajectory of the target vehicle is influenced by its own motion state, the gap between vehicles in front and behind in the target lane, its relative position, and the nearest vehicle on the path ahead. In this embodiment, a key spatial region is defined, which includes the candidate target lane gap region. The candidate target lane gap area The candidate target lane gap area is determined based on the positions of adjacent vehicles in the candidate target lane, primarily defined by the vehicles before and after the gap and the lane lines. Specifically, the candidate target lane gap area is determined based on the high-precision map matching results and the positional relationships of key vehicles. That is, the lane where the EGO is located is used as the candidate target lane for the TV, and the longitudinal area between the EGO and its nearest on-path vehicle CIPV is defined as the candidate target lane gap area. .

[0072] Furthermore, in this embodiment, the target vehicle determined in step 1.3 is... and bicycle In addition to the surrounding traffic participant information obtained in step 1.1, differentiated visual style parameters are set for the target vehicle, the vehicle itself, surrounding traffic participants (surrounding vehicles), historical trajectory, candidate target lane gap area and road direction of travel in the traffic context map, so as to enhance the visual language model's ability to understand different traffic participants and key spatial areas in the Cut-in scene.

[0073] Specifically, in this embodiment, the target vehicle First-person view style Mark the vehicle Adopting a second visual style The signage is provided, and surrounding road users are encouraged to use a third-vision style. Display it. For the first... The visual style of a traffic participant can be represented as follows:

[0074] In the formula, Indicates the first Visual styles of traffic participants in a traffic context map; Indicates the visual style of the target vehicle; Indicates the visual style of the vehicle; It indicates the visual style of surrounding traffic participants.

[0075] The visual style can be determined by display color, outline style, transparency, vehicle label, and drawing level, and can be represented as follows:

[0076] In the formula, Indicates vehicle The display color, Indicates the vehicle's outline style. Indicates transparency. Indicates vehicle label. Indicates the drawing level.

[0077] For the target vehicle Set its label as:

[0078] For bicycles Set its label as:

[0079] For surrounding traffic participants, no specific interactive role labels are set; they can be drawn using the standard vehicle display style.

[0080] To represent the evolution of vehicle trajectories over time, the historical trajectories of traffic participants are analyzed by displaying intensity weights. The trajectory points or lines are drawn from light to dark to represent the temporal evolution relationship.

[0081] Set historical moment The corresponding trajectory display intensity weight is Then it can be expressed as:

[0082] In the formula, Representing historical moments The corresponding trajectory displays the intensity; This indicates the display intensity corresponding to the earliest historical moment; This indicates the display intensity in the vicinity of the current moment; Indicates the length of the historical observation time window. When... When the size is small, the trajectory is displayed lightly, and as... As the trajectory deepens, it represents the evolution of traffic participants' movements from historical moments to the current moment.

[0083] To more clearly represent the gap area of ​​the candidate target lane, in this embodiment, the gap area of ​​the candidate target lane is colored differently from the lane.

[0084] Direction of the road Displayed using directional arrows or text icons, denoted as:

[0085] In the formula, This indicates the location of the road direction sign within the traffic context map. Indicates the direction angle of the road's movement. Textual signs indicating the direction of travel on a road.

[0086] Furthermore, in this embodiment, the visual styles of different objects are also set differently. For example, the target vehicle, the vehicle itself, surrounding vehicles, historical trajectories, and candidate target lane gap areas are represented by different colors, line types, or transparency, which can more clearly highlight the key vehicles and key areas in the Cut-in scene.

[0087] Step 2.7. Traffic Context Graph Generation: The historical trajectories of traffic participants obtained in step 2.3, the current vehicle posture obtained in step 2.4, the lane centerline, lane boundaries, and road direction markers obtained in step 2.5, and the candidate target lane gap regions and visual style parameters obtained in step 2.6 are combined and drawn to generate a traffic context graph for visual language model understanding. .

[0088] Specifically, the traffic context graph can be represented as:

[0089] In the formula, This represents the image drawing function; Indicates the first The historical trajectory of each traffic participant is mapped onto the bird's-eye view; Indicates the first A bird's-eye view of the vehicle frame for each traffic participant at the current moment; This represents the lane centerline mapped to the traffic context graph; This represents the lane boundaries mapped to the traffic context graph; Road direction signs; Indicates the first The visual style corresponding to each traffic participant.

[0090] The final generated traffic context graph can be represented as:

[0091] In the formula, and These represent the height and width of the traffic context map, respectively. This indicates the number of color channels in the image.

[0092] Thus, a bird's-eye view traffic context map corresponding to the current highway cut-in scene is obtained. The traffic context map displays the target vehicle, the vehicle itself, surrounding vehicles, historical trajectories, lane geometry, candidate target lane gap areas, and road direction of travel through differentiated visual styles, providing an image foundation for subsequent generation of structured scene description feature files and visual language models to understand the cut-in traffic context.

[0093] Step 3. Constructing a structured scene description feature file: Based on the basic data set of highway cut-in scenarios obtained in step 1.5 and the traffic context graph obtained in step 2.7 Generate traffic context Figure 1 A corresponding structured scene description feature file It is used to describe in text form the vehicle status, road environment status, target vehicle movement trend, and key interaction constraint information in the current Cut-in scene.

[0094] Let the first The traffic context graph corresponding to each Cut-in scenario is as follows: The corresponding structured scene description feature file is Then the one-to-one correspondence between the two can be expressed as:

[0095] In the formula, This represents the function for generating structured scene description feature files; Indicates the first The basic data set for each Cut-in scenario; Indicates the first Traffic context graph corresponding to each Cut-in scenario; Indicates the first The structured scene description feature file corresponding to each Cut-in scene.

[0096] Specifically, step 3 includes the following steps: Step 3.1. Generation of basic scene information description: Based on the road environment information and high-precision map matching results obtained in step 1.2, and the road forward direction indicator obtained in step 2.7. Generate basic scene information description The basic information of the scenario includes road type, lane type, speed limit information, weather information, time information, lane center line, lane boundary, road direction of travel, and intersection or special road target information.

[0097] Specifically, the basic information of the scene can be represented as follows:

[0098] In the formula, Indicates the road type; Indicates lane type; Indicates the road speed limit; Indicates weather information; Indicates time information; Represents the set of lane centerlines; Represents the set of lane boundaries; Indicates the direction of travel; This indicates target information at intersections or on special roads.

[0099] Step 3.2. Generation of current state descriptions for the target vehicle and the driver vehicle: The target vehicle is obtained based on the vehicle motion state data acquired in step 1.1. bicycle In addition to the nearest vehicle in front of the vehicle, CIPV, generate the state descriptions of TV, EGO, and CIPV at the current moment. , and .

[0100] For vehicles Its current state can be described as follows:

[0101] In the formula, Indicates vehicle The lane matched at the current moment; Indicates vehicle The vertical coordinate; Indicates vehicle The lateral coordinates relative to the lane centerline; Indicates vehicle Lateral velocity, longitudinal velocity; Indicates vehicle The heading angle; and Representing vehicles Length and width.

[0102] Therefore, the target vehicle The current state can be described as follows:

[0103] bicycle The current state can be described as follows:

[0104] The vehicle in front of you (the nearest vehicle on your path) The current state can be described as follows:

[0105] Step 3.3. Generation of historical motion trend description of the target vehicle: Based on the target vehicle obtained in step 1.1 In the historical observation time window The motion state data (lateral position change and longitudinal velocity change) within the vehicle are used to generate the lateral movement trend of the target vehicle. and speed trend .

[0106] Target vehicle The lateral displacement change within the historical observation window can be expressed as:

[0107] In the formula, Indicates the start time of the historical observation time window Horizontal position; Indicates the current time Horizontal position; express The amount of lateral displacement change within the historical observation time window.

[0108] Target vehicle The lateral movement trend can be represented as:

[0109] In the formula, Indicates the target vehicle The horizontal movement trend; This indicates the threshold for determining lateral displacement. It represents stable driving.

[0110] Target vehicle The longitudinal velocity change within the historical observation window can be expressed as:

[0111] In the formula, Indicates the start time of the historical observation time window longitudinal velocity; Indicates the current time longitudinal velocity; express The change in longitudinal velocity.

[0112] Target vehicle The speed trend can be expressed as:

[0113] In the formula, Indicates the target vehicle The speed trend; The threshold for judging speed change is indicated. Represents an accelerated state. This represents a deceleration state. This represents a uniform steady state.

[0114] Step 3.4. Generation of Cut-in Key Interaction Constraint Indicators: Based on the Cut-in interaction feature set obtained in step 1.4 ,generate , , Description of longitudinal distances and collision times between the three types of key vehicles.

[0115] Step 3.5. Output the structured scene description feature file: Based on basic scene information Current status of the target vehicle Current status of the vehicle Current status of the vehicle in front of you Lateral movement trend of the target vehicle and speed trend And the longitudinal distance and collision time description between key vehicles (or the Cut-in key interaction constraint Cut-in scene interaction relationship feature set obtained in step 1.4 can be used directly). Generate structured scene description feature files. .

[0116] Specifically, the structured scene description feature file can be represented as:

[0117] In the formula, This describes the basic information of the scene. This describes the current state of the target vehicle. This describes the current status of the vehicle. Indicates the vehicles ahead in the lane. A description of the current state; express The horizontal movement trend; express The speed trend; Indicates by , , The set of interaction relationship features consisting of the longitudinal distance and collision time between them.

[0118] Thus, we obtain the traffic context diagram. One-to-one structured scene description feature file This provides a text input foundation for subsequent visual language models to understand the road environment, target vehicle status, and key interaction constraints in highway cut-in scenarios.

[0119] Step 4. Construct Cut-in scene task prompts for visual language models: Step 4.1. Construct task background and prediction target prompts It is used to prompt task understanding of the Visual Language Model (VLM).

[0120] In this embodiment, the task background and prediction target prompt words It can be represented as: You will receive a bird's-eye view of the traffic context for the lane change trajectory prediction task, and your goal is to infer the target vehicle. The most likely motion state within the next 5 seconds. This scenario focuses on... Will it cut in? Are the available lanes and candidate gaps sufficient to accommodate the situation? ,as well as Is it possible to merge with [other lanes] after completing the lane change? Or the car ahead on the shortest path This poses a risk of collision.

[0121] Step 4.2. Construct information usage constraints and truth-leaking-free requirement instructions. This tells the VLM model which information it can and cannot see, thus preventing future truth leaks.

[0122] In this embodiment, the information uses constraints and truth-leaking-free instructions. It can be represented as: The model may only use visible historical trajectories, current frame states, lane geometry information, and provided interaction metrics for inference. The model may not use future real-world trajectory labels, future action labels, or future lane labels, nor may it fabricate precise future coordinates.

[0123] Step 4.3. Constructing visual legends and image element explanation instructions It assists the VLM model in scene understanding.

[0124] In this embodiment, the visual legend and image element interpretation instructions It can be represented as: The target vehicles that need to be predicted are shown in red and labeled "TV".

[0125] It is a bicycle, displayed in blue and marked "EGO".

[0126] Other vehicles are displayed in gray without specifying their roles.

[0127] Gradient dots and lines represent historical trajectories; the darker the color, the newer the dot and the closer it is to the current frame.

[0128] The upward direction in the diagram indicates the direction the road is moving forward.

[0129] Step 5. Extract task-relevant dense features directly from the visual language model: Step 5.1. Deploy the visual large language model locally to obtain... Local visual language model.

[0130] Furthermore, the visual language model includes, but is not limited to, Tongyi Qianwen, DeepSeek, and other visual language models (VLMs). In this embodiment, Qianwen is specifically used.

[0131] Step 5.2. Transfer the traffic context graph obtained in Step 2. The structured scene description feature file obtained in step 3 The Cut-in scene task prompts obtained in step 4 The text and image inputs are fed into the local visual language model to construct a task-conditional image-text joint input. , can be represented as:

[0132] Step 5.3. Use the image and text processor corresponding to the visual language model. Combined text and image input for the current traffic scenario Preprocessing is performed to obtain the multimodal token sequence corresponding to the current traffic scenario. , can be represented as:

[0133] In the formula, This indicates a text and image processor used for processing traffic context graphs. Perform image preprocessing and analyze the task description text. , , , Perform template-based processing and word segmentation encoding; Represents a multimodal token sequence. Indicates the first An image token or a text token. This indicates the length of the multimodal token sequence corresponding to the current traffic scenario.

[0134] Step 5.4. Obtain the multimodal token sequence corresponding to the current traffic scenario. Input local visual language model The forward propagation is performed, and the intermediate hidden states of the visual language model are output during the forward propagation process. This process does not invoke the text generation module, nor does it require the visual language model to generate natural language answers, predicted labels, or pseudo-labels. Instead, it only acquires the internal multimodal hidden representations formed by the visual language model when understanding the current combined text and image input.

[0135] In the formula, The visual language model is represented by the first The hidden state obtained from the current traffic scenario input is layer 1. This indicates the number of layers in the visual language model. This represents the hidden state dimension.

[0136] Step 5.5. Hide the state from the intermediate layer Select a preset layer set The corresponding hidden state is determined, and the hidden representation of the image token is extracted based on the image token location set. The last valid text token hidden representation corresponding to the current traffic scenario .

[0137] Specifically, in this embodiment, from Extract the preset layer set from the middle The features of the two layers are then used to extract the hidden representations of the image tokens in layers 48 and 96 based on the image token location set. , spliced ​​together Based on the valid length of the text token sequence, extract the hidden representation of the last valid text token at layers 48 and 96 corresponding to the current traffic scenario. , spliced ​​together .

[0138] Step 5.6. Hide the image token state for the selected layer. Max pooling and average pooling are performed separately to obtain the image token hidden state max pooling. and average pooling It can be represented as:

[0139]

[0140] In the formula, This represents the average pooling features of the image tokens in the current traffic scene, used to characterize the overall visual context. This represents the current traffic scene's max-pooling features in the image token model, used to characterize significant visual responses. Represents max pooling. This represents average pooling.

[0141] Step 5.7. Average pool the image token features. Image token max pooling feature Hidden representation with the last valid text token The data is then stitched together to obtain task-relevant dense features. , can be represented as:

[0142] In the formula, This represents the task-related dense features of the current traffic scene extracted from the local visual language model. These task-related dense features simultaneously include spatial structure information in the traffic context image, traffic semantic information in the task description text, candidate lane merging gap information, interaction risk information, and task constraint information without future truth leakage.

[0143] Step 6. Trajectory Prediction: Step 6.1. Based on the historical motion state set from Step 1.1 and the high-precision map information obtained in step 1.2 Interactive features are obtained using existing mature trajectory prediction encoders. , can be represented as:

[0144] In the formula, Encoders representing trajectory prediction models include, but are not limited to, map encoders, trajectory encoders, and interactive encoders.

[0145] Step 6.2. The interaction features obtained in Step 6.1 And the dense features related to the current traffic scenario task in step 5.7 After concatenation, an MLP is used for mapping, and then a multi-head self-attention mechanism is used for fusion to obtain the decoded features. , can be represented as:

[0146] In the formula, This represents a multi-head self-attention mechanism. MLP stands for Multilayer Perceptron, which consists of two linear layers and a ReLU activation function layer.

[0147] Step 6.3. Decode the future trajectory of the target vehicle using an existing, mature trajectory prediction decoder. , can be represented as:

[0148] In the formula, This refers to trajectory prediction decoders, including but not limited to Transformer, LSTM, GRU, MLP decoders, etc.

[0149] To verify the technical advantages of the method described in this invention, this invention was trained and validated on a large cut-in dataset (ESP dataset), and compared with MTR, DenseTNT, HPNet, SIMPL and advanced DeMo prediction methods. The test results of each method are shown in Table 1.

[0150] Table 1 shows the validation results for a large cut-in dataset. MTR 0.25 2.03 3.48 0.46 DenseTNT 0.24 2.14 3.24 0.41 HPNet 0.24 2.11 3.32 0.43 SIMPL 0.23 2.25 3.56 0.45 DeMo 0.28 1.62 2.83 0.39 This invention (ours) 0.31 1.22 2.05 0.25 The comparative analysis of experimental data in Table 1 shows that the method provided by this invention significantly outperforms mainstream trajectory prediction algorithms, achieving the highest mAP6 value while miniADE6, miniFDE6, and MR6 are all at their lowest levels. Compared to the suboptimal baseline DeMo, the average accuracy mAP6 of the multi-trajectory system is improved by 10.7%, the average displacement error minADE6 of the best trajectory among the six trajectories is reduced by 24.7%, the endpoint error minFDE6 of the best trajectory among the six trajectories is reduced by 27.6%, and the failure rate (false positive rate) MR6, where none of the best trajectories hit the ground truth, is reduced by 35.9%. These results fully demonstrate that this method can generate higher-quality multimodal candidate trajectories, more accurately capture the future movement patterns of traffic participants, and effectively reduce trajectory prediction errors and the probability of prediction failure.

[0151] The present invention also provides an electronic device, comprising: one or more processors and a memory; wherein the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the above-described method for predicting vehicle trajectories in cut-in scenarios based on traffic context graphs and VLMs.

[0152] The present invention also provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method for predicting vehicle trajectories in cut-in scenarios based on traffic context graphs and VLMs.

[0153] Those skilled in the art will understand that all or part of the functions of the various methods / modules in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the above functions can be implemented by executing the program with a computer. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be implemented.

[0154] In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the programs can also be stored in storage media such as servers, other computers, disks, optical discs, flash drives, or portable hard drives. They can be downloaded or copied to the memory of the local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be implemented.

[0155] The above-described specific examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art can make various simple deductions, modifications, or substitutions based on the principles of this invention. Therefore, the scope of protection of this invention should be determined by the scope of the claims.

Claims

1. A method for predicting vehicle trajectories in cut-in scenarios based on traffic context graphs and VLM, characterized in that, The method includes the following steps: Step 1. Obtain historical motion state data of each vehicle in the highway cut-in scenario, as well as surrounding road environment information, and combine this with high-precision map information to determine key interactive vehicles; the key interactive vehicles include the target vehicle. bicycle The vehicle in front of the bicycle lane The target vehicle is a vehicle intended for cut-in. Step 2. Map and draw the vehicle's historical trajectory, current vehicle posture, lane geometry, candidate target lane gap area, and road direction of travel onto the traffic context graph, and set different visual styles for different objects. Step 3. Based on the data obtained in Step 1, generate a structured scene description feature file corresponding to the traffic context graph; Step 4. Construct Cut-in scene task prompts for visual language models; Step 5. Construct a combined image and text input based on the data from Steps 2 to 4. The image processing unit corresponding to the visual language model is used for preprocessing, followed by forward propagation. During forward propagation, the intermediate hidden states of the visual language model are output. Then, a preset set of layers is selected from the intermediate hidden states. The corresponding hidden state is determined, and the hidden representation of the image token is extracted. The last valid text token hidden representation corresponding to the current traffic scenario ; after that Max pooling and average pooling are performed separately; finally, the image tokens are average pooled for feature extraction. Image token max pooling feature and The features are then spliced ​​together to obtain task-relevant dense features. Step 6. Encode the historical motion state and high-precision map information to obtain the interaction features; stitch the interaction features with the task-related dense features, and then perform mapping, fusion and decoding to obtain the future trajectory of the target vehicle.

2. The method for predicting vehicle trajectories in a cut-in scene based on traffic context graph and VLM as described in claim 1, characterized in that, The historical motion state data in step 1 includes the vehicle's position coordinates, speed, acceleration, heading angle, vehicle length, and width.

3. The method for predicting vehicle trajectories in a cut-in scene based on traffic context graph and VLM as described in claim 2, characterized in that, After identifying the key interaction vehicles in Step 1, a set of Cut-in scene interaction relationship features is constructed based on these vehicles. The Cut-in scene interaction relationship features include: and The longitudinal distance and collision time, and The longitudinal distance and collision time, and and The longitudinal distance and collision time.

4. The method for predicting vehicle trajectories in a cut-in scene based on traffic context graph and VLM as described in claim 3, characterized in that, Step 2 will include historical observation time windows. The historical trajectories of each vehicle are mapped and plotted onto the bird's-eye view traffic context map; based on the vehicle's current position, heading angle, vehicle length, and vehicle width obtained in step 1, the target vehicle is... bicycle The surrounding vehicles are drawn as bird's-eye view vehicle frames; the lane center lines and lane boundaries around the current cut-in scene are mapped and drawn onto the bird's-eye view traffic context map, and the road forward direction indicator is determined according to the direction of the drawn lane center lines; at the same time, the longitudinal area between the vehicle's EGO and the nearest vehicle on the path in front is determined as the candidate target lane gap area.

5. The method for predicting vehicle trajectories in a cut-in scene based on traffic context graph and VLM according to claim 4, characterized in that, In step 2, differentiated visual style parameters are set for the target vehicle, the vehicle itself, surrounding vehicles, historical trajectory, candidate target lane gap area and road direction of travel in the bird's-eye view traffic context map. For the historical trajectory of the vehicle, trajectory points or trajectory lines are drawn from light to dark by display intensity weight. The gap area between candidate target lanes is colored differently from the lane; The direction of travel on the road is indicated by directional arrows or text labels.

6. The method for predicting vehicle trajectories in a cut-in scene based on traffic context graph and VLM according to claim 5, characterized in that, In step 3, based on the road environment information and high-precision map matching results obtained in step 1, and the road forward direction identifier obtained in step 2, a basic scene information description is generated. The basic scene information includes road type, lane type, speed limit information, weather information, time information, lane center line, lane boundary, and road forward direction. Based on the vehicle motion state data obtained in step 1, a state description of TV, EGO, and CIPV at the current moment, as well as the lateral movement trend and speed trend of the target vehicle are generated. Based on all the generated data and the obtained Cut-in scene interaction relationship feature set, a structured scene description feature file is generated.

7. The method for predicting vehicle trajectories in a cut-in scene based on traffic context graph and VLM as described in claim 6, characterized in that, Step 4 involves constructing the task background and prediction target prompts. It is used to prompt visual language models for task understanding; it constructs instructions that use constraints and require no truth leakage. It is used to tell the visual language model what information it can and cannot see; to construct visual legends and image element interpretation instructions. It assists the VLM model in scene understanding.

8. The method for predicting vehicle trajectories in a cut-in scene based on traffic context graph and VLM according to claim 7, characterized in that, In step 5, hide the state from the intermediate layer. Extract the preset layer set from the middle The features of the two layers are then used to extract the hidden representations of the image tokens in layers 48 and 96 based on the image token location set. , spliced ​​together Based on the valid length of the text token sequence, extract the hidden representation of the last valid text token at layers 48 and 96 corresponding to the current traffic scenario. , spliced ​​together .

9. The method for predicting vehicle trajectories in a cut-in scene based on traffic context graph and VLM as described in claim 8, characterized in that, In step 6, the interaction features are concatenated with the dense features related to the current traffic scene task, then mapped using a multilayer perceptron, and finally fused using a multi-head self-attention mechanism to obtain the decoded features. The future trajectory of the target vehicle is then obtained by using a trajectory prediction decoder.

10. The method for predicting vehicle trajectories in a cut-in scene based on traffic context graph and VLM according to claim 9, characterized in that, Step 2 displays the intensity weights Represented as: ; In the formula, Representing historical moments The corresponding trajectory displays the intensity; This indicates the display intensity corresponding to the earliest historical moment; This indicates the display intensity in the vicinity of the current moment; Indicates the length of the historical observation time window.

Citation Information

Patent Citations

  • Automatic driving track prediction method based on vision-language model, near-end strategy optimization and diffusion model

    CN121210643A

  • Automatic driving interpretation text determination method based on large visual language model

    CN119142366A

  • Automatic driving method and device based on large language model and diffusion model

    CN120375325A