Method for generating travel trajectory, electronic device, storage medium and program product

By utilizing environmental images and semantic features to generate driving trajectories, the problem of limited application scope of high-definition maps is solved. It enables accurate trajectory generation and dynamic information processing in any road scenario, reducing costs and improving safety.

CN120902773BActive Publication Date: 2025-12-05INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511449183.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-12-05
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing trajectory prediction technologies rely on high-definition maps, which limit their application scope and cannot cover all roads. Furthermore, high-definition maps are expensive to produce and cannot be updated with dynamic information in real time, resulting in inaccurate driving trajectories and potential safety hazards.

Method used

By acquiring environmental images and semantic feature extraction instructions for the target vehicle, a pre-trained language model is used to generate target semantic features and image features. Based on these features, a driving trajectory that meets the conditions is generated to guide vehicle driving operations, avoiding reliance on high-definition maps.

Benefits of technology

It enables the generation of accurate driving trajectories in any road scenario, reduces costs, avoids safety issues caused by untimely updates of high-definition maps, and can process dynamic environmental information in real time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120902773B_ABST
    Figure CN120902773B_ABST
Patent Text Reader

Abstract

The application discloses a driving track generation method, an electronic device, a storage medium and a program product, relates to the technical field of automatic driving, and comprises the following steps: after obtaining an environment image of a target vehicle, extraction indication information of semantic features and a basic driving track, first, inputting the environment image and the extraction indication information of the semantic features into a language model, so that the language model can extract target semantic features required for generating the driving track. Then, the target image features are directly extracted from the environment image. Further, the target semantic features and the target image features can be used as guiding conditions for track generation, and the basic driving track is adjusted to obtain a target driving track meeting target conditions. In this way, in the process of generating the driving track, the driving track is generated by using the real-time obtained environment image, rather than using a high-definition map to generate the driving track, and the application range is wider.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic driving, in particular to a driving trajectory generation method, an electronic device, a storage medium and a program product. BACKGROUND

[0002] In the technical field of automatic driving, the accuracy of trajectory prediction is directly related to driving safety. The trajectory prediction technology mainly predicts the future driving path based on the current vehicle state, surrounding environment information and driving intention, so that the current vehicle can drive based on the predicted driving path.

[0003] The current trajectory prediction technology generally depends on a high-definition map. When predicting a trajectory, the real-time captured surrounding environment information of the vehicle and the pre-stored high-definition map are fused to accurately position the vehicle on the accurate position of the high-definition map and obtain accurate traffic information, such as road topology, lane lines and traffic light positions. However, the production cost of the high-definition map is high, the maintenance cycle is long, it cannot cover all roads, and the application conditions are relatively harsh, that is, it can only be applied to road scenes where a high-definition map has been constructed. SUMMARY

[0004] The present application provides a driving trajectory generation method, device, electronic device, storage medium and program product to solve the problem of limited application range of high-definition maps.

[0005] The present application provides a driving trajectory generation method, comprising:

[0006] obtaining an environment image of a target vehicle, extraction indication information of semantic features, and a pre-generated basic driving trajectory;

[0007] inputting the environment image and the extraction indication information of the semantic features into a pre-trained language model to obtain target semantic features output by the language model;

[0008] extracting target image features from the environment image;

[0009] generating a target condition based on the target image features and the target semantic features;

[0010] generating a target driving trajectory meeting the target condition based on the target condition and the basic driving trajectory, to guide the driving operation of the target vehicle.

[0011] The present application also provides a driving trajectory generation device, comprising:

[0012] an acquisition module configured to obtain an environment image of a target vehicle, extraction indication information of semantic features, and a pre-generated basic driving trajectory;

[0013] The semantic feature extraction module is used to input environmental images and semantic feature extraction instructions into a pre-trained language model to obtain the target semantic features output by the language model.

[0014] The image feature extraction module is used to extract target image features from environmental images;

[0015] The generation module is used to generate target conditions based on target image features and target semantic features; and to generate a target driving trajectory that meets the target conditions based on the target conditions and the basic driving trajectory, so as to guide the driving operation of the target vehicle.

[0016] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described methods for generating a driving trajectory.

[0017] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described methods for generating driving trajectories.

[0018] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described methods for generating driving trajectories.

[0019] This application, after acquiring the environmental image of the target vehicle, semantic feature extraction instructions, and the basic driving trajectory, firstly, inputs the environmental image and semantic feature extraction instructions into a language model, enabling the language model to extract the target semantic features required to generate the driving trajectory. Then, target image features are directly extracted from the environmental image. Furthermore, the target semantic features and target image features can be used as guiding conditions for trajectory generation, adjusting the basic driving trajectory to obtain a target driving trajectory that meets the target conditions. Thus, by using real-time acquired environmental images to generate the driving trajectory, this method can be applied to any road scenario, broadening its application range. Attached Figure Description

[0020] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the architecture of a driving trajectory generation system provided in an embodiment of this application;

[0022] Figure 2A flowchart illustrating a method for generating a driving trajectory provided in an embodiment of this application;

[0023] Figure 3 A data flow diagram of a language model provided in an embodiment of this application;

[0024] Figure 4 This is a schematic diagram of a data flow for generating target image features, provided in an embodiment of this application.

[0025] Figure 5 A schematic diagram of the data flow of a diffusion model provided in an embodiment of this application;

[0026] Figure 6 This application provides a schematic diagram of a data stream for generating a target driving trajectory.

[0027] Figure 7 A schematic flowchart of a driving trajectory generation device provided in an embodiment of this application;

[0028] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0030] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0031] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] The method for generating driving trajectories provided in this application can be implemented by a driving trajectory generation system, such as... Figure 1 As shown, the driving trajectory generation system may include an image acquisition device, an in-vehicle terminal, and a driving trajectory generation platform.

[0033] Image acquisition devices can be surround-view cameras (including front-view, surround-view, and rear-view cameras), radar, etc., used to collect environmental images around the vehicle and transmit them to the vehicle terminal.

[0034] The vehicle-mounted terminal can be used to receive environmental images sent by the image acquisition device and generate driving trajectories, or it can send the received environmental images to the driving trajectory generation platform.

[0035] The driving trajectory generation platform can be a server, server cluster, etc., which can be used to generate driving trajectories based on environmental images sent by the vehicle terminal and send the generated driving trajectory to the vehicle terminal.

[0036] Embodiments of this application provide a method for generating a driving trajectory, which can be executed by the aforementioned vehicle-mounted terminal or driving trajectory generation platform (the following description uses a vehicle-mounted terminal as an example). Figure 2 As shown, the specific processing steps of the method for generating the driving trajectory may include:

[0037] Step S201: Obtain the environmental image of the target vehicle, the semantic feature extraction instruction information, and the pre-generated basic driving trajectory.

[0038] The environmental image can include images from multiple angles, such as front view, rear view, and side view. The image may include elements such as pedestrians, sky, traffic lights, and roads.

[0039] The semantic feature extraction instructions can include extraction prompts for various semantic types. For example, semantic types could include traffic congestion level, time period (morning, noon, evening, night, etc.), weather conditions (rain, sunny, etc.), traffic light status (red, green, yellow), longitudinal driving intention (speed, acceleration), and lateral driving intention (straight ahead, left turn). These extraction prompts guide the semantic feature extraction process. Examples include, "What are the weather conditions?" and "What is the level of traffic congestion?".

[0040] The basic driving trajectory can be a driving trajectory generated by adding noise to multiple typical driving trajectories included in the Trajectory AnchorVocabulary beforehand, or it can be a pre-constructed driving trajectory, such as a straight driving trajectory.

[0041] Specifically, the image acquisition device of the target vehicle can periodically acquire environmental images around the target vehicle and transmit them to the vehicle-mounted terminal. After receiving the environmental images, the vehicle-mounted terminal can read semantic feature extraction prompts and pre-generated basic driving trajectories from a preset storage location (which can be its own memory or the driving trajectory generation platform).

[0042] Step S202: Input the environmental image and semantic feature extraction instructions into the pre-trained language model to obtain the target semantic features output by the language model.

[0043] The language model can include a Mixture of Experts (MoE) model and a Vision-Language Model (VLM). For example, a Vision-Language Model can be QwenVL2.5, InternVL3, etc.

[0044] Specifically, the vehicle-mounted terminal can input environmental images and semantic feature extraction instructions into the language model, and the language model can extract target semantic features from the environmental images based on the extraction instructions.

[0045] Step S203: Extract target image features from the environmental image.

[0046] Specifically, the vehicle-mounted terminal can use a pre-trained image feature extraction model to extract target image features from environmental images. These target image features can include one or more types of image features, such as basic image features, bird's-eye view features, map features, etc. Basic image features can include texture, color, etc.

[0047] Step S204: Generate target conditions based on target image features and target semantic features.

[0048] Specifically, the vehicle-mounted terminal can stitch together or aggregate the target image features and target semantic features to obtain the target conditions, which can be in the form of vectors, matrices, etc.

[0049] Step S205: Based on the target conditions and the basic driving trajectory, generate a target driving trajectory that meets the target conditions to guide the driving operation of the target vehicle.

[0050] Specifically, the vehicle-mounted terminal can use the target image features and target semantic features indicated in the target conditions as a reference to adjust the basic driving trajectory to obtain a target driving trajectory that meets the target conditions. In this way, the vehicle-mounted terminal can display the target driving trajectory, allowing the driver to perform driving operations based on the target driving trajectory, or directly control the driving process of the target vehicle based on the target driving trajectory.

[0051] The method for generating a driving trajectory according to embodiments of this application, after acquiring the environmental image of the target vehicle, semantic feature extraction instructions, and a basic driving trajectory, firstly inputs the environmental image and semantic feature extraction instructions into a language model, enabling the language model to extract the target semantic features required to generate the driving trajectory. Then, the target image features are directly extracted from the environmental image. Furthermore, the target semantic features and target image features can be used as guiding conditions for trajectory generation, adjusting the basic driving trajectory to obtain a target driving trajectory that meets the target conditions. Thus, by using real-time acquired environmental images to generate the driving trajectory, the application scope is broadened.

[0052] Furthermore, due to frequent road changes—for example, temporary construction, traffic sign adjustments, or new road construction—high-definition maps may not be updated in a timely manner, leading to inaccurate driving trajectories and even serious safety issues. Additionally, high-definition maps cannot provide dynamic information such as traffic congestion ahead, temporary traffic control at an intersection, pedestrians crossing the road, or an ambulance approaching with its siren blaring, resulting in inaccurate driving trajectories and safety problems. However, this solution generates driving trajectories using real-time captured environmental images; therefore, driving decisions based on accurate trajectories can avoid such safety issues.

[0053] The production of high-definition maps is complex, requiring specialized high-precision surveying vehicles for large-scale map data collection, supplemented by extensive manual annotation and proofreading, resulting in high costs. This solution, however, eliminates the need for high-definition map maintenance, thus reducing costs.

[0054] In some alternative implementations, such as Figure 3 As shown, the language model may include a first image encoder, a text encoder, a language model decoder (LM decoder), and semantic feature generators of various semantic types. The first image encoder, text encoder, and language model decoder can be components of the aforementioned visual-language model. Accordingly, in step S202, the vehicle terminal can specifically determine the target semantic features using the following steps:

[0055] Step 1: Input the environmental image into the first image encoder to obtain the image features output by the first image encoder.

[0056] Step 2: Input the semantic feature extraction instruction information into the text encoder to obtain the text features output by the text encoder.

[0057] Step 3: Input the image features and text features output by the first image encoder into the language model decoder to obtain the fused features output by the language model decoder.

[0058] Step four: Input the fused features into multiple semantic feature generators to obtain semantic features output by each semantic feature generator corresponding to its own semantic type.

[0059] Step 5: Concatenate the semantic features corresponding to the various semantic types to obtain the target semantic features.

[0060] The first image encoder, also known as the vision encoder, is typically based on a neural network architecture (Transformer) with an attention mechanism, such as the Visual Transformer (ViT). The image features output by the first image encoder can be a basic image feature.

[0061] Semantic feature generators of various semantic types can all be Mixture of Experts (MoE) output heads, including Traffic Congestion Head, Time of Day Head, Weather Condition Head, Traffic Light Status Head, Longitudinal Action Head, and Lateral Action Head. The Traffic Congestion Head can analyze vehicle density, speed, and other information in an image to determine the degree of traffic congestion. The Time of Day Head can determine the current weather conditions based on visual cues such as light, raindrops, snowflakes, or fog in the image. The Traffic Light Status Head can identify traffic lights in an image and determine their current status. The Longitudinal Action Head can infer the vehicle's longitudinal driving intention, such as accelerating, decelerating, or maintaining the current speed. The Lateral Action Head can infer the vehicle's lateral driving intention, such as going straight, turning left, or turning right.

[0062] Specifically, in step one, the vehicle-mounted terminal can input the environmental image into the first image encoder, which encodes the environmental image to obtain the image features output by the first image encoder. Alternatively, when the environmental image includes multiple environmental sub-images from different angles, the vehicle-mounted terminal can input each environmental sub-image into the first image encoder separately to obtain the sub-image features output by the first image encoder corresponding to each environmental sub-image. Then, the vehicle-mounted terminal can stitch together all the sub-image features to obtain the image features corresponding to the environmental image, which is the image features output by the first image encoder.

[0063] In step two, the vehicle terminal can input the semantic feature extraction instruction information into the text encoder. The text encoder encodes the extraction instruction information to obtain text features. These text features encode the semantic information of the question and can provide guidance for the subsequent language model decoder.

[0064] In step three, the vehicle terminal can concatenate the text features with the image features determined in step one and input the result into the language model decoder. For example, the text features can be concatenated after the image features. The language model decoder can extract and fuse image information related to the text features from the image features based on the text features, obtaining the fused features output by the language model decoder. This fused feature can include information corresponding to all semantic types indicated in the extracted instruction information.

[0065] In step four, the vehicle terminal can input the fused features into multiple semantic feature generators to obtain semantic features output by each generator corresponding to its own semantic type. For example, the traffic congestion output head can output traffic congestion level features, the time output head can input time period features, the weather condition output head can input weather condition features, the traffic light status output head can output traffic light status features, the longitudinal motion output head can output the longitudinal motion features of the target vehicle, and the lateral motion output head can output the lateral motion features of the target vehicle.

[0066] Step 5: The vehicle terminal can concatenate all the semantic features to obtain target semantic features that include semantic information of multiple semantic types.

[0067] Generally, processing images solely based on convolutional neural networks or Transformer models is insufficient for comprehensive and high-precision perception of complex road environments. For example, in low-light conditions such as nighttime, rain, snow, or fog, or in situations with occlusion and glare, image quality deteriorates, making it difficult for models to accurately identify road boundaries, lane lines, drivable areas, and the location of dynamic obstacles, resulting in significant errors in the generated driving trajectory. This solution, however, uses semantic feature extraction prompts as guidance. It extracts various key semantic information from environmental images to obtain target semantic features. Furthermore, it uses these target semantic features, rich in semantic information (e.g., weather conditions, time of day, traffic congestion), to generate the target driving trajectory, improving the accuracy of driving trajectory generation in map-free scenarios.

[0068] In some optional implementations, in step S203 above, the vehicle-mounted terminal can sequentially extract image features of different dimensions from the environmental image, such as basic image features, bird's-eye view features, and map features. The environmental image may include environmental sub-images corresponding to different shooting angles. Accordingly, such as... Figure 4 As shown, the vehicle-mounted terminal can extract target image features from environmental images using the following specific steps:

[0069] Step 1: Input each environmental sub-image into the pre-constructed second image encoder to obtain the sub-image features output by the second image encoder corresponding to each environmental sub-image.

[0070] The sub-image features corresponding to each of the environmental sub-images constitute the basic image features.

[0071] Step 2: Input the sub-image features corresponding to each environmental sub-image into the pre-built view transformation model to obtain the bird's-eye view features output by the view transformation model.

[0072] Step 3: Input the bird's-eye view features into the pre-built mapping model to obtain the map features output by the mapping model.

[0073] Among them, the sub-image features, bird's-eye view features, and map features corresponding to all environmental sub-images constitute the target image features.

[0074] The second image encoder can be an image encoder based on a Residual Network (ResNet) architecture. The view transformation model can be a Bird's Eye View (BEV) model, such as BEVFormer, PETRv2, BEVDepth, etc. The mapping model can be MapTR, MapTRv2, StreamMapNet, etc.

[0075] Specifically, the vehicle-mounted terminal can input each environmental sub-image into a pre-built second image encoder to obtain the sub-image features corresponding to each environmental sub-image output by the second image encoder. Then, the vehicle-mounted terminal can input the sub-image features corresponding to each environmental sub-image into a view transformation model. After the view transformation model performs a perspective transformation operation, it obtains bird's-eye view features, which include environmental depth, semantic, and geometric information. Finally, the vehicle-mounted terminal can input the bird's-eye view features into a pre-built mapping model to obtain the map features output by the mapping model.

[0076] In this way, by extracting image features from different dimensions, the target image features can include basic image information, top-down information, and map information. When used as target conditions, this can generate a more accurate target driving trajectory. In addition, extracting map features based on real-time captured environmental images allows the subsequent use of map features consistent with the actual situation to guide the generation process of the driving trajectory, resulting in a more accurate driving trajectory and avoiding safety issues.

[0077] In some optional implementations, the mapping model described above may include a queryer and a third image encoder. Accordingly, in step three of step S203 above, the vehicle-mounted terminal may specifically employ the following steps to obtain map features based on bird's-eye view features:

[0078] Step 1: Input the bird's-eye view features into the query tool to obtain the structured information corresponding to at least one map element output by the query tool.

[0079] Step two: Input the structured information of the target map element into the third image encoder to obtain the image encoding vector corresponding to the target map element output by the third image encoder.

[0080] Step 3: After determining the map encoding vector corresponding to at least one map element, construct map features based on the map encoding vector corresponding to at least one map element.

[0081] Map elements can include lane lines, road boundaries, drivable areas, obstacles, etc. Structured information can be a series of coordinate points. For example, a lane line can be represented as an ordered two-dimensional point sequence L1={(x1,y1),(x2,y2),...,(xn,yn)}, which defines the shape and position of the lane line. Road boundaries can be represented as a series of coordinate points RB1={(x'1,y'1),(x'2,y'2),...,(xm,y'm)}. Drivable areas are typically modeled as a binary semantic segmentation mask, i.e., a two-dimensional matrix with the same size as the BEV feature. Each pixel value in this matrix indicates whether the location belongs to the drivable area; for example, a pixel value of 0 indicates that the location belongs to a non-drivable area, and a pixel value of 1 indicates that the location belongs to the drivable area. The structured information of an obstacle (such as other vehicles) can be represented by its center point coordinates (xc, yc) and orientation angle (θ) in the BEV space: Obstacle1=(xc, yc, θ).

[0082] Specifically, the vehicle-mounted terminal can input bird's-eye view features into a query generator. The query generator can use its own query mechanism to predict at least one map element in parallel and determine the structured information of each map element. Further, the vehicle-mounted terminal can input the structured information of each map element into a third image encoder. The third image encoder can then use a preset encoding method (e.g., VectorNet) to encode each map element separately, obtaining a graph encoding vector (generally called a subgraph encoding) corresponding to each map element. Each subgraph encoding is then treated as a node, and a self-attention mechanism is used to capture the relationships between nodes. Based on these relationships, a map feature containing all map elements is constructed.

[0083] In some optional implementations, in step S201 above, the vehicle terminal may specifically generate the basic driving trajectory using the following steps:

[0084] Step 1: Obtain multiple driving trajectories.

[0085] Step 2: After performing noise addition operations on multiple driving trajectories, the basic driving trajectory is obtained.

[0086] Specifically, multiple driving trajectories can be driving trajectories in the trajectory anchor point vocabulary. After acquiring multiple driving trajectories, the vehicle terminal can add noise to each of the multiple driving trajectories. The multiple driving trajectories with added noise constitute the basic driving trajectory.

[0087] In some optional implementations, when the basic driving trajectory is a driving trajectory generated by pre-noising multiple typical driving trajectories included in the trajectory anchor point vocabulary, the vehicle terminal can specifically perform noise addition operations on the multiple driving trajectories separately to obtain the basic driving trajectory, including:

[0088] Step 1: In the current noise-adding cycle, random sampling is performed based on the second preset variance and the pre-constructed Gaussian distribution function to obtain the noise added to the current noise-adding cycle and determine the driving trajectory to be processed in the current noise-adding cycle.

[0089] Step 2: Extract the noise scheduling parameters and signal retention ratio corresponding to the current noise addition round from the preset parameter sequence.

[0090] Step 3: Based on the noise added in the current noise-added round, as well as the noise scheduling parameters and signal retention ratio corresponding to the current noise-added round, perform noise-added operation on the driving trajectory to be processed to obtain the noise-added driving trajectory in the current noise-added round.

[0091] Step 4: Once it is determined that the current noise-adding round is the last noise-adding round, stop the noise-adding process.

[0092] Alternatively, in step 5, once it is determined that the current noise-adding round is not the last noise-adding round, proceed to the next noise-adding round, and stop the noise-adding process only after the noise-adding processing of the last noise-adding round is completed.

[0093] Specifically, when the current noise-adding round is the first noise-adding round, the driving trajectory to be processed is the first driving trajectory, which can be any one of multiple driving trajectories. Alternatively, when the current noise-adding round is not the first noise-adding round, the driving trajectory to be processed is the second driving trajectory, which is the driving trajectory obtained after the first driving trajectory has undergone noise-adding operations in previous noise-adding rounds before the current noise-adding round. The pre-constructed Gaussian distribution function can be a standard Gaussian distribution function, that is, a Gaussian distribution function with a mean of 1 and a variance of 1. The multiple noise-adding driving trajectories determined in the last noise-adding round constitute the basic driving trajectory.

[0094] Specifically, the vehicle terminal can use a Markov chain to perform a preset number of noise-adding cycles on each driving trajectory in the trajectory anchor point vocabulary.

[0095] Taking the first driving trajectory as an example, in the current noise-adding cycle, the vehicle terminal can input the second preset variance into the Gaussian distribution function, perform a sampling operation, and obtain the noise added in the current noise-adding cycle output by the Gaussian distribution function. If the current noise-adding cycle is the first noise-adding cycle, the vehicle terminal can determine the first driving trajectory as the driving trajectory to be processed in the current noise-adding cycle. Alternatively, if the current noise-adding cycle is not the first noise-adding cycle, the vehicle terminal can determine the second driving trajectory as the driving trajectory to be processed in the current noise-adding cycle. The second driving trajectory has already undergone one or more rounds of noise-adding operations.

[0096] In addition, the vehicle-mounted terminal can pre-store a preset parameter sequence, which may include a preset number of noise scheduling parameters and signal retention ratios corresponding to each noise-adding cycle. The vehicle-mounted terminal can extract the noise scheduling parameters and signal retention ratios corresponding to the current noise-adding cycle from the preset parameter sequence, i.e., extract the noise scheduling parameters and signal retention ratios corresponding to the current noise-adding cycle. Since the noise scheduling parameters indicate the noise intensity during the noise-adding process, the signal retention ratio indicates the proportion of the original signal (i.e., the driving trajectory to be processed) retained, and the added noise indicates the noise level, the vehicle-mounted terminal can perform noise-adding operations on the driving trajectory to be processed based on the added noise of the current noise-adding cycle, as well as the noise scheduling parameters and signal retention ratios corresponding to the current noise-adding cycle, to obtain the noise-added driving trajectory in the current noise-adding cycle.

[0097] When it is determined that the current noise-adding rounds are in the same order as the preset order (and the same number of rounds), the current noise-adding round is designated as the last noise-adding round, and noise-adding processing can be stopped. Alternatively, when it is determined that the current noise-adding rounds are in a different order than the preset order, the next noise-adding round begins, and noise-adding processing continues until the last noise-adding round is completed, at which point it stops.

[0098] For each driving trajectory in the trajectory anchor point vocabulary, a similar method can be used to perform a preset number of noise-adding rounds. In this way, the multiple noisy driving trajectories determined in the last noisy round constitute the basic driving trajectory. For example... Figure 5 The noise-adding process shown in the diagram generates a base driving trajectory with more noise after multiple driving trajectories are added.

[0099] For example, step 3 above can be expressed as follows:

[0100] (1)

[0101] in, The driving trajectory after noise addition. The noise added in the current noise-adding round, The driving trajectory to be processed (i.e., the driving trajectory after the previous noise-added cycle). To preserve the signal proportion for the current noisy round, The noise scheduling parameters for the current noise-adding round. The cumulative product of the signal retention ratios of the current noise-adding round and the historical noise-adding rounds before the current noise-adding round. , The signal retention ratio for the second round of noise addition.

[0102] In some optional implementations, when the basic driving trajectory is a driving trajectory generated by pre-noising multiple typical driving trajectories included in the trajectory anchor point vocabulary, in step S205 above, the vehicle terminal can specifically perform the following steps to denoise the basic driving trajectory and then generate the target driving trajectory, including:

[0103] Step 1: In the current denoising cycle, determine the initial driving trajectory of the current denoising cycle.

[0104] Step 2: Determine the final trajectory of the current denoising round based on the current denoising round's sequence, initial trajectory, and target conditions.

[0105] Step 3: Once it is determined that the current denoising cycle is the last denoising cycle, stop the denoising process and determine the final driving trajectory determined in the last denoising cycle as the target driving trajectory.

[0106] Alternatively, in step four, once it is determined that the current denoising cycle is not the last denoising cycle, proceed to the next denoising cycle until the denoising process of the last denoising cycle is completed, then stop the denoising process and determine the final driving trajectory determined by the last denoising cycle as the target driving trajectory.

[0107] Specifically, when the current denoising cycle is the first denoising cycle, the initial driving trajectory is the basic driving trajectory; or when the current denoising cycle is not the first denoising cycle, the initial driving trajectory is the final driving trajectory determined by the previous denoising cycle before the current denoising cycle.

[0108] Specifically, the vehicle terminal can use a pre-built diffusion model to perform a preset number of denoising cycles on the basic driving trajectory. For example, the diffusion model can be DiffusionDrive, HydraMDP, etc.

[0109] In the first denoising cycle, the vehicle terminal can directly determine the basic driving trajectory as the initial driving trajectory. In subsequent denoising cycles, the vehicle terminal can determine the final driving trajectory generated after the denoising process of the previous cycle as the initial driving trajectory for the current denoising cycle. Since the number of denoising processes affects the denoising process, the order of the current denoising cycles can also be used as a reference condition. Based on the order of the current denoising cycles and the target conditions, the initial driving trajectory of the current denoising cycle can be adjusted to obtain the final driving trajectory of the current denoising cycle. After determining that the order of denoising processes equals the preset number, the current denoising cycle can be determined as the last denoising cycle. At this point, denoising processing can be stopped, and the final driving trajectory determined in the last denoising cycle can be determined as the target driving trajectory. Alternatively, if the number of noise reduction processes is less than the preset number, it can be determined that the current noise reduction round is not the last one, and the next noise reduction round can be processed. This continues until the last noise reduction round is completed, at which point the noise reduction process stops, and the final driving trajectory determined by the last noise reduction round is taken as the target driving trajectory. Figure 5 The noise reduction process shown in the figure transforms the noisy base trajectory into a clear target trajectory after noise reduction.

[0110] In this way, each denoising step builds upon the result of the previous step, and each step incorporates target conditions. This ensures that each optimization is aligned with the driving scenario, resulting in a more accurate and realistic target driving trajectory. For example, if the target condition indicates the vehicle is turning, noise that would cause the trajectory to be straight can be removed during denoising, guiding the trajectory towards the correct curve. Another example is if the target condition indicates rainy weather and traffic congestion, which would tend to generate a slower, smoother trajectory. Yet another example is if the target condition indicates an open road segment and the driving intention is to overtake, which could generate a more aggressive and faster lane-changing trajectory.

[0111] In some optional implementations, in step two of step S205 above, the vehicle terminal may use the following specific steps to determine the final driving trajectory of the current denoising cycle based on the current denoising cycle sequence, the initial driving trajectory of the current denoising cycle, and the target conditions:

[0112] Step 1: Input the current denoising cycle order, the initial driving trajectory of the current denoising cycle, and the target conditions into the pre-built noise prediction model to obtain the noise removal prediction output of the noise prediction model corresponding to the current denoising cycle.

[0113] Step 2: Extract the noise scheduling parameters and signal retention ratio corresponding to the current denoising round from the preset parameter sequence.

[0114] Step 3: Determine the average driving trajectory of the current denoising round based on the initial driving trajectory, predicted noise removal, noise scheduling parameters, and signal retention ratio of the current denoising round.

[0115] Step 4: Obtain the random noise corresponding to the current denoising round.

[0116] Step 5: Generate the final driving trajectory of the current denoised cycle based on the average driving trajectory, the first preset variance, and random noise.

[0117] Specifically, the vehicle-mounted terminal can input the current denoising cycle order, the initial driving trajectory of the current denoising cycle, and the target conditions into a pre-constructed noise prediction model (which can be a component of the aforementioned diffusion model, or a denoising network within the diffusion model) to obtain the predicted noise removal output of the noise prediction model corresponding to the current denoising cycle. The preset parameter sequence can include a preset number of noise scheduling parameters and signal retention ratios corresponding to each denoising cycle. Then, the vehicle-mounted terminal can extract the noise scheduling parameters and signal retention ratios corresponding to the current denoising cycle order from the preset parameter sequence, and determine them as the noise scheduling parameters and signal retention ratios corresponding to the current denoising cycle. During the denoising process, the noise scheduling parameters can indicate the noise intensity eliminated during denoising. Furthermore, the vehicle-mounted terminal can perform denoising processing on the initial driving trajectory of the current denoising cycle based on the predicted noise removal, noise scheduling parameters, and signal retention ratio to obtain the average driving trajectory of the current denoising cycle. Additionally, the vehicle-mounted terminal can randomly sample noise from standard Gaussian noise to obtain the random noise corresponding to the current denoising cycle. Finally, the vehicle terminal can adjust the average driving trajectory based on the first preset variance and random noise to obtain the final driving trajectory of the current denoised cycle.

[0118] In this way, in each denoising cycle, the predicted noise removal is determined by the target conditions, the order of the current denoising cycle, and the initial driving trajectory of the current denoising cycle. This ensures that the denoising direction is consistent with the target conditions. Furthermore, by using historical denoising information as a reference (the order of the current denoising cycle and the initial driving trajectory of the current denoising cycle), the predicted noise removal takes into account the characteristics of historical denoising operations. This makes the final driving trajectory after subsequent denoising based on this prediction more accurate and in line with the actual scenario, thus avoiding various safety issues.

[0119] For example, step 3 can be expressed as follows:

[0120] (2)

[0121] in, Let t1 be the initial driving trajectory for the current denoising cycle, and t1 be the current denoising cycle. This represents the average driving trajectory of the current noise-reduced wheel. The proportion of signal retained in the current denoising round. These are the noise scheduling parameters for the current denoising round. To predict and remove noise, The cumulative product of the signal retention ratios of the current denoising round and the historical denoising rounds preceding the current denoising round. , The signal retention ratio for the s1th noise-adding round.

[0122] Step 5 can be expressed as follows:

[0123] (3)

[0124] in, This represents the final driving trajectory of the current noise-reducing cycle. The average driving trajectory, The first preset variance, It is random noise.

[0125] In this way, random noise can simulate the uncertainties in the actual traffic environment. By adding random noise to the average driving trajectory, the generated final driving trajectory can be made more realistic and accurate.

[0126] Based on the above embodiments, such as Figure 6 As shown, the vehicle-mounted terminal inputs the environmental image into the second image encoder to obtain the basic image features output by the second image encoder. These basic image features are then input into the view transformation model to obtain the bird's-eye view features. Additionally, the vehicle-mounted terminal can input the extracted prompts from the environmental image and semantic features into the language model to obtain the target semantic features. Finally, the vehicle-mounted terminal can input the target semantic features, bird's-eye view features, map features, and trajectory anchor point vocabulary into the diffusion model to obtain the target driving trajectory.

[0127] It should be noted that the letters in the above formulas (1) to (3) are all dimensionless vectors or matrices.

[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0129] Embodiments of this application also provide a device for generating a driving trajectory, such as... Figure 7 As shown, it includes:

[0130] The acquisition module 710 is used to acquire environmental images of the target vehicle, semantic feature extraction indication information, and pre-generated basic driving trajectory;

[0131] The semantic feature extraction module 720 is used to input environmental images and semantic feature extraction instruction information into a pre-trained language model to obtain the target semantic features output by the language model.

[0132] Image feature extraction module 730 is used to extract target image features from environmental images;

[0133] The generation module 740 is used to generate target conditions based on target image features and target semantic features; and to generate a target driving trajectory that meets the target conditions based on the target conditions and the basic driving trajectory, so as to guide the driving operation of the target vehicle.

[0134] In some alternative implementations, the generation module 740 is specifically used for:

[0135] In the current denoising cycle, the initial driving trajectory of the current denoising cycle is determined. When the current denoising cycle is the first denoising cycle, the initial driving trajectory is the basic driving trajectory. Alternatively, when the current denoising cycle is not the first denoising cycle, the initial driving trajectory is the final driving trajectory determined by the previous denoising cycle before the current denoising cycle.

[0136] Based on the current denoising cycle sequence, the initial trajectory of the current denoising cycle, and the target conditions, determine the final trajectory of the current denoising cycle.

[0137] Once it is determined that the current denoising cycle is the last denoising cycle, the denoising process is stopped, and the final driving trajectory determined by the last denoising cycle is set as the target driving trajectory.

[0138] Alternatively, once it is determined that the current denoising cycle is not the last denoising cycle, the next denoising cycle is started, and the denoising process continues until the last denoising cycle is completed. Then, the denoising process is stopped, and the final driving trajectory determined by the last denoising cycle is set as the target driving trajectory.

[0139] In some alternative implementations, the generation module 740 is specifically used for:

[0140] The order of the current denoising round, the initial driving trajectory of the current denoising round, and the target conditions are input into the pre-built noise prediction model to obtain the noise prediction model output corresponding to the current denoising round.

[0141] Extract the noise scheduling parameters and signal retention ratio corresponding to the current denoising round from the preset parameter sequence;

[0142] The average driving trajectory of the current denoising round is determined based on the initial driving trajectory, predicted noise removal, noise scheduling parameters, and signal retention ratio of the current denoising round.

[0143] Obtain the random noise corresponding to the current denoising round;

[0144] Based on the average driving trajectory, the first preset variance, and random noise, the final driving trajectory of the current denoised cycle is generated.

[0145] In some optional implementations, the average trajectory of the current denoising round is determined based on the initial trajectory of the current denoising round, the predicted noise removal, the noise scheduling parameters, and the signal retention ratio, using the following expression:

[0146]

[0147] in, Let t1 be the initial driving trajectory for the current denoising cycle, and t1 be the current denoising cycle. This represents the average driving trajectory of the current noise-reduced wheel. The proportion of signal retained in the current denoising round. These are the noise scheduling parameters for the current denoising round. To predict and remove noise, It is the cumulative product of the signal retention ratio of the current denoising round and the historical denoising rounds before the current denoising round.

[0148] In some optional implementations, the final driving trajectory of the current denoised round is generated based on the average driving trajectory, a first preset variance, and random noise, using the following expression:

[0149]

[0150] in, This represents the final driving trajectory of the current noise-reducing cycle. The average driving trajectory, The first preset variance, It is random noise.

[0151] In some optional implementations, obtaining a pre-generated basic driving trajectory includes:

[0152] Obtain multiple driving trajectories;

[0153] After performing noise addition operations on multiple driving trajectories, the basic driving trajectory is obtained.

[0154] In some alternative implementations, the generation module 740 is specifically used for:

[0155] In the current noise-adding round, random sampling is performed based on the second preset variance and the pre-constructed Gaussian distribution function to obtain the noise added in the current noise-adding round and determine the driving trajectory to be processed in the current noise-adding round. When the current noise-adding round is the first noise-adding round, the driving trajectory to be processed is the first driving trajectory, which is any one of multiple driving trajectories. Alternatively, when the current noise-adding round is not the first noise-adding round, the driving trajectory to be processed is the second driving trajectory, which is the driving trajectory obtained after the first driving trajectory has undergone noise-adding operations in the previous noise-adding rounds before the current noise-adding round.

[0156] Extract the noise scheduling parameters and signal retention ratio corresponding to the current noise addition round from the preset parameter sequence;

[0157] Based on the noise added in the current noise-adding round, as well as the noise scheduling parameters and signal retention ratio corresponding to the current noise-adding round, noise-adding operation is performed on the driving trajectory to be processed to obtain the noise-adding driving trajectory in the current noise-adding round.

[0158] Once it is determined that the current noise-adding round is the last noise-adding round, the noise-adding process is stopped.

[0159] Alternatively, once it is determined that the current noise-adding round is not the last noise-adding round, proceed to the next noise-adding round, and continue until the noise-adding process of the last noise-adding round is completed, and then stop the noise-adding process;

[0160] Among them, the multiple driving trajectories determined in the last noise-added cycle constitute the basic driving trajectory.

[0161] In some optional implementations, based on the noise added in the current noise-adding round, and the noise scheduling parameters and signal retention ratio corresponding to the current noise-adding round, a noise-adding operation is performed on the driving trajectory to be processed to obtain the noise-adding driving trajectory in the current noise-adding round, using the following expression:

[0162]

[0163] in, The driving trajectory after noise addition. The noise added in the current noise-adding round, The driving trajectory to be processed. To preserve the signal proportion for the current noisy round, The noise scheduling parameters for the current noise-adding round. The cumulative product of the signal retention ratios of the current noise-adding round and the historical noise-adding rounds preceding the current noise-adding round.

[0164] In some optional implementations, the language model includes a first image encoder, a text encoder, a language model decoder, and a semantic feature generator for various semantic types; the semantic feature extraction module 720 is specifically used for:

[0165] The environmental image is input into the first image encoder to obtain the image features output by the first image encoder;

[0166] The semantic feature extraction instructions are input into the text encoder to obtain the text features output by the text encoder.

[0167] The image features and text features output by the first image encoder are input into the language model decoder to obtain the fused features output by the language model decoder;

[0168] The fused features are input into multiple semantic feature generators to obtain semantic features output by each semantic feature generator corresponding to its own semantic type.

[0169] The semantic features corresponding to multiple semantic types are concatenated to obtain the target semantic features.

[0170] For a description of the features in the embodiment corresponding to the device for generating the driving trajectory, please refer to the relevant description in the embodiment corresponding to the method for generating the driving trajectory, which will not be repeated here.

[0171] Embodiments of this application also provide an electronic device, such as... Figure 8 As shown, it includes a memory 10 and a processor 20. The memory 10 stores a computer program, and the processor 20 is configured to run the computer program to perform the steps in any of the above-described embodiments of the method for generating driving trajectories.

[0172] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the driving trajectory generation method when it is run.

[0173] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0174] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the driving trajectory generation method.

[0175] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described methods for generating driving trajectories.

[0176] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0177] The foregoing has provided a detailed description of the method, apparatus, electronic device, storage medium, and program product for generating driving trajectories provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A method of generating a travel trajectory, characterized by, The method comprises the following steps: obtaining an environment image of a target vehicle, extraction indication information of semantic features, and a pre-generated basic driving track; inputting the environment image and the extraction indication information of the semantic features into a pre-trained language model to obtain target semantic features output by the language model; extracting target image features from the environment image; generating a target condition based on the target image features and the target semantic features; generating a target driving track that meets the target condition based on the target condition and the basic driving track, comprising: in a current denoising round, determining an initial driving track of the current denoising round, wherein when the current denoising round is the first denoising round, the initial driving track is the basic driving track, or when the current denoising round is a non-first denoising round, the initial driving track is a final driving track determined by a previous denoising round before the current denoising round; inputting the order of the current denoising round, the initial driving track of the current denoising round, and the target condition into a pre-constructed noise prediction model to obtain predicted removed noise corresponding to the current denoising round output by the noise prediction model; extracting noise scheduling parameters and signal retention ratios corresponding to the current denoising round from a preset parameter sequence; determining an average driving track of the current denoising round according to the initial driving track of the current denoising round, the predicted removed noise, the noise scheduling parameters, and the signal retention ratios; obtaining random noise corresponding to the current denoising round; generating a final driving track of the current denoising round according to the average driving track, a first preset variance, and the random noise; when it is determined that the current denoising round is the last denoising round, stopping the denoising process and determining the final driving track determined by the last denoising round as the target driving track.

2. The travel trajectory generation method according to claim 1, characterized by, The method further comprises the following steps: when it is determined that the current denoising round is not the last denoising round, entering the processing of a next denoising round until the denoising process of the last denoising round is completed, then stopping the denoising process and determining the final driving track determined by the last denoising round as the target driving track.

3. The travel trajectory generation method according to claim 1, characterized by, The determination of the average driving track of the current denoising round according to the initial driving track of the current denoising round, the predicted removed noise, the noise scheduling parameters, and the signal retention ratios adopts the following expression: wherein, is an initial driving trajectory for the current denoising round, t1 is the current denoising round, is an average driving trajectory for the current denoising round, is a signal reservation ratio for the current denoising round, is the noise scheduling parameter for the current denoising round, is the predicted denoising, is a cumulative product of signal reservation ratios for the current denoising round and a historical denoising round before the current denoising round.

4. The travel trajectory generation method according to claim 1, characterized by, The generation of the final driving track of the current denoising round according to the average driving track, the first preset variance, and the random noise adopts the following expression: wherein, is the final driving trajectory for the current denoising iteration, is the average driving trajectory, is the first preset variance, is the random noise.

5. The travel trajectory generation method according to any one of claims 1 to 4, characterized by, The method for obtaining the pre-generated basic driving track comprises the following steps: obtaining a plurality of driving tracks; performing a noise adding operation on each of the plurality of driving tracks to obtain the basic driving track.

6. The travel trajectory generation method according to claim 5, characterized by, The method for obtaining the basic driving track by performing a noise adding operation on each of the plurality of driving tracks comprises the following steps: In the current noise adding round, based on the second preset variance and the pre-constructed Gaussian distribution function, random sampling is performed to obtain noise of the current noise adding round, and a to-be-processed driving track of the current noise adding round is determined, wherein when the current noise adding round is the first noise adding round, the to-be-processed driving track is a first driving track, and the first driving track is any one of the plurality of driving tracks, or when the current noise adding round is a non-first noise adding round, the to-be-processed driving track is a second driving track, and the second driving track is a driving track obtained after a noise adding operation of a historical noise adding round before the first driving track; The noise scheduling parameter and the signal retention ratio corresponding to the current noise adding round are extracted from the preset parameter sequence; Based on the noise of the current noise adding round, and the noise scheduling parameter and the signal retention ratio corresponding to the current noise adding round, a noise adding operation is performed on the to-be-processed driving track to obtain a driving track after noise adding in the current noise adding round; When it is determined that the current noise adding round is the last noise adding round, the noise adding process is stopped; Or, when it is determined that the current noise adding round is not the last noise adding round, the next noise adding round is entered, and the noise adding process is stopped after the last noise adding round is completed; The driving tracks after noise adding determined by the last noise adding round constitute the basic driving track.

7. The travel trajectory generation method according to claim 6, characterized by, The noise adding operation on the to-be-processed driving track based on the noise of the current noise adding round, and the noise scheduling parameter and the signal retention ratio corresponding to the current noise adding round, to obtain the driving track after noise adding in the current noise adding round, adopts the following expression: wherein, is the current noise added driving track, is the noise added noise of the current noise added round, is the driving track to be processed, is the signal reservation ratio of the current noise added round, is the noise scheduling parameter of the current noise added round, is the cumulative product of the signal reservation ratios of the current noise added round and the historical noise added round before the current noise added round.

8. The travel trajectory generation method according to any one of claims 1 to 4, characterized by The language model comprises a first image encoder, a text encoder, a language model decoder, and a plurality of semantic feature generators of different semantic types; The language model comprises a first image encoder, a text encoder, a language model decoder, and a plurality of semantic feature generators of different semantic types; The image feature output by the first image encoder is input into the text encoder to obtain a text feature output by the text encoder; The image feature output by the first image encoder is input into the text encoder to obtain a text feature output by the text encoder; The image feature output by the first image encoder is input into the text encoder to obtain a text feature output by the text encoder; The fusion feature is input into a plurality of the semantic feature generators respectively to obtain a plurality of semantic features corresponding to the semantic types of the semantic feature generators respectively; The semantic features corresponding to the different semantic types are spliced to obtain the target semantic feature.

9. An electronic device, comprising: The memory is configured to store a computer program. The processor is configured to execute the computer program to implement the steps of the driving track generation method according to any one of claims 1 to 8. ​

Citation Information

Patent Citations

  • Automatic driving dangerous scene detection method based on large language model

    CN119091417A

  • Driving track planning method, and driving track planning model training method and device

    CN120609378A