Track generation method and device, equipment, medium and product

By generating and correcting candidate driving trajectories using a target trajectory generation model and a driving correction model, the problem of inaccurate trajectory prediction in complex scenarios by end-to-end models is solved, and higher trajectory correction accuracy is achieved.

CN121277162APending Publication Date: 2026-01-06BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510168766.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-04
Filing Date
2025-02-14
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing end-to-end models cannot effectively handle complex driving scenarios, resulting in inaccurate trajectory prediction.

Method used

By acquiring vehicle driving-related data and utilizing pre-created target trajectory generation and driving correction models, candidate driving trajectories are generated and corrected, thereby improving the accuracy of trajectory prediction.

Benefits of technology

It effectively improves the accuracy and effectiveness of trajectory correction in complex driving scenarios and solves the problem of poor trajectory prediction in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121277162A_ABST
    Figure CN121277162A_ABST
Patent Text Reader

Abstract

The invention discloses a track generation method and device, equipment, a medium and a product. The method comprises the following steps: acquiring driving related data corresponding to a current vehicle; inputting first driving related data in the driving related data into a pre-created target trajectory generation model to obtain a candidate driving trajectory corresponding to the current vehicle; inputting second driving related data in the driving related data into a pre-created target driving correction model to obtain corresponding driving correction information; and correcting the candidate driving track based on the driving correction information to obtain a corresponding target driving track. According to the invention, the problem of poor correction effect caused by the adoption of a manually set rule-based track optimization strategy for correction in the prior art is effectively avoided, and the correction accuracy and effectiveness of the driving track in a complex driving scene are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 2024108968033, filed on July 4, 2024, entitled “Trajectory Generation Method, Apparatus, Device, Medium and Product”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This invention relates to the field of intelligent driving technology, and in particular to a trajectory generation method, apparatus, device, medium, and product. Background Technology

[0003] In the field of autonomous driving, trajectory prediction technology is playing an increasingly important role. This technology can predict a vehicle's trajectory over a certain period of time, thus avoiding dangerous interactions and providing analytical support to decision-making and planning systems.

[0004] In existing technologies, end-to-end models can be used for trajectory correction, but these models cannot handle complex driving scenarios (such as the sudden appearance of cattle or sheep on the road). Therefore, how to deal with complex driving scenarios and perform trajectory prediction is an urgent problem to be solved. Summary of the Invention

[0005] This invention provides a trajectory generation method, apparatus, device, medium, and product to solve the technical problem that existing technologies cannot predict trajectories for complex scenarios.

[0006] According to one aspect of the present invention, a trajectory generation method is provided, comprising:

[0007] Obtain the driving-related data for the current vehicle;

[0008] The first driving-related data in the driving-related data is input into a pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle;

[0009] The second driving-related data from the driving-related data is input into the pre-created target driving correction model to obtain the corresponding driving correction information;

[0010] The candidate driving trajectory is corrected based on the driving correction information to obtain the corresponding target driving trajectory.

[0011] According to another aspect of the present invention, a trajectory generation apparatus is provided, comprising:

[0012] The acquisition module is used to acquire driving-related data for the current vehicle.

[0013] The first generation module is used to input the first driving-related data from the driving-related data into a pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle.

[0014] The second generation module is used to input the second driving-related data from the driving-related data into the pre-created target driving correction model to obtain the corresponding driving correction information.

[0015] The third generation module is used to correct the candidate driving trajectory based on the driving correction information to obtain the corresponding target driving trajectory.

[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the trajectory generation method according to any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the trajectory generation method according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the trajectory generation method described in any embodiment of the present invention.

[0022] The technical solution of this invention obtains first and second driving-related data of the vehicle during its current driving process, inputs the first driving-related data into a pre-created target trajectory generation model to obtain corresponding candidate driving trajectories, and inputs the second driving-related data into a target driving correction model to obtain corresponding driving correction information. The pre-generated candidate driving trajectories are automatically corrected using the driving correction information to obtain the corresponding target driving trajectory. This effectively avoids the problem of poor correction effect caused by manually set rule-based trajectory optimization strategies in the prior art, and effectively improves the accuracy and effectiveness of driving trajectory correction in complex driving scenarios.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a trajectory generation method provided in an embodiment of the present invention;

[0026] Figure 2 This is a flowchart of a training method for a target trajectory generation model provided in an embodiment of the present invention;

[0027] Figure 3 A flowchart of another trajectory generation method provided in an embodiment of the present invention;

[0028] Figure 4 This is a flowchart of another trajectory generation method provided in an embodiment of the present invention;

[0029] Figure 5 This is a flowchart of the training process for a target driving correction model provided in an embodiment of the present invention;

[0030] Figure 6 This is a schematic diagram illustrating an implementation of information correction provided in an embodiment of the present invention;

[0031] Figure 7 This is a flowchart illustrating the implementation of timing fusion according to an embodiment of the present invention;

[0032] Figure 8 This is a flowchart of another trajectory generation method provided in an embodiment of the present invention;

[0033] Figure 9 This is a schematic diagram of the structure of a trajectory generation device provided in an embodiment of the present invention;

[0034] Figure 10 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0035] To enable those skilled in the art to better understand the present invention, the following will be described in conjunction with embodiments of the present invention. Attached imageThe technical solutions in the embodiments of the present invention have been clearly and completely described. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0036] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0037] In one embodiment, Figure 1 This is a flowchart of a trajectory generation method provided by an embodiment of the present invention. This embodiment is applicable to generating driving trajectories in complex driving scenarios under autonomous driving conditions and automatically correcting the driving trajectories. This method can be executed by a trajectory generation device, which can be implemented in hardware and / or software and can be configured in the vehicle. Figure 1 As shown, the method includes:

[0038] S110. Obtain the driving-related data corresponding to the current vehicle.

[0039] Among them, driving-related data is used to characterize various information generated or collected by the vehicle during autonomous driving, which can characterize the surrounding environment, the vehicle's own state, and driving operations. In one embodiment, driving-related data includes first driving-related data and second driving-related data.

[0040] The first driving-related data may include: sensor information, navigation planning information, and / or driving rule information. Sensor information includes: environmental perception information and / or state information. Environmental perception information includes: current frame data and / or current point cloud data. State information includes: current vehicle position information and / or current vehicle attitude information. The current vehicle position information can be location information collected by GPS deployed inside the vehicle. The current vehicle attitude information may include: current vehicle orientation, pitch angle, accelerator pedal opening, and gear information, etc. Navigation planning information can be obtained by extracting navigation data (total navigation data from the starting point to the destination) based on the current vehicle position information (navigation data from the current vehicle's position to the destination, or navigation data from a preset distance before the current vehicle's position to the destination). Driving rule information may include: traffic rule information and / or speed limit information. Speed ​​limit information may include speed limits, acceleration, deceleration, etc. Traffic rule information is used to determine whether driving behavior complies with traffic regulations; for example, bus lanes are prohibited, and speed limits are required in school zones.

[0041] In one example, the second driving-related data includes: environmental perception information, navigation planning information, and driving prompt information. Environmental perception information refers to information collected by the vehicle during autonomous driving that characterizes the surrounding environment. In this embodiment, various sensors in the vehicle (e.g., cameras, millimeter-wave radar, and lidar) can be used to acquire surrounding environmental information, including: road conditions, obstacles, other vehicles, pedestrians, and weather conditions. Road conditions may include: road type, lane lines, traffic signs, and markings. Obstacles may include: obstacle type, obstacle location, obstacle shape, and obstacle speed. Other vehicles may include: the location, speed, and direction of travel of other vehicles. Pedestrians may include: the location and direction of travel of pedestrians. Weather conditions may include: rain, snow, and fog. Navigation planning information refers to the vehicle's real-time geographical location, map data, and route planning information. Real-time geographical location can be obtained through GPS or other positioning systems; map data can include relevant data from high-precision maps; route planning information refers to the data and description of the optimal or near-optimal driving route from the vehicle's starting point to its destination, such as: the geometry of the driving trajectory (e.g., curves, slopes, and straight sections), road nodes and segments (intersections, highway entrances and exits), traffic condition predictions (e.g., traffic congestion, construction areas, and accidents), obstacle avoidance strategies, and speed limits. Driving prompt information refers to a description of the vehicle's current driving environment and the identification of key objects, used to guide the target driving correction model in its thinking and reasoning. The current driving environment can include road conditions and weather conditions, and key objects can include obstacles.

[0042] In this embodiment, a prompt information database for an intelligent driving system can be created in advance, and corresponding driving prompt information can be obtained from the prompt information database based on the current environmental perception information of the vehicle.

[0043] S120. Input the first driving-related data from the driving-related data into the pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle.

[0044] Among them, the candidate driving trajectory is the predicted trajectory within a certain period of time in the future. For example, the candidate driving trajectory can be the predicted trajectory within the next 8 seconds.

[0045] It should be noted that the target trajectory generation model is obtained through iterative training based on a pre-trained initial trajectory generation model.

[0046] The target trajectory generation model may include a target backbone network, a target encoder, and a target decoder. The target trajectory generation model may also include a target backbone network, a target encoder, a target decoder, and a target memory module, wherein the target memory module is used to store historical features.

[0047] Specifically, the target trajectory generation model includes: a target backbone network, a target encoder, and a target decoder; the first driving-related data from the driving-related data is input into the pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle, including:

[0048] Input the first driving-related data into the target backbone network to obtain the target fusion features;

[0049] Input the first driving-related data into the target encoder to obtain the target encoded features;

[0050] The target fusion features and target encoding features are input into the target decoder to obtain the candidate driving trajectory corresponding to the current vehicle.

[0051] Specifically, the target trajectory generation model includes: a target backbone network, a target encoder, a target decoder, and a target memory module; the first driving-related data from the driving-related data is input into the pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle, including:

[0052] Input the first driving-related data into the target backbone network to obtain the target fusion features;

[0053] Input the first driving-related data into the target encoder to obtain the target encoded features;

[0054] The target fusion features, the historical features output by the target memory module, and the target encoding features are input into the target decoder to obtain the candidate driving trajectory corresponding to the current vehicle.

[0055] It should be noted that the first driving-related data includes sensor information and navigation planning information. The sensor information includes environmental perception information and state information. The target trajectory generation model includes a target backbone network, a target encoder, and a target decoder. The first driving-related data is input into the pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle, including:

[0056] Environmental perception information is input into the target backbone network to obtain target fusion features;

[0057] The state information and navigation planning information are input into the target encoder to obtain the target encoding features;

[0058] The target encoded features and target fusion features are input into the target decoder to obtain the candidate driving trajectory corresponding to the current vehicle.

[0059] It should be noted that the first driving-related data includes sensor information and navigation planning information. The sensor information includes environmental perception information and state information. The target trajectory generation model includes a target backbone network, a target encoder, a target decoder, and a target memory module. The first driving-related data is input into the pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle, including:

[0060] Environmental perception information is input into the target backbone network to obtain target fusion features;

[0061] The state information and navigation planning information are input into the target encoder to obtain the target encoding features;

[0062] The target encoding features, historical features output by the target memory module, and target fusion features are input into the target decoder to obtain the candidate driving trajectory corresponding to the current vehicle.

[0063] It should be noted that the first driving-related data includes sensor information, driving rule information, and navigation planning information. Sensor information includes environmental perception information and state information. The target trajectory generation model includes a target backbone network, a target encoder, and a target decoder. The first driving-related data is input into the pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle, including:

[0064] Environmental perception information is input into the target backbone network to obtain target fusion features;

[0065] The status information, driving rule information, and navigation planning information are input into the target encoder to obtain the target encoding features;

[0066] The target encoded features and target fusion features are input into the target decoder to obtain the candidate driving trajectory corresponding to the current vehicle.

[0067] It should be noted that the first driving-related data includes sensor information, driving rule information, and navigation planning information. Sensor information includes environmental perception information and state information. The target trajectory generation model includes a target backbone network, a target encoder, a target decoder, and a target memory module. The first driving-related data is input into the pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle, including:

[0068] Environmental perception information is input into the target backbone network to obtain target fusion features;

[0069] The status information, driving rule information, and navigation planning information are input into the target encoder to obtain the target encoding features;

[0070] The target fusion features, the historical features output by the target memory module, and the target encoding features are input into the target decoder to obtain the candidate driving trajectory corresponding to the current vehicle.

[0071] The technical solution of this embodiment obtains candidate driving trajectories for the current vehicle by inputting the first driving-related data of the current vehicle into the target trajectory generation model. This reduces information loss and improves the accuracy of the predicted trajectory. Based on the candidate driving trajectory output by the target trajectory generation model, the vehicle can be controlled in parallel in both the lateral and longitudinal directions, rather than in a serial manner. That is, left-right control and forward-backward control are not separated, which allows the vehicle to change lanes or avoid obstacles more smoothly during driving.

[0072] Optionally, driving-related data can be input into the target trajectory generation model to obtain candidate driving trajectories corresponding to the current vehicle, including:

[0073] By inputting driving-related data into the target trajectory generation model, candidate driving trajectories, obstacle information, and road structure corresponding to the current vehicle are obtained.

[0074] It should be noted that the output of the target trajectory generation model includes not only candidate driving trajectories but also obstacle information and road structure, with the road structure potentially including lane markings. This obstacle information and road structure are used to assist in displaying obstacles, thus alerting the driver to their location and helping to avoid collisions, thereby improving driving safety.

[0075] In this embodiment of the invention, the target trajectory generation model can output multiple trajectories. The multiple trajectories are verified through post-processing, and the trajectories that pass the verification are determined as candidate driving trajectories.

[0076] Optionally, obstacle information includes: first type obstacle information and second type obstacle information.

[0077] The first type of obstacle information can be information about obstacles with fixed shapes. Examples include pedestrians, bicycles, and cars. The second type of obstacle information can be information about obstacles with non-fixed shapes. Examples include excavators, fences, cranes, and streetlights. The second type of obstacle information can also be OCC (Occupancy Network) information. It enables accurate identification and segmentation of obstacles in complex road environments, thereby improving the safety and flexibility of intelligent driving systems. OCC information can be a three-dimensional occupancy grid representing the spatial distribution of obstacles in an image.

[0078] In a specific example, frame data captured by a camera and point cloud data captured by a LiDAR are input into the target backbone network. Features from multiple sensors are extracted and fused, and projected onto the BEV space. State information and navigation planning information are input into the target encoder. After being encoded by a transformer, the data, along with the BEV features, are decoded to obtain information about the first obstacle, the second obstacle, and the road structure, and a candidate driving trajectory is planned. In summary, this embodiment of the invention achieves multi-task output through an integrated model. Therefore, the target trajectory generation model provided by this embodiment differs from other multi-segment or ensemble models in the prior art.

[0079] S130. Input the second driving-related data from the driving-related data into the pre-created target driving correction model to obtain the corresponding driving correction information.

[0080] The driving correction information includes at least one of the following: driving reference position, driving scenario, and driving control correction information. The driving reference position, also known as the longitudinal and lateral trajectory reference signals, refers to multiple reference points in the predicted driving trajectory of the current vehicle. In this embodiment, the predicted driving trajectory of the current vehicle can be given in the form of reference points. For example, if the driving control correction information indicates a detour, the detour trajectory can be broken down into multiple key points, and then automatic driving can be performed according to the trajectory composed of these key points. The driving scenario refers to the road type where the current vehicle is currently located, such as an overpass, a slope, or a main and auxiliary road. The driving control correction information refers to the suggested driving operations for the current vehicle. For example, driving operations (also known as macro driving decision information) can include, but are not limited to, longitudinal and lateral driving. Longitudinal driving can include operations such as acceleration and deceleration; lateral driving can include operations such as turning, lane changing, left turning, and right turning. The target driving correction model is used to generate driving correction information based on the second driving-related data of the current vehicle, and is a model for dynamically adjusting the driving operations of the current vehicle. For example, the target driving correction model can be a Visual Language Model (VLM).

[0081] In this embodiment, the second driving-related data of the current vehicle is input into the target driving correction model to obtain suggestion information on whether acceleration, deceleration and steering operations are needed, and this information is used as driving correction information.

[0082] S140. Correct the candidate driving trajectory based on the driving correction information to obtain the corresponding target driving trajectory.

[0083] In this embodiment, driving correction information is input into a pre-created target trajectory generation model, so that the target trajectory generation model corrects the candidate driving trajectory based on the driving correction information to obtain the corresponding target driving trajectory, so that the current vehicle can drive automatically according to the target driving trajectory.

[0084] In one example, some information from the driving correction information can be input into the target trajectory generation model. For example, macroscopic driving decision information, such as lateral lane changes, detours, and turns; longitudinal information, including set speed, acceleration, deceleration, and stopping / waiting positions; feature vectors encoding complex driving decisions; and lateral and longitudinal trajectory reference information, where the lateral trajectory reference information mainly refers to the sampling points of the driving path, and the longitudinal trajectory reference information mainly refers to the target speed of each point on the sampling points. This allows the target trajectory generation model to correct the candidate driving trajectory based on the driving correction information and obtain the corresponding target driving trajectory.

[0085] In one example, the correction process includes three implementation methods, which are increasingly integrated. Implementation method one: The target driving correction model outputs macroscopic, long-term driving decision suggestions (e.g., lateral and longitudinal), and directly uses these suggestions as input data to the target trajectory generation model. This ensures that the trajectory output by the target trajectory generation model better aligns with the macroscopic suggestions, generating a target driving trajectory that conforms to macroscopic driving decisions. Implementation method two: The target driving correction model outputs macroscopic, long-term driving decision suggestions (e.g., lateral and longitudinal), and presents these suggestions as feature vectors (i.e., encoding the driving decision suggestions as feature vectors). These encoded feature vectors are then used as input data to the target trajectory generation model, ensuring that the target trajectory generation model outputs more accurate driving decisions and trajectories. The third implementation method involves using a learned model router to select whether the target trajectory generation model or the target driving correction model outputs the corresponding target driving trajectory. This allows for the direct use of the target trajectory generation model to output more accurate driving decisions and driving trajectories in complex scenarios, effectively using the driving correction information as the corresponding target driving trajectory and avoiding deviations in the driving decisions and driving trajectories output by the target driving correction model.

[0086] In one example, the target trajectory generation model is a high-frequency system (e.g., operating at a frequency between 30-100Hz) that continuously outputs candidate driving trajectories; the target driving correction model is a low-frequency system (e.g., performing two inferences every second or two seconds).

[0087] The technical solution of this embodiment acquires first and second driving-related data of the vehicle during its current driving process, inputs the first driving-related data into a pre-created target trajectory generation model to obtain the corresponding candidate driving trajectory, and inputs the second driving-related data into a target driving correction model to obtain the corresponding driving correction information. The pre-generated candidate driving trajectory is automatically corrected using the driving correction information to obtain the corresponding target driving trajectory. This effectively avoids the problem of poor correction effect caused by the use of manually set rule-based trajectory optimization strategies in the prior art, and effectively improves the accuracy and effectiveness of driving trajectory correction in complex driving scenarios.

[0088] In one embodiment, Figure 2This is a flowchart of a training method for a target trajectory generation model provided by an embodiment of the present invention. This embodiment is an optimization based on the above embodiment. In this embodiment, the training method of the target trajectory generation model is as follows: acquiring a perception sample set, a regulation control sample set, and an initial trajectory generation model; iteratively training the parameters of the initial trajectory generation model based on the perception sample set, and determining the trained initial trajectory generation model as the first model; iteratively training the parameters of the first model based on the regulation control sample set, and determining the trained first model as the second model; iteratively training the parameters of the second model based on the perception sample set and the regulation control sample set, and determining the trained second model as the target trajectory generation model. Figure 2 As shown, the method specifically includes the following steps:

[0089] S201. Obtain the perception sample set, the control sample set, and the initial trajectory generation model.

[0090] The perception sample set can include multiple driving-related samples carrying labels. For example, it could include multiple driving-related samples and obstacle and road structure labels carried by each driving-related sample, where the road structure labels can be lane line labels. The perception sample set can also include multiple driving-related samples carrying labels and multiple driving-related samples without labels. The multiple driving-related samples carrying labels carry obstacle and road structure labels. It should be noted that if the perception sample set includes multiple driving-related samples with and without labels, the training of the initial trajectory generation model can be divided into two stages. The first stage can use supervised learning to train the initial trajectory generation model based on the multiple driving-related samples with labels. The second stage uses reinforcement learning to train the model based on the multiple driving-related samples without labels. This method of training with supervised learning followed by reinforcement learning can improve performance, and since some driving-related samples do not need to carry labels, it can also improve the generation efficiency of the perception sample set. It should be noted that the method of training based on multiple driving-related samples without labels can be achieved using existing reinforcement learning methods, which will not be elaborated here.

[0091] Because different drivers have different driving styles, and even the same driver's driving style varies at different times, it is necessary to learn the input-output causal relationship in addition to the input-output correlation. This invention employs supervised learning for initial training and reinforcement learning for later training, which enhances the learning of causal relationships and makes the predicted trajectory more consistent with the driver's driving style.

[0092] Obstacle labels can include: first-type obstacle labels and second-type obstacle labels.

[0093] The regulatory control sample set includes multiple driving-related samples with labels. For example, it could be multiple driving-related samples and the trajectory of each driving-related sample at the next time step. The regulatory control sample set can also include multiple driving-related samples with labels and multiple driving-related samples without labels. It should be noted that if the regulatory control sample set includes multiple driving-related samples with labels and multiple driving-related samples without labels, the training of the first model can be divided into two stages. The first stage can be trained using supervised learning methods based on the multiple driving-related samples with labels. The second stage can be trained using reinforcement learning methods based on the multiple driving-related samples without labels. The training method of first supervised learning and then reinforcement learning can improve the performance, and since some driving-related samples do not need to carry labels, it can also improve the efficiency of generating the perception sample set. It should be noted that the method of training based on multiple driving-related samples without labels using reinforcement learning methods can be implemented using existing reinforcement learning methods, which will not be elaborated here. It should be noted that the reinforcement learning method provided in this embodiment of the invention is a reinforcement learning method performed through comparison. Reinforcement learning is learning through interaction.

[0094] The initial trajectory generation model may include: an initial backbone network, an initial encoder, an initial decoder, and a target memory module. Alternatively, the initial trajectory generation model may include: an initial backbone network, an initial encoder, and an initial decoder.

[0095] Specifically, the methods for obtaining the perception sample set, the planning and control sample set, and the initial trajectory generation model are as follows: Obtain historical driving-related data and generate multiple driving-related samples based on this data. Label these driving-related samples, and generate the perception sample set based on the driving-related samples with added obstacle and road structure labels. Generate the planning and control sample set based on the driving-related samples with the trajectory for the next time step. Create the initial trajectory generation model.

[0096] S202. Based on the perception sample set, iteratively train the parameters of the initial trajectory generation model, and determine the trained initial trajectory generation model as the first model.

[0097] The parameters of the initial trajectory generation model include: the parameters of the initial backbone network, the parameters of the initial encoder, and the parameters of the initial decoder. The perception sample set includes: multiple driving-related samples and obstacle labels and road structure labels carried by each driving-related sample.

[0098] Specifically, the method of iteratively training the parameters of the initial trajectory generation model based on the perception sample set and determining the trained initial trajectory generation model as the first model can be as follows: Input driving-related samples from the perception sample set into the initial trajectory generation model to obtain first predicted obstacle information and first predicted road structure; train the parameters of the initial trajectory generation model based on the differences between the first predicted obstacle information and obstacle labels, and the differences between the first predicted road structure and road structure labels, and determine the trained initial trajectory generation model as the first model. Alternatively, the method of iteratively training the parameters of the initial trajectory generation model based on the perception sample set and determining the trained initial trajectory generation model as the first model can be as follows: Input driving-related samples from the perception sample set into the initial trajectory generation model to obtain first predicted obstacle information and first predicted lane lines; train the parameters of the initial trajectory generation model based on the differences between the first predicted obstacle information and obstacle labels, and the differences between the first predicted lane lines and lane line labels, and determine the trained initial trajectory generation model as the first model.

[0099] Optionally, the perception sample set includes: multiple driving-related samples and obstacle labels and road structure labels carried by each driving-related sample;

[0100] Based on the perceptual sample set, the parameters of the initial trajectory generation model are iteratively trained, and the trained initial trajectory generation model is determined as the first model, including:

[0101] Input driving-related samples from the perception sample set into the initial trajectory generation model to obtain the first predicted obstacle information and the first predicted road structure.

[0102] Based on the differences between the first predicted obstacle information and the obstacle label, and the differences between the first predicted road structure and the road structure label, the parameters of the initial trajectory generation model are trained, and the trained initial trajectory generation model is determined as the first model.

[0103] It should be noted that when training the initial trajectory generation model, it is trained only based on obstacle information and road structure. The weights of the loss function corresponding to the predicted trajectory can be set to zero.

[0104] Specifically, the initial trajectory generation model includes an initial backbone network, an initial encoder, and an initial decoder. Driving-related samples include sensor data samples and navigation planning samples. The sensor data samples include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the first predicted obstacle information and the first predicted road structure can be as follows: input the environmental perception samples into the initial backbone network to obtain the first fusion feature; input the state samples and navigation planning samples into the initial encoder to obtain the first encoded feature; and input the first encoded feature and the first fusion feature into the initialized initial decoder to obtain the first predicted obstacle and the first predicted road structure.

[0105] The initial trajectory generation model includes an initial backbone network, an initial encoder, and an initial decoder. Driving-related samples include sensor data samples, which in turn include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the first predicted obstacle information and the first predicted road structure can be as follows: input the environmental perception samples into the initial backbone network to obtain the first fusion feature; input the state samples into the initial encoder to obtain the first encoded feature; and input the first encoded feature and the first fusion feature into the initialized initial decoder to obtain the first predicted obstacle and the first predicted road structure.

[0106] The initial trajectory generation model includes an initial backbone network, an initial encoder, an initial decoder, and a target memory module. Driving-related samples include sensor data samples and navigation planning samples. The sensor data samples include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the first predicted obstacle information and the first predicted road structure can be as follows: input the environmental perception samples into the initial backbone network to obtain the first fusion feature; determine the second fusion feature based on the first fusion feature and the historical features output by the target memory module; input the state samples and navigation planning samples into the initial encoder to obtain the first encoded feature; and input the first encoded feature and the second fusion feature into the initialized initial decoder to obtain the first predicted obstacle and the first predicted road structure.

[0107] The initial trajectory generation model includes an initial backbone network, an initial encoder, an initial decoder, and a target memory module. Driving-related samples include sensor data samples, which in turn include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the first predicted obstacle information and the first predicted road structure can be as follows: input the environmental perception samples into the initial backbone network to obtain the first fusion feature; determine the second fusion feature based on the first fusion feature and the historical features output by the target memory module; input the state samples into the initial encoder to obtain the first encoded feature; and input the first encoded feature and the second fusion feature into the initialized initial decoder to obtain the first predicted obstacle and the first predicted road structure.

[0108] Specifically, based on the differences between the first predicted obstacle information and obstacle labels, and the differences between the first predicted road structure and road structure labels, the parameters of the initial trajectory generation model can be trained as follows: based on the first loss function, the differences between the first predicted obstacle information and obstacle labels, and the differences between the first predicted road structure and road structure labels, the parameters of the initial trajectory generation model are trained. The first loss function can be any loss function in the prior art, and this embodiment of the invention does not impose any limitations on it.

[0109] Optionally, the initial trajectory generation model includes: an initial backbone network, an initial encoder, an initial decoder, and a target memory module;

[0110] Inputting driving-related samples from the perception sample set into the initial trajectory generation model yields first predicted obstacle information and first predicted road structure information, including:

[0111] The initial decoder is initialized based on a preset instance;

[0112] Environmental perception samples are input into the initial backbone network to obtain the first fusion feature, and the first fusion feature is projected onto the BEV space.

[0113] Based on the first fusion feature projected onto the BEV space and the BEV feature output by the target memory module, the first BEV feature is determined, and the BEV feature stored in the target memory module is updated based on the first BEV feature.

[0114] Input the state samples and navigation planning samples into the initial encoder to obtain the first encoded features;

[0115] The first encoded feature and the first BEV feature are input into the initialized initial decoder to obtain the first predicted obstacle information and the first predicted road structure.

[0116] The preset instance is used to initialize the initial decoder. For example, if it is necessary to predict obstacle information, road structure, and trajectory, the initial information of obstacle information, road structure, and trajectory is initialized in advance. During training, after the feature information is input into the decoder, it will gradually evolve into the actual obstacle information, road structure, and trajectory.

[0117] S203. Based on the regulatory sample set, iteratively train the parameters of the first model, and determine the trained first model as the second model.

[0118] Optionally, the control sample set includes: multiple driving-related samples and the trajectory of each driving-related sample at the next moment;

[0119] Based on the regulatory sample set, the parameters of the first model are trained iteratively, and the trained first model is determined as the second model, including:

[0120] Input the driving-related samples from the control sample set into the first model to obtain the first predicted trajectory;

[0121] Based on the difference between the first predicted trajectory and the trajectory at the next moment, the parameters of the first model are trained, and the trained first model is determined as the second model.

[0122] It should be noted that when training the first model, it is trained only based on the trajectory. The weights of the loss functions corresponding to obstacle information and lane line information can be set to zero.

[0123] The initial trajectory generation model includes an initial backbone network, an initial encoder, and an initial decoder. Driving-related samples include sensor data samples and navigation planning samples. Sensor data samples include environmental perception samples and state samples. The method to input the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the first predicted trajectory can be as follows: input the environmental perception samples into the initial backbone network to obtain the first fusion feature; input the state samples and navigation planning samples into the initial encoder to obtain the first encoded feature; and input the first encoded feature and the first fusion feature into the initialized initial decoder to obtain the first predicted trajectory.

[0124] The initial trajectory generation model includes an initial backbone network, an initial encoder, and an initial decoder. Driving-related samples include sensor data samples, which in turn include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the first predicted trajectory can be as follows: input the environmental perception samples into the initial backbone network to obtain the first fused features; input the state samples into the initial encoder to obtain the first encoded features; and input the first encoded features and the first fused features into the initialized initial decoder to obtain the first predicted trajectory.

[0125] The initial trajectory generation model includes an initial backbone network, an initial encoder, an initial decoder, and a target memory module. Driving-related samples include sensor data samples and navigation planning samples. The sensor data samples include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the first predicted trajectory is as follows: input the environmental perception samples into the initial backbone network to obtain the first fusion feature; determine the second fusion feature based on the first fusion feature and the historical features output by the target memory module; input the state samples and navigation planning samples into the initial encoder to obtain the first encoding feature; and input the first encoding feature and the second fusion feature into the initialized initial decoder to obtain the first predicted trajectory.

[0126] The initial trajectory generation model includes an initial backbone network, an initial encoder, an initial decoder, and a target memory module. Driving-related samples include sensor data samples, which in turn include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the first predicted trajectory can be as follows: input the environmental perception samples into the initial backbone network to obtain the first fusion feature; determine the second fusion feature based on the first fusion feature and the historical features output by the target memory module; input the state samples into the initial encoder to obtain the first encoded feature; and input the first encoded feature and the second fusion feature into the initialized initial decoder to obtain the first predicted trajectory.

[0127] Specifically, based on the difference between the first predicted trajectory and the trajectory at the next time step, the parameters of the first model can be trained by using a second loss function and the difference between the first predicted trajectory and the trajectory at the next time step. The second loss function can be any loss function existing in the art, and this embodiment of the invention does not impose any limitations on it.

[0128] S204. Based on the perception sample set and the control sample set, iteratively train the parameters of the second model, and determine the trained second model as the target trajectory generation model.

[0129] Specifically, the method for iteratively training the parameters of the second model based on the perception sample set and the regulatory control sample set, and determining the trained second model as the target trajectory generation model, can be as follows: Generate a fusion sample set based on the perception sample set and the regulatory control sample set; input the driving-related samples in the fusion sample set into the second model to obtain the second predicted obstacle, the second predicted road structure, and the second predicted trajectory; train the parameters of the second model based on the differences between the second predicted obstacle and the obstacle label, the differences between the second predicted road structure and the road structure label, and the differences between the second predicted trajectory and the trajectory at the next moment; and determine the trained second model as the target trajectory generation model.

[0130] Optionally, based on the perception sample set and the regulatory sample set, the parameters of the second model are iteratively trained, and the trained second model is determined as the target trajectory generation model, including:

[0131] A fusion sample set is generated based on the perception sample set and the planning and control sample set. The fusion sample set includes: multiple driving-related samples and the obstacle label, road structure label, and trajectory of each driving-related sample at the next moment.

[0132] Input the driving-related samples from the fusion sample set into the second model to obtain the second predicted obstacle, the second predicted road structure, and the second predicted trajectory.

[0133] Based on the differences between the second predicted obstacle and the obstacle label, the differences between the second predicted road structure and the road structure label, and the differences between the second predicted trajectory and the trajectory at the next moment, the parameters of the second model are trained, and the trained second model is determined as the target trajectory generation model.

[0134] Specifically, the method for generating a fused sample set based on the perception sample set and the planning control sample set can be as follows: Add the obstacle labels and road structure labels corresponding to the driving-related samples in the perception sample set to the corresponding driving-related samples in the planning control sample set; the planning control sample set after adding the obstacle labels and road structure labels is then determined as the fused sample set. Alternatively, the method can be as follows: Add the trajectory of the next moment corresponding to the driving-related samples in the planning control sample set to the corresponding driving-related samples in the perception sample set; the perception sample set after adding the trajectory of the next moment is then determined as the fused sample set.

[0135] The initial trajectory generation model includes an initial backbone network, an initial encoder, and an initial decoder. Driving-related samples include sensor data samples and navigation planning samples. The sensor data samples include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the second predicted obstacle information, the second predicted road structure, and the second predicted trajectory can be as follows: input the environmental perception samples into the initial backbone network to obtain the first fusion feature; input the state samples and navigation planning samples into the initial encoder to obtain the first encoded feature; input the first encoded feature and the first fusion feature into the initialized initial decoder to obtain the second predicted obstacle information, the second predicted road structure, and the second predicted trajectory.

[0136] The initial trajectory generation model includes an initial backbone network, an initial encoder, and an initial decoder. Driving-related samples include sensor data samples, which in turn include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the second predicted obstacle information, the second predicted road structure, and the second predicted trajectory can be as follows: inputting environmental perception samples into the initial backbone network to obtain the first fusion feature; inputting state samples into the initial encoder to obtain the first encoded feature; and inputting the first encoded feature and the first fusion feature into the initialized initial decoder to obtain the second predicted obstacle information, the second predicted road structure, and the second predicted trajectory.

[0137] The initial trajectory generation model includes an initial backbone network, an initial encoder, an initial decoder, and a target memory module. Driving-related samples include sensor data samples and navigation planning samples. The sensor data samples include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the second predicted obstacle information, the second predicted road structure, and the second predicted trajectory can be as follows: Input the environmental perception samples into the initial backbone network to obtain the first fusion feature; determine the second fusion feature based on the first fusion feature and the historical features output by the target memory module; input the state samples and navigation planning samples into the initial encoder to obtain the first encoding feature; input the first encoding feature and the second fusion feature into the initialized initial decoder to obtain the second predicted obstacle information, the second predicted road structure, and the second predicted trajectory.

[0138] The initial trajectory generation model includes an initial backbone network, an initial encoder, an initial decoder, and a target memory module. Driving-related samples include sensor data samples, which in turn include environmental perception samples and state samples. The method for inputting the driving-related samples from the perception sample set into the initial trajectory generation model to obtain the second predicted obstacle information, the second predicted road structure, and the second predicted trajectory can be as follows: Environmental perception samples are input into the initial backbone network to obtain a first fusion feature; based on the first fusion feature and the historical features output by the target memory module, a second fusion feature is determined; state samples are input into the initial encoder to obtain a first encoding feature; and the first encoding feature and the second fusion feature are input into the initialized initial decoder to obtain the second predicted obstacle information, the second predicted road structure, and the second predicted trajectory.

[0139] Specifically, based on the differences between the second predicted obstacles and obstacle labels, the differences between the second predicted road structure and road structure labels, and the differences between the second predicted trajectory and the trajectory at the next time step, the parameters of the second model can be trained as follows: A third loss function is determined based on the first and second loss functions; the parameters of the second model are then trained based on the third loss function, the differences between the second predicted obstacles and obstacle labels, the differences between the second predicted road structure and road structure labels, and the differences between the second predicted trajectory and the trajectory at the next time step. Alternatively, the parameters of the second model can also be trained based on the differences between the second predicted obstacles and obstacle labels, the differences between the second predicted road structure and road structure labels, and the differences between the second predicted trajectory and the trajectory at the next time step. The third loss function can be any loss function in the prior art, and this embodiment of the invention does not impose any limitations on it.

[0140] Optionally, driving-related samples include: sensor data samples and navigation planning samples. Sensor data samples include: environmental perception samples and state samples. Environmental perception samples include: frame samples and point cloud samples.

[0141] In a specific example Figure 3 A flowchart illustrating another trajectory generation method provided in an embodiment of the present invention. For example... Figure 3As shown, the target trajectory generation model includes a target backbone network, a target encoder, a target decoder, and a target memory module. Frame data from the camera and point cloud data from the LiDAR are input into the target backbone network to obtain target fusion features, which are then projected into the BEV space. Based on the projected target fusion features and the BEV features output by the target memory module, the target BEV features are determined, and the BEV features stored in the target memory module are updated accordingly. The current vehicle position information from GPS, the current vehicle attitude information from sensors, navigation planning information, and driving rule information are input into the target encoder to obtain target encoded features. The target encoded features and target BEV features are input into the target decoder to obtain the candidate driving trajectory, first obstacle information, second obstacle information, and road structure corresponding to the current vehicle. It should be noted that the target trajectory generation model is a pre-trained initial trajectory generation model (the initial trajectory generation model includes: an initial backbone network, an initial encoder, an initial decoder, and a target memory module). The training methods for the initial trajectory generation model include: label-based supervised learning training and label-and-reward-based training (i.e., supervised learning in the first half and reinforcement learning in the second half). Here, the reward is usually denoted as Rt, representing the reward value returned at time step t. For example, the training of the initial trajectory generation model can be divided into two stages: in the first stage, the parameters of the initial trajectory generation model are trained using label-based supervised learning; in the second stage, the model obtained after the first stage of training is trained using the method of supervised learning in the first half and reinforcement learning in the second half. Specifically, the first stage of the training process is as follows: acquire a perception sample set, which includes: multiple driving-related samples and the first type of obstacle label, the second type of obstacle label, and the road structure label carried by each driving-related sample. The driving-related samples include: sensor samples and navigation planning samples. The sensor data samples include: environmental perception samples and state samples. The environmental perception samples include: frame samples and point cloud samples. The state samples include: position samples and attitude samples.The initial decoder is initialized based on a preset instance. Frame samples and point cloud samples are input into the initial backbone network to obtain the first fused feature, which is then projected onto the BEV space. Based on the first fused feature projected into the BEV space and the BEV feature output by the target memory module, the first BEV feature is determined, and the BEV feature stored in the target memory module is updated based on the first BEV feature. Position samples, attitude samples, and navigation planning samples are input into the initial encoder to obtain the first encoded feature. The first encoded feature and the first BEV feature are input into the initialized initial decoder to obtain the first predicted first type of obstacle information, the first predicted second type of obstacle information, and the first predicted road structure. Based on the differences between the first predicted first type of obstacle information and the first type of obstacle label, the differences between the first predicted second type of obstacle information and the second type of obstacle label, and the differences between the first predicted road structure and the road structure label, the parameters of the initial trajectory generation model are trained, and the trained initial trajectory generation model is determined as the first model.

[0142] A control sample set is acquired, comprising labeled and unlabeled samples, with the label representing the trajectory at the next time step. The first model includes a first backbone network, a first encoder, a first decoder, and a target memory module. The first half of the second-stage training process involves: inputting frame samples and point cloud samples into the first backbone network to obtain second fused features; projecting the second fused features into the BEV space; determining the second BEV features based on the second fused features projected into the BEV space and the BEV features output by the target memory module; updating the BEV features stored in the target memory module based on the second BEV features; inputting position samples, attitude samples, and navigation planning samples into the first encoder to obtain second encoded features; inputting the second encoded features and the second BEV features into the first decoder to obtain the first predicted trajectory; and training the parameters of the first model based on the difference between the first predicted trajectory and the trajectory at the next time step, resulting in the model trained in the first half.

[0143] The model trained in the first half includes: a second backbone network, a second encoder, a second decoder, and a target memory module;

[0144] The latter half of the second phase of training is as follows: Frame samples and point cloud samples are input into the second backbone network to obtain the third fused feature, which is then projected into the BEV space. Based on the third fused feature projected into the BEV space and the BEV feature output by the target memory module, the third BEV feature is determined, and the BEV feature stored in the target memory module is updated accordingly. Position samples, pose samples, and navigation planning samples are input into the second encoder to obtain the third encoded feature. The third encoded feature and the third BEV feature are input into the second decoder to obtain the first predicted trajectory. The model trained in the first half is then trained again based on the first predicted trajectory and the returned reward value. After the second half of training is completed, the target trajectory generation model is obtained.

[0145] The technical solution of this embodiment divides the training of the target trajectory generation model into three stages: First stage: Based on the perception sample set, iteratively train the parameters of the initial trajectory generation model, and determine the trained initial trajectory generation model as the first model; Second stage: Based on the traffic control sample set, iteratively train the parameters of the first model, and determine the trained first model as the second model; Third stage: Based on the perception sample set and the traffic control sample set, iteratively train the parameters of the second model, and determine the trained second model as the target trajectory generation model. After obtaining the target trajectory generation model, the driving-related data of the current vehicle is input into the target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle. This reduces information loss and thus improves the accuracy of the predicted trajectory. Based on the candidate driving trajectory output by the target trajectory generation model, parallel lateral and longitudinal control of the vehicle is possible, rather than serial control, i.e., left-right control and forward-backward control are not separated. This allows the vehicle to change lanes or avoid obstacles more smoothly during driving.

[0146] In one embodiment, Figure 4 This is a flowchart of another trajectory generation method provided by an embodiment of the present invention. This embodiment further explains the generation process of the target driving trajectory based on the above embodiments. In this embodiment, the target driving correction model includes: a target streaming encoder, a target navigation encoder, a target modality alignment module, and a target driving decision model. The target streaming encoder is used to encode the video stream of the current vehicle; the target navigation encoder is used to encode the navigation planning information of the current vehicle; the target modality alignment module is used for feature space unification / alignment (mapping to text feature space) of multimodal information; and the target driving decision model is used to output corresponding driving correction information.

[0147] The first driving-related data includes: sensor information and navigation planning information. The sensor information includes: environmental perception information and status information. The target trajectory generation model includes: target backbone network, target encoder, target decoder and target memory module. The target memory module is used to store BEV features in the time dimension and spatial dimension.

[0148] The second driving-related data includes: environmental perception information, navigation planning information, and driving prompt information.

[0149] like Figure 4 As shown, the method includes:

[0150] S410: Obtain the first driving-related data and the second driving-related data corresponding to the current vehicle.

[0151] S420. Input environmental perception information into the target backbone network to obtain target fusion features, and project the target fusion features onto the BEV space.

[0152] The environmental perception information includes current frame data and current point cloud data. Specifically, the environmental perception information is input into the target backbone network to obtain target fusion features, and the target fusion features are projected onto the BEV space, including: inputting the current frame data and current point cloud data into the target backbone network to obtain target fusion features, and projecting the target fusion features onto the BEV space.

[0153] It should be noted that projecting the target fusion features into the BEV space can be achieved using projection methods in existing technologies.

[0154] Specifically, the target BEV feature can be determined by superimposing the target fusion feature projected into the BEV space and the BEV feature output by the target memory module. The resulting BEV feature is then determined as the target BEV feature. It should be noted that when superimposing the target fusion feature projected into the BEV space and the BEV feature output by the target memory module, the weights of the target fusion feature projected into the BEV space and the BEV feature output by the target memory module can be preset. Then, based on these weights, a weighted sum is performed on the target fusion feature projected into the BEV space and the BEV feature output by the target memory module.

[0155] S430. Based on the target fusion features projected into the BEV space and the BEV features output by the target memory module, determine the target BEV features, and update the BEV features stored in the target memory module according to the target BEV features.

[0156] Specifically, updating the BEV features stored in the target memory module based on the target BEV features can be done in two ways: First, according to the storage rules corresponding to the target memory module, store the target BEV features in the target memory module, and then delete some / all of the historically stored BEV features from the target memory module according to the storage rules. Alternatively, updating the BEV features stored in the target memory module based on the target BEV features can be done as follows: delete all historically stored BEV features from the target memory module, and then store some / all of the target BEV features in the target memory module according to the storage rules corresponding to the target memory module.

[0157] S440. Input the status information and navigation planning information into the target encoder to obtain the target coding features.

[0158] The status information may include the current vehicle's position information and the current vehicle's attitude information. Specifically, the status information and navigation planning information are input into the target encoder to obtain target coding features, including: inputting the current vehicle's position information, the current vehicle's attitude information, and the navigation planning information into the target encoder to obtain target coding features.

[0159] S450. Input the target encoded features and target BEV features into the target decoder to obtain the candidate driving trajectory corresponding to the current vehicle.

[0160] The target backbone network can be a convolutional neural network (CNN), and the target encoder can be a Transformer model, which is a deep learning model based on an attention mechanism.

[0161] It's important to note that projecting the target fusion features into the BEV space ensures that all vehicle features are projected onto the same dimension. Since different vehicles have different heights, and the positions of cameras and LiDAR sensors may also differ, projecting the target fusion features into the BEV space ensures that all vehicle features are projected onto the same dimension, eliminating the need to consider vehicle height, camera position, or LiDAR position. This, in turn, improves the model's training speed.

[0162] To enhance the model's representational capabilities, the target trajectory generation model in this embodiment includes a target memory module, which stores BEV features in both the temporal and spatial dimensions. For example, the target memory module can store BEV features from 20 seconds ago, as well as BEV features within a 200-meter distance range. Recording only one dimension of BEV features may result in incomplete information, thus affecting the accuracy of the predicted trajectory. For instance, if the vehicle is parked for several minutes, the BEV features stored in the target memory model will all be those of the parked state. Predicting the trajectory based on these parked BEV features will negatively impact the accuracy of the predicted trajectory. In this embodiment, the target memory module stores BEV features in both the temporal and spatial dimensions, ensuring information integrity and preventing the aforementioned issues, thereby further improving the accuracy of the predicted trajectory.

[0163] In this embodiment of the invention, the target trajectory generation model includes: a target backbone network, a target encoder, a target decoder, and a target memory module. The target memory module memorizes intermediate temporal features, including not only time-dimensional memory but also spatial-dimensional memory. This not only enables the model to cope with special scenarios but also allows it to directly associate actions with input information, generating more refined and human-like actions.

[0164] S460. Input the environmental perception information into the target streaming encoder to obtain the corresponding image labeling information.

[0165] The target streaming encoder includes a target image feature extractor and a target temporal fusion module. For example, the target streaming encoder can also be called a Streaming Video Encoder; the target image feature extractor can be a ViT suitable for multi-resolution image processing, such as a Multi-resolution ViT (Multi-resolution ViT), a visual processing model based on the Transformer architecture, which can capture the relationships and features between different parts of an image by utilizing multi-head attention mechanisms and other operations; the target temporal fusion module can be a Temporal Encoder, used to capture the relationships between image information at different time points.

[0166] In one embodiment, S460 includes S4601-S4602:

[0167] S4601. Input the environmental perception information into the target image feature extractor to obtain the corresponding environmental perception coding information.

[0168] Environmental awareness information can be presented as a video stream or as a series of consecutive frames. Environmental awareness coding information refers to the relevant information obtained by feature extraction and encoding of environmental awareness information as images or video streams. For example, environmental awareness coding information can include global features and image features for each frame. For instance, if the environmental awareness information consists of N frames, inputting these N frames into a target image feature extractor can output N*H*W image features and N global features. The image features focus more on the corresponding location within the image; the global features focus more on key features within the image.

[0169] In this embodiment, when the environmental perception information is presented as a video stream, the video stream needs to be divided into multiple frames. These frames are then input into a target image feature extractor to obtain the corresponding environmental perception encoding information. During the encoding of the environmental perception information, a class token can be added simultaneously to introduce global information, thereby improving model performance.

[0170] S4602. Input the environmental perception coding information into the target temporal fusion module to obtain the corresponding image labeling information.

[0171] Here, image tagging information refers to high-dimensional visual features (which can be simply referred to as image tokens) representing local and global information obtained from environmental perception information through the feature extraction module. The target temporal fusion module is implemented based on pooling layers and temporal attention mechanisms. That is, on the basis of pooling, an SE structure is added, which can perform weighted fusion of multiple frames of images in a temporal sequence, effectively reducing the number of tags and thus improving training and inference speeds. At the same time, classification tokens (class tokens) can be added to introduce global information, thereby improving model performance. In the embodiment, when the environmental perception information consists of N frames of images, N*H*W image features and N global features can be used as input data for the target temporal fusion module to obtain the corresponding image tagging information, i.e., H*W+N.

[0172] S470. Input the navigation planning information into the target navigation encoder to obtain the corresponding navigation mark information.

[0173] The navigation planning information can be presented as an image; navigation token information refers to the high-dimensional visual features (which can be called navigation tokens) representing local and global information obtained by the feature extraction module from the navigation planning information; the target navigation encoder is used to convert the input navigation planning information presented as an image into a series of navigation token information. This navigation token information can capture the feature information in the navigation planning information. For example, the target navigation encoder can be a ViT Encoder (Vision Transformer Encoder). Generally speaking, the ViT Encoder can better process and understand image information, which can improve the perception and decision-making capabilities of autonomous driving systems in complex scenarios. Furthermore, the ViT Encoder can capture local features and global information in the image, providing more valuable image information for subsequent processing steps.

[0174] S480. Input the image marker information and navigation marker information as driving marker information into the target modality alignment module to obtain the mapped multimodal features.

[0175] The target modality alignment module maps video and navigation features to a text feature space, enabling unified processing and interaction. The mapped multimodal features refer to those mapped to the text feature space (aligned with text features), converting image and navigation tag information into features with similar form and dimensions to the text representation in the LLM model. In this embodiment, the image and navigation tag information output by the target streaming encoder can be mapped to have similar form and dimensions to the text representation in the LLM model, allowing for fusion, interaction, and decoding with text features in subsequent processing.

[0176] S490. Input the driving prompt information into the pre-created text feature extraction module to obtain the corresponding prompt text features.

[0177] The text feature extraction module includes a tokenizer and a vocabulary. Driving prompts can include general prompt text and scenario-specific prompt text. When the driving prompts are text data, the prompt text features refer to the process of splitting the text into smaller word units (sub-words) using the tokenizer and indexing them into corresponding text features (word embeddings) using the vocabulary. In this embodiment, the text feature extraction module can uniformly convert different types of driving prompts into a sequence of tokens, which can then be input into the target driving decision model for processing and interaction, thereby enabling the planning of the vehicle's driving trajectory.

[0178] S4100. Input the prompt text features corresponding to the driving prompt information and the mapped multimodal features into the target driving decision model to obtain the corresponding driving correction information.

[0179] In this embodiment, the target driving decision model is used to determine the specific scenario in which the current vehicle is located based on driving prompt information, and to obtain suggestions for correcting the current vehicle's driving trajectory based on environmental perception information and navigation planning information. For example, the target driving decision model can be a large language model (LLM model for short).

[0180] In one embodiment, S4100 includes S41001-S41002:

[0181] S41001. Store the environmental perception coding information, prompt text features, and driving correction information of the historical frames into the memory storage space.

[0182] The memory storage space is a space that uses a memory mechanism to store scene information and follows a queue mechanism, i.e., first-in, first-out (FIFO). In this embodiment, the available storage space of the memory storage space is limited, meaning it is used to store image-related information of a fixed length. For example, it can be used to store information related to N frames of images, namely, the environmental perception coding information, prompt text features, and driving correction information corresponding to the N frames of images. In the actual storage process, before storing the environmental perception coding information, prompt text features, and driving correction information corresponding to the N+1th frame of image into the memory storage space, the environmental perception coding information, prompt text features, and driving correction information corresponding to the first frame of image in the memory storage space must first be deleted from the memory storage space. Then, the environmental perception coding information, prompt text features, and driving correction information corresponding to the N+1th frame of image are stored as the environmental perception coding information, prompt text features, and driving correction information corresponding to the Nth frame of image in the memory storage space.

[0183] S41002. Extract the environmental perception coding information, prompt text features, and driving correction information of the current frame from the memory storage space, and fuse the prompt text features of the current frame with the prompt text features and driving correction information of the historical frames, and fuse the environmental perception coding information of the current frame with the environmental perception coding information of the historical frames, and input the fused information into the target driving decision model to obtain the corresponding driving correction information.

[0184] In this embodiment, a cross-attention mechanism can be used to fuse the prompt text features of the current frame and the prompt text features of historical frames, as well as the environmental awareness coding information of the current frame and the environmental awareness coding information of historical frames, and the driving correction information of the current frame and the driving correction information of historical frames. Specifically, the prompt text features of the current frame are grouped into a sequence as a query term, and the prompt text features of historical frames, driving correction information, and driving prompt information are grouped into a sequence as key and value. An attention mechanism is used to model the relationship between the sequences and aggregate effective information to obtain the fused prompt text features. In one example, this is implemented based on a pooling layer and a temporal attention mechanism, that is, adding an SE structure on top of Pooling, which can perform weighted fusion of the environmental awareness coding information of the current frame and the environmental awareness coding information of historical frames.

[0185] S4110. Correct the candidate driving trajectory based on the driving correction information to obtain the corresponding target driving trajectory.

[0186] The technical solution of this embodiment, based on the above embodiments, configures a target memory module in the target trajectory generation model to store BEV features in the time and space dimensions to ensure the integrity of information and thus improve the accuracy of trajectory prediction. At the same time, the target driving correction model is obtained by iteratively training through specific scenario prompt text samples, which can ensure that the target driving correction model can accurately provide driving correction information for different complex driving scenarios, thereby ensuring the driving safety of the vehicle.

[0187] In one embodiment, Figure 5 This is a flowchart illustrating the training process of a target driving correction model provided in an embodiment of the present invention. This embodiment describes the iterative training process of the driving correction model used in the above-mentioned information correction process. By iteratively training a pre-created initial driving correction model, the corresponding target driving correction model can be obtained.

[0188] like Figure 5 As shown, the training process of the target driving correction model includes:

[0189] S510, Obtain a driving-related sample set.

[0190] The driving-related sample set includes driving-related samples from multiple vehicles. Each driving-related sample refers to various information generated or collected by each vehicle during its operation, characterizing the surrounding environment, the vehicle's own state, and driving operations. In this embodiment, driving-related samples for multiple vehicles can be obtained from a sample database to form a corresponding driving-related sample set. The driving-related sample set includes: an environmental perception sample set, a navigation planning sample set, and a driving prompt sample set; the driving prompt sample set includes: a general prompt text sample set and a scenario-specific prompt text sample set. The environmental perception sample set includes environmental perception samples from multiple vehicles. A single environmental perception sample refers to a single frame of information collected by a vehicle during its driving process, which can characterize the surrounding environment of the vehicle. The navigation planning sample set includes navigation planning samples from multiple vehicles. Each navigation planning sample refers to a vehicle's real-time geographical location, map data, and route planning information. The driving prompt sample set includes driving prompt samples from multiple vehicles. Each driving prompt sample refers to a textual instruction that defines and standardizes the model's output. It is a textual input that requires the model to describe a vehicle's current driving environment and identify key objects, and is used to guide the target driving correction model in thinking and reasoning.

[0191] S520. Based on the driving-related sample set, the pre-constructed initial driving correction model is iteratively trained to obtain the corresponding target driving correction model.

[0192] The initial driving correction model refers to an untrained driving correction model, such as a VLM model, in which case the initial driving correction model is the initial VLM model. In one embodiment, the initial driving correction model includes: an initial streaming encoder, an initial transform encoder, an initial modal alignment module, and an initial driving decision model. The initial streaming encoder is an untrained streaming encoder, such as a Multi-resolution ViT encoder, in which case the initial streaming encoder is the initial Multi-resolution ViT. The initial transform encoder is an untrained transform encoder, such as a ViT Encoder, in which case the initial transform encoder is the initial ViT Encoder. The initial modal alignment module is an untrained modal alignment module, such as a Projection module, in which case the initial modal alignment module is the initial Projection. The initial driving decision model is an untrained driving decision model, such as an LLM model, in which case the initial driving decision model is the initial LLM model.

[0193] The training process of the target driving correction model includes three stages: In the first stage, the parameters of the initial modal alignment module are iteratively trained, and the parameters of the initial streaming encoder and the initial driving decision model remain unchanged, and the initial transform encoder is not set; In the second stage, the parameters of the initial streaming encoder, the initial driving decision model, and the intermediate modal alignment module are iteratively trained, and the initial transform encoder is not set; In the third stage, the initial transform encoder, as well as the candidate streaming encoder, the candidate driving decision model, and the candidate modal alignment module, are iteratively trained to obtain the corresponding target navigation encoder, target streaming encoder, target driving decision model, and target modal alignment module, which constitute the corresponding target driving correction model.

[0194] In one embodiment, S520 includes S5201-S5203:

[0195] S5201. Based on the environmental perception sample set and the prompt text sample set, the initial modal alignment module in the initial driving correction model is iteratively trained to obtain the corresponding intermediate modal alignment module.

[0196] The environmental perception sample set and the prompt text sample set are derived from video / image-text pairs constructed from publicly available datasets. In the first stage of training, these sets are considered general data. During this stage, the environmental perception sample set is input into the initial streaming encoder, and the text features of the corresponding prompt text information are input into the initial driving decision model. The parameters of the initial modality alignment module are iteratively trained to obtain the corresponding intermediate modality alignment module. During training, the parameters of the initial modality alignment module need to be adjusted to align the image features with the word embedding space of the pre-trained driving decision model. This achieves effective fusion and interaction of image and text features, helping the driving decision model better understand and process multimodal information, thus achieving multimodal feature alignment.

[0197] S5202. Based on the environmental perception sample set and the prompt text sample set, iteratively train the initial streaming encoder and initial driving decision model in the initial driving correction model, as well as the intermediate modality alignment module, to obtain the corresponding candidate streaming encoder, candidate driving decision model and candidate modality alignment module.

[0198] The environmental perception sample set and the prompt text sample set are derived from video / image-text pairs constructed from publicly available datasets. In the second phase of the training process, the environmental perception sample set and the prompt text sample set include general knowledge data and driving-related data. In this second phase, the environmental perception sample set is input into the initial streaming encoder, and the general prompt text features corresponding to the general prompt text information in the general prompt text sample set are input into the initial driving decision model. The parameters of the initial streaming encoder, the initial driving decision model, and the intermediate modality alignment module are iteratively trained to obtain the corresponding candidate streaming encoder, candidate driving decision model, and candidate modality alignment module.

[0199] S5203. Based on the environmental perception sample set, navigation planning sample set, and specific scene prompt text sample set, the initial transformation encoder, candidate streaming encoder, candidate driving decision model, and candidate modality alignment module in the initial driving correction model are iteratively trained to obtain the corresponding target driving correction model.

[0200] The environmental perception sample set and the prompt text sample set are derived from video / image-text pairs constructed from publicly available datasets. In the third stage of the training process, the environmental perception sample set and the prompt text sample set include driving data related to downstream business operations. In this third stage, the environmental perception sample set is input into the candidate streaming encoder, the navigation planning sample set is input into the initial transformation encoder, and the specific scene prompt text features corresponding to the specific scene prompt text information in the specific scene prompt text sample set are input into the candidate driving decision model. The initial transformation encoder, as well as the candidate streaming encoder, the candidate driving decision model, and the candidate modality alignment module, are iteratively trained to obtain the corresponding target navigation encoder, target streaming encoder, target driving decision model, and target modality alignment module, constituting the corresponding target driving correction model.

[0201] During training, the driving correction model can be iteratively trained using multiple rounds of prompt text sample sets. Specifically, in the third stage of the training process, the driving correction model can output text in a fixed format, including, for example, the driving scenario, driving control correction information, and driving reference position. In this third stage, through a multi-round dialogue format using thought chains, the model can be guided to focus on more relevant elements based on the results of each round's output, thus correcting and supplementing the content of each round's dialogue output. For example, the output can be adjusted according to changes in the scenario to correct the previous round's output and further supplement information. For instance, additional questions can be asked for specific scenarios, such as whether there are stop signs at a T-junction, to supplement the driving correction information output by the driving correction model.

[0202] It should be noted that navigation planning information may not be configured during the training of the target driving correction model. Of course, the accuracy of the driving correction information output by the target driving correction model will be reduced accordingly.

[0203] The technical solution of this embodiment uses general knowledge data and intelligent driving-related data to continuously iterate and train the driving decision model in the first and second stages, and then uses driving data closely related to downstream tasks to continuously iterate and train the driving decision model for specific scenarios in the third stage, so as to correct the driving information output from end to end. This ensures that the driving decision model can provide relatively accurate driving correction information for different complex driving scenarios, thus ensuring the safety of autonomous driving.

[0204] In one embodiment, Figure 6 This is a schematic diagram illustrating the implementation of information correction according to an embodiment of the present invention. In this embodiment, the target driving correction model is a VLM model, the target streaming encoder is a Streaming Video Encoder, the target image feature extractor is a Multi-resolution ViT, the target temporal fusion module can be a Temporal Encoder, the target navigation encoder can be a ViT Encoder, the text feature extraction module is a Tokenizer and a vocabulary, the target modality alignment module is Projection, the intelligent driving prompt text library is the intelligent driving system Prompt library, and the memory storage space is a MemoryBank, to illustrate the implementation process of information correction. Figure 6 As shown, the relationships between the various modules included in the VLM model are as follows:

[0205] First, the environmental perception information corresponding to the current vehicle is input into the Streaming VideoEncoder as a video stream. The Multi-resolution ViT in the Streaming Video Encoder is used to extract image features to obtain the corresponding image features and global features. Then, the image features and global features are input into the TemporalEncoder for temporal encoding to obtain the corresponding image labeling information.

[0206] Figure 7 This is a flowchart illustrating a time-series fusion implementation according to an embodiment of the present invention. In this embodiment, the time-series fusion process is explained by taking the driving-related data of each second as the driving-related data for two corresponding frames. Figure 7As shown, driving-related data is presented over four seconds (i.e., seconds 0, 1, 2, and 3), corresponding to four frames of driving-related data. Specifically, F00 represents the environmental perception information presented as a video stream at second 0; F01 represents the navigation planning information at second 0; F10 represents the environmental perception information presented as a video stream at second 1; F11 represents the navigation planning information at second 1; F20 represents the environmental perception information presented as a video stream at second 2; F21 represents the navigation planning information at second 2; F30 represents the environmental perception information presented as a video stream at third 3; and F31 represents the navigation planning information at third 3. The ViT includes a Multi-resolution ViT and a Temporal Encoder, as well as a ViT Encoder, which respectively input the environmental perception information presented as a video stream per second into the Multi-resolution ViT and Temporal Encoder, and input the navigation planning information per second into the ViT. The Encoder obtains the corresponding image marker information and navigation marker information, namely Embed_F00, Embed_F01, Embed_F10, Embed_F11, Embed_F20, Embed_F21, Embed_F30, and Embed_F31. Among them, Embed_F00 represents the image marker information at second 0; Embed_F01 represents the navigation marker information at second 0; Embed_F10 represents the image marker information at second 1; Embed_F11 represents the navigation marker information at second 1; Embed_F20 represents the image marker information at second 2; Embed_F21 represents the navigation marker information at second 2; Embed_F30 represents the image marker information at second 3; and Embed_F31 represents the navigation marker information at second 3. Then, the image marker information and navigation marker information per second are input into the Projection and LLM models to obtain the corresponding driving correction information, which is used as the corresponding output information (i.e., Output0, Output1, Output2 and Output3).

[0207] Then, the navigation planning information corresponding to the current vehicle is input into the ViT Encoder as an image to obtain the corresponding navigation marker information.

[0208] Then, the image tag information and navigation tag information are input into Projection as driving tag information (referred to as driving token) to map the image tag information and navigation tag information to the same space as the text features, so as to obtain the corresponding multimodal features mapped to the text feature space (aligned with the text features) (referred to as mapped multimodal features).

[0209] Then, the system searches for specific scenario prompt text information that matches the current vehicle's driving scenario in the Prompt library of the intelligent driving system, inputs the specific scenario prompt text information into the Tokenizer, and consults the vocabulary to combine multiple text tag information into the corresponding prompt text feature.

[0210] Then, the environmental awareness coding information, prompt text features, and driving correction information of the current frame are extracted from the pre-built Memory Bank. The prompt text features of the current frame, as well as the prompt text features and driving correction information of the historical frames, are fused together. The environmental awareness coding information of the current frame and the environmental awareness coding information of the historical frames are also fused together.

[0211] Finally, the fused and mapped multimodal features and prompt text features are input into the LLM model to output the corresponding driving correction information.

[0212] It should be noted that the driving rule information in this invention is a portion of the driving correction information, which is similar to driving control correction information.

[0213] In one embodiment, Figure 8 This is a flowchart of another trajectory generation method provided by an embodiment of the present invention. Based on the above embodiments, this embodiment adjusts the output frequency of the target driving correction model according to the current vehicle driving scenario and / or candidate driving trajectories. For example... Figure 8 As shown, the method includes:

[0214] S81. Obtain the driving-related data corresponding to the current vehicle.

[0215] S82. Input the first driving-related data from the driving-related data into the pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle.

[0216] S83. Determine the output frequency of the target driving correction model based on the current driving scenario of the vehicle and / or the candidate driving trajectory.

[0217] In one exemplary embodiment, the output frequency of the target driving correction model is also the processing frequency of the target driving correction model. The driving scenario refers to the road type and / or lane position of the current vehicle, etc.

[0218] In an exemplary embodiment, the output frequency of the target driving correction model can be adjusted according to the current driving scenario and / or candidate driving trajectory. For example, when the driving scenario is simple (e.g., with few obstacles) or the candidate driving trajectory represents a simple driving state, the output frequency can be appropriately reduced; or, when the driving scenario is complex (e.g., with many obstacles) or the candidate driving trajectory represents a complex driving state, the output frequency can be appropriately increased (or the maximum output frequency can be used).

[0219] In one embodiment, determining the output frequency of the target driving correction model based on the current vehicle's driving scenario and / or the candidate driving trajectory may include at least one of the following:

[0220] When the driving scenario is an urban road, the output frequency is determined to be a first output frequency, and the first output frequency is less than or equal to the maximum output frequency of the target driving correction model;

[0221] When the driving scenario is a highway, the output frequency is determined to be a second output frequency, which is the difference between the maximum output frequency and the first preset frequency, and is less than the first output frequency;

[0222] When the driving scenario is the innermost lane, the output frequency is determined to be the third output frequency, which is the difference between the maximum output frequency and the second preset frequency.

[0223] When the driving scenario is the outermost lane, the output frequency is determined to be the fourth output frequency, which is greater than the third output frequency and less than or equal to the maximum output frequency.

[0224] When the driving scenario is a construction section scenario, the output frequency is determined to be the maximum output frequency;

[0225] When the candidate driving trajectory is a straight trajectory, the output frequency is determined to be the fifth output frequency, and the fifth output frequency is the difference between the maximum output frequency and the third preset frequency;

[0226] When the candidate driving trajectory is a non-linear trajectory, the output frequency is determined to be the sixth output frequency, which is greater than the fifth output frequency and less than or equal to the maximum output frequency.

[0227] In an exemplary embodiment, when the driving scenario is an urban road, there are many pedestrians, vehicles or other obstacles in the urban road, and sudden situations (such as the sudden appearance of pedestrians) are likely to occur. In this case, the target driving correction model can be run at a higher output frequency. The output frequency of the target driving correction model can be determined as the first output frequency. The first output frequency can be the maximum output frequency of the target driving correction model, or it can be slightly lower than the maximum output frequency, so as to obtain driving correction information at a higher frequency, ensure the effect of intelligent driving, and improve safety.

[0228] In an exemplary embodiment, when the driving scenario is a highway, the output frequency of the target driving correction model can be appropriately reduced. The output frequency of the target driving correction model is determined as a second output frequency, which is the difference between the maximum output frequency and a first preset frequency. This reduces the computational load and reserves some computing power for other modules (such as the safety perception module, automatic emergency braking module, etc.), ensuring the operation of other modules. The first preset frequency can be the optimal preset frequency obtained by pre-testing in a highway scenario, representing a reduction from the maximum output frequency.

[0229] In one exemplary embodiment, when the driving scenario is the innermost lane, there are fewer unexpected situations such as pedestrians. Therefore, the output frequency of the target driving correction model can be appropriately reduced. The output frequency of the target driving correction model is determined as a third output frequency, which is the difference between the maximum output frequency and a second preset frequency, thereby reducing the computational load. The second preset frequency can be the optimal preset frequency obtained by pre-testing in the innermost lane scenario, representing a reduction from the maximum output frequency. It should be noted that when the driving scenario is the innermost lane, different second preset frequencies can also be set based on the road type. For example, the second preset frequency for the innermost lane on urban roads can be lower than the second preset frequency for the innermost lane on highways.

[0230] In an exemplary embodiment, when the driving scenario is the outermost lane, there are more unexpected situations such as pedestrians. The output frequency of the target driving correction model can be appropriately increased. The output frequency of the target driving correction model is determined as the fourth output frequency, which is greater than the third output frequency. The fourth output frequency can be the maximum output frequency of the target driving correction model or slightly less than the maximum output frequency, so as to obtain driving correction information at a higher frequency, ensure the effect of intelligent driving, and improve safety.

[0231] In an exemplary embodiment, when the driving scenario is a construction section, the road conditions are relatively complex. The target driving correction model can be run at the maximum output frequency to obtain driving correction information at the maximum output frequency, thereby ensuring the effect of intelligent driving and improving safety.

[0232] In an exemplary embodiment, when all determined candidate driving trajectories are straight lines, it indicates that the current driving state is relatively simple. The output frequency of the target driving correction model can be appropriately reduced, and the output frequency of the target driving correction model is determined as the fifth output frequency. The fifth output frequency is the difference between the maximum output frequency and the third preset frequency, thereby reducing the computational load. The third preset frequency can be the optimal preset frequency obtained by testing beforehand when the candidate driving trajectories are straight lines, representing a reduction from the maximum output frequency. It should be noted that when the candidate driving trajectories are straight lines, a second preset frequency can also be set in conjunction with the driving scenario.

[0233] In an exemplary embodiment, when the determined candidate driving trajectory is a non-linear trajectory, it indicates that the current driving state is relatively complex. The output frequency of the target driving correction model can be appropriately increased. The output frequency of the target driving correction model is determined to be the sixth output frequency, which is greater than the fifth output frequency. The sixth output frequency can be the maximum output frequency of the target driving correction model or slightly less than the maximum output frequency, so as to obtain driving correction information at a higher frequency, ensure the effect of intelligent driving, and improve safety.

[0234] S84. Based on the output frequency, input the second driving-related data from the driving-related data into the pre-created target driving correction model to obtain the corresponding driving correction information.

[0235] The target driving correction model is run according to the determined output frequency. Specifically, the process of obtaining driving correction information through the target driving correction model can be referred to the above embodiment, and will not be repeated here.

[0236] S85. Correct the candidate driving trajectory based on the driving correction information to obtain the corresponding target driving trajectory.

[0237] The technical solution of this embodiment determines the output frequency of the target driving correction model based on the current driving scenario and / or candidate driving trajectory, and runs the target driving correction model according to the output frequency to obtain driving correction information. This realizes the adjustment of the computational load consumed by the target driving correction model, which can reduce the load while ensuring driving safety, so that the system can better achieve a balance between response frequency and function (intelligent driving) and meet the normal operation of the vehicle system.

[0238] In one embodiment, Figure 9 This is a schematic diagram of a trajectory generation device provided in an embodiment of the present invention. In this embodiment, the trajectory generation device is integrated into the vehicle's internal components. Figure 9As shown, the device includes: an acquisition module 810, a first generation module 820, a second generation module 830, and a third generation module 840.

[0239] Among them, the acquisition module 810 is used to acquire driving-related data corresponding to the current vehicle;

[0240] The first generation module 820 is used to input the first driving-related data from the driving-related data into the pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle.

[0241] The second generation module 830 is used to input the second driving-related data from the driving-related data into the pre-created target driving correction model to obtain the corresponding driving correction information.

[0242] The third generation module 840 is used to correct the candidate driving trajectory based on the driving correction information to obtain the corresponding target driving trajectory.

[0243] In one embodiment, the second driving-related data includes: sensor information and navigation planning information. The sensor information includes: environmental perception information and state information. The target trajectory generation model includes: a target backbone network, a target encoder, a target decoder, and a target memory module. The target memory module is used to store BEV features in the time and spatial dimensions.

[0244] The first generation module 820 is specifically used for:

[0245] Environmental perception information is input into the target backbone network to obtain target fusion features, and the target fusion features are projected onto the BEV space.

[0246] Based on the target fusion features projected into the BEV space and the BEV features output by the target memory module, the target BEV features are determined, and the BEV features stored in the target memory module are updated based on the target BEV features.

[0247] The state information and navigation planning information are input into the target encoder to obtain the target encoding features;

[0248] Input the target encoded features and target BEV features into the target decoder to obtain the candidate driving trajectory corresponding to the current vehicle.

[0249] In one embodiment, the training process of the target trajectory generation model includes:

[0250] Acquire the perception sample set, the control sample set, and the initial trajectory generation model;

[0251] Based on the perceptual sample set, the parameters of the initial trajectory generation model are trained iteratively, and the trained initial trajectory generation model is determined as the first model.

[0252] Based on the regulatory sample set, the parameters of the first model are trained iteratively, and the trained first model is determined as the second model;

[0253] Based on the perception sample set and the control sample set, the parameters of the second model are trained iteratively, and the trained second model is determined as the target trajectory generation model.

[0254] In one embodiment, the perception sample set includes: multiple driving-related samples and obstacle tags and road structure tags carried by each driving-related sample;

[0255] Based on the perceptual sample set, the parameters of the initial trajectory generation model are trained iteratively. The trained initial trajectory generation model is then determined as the first model, specifically used for:

[0256] Input driving-related samples from the perception sample set into the initial trajectory generation model to obtain the first predicted obstacle information and the first predicted road structure.

[0257] Based on the differences between the first predicted obstacle information and the obstacle label, and the differences between the first predicted road structure and the road structure label, the parameters of the initial trajectory generation model are trained, and the trained initial trajectory generation model is determined as the first model.

[0258] In one embodiment, the control sample set includes: multiple driving-related samples and the trajectory of each driving-related sample at the next moment;

[0259] Based on the regulatory sample set, the parameters of the first model are trained iteratively, and the trained first model is determined as the second model, specifically used for:

[0260] Input the driving-related samples from the control sample set into the first model to obtain the first predicted trajectory;

[0261] Based on the difference between the first predicted trajectory and the trajectory at the next moment, the parameters of the first model are trained, and the trained first model is determined as the second model.

[0262] In one embodiment, based on the perception sample set and the regulation sample set, the parameters of the second model are iteratively trained, and the trained second model is determined as the target trajectory generation model, specifically used for:

[0263] A fusion sample set is generated based on the perception sample set and the planning and control sample set. The fusion sample set includes: multiple driving-related samples and the obstacle label, road structure label, and trajectory of each driving-related sample at the next moment.

[0264] Input the driving-related samples from the fusion sample set into the second model to obtain the second predicted obstacle, the second predicted road structure, and the second predicted trajectory.

[0265] Based on the differences between the second predicted obstacle and the obstacle label, the differences between the second predicted road structure and the road structure label, and the differences between the second predicted trajectory and the trajectory at the next moment, the parameters of the second model are trained, and the trained second model is determined as the target trajectory generation model.

[0266] In one embodiment, driving-related samples include: sensor samples and navigation planning samples; sensor data samples include: environmental perception samples and state samples; environmental perception samples include: frame samples and point cloud samples.

[0267] In one embodiment, the initial trajectory generation model includes: an initial backbone network, an initial encoder, an initial decoder, and a target memory module;

[0268] The driving-related samples from the perception sample set are input into the initial trajectory generation model to obtain the first predicted obstacle information and the first predicted road structure information, which are specifically used for:

[0269] The initial decoder is initialized based on a preset instance;

[0270] Environmental perception samples are input into the initial backbone network to obtain the first fusion feature, and the first fusion feature is projected onto the BEV space.

[0271] Based on the first fusion feature projected onto the BEV space and the BEV feature output by the target memory module, the first BEV feature is determined, and the BEV feature stored in the target memory module is updated based on the first BEV feature.

[0272] Input the state samples and navigation planning samples into the initial encoder to obtain the first encoded features;

[0273] The first encoded feature and the first BEV feature are input into the initialized initial decoder to obtain the first predicted obstacle information and the first predicted road structure.

[0274] In one embodiment, the first driving-related data from the driving-related data is input into a pre-created target trajectory generation model to obtain the candidate driving trajectory corresponding to the current vehicle, specifically for:

[0275] By inputting driving-related data into the target trajectory generation model, candidate driving trajectories, obstacle information, and road structure corresponding to the current vehicle are obtained.

[0276] In one embodiment, the obstacle information includes: first type obstacle information and second type obstacle information.

[0277] In one embodiment, the target driving correction model includes: a target streaming encoder, a target navigation encoder, a target modal alignment module, and a target driving decision model; driving-related data includes: environmental perception information, navigation planning information, and driving prompt information; the second generation module 830 includes:

[0278] The first generation unit is used to input environmental perception information into the target streaming encoder to obtain corresponding image labeling information;

[0279] The second generation unit is used to input navigation planning information into the target navigation encoder to obtain the corresponding navigation mark information;

[0280] The third generation unit is used to input image labeling information and navigation labeling information as driving labeling information into the target modality alignment module to obtain the mapped multimodal features;

[0281] The fourth generation unit is used to input the prompt text features corresponding to the driving prompt information and the mapped multimodal features into the target driving decision model to obtain the corresponding driving correction information.

[0282] In one embodiment, the target streaming encoder includes: a target image feature extractor and a target temporal fusion module; the first generation unit includes:

[0283] The first generation subunit is used to input environmental perception information into the target image feature extractor to obtain the corresponding environmental perception coding information;

[0284] The second generation subunit is used to input environmental perception coding information into the target temporal fusion module to obtain the corresponding image labeling information.

[0285] In one embodiment, the driving correction information includes at least one of the following: driving reference position, driving scenario, and driving control correction information.

[0286] In one embodiment, the second generation module 830 further includes:

[0287] The conversion unit is used to input the driving prompt information into a pre-created prompt text feature extraction module to obtain the corresponding prompt text features before inputting the prompt text features corresponding to the driving prompt information and the mapped multimodal features into the target driving decision model to obtain the corresponding driving correction information.

[0288] In one embodiment, the fourth generation unit includes:

[0289] The storage subunit is used to store the environmental perception coding information, prompt text features, and driving correction information of historical frames into the memory storage space;

[0290] The third generation subunit is used to extract the environmental perception coding information, prompt text features and driving correction information of the current frame from the memory storage space, and to fuse the prompt text features of the current frame with the prompt text features and driving correction information of the historical frames, as well as to fuse the environmental perception coding information of the current frame with the environmental perception coding information of the historical frames, and to input the fused information into the target driving decision model to obtain the corresponding driving correction information.

[0291] In one embodiment, the training process of the target driving correction model includes:

[0292] Obtain a driving-related sample set;

[0293] The initial driving correction model is iteratively trained based on a driving-related sample set to obtain the corresponding target driving correction model.

[0294] In one embodiment, the initial driving correction model includes: an initial streaming encoder, an initial transform encoder, an initial modal alignment module, and an initial driving decision model; the driving-related sample set includes: an environmental perception sample set, a navigation planning sample set, and a driving prompt sample set; the driving prompt sample set includes: a prompt text sample set and a scenario-specific prompt text sample set;

[0295] The initial driving correction model is iteratively trained based on a driving-related sample set to obtain the corresponding target driving correction model, which is specifically used for:

[0296] The initial modality alignment module in the initial driving correction model is iteratively trained based on the environmental perception sample set and the prompt text sample set to obtain the corresponding intermediate modality alignment module;

[0297] Based on the environmental perception sample set and the prompt text sample set, the initial streaming encoder and initial driving decision model in the initial driving correction model, as well as the intermediate modality alignment module, are iteratively trained to obtain the corresponding candidate streaming encoder, candidate driving decision model and candidate modality alignment module.

[0298] Based on the environmental perception sample set, navigation planning sample set, and specific scenario prompt text sample set, the initial transformation encoder in the initial driving correction model, as well as the candidate streaming encoder, candidate driving decision model, and candidate modality alignment module, are iteratively trained to obtain the corresponding target driving correction model.

[0299] In one embodiment, the device further includes:

[0300] The output frequency determination module is used to determine the output frequency of the target driving correction model based on the current driving scenario of the vehicle and / or the candidate driving trajectory.

[0301] The second generation module is specifically used for:

[0302] Based on the output frequency, the second driving-related data in the driving-related data is input into the pre-created target driving correction model to obtain the corresponding driving correction information.

[0303] In one embodiment, the output frequency determination module is configured to perform at least one of the following:

[0304] When the driving scenario is an urban road, the output frequency is determined to be a first output frequency, and the first output frequency is less than or equal to the maximum output frequency of the target driving correction model;

[0305] When the driving scenario is a highway, the output frequency is determined to be a second output frequency, which is the difference between the maximum output frequency and the first preset frequency, and is less than the first output frequency;

[0306] When the driving scenario is the innermost lane, the output frequency is determined to be the third output frequency, which is the difference between the maximum output frequency and the second preset frequency.

[0307] When the driving scenario is the outermost lane, the output frequency is determined to be the fourth output frequency, which is greater than the third output frequency and less than or equal to the maximum output frequency.

[0308] When the driving scenario is a construction section scenario, the output frequency is determined to be the maximum output frequency;

[0309] When the candidate driving trajectory is a straight trajectory, the output frequency is determined to be the fifth output frequency, and the fifth output frequency is the difference between the maximum output frequency and the third preset frequency;

[0310] When the candidate driving trajectory is a non-linear trajectory, the output frequency is determined to be the sixth output frequency, which is greater than the fifth output frequency and less than or equal to the maximum output frequency.

[0311] The trajectory generation device provided in the embodiments of the present invention can execute the trajectory generation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0312] In one embodiment, Figure 10 This is a structural block diagram of an electronic device provided in an embodiment of the present invention, such as... Figure 10The diagram illustrates a schematic representation of an electronic device 10 that can be used to implement embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0313] like Figure 10 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0314] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0315] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as trajectory generation methods.

[0316] In some embodiments, the trajectory generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the trajectory generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the trajectory generation method by any other suitable means (e.g., by means of firmware).

[0317] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0318] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0319] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0320] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0321] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0322] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0323] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the trajectory generation method provided in any embodiment of this application.

[0324] In the implementation of the computer program product, computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0325] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0326] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A trajectory generation method characterized by, The method comprises: obtaining driving-related data corresponding to a current vehicle; inputting first driving-related data in the driving-related data into a pre-created target trajectory generation model to obtain a candidate driving trajectory corresponding to the current vehicle; inputting second driving-related data in the driving-related data into a pre-created target driving correction model to obtain corresponding driving correction information; correcting the candidate driving trajectory based on the driving correction information to obtain a corresponding target driving trajectory.

2. The method of claim 1, wherein, The first driving-related data comprises sensor information and navigation planning information, the sensor information comprises environmental perception information and state information, the target trajectory generation model comprises a target backbone network, a target encoder, a target decoder, and a target memory module, and the target memory module is used to store BEV features in time and space dimensions; The method comprises: inputting the environmental perception information into the target backbone network to obtain target fusion features, and projecting the target fusion features to a BEV space; determining target BEV features according to the target fusion features projected to the BEV space and the BEV features output by the target memory module, and updating the BEV features stored in the target memory module according to the target BEV features; inputting the state information and the navigation planning information into the target encoder to obtain target encoding features; inputting the target encoding features and the target BEV features into the target decoder to obtain the candidate driving trajectory corresponding to the current vehicle.

3. The method of claim 1, wherein, The training process of the target trajectory generation model comprises: obtaining a perception sample set, a regulation and control sample set, and an initial trajectory generation model; iteratively training parameters of the initial trajectory generation model based on the perception sample set, determining the initial trajectory generation model after training as a first model; iteratively training parameters of the first model based on the regulation and control sample set, determining the first model after training as a second model; iteratively training parameters of the second model based on the perception sample set and the regulation and control sample set, and determining the second model after training as a target trajectory generation model.

4. The method of claim 3, wherein, The perception sample set comprises a plurality of driving-related samples, obstacle labels carried by each driving-related sample, and road structure labels; The method comprises: inputting the driving-related samples in the perception sample set into the initial trajectory generation model to obtain first predicted obstacle information and first predicted road structures; training parameters of the initial trajectory generation model according to differences between the first predicted obstacle information and the obstacle labels and differences between the first predicted road structures and the road structure labels, and determining the initial trajectory generation model after training as a first model.

5. The method of claim 4, wherein, The regulation and control sample set comprises a plurality of driving-related samples and trajectories corresponding to each driving-related sample at a next moment. Based on the regulation sample set, the parameters of the first model are iteratively trained, and the trained first model is determined as a second model, including: Inputting the driving-related samples in the regulation sample set into the first model to obtain a first predicted trajectory; According to the difference between the first predicted trajectory and the trajectory at the next time, the parameters of the first model are trained, and the trained first model is determined as a second model.

6. The method of claim 5, wherein, Based on the perception sample set and the regulation sample set, the parameters of the second model are iteratively trained, and the trained second model is determined as a target trajectory generation model, including: According to the perception sample set and the regulation sample set, a fusion sample set is generated, wherein the fusion sample set includes a plurality of driving-related samples and obstacle labels, road structure labels, and trajectories at the next time corresponding to each driving-related sample; Inputting the driving-related samples in the fusion sample set into the second model to obtain second predicted obstacles, second predicted road structures, and a second predicted trajectory; According to the difference between the second predicted obstacles and the obstacle labels, the difference between the second predicted road structures and the road structure labels, and the difference between the second predicted trajectory and the trajectory at the next time, the parameters of the second model are trained, and the trained second model is determined as a target trajectory generation model.

7. The method according to any one of claims 4-6, characterized in that, The driving-related sample includes a sensor data sample and a navigation planning sample, and the sensor data sample includes an environment perception sample and a state sample.

8. The method of claim 7, wherein, The initial trajectory generation model includes an initial backbone network, an initial encoder, an initial decoder, and a target memory module. Inputting the driving-related samples in the perception sample set into the initial trajectory generation model to obtain first predicted obstacle information and first predicted road structures, including: Initializing the initial decoder based on a preset instance; Inputting the environment perception sample into the initial backbone network to obtain first fusion features, and projecting the first fusion features to a BEV space; According to the first fusion features projected to the BEV space and the BEV features output by the target memory module, determining first BEV features, and updating the BEV features stored in the target memory module according to the first BEV features; Inputting the state sample and the navigation planning sample into the initial encoder to obtain first encoding features; Inputting the first encoding features and the first BEV features into the initialized initial decoder to obtain first predicted obstacle information and first predicted road structures.

9. The method of claim 1, wherein, Inputting the first driving-related data in the driving-related data into the pre-created target trajectory generation model to obtain a candidate driving trajectory corresponding to the current vehicle, including: Inputting the first driving-related data into the pre-created target trajectory generation model to obtain a candidate driving trajectory corresponding to the current vehicle, obstacle information, and road structure.

10. The method of claim 9, wherein, The obstacle information includes first-type obstacle information and second-type obstacle information.

11. The method of claim 1, wherein, The target driving correction model comprises a target stream encoder, a target navigation encoder, a target modality alignment module, and a target driving decision model; and the second driving-related data comprises environmental perception information, navigation planning information, and driving prompt information. The inputting of the second driving-related data in the driving-related data into the pre-created target driving correction model to obtain corresponding driving correction information comprises: inputting the environmental perception information into the target stream encoder to obtain corresponding image label information; inputting the navigation planning information into the target navigation encoder to obtain corresponding navigation label information; inputting the image label information and the navigation label information as driving label information into the target modality alignment module to obtain mapped multi-modal features; inputting corresponding prompt text features of the driving prompt information and the mapped multi-modal features into the target driving decision model to obtain corresponding driving correction information.

12. The method of claim 11, wherein, The target stream encoder comprises a target image feature extractor and a target time sequence fusion module; and the inputting of the environmental perception information into the target stream encoder to obtain corresponding image label information comprises: inputting the environmental perception information into the target image feature extractor to obtain corresponding environmental perception encoding information; inputting the environmental perception encoding information into the target time sequence fusion module to obtain corresponding image label information.

13. The method of claim 1, wherein, The driving correction information at least comprises one of the following: driving reference position, driving scene, and driving control correction information.

14. The method of claim 11, wherein, Before the inputting of the corresponding prompt text features of the driving prompt information and the mapped multi-modal features into the target driving decision model to obtain corresponding driving correction information, the method further comprises: inputting the driving prompt information into a pre-created text feature extraction module to obtain corresponding prompt text features.

15. The method of claim 12, wherein, The inputting of the corresponding prompt text features of the driving prompt information and the mapped multi-modal features into the target driving decision model to obtain corresponding driving correction information comprises: storing environmental perception encoding information, prompt text features, and driving correction information of a historical frame into a memory storage space; extracting environmental perception encoding information, prompt text features, and driving correction information of a current frame from the memory storage space, fusing the prompt text features of the current frame, the prompt text features, and the driving correction information of the historical frame, and fusing the environmental perception encoding information of the current frame and the environmental perception encoding information of the historical frame, and inputting the fused information into a target driving decision model to obtain corresponding driving correction information.

16. The method of any one of claims 1 or 11-15, wherein, The training process of the target driving correction model comprises: obtaining a driving-related sample set; iteratively training a pre-created initial driving correction model based on the driving-related sample set to obtain a corresponding target driving correction model.

17. The method of claim 16, wherein, The initial driving correction model comprises: an initial streaming encoder, an initial transformation encoder, an initial modal alignment module, and an initial driving decision model; the driving-related sample set comprises: an environment perception sample set, a navigation planning sample set, and a driving prompt sample set; the driving prompt sample set comprises: a prompt text sample set and a specific scene prompt text sample set; The initial driving correction model is iteratively trained based on the driving-related sample set to obtain a corresponding target driving correction model, comprising: The initial modal alignment module in the initial driving correction model is iteratively trained based on the environment perception sample set and the prompt text sample set to obtain a corresponding intermediate modal alignment module; The initial streaming encoder and the initial driving decision model in the initial driving correction model, and the intermediate modal alignment module are iteratively trained based on the environment perception sample set and the prompt text sample set to obtain a corresponding candidate streaming encoder, a candidate driving decision model, and a candidate modal alignment module; The initial transformation encoder in the initial driving correction model, and the candidate streaming encoder, the candidate driving decision model, and the candidate modal alignment module are iteratively trained based on the environment perception sample set, the navigation planning sample set, and the specific scene prompt text sample set to obtain a corresponding target driving correction model.

18. The method according to any one of claims 1 to 17, characterized in that, Before the second driving-related data in the driving-related data is input into the pre-created target driving correction model, further comprising: determining an output frequency of the target driving correction model according to the driving scene of the current vehicle and / or the candidate driving trajectory; inputting the second driving-related data in the driving-related data into the pre-created target driving correction model to obtain corresponding driving correction information, comprising: inputting the second driving-related data in the driving-related data into the pre-created target driving correction model according to the output frequency to obtain corresponding driving correction information.

19. The method of claim 18, wherein, The determination of the output frequency of the target driving correction model according to the driving scene of the current vehicle and / or the candidate driving trajectory comprises at least one of the following: when the driving scene is an urban road, determining that the output frequency is a first output frequency, the first output frequency being less than or equal to a maximum output frequency of the target driving correction model; when the driving scene is a highway, determining that the output frequency is a second output frequency, the second output frequency being a difference between the maximum output frequency and a first preset frequency, and being less than the first output frequency; when the driving scene is an innermost lane, determining that the output frequency is a third output frequency, the third output frequency being a difference between the maximum output frequency and a second preset frequency; when the driving scene is an outermost lane, determining that the output frequency is a fourth output frequency, the fourth output frequency being greater than the third output frequency and less than or equal to the maximum output frequency; when the driving scene is a construction road section scene, determining that the output frequency is the maximum output frequency; When the candidate driving trajectory is a straight line trajectory, the output frequency is determined as a fifth output frequency, the fifth output frequency being a difference between the maximum output frequency and a third preset frequency; When the candidate driving trajectory is a non-straight line trajectory, the output frequency is determined as a sixth output frequency, the sixth output frequency being greater than the fifth output frequency and less than or equal to the maximum output frequency.

20. A trajectory generation device characterized by comprising: Comprise: An acquisition module, configured to acquire driving related data corresponding to a current vehicle; A first generation module, configured to input first driving related data in the driving related data into a pre-created target trajectory generation model to obtain a candidate driving trajectory corresponding to the current vehicle; A second generation module, configured to input second driving related data in the driving related data into a pre-created target driving correction model to obtain corresponding driving correction information; A third generation module, configured to correct the candidate driving trajectory based on the driving correction information to obtain a corresponding target driving trajectory.

21. An electronic device, comprising: The electronic device comprises: At least one processor; and A memory connected in communication with the at least one processor; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the trajectory generation method in any one of claims 1-19.

22. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the processor to execute the trajectory generation method in any one of claims 1-19 when executed.

23. A computer program product, characterised in that, The computer program product comprises a computer program, and the computer program implements the trajectory generation method according to any one of claims 1-19 when executed by the processor.