Method and system for generating driving track of autonomous vehicle
By acquiring image and point cloud data and combining them with multimodal BEV features to generate vehicle driving trajectories, the problems of large errors and low efficiency in traditional autonomous driving are solved, and more stable and accurate trajectory generation is achieved.
Patent Information
- Application Number
- CN202511065630.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-09-16
AI Technical Summary
In traditional end-to-end autonomous driving technology, vehicle trajectory generation has large errors and low efficiency, and the supervision cost is high.
Image data and point cloud data are acquired through vehicle sensors, combined with multimodal BEV features to determine environmental information, including the vehicle's risk perception status, sparse environmental scenes, and the fused interaction information of the intelligent agent, and the vehicle's driving trajectory is generated using the cross-attention mechanism.
The error of vehicle trajectory generation is reduced, the stability and accuracy of trajectory generation are improved, the supervision cost is reduced, and the generation speed is increased.
Smart Images

Figure CN120646020A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method and system for generating a driving trajectory of an autonomous driving vehicle. Background Art
[0002] Traditional end-to-end autonomous driving technology typically relies on single-modal image data to predict vehicle trajectories, resulting in large errors in the predicted vehicle trajectories. Furthermore, during the trajectory prediction process, traditional autonomous driving technology requires manual annotation for model supervision of intermediate subtasks (such as detection, segmentation, and mapping), resulting in high supervision costs and reduced trajectory generation efficiency. Summary of the Invention
[0003] In view of this, the present invention proposes a method and system for generating a driving trajectory of an autonomous vehicle, which solves the problems of large errors and low efficiency in vehicle driving trajectory generation in traditional end-to-end autonomous driving technology.
[0004] In one aspect, an embodiment of the present invention provides a method for generating a driving trajectory of an autonomous vehicle, the method comprising: Acquire image data and point cloud data based on vehicle sensors; Determining environmental information based on the image data and the point cloud data, the environmental information including risk perception status information of the vehicle; Based on environmental information, determine the joint representation information of scene and behavior; Generate the vehicle's driving trajectory based on the joint representation information of scene and behavior.
[0005] In some embodiments, determining environmental information based on image data and point cloud data includes: Determine multimodal BEV features based on image data and point cloud data; Based on multimodal BEV features, environmental information is determined.
[0006] In some embodiments, determining risk perception status information of the vehicle based on the image data and the point cloud data includes: Based on multimodal BEV features, determine the local BEV feature information centered on the vehicle and the state information of the intelligent agents around the vehicle; The risk perception status information of the vehicle is determined based on the local BEV feature information and the status information of the intelligent entities around the vehicle.
[0007] In some embodiments, determining the risk perception state information of the vehicle based on the local BEV characteristic information and the state information of the intelligent agents surrounding the vehicle includes: Determining interaction risk information between the vehicle and the surrounding intelligent agents based on the vehicle's position information in the multimodal BEV signature and the position information of the vehicle's surrounding intelligent agents in the multimodal BEV signature; Determine risk-enhanced vehicle feature information based on local BEV feature information and interactive risk information; The risk-enhanced vehicle feature information, vehicle location information, and vehicle navigation instructions are integrated to obtain the vehicle's risk perception status information.
[0008] In some embodiments, the environmental information further includes sparse environmental scene information; and determining the sparse environmental scene information based on the image data and the point cloud data further includes: determining an enhanced multimodal BEV feature based on regions of the multimodal BEV feature that are associated with the navigation instruction; Through the spatial attention mechanism, target elements are extracted from the enhanced multimodal BEV features to obtain sparse environmental scene information.
[0009] In some embodiments, the environmental information further includes fused interaction information of intelligent bodies around the vehicle; and determining the fused interaction information of intelligent bodies around the vehicle based on the image data and the point cloud data includes: Determine the state information of the intelligent bodies around the vehicle based on image data and point cloud data; Based on the state information of the intelligent agents around the vehicle and the pre-configured behavior pattern embedding information, the fusion interaction information of the intelligent agents is generated.
[0010] In some embodiments, generating fused interaction information of the intelligent agents based on the state information of the intelligent agents around the vehicle and the pre-configured behavior pattern embedding information includes: Determine the multimodal trajectory characteristics of the intelligent agents surrounding the vehicle based on the state information of the intelligent agents surrounding the vehicle and the pre-configured behavior pattern embedding information; Based on the multimodal trajectory characteristics of the intelligent agents around the vehicle, the fused interaction information of the intelligent agents around the vehicle is generated.
[0011] In some embodiments, the environmental information further includes sparse environmental scene information and fused interaction information of intelligent agents surrounding the vehicle; and determining the scene and behavior joint representation information based on the environmental information includes: Determine the interaction between the vehicle and the scene based on the sparse environmental scene information and the risk perception status information of the vehicle; Based on the interaction information between the vehicle and the scene and the fused interaction information of the intelligent agents surrounding the vehicle, the joint representation information of the scene and behavior is determined.
[0012] In some embodiments, generating the vehicle's driving trajectory based on the scene and behavior joint representation information includes: Configuring an initial trajectory point sequence of the vehicle in the future; Based on the cross-attention mechanism, the vehicle's driving trajectory in the future is generated according to the joint representation information of the scene and behavior and the initial trajectory point sequence.
[0013] On the other hand, an embodiment of the present invention further provides a driving trajectory generation system for an autonomous driving vehicle, comprising: at least one processor; and A memory storing a computer program that can be run on the processor, wherein the processor executes the steps of the method for generating a driving trajectory of an autonomous driving vehicle as described above when executing the program.
[0014] The present invention has at least the following beneficial effects: The present invention provides a method for generating a driving trajectory of an autonomous vehicle. The method for generating a driving trajectory of an autonomous vehicle provided by the present invention determines environmental information including risk perception status information of the vehicle through image data and point cloud data acquired by vehicle sensors, and determines scene and behavior joint representation information through the environmental information, and generates the vehicle's driving trajectory through the scene and behavior joint representation information, thereby reducing the error of the generated vehicle driving trajectory, improving the stability and accuracy of the generated vehicle driving trajectory, and improving the generation speed of the vehicle driving trajectory. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 A flowchart of a method for generating a driving trajectory of an autonomous vehicle provided by an embodiment of the present invention; Figure 2 A flowchart of a method for generating a driving trajectory of an autonomous vehicle provided by an embodiment of the present invention; Figure 3 A flowchart of a method for determining environmental information in a method for generating a driving trajectory of an autonomous vehicle provided in an embodiment of the present invention; Figure 4 A flowchart of a method for generating a driving trajectory of an autonomous vehicle provided by an embodiment of the present invention; Figure 5 A flowchart of a method for determining vehicle risk perception status information in a method for generating a driving trajectory of an autonomous vehicle provided by an embodiment of the present invention; Figure 6 A flowchart of a method for generating fused interaction information of intelligent agents in a method for generating a driving trajectory of an autonomous vehicle provided by an embodiment of the present invention; Figure 7 A flowchart of a method for generating fused interaction information of intelligent agents in a method for generating a driving trajectory of an autonomous vehicle provided by an embodiment of the present invention; Figure 8 A flowchart of a method for determining scene and behavior joint representation information in a method for generating a driving trajectory of an autonomous vehicle provided in an embodiment of the present invention; Figure 9 A flowchart of a method for determining scene and behavior joint representation information in a method for generating a driving trajectory of an autonomous vehicle provided by an embodiment of the present invention; Figure 10 A flowchart of a method for generating a vehicle's driving trajectory based on scene and behavior joint representation information in a method for generating a driving trajectory of an autonomous vehicle provided by an embodiment of the present invention; Figure 11 A schematic structural diagram of a driving trajectory generation system for an autonomous driving vehicle provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the embodiments of the present invention are further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0018] It should be noted that all expressions using "first" and "second" in the embodiments of the present invention are for distinguishing two non-identical entities with the same name or non-identical parameters. It can be seen that "first" and "second" are only for the convenience of expression and should not be understood as limitations on the embodiments of the present invention. Subsequent embodiments will not explain this one by one.
[0019] The present invention is described in detail below with reference to the embodiments and accompanying drawings.
[0020] A first aspect of an embodiment of the present invention provides a method for generating a driving trajectory of an autonomous vehicle, such as Figure 1 As shown, the method specifically includes steps S10 to S40.
[0021] S10. Acquire image data and point cloud data based on vehicle sensors.
[0022] Sensors may include visual sensors (e.g., cameras) and radar sensors (e.g., lidar, millimeter-wave radar, etc.). Image data may include images of the vehicle's surroundings captured by the vehicle's visual sensors. These images may include objects such as road markings, pedestrians, and surrounding vehicles. Point cloud data may include point cloud data of the vehicle's surroundings captured by the vehicle's lidar. These point cloud data may include objects such as road markings, pedestrians, and surrounding vehicles.
[0023] In an embodiment of the present invention, an image of the vehicle's surrounding environment may be acquired through a vehicle visual sensor, and point cloud data of the vehicle's surrounding environment may be acquired through a vehicle radar sensor.
[0024] S20: Determine environmental information based on the image data and the point cloud data.
[0025] In an embodiment of the present invention, multimodal BEV features can be determined based on image data and point cloud data, and environmental information can be determined based on multimodal BEV (Bird's Eye View) features. The environmental information may include risk perception status information of the vehicle.
[0026] More specifically, the image data and point cloud data can be converted into BEV spatial features respectively, and then the two can be fused to obtain multimodal BEV features. The multimodal BEV features simultaneously contain image information and point cloud information of the vehicle's surrounding environment, laying the foundation for improving the accuracy of vehicle trajectory generation.
[0027] Through this multimodal BEV feature, the local BEV feature information centered on the vehicle and the status information of the intelligent bodies around the vehicle can be determined. Through the local BEV feature information and the status information of the intelligent bodies around the vehicle, the risk perception status information of the vehicle can be determined, thereby reducing the error of the subsequently generated vehicle driving trajectory and improving the stability and accuracy of the generated vehicle driving trajectory.
[0028] S30: Determine scene and behavior joint representation information based on environmental information.
[0029] In an embodiment of the present invention, the scene and behavior joint representation information can be determined based on one or more of the risk perception status information of the vehicle in the environmental information, the sparse environmental scene information and the fused interaction information of the intelligent bodies around the vehicle, thereby providing an environmental basis for the generation of subsequent vehicle driving trajectories.
[0030] S40: Generate a vehicle's driving trajectory based on the scene and behavior joint representation information.
[0031] In an embodiment of the present invention, an initial trajectory point sequence of a vehicle in the future time can be pre-configured. Through a cross-attention mechanism, the vehicle's driving trajectory in the future time can be generated based on the joint representation information of the scene and behavior and the initial trajectory point sequence.
[0032] In an embodiment of the present invention, image data and point cloud data acquired by vehicle sensors are used to determine environmental information including risk perception status information of the vehicle, and scene and behavior joint representation information is determined through the environmental information, and the vehicle's driving trajectory is generated through the scene and behavior joint representation information. As a result, the error of the generated vehicle driving trajectory is reduced, the stability and accuracy of the generated vehicle driving trajectory are improved, and the speed of generating the vehicle driving trajectory is increased.
[0033] In some embodiments of the present invention, the environmental information may include sparse environmental scene information in addition to the risk perception status information of the vehicle.
[0034] In some embodiments of the present invention, the environmental information may include not only the risk perception status information of the vehicle but also the fusion interaction information of the intelligent agents surrounding the vehicle.
[0035] In some embodiments of the present invention, in addition to the risk perception status information of the vehicle, the environmental information may also include sparse environmental scene information and fused interaction information of intelligent agents surrounding the vehicle.
[0036] In some embodiments of the present invention, Figure 2 As shown, step S20 (determining environmental information based on image data and point cloud data) may include steps S21 and S22. That is, the environmental information in the method for generating a driving trajectory of an autonomous vehicle provided in an embodiment of the present invention may be determined through steps S21 and S22.
[0037] S21. Determine multimodal BEV features based on the image data and the point cloud data.
[0038] S22. Determine environmental information based on the multimodal BEV features.
[0039] In an embodiment of the present invention, multimodal BEV features can be determined through image data and point cloud data, and through multimodal BEV features, one or more of the vehicle's risk perception status information, sparse environmental scene information and / or fused interaction information of intelligent bodies surrounding the vehicle can be determined.
[0040] In some examples, this multimodal BEV feature can be used to construct a perception region centered on the vehicle itself, using the BEV center point as a reference. Based on the multimodal BEV features within this perception region, environmental information surrounding the vehicle (such as road shape, lane markings, nearby vehicles, people, obstacles, etc.) is extracted to generate a BEV perception feature. This BEV perception feature can be used to determine whether there is interaction risk between the vehicle and surrounding intelligent agents, thereby obtaining information about the vehicle's risk perception status.
[0041] In some examples, a spatial attention mechanism can identify key scene elements (e.g., intersections, neighboring vehicles) within multimodal BEV features, compressing tens of thousands of BEV raster features into a lower-dimensional sparse environmental scene vector (i.e., sparse environmental scene information). This process mimics the attention mechanism of human drivers, retaining only environmental elements relevant to trajectory decisions (e.g., drivable area boundaries, traffic participant locations), filtering out irrelevant background information, and preventing redundant information from interfering with the decision-making model.
[0042] In some examples, multimodal BEV features can be used to extract the features and positions of all intelligent agents around the vehicle. The features and positions of each intelligent agent can be used to perform trajectory regression and future behavior modeling on the agent, thereby extracting a structured interaction vector (i.e., the fused interaction information of the agents).
[0043] In some examples, such as Figure 3 As shown, the environmental information in the method for generating a driving trajectory of an autonomous driving vehicle provided by an embodiment of the present invention can be determined through steps S210 to S240.
[0044] S210: Process the image data based on a convolutional neural network to obtain image features, and convert the image features into BEV spatial features.
[0045] In this embodiment of the present invention, an image sequence from a multi-view camera on a vehicle is received and processed using a convolutional neural network (e.g., a ResNet50-FPN architecture) with hierarchical semantic extraction capabilities to extract image features. These features encode color, texture, and structure information while preserving spatial details. These image features are then converted to BEV (bird's-eye view) space using a camera-to-plane mapping function to obtain BEV spatial features.
[0046] S220. Process the point cloud data based on a sparse three-dimensional convolutional neural network to obtain BEV point cloud spatial features.
[0047] In an embodiment of the present invention, point cloud data can be input into a sparse 3D (three-dimensional) convolutional network (such as SparseConv), through which the point cloud can be compressed and projected in the height direction to generate dense BEV point cloud spatial features, which have precise spatial geometric information.
[0048] S230. Based on the cross-attention mechanism, the BEV point cloud spatial features are used as query information, the BEV spatial features are processed to obtain multimodal BEV features.
[0049] In an embodiment of the present invention, BEV point cloud spatial features and BEV spatial features can be input into a designed multimodal BEV feature converter to align and fuse the BEV point cloud spatial features and BEV spatial features. The multimodal BEV feature converter is composed of several layers of deformable cross-attention modules. The BEV point cloud spatial features act as query tokens in the aggregation of deformable cross-attention, guiding the alignment and fusion of the BEV spatial features of the image data in terms of position. The multimodal BEV feature converter also includes a temporal attention module. The temporal attention mechanism of the temporal attention module is used to encode the dynamic evolution features between consecutive frames, thereby constructing a multimodal BEV feature that fuses the time series. The multimodal BEV features of each frame are equivalent to a global BEV feature map. The size information of each frame of the BEV feature map can be pre-set according to usage requirements. For example, the resolution can be set to 0.5m / pixel and the coverage range can be set to 100m×100m.
[0050] S240: Determine the risk perception state information of the vehicle, the sparse environment scene information, and the fused interaction information of the intelligent bodies surrounding the vehicle through multimodal BEV characteristics.
[0051] In embodiments of the present invention, image data can be processed using a convolutional neural network to obtain image features. These image features can then be converted to BEV space using a camera-to-plane mapping function to obtain BEV spatial features. Point cloud data can then be processed using a sparse three-dimensional convolutional neural network to obtain BEV point cloud spatial features. Using a cross-attention mechanism, BEV point cloud spatial features can be used as query information to process these features and obtain multimodal BEV features. The resulting multimodal BEV features can be used to obtain information about the vehicle's risk perception status, sparse environmental scene information, and fused interaction information about the intelligent agents surrounding the vehicle. In embodiments of the present invention, a lightweight pyramid fusion structure and an adaptive time window mechanism are constructed using a convolutional neural network, a sparse three-dimensional convolutional neural network, a cross-attention mechanism, and a spatial attention mechanism. This reduces computational complexity while improving the temporal and spatial alignment of asynchronous image and point cloud data, thereby reducing the error and collision rate of the subsequently generated vehicle trajectory, improving the stability and accuracy of the generated vehicle trajectory, and increasing the speed of vehicle trajectory generation.
[0052] In some embodiments of the present invention, reference Figure 2 and Figure 4 In the method for generating a driving trajectory for an autonomous vehicle provided in an embodiment of the present invention, step S22 (determining environmental information based on multimodal BEV characteristics) may include determining the vehicle's risk perception status information using the multimodal BEV characteristics. Determining the vehicle's risk perception status information using the multimodal BEV characteristics may include steps S221 and S222.
[0053] S221. Based on the multimodal BEV features, determine the local BEV feature information centered on the vehicle and the status information of the intelligent bodies around the vehicle.
[0054] S222: Determine the risk perception status information of the vehicle based on the local BEV feature information and the status information of the intelligent bodies surrounding the vehicle.
[0055] In step S221, the multimodal BEV features are used to determine the state information of the surrounding intelligent entities. The state information of the intelligent entities may include their characteristics, location information, and classification information. The classification information includes the type of intelligent entity (e.g., vehicle, person, obstacle, etc.) and the classification confidence.
[0056] Using multimodal BEV features, a perception region centered on the vehicle itself can be constructed, with the BEV center point as a reference. The shape of this perception region can be square, circular, rectangular, or other shapes, with no specific restrictions. The multimodal BEV features within this perception region can be used to extract environmental information surrounding the vehicle (such as road shape, lane markings, nearby vehicles, people, obstacles, etc.) to generate BEV perception features.
[0057] In step S222 , risk-enhanced BEV feature information may be determined based on the local BEV feature information and the status information of the intelligent bodies surrounding the vehicle, and the risk perception status information of the vehicle may be determined based on the risk-enhanced BEV feature information.
[0058] More specifically, it can be assumed that the vehicle is always located at the center of the multimodal BEV signature. This allows the vehicle's position coordinates to be determined. Based on the vehicle's position coordinates and the position coordinates of the surrounding intelligent entities, the distance between the vehicle and the surrounding intelligent entities (such as surrounding vehicles, pedestrians, or other obstacles) can be calculated. This allows the minimum distance between the vehicle and the nearest intelligent entity to be determined. This minimum distance is compared with a set risk threshold. If the minimum distance is lower than the set risk threshold, a potential interaction risk between the vehicle and the surrounding intelligent entities is determined. The inverse of the minimum distance is used as a risk coefficient, and a risk vector is generated based on this risk coefficient. This risk vector is then concatenated with the local BEV signature to generate risk-enhanced BEV signature information containing risk information.
[0059] Subsequently, based on the risk-enhanced BEV characteristics, the vehicle's location coordinates, and the vehicle navigation instructions, multimodal vehicle risk perception state information containing risk information, vehicle location information, and navigation semantic instructions can be determined, thereby improving the adaptability of the solution of the present invention in complex traffic interaction scenarios (such as congestion or conflict scenarios), further reducing the error and collision rate of the subsequently generated vehicle driving trajectory, and improving the stability and accuracy of the generated vehicle driving trajectory.
[0060] In order to better understand the embodiments of the present invention, the technical solutions described in the embodiments of the present invention are explained below through specific examples. It should be understood that the following examples are only used to explain the present invention, and are not used to limit the present invention.
[0061] like Figure 5 As shown, the process of determining the vehicle risk perception status information in the driving trajectory generation method of the autonomous driving vehicle provided by an embodiment of the present invention is as follows.
[0062] 1. Local BEV feature extraction As mentioned above, the multimodal BEV features of each frame are equivalent to a global BEV feature map. Within the global BEV feature map, a square perception region centered on the ego vehicle is constructed with the BEV center point as a reference. The size of this square perception region can be 1 / 4 the size of the BEV feature map. This region is used to extract important environmental semantics around the ego vehicle (such as road shape, lane markings, obstacles, etc.).
[0063] The multimodal BEV features of the region can be enhanced through multiple (e.g., 2, 3, or 4, etc.) convolutional layers, and then global average pooling is used to generate a local BEV feature vector.
[0064] 2. Vehicle Position Coding Assuming that the ego vehicle is always located at the center of the BEV, the coordinates of the vehicle can be expressed as (0.5, 0.5), which can be encoded through a multi-layer perceptron module to capture the relative geometric position information of the ego vehicle in the global BEV feature map and form a position feature vector.
[0065] Through the above-mentioned local BEV feature extraction and vehicle position encoding, the solution of the embodiment of the present invention can distinguish the direction and position in the BEV feature map.
[0066] 3. Risk Perception Modeling Based on the positions of the surrounding intelligent entities and the position of the vehicle itself, the Euclidean distance between the vehicle and all surrounding intelligent entities can be calculated, and the minimum distance to the intelligent entity closest to the vehicle can be screened out. This minimum distance is compared with a risk threshold (for example, 0.1 (equivalent to 10m), 0.2 (equivalent to 20m), 0.3 (equivalent to 30m), or 0.4 (equivalent to 40m)). If the minimum distance is lower than the risk threshold, it is considered that there is a potential interaction risk. The inverse of the minimum distance is used as the risk coefficient, and a risk vector is generated based on this risk coefficient. The local BEV feature vector is spliced with the risk vector, and the spliced vector is input into the risk perception network module for fusion, and a risk-enhanced BEV feature vector is output, thereby improving the response capability of this solution in congested or conflict scenarios.
[0067] Among them, the structure of the risk perception network module in the present invention can be composed of multiple cascaded convolutional layers, or a multi-layer perceptron structure, which is not specifically limited here.
[0068] 4. Adaptive Feature Fusion The risk-enhanced BEV feature vector, the ego-vehicle position feature vector, and the navigation semantic vector generated by the navigation instructions are concatenated to form a multi-source semantic vector sequence.
[0069] A fully connected layer outputs a fusion weight, which includes a first weight corresponding to the risk-enhanced BEV feature vector, a second weight corresponding to the ego-vehicle position feature vector, and a third weight corresponding to the navigation semantic vector. The first, second, and third weights respectively represent the degree of attention paid to the corresponding features. Therefore, to enhance the dominant role of local BEV features in trajectory generation, a positive bias (e.g., +0.3, +0.4, +0.5, or +0.6) can be added to the first weight corresponding to this feature to obtain a modified first weight. The modified first, second, and third weights are normalized using the softmax function to obtain the final fusion weight.
[0070] Finally, the risk-enhanced BEV feature vector, the ego vehicle position feature vector, and the navigation semantic vector generated by the navigation command in the multi-source semantic vector sequence are weighted and summed according to their corresponding weights to obtain the ego vehicle state vector, i.e., the vehicle's risk-aware state information. This risk-aware state information can be used as a query vector for ego vehicle trajectory prediction or scenario interaction.
[0071] In some embodiments of the present invention, determining the sparse environmental scene information based on the image data and the point cloud data may include: determining the sparse environmental scene information based on multimodal BEV features.
[0072] In some embodiments of the present invention, reference Figure 2 and Figure 4 In the method for generating a driving trajectory of an autonomous vehicle provided in an embodiment of the present invention, step S22 (determining environmental information based on multimodal BEV features) may include determining sparse environmental scene information based on the multimodal BEV features. Determining sparse environmental scene information using the multimodal BEV features may include steps S223 and S224.
[0073] S223: Determine an enhanced multimodal BEV feature based on the region related to the navigation instruction in the multimodal BEV feature.
[0074] S224. Through the spatial attention mechanism, target elements are extracted from the enhanced multimodal BEV features to obtain sparse environmental scene information.
[0075] In step S223, based on the channel attention mechanism and the vehicle navigation instructions, the area related to the navigation instructions in the multimodal BEV feature can be enhanced to obtain an enhanced multimodal BEV feature.
[0076] More specifically, the vehicle navigation instructions can be encoded into a navigation instruction vector, and the navigation instruction vector and the multimodal BEV features can be calculated through the channel attention mechanism, thereby dynamically enhancing the area of the multimodal BEV features related to the navigation instructions to obtain enhanced multimodal BEV features.
[0077] In step S224, spatial and semantic joint features can be generated based on the enhanced multimodal BEV features and pre-configured location information; through the spatial attention mechanism, target elements are extracted from the spatial and semantic joint features to obtain sparse environmental scene information.
[0078] More specifically, the pre-configured position information may be a position encoding matrix, and by concatenating the position encoding matrix with the enhanced multimodal BEV feature, a spatial and semantic joint feature may be established.
[0079] The spatial attention mechanism is used to automatically identify target elements in the joint spatial and semantic features, i.e., key scene elements (such as intersections and neighboring vehicles), and compress tens of thousands of BEV grid features into sparse environmental semantic vectors (i.e., sparse environmental scene information) of low dimension (e.g., 16*512 dimensions).
[0080] The embodiments of the present invention compress dense environmental scenes into sparse environmental scenes, thereby reducing the amount of computation required and increasing the speed of calculation when generating vehicle trajectories. Furthermore, the embodiments of the present invention can simulate the attention mechanism of human drivers, retaining only environmental elements relevant to trajectory decisions (such as the boundaries of the drivable area and the positions of traffic participants), filtering out irrelevant background information, and eliminating the need for manual labeling of obstacle bounding boxes and lane lines. This reduces supervision costs and avoids interference of redundant information on subsequent trajectory generation, further reducing errors in vehicle trajectory generation and improving the stability and accuracy of vehicle trajectory generation.
[0081] In some embodiments of the present invention, determining the fused interaction information of the intelligent bodies around the vehicle based on the image data and the point cloud data may include: determining the fused interaction information of the intelligent bodies around the vehicle based on multimodal BEV features.
[0082] In some embodiments of the present invention, reference Figure 2 and Figure 4 In the method for generating a driving trajectory for an autonomous vehicle provided in an embodiment of the present invention, step S22 (determining environmental information based on multimodal BEV features) may include determining fused interaction information of intelligent agents surrounding the vehicle based on the multimodal BEV features. Determining the fused interaction information of intelligent agents surrounding the vehicle based on the multimodal BEV features may include steps S225 and S226.
[0083] S225 : Determine status information of intelligent entities surrounding the vehicle based on multimodal BEV characteristics.
[0084] S226: Generate fused interaction information of the intelligent agents based on the state information of the intelligent agents surrounding the vehicle and the pre-configured behavior pattern embedding information. In some embodiments of the present invention, the state information of the intelligent agents may include characteristics and location information of the intelligent agents, as well as classification information. The classification information includes the type of intelligent agent (e.g., vehicle, person, obstacle, etc.) and the classification confidence.
[0085] In an embodiment of the present invention, multimodal BEV features can be input into a Transformer decoder, and trajectory regression and future behavior modeling can be performed on each agent through the features and position information of the agent output by the decoder.
[0086] The decoder has multiple layers. Initial decoding layers use the initial reference position, while intermediate decoding layers use the intermediate reference positions generated by the intermediate layers. All reference positions are converted to continuous spatial coordinates. For each layer's output agent feature, the classification and regression heads output the agent feature's category, the corresponding classification confidence score, and regression parameters such as relative position and velocity. The agent's relative position coordinates output by the regression head are then decoded. The predicted position coordinates are added to the reference point (either the initial reference position coordinates or the intermediate reference position coordinates) and mapped to the range [0, 1] using a sigmoid activation to obtain normalized coordinates. These coordinates represent the agent's relative position in BEV space. These normalized coordinates can also be mapped back to actual physical space (for example, by rescaling based on the point cloud extent). The resulting coordinates [x, y] represent the actual position of the predicted agent feature in BEV space. The output categories and position coordinates of all layers can be cached for subsequent supervision.
[0087] In an embodiment of the present invention, the state information of the intelligent agents surrounding the vehicle can be determined through multimodal BEV features. The state information of the intelligent agents and the self-attention mechanism can be used to generate the multimodal trajectory features of the intelligent agents. The fused interaction information of the intelligent agents includes multiple possible intended trajectories of the vehicle in the future, enabling downstream planning to have probabilistic and diverse expression capabilities. The multimodal trajectory features of the intelligent agents can be used to integrate the features of multiple intelligent agents surrounding the vehicle to obtain the fused interaction information of the intelligent agents. This fused interaction information of the intelligent agents can provide spatial semantic support for subsequent interaction modeling with the vehicle itself and improve the stability of subsequent vehicle trajectory generation.
[0088] In some embodiments of the present invention, Figure 4 and 6As shown, S226 (generating fusion interaction information of intelligent agents based on the state information of the intelligent agents around the vehicle and the pre-configured behavior pattern embedding information) may include steps S2261 and S2262.
[0089] Step S2261: Determine the multimodal trajectory characteristics of the intelligent agents surrounding the vehicle based on the state information of the intelligent agents surrounding the vehicle and the pre-configured behavior pattern embedding information.
[0090] Step S2262: Generate fusion interaction information of the intelligent bodies around the vehicle based on the multimodal trajectory characteristics of the intelligent bodies around the vehicle.
[0091] In an embodiment of the present invention, based on the pre-configured behavior pattern embedding information of the agent's features and position information, the agent's multimodal behavior trajectory in the future (i.e., multimodal trajectory features) can be obtained. Specifically, for the agent features output by the last layer of the decoder, the behavior pattern embedding (such as going straight / changing lanes / slowing down) defined in the motion pattern query features can be added to the agent features output by the last layer of the decoder to obtain generated multimodal intention query information that contains both behavior information and navigation semantic information. Based on the agent's position information, a position vector is generated as the query position embedding information of the agent. By calling the action decoder to perform self-attention decoding on the agent's multimodal intention query information and query position embedding information, the agent's multimodal behavior trajectory in the future, i.e., multimodal trajectory features, can be obtained. By processing the agent's multimodal trajectory features, the features of all agents can be unified to obtain the agent's fused interaction information.
[0092] In the embodiments of the present invention, by considering multiple possible future intention trajectories of each intelligent agent, downstream vehicle trajectory planning is endowed with probabilistic and diverse expression capabilities, thereby improving the robustness of the solution of the present invention and enabling the solution of the present invention to adapt to complex traffic interaction scenarios, further reducing the error and collision rate of subsequently generated vehicle driving trajectories, and improving the stability and accuracy of the generated vehicle driving trajectories.
[0093] In the embodiment of the present invention, Figure 6 and 7 As shown, step S2262 (generating fusion interaction information of the intelligent agent based on the multimodal trajectory characteristics of the intelligent agent) may include steps S700~S730.
[0094] S700: Through a multi-layer dimensionality reduction perceptron, the multimodal trajectory features of all surrounding intelligent agents of the vehicle are reduced in dimensionality to obtain a unified query vector.
[0095] S710: Filter out a first target agent from all surrounding agents based on the classification information.
[0096] S720: Based on the unified query vector of the first target agent, fill in the unified query vectors of other agents in all surrounding agents.
[0097] S730. Through a multi-layer dimensionality-raising perceptron, the unified query vector of all surrounding intelligent agents is upgraded to obtain the fused interaction information of all surrounding intelligent agents.
[0098] In step S700, a multi-layer dimensionality reduction perceptron is used to implement vector dimensionality reduction. Through the multi-layer dimensionality reduction perceptron, the multimodal trajectory features of the intelligent agent can be reduced in dimension and integrated into a unified query vector.
[0099] In step S710, the classification information of each agent includes the category of the agent and the corresponding classification confidence. Based on the classification information of each agent, the agent with a classification confidence greater than a threshold in the classification information can be screened out from all surrounding agents of the vehicle as the first target agent, and the remaining agents among all surrounding agents are regarded as other agents. The threshold value can be set according to the use requirements. For example, the threshold value can be set to 0.5, 0.6, 0.7 or 0.8, etc., which is not specifically limited here. Accordingly, the number of first target agents can be one or more, depending on the specific screening situation.
[0100] In step S720, the unified query vector of the first target agent can be used to fill the unified query vectors of other agents in all surrounding agents, so that the unified query vectors of other agents are consistent with the unified query vector of the first target agent, thereby improving the stability of subsequent vehicle trajectory generation.
[0101] If there are multiple first target agents, a range can be defined to identify the other agents surrounding each first target agent. For example, with the first target agent as the center, other agents within a certain distance range of the first target agent can be used as the first surrounding agents of the first target agent, and the unified query vector of the first surrounding agents can be filled based on the unified query vector of the first target agent.
[0102] In step S730, the multi-layer dimensionality-raising perceptron (MDR) performs vector dimensionality-raising. After integrating the unified query vectors of all surrounding agents through steps S700 to S720, the MLR can be used to perform dimensionality-raising on the unified query vectors of all surrounding agents, resulting in a fused interaction vector (i.e., fused interaction information) for all surrounding agents.
[0103] Reference Figure 1 and Figure 8In some embodiments of the present invention, environmental information includes: sparse environmental scene information, fused interaction information of intelligent agents, and risk perception status information of vehicles. Step S30 (determining scene and behavior joint representation information based on environmental information) may include steps S31 and S32.
[0104] S31. Determine vehicle-scene interaction information based on sparse environment scene information and risk perception state information of the vehicle.
[0105] S32. Determine scene and behavior joint representation information based on the vehicle-scene interaction information and the fused interaction information of the intelligent agent.
[0106] In step S31, the vehicle-scene interaction information is obtained by combining the sparse environmental scene information and the vehicle's risk perception status information. This vehicle-scene interaction information includes the relationship between the vehicle's motion and the scene topology, such as the matching relationship between the vehicle speed and the road curvature, and the matching relationship between the vehicle and the lane.
[0107] More specifically, a cross-attention mechanism can be used to calculate the sparse environment scene vector (i.e., sparse environment scene information) and the vehicle's risk perception state vector (i.e., risk perception state information) to obtain the vehicle-scene interaction vector (i.e., vehicle-scene interaction information). The sparse environment scene vector serves as the query vector (Q) during the calculation, and the risk perception state vector serves as the key-value vector (K, V).
[0108] In step S32, the scene and behavior joint representation information can be obtained through the fusion interaction information between the vehicle and the scene and the intelligent body, which provides an environmental basis for the subsequent generation of the vehicle's driving trajectory.
[0109] More specifically, the vehicle-scene interaction vector and the agent's fused interaction vector can be fused using a masked attention mechanism to obtain fused scene-behavior interaction information. The vehicle-scene interaction vector serves as the key vector for the masked attention calculation, while the agent's fused interaction vector serves as the query vector for the masked attention calculation. The fused scene-behavior interaction information is then calculated using a self-attention mechanism to obtain a joint scene-behavior representation vector (i.e., scene-behavior joint representation information).
[0110] In order to better understand the embodiments of the present invention, the technical solutions described in the embodiments of the present invention are explained below through specific examples. It should be understood that the following examples are only used to explain the present invention, and are not used to limit the present invention.
[0111] The embodiment of the present invention can be implemented by a cascade scene decoder. Figure 9As shown, the process of determining the scene and behavior joint representation information in the driving trajectory generation method of the autonomous driving vehicle provided by the embodiment of the present invention is as follows.
[0112] First, the ego vehicle-scene interaction is performed. A cross-attention calculation is performed on the ego vehicle state vector and the sparse environment semantic vector to obtain the vehicle-scene interaction vector (i.e., vehicle-scene interaction information). This vehicle-scene interaction vector includes the relationship between the ego vehicle's motion and the scene topology, such as the matching relationship between vehicle speed and road curvature, and the matching relationship between the ego vehicle and the lane.
[0113] Then, agent-scene interaction is performed. The interaction vectors of the surrounding agents are fused with the vehicle-scene interaction vector via a masked attention mechanism. Because the unified query vector of the first target agent among the surrounding agents is unified with the unified query vectors of the other agents, the masked attention mechanism ensures that the interaction vectors of the agents and the vehicle-scene interaction vectors are fused with the interaction features of agents with high conflict risk, resulting in a fused interaction feature of the scene and behavior.
[0114] Finally, the representation is refined, and the scene and behavior fusion interaction features are calculated through the self-attention mechanism to eliminate feature conflicts and output the joint representation information of the scene and behavior.
[0115] In embodiments of the present invention, scene-behavior joint representation information can be determined based on vehicle-scene interaction information and the fused interaction information of the intelligent agent. Vehicle-scene interaction information can also be determined based on sparse environmental scene information and the vehicle's risk perception status information. This joint scene-behavior representation information is highly interpretable, providing an environmental basis for driving decisions, ensuring accurate scene understanding and rational behavior, and enabling coordinated optimization of perception and decision-making. This further reduces the error and collision rate of the subsequently generated vehicle trajectory, while improving the stability and accuracy of the generated vehicle trajectory.
[0116] Reference Figure 1 and Figure 10 In some embodiments of the present invention, step S40 (generating the vehicle's driving trajectory based on the scene and behavior joint representation information) may include steps S41 and S42.
[0117] S41. Configure the initial trajectory point sequence of the vehicle in the future.
[0118] S42. Based on the cross-attention mechanism, the vehicle's driving trajectory in the future is generated according to the joint representation information of the scene and behavior and the initial trajectory point sequence.
[0119] In step S41, an initial trajectory point sequence for the vehicle in the future can be configured. This initial trajectory point sequence can include the vehicle's position in the future. For example, if the vehicle currently captures an image or point cloud frame every 0.5 seconds, the initial trajectory point sequence can include six vehicle positions in the next three seconds.
[0120] In step S42, the vehicle's future trajectory can be generated based on the cross-attention mechanism, the joint scene and behavior representation information, and the initial trajectory point sequence. Because this joint scene and behavior representation information is generated by taking into account the interaction vectors of the vehicle's surrounding agents, sparse environmental scene information, and the vehicle's risk perception state, the generated vehicle trajectory is guaranteed to have multiple trajectory modes (e.g., straight ahead, left deviation, right deviation), adapting to complex traffic scenarios and improving the quality of the generated trajectory.
[0121] In some examples, a multi-head attention mechanism can also be used to enable each trajectory point to dynamically focus on relevant scene elements (such as focusing on the intersection location in the first second and focusing on vehicles in the merging area in the third second), thereby further improving the dynamic adaptability of the generated driving trajectory.
[0122] In some examples, the vehicle may select a trajectory that matches the navigation semantic instruction from the generated driving trajectory containing multiple trajectory modalities in combination with the navigation semantic instruction as the final driving trajectory.
[0123] In some examples, when a sudden change in the environment is detected through point cloud data and / or image data, the vehicle's driving trajectory can be regenerated based on the above scheme, and the vehicle can travel according to the newly generated driving trajectory in the future to improve the dynamic adaptability of the present invention.
[0124] In some examples, the output format of the generated vehicle driving trajectory can be compatible with a vehicle control interface, thereby enabling the vehicle to directly drive the vehicle steering and power systems according to the generated vehicle driving trajectory.
[0125] By integrating deep image and point cloud data, modeling sparse environmental scene information, integrating interactive information about surrounding intelligent agents, and modeling the vehicle's risk perception status, and introducing navigation semantic instructions, the present invention achieves accurate modeling and safe decision-making for the vehicle and surrounding intelligent agents in complex traffic scenarios. This improves the vehicle's stability and safety in diverse dynamic traffic scenarios. The present invention's solution automatically completes key perception and prediction tasks and updates trajectory planning in real time, ensuring stable driving and rational behavior even when the external environment undergoes drastic changes.
[0126] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a driving trajectory generation system for an autonomous driving vehicle, such as Figure 11As shown, the driving trajectory generation system 110 of the autonomous driving vehicle includes: at least one processor 111; and a memory 112, the memory 112 stores a computer program that can be run on the processor 111, and when the processor 111 executes the program, it performs the steps of the driving trajectory generation method described in the embodiment of the present invention.
[0127] The memory, as a non-volatile storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the methods described in the embodiments of this application. The processor executes the non-volatile software programs, instructions, and modules stored in the memory to execute various functional applications and data processing of the device, thereby implementing the methods of the above embodiments.
[0128] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the device, etc. In addition, the memory may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the local module via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0129] Finally, it should be noted that those skilled in the art will understand that all or part of the processes in the above-described method embodiments can be implemented using a computer program to instruct the relevant hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes in the above-described method embodiments. The program storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM). The above-described computer program embodiments can achieve the same or similar effects as any of the corresponding aforementioned method embodiments.
[0130] It will also be appreciated by those skilled in the art that the various exemplary logic blocks, modules, circuits and algorithmic steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software or a combination of the two. In order to clearly illustrate this interchangeability of hardware and software, a general description has been given of the functions of various schematic components, blocks, modules, circuits and steps. Whether this function is implemented as software or hardware depends on specific applications and the design constraints imposed on the entire system. Those skilled in the art can implement the function in various ways for each specific application, but this implementation decision should not be interpreted as causing a departure from the disclosed scope of the embodiments of the present invention.
[0131] The above are exemplary embodiments disclosed in the present invention, but it should be noted that various changes and modifications can be made without departing from the scope of the disclosure of the embodiments of the present invention as defined in the claims. The functions, steps and / or actions of the method claims according to the disclosed embodiments described herein do not need to be performed in any particular order. The serial numbers of the embodiments disclosed in the above embodiments of the present invention are for description only and do not represent the advantages and disadvantages of the embodiments. In addition, although the elements disclosed in the embodiments of the present invention can be described or required in individual form, they can also be understood as multiple unless expressly limited to the singular.
[0132] It should be understood that, as used herein, the singular forms "a" and "an" are intended to include the plural forms as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" is intended to include any and all possible combinations of one or more of the associated listed items.
[0133] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to limit the scope of the disclosure of the present invention (including the claims) to these examples. Within the spirit of the present invention, the technical features of the above embodiments or different embodiments may be combined, and many other variations exist in different aspects of the above embodiments, which are not provided in detail for the sake of clarity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for generating a driving trajectory of an autonomous vehicle, characterized in that: include: Acquire image data and point cloud data based on vehicle sensors; determining environmental information based on the image data and the point cloud data, the environmental information including risk perception status information of the vehicle; Determining scene and behavior joint representation information based on the environmental information; Based on the scene and behavior joint representation information, a driving trajectory of the vehicle is generated.
2. The method according to claim 1, characterized in that Determining environmental information based on the image data and the point cloud data includes: determining a multimodal BEV feature based on the image data and the point cloud data; Based on the multimodal BEV characteristics, the environmental information is determined.
3. The method according to claim 2, characterized in that Determining the risk perception status information of the vehicle based on the image data and the point cloud data includes: Determining local BEV feature information centered on the vehicle and state information of intelligent entities surrounding the vehicle based on the multimodal BEV feature; Based on the local BEV feature information and the status information of the vehicle surrounding intelligent body, risk perception status information of the vehicle is determined.
4. The method according to claim 3, characterized in that Determining the risk perception state information of the vehicle based on the local BEV feature information and the state information of the intelligent body surrounding the vehicle includes: Determining interaction risk information between the vehicle and the vehicle surrounding intelligent entities based on the position information of the vehicle in the multimodal BEV feature and the position information of the vehicle surrounding intelligent entities in the multimodal BEV feature; determining risk-enhanced vehicle characteristic information based on the local BEV characteristic information and the interactive risk information; The risk-enhanced vehicle feature information, the vehicle's location information, and the vehicle navigation instructions are fused to obtain risk perception status information of the vehicle.
5. The method according to claim 2, characterized in that The environmental information also includes sparse environmental scene information; determining the sparse environmental scene information based on the image data and the point cloud data further includes: determining an enhanced multimodal BEV signature based on a region of the multimodal BEV signature associated with the navigation instruction; Through the spatial attention mechanism, target elements are extracted from the enhanced multimodal BEV features to obtain sparse environmental scene information.
6. The method according to claim 1, characterized in that The environmental information also includes fusion interaction information of the intelligent bodies around the vehicle; Determining fused interaction information of the intelligent bodies surrounding the vehicle based on the image data and the point cloud data includes: Determining state information of the intelligent body surrounding the vehicle based on the image data and the point cloud data; Based on the state information of the vehicle surrounding intelligent agents and pre-configured behavior pattern embedding information, fusion interaction information of the intelligent agents is generated.
7. The method according to claim 6, characterized in that Generating fusion interaction information of the intelligent agents based on the state information of the intelligent agents around the vehicle and the pre-configured behavior pattern embedding information includes: Determining multimodal trajectory features of the vehicle surrounding intelligent agent based on state information of the vehicle surrounding intelligent agent and pre-configured behavior pattern embedding information; Based on the multimodal trajectory features of the vehicle surrounding intelligent agents, fused interaction information of the vehicle surrounding intelligent agents is generated.
8. The method according to claim 1, characterized in that The environmental information also includes sparse environmental scene information and fused interaction information of the intelligent bodies surrounding the vehicle; Determining the scene and behavior joint representation information based on the environmental information includes: Determining vehicle-scene interaction information based on the sparse environment scene information and the risk perception state information of the vehicle; Based on the vehicle-scene interaction information and the fused interaction information of the intelligent bodies surrounding the vehicle, scene and behavior joint representation information is determined.
9. The method according to claim 1, characterized in that Generating the vehicle's driving trajectory based on the scene and behavior joint representation information includes: Configuring an initial trajectory point sequence of the vehicle in the future; Based on the cross-attention mechanism, the vehicle's driving trajectory in the future is generated according to the joint representation information of the scene and behavior and the initial trajectory point sequence.
10. A driving trajectory generation system for an autonomous vehicle, comprising: at least one processor; as well as A memory storing a computer program that can be run on the processor, wherein the processor performs the steps of the method according to any one of claims 1 to 9 when executing the program.
Citation Information
Patent Citations
Automatic driving vehicle track prediction method and device and electronic equipment
CN113705636A
Automatic driving multi-mode cooperative sensing method and system based on BEV visual angle
CN116977963A
Automatic driving vehicle driving track planning method, device and equipment and storage medium
CN117492447A
Three-dimensional space occupation identification method and system in automatic driving scene
CN117557985A
Multi-modal fusion method for heterogeneous data of intelligent networked vehicle multi-source sensor
CN118445748A