Map feature extraction method based on attention mechanism

By adopting the map feature extraction method based on attention mechanism in trajectory prediction, the problem of neglecting relationship between node-level features and lane-level features is solved, and a more complete map topology and more accurate trajectory prediction are achieved.

CN119963850APending Publication Date: 2025-05-09SHENYANG UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510029448.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art ignores the relationship between node-level features and lane-level features in trajectory prediction, resulting in incomplete generation of the topological structure of the map.

Method used

The map feature extraction method based on attention mechanism is adopted, through Self-Attention and Cross-Attention mechanisms, the interaction information between node-level features and lane-level features is extracted, and a global interaction diagram of lane sections is constructed to capture the global connection between different lane sections.

Benefits of technology

By constructing the interactive relationship between node-level and lane-level features, the integrity of map features and the accuracy of topological structure are enhanced, thereby improving the performance of trajectory prediction algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963850A_ABST
    Figure CN119963850A_ABST
Patent Text Reader

Abstract

The invention discloses a map feature extraction method based on an attention mechanism, and the method comprises the steps: carrying out the preprocessing of a given high-precision map, carrying out the vectorization of the high-precision map, carrying out the partitioning of node elements in the map according to lane segments, and carrying out the mapping of the feature of each node element into a high-dimensional vector through MLP; performing feature processing on the lane section, the local node features and nodes belonging to the same lane section so as to output the local node features, the local lane section features and the global lane section features; according to the method, map node-level features and lane-level features are constructed at the same time, so that an intelligent agent can capture more comprehensive map connection features as much as possible. In general work, lane features are extracted from lane segment nodes, then global lane segment features are constructed, and lane-level features are used in subsequent space-time interaction modeling. The node level and the lane level are used for jointly representing the features of the map, so that the topological structure of the map is more complete.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of high-precision map feature extraction in trajectory prediction, and more specifically, to a map feature extraction method based on an attention mechanism. Background Art

[0002] For trajectory prediction technology itself, the quality of prediction results depends not only on high-precision prediction algorithms, but also on high-precision perception input and the model's understanding of road topology. Due to the complexity of the road topology of the urban structured road environment, the diversity of road elements and traffic participants, the complexity of the interaction between traffic participants and between traffic participants and road elements, and the high uncertainty of the future movement of objects, how to accurately predict the future trajectory and behavior intention of multiple types of moving objects around is a very challenging problem, and it is also the direction that many scholars and engineers are constantly exploring.

[0003] Due to the limitations of using CNN and rasterized scene encoding, many recent works have been exploring the expression methods of map vectorization. Map vectorization refers to the conversion of high-precision maps from the original raster form and unstructured form to a structured vector form for computer mathematical calculations. Vectorized map representation can retain enough map elements such as road centerlines, intersections, lane lines, etc. expressed in geometric shapes such as line segments, curves or polygons. They can be used to accurately describe the topological structure and geometric properties of the road network. Due to its structured characteristics, vectorized maps can more accurately represent information such as road width, length, curvature, slope, etc. Compared with raster maps, vectorized high-precision maps usually contain less redundant data and can be used more efficiently in trajectory prediction models. While reducing the consumption of computing resources, they help to extract general features from local features, thereby having better generalization capabilities in different regions or under different conditions. In addition, vector maps can support representations of different scales (such as node level, lane level, etc.) by adjusting the fine granularity.

[0004] However, previous works only encode the relationship between node-level features or lane-level features, while ignoring the relationship between node-level features and lane-level features, which results in incomplete topological structure generation of the map. Summary of the invention

[0005] The purpose of the present invention is to provide a map feature extraction method based on an attention mechanism, which aims to solve the problem that previous work only encodes the relationship between node-level features or lane-level features, while ignoring the relationship between node-level features and lane-level features, thereby causing incomplete generation of the topological structure of the map.

[0006] The above technical objectives of the present invention are achieved through the following technical solutions: A map feature extraction method based on an attention mechanism, comprising the following steps: 1) Preprocess a given high-precision map. Lane segment , in order to extract the vectorized nodes Spatial and semantic features, this paper first uses MLP to directly vectorize each node Encode and get each vector node Characteristics of Dimensions .

[0007] 2) In order for each node to contain the interactive information with all nodes in the same lane segment, the present invention uses Self-Attention and residual connection to extract interactive features. Specifically, three parameter matrices are used to transform each input Mapped to three different spaces, we get the query vector , the key vector Sum value vector For the same lane segment Node feature sequence The linear mapping process can be written as: in , , are respectively the learnable parameter matrices of the linear mapping, shared in the lane segment node interaction feature computation of all lane segments. , as well as are matrices consisting of query vector, key vector and value vector respectively. The attention score is calculated using the scaled dot product, and the vector sequence output by self-attention is shown in the formula: in , represents the lane segment node feature sequence after interaction; then the residual connection and feedforward neural network are used to output the final interaction features: Where add(·) represents vector addition, LN(·) represents layer normalization, and the same calculation formula is used in the residual connection and layer normalization calculation of subsequent modules. Indicates lane segment Intra-node interaction feature sequence.

[0008] 3) Then, in order to facilitate the establishment of global road interaction, the present invention first extracts the local features of all lane segments . Local features of a lane segment It is composed of lane segments Node feature sequence Therefore, the present invention uses the Cross-Attention mechanism to learn the local features of the lane segment. Specifically, by setting a learnable query vector for each lane segment The lane segment features are fully learned from the lane segment node feature sequence, where Represents a randomly initialized value; it is composed of the node feature sequence in each lane segment Generates a tensor of key / value pairs , for a lane segment feature: in , , are respectively the learnable parameter matrices of the linear mapping, shared in all interactive computations for extracting lane segment local features from lane segment nodes. , are matrices consisting of key vectors and value vectors respectively. The scaled dot product model is used to calculate the attention score, and the vector sequence output by Self-Attention is shown in the formula: For all calculated lane segment features After residual connection and layer normalization, the local feature sequence of the lane segment is output .

[0009] 4) In order to capture the global connection between different lane segments, the present invention uses the Self-Attention mechanism and residual connection to realize the lane segment global interaction graph, which is used to extract the lane segment global interaction features. Perform Self-Attention calculation to obtain the global interactive feature sequence of lane segments .

[0010] Specifically, three parameter matrices are used to transform the local features of each input lane segment into Mapped to three different spaces, we get the query vector , the key vector Sum value vector For the entire input sequence The linear mapping process can be written as: in , , They are the learnable parameter matrices of the linear mapping, shared in computing the interaction of each lane segment with other lane segments; , and are matrices consisting of query vector, key vector and value vector respectively. The attention score is calculated using the scaled dot product, and the vector sequence output by Self-Attention is as shown in the formula: in It represents the lane segment interaction feature sequence after Self-Attention calculation. Then the residual connection and layer normalization output the lane segment global feature sequence and the final interaction feature .

[0011] 5) The global map contains nodes, then for all lane segment node feature sequences Can form a global map node feature sequence , its shape ,in In order to learn the relationship between the node and the remote lane segment, the present invention uses the Cross-Attention mechanism to learn the relationship between the lane segment Node features and lane segment global feature sequence Interactive calculations are performed. Each node feature Generate query vector , lane segment feature sequence Generates a tensor of key / value pairs For the entire map node feature sequence , the calculation relationship between node and lane follows the formula: in , , They are the learnable parameter matrices of the linear mapping, which are shared in the interactive calculation of all lane segment node features and the global lane segment features; , , are matrices consisting of query vector, key vector and value vector respectively. The scaled dot product model is used to calculate the attention score. The vector sequence output by Self-Attention is shown in the formula: in is the result of the calculation. After residual connection and layer normalization, the node feature sequence after node-lane interaction can be obtained .

[0012] The present invention has the following beneficial effects: In the present invention, map node-level features and lane-level features are constructed at the same time, allowing the intelligent agent to capture more comprehensive map connection features as much as possible. Compared with the prior art, the present invention extracts lane features from lane segment nodes, and then constructs global lane segment features. The subsequent spatiotemporal interaction modeling uses lane-level features, and by reasonably dividing the feature scale of the high-precision map, the output high-precision map features contain both map detail information and the topological information of the overall map structure, thereby enhancing the performance of the trajectory prediction algorithm using this method. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a flow chart in an embodiment of the present invention; Figure 2 is a schematic diagram of a network structure for extracting a feature map in an embodiment of the present invention; Figure 3 It is a visual comparison between lane node interaction coding in an embodiment of the present invention and other methods. DETAILED DESCRIPTION

[0014] The following is combined with Figure 1-2 The present invention is described in further detail.

[0015] Embodiment: A map feature extraction method based on attention mechanism, such as Figure 1 , Figure 2 As shown, the following steps are included: 1) Preprocess a given high-precision map. Lane segment , in order to extract the vectorized nodes Spatial and semantic features, this paper first uses MLP to directly vectorize each node Encode and get each vector node Characteristics of Dimensions .

[0016] 2) In order for each node to contain the interactive information with all nodes in the same lane segment, the present invention uses Self-Attention and residual connection to extract interactive features. Specifically, three parameter matrices are used to transform each input Mapped to three different spaces, we get the query vector , the key vector Sum value vector For the same lane segment Node feature sequence The linear mapping process can be written as: in , , are respectively the learnable parameter matrices of the linear mapping, shared in the lane segment node interaction feature computation of all lane segments. , as well as are matrices consisting of query vector, key vector and value vector respectively. The attention score is calculated using the scaled dot product, and the vector sequence output by self-attention is shown in the formula: in , represents the lane segment node feature sequence after interaction; then the residual connection and feedforward neural network are used to output the final interaction features: Where add(·) represents vector addition, LN(·) represents layer normalization, and the same calculation formula is used in the residual connection and layer normalization calculation of subsequent modules. Indicates lane segment Intra-node interaction feature sequence.

[0017] 3) Then, in order to facilitate the establishment of global road interaction, the present invention first extracts the local features of all lane segments . Local features of a lane segment It is composed of lane segments Node feature sequence Therefore, the present invention uses the Cross-Attention mechanism to learn the local features of the lane segment. Specifically, by setting a learnable query vector for each lane segment The lane segment features are fully learned from the lane segment node feature sequence, where Represents a randomly initialized value; it is composed of the node feature sequence in each lane segment Generates a tensor of key / value pairs , for a lane segment feature: in , , are respectively the learnable parameter matrices of the linear mapping, shared in all interactive computations for extracting lane segment local features from lane segment nodes. , are matrices consisting of key vectors and value vectors respectively. The scaled dot product model is used to calculate the attention score, and the vector sequence output by Self-Attention is shown in the formula: For all calculated lane segment features After residual connection and layer normalization, the local feature sequence of the lane segment is output .

[0018] 4) In order to capture the global connection between different lane segments, the present invention uses the Self-Attention mechanism and residual connection to realize the lane segment global interaction graph, which is used to extract the lane segment global interaction features. Perform Self-Attention calculation to obtain the global interactive feature sequence of lane segments .

[0019] Specifically, three parameter matrices are used to transform the local features of each input lane segment into Mapped to three different spaces, we get the query vector , the key vector Sum value vector For the entire input sequence The linear mapping process can be written as: in , , They are the learnable parameter matrices of the linear mapping, shared in computing the interaction of each lane segment with other lane segments; , and are matrices consisting of query vector, key vector and value vector respectively. The attention score is calculated using the scaled dot product, and the vector sequence output by Self-Attention is as shown in the formula: in It represents the lane segment interaction feature sequence after Self-Attention calculation. Then the residual connection and layer normalization output the lane segment global feature sequence and the final interaction feature .

[0020] 5) The global map contains nodes, then for all lane segment node feature sequences Can form a global map node feature sequence , its shape ,in In order to learn the relationship between the node and the remote lane segment, the present invention uses the Cross-Attention mechanism to learn the relationship between the lane segment Node features and lane segment global feature sequence Interactive calculations are performed. Each node feature Generate query vector , lane segment feature sequence Generates a tensor of key / value pairs For the entire map node feature sequence , the calculation relationship between node and lane follows the formula: in , , They are the learnable parameter matrices of the linear mapping, which are shared in the interactive calculation of all lane segment node features and the global lane segment features; , , are matrices consisting of query vector, key vector and value vector respectively. The scaled dot product model is used to calculate the attention score. The vector sequence output by Self-Attention is shown in the formula: in is the result of the calculation. After residual connection and layer normalization, the node feature sequence after node-lane interaction can be obtained .

[0021] The following is a verification experiment of this embodiment: The method is compared on a public dataset: Argoverse1. The statistics of the dataset are shown in Table 1: Table 1 Data representation The implementation details are as follows: Our experimental configuration uses Ubuntu 20.04 as the operating system, and selects 4 NVIDIA RTX 4090 graphics cards as computing resources. In the model training process, this paper uses the AdamW optimizer, the initial learning rate is set to 0.0003, the batch_size is set to 16, and 50 epochs are trained. In order to evaluate the map feature extraction method proposed in this invention, we connect it to the trajectory prediction method. We use three indicators, ADE, FDE, and MR, and test it on the Argoverse1 dataset. It is compared and analyzed with a series of advanced trajectory prediction methods. The specific data is shown in Table 2.

[0022] Table 2 Comparison of minADE, minFDE and MR indicators on the Argoverse1 dataset Note: The bold font is the optimal value for each row.

[0023] The numerical comparison results are as follows: As shown in the table, we tested the performance of the comparison method on the validation set of the Argoverse1 trajectory prediction dataset, where BaseLine 1 and BaseLine 2 are the baselines provided by the dataset. Since Argoverse 1 only supports the prediction of single-target agents, this paper lists the prediction performance of single-target agents. In the table, K represents the number of modes, and this paper sets 6 and 1 for testing respectively. The results show that in multi-trajectory prediction, the model proposed in this paper can achieve the best results in the table on the minADE indicator, and MR and minFDE can achieve relatively good results; in single-modal prediction, the performance of the model proposed in this paper has achieved the best results in the table on the minADE indicator, and the minFDE indicator is also relatively good. The main reason is that the road encoding more fully models the interactive information, making the prediction results more accurate.

[0024] We generate visualization results of the trajectory prediction results generated by our HD map feature extraction method. By analyzing the visualization results, such as Figure 3 shown. Figure 3 The visual comparison of lane node interaction encoding and other methods is shown. The visual comparison results are as follows: Figure 3The pink trajectory in the figure is the trajectory prediction result after using the present invention, and the dark green trajectory is the trajectory prediction result without using the present invention. It can be seen that after using the present invention, the trajectory prediction result is obviously less offset than the true value, and the future trajectory of the vehicle can be predicted more accurately.

[0025] The parameter configuration is as follows: In the parameter configuration of the model, we set the node feature, agent feature, and lane feature vector dimensions to 128, the initial learning rate to 0.0003, the weight decay to 0.0001, the dropout rate to 0.0001, and the learning rate decay to cosine decay (CosineDecay). According to the principle that the larger the batch_size, the better when the performance is satisfied, the batch_size is set to 16 and trained for 50 epochs.

[0026] Through the above verification experiments, a neural network based on the attention mechanism is used to achieve efficient HD map feature extraction and build an interactive process between lane-level features and node-level features. This method reasonably divides the HD map feature scale so that the output HD map features contain both map detail information and topological information of the overall map structure, thereby enhancing the performance of the trajectory prediction algorithm using this method. In this verification experiment, we tested the method on a widely used benchmark dataset and compared the numerical results and visual effects of the method with existing advanced methods. The results show that this method performs outstandingly in improving the accuracy of trajectory prediction results.

[0027] This specific embodiment is only an explanation of the present invention, and it is not a limitation of the present invention. After reading this specification, those skilled in the art can make modifications to this embodiment without creative contribution as needed, but they are protected by the patent law as long as they are within the scope of the claims of the present invention.

Claims

1. A map feature extraction method based on attention mechanism, characterized in that: The following steps are involved: Step 1: Preprocess the given HD map, vectorize the HD map, and divide the node elements in the map into blocks according to lane segments. Then, use MLP to map the features of each node element into a high-dimensional vector. Step 2: For the nodes in the same lane segment, generate query vector Q, key vector K, value vector V for their features, and calculate the attention score to output the local node features; Step 3: Introduce a learnable tensor and generate a query vector Q for it. At the same time, generate a key vector K and a value vector V for the local node features, then perform attention calculation and output the local lane segment features. Step 4: For all lane segments, their features are superimposed, and query vector Q, key vector K, value vector V are generated for their features, and attention scores are calculated to output global lane segment features; Step 5: Generate a query vector Q for the local node feature, and generate a key vector K and a value vector V for the global lane segment feature. Then perform attention calculation and output the global lane segment feature.

2. The map feature extraction method based on the attention mechanism according to claim 1 is characterized in that: In step 1, a given high-precision map is preprocessed. Lane segment , in order to extract the vectorized nodes Spatial and semantic features, using MLP to directly vectorize each node Encode and get each vector node Characteristics of Dimensions .

3. The map feature extraction method based on the attention mechanism according to claim 1, characterized in that: In step 2, the features are obtained through the processing of step 1 , using Self-Attention and residual connections to extract interactive features, by using three parameter matrices to transform each input Mapped to three different spaces, we get the query vector , the key vector Sum value vector , for the same lane segment Node feature sequence The linear mapping process is: in , , They are respectively the learnable parameter matrices of the linear mapping, shared in the lane segment node interaction feature computation of all lane segments; , as well as are matrices consisting of query vector, key vector and value vector respectively; The attention score is calculated using the scaled dot product, and the vector sequence output by self-attention is as shown in the formula: in , represents the lane segment node feature sequence after interaction; Then use residual connection and feedforward neural network to output the final interaction features: Where add(·) represents vector addition, LN(·) represents layer normalization, and the same calculation formula is used in the residual connection and layer normalization calculation of subsequent modules. Indicates lane segment Intra-node interaction feature sequence.

4. The map feature extraction method based on the attention mechanism according to claim 1 is characterized in that: In step 3, local features of all lane segments are extracted , a lane segment local feature It is composed of lane segments Node feature sequence Jointly decide and use the Cross-Attention mechanism to learn the local features of lane segments , by setting a learnable query vector for each lane segment , fully learn the lane segment features from the lane segment node feature sequence, where Represents a randomly initialized value; The node feature sequence in each lane segment Generates a tensor of key / value pairs , for a lane segment feature: in , , They are respectively the learnable parameter matrices of the linear mapping, shared in all interactive computations for extracting lane segment local features from lane segment nodes; , are matrices consisting of key vectors and value vectors respectively; The attention score is calculated using the scaled dot product model. The vector sequence output by Self-Attention is shown in the formula: For all calculated lane segment features After residual connection and layer normalization, the local feature sequence of the lane segment is output .

5. The map feature extraction method based on the attention mechanism according to claim 1, characterized in that: In step 4, the Self-Attention mechanism and residual connection are used to realize the lane segment global interaction graph, extract the lane segment global interaction features, and then perform local feature sequence analysis on the lane segment. Perform Self-Attention calculation to obtain the global interactive feature sequence of lane segments ,include: The local features of each input lane segment are transformed using three parameter matrices Mapped to three different spaces, we get the query vector , the key vector Sum value vector , for the entire input sequence The linear mapping process is: in , , They are the learnable parameter matrices of the linear mapping, shared in computing the interaction of each lane segment with other lane segments; , and are matrices consisting of query vector, key vector and value vector respectively; The attention score is calculated using the scaled dot product, and the vector sequence output by Self-Attention is as shown in the formula: in Represents the lane segment interaction feature sequence after Self-Attention calculation; Then the residual connection and layer normalization output lane segment global feature sequence final interaction feature .

6. The map feature extraction method based on the attention mechanism according to claim 1 is characterized in that: In step 5, the global map contains nodes, then for all lane segment node feature sequences Can form a global map node feature sequence , its shape ,in ; In order to learn the relationship between the node and the remote lane segment, the Cross-Attention mechanism is used to learn the relationship between the lane segment Node features and lane segment global feature sequence Interactive calculations are performed, where each node feature Generate query vector , lane segment feature sequence Generates a tensor of key / value pairs ; For the entire map node feature sequence , the calculation relationship between node and lane follows the formula: in , , They are the learnable parameter matrices of the linear mapping, which are shared in the interactive calculation of all lane segment node features and the global lane segment features; , , are matrices consisting of query vector, key vector and value vector respectively; The attention score is calculated using the scaled dot product model. The vector sequence output by Self-Attention is shown in the formula: in is the result of the calculation. After residual connection and layer normalization, the node feature sequence after node-lane interaction can be obtained. .