Vectorized map element fusion method and system for public source mapping
The vectorized map elements of mass-source map building are integrated end-to-end through deep learning models, solving the limitations of the existing technology of relying on graph optimization strategies and achieving a more efficient and reliable map fusion process.
Patent Information
- Application Number
- CN202510209829.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
The existing multi-track fusion method of multi-source map construction depends on the graph optimization strategy, resulting in complex constraint relationships, high dependence on observed data, and manual design constraints introducing subjectivity and uncertainty, affecting the accuracy and reliability of the map.
Using a deep learning model, we use multiple vectorized online map construction results to extract coordinates and semantic type information of map elements, generate feature representations, and encode them using TransFormer encoder to predict map elements after multiple fusions, calculate the total loss function to optimize model parameters, and realize end-to-end fusion of vectorized map elements.
It has escaped the limitations of the graph optimization strategy, avoided the quality risks of map construction brought by human factors, reduced the consumption of storage space and computing resources, and improved the production efficiency of crowd-source maps.
Smart Images

Figure CN120147151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method and system for fusing vectorized map elements in crowdsourcing mapping. Background Art
[0002] As one of the development trends of "light maps" for intelligent driving, crowdsourcing mapping collects the results of multiple online perception-based mapping on the vehicle side, fuses the results of multiple trips in the cloud, and finally generates a map for use in vehicle-side intelligent driving.
[0003] Most of the existing multi-trip fusion methods for crowdsourcing mapping use graph optimization strategies in SLAM. The poses of map elements are represented as vertices of a graph, and the associations between vertices are represented as edges of the graph. By optimizing the constraint relationships carried by the edges of the graph, the optimal estimated values of the vertices are solved. In actual use, the graph optimization method faces many complex constraint relationships and has a high dependence on the accuracy and integrity of the observation data. In addition, manually designing these constraint relationships inevitably introduces subjectivity and uncertainty, which may lead to inconsistencies and instabilities in the final optimization results, thus affecting the accuracy and reliability of the map.
[0004] Currently, a few multi-trip result fusion methods based on deep learning models have emerged in the field of crowdsourcing mapping. These methods first convert vectorized map elements into a raster form, and then use a deep neural network to process the rasterized map elements from the perspective of image processing. Although these methods have certain innovation compared with the graph optimization strategy, the rasterized map elements require a large amount of storage space, thus increasing the burden on computing resources. The rasterized map also needs to go through a post-processing step to restore its vector form for use in vehicle-side intelligent driving, which is not only technically complex but also time-consuming and laborious. Summary of the Invention
[0005] In order to overcome the above technical defects, the present invention provides a method and system for fusing vectorized map elements in crowdsourcing mapping.
[0006] To solve the above problems, the present invention is implemented according to the following technical solutions:
[0007] In a first aspect, the present invention provides a method for fusing vectorized map elements in crowdsourcing mapping, including the following steps:
[0008] Read the results of multiple trips of vectorized online mapping, and extract all map element sets in the results of the multiple trips of vectorized online mapping;
[0009] Encode the coordinates and semantic type information of the map elements to generate corresponding map element feature representations;
[0010] Based on the above map element feature representation, predict the map elements after multiple rounds of fusion through a vectorized map element fusion model;
[0011] By calculating the total loss function of the vectorized map element fusion model and optimizing the model parameters of the vectorized map element based on the total loss function, the training of the vectorized map element fusion model is completed;
[0012] By filtering out the representation of the fused map elements with a semantic type prediction probability lower than the confidence threshold, the vectorized map elements after multiple rounds of fusion are obtained.
[0013] Combined with the first aspect, the present invention also provides a first specific implementation manner of the first aspect. Specifically, the step of encoding the coordinates and semantic type information of the map elements to generate corresponding map element feature representations includes:
[0014] Perform sine-cosine encoding on the coordinate information of the map elements to obtain a coordinate feature representation;
[0015] Perform one-hot encoding on the semantic type information of the map elements to obtain a semantic feature representation;
[0016] Fuse the coordinate feature representation and the semantic feature representation through a fully connected layer to obtain a comprehensive feature representation of the map elements;
[0017] Use a TransFormer encoder to encode the comprehensive feature representation of the map elements to obtain the final map element feature representation.
[0018] Combined with the first aspect, the present invention also provides a second specific implementation manner of the first aspect. Specifically, the step of predicting the map elements after multiple rounds of fusion through a vectorized map element fusion model based on the map element feature representation includes:
[0019] Initialize the query vector, position encoding, and reference point coordinates representing the fused map elements;
[0020] Input the query vector representing the initialized fused map elements into a map decoder to obtain a refined query representing the fused map elements;
[0021] Use an aggregation branch and a classification branch to process the refined query vector to predict the aggregation type and semantic type of each fused map element;
[0022] Use a regression branch to process the refined query vector to predict the coordinates of each node of each fused map element.
[0023] Combined with the first aspect, the present invention also provides a third specific implementation manner of the first aspect. Specifically, the step of training the vectorized map element fusion model by calculating the total loss function of the vectorized map element fusion model and optimizing the vectorized map element model parameters based on the total loss function includes:
[0024] Calculate the pairwise matching cost between the ground truth and the prediction. The pairwise matching cost includes but is not limited to semantic type matching cost, coordinate matching cost, or aggregation type matching cost;
[0025] According to the pairwise matching cost, determine the optimal matching relationship between the ground truth and the predicted value. The optimal matching relationship is obtained by solving the maximum matching problem of a bipartite graph;
[0026] Based on the optimal matching relationship, calculate the total loss function of the vectorized map element fusion model;
[0027] Take the total loss function as the optimization objective, and update the vectorized map element fusion model parameters through backpropagation to obtain the optimized vectorized map element fusion model after training.
[0028] Combined with the first aspect, the present invention also provides a fourth specific implementation manner of the first aspect. Specifically, the formula for calculating the matching process between the ground truth and the prediction is
[0029] where, is a vector space composed of all permutations of N u representing the fused map elements, is the ground truth and the θ i th prediction the pairwise matching cost between them; is the permutation of N u representing the fused map elements corresponding to the minimum matching cost, that is, the optimal matching result.
[0030] Combined with the first aspect, the present invention also provides a fifth specific implementation manner of the first aspect. Specifically, the total loss function of the vectorized map element fusion model is
[0031] where, represents the total loss function of the vectorized map element fusion model, represents the classification loss function of the vectorized map element fusion model, represents the regression loss function of the vectorized map element fusion model, represents the aggregation loss function of the vectorized map element fusion model, λ 1, λ 2 and λ 3 represent the balance coefficients of the loss function terms.
[0032] In a second aspect, the present invention also provides a vectorized map element fusion system for crowdsourced map building, including:
[0033] A construction module, which is used to construct a vectorized map element fusion model. The construction module includes a vectorized data reading module, a map embedding encoding module, a fusion element decoding module, a loss function calculation module, and an inference result module;
[0034] The vector data reading module is used to read the results of multiple trips of vectorized online map building;
[0035] Among them, when the vectorized map element fusion model is in the training stage, the vector data reading module is also used to read the ground truth labels of the results of multiple trips of vectorized online map building;
[0036] The map embedding encoding module is used to embed all the information of the map elements;
[0037] The fusion element decoding module is used to predict the map elements after multiple trips of vectorized element fusion;
[0038] The loss function calculation module is used to calculate the total loss function when the vectorized map element fusion model is in the training stage;
[0039] The inference result processing module is used to filter out the representations of the fused map elements with a semantic type prediction probability lower than the confidence threshold when the vectorized map element fusion model is in the testing stage;
[0040] A training module, which, on the premise that the results of multiple trips of vectorized online map building and their ground truth labels are ready, inputs the results of multiple trips of vectorized online map building into the vectorized map element fusion model constructed by the construction module, and calculates the total loss function of the vectorized map element fusion model;
[0041] Among them, the training module is also used to train the vectorized map fusion model and optimize the weight parameters of the vectorized map element fusion model;
[0042] An inference module, which inputs the results of multiple trips of vectorized online map building into the vectorized map element fusion model with weight parameters added, and outputs the semantic type of each fused map element and the coordinates of each node.
[0043] Compared with the prior art, the beneficial effects of the present invention are:
[0044] The vectorized map element fusion method for crowdsourcing mapping in the present invention breaks away from the limitations of the graph optimization strategy commonly used in multi-pass fusion of existing crowdsourcing mapping, constructs the process of a vectorized map element fusion model from the perspective of deep learning, realizes the input of multi-pass vectorized online mapping results, and outputs the fused results of multi-pass vectorized online mapping end-to-end, avoiding the mapping quality risks brought by traditional rule design and the introduction of human factors. In addition, the vectorized map element fusion method for crowdsourcing mapping abandons the rasterization method of vectorized map elements, and always maintains the vectorized form of map elements throughout the entire map element fusion process, greatly reducing the occupation of storage space and the consumption of computer resources, and is conducive to improving the production efficiency of crowdsourcing maps. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The following further elaborates on the specific embodiments of the present invention in conjunction with the drawings, where:
[0046] Figure 1 is a flowchart of a vectorized map element fusion method for crowdsourcing mapping in the present invention.
[0047] Figure 2 is a framework diagram of a vectorized map element fusion system for crowdsourcing mapping in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The following describes the preferred embodiments of the present invention in conjunction with the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0049] Embodiment 1
[0050] Embodiment 1 provides a vectorized map element fusion method for crowdsourcing mapping, as Figure 1 is a framework diagram of a vectorized map element fusion method for crowdsourcing mapping in Embodiment 1 of the present invention.
[0051] The vectorized map element fusion method for crowdsourcing mapping includes the following steps:
[0052] Step S100: Read the results of multi-pass vectorized online mapping, and extract all map element sets in the results of the multi-pass vectorized online mapping;
[0053] Step S200: Encode the coordinate and semantic type information of the map elements to generate corresponding map element feature representations;
[0054] Step S300: Based on the map element feature representations, predict the map elements after multi-pass fusion through a vectorized map element fusion model;
[0055] Step S400: Complete the training of the vectorized map element fusion model by calculating the total loss function of the vectorized map element fusion model and optimizing the vectorized map element model parameters based on the total loss function;
[0056] Step S500: Obtain the vectorized map elements after multiple passes of fusion by filtering out the map elements representing the fused map elements with a semantic type prediction probability lower than the confidence threshold.
[0057] In the vectorized map element fusion method for crowdsourcing mapping in this embodiment, by reading the results of multiple passes of vectorized online mapping, extracting the set of map elements, and encoding the coordinate and semantic type information of the map elements, the corresponding map element feature representations are generated. Based on the map element feature representations, the vectorized map element fusion model predicts the fused map elements, calculates the total loss function during the training phase, filters out the map elements with low confidence during the testing phase, and finally outputs the semantic type and coordinate information of the fused map elements.
[0058] The vectorized map element fusion model is a deep learning model based on Transformer that fuses the results of multiple passes of vectorized online mapping to obtain the results of multiple passes of fusion of map elements belonging to the same aggregation type. The results of multiple passes of vectorized online mapping refer to the results of vectorized online mapping of a single vehicle for multiple passes or multiple vehicles for multiple passes in the same spatial region. Each pass of vectorized online mapping results includes, but is not limited to, map elements of semantic types such as lane dividers, road boundaries, crosswalks, stop lines, and lane centerlines. Each map element includes, but is not limited to, information such as node coordinates, semantic types, and aggregation types. For convenience, the present invention assumes that all coordinates are located in the UTM coordinate system and all coordinates are 3D coordinates. In the 3D setting, the same spatial region is a cube, and the range of the cube is denoted as [(X min , Y min , Z min ), (X max , Y max , Z max )]. Among them, X min , Y min , Z min respectively represent the minimum values of all vertices of the cube on the x-axis, y-axis, and z-axis in the UTM coordinate system, and X max , Y max , Z max respectively represent the maximum values of all vertices of the cube on the x-axis, y-axis, and z-axis in the UTM coordinate system.
[0059] Step S100: Read the results of multiple passes of vectorized online mapping and extract all sets of map elements in the results of the multiple passes of vectorized online mapping.
[0060] In step S100, the vectorized online mapping result is generated by the in-vehicle online mapping model and has completed the continuous vectorized map after multi-frame fusion. The online mapping model is deployed on the vehicle side and uses the real-time data collected by the in-vehicle surround-view camera or in-vehicle lidar as input. The online mapping model generates a single-frame vectorized map of the area around the vehicle based on the input data. The online mapping model used is the MapTR series model. Multi-frame fusion is to fuse the continuous single-trip single-frame vectorized maps of the area around the vehicle to obtain a continuous vectorized map within a section of the vehicle's moving area.
[0061] The vectorized map element fusion model reads the pre-collected and prepared multi-trip vectorized online mapping results, extracts the set of all map elements in the multi-trip vectorized online mapping results, denoted as Among them, represents the i-th map element in the multi-trip vectorized online mapping results, 0 ≤ i ≤ (N v -1), N v represents the total number of map elements in the multi-trip vectorized online mapping results. represents the j-th node coordinate of the i-th map element in the multi-trip vectorized online mapping results, represents the total number of node coordinates of the i-th map element in the multi-trip vectorized online mapping results. t i represents the semantic type of the i-th map element in the multi-trip vectorized online mapping results, t i ∈{0, 1, …, t, …, (T - 1)}, T represents the total number of map element semantic types.
[0062] When the vectorized map element fusion model is in the training stage, the vectorized map element fusion model reads the ground truth labels of the multi-trip vectorized online mapping results. The ground truth labels of the multi-trip vectorized online mapping results include the aggregation type of each map element and the map elements after multi-trip fusion. Denote the set of the aggregation types of all map elements in the multi-trip vectorized online mapping results as c i represents the aggregation type of the i-th map element in the multi-trip vectorized online mapping results, c i ∈{0, 1, …, c, …, (C - 1)}, C represents the total number of map element aggregation types. Each map element after multi-trip fusion includes information such as node coordinates, semantic type, aggregation type, etc. Denote the set of all map elements after multi-trip fusion as Among them, represents the i-th map element after multi-trip fusion, represents the total number of map elements after multi-trip fusion. Denote the j-th node coordinate of the i-th map element after multiple rounds of fusion. Denote the total number of node coordinates of the i-th map element after multiple rounds of fusion. Denote the aggregation type of the i-th map element after multiple rounds of fusion. Denote the aggregation type of the i-th map element after multiple rounds of fusion.
[0063] Step S200: Encode the coordinates and semantic type information of the map element to generate the corresponding map element feature representation.
[0064] In step S200, encoding the coordinates and semantic type information of the map element is a key step in realizing map element fusion. By encoding this information into a feature representation, a structured and semantically rich feature representation can be provided for the subsequent fusion model.
[0065] The step of encoding the coordinates and semantic type information of the map element to generate the corresponding map element feature representation includes:
[0066] Step S201: Perform sine-cosine encoding on the coordinate information of the map element to obtain the coordinate feature representation.
[0067] In step S201, uniform sampling is performed on the multi-segment lines formed by connecting the node coordinates of each map element in the multi-round vectorized online mapping result to obtain a map element with a fixed total number of node coordinates of N p To enhance the sensitivity of the subsequent fusion element decoding module to the node position differences, sine-cosine encoding is used to transform each node coordinate of each map element, and the transformation of each node coordinate of each map element satisfies the following formula: Where, Denote The value at the 2k position after passing through the sine function sin, Denote The value at the (2k + 1) position after passing through the cosine function cos, k = 0, 1, …, (d / 2 - 1), d denotes The dimension after sine-cosine encoding, generally 256, τ denotes the temperature coefficient, which can take values such as 1000, 10000, etc. Denote the coordinate value of the map element node coordinate normalized to the range [0, 2π) relative to the same spatial region, and the normalization process satisfies the following relational expression:
[0068] Step S202: Perform one-hot encoding on the semantic type information of the map element to obtain the semantic feature representation.
[0069] In step S202, the semantic type information of each map element in the multi-pass vectorized online mapping result is one-hot encoded. For example, given T semantic types of map elements, a zero vector with a dimension of T is initialized. If the map element belongs to the t-th semantic type, the value at the c-th position of the T-dimensional zero vector is set to 1, and the values at other positions remain 0.
[0070] Step S203: Feature fusion is performed on the coordinate feature representation and the semantic feature representation through a fully connected layer to obtain a comprehensive feature representation of the map element.
[0071] In step S203, the results after sine-cosine encoding and one-hot encoding of each map element are concatenated to obtain a first feature representation of each map element. Denote the first feature representations of all map elements as The dimension of the first feature representation of all map elements is denoted as N v ×(N p ·3d + C). The first feature representations V of all map elements Ⅰ are input into a fully connected layer with D input neurons to obtain the comprehensive feature representations of all map elements The dimension of the comprehensive feature representation of the map element is denoted as N v ×D, and D is generally 256. The dimension of the first feature representation of the map element is usually large and not a power of 2. Through the fully connected layer, it can be converted into a comprehensive feature representation with a fixed dimension.
[0072] Step S204: Use the TransFormer encoder to encode the comprehensive feature representation of the map element to obtain the final feature representation of the map element.
[0073] To enhance the abstract representation ability of map elements in the high-dimensional feature space, the TransFormer encoder is used to extract features from the comprehensive feature representation of map elements. The TransFormer encoder consists of 6 encoder layers, and each layer is stacked by a self-attention layer, a linear layer, a dropout layer, and a normalization layer. The interaction process formula of the self-attention layer in each encoder layer:
[0074] Among them, respectively represent the linear mapping matrices corresponding to the query, key, and value in the self-attention mechanism. ⊙ represents the Hadamard product, ┬ represents the transpose of the matrix, and Softmax represents the Softmax normalization function. represents the input object of a certain self-attention layer. For the first encoder layer in the TransFormer encoder, V inThe fused feature representation V of all map elements Ⅱ . Represents the output object of a self-attention layer. For the 6th encoder layer in the TransFormer encoder, V out Is the embedding of all map elements Represents the self-attention mask matrix of the encoder layer.
[0075] When the vectorized map element fusion model is in the test phase, all elements in M are filled with 1;
[0076] When the vectorized map element fusion model is in the training phase, the filling condition of the element at the m-th row and n-th column position in M satisfies the following expression:
[0077] Where, Γ represents the set of all map element aggregation types extracted from multiple trips of vectorized online mapping results The operation of extracting the aggregation type from. Γ(v m ) represents the aggregation type of the m-th map element in the multiple trips of vectorized online mapping results, and Γ(v n ) represents the aggregation type of the n-th map element in the multiple trips of vectorized online mapping results. If the aggregation types of two map elements are the same, then the element at the corresponding position of the encoder layer self-attention mask matrix is 1, otherwise it is 0.
[0078] Step S300: Based on the map element feature representation, predict the map elements after multiple trips of fusion through the vectorized map element fusion model.
[0079] The step of predicting the map elements after multiple trips of fusion through the vectorized map element fusion model based on the map element feature representation includes:
[0080] Step S301: Initialize the query vector, position encoding, and reference point coordinates representing the fused map elements;
[0081] In step S301, randomly initialize the query and its position encoding representing the fused map elements, denoted as And Both are learnable, and their values can be iteratively updated in the map decoder. Where Indicates that the dimensions of U and U' are both N u ×D. N uTo represent the number of queries for the fused map elements, it generally needs to cover the number of map elements after multiple rounds of fusion, such as taking 200, 900, etc. Those skilled in the art should understand that the queries for the fused map elements in this embodiment are instance-level queries, that is, one query corresponds to one map element; in addition, the queries can also be designed as point-level queries, that is, one query corresponds to one point of one map element; or designed as hierarchical queries, that is, one query combines the information of instance-level queries and point-level queries.
[0082] Input the position encoding U' of the query for the fused map elements into a fully connected layer with 3 neurons, and after passing through the step Sigmoid activation function, obtain N u 3D coordinates. Then, for each query representing the fused map elements, copy the N p times of 3D coordinates to obtain the initial reference point coordinates of each node of the fused map elements, denoted as
[0083] Step S302: Input the query vector of the initialized fused map elements into the map decoder to obtain the refined query of the fused map elements.
[0084] In step S302, the map decoder includes 6 decoder layers, and each layer is stacked by a self-attention layer, a cross-attention layer, a linear layer, a dropout layer, and a normalization layer. The interaction process of the self-attention layer in each decoder layer is the same as the self-attention mechanism of the TransFormer encoder. The interaction process of the cross-attention layer in each decoder layer is as shown in the formula:
[0085] Among them, respectively represent the linear mapping matrices corresponding to the query, key, and value in the cross-attention mechanism. denotes element-wise addition of matrices, denotes the transpose of the matrix, and Softmax represents the Softmax normalization function. represents the input object of a certain cross-attention layer. For the first decoder layer in the map decoder, U in is the sum of the randomly initialized query U of the fused map elements and its position encoding U', that is represents the output object of a certain cross-attention layer. For the l-th decoder layer in the map decoder, U out is the refined query of the fused map elements in the l-th layer, denoted as is the embedding of all map elements. represents the cross-attention weighted matrix of the decoder layer. U in and V respectively pass through the aggregation branches corresponding to their respective decoder layers to obtain Uin High-dimensional embedding and the high-dimensional embedding of V By calculating with matrix product, and passing through the Sigmoid activation function, M' is obtained, as shown in the formula:
[0086] where σ represents the Sigmoid activation function, and ⊤ represents the transpose of the matrix.
[0087] Step S303: Use the aggregation branch and the classification branch to process the refined query vector, and predict the aggregation type and semantic type of each map element representing the fusion;
[0088] All map element embeddings V output by the TransFormer encoder pass through the aggregation branch and are further mapped to a new feature space to obtain the high-dimensional embedding of map elements By calculating the matrix product of the high-dimensional embedding of map elements with itself, and passing through the Sigmoid activation function, the similarity score matrix of each map element embedding with other map element embeddings in the new feature space is obtained. The specific calculation formula is as follows:
[0089] where σ represents the Sigmoid activation function, ⊤ represents the transpose of the matrix, represents the similarity score matrix of the map element embedding itself.
[0090] Each decoder layer outputs a refined query representing the map element after fusion After passing through the classification branch, the predicted semantic type corresponding to each node of the map element representing the fusion is obtained, denoted as
[0091] Step S304: Use the regression branch to process the refined query vector, and predict the coordinates of each node of each map element representing the fusion.
[0092] Each decoder layer outputs a refined query representing the map element after fusion It will also pass through the regression branch to predict the offset of the reference point coordinates corresponding to each node of the map element representing the fusion for each decoder layer, denoted as The update strategy of the reference point coordinates corresponding to each node of the map element representing the fusion between adjacent decoder layers is shown in the formula:
[0093] where σ represents the Sigmoid activation function, and σ -1 represents the inverse Sigmoid activation function. It is the reference point coordinates of each node of the fused map element for the l-th decoder layer representation.
[0094] Scale R l The process of scaling to the UTM coordinate system is as shown in the formula:
[0095] where, are respectively the x coordinate, y coordinate, and z coordinate of the reference point corresponding to each node of the fused map element for the l-th decoder layer representation. are respectively the x coordinate, y coordinate, and z coordinate of each node of the fused map element for the l-th decoder layer representation in the UTM coordinate system.
[0096] In a specific implementation, the aggregation branch is a combination of a fully connected layer with D neurons and a ReLU activation function, the classification branch is a combination of a fully connected layer with T neurons and a ReLU activation function, and the regression branch is a combination of a fully connected layer with 3N p neurons and a ReLU activation function. A corresponding aggregation branch, classification branch, and regression branch are provided in each decoder layer of the map decoder.
[0097] Step S400: By calculating the total loss function of the vectorized map element fusion model and optimizing the model parameters of the vectorized map element based on the total loss function, the training of the vectorized map element fusion model is completed.
[0098] The step of calculating the total loss function of the vectorized map element fusion model and optimizing the model parameters of the vectorized map element based on the total loss function to complete the training of the vectorized map element fusion model includes:
[0099] Step S401: Calculate the pairwise matching cost between the ground truth and the prediction. The pairwise matching cost includes but is not limited to semantic type matching cost, coordinate matching cost, or aggregation type matching cost;
[0100] Before formally calculating the total loss function, it is necessary to perform positive and negative sample matching on the prediction result and the ground truth. The matching process formula:
[0101] where, is a vector space composed of all permutations of N u representing the fused map elements, is the ground truth and the θ i th prediction The pairwise matching cost between them. In this embodiment, the ground truth is actually the i-th map element after multiple passes of fusion, that is, Prediction from the semantic types corresponding to all nodes representing the fused map elements and the x - coordinate in the UTM coordinate system y - coordinate z - coordinate
[0102] Therefore, the pairwise matching cost should further include the matching cost between the ground - truth and the predicted semantic types, and the matching cost between the ground - truth and the predicted coordinates. The matching - cost formula is as follows:
[0103] where, represents the semantic type of the ground - truth and the i θ - th predicted semantic type, as the focal loss. represents the coordinates of the ground - truth and the i θ - th predicted coordinates, as the L1 loss.
[0104] As those skilled in the art should understand, the pairwise matching cost in addition to the semantic - type matching cost and the coordinate matching cost, may also have other matching - cost terms that are beneficial to constraining the deviation between the prediction and the ground - truth. In this embodiment, the semantic type of the ground - truth is actually the semantic type of the i - th map element after multiple - pass fusion the coordinates of the ground - truth are actually the coordinates of all nodes of the i - th map element after multiple - pass fusion the predicted semantic type comes from the semantic types corresponding to all nodes representing the fused map elements the predicted coordinates come from the x - coordinates corresponding to all nodes representing the fused map elements in the UTM coordinate system y - coordinates z - coordinates
[0105] Step S402: Determine the optimal matching relationship between the ground - truth and the predicted value according to the pairwise matching cost. The optimal matching relationship is obtained by solving the maximum - matching problem of a bipartite graph;
[0106] In step S402, in the bipartite graph, for each pair of nodes between the ground - truth and the predicted value, the permutation of N u representing the fused map elements corresponding to the minimum matching cost, that is, the optimal matching result, can be specifically solved by the Hungarian algorithm.
[0107] Step S403: Calculate the total loss function of the vectorized map element fusion model based on the optimal matching relationship;
[0108] After obtaining the positive and negative sample matching results of the prediction result and the ground truth, the final total loss function can be calculated to supervise the vectorized map element fusion model. The formula for the total loss function of the vectorized map element fusion model
[0109] where, represents the total loss function of the vectorized map element fusion model, represents the classification loss function of the vectorized map element fusion model, represents the regression loss function of the vectorized map element fusion model, represents the aggregation loss function of the vectorized map element fusion model, λ 1 、λ 2 and λ 3 represent the balance coefficients of the loss function terms, and can take values such as 0.1, 0.01, etc.
[0110] As those skilled in the art should understand, the total loss function of the vectorized map element fusion model In addition to the classification loss function, regression loss function, and aggregation loss function, there can also be other loss function terms that are beneficial to reducing the deviation between the prediction and the ground truth. Specifically, The formula for The formula for and The formula for is:
[0111] where, comes from the semantic types corresponding to all nodes of the fused map element under the optimal matching result comes from the x coordinate of all nodes of the fused map element corresponding to the UTM coordinate system under the optimal matching result y coordinate z coordinate represents the loss function for supervising the accuracy of the self-similarity score of the map element embedding, which can be cross-entropy loss or focal loss. represents the self-attention mask matrix of the encoder layer. Diag -1 represents the operation of extracting non-diagonal position elements from the matrix.
[0112] Step S404: Using the total loss function as the optimization objective, update the parameters of the vectorized map element fusion model through backpropagation, thereby obtaining the trained optimized vectorized map element fusion model.
[0113] In step S404, in order to obtain the optimized vectorized map element fusion model, the total loss function is used as the optimization objective, and the gradient of the total loss function with respect to the parameters of the vectorized map element fusion model is calculated through backpropagation. The core of this process is to use the chain rule to calculate the gradient layer by layer, thereby updating the parameters of the vectorized map element fusion model. First, calculate the gradient of the total loss function with respect to the parameters of each layer layer by layer starting from the output layer. Then, use the SGD optimizer to update the model according to the calculated gradient. The SGD optimizer is used to adjust the step size and direction of parameter update to accelerate convergence and avoid falling into the local optimal stage. Finally, repeat the process of calculating the loss, backpropagation, and parameter update. The vectorized map element fusion model is gradually optimized until the total loss function converges to the minimum value or reaches the preset number of iterations.
[0114] Step S500: Obtain the vectorized map elements after multiple passes of fusion by filtering out the representation of the fused map elements whose semantic type prediction probability is lower than the confidence threshold.
[0115] According to the predefined confidence threshold (such as taking 0.5), retain the representation of the fused map elements whose semantic type prediction probability is greater than the confidence threshold. Each representation of the fused map element includes the semantic type and the coordinates of each node in the UTM coordinate system. In the test phase of the vectorized map element fusion model, the vectorized map element fusion model will output the semantic type and its prediction probability of each fused map element, as well as the coordinates of each node. To ensure the high reliability of the output map elements, these elements need to be filtered by the confidence threshold: first, set a predefined confidence threshold, such as 0.5. Then, for each fused map element, check whether its semantic type prediction probability is lower than the confidence threshold; if it is lower than the threshold, filter out the map element. Finally, retain the map elements whose semantic type prediction probability is greater than or equal to the confidence threshold; these retained map elements include their semantic types and the coordinates of each node in the UTM coordinate system.
[0116] The vectorized map element fusion model outputs the semantic type of each representation of the fused map element and the coordinates of each node. These coordinates are represented in the UTM coordinate system, ensuring the precise position of the map elements. The output results are usually represented in a structured form, for example: Semantic type: lane line, road boundary, crosswalk, etc. Node coordinates: The position of each node is represented by the (x, y) coordinate pair in the UTM coordinate system.
[0117] Embodiment 2
[0118] A vectorized map element fusion system for crowdsourcing mapping, as Figure 2 shown, the vectorized map element fusion system for crowdsourcing mapping includes:
[0119] A construction module, which is used to construct a vectorized map element fusion model. The construction module includes a vectorized data reading module, a map embedding encoding module, a fusion element decoding module, a loss function calculation module, and an inference result module;
[0120] The vector data reading module is used to read the results of multiple trips of vectorized online mapping;
[0121] Among them, when the vectorized map element fusion model is in the training stage, the vector data reading module is also used to read the ground truth labels of the results of multiple trips of vectorized online mapping;
[0122] The map embedding encoding module is used to embed all the information of the map elements;
[0123] The map embedding encoding module first uses sine-cosine encoding and one-hot encoding to obtain the first feature. The first feature is fused through a fully connected layer to obtain a comprehensive feature. Finally, the TransFormer encoder is used to encode the comprehensive feature representation to generate the final map element feature representation.
[0124] The fusion element decoding module is used to predict the map elements after multiple trips of vectorized element fusion;
[0125] The fusion element decoding module first initializes the query, position encoding, and reference point coordinates representing the fused map elements. Then, through the map decoder, it obtains the refined query representing the fused map elements. Finally, the aggregation branch and classification branch are used to predict the aggregation type and semantic type of each fused map element, and the regression branch is used to predict the coordinates of each node of each fused map element.
[0126] The loss function calculation module is used to calculate the total loss function when the vectorized map element fusion model is in the training stage;
[0127] The loss function calculation module first calculates the pairwise matching cost between the ground truth and the prediction. After obtaining the positive and negative sample matching results of the prediction and the ground truth, it calculates the total loss function of the vectorized map element fusion model.
[0128] The inference result processing module is used to filter out the fused map elements whose predicted probability of semantic type is lower than the confidence threshold when the vectorized map element fusion model is in the test stage;
[0129] A training module, on the premise that multiple trips of vectorized online mapping results and their ground truth labels are ready, inputs the multiple trips of vectorized online mapping results into the vectorized map element fusion model constructed by the construction module, and calculates the total loss function of the vectorized map element fusion model;
[0130] Wherein, the training module is further configured to train the vectorized map fusion model and optimize to obtain the weight parameters of the vectorized map element fusion model;
[0131] An inference module, which inputs multiple trips of vectorized online mapping results into the vectorized map element fusion model with weight parameters added, and outputs the semantic type of each fused map element and the coordinates of each node.
[0132] As mentioned above, it is only a preferred embodiment of the present invention, and there is no limitation in any form to the present invention. Therefore, any modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for fusion of vectorized map elements for multi-source mapping, characterized in that: The following steps are involved: Reading multiple vectorization online mapping results, and extracting all map element sets in the multiple vectorization online mapping results; Encoding the coordinates and semantic type information of the map element to generate a corresponding map element feature representation; Based on the map element feature representation, predicting the map elements after multiple fusions through a vectorized map element fusion model; The training of the vectorized map element fusion model is completed by calculating the total loss function of the vectorized map element fusion model and optimizing the parameters of the vectorized map element model based on the total loss function; By filtering out the fused map elements whose semantic type prediction probability is lower than the confidence threshold, vectorized map elements after multiple fusions are obtained.
2. The method for fusion of vectorized map elements for multi-source mapping according to claim 1, characterized in that: The step of encoding the coordinates and semantic type information of the map element to generate a corresponding map element feature representation comprises: Performing sine and cosine encoding on the coordinate information of the map element to obtain a coordinate feature representation; Performing one-hot encoding on the semantic type information of the map element to obtain a semantic feature representation; The coordinate feature representation and the semantic feature representation are subjected to feature fusion through a fully connected layer to obtain a comprehensive feature representation of the map element; The comprehensive feature representation of the map element is encoded using a TransFormer encoder to obtain a final map element feature representation.
3. The method for fusion of vectorized map elements for multi-source mapping according to claim 1 is characterized in that: The step of predicting the map elements after multiple fusions through a vectorized map element fusion model based on the map element feature representation comprises: Initialize the query vector, position encoding and reference point coordinates representing the fused map elements; Inputting the query vector representing the fused map element after the initialization into the map decoder to obtain a refined query representing the fused map element; Processing the refined query vector using an aggregation branch and a classification branch to predict an aggregation type and a semantic type of each fused map element; The refined query vector is processed using a regression branch to predict the coordinates of each node representing a fused map element.
4. The method for fusion of vectorized map elements for multi-source mapping according to claim 1 is characterized in that: The step of calculating the total loss function of the vectorized map element fusion model and optimizing the vectorized map element model parameters based on the total loss function to complete the training of the vectorized map element fusion model includes: Calculate the pairwise matching cost between the true value and the prediction, wherein the pairwise matching cost includes but is not limited to a semantic type matching cost, a coordinate matching cost or an aggregate type matching cost; Determine an optimal matching relationship between a true value and a predicted value according to the pairwise matching cost, wherein the optimal matching relationship is obtained by solving a maximum matching problem of a bipartite graph; Based on the optimal matching relationship, calculating the total loss function of the vectorized map element fusion model; The total loss function is used as an optimization target, and the vectorized map element fusion model parameters are updated through back propagation, so as to obtain a trained optimized vectorized map element fusion model.
5. The method for fusion of vectorized map elements for multi-source mapping according to claim 4 is characterized in that: The formula for calculating the matching process between the true value and the prediction is: in, Because N u A vector space representing all permutations of fused map elements, True value With the θ i Predictions The pairwise matching cost between ; To minimize the matching cost, N u represents the arrangement of the fused map elements, i.e. the optimal matching result.
6. The method for fusion of vectorized map elements for multi-source mapping according to claim 4 is characterized in that: The total loss function of the vectorized map element fusion model is: in, represents the total loss function of the vectorized map element fusion model, represents the classification loss function of the vectorized map element fusion model, represents the regression loss function of the vectorized map element fusion model, represents the aggregation loss function of the vectorized map element fusion model, and λ1, λ2, and λ3 represent the balance coefficients of the loss function terms.
7. A vectorized map element fusion system for multi-source mapping, implemented based on the vectorized map element fusion method for multi-source mapping according to any one of claims 1 to 6, characterized in that: include: A construction module, wherein the construction module is used to construct a vectorized map element fusion model, and the construction module includes a vectorized data reading module, a map embedding encoding module, a fusion element decoding module, a loss function calculation module, and an inference result module; The vector data reading module is used to read multiple vectorization online mapping results; Wherein, when the vectorized map element fusion model is in the training stage, the vector data reading module is also used to read the true value labels of multiple vectorized online mapping results; The map embedding coding module is used to embed all map element information; The fused element decoding module is used to predict the map elements after multiple passes of vectorized element fusion; The loss function calculation module is used for calculating the total loss function of the vectorized map element fusion model during the training phase; The inference result processing module is used for filtering out the fused map elements whose semantic type prediction probability is lower than the confidence threshold when the vectorized map element fusion model is in the testing phase; A training module, which, on the premise that the multi-pass vectorization online mapping results and their true value labels are ready, inputs the multi-pass vectorization online mapping results into the vectorization map element fusion model constructed by the construction module, and calculates the total loss function of the vectorization map element fusion model; The training module is also used to train the vectorized map fusion model and optimize the weight parameters of the vectorized map element fusion model; The inference module inputs the vectorized map element fusion model with weight parameters into the vectorized map element fusion model, and outputs the semantic type of each fused map element and the coordinates of each node.