Vector map generation method and device, equipment, storage medium and program product
By introducing the Hungarian matching algorithm in the initial model training stage to generate a prediction model, the problems of error accumulation and information loss in large-scale vector map data processing are solved, and the generation of high-quality vector maps and a simplified process are achieved.
Patent Information
- Application Number
- CN202510781789.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies are prone to error accumulation and information loss when processing large-scale vector map data, resulting in poor quality of the generated vector maps.
The initial model is used to process the vector map data to be trained, and the model is trained in combination with the Hungarian matching algorithm to generate a prediction model. The vector map data is then directly processed based on the model to generate a high-quality vector map.
The robustness and generalization ability of the prediction model are improved, the vector map generation process is simplified, and the generated vector maps are of higher quality without the need for complex pre-processing and post-processing steps.
Smart Images

Figure CN120702447A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of map generation technology, and in particular to a vector map generation method, apparatus, device, storage medium, and program product. Background Art
[0002] With the rapid development of digital technology, maps have been widely used in many fields such as navigation, geographic information systems, and intelligent transportation. Among them, vector maps, with their precise geometric shapes and rich attribute information, provide users with an efficient map representation method.
[0003] In the prior art, key features such as road centerlines and building outlines are extracted from multi-source vector map data, and then matching and fusion are performed based on these key features to generate a corresponding vector map.
[0004] However, in the above-mentioned method, when processing large-scale vector map data, problems of error accumulation and information loss are prone to occur, resulting in poor quality of the generated vector map. Summary of the Invention
[0005] The embodiments of the present application provide a vector map generation method, apparatus, device, storage medium, and program product, which can improve the quality of vector maps and simplify the vector map generation process.
[0006] In a first aspect, an embodiment of the present application provides a vector map generation method, comprising:
[0007] Acquire vector map data to be trained; and determine, based on the initial model, a predicted vector result for each vector line to be trained in the vector map data to be trained;
[0008] Based on the Hungarian matching algorithm, the initial model is trained according to the predicted vector result of each vector line to be trained to obtain a prediction model;
[0009] The vector map data to be processed is processed based on the prediction model to generate a vector map corresponding to the vector map data to be processed.
[0010] In a possible implementation, the Hungarian matching algorithm is used to train the initial model according to the predicted vector result of each vector line to be trained to obtain a prediction model, including:
[0011] Based on the Hungarian matching algorithm, the predicted vector results of each of the vector lines to be trained and the real vector results of each of the vector lines to be trained are processed to obtain a matching index; wherein the matching index includes the predicted vector result of each of the vector lines to be trained and the real vector result that matches the predicted vector result;
[0012] Determining a multi-task loss function corresponding to the matching index based on the predicted vector result of each to-be-trained vector line and the corresponding matched real vector result in the matching index; wherein the multi-task loss function represents the difference between the predicted vector result of each to-be-trained vector line and the corresponding matched real vector result;
[0013] The initial model is trained according to the multi-task loss function to obtain the prediction model.
[0014] In a possible implementation, the processing of the predicted vector result of each to-be-trained vector line and the real vector result of each to-be-trained vector line based on the Hungarian matching algorithm to obtain a matching index includes:
[0015] Determining a cost matrix corresponding to each of the vector lines to be trained based on the predicted vector results of each of the vector lines to be trained and the actual vector results of each of the vector lines to be trained; wherein the cost matrix represents a matching cost between each predicted vector result of each of the vector lines to be trained and the actual vector results of each of the vector lines to be trained;
[0016] Based on the Hungarian matching algorithm, the cost matrix is calculated to obtain the matching index.
[0017] In a possible implementation, processing the predicted vector result of each to-be-trained vector line in the matching index and the corresponding matched real vector result to determine the multi-task loss function corresponding to the matching index includes:
[0018] Determining a classification loss corresponding to the matching index based on predicted category information in the predicted vector result of each to-be-trained vector line in the matching index and true category information in the corresponding matched true vector result; wherein the classification loss represents the difference between the predicted category distribution and the true category distribution of all to-be-trained vector lines;
[0019] Determine, based on the predicted coordinate information in the predicted vector result of each vector line to be trained in the matching index and the actual coordinate information in the corresponding matched actual vector result, a vector line loss and a direction loss corresponding to the matching index; wherein the vector line loss represents the difference between the predicted geometric shapes and the actual geometric shapes of all vector lines to be trained; and the direction loss represents the difference between the predicted directions and the actual directions of all vector lines;
[0020] Determine a multi-task loss function corresponding to the matching index according to the classification loss, vector line loss, and direction loss corresponding to the matching index.
[0021] In a possible implementation, determining a predicted vector result of each vector line to be trained in the vector map data to be trained based on the initial model includes:
[0022] Based on the embedding layer of the initial model, encoding the vector line to be trained is performed to obtain an initial embedding representation of the vector line to be trained;
[0023] Based on the encoder and decoder of the initial model, the initial embedded representation of the vector line to be trained is processed to obtain vector line features of the vector line to be trained; wherein the vector line features represent the geometric shape and semantic structure of the vector line to be trained;
[0024] Based on the output layer of the initial model, the vector line features of the vector line to be trained are processed to obtain a predicted vector result of the vector line to be trained.
[0025] In a possible implementation, encoding the vector line to be trained based on the embedding layer of the initial model to obtain an initial embedding representation of the vector line to be trained includes:
[0026] Processing the coordinate information of the vector line to be trained based on the multi-layer perceptron of the embedding layer to obtain a coordinate embedding representation of the vector line to be trained;
[0027] Based on the category embedding layer of the embedding layer, the category information of the vector line to be trained is processed to obtain a category embedding representation of the vector line to be trained;
[0028] Processing the category embedding representation of the vector line to be trained based on the conditional adjustment layer of the embedding layer to obtain a conditional embedding representation of the vector line to be trained;
[0029] Based on the embedding layer, the coordinate embedding representation and the conditional embedding representation of the vector line to be trained are combined to obtain a combined embedding representation of the vector line to be trained;
[0030] Based on the embedding layer, position encoding processing is performed on the attribute information of the vector line to be trained to obtain the attribute embedding representation of the vector line to be trained; wherein the combined embedding representation and the attribute embedding representation constitute the initial embedding representation.
[0031] In a possible implementation, the encoder and decoder based on the initial model process the initial embedding representation of the vector line to be trained to obtain the vector line features of the vector line to be trained, including:
[0032] Based on the multi-head self-attention mechanism of the encoder, the initial embedded representation of each of the vector lines to be trained is processed to obtain a set of attention features corresponding to each of the vector lines to be trained; wherein the set of attention features includes at least one self-attention feature; each of the attention features represents a global dependency relationship between each of the vector lines to be trained;
[0033] Based on the multi-head attention mechanism of the decoder, the attention feature set is processed to obtain the vector line features of the vector line to be trained.
[0034] In one possible implementation, the method further includes:
[0035] Normalization processing is performed on all vector points in the vector line to be trained, and at the same time, resampling processing is performed on the vector line to be trained to obtain a processed vector line to be trained.
[0036] In a second aspect, an embodiment of the present application provides a vector map generation device, comprising:
[0037] A processing module, configured to obtain vector map data to be trained; and determine, based on an initial model, a predicted vector result for each vector line to be trained in the vector map data to be trained;
[0038] A training module, configured to train the initial model based on the Hungarian matching algorithm and according to the predicted vector result of each vector line to be trained, to obtain a prediction model;
[0039] The generation module is used to process the vector map data to be processed based on the prediction model, and generate a vector map corresponding to the vector map data to be processed.
[0040] In one possible embodiment, the training module is specifically used to: based on the Hungarian matching algorithm, process the predicted vector results of each vector line to be trained and the real vector results of each vector line to be trained to obtain a matching index; wherein the matching index includes the predicted vector results of each vector line to be trained and the real vector results that match the predicted vector results; according to the predicted vector results of each vector line to be trained in the matching index and the corresponding matched real vector results, determine the multi-task loss function corresponding to the matching index; wherein the multi-task loss function represents the gap between the predicted vector results of each vector line to be trained and the corresponding matched real vector results; according to the multi-task loss function, train the initial model to obtain the prediction model.
[0041] In one possible implementation, the training module is specifically configured to: determine a cost matrix corresponding to each vector line to be trained based on a predicted vector result of each vector line to be trained and a true vector result of each vector line to be trained; wherein the cost matrix represents the matching cost of each predicted vector result of each vector line to be trained and the true vector result of each vector line to be trained; and calculate and process the cost matrix based on the Hungarian matching algorithm to obtain the matching index.
[0042] In a possible embodiment, the training module is further specifically used to: determine the classification loss corresponding to the matching index based on the predicted category information in the predicted vector result of each vector line to be trained in the matching index and the real category information in the corresponding matched real vector result; wherein the classification loss represents the difference between the predicted category distribution and the real category distribution of all vector lines to be trained; determine the vector line loss and direction loss corresponding to the matching index based on the predicted coordinate information in the predicted vector result of each vector line to be trained in the matching index and the real coordinate information in the corresponding matched real vector result; wherein the vector line loss represents the difference between the predicted geometric shapes and the real geometric shapes of all vector lines to be trained; the direction loss represents the difference between the predicted directions and the real directions of all vector lines; and determine the multi-task loss function corresponding to the matching index based on the classification loss, vector line loss and direction loss corresponding to the matching index.
[0043] In one possible implementation, the processing module is specifically used to: encode the vector line to be trained based on the embedding layer of the initial model to obtain the initial embedding representation of the vector line to be trained; process the initial embedding representation of the vector line to be trained based on the encoder and decoder of the initial model to obtain the vector line features of the vector line to be trained; wherein the vector line features represent the geometric shape and semantic structure of the vector line to be trained; and process the vector line features of the vector line to be trained based on the output layer of the initial model to obtain a predicted vector result of the vector line to be trained.
[0044] In one possible implementation, the processing module is specifically used to: process the coordinate information of the vector line to be trained based on the multi-layer perceptron of the embedding layer to obtain the coordinate embedding representation of the vector line to be trained; process the category information of the vector line to be trained based on the category embedding layer of the embedding layer to obtain the category embedding representation of the vector line to be trained; process the category embedding representation of the vector line to be trained based on the conditional adjustment layer of the embedding layer to obtain the conditional embedding representation of the vector line to be trained; combine the coordinate embedding representation and the conditional embedding representation of the vector line to be trained based on the embedding layer to obtain the combined embedding representation of the vector line to be trained; perform position encoding processing on the attribute information of the vector line to be trained based on the embedding layer to obtain the attribute embedding representation of the vector line to be trained; wherein, the combined embedding representation and the attribute embedding representation constitute the initial embedding representation.
[0045] In one possible embodiment, the processing module is further specifically used to: process the initial embedding representation of each of the vector lines to be trained based on the multi-head self-attention mechanism of the encoder to obtain a set of attention features corresponding to each of the vector lines to be trained; wherein the attention feature set includes at least one self-attention feature; each of the attention features represents the global dependency relationship between each of the vector lines to be trained; based on the multi-head attention mechanism of the decoder, process the attention feature set to obtain the vector line features of the vector lines to be trained.
[0046] In a possible implementation, the apparatus is further configured to: perform normalization processing on all vector points in the vector line to be trained, and simultaneously perform resampling processing on the vector line to be trained to obtain a processed vector line to be trained.
[0047] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor;
[0048] The memory stores computer-executable instructions;
[0049] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.
[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementations of the first aspect.
[0051] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.
[0052] The vector map generation method, apparatus, device, storage medium, and program product provided in the embodiments of the present application process the vector map data to be trained through an initial model to obtain a predicted vector result for each vector line to be trained. By introducing the Hungarian matching algorithm, the initial model is trained based on the predicted vector result for each vector line to be trained to obtain a prediction model. Based on the prediction model, the vector map data to be processed is processed to generate a corresponding vector map. Furthermore, by introducing the Hungarian matching algorithm in the training stage of the initial model to obtain the prediction model, the robustness and generalization ability of the prediction model are improved, and the generation quality of the vector map obtained based on the prediction model is further improved. Moreover, based on the end-to-end fusion generation method, it is possible to directly learn and generate a fused vector map from multi-source vector map data without the need for complex pre-processing and post-processing steps, thereby greatly simplifying the vector map generation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0054] Figure 1 A schematic diagram comparing a high-precision map and a common navigation map provided in this application;
[0055] Figure 2 A flowchart of a vector map generation method provided in an embodiment of the present application;
[0056] Figure 3 A visualization example diagram of the true map data provided in an embodiment of the present application;
[0057] Figure 4 A flowchart of another vector map generation method provided in an embodiment of the present application;
[0058] Figure 5 A model architecture diagram of a Map Transformer provided in an embodiment of the present application;
[0059] Figure 6 A flowchart of another vector map generation method provided in an embodiment of the present application;
[0060] Figure 7 A schematic diagram of an example of a data processing platform provided in an embodiment of the present application;
[0061] Figure 8 A schematic diagram of the structure of a vector map generation device provided in an embodiment of the present application;
[0062] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0063] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0064] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0065] It should be noted that the present application can be used in the field of map generation technology, and can also be used in any field other than map generation technology, and the application field of the present application is not limited.
[0066] Figure 1 A schematic diagram comparing a high-precision map and a common navigation map provided in this application, such as Figure 1 As shown, compared with conventional navigation maps, high-precision maps are for people to see, while high-precision maps are for machines to see. The accuracy of navigation maps has errors of more than meters (5-10 meters), while the accuracy of high-precision maps is at the centimeter level. High-precision maps have richer road elements and traffic-related dynamic elements. Among them, vector maps, with their precise geometric shapes and rich attribute information, provide users with an efficient map representation method.
[0067] Based on the above scenarios, we can see that by formulating a series of data conversion and matching rules, multi-source vector map data is processed and integrated layer by layer to generate vector maps; however, for complex data and large-scale data sets, the cost of formulating and maintaining rules is extremely high, and it is difficult to adapt to the diversity and variability of data, resulting in poor quality of the generated vector maps.
[0068] In another example, key features in vector maps, such as road centerlines and building outlines, are extracted and then matched and fused based on these features. However, when processing large-scale vector map data, this approach has poor adaptability to different data sources and is prone to error accumulation and information loss, resulting in poor quality of the generated vector maps.
[0069] In another example, deep learning technology was applied to the vector map fusion task for raster map design. However, insufficient mining of the special structure and semantic information of the vector map resulted in poor quality of the generated vector map.
[0070] The vector map generation method provided in this application processes the vector map data to be trained through an initial model to obtain a predicted vector result for each vector line to be trained, and combines the Hungarian matching algorithm to train the initial model to obtain a prediction model. Based on the prediction model, the vector map data to be processed is processed to generate a corresponding vector map, thereby solving the technical problem of poor quality of the generated vector map.
[0071] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0072] Figure 2 A flow chart of a vector map generation method provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the method includes:
[0073] 201. Obtain vector map data to be trained; and determine a predicted vector result for each vector line to be trained in the vector map data to be trained based on an initial model.
[0074] For example, the execution subject of this embodiment may be an electronic device, hereinafter referred to as a device. The device may be a virtual device or a physical device that executes the vector map generation method. The device may obtain vector map data to be trained from a multi-source database, and the vector map data to be trained includes a plurality of vector lines to be trained, for example, Figure 3 This is a visualization example diagram of the true map data provided in the embodiment of the present application, such as Figure 3As shown, multi-source vector map data to be trained is collected, including multi-source reported vector data obtained from satellite remote sensing systems, field measurements, etc., and also includes ground-truth map data, that is, the real vector results corresponding to the multi-source reported vector data. These ground-truth map data and multi-source reported vector data are stored in a data file according to certain rules, such as saving the coordinate points and attribute information of the vector lines in txt format; the multi-source vector map data to be trained and the ground-truth map data stored in the data file are preprocessed, including removing duplicate data, correcting obvious errors, unifying data formats, etc., to provide high-quality input data for subsequent model training, so as to obtain data of multiple vector lines to be trained, including lane markings, stop locations, and crosswalks.
[0075] The device calls a preset initial model, such as a model based on the Transformer architecture, a Recurrent Neural Network (RNN), or a Long Short-Term Memory (LSTM), and inputs each vector line to be trained in the vector map data to be trained into the initial model, so that the initial model processes each vector line to be trained in the vector map data to be trained and obtains a predicted vector result for each vector line to be trained, such as the predicted shape, position, lane lines, buildings, and other vector attribute information corresponding to each vector line to be trained. Among them, the Transformer architecture is introduced to process vector map data, and its powerful parallel computing capabilities and modeling capabilities for sequence data are used to effectively capture the complex relationships between different vector elements, improve the accuracy and efficiency of fusion generation, and thus help improve the quality of subsequently generated vector maps.
[0076] 202. Based on the Hungarian matching algorithm, the initial model is trained according to the predicted vector results of each vector line to be trained to obtain a prediction model.
[0077] Exemplarily, the device calls a preset Hungarian matching algorithm, such as the assignment problem algorithm or the Hopcroft-Karp algorithm, to perform matching calculation processing on the predicted vector results of each vector line to be trained, obtain the matching calculation results of each vector line to be trained, and optimize the model parameters of the initial model based on the matching calculation results of each vector line to be trained to obtain a prediction model.
[0078] For example, during the initial model training process, an optimizer (such as the AdamW optimizer) and a learning rate scheduler (such as the cosine annealing learning rate scheduler) are used to optimize the initial model, gradually reduce the loss value, and improve the model performance to obtain a more accurate prediction model.
[0079] For another example, based on the preset Hungarian matching algorithm KM (Kuhn-Munkres Algorithm), the predicted vector results of all the vector lines to be trained are initialized to generate the corresponding initial weight matrix, including at least one row and at least one column. By constructing the initial matching, an initial matching is found for each row, for example, row 1: select column 1 (weight 4), row 2: select column 3 (weight 5), row 3: unmatched, and the initial matching is: (1,1), (2,3); then the current matching is not a perfect matching because row 3 is unmatched. By finding the augmenting path algorithm, the matching is updated and the matching results are obtained: (1,3), (2,3), (3,2). In this way, the Kuhn-Munkres algorithm can effectively adjust the matching and finally achieve the maximum weight matching to obtain the matching calculation results of all the vector lines to be trained. Based on the matching calculation results, the model parameters and hyperparameters of the initial model are optimized through the cosine annealing learning rate scheduler to obtain the prediction model.
[0080] It is worth adding that during the training process, the initial model is regularly evaluated and verified. Using the validation dataset, the initial model is switched to the evaluation mode to calculate the validation loss and related indicators, such as classification accuracy, average vector line distance, etc. The hyperparameters, training strategies, or data preprocessing methods of the initial model are adjusted according to the validation results to further improve the generalization and robustness of the model. At the same time, visualization tools (such as Tensor Board) are used to record the loss curve and learning rate changes during the training and validation process to facilitate the analysis of the training status of the initial model in order to obtain a prediction model. After the training is completed, the parameters and status of the prediction model are saved. The prediction model can be saved at certain intervals (such as every 10 training cycles) for subsequent testing and deployment.
[0081] 203. Process the vector map data to be processed based on the prediction model to generate a vector map corresponding to the vector map data to be processed.
[0082] Exemplarily, the device integrates the trained model, i.e., the prediction model, into an actual application system, such as navigation software, geographic information system, etc., and processes the acquired vector map data to be processed based on the prediction model to generate a vector map corresponding to the vector map data to be processed, so as to realize the end-to-end fusion generation function of the vector map and provide users with high-quality map services.
[0083] In this embodiment, a vector map generation method is provided. By introducing the Hungarian matching algorithm during the training phase of the initial model to obtain a prediction model, the robustness and generalization ability of the prediction model are improved, and the quality of the vector map obtained based on the prediction model is further improved. At the same time, based on the end-to-end fusion generation method, it is possible to directly learn and generate a fused vector map from multi-source vector map data without the need for complex pre-processing and post-processing steps, greatly simplifying the vector map generation process.
[0084] Figure 4 A flow chart of another vector map generation method provided in an embodiment of the present application is shown as follows: Figure 4 As shown, the method includes:
[0085] 301. Obtain vector map data to be trained; perform normalization processing on all vector points in the training vector line, and at the same time, perform resampling processing on the training vector line to obtain a processed vector line to be trained.
[0086] For example, the device can obtain vector map data to be trained from a multi-source database. The vector map data to be trained includes multiple vector lines to be trained, and each vector line to be trained includes multiple vector points. For each vector line to be trained, all vector points in each vector line to be trained are normalized based on a preset normalization algorithm, mapping coordinate values in different ranges to a unified interval, eliminating scale differences in the data and providing a good data foundation for subsequent model training and inference. At the same time, each vector line to be trained is resampled based on a preset resampling algorithm to obtain each processed vector line to be trained. For example, the number of sampling points for each lane line is fixed to a preset value to ensure dimensional consistency of the input data while retaining the main geometric features of the vector line.
[0087] For example, use the custom Vector Map Dataset class to load data. This class can read the vec file and label file from the specified folder and parse out the relevant information of all the vector lines to be trained, such as vector lines of different passes, vector lines with real labels, and the boundary range of the calculated map. During the parsing process, the normalization formula is used:
[0088]
[0089] Normalize the coordinate points of the vector line and map them to the interval [0, 1] to obtain the x-axis normalized coordinate x of each vector point in each vector line to be trained in the preset coordinate system. norm and the y-axis normalized coordinate y norm; Wherein, x and y are the x-axis coordinate and y-axis coordinate of each vector point in each vector line to be trained in the preset coordinate system, respectively. min 、x max They are the minimum and maximum values of the x-axis coordinates of all vector points in each vector line to be trained in the preset coordinate system, and the y min 、y max are the minimum and maximum values of the y-axis coordinates of all vector points in each vector line to be trained in the preset coordinate system. According to the total length L of each vector line to be trained and the target number of sampling points N, the equally spaced step length Δ=L / (N-1) corresponding to each vector line to be trained is determined; according to the equally spaced step length corresponding to each vector line to be trained, each vector line to be trained is segmented to obtain at least one line segment corresponding to each vector line to be trained; for each line segment (P i P (i+1) ) performs parameterized interpolation processing to obtain at least one interpolation point corresponding to the vector line to be trained, where the interpolation point is: P(t) = P i +t·(P (i+1) -P i ), t∈[0,1]; according to the target number of sampling points and the equally spaced step size corresponding to each vector line to be trained, the sampling point position corresponding to the vector line to be trained is determined, wherein the k-th sampling point position is determined by the cumulative distance s=kΔ, wherein k∈{0,1,...,N-1}; according to the sampling point position corresponding to each vector line to be trained, at least one interpolation point corresponding to each vector line to be trained and all vector points in the vector line to be trained are resampled to obtain the sampled vector line to be trained corresponding to each vector line to be trained for processing.
[0090] 302. Based on the embedding layer of the initial model, encode the vector line to be trained to obtain an initial embedding representation of the vector line to be trained.
[0091] Exemplarily, the initial model can be a Transformer-based model architecture, including an embedding layer, a Transformer encoder and decoder, and an output layer. Based on the embedding layer, the coordinate information and category information of each vector point in each vector line to be trained are encoded to obtain a combined embedding representation in the initial embedding representation of each vector line to be trained; the trip information and lane line information of each vector point in each vector line to be trained are encoded to obtain an attribute embedding representation in the initial embedding representation of each vector line to be trained. Among them, the coordinate information (such as longitude and latitude) and category information (such as road type and name) are represented in the form of a combined embedding representation, in accordance with the general data representation method of the vector data model, which is convenient for subsequent data processing; the trip information (such as vehicle driving path and time) and lane line information (such as lane position and direction) are represented in the form of an attribute embedding representation, which can more intuitively reflect the driving conditions of vehicles in different lanes, and facilitate traffic flow analysis or simulation.
[0092] For example, the initial model can be a Transformer-based Map Transformer model, with hyperparameters such as hidden layer dimension, number of attention heads, number of encoder and decoder layers, and number of queries set. For example, the hidden layer dimension is set to 256, the number of attention heads is set to 8, the number of encoder and decoder layers is set to 6, and the number of queries is set to 50. In the embedding layer, the coordinate information and category information of each vector point in each vector line to be trained are encoded respectively through a multi-layer perceptron (MLP) and an embedding layer. Then, the category information of each vector point is adjusted in combination with the conditional layer to form a comprehensive embedding representation, that is, a combined embedding representation of each vector line to be trained is obtained.
[0093] In one example, step 302 includes the following steps:
[0094] In the first step of step 302 , the multi-layer perceptron based on the embedding layer processes the coordinate information of the vector line to be trained to obtain the coordinate embedding representation of the vector line to be trained.
[0095] The second step of step 302 is to process the category information of the vector line to be trained based on the category embedding layer of the embedding layer to obtain the category embedding representation of the vector line to be trained.
[0096] The third step of step 302 is to process the category embedding representation of the vector line to be trained based on the conditional adjustment layer of the embedding layer to obtain the conditional embedding representation of the vector line to be trained.
[0097] The fourth step of step 302 is to combine the coordinate embedding representation and the conditional embedding representation of the vector line to be trained based on the embedding layer to obtain a combined embedding representation of the vector line to be trained.
[0098] In the fifth step of step 302, based on the embedding layer, position encoding processing is performed on the attribute information of the vector line to be trained to obtain the attribute embedding representation of the vector line to be trained; wherein, the combined embedding representation and the attribute embedding representation constitute the initial embedding representation.
[0099] For example, Figure 5 A model architecture diagram of a Map Transformer provided in an embodiment of the present application, such as Figure 5 As shown, the embedding layer includes the input projection layer, the category embedding layer and the conditional adjustment layer. Based on the multi-layer perceptron in the input projection layer, the shape of the input tensor (input_tensor) can be defined by the reshape tensor operation for the coordinate information of each vector line to be trained. The reshape tensor operation can be expressed as input_tensor[...,:2].reshape(B,M×L max ,N×2), to process the coordinate information and obtain the coordinate embedding representation coord_embed of each vector line to be trained, where B represents the batch size of all vector lines to be trained, M represents the preset embedding representation dimension, and Lmax represents the maximum length of the coordinate sequence corresponding to the vector line to be trained; N is the number of vector points in the vector line to be trained; based on the category embedding layer, the category information of each vector line to be trained can be used to define the shape of the input tensor (input_tensor) through the view tensor operation, where the view operation can be expressed as input_tensor[...,2].view(B,M×L max,N), to process the category information and obtain the category embedding representation class_embed of each vector line to be trained; based on the conditional adjustment layer, the category embedding representation of each vector line to be trained is processed, for example, through the connection operation in the conditional adjustment layer, the category embedding representation of each vector line to be trained and the preset conditional vector are vector-connected to obtain the conditional embedding representation conditional_embed of each vector line to be trained; through the tensor combination or vector splicing in the embedding layer, the coordinate embedding representation and the conditional embedding representation of each vector line to be trained are combined to obtain the combined embedding representation combined_embed of each vector line to be trained. Through the preset position encoding formula in the embedding layer, the attribute information of each vector line to be trained is position-encoded to obtain the attribute embedding representation of each vector line to be trained; wherein, each attribute embedding representation includes a trip embedding representation (Trip Embed) representing the trip information of the vector line to be trained and a lane embedding representation (Line Embed) representing the lane information of the vector line to be trained, so as to distinguish different acquisition trips and lane line instances. Finally, the combined embedding representation and attribute embedding representation of each vector line to be trained constitute the initial embedding representation of each vector line to be trained.
[0100] In one example, based on several hidden layers in a multi-layer perceptron, an activation function (such as ReLU) is used to perform nonlinear processing on the coordinate information of each vector line to be trained to obtain the initial coordinate vector of the vector line to be trained, and a vector dimensionality reduction algorithm is used to reduce the dimensionality of the initial coordinate vector to obtain the coordinate embedding representation of each vector line to be trained; wherein, the number of hidden layers can be adjusted according to the complexity of the current data processing task. Based on the category embedding layer, the category information of each vector line to be trained is converted into an integer index, such as the first category is the integer 0 and the second category is the integer 1; the preset embedding matrix is called, in which each row corresponds to the embedding vector of a category, and for each category index, the corresponding embedding vector is found, which is the category embedding representation of the corresponding vector line to be trained; wherein, if a vector line to be trained has multiple categories, the embedding vectors corresponding to these categories can be combined (such as average, summation) to obtain the final category embedding representation. Based on the conditional adjustment layer, the category embedding representation of each vector line to be trained and the preset conditional vector are vector-connected through the connection operation in the conditional adjustment layer to obtain Connect the vectors and perform linear processing on the connected vectors through the preset activation function in the conditional adjustment layer to obtain the conditional embedding representation of each vector line to be trained. Based on the combination operation in the embedding layer, such as averaging, summing or weighted summing, the coordinate embedding representation and conditional embedding representation of each vector line to be trained are combined to obtain the combined embedding representation of the vector line to be trained. Based on the embedding layer, according to the pass information in the attribute information of each vector line to be trained, the time step corresponding to each vector line to be trained (i.e., the time information t corresponding to each pass) is determined, and the preset sine and cosine functions are used to perform position encoding processing on the time step t corresponding to each vector line to be trained to obtain the corresponding position encoding vector, and the pass information and lane line information in the attribute information of each vector line to be trained are vector encoded to obtain the pass embedding representation and lane line embedding representation of each vector line to be trained respectively. Through simple vector addition in the embedding layer, the pass embedding representation, lane line embedding representation, and position encoding vector of each vector line to be trained are vector added to obtain the attribute embedding representation of each vector line to be trained.
[0101] For example, for the input coordinate information and category information of each vector line to be trained, the coordinate information includes the coordinates of each vector point in each vector line to be trained. The calculation formula of the embedding layer is as follows:
[0102] coord_embed = MLP coord (input_tensor[...,:2].reshape(B,M×L max ,N×2));
[0103] class_embed=Embeddingclass (input_tensor[...,2].view(B,M×L max ,N));
[0104] conditional_embed=Linear(class_embed);
[0105] combined_embed=coord_embed+conditional_embed;
[0106] Among them, B represents the batch size of all vector lines to be trained, M represents the preset embedding representation dimension, and L max Represents the maximum length of the coordinate sequence corresponding to the vector line to be trained; N is the number of vector points in the vector line to be trained, MLP coord It is a multi-layer perceptron with coordinate embedding. class It is a category embedding layer, and Linear is a conditional adjustment layer. Through the position encoding formula:
[0107]
[0108] The corresponding learnable embedding vector and the trip information and lane line information in the attribute information of each vector line to be trained are processed to obtain the joint position encoding of the trip (Trip) and lane line (Line) of each vector line to be trained, that is, the attribute embedding representation of each vector line to be trained is obtained. are the corresponding learnable embedding vectors.
[0109] 303. Based on the encoder and decoder of the initial model, the initial embedded representation of the vector line to be trained is processed to obtain vector line features of the vector line to be trained; wherein the vector line features represent the geometric shape and semantic structure of the vector line to be trained.
[0110] Exemplarily, an encoder based on the initial model performs encoding fusion processing on the combined embedding representation and the attribute embedding representation in the initial embedding representation of each vector line to be trained to obtain the encoded embedding representation corresponding to each vector line to be trained. A decoder based on the initial model performs decoding processing on the corresponding encoded embedding representation of each vector line to be trained to obtain the vector line features of each vector line to be trained to characterize the geometric shape and semantic structure of each vector line to be trained.
[0111] In one example, step 303 includes the following steps:
[0112] The first step of step 303 is to process the initial embedding representation of each vector line to be trained based on the multi-head self-attention mechanism of the encoder to obtain a set of attention features corresponding to each vector line to be trained; wherein the attention feature set includes at least one self-attention feature; each attention feature represents the global dependency relationship between each vector line to be trained.
[0113] The second step of step 303 is to process the attention feature set based on the multi-head attention mechanism of the decoder to obtain the vector line features of the vector line to be trained.
[0114] Exemplarily, the encoder of the initial model deploys a multi-head self-attention mechanism, while the decoder deploys a multi-head attention mechanism. Based on the multi-head self-attention mechanism, attention processing is performed on the combined embedding representation and attribute embedding representation in the initial embedding representation of all vector lines to be trained, resulting in a set of attention features corresponding to all vector lines to be trained. This set of attention features includes one or more attention features that characterize the global dependencies between the vector lines to be trained. Based on the multi-head attention mechanism, all attention features are processed to obtain vector line features for each vector line to be trained, enabling the initial model to understand the complex geometric and semantic structures of the vector map.
[0115] Among them, based on the multi-head self-attention mechanism, for the input data X∈R (B×S×d) , calculate the attention weight:
[0116]
[0117] Where Q = XW Q , K=XW K , V=XW V is a linear projection. X is the input matrix including all embedded representations; Q is the query matrix, W Q is the query weight matrix; K is the key matrix, W K is the key weight matrix; V is the value matrix, that is, the linear projection matrix, W V is the value weight matrix; d is the vector dimension; B represents the batch size of all vector lines to be trained, and S represents the length of the corresponding coordinate sequence.
[0118] For example, combining Figure 5Based on the encoder's multi-head self-attention mechanism, the coordinate embedding (Coord Embed), class embedding (Class Embed), trip embedding (Trip Embed), lane embedding (Line Embed), and at least one learnable embedding (Object Query) in the combined embedding representation of each vector line to be trained are processed to obtain a set of attention features for all vector lines to be trained. This captures the vector element relationships between all vector lines to be trained. Based on the decoder's multi-head attention mechanism, the attention feature set is processed to obtain the vector line features of each vector line to be trained.
[0119] 304. Based on the output layer of the initial model, the vector line features of the vector line to be trained are processed to obtain a predicted vector result of the vector line to be trained.
[0120] For example, in combination Figure 5 , the initial model includes an output layer, i.e. Figure 5 The prediction heads (predictionheads) in the algorithm include a feedforward neural network (FNN), and the output layer includes a classification head and a vector line head, which are respectively used to predict the category and coordinate points of the vector elements. That is, the conditional adjustment layer based on the classification head processes the vector line features of each vector line to be trained to obtain the predicted category information of each vector line to be trained; the multi-layer perceptron based on the vector line head processes the vector line features of each vector line to be trained to obtain the predicted coordinate information of each vector line to be trained, wherein the predicted category information and the predicted coordinate information constitute the predicted vector result.
[0121] For example, through the formula of the output layer: outputs_class = Linear(hs), outputs_polylines = MLP polyline (hs).sigmoid(), processes the output of the encoder and decoder of the Transformer model architecture, that is, the vector line features of each vector line to be trained, to obtain the final predicted vector result, where MLP polyline It is a multi-layer perceptron used to predict the coordinates of vector lines. outputs_class is the predicted category information of each vector line to be trained, and outputs_polylines is the predicted coordinate information of each vector line to be trained.
[0122] Furthermore, the Transformer architecture is introduced to process vector map data. By leveraging its powerful parallel computing capabilities and ability to model sequential data, the complex relationships between different vector lines can be effectively captured, thereby improving the accuracy and efficiency of the fusion-generated predicted vector results.
[0123] 305. Based on the Hungarian matching algorithm, the predicted vector results of each to-be-trained vector line and the real vector results of each to-be-trained vector line are processed to obtain a matching index; wherein the matching index includes the predicted vector result of each to-be-trained vector line and the real vector result that matches the predicted vector result.
[0124] Exemplarily, each vector line to be trained has a real vector result, including real coordinate information and real category information. A matcher based on the Hungarian matching algorithm matches the predicted vector results of all vector lines to be trained with the real vector results of all vector lines to be trained, obtaining a matching index. The matching index is a data structure comprising a matching set and a matching relationship. The matching set includes multiple matching pairs, each matching pair including the predicted vector result of each vector line to be trained and the real vector result that matches the predicted vector result. The matching relationship is the matching relationship between the predicted vector result and the corresponding real vector result in each matching pair.
[0125] For example, multiple vector lines to be trained are divided into batches, and the data is fed into the model for training using a DataLoader. In each training batch, data such as the input tensor, mask matrix, and true label tensor are transferred to the device (such as a GPU). The forward method of the initial model is called to obtain the prediction results, including the predicted classification results and vector line coordinate results. Based on the true labels, a target dictionary is constructed, which contains the true category label and true vector line coordinates of each vector line to be trained. The matcher of the Hungarian matching algorithm is used to match the predicted result of each vector line to be trained with the true label of each other vector line to be trained to obtain a matching index.
[0126] In one example, step 305 includes the following steps:
[0127] The first step of step 305 is to determine a cost matrix corresponding to each to-be-trained vector line based on the predicted vector result of each to-be-trained vector line and the actual vector result of each to-be-trained vector line; wherein the cost matrix represents the matching cost between each predicted vector result of each to-be-trained vector line and the actual vector result of each to-be-trained vector line.
[0128] The second step of step 305 is to calculate and process the cost matrix based on the Hungarian matching algorithm to obtain a matching index.
[0129] Exemplarily, a pre-set matching cost calculation method is used to calculate the predicted vector results and the actual vector results for each training vector line. This results in a cost matrix corresponding to all training vector lines. This cost matrix represents the cost of matching each pairwise between the predicted vector results and the actual vector results for each training vector line. This cost matrix is then processed using the Hungarian matching algorithm to determine the optimal match and obtain a matching index.
[0130] For example, based on the predicted vector results of all vector lines to be trained (including the predicted category information and predicted coordinate information of each vector line to be trained) and the real vector results, the matrix set corresponding to all vector lines to be trained is determined; wherein the real results include the real category information and the real coordinate information; the matrix set includes the classification costs, the distance costs and the direction costs of the vector lines to be trained that are matched between the predicted results and the real results of all vector lines to be trained; according to the preset weight set and the matrix set corresponding to the vector map data, the cost matrix corresponding to all vector lines to be trained is constructed; wherein the preset weight set includes the classification cost weight coefficient, the distance cost weight coefficient of the vector line to be trained and the direction cost weight coefficient.
[0131] For another example, for the predicted vector results of all vector lines to be trained, a predicted set {P1, P2, ..., Pm} is constructed, and for the real vector results of all vector lines to be trained, a real set {T1, T2, ..., Tn} is constructed, and a cost matrix C is constructed, wherein the cost element C(i, j) can be calculated by the formula: C(i, j) = α·cost class (Pi,Tj)+β·cost polyline (Pi,Tj)+γ·cost dir (Pi, Tj) is calculated; where the element C(i, j) represents the matching cost between the predicted vector result Pi of the i-th training vector line and the true vector result Tj of the j-th training vector line, and α, β, and γ are the classification cost weight coefficients, the training vector line distance cost weight coefficients, and the direction cost weight coefficients, respectively. The cost matrix C is processed using the Hungarian algorithm to find the matching pair that minimizes the total cost. Specifically, a matching set is found, namely the matching indices {(i1, j1), (i2, j2), ..., (ik, jk)}, and the following formula is calculated:
[0132]
[0133] The value of the above formula is minimized, and each predicted element and true element are matched at most once.
[0134] In the calculation of the loss function, the Hungarian matching algorithm is used to match the predicted results and the true labels, ensuring that the model can more accurately learn the correspondence between different vector elements during the training process, improving the robustness and generalization ability of the model, and obtaining a prediction model with better prediction performance.
[0135] 306. Determine a multi-task loss function corresponding to the matching index based on the predicted vector result of each to-be-trained vector line in the matching index and the corresponding matched real vector result; wherein the multi-task loss function represents the difference between the predicted vector result of each to-be-trained vector line and the corresponding matched real vector result.
[0136] Exemplarily, according to a preset loss function calculation method, the value of the loss function is calculated for the predicted vector result of each vector line to be trained in the matching index and the corresponding matched real vector result, and a multi-task loss function corresponding to the matching index is obtained to characterize the gap between the predicted vector result of each vector line to be trained and the corresponding matched real vector result.
[0137] In one example, step 306 includes the following steps:
[0138] The first step of step 306 is to determine the classification loss corresponding to the matching index based on the predicted category information in the predicted vector result of each vector line to be trained in the matching index and the actual category information in the corresponding matched actual vector result; wherein the classification loss represents the difference between the predicted category distribution and the actual category distribution of all vector lines to be trained.
[0139] The second step of step 306 is to determine the vector line loss and direction loss corresponding to the matching index based on the predicted coordinate information in the predicted vector result of each vector line to be trained in the matching index and the actual coordinate information in the corresponding matched actual vector result; wherein the vector line loss represents the difference between the predicted geometric shapes and the actual geometric shapes of all vector lines to be trained; and the direction loss represents the difference between the predicted directions and the actual directions of all vector lines.
[0140] The third step of step 306 is to determine the multi-task loss function corresponding to the matching index according to the classification loss, vector line loss and direction loss corresponding to the matching index.
[0141] Exemplarily, according to a preset classification loss calculation method, the predicted category information in the prediction result of each vector line to be trained in the matching index and the actual category information in the corresponding actual result are calculated to determine the classification loss corresponding to the matching index; wherein the classification loss represents the difference between the predicted category distribution and the actual category distribution of all vector lines to be trained. According to a preset vector line loss calculation method, the symmetric average distance corresponding to the matching index is determined based on the predicted coordinate information in the prediction result of each vector line to be trained in the matching index and the actual coordinate information in the corresponding actual result; wherein the symmetric average distance is the symmetric average distance between each predicted coordinate point in the predicted coordinate information of each vector line to be trained and each actual coordinate point in the corresponding actual coordinate information; the symmetric average distance corresponding to the matching index is determined as the vector line loss corresponding to the matching index; wherein the vector line loss represents the difference between the predicted geometric shapes and the actual geometric shapes of all vector lines to be trained. According to a preset direction loss calculation method, the predicted coordinate information in the prediction result of each vector line to be trained in the matching index is calculated with the actual coordinate information in the corresponding actual result to obtain a set of start and end point vectors corresponding to the matching index; wherein the start and end point distance set includes the predicted direction vector and the actual direction vector of each vector line to be trained; the predicted direction vector is the start and end point direction vector corresponding to the first and last coordinate points in the predicted coordinate information; the actual direction vector is the start and end point direction vector corresponding to the first and last coordinate points in the actual coordinate information; the similarity set corresponding to the set of start and end point vectors is calculated, and the similarity set includes the similarity between the predicted direction vector and the actual direction vector of each vector line to be trained, and the similarity set corresponding to the set of start and end point vectors is determined as the direction loss corresponding to the matching index; wherein the direction loss represents the difference between the predicted direction and the actual direction of all vector lines to be trained. The classification loss, vector line loss, and direction loss corresponding to the matching index are calculated, such as by weighted summation, to obtain a multi-task loss function corresponding to the matching index.
[0142] For example, the Set Criterion class, a key component for calculating loss functions, is defined to implement multi-task loss functions, including classification loss, vector line loss, and direction loss. The classification loss uses cross-entropy loss, the vector line loss is calculated by calculating the average distance between the predicted and true vector lines, and the direction loss is calculated based on the cosine similarity of the direction vectors at the start and end points of the vector line.
[0143] Specifically, classification loss usually uses cross entropy loss to measure the difference between the category distribution predicted by the model and the true category distribution. The formula for classification loss in the Map Transformer model is:
[0144]
[0145] Among them, y i is the one-hot encoding of the true label, indicating whether the i-th category is the true category; p i Is the probability of belonging to the i-th category predicted by the model. Define the calculation formula of the symmetric average distance between the predicted point set P and the true value point set Q:
[0146]
[0147] Among them, p and q are respectively the predicted coordinate information of each vector point in each vector line to be trained in the prediction point set P and the true coordinate information of each vector point in each vector line to be trained in the true value point set Q. Calculate the predicted direction vector d pred =P N -P0, true direction vector d gt =Q N -Q0; and calculate the cosine similarity L of the direction vectors of the first and last points direction , as shown below:
[0148]
[0149] The multi-task loss function Ltotal is expressed as: Ltotal = α1L class +β1L polyline +γ1L direction , where α1, β1, and γ1 are the classification loss weight coefficients, vector line loss weight coefficients, and direction loss weight coefficients, respectively, which are used to balance the contributions of different losses.
[0150] By introducing the Hungarian matching algorithm and multi-task loss function, the model can more comprehensively learn the characteristics and relationships of vector elements during training, improving its adaptability to different data distributions and noise. This enables the model to stably output accurate fusion results when faced with diverse data in real applications.
[0151] 307. According to the multi-task loss function, the initial model is trained to obtain a prediction model.
[0152] Exemplarily, the initial model is trained according to the multi-task loss function, and the model parameters of the initial model are continuously optimized to obtain an optimized model, which is the prediction model.
[0153] For example, the parameters of the initial model are updated by backpropagating the total loss of the multi-task loss function. During the training process, the initial model is optimized using an optimizer (such as the AdamW optimizer) and a learning rate scheduler (such as the cosine annealing learning rate scheduler), gradually reducing the loss value of the multi-task loss function and improving the performance of the initial model to obtain a prediction model.
[0154] 308. Process the vector map data to be processed based on the prediction model to generate a vector map corresponding to the vector map data to be processed.
[0155] For example, this step may refer to step 203 and will not be described in detail here.
[0156] In this embodiment, based on the above embodiments, on the one hand, a prediction model is obtained based on the powerful modeling capabilities of the Transformer architecture, ensuring that the generated vector map has high quality in terms of geometric shape, topological relationship and semantic information; on the other hand, the end-to-end architecture makes the maintenance and update of the model more convenient, and can quickly adapt to new data sources and application requirements, simplifying the entire generation process, reducing the data processing and conversion links, and reducing the complexity and error probability of the system.
[0157] Figure 6 A flow chart of another vector map generation method provided in an embodiment of the present application is shown as follows: Figure 6 As shown, the method includes:
[0158] 401. Obtain vector map data to be processed; perform normalization processing on all vector points in the vector line to be processed, and at the same time, perform resampling processing on the vector line to be processed to obtain a processed vector line to be processed.
[0159] For example, the device can obtain unprocessed vector map data from a multi-source database. The unprocessed vector map data includes multiple unprocessed vector lines, each of which includes multiple vector points. For each unprocessed vector line, all vector points in each unprocessed vector line are normalized based on a preset normalization algorithm, mapping coordinate values in different ranges to a unified interval, eliminating scale differences in the data and providing a good data foundation for subsequent model processing. At the same time, each unprocessed vector line is resampled based on a preset resampling algorithm to obtain each processed unprocessed vector line. For example, the number of sampling points for each lane line is fixed to a preset value to ensure dimensional consistency of the input data while retaining the main geometric features of the vector line.
[0160] 402. Based on the embedding layer of the prediction model, encode the vector line to be processed to obtain an initial embedding representation of the vector line to be processed.
[0161] For example, the prediction model is trained based on an initial model, which can be a Transformer-based model architecture, including an embedding layer, a Transformer encoder and decoder, and an output layer. Based on the embedding layer, the coordinate information and category information of each vector point in each vector line to be processed are encoded to obtain a combined embedding representation of each vector line to be processed. The trip information and lane information of each vector point in each vector line to be processed are also encoded to obtain an attribute embedding representation of each vector line to be processed.
[0162] In one example, step 402 includes the following steps:
[0163] The first step of step 402 is to process the category information of the vector line to be processed based on the category embedding layer of the embedding layer to obtain a category embedding representation of the vector line to be processed.
[0164] The second step of step 402 is to process the category embedding representation of the vector line to be processed based on the conditional adjustment layer of the embedding layer to obtain the conditional embedding representation of the vector line to be processed.
[0165] In the third step of step 402 , based on the embedding layer, the coordinate embedding representation and the conditional embedding representation of the vector line to be processed are combined to obtain a combined embedding representation of the vector line to be processed.
[0166] The fourth step of step 402 is to perform position encoding processing on the attribute information of the vector line to be processed based on the embedding layer to obtain the attribute embedding representation of the vector line to be processed; wherein, the combined embedding representation and the attribute embedding representation constitute the initial embedding representation.
[0167] Exemplarily, the embedding layer includes an input projection layer, a category embedding layer, and a conditional adjustment layer. Based on the multi-layer perceptron in the input projection layer, the coordinate information of each input vector line to be processed is processed to obtain a coordinate embedding representation of each vector line to be processed; based on the category embedding layer, the category information of each vector line to be processed is processed to obtain a category embedding representation of each vector line to be processed; based on the conditional adjustment layer, the category embedding representation of each vector line to be processed is processed to obtain a conditional embedding representation of each vector line to be processed; by combining tensors or vector concatenation in the embedding layer, the coordinate embedding representation and the conditional embedding representation of each vector line to be processed are combined to obtain a combined embedding representation of each vector line to be processed. The attribute information of each vector line to be processed is position-encoded using a preset position encoding formula in the embedding layer to obtain an attribute embedding representation of each vector line to be processed; wherein, each attribute embedding representation includes a trip embedding representation representing the trip information of the vector line to be processed and a lane line embedding representation representing the lane line information of the vector line to be processed, so as to distinguish different acquisition trips and lane line instances.
[0168] 403. The encoder and decoder based on the prediction model process the initial embedding representation of the vector line to be processed to obtain vector line features of the vector line to be processed; wherein the vector line features represent the geometric shape and semantic structure of the vector line to be processed.
[0169] Exemplarily, an encoder based on a prediction model performs encoding fusion processing on the combined embedding representation and the attribute embedding representation in the initial embedding representation of each vector line to be processed to obtain the encoded embedding representation corresponding to each vector line to be processed. A decoder based on the initial model performs decoding processing on the corresponding encoded embedding representation of each vector line to be processed to obtain the vector line features of each vector line to be processed to characterize the geometric shape and semantic structure of each vector line to be processed.
[0170] In one example, step 403 includes the following steps:
[0171] The first step of step 403 is to process the initial embedding representation of each vector line to be processed based on the multi-head self-attention mechanism of the encoder to obtain an attention feature set corresponding to each vector line to be processed; wherein the attention feature set includes at least one attention feature; each attention feature represents the global dependency relationship between each vector line to be processed.
[0172] The second step of step 403 is to process the attention feature set based on the multi-head attention mechanism of the decoder to obtain the vector line features of the vector line to be processed.
[0173] Exemplarily, the encoder of the prediction model deploys a multi-head self-attention mechanism, while the decoder deploys a multi-head attention mechanism. Based on the multi-head self-attention mechanism, attention processing is performed on the combined embedding representation and attribute embedding representation in the initial embedding representation of all vector lines to be processed, resulting in a set of attention features corresponding to all vector lines to be processed. This set of attention features includes one or more attention features that characterize the global dependencies between the individual vector lines to be processed. Based on the multi-head attention mechanism, all attention features are processed to obtain vector line features for each vector line to be processed, enabling the initial model to understand the complex geometric and semantic structures in the vector map.
[0174] 404. Based on the output layer of the prediction model, the vector line features of the vector line to be processed are processed to obtain a prediction vector result of the vector line to be processed.
[0175] Exemplarily, the prediction model includes an output layer, in which a feedforward neural network is set up, and the output layer includes a classification head and a vector line head, which are respectively used to predict the category and coordinate points of the vector elements, that is, the conditional adjustment layer based on the classification head processes the vector line features of each vector line to be processed to obtain the predicted category information of each vector line to be processed; the multi-layer perceptron based on the vector line head processes the vector line features of each vector line to be processed to obtain the predicted coordinate information of each vector line to be processed, wherein the predicted category information and the predicted coordinate information constitute the predicted vector result.
[0176] 405. Generate a vector map corresponding to the vector map data to be processed according to the predicted vector results of each vector line to be processed.
[0177] For example, Figure 7 A schematic diagram of an example of a data processing platform provided in an embodiment of the present application, such as Figure 7 As shown in the figure, the trained model is integrated into actual application systems, such as navigation software, geographic information systems, etc. Based on the preset vector map generation tool, the predicted vector results of all vector lines to be processed are processed to obtain the vector maps corresponding to all vector lines to be processed, so as to realize the end-to-end fusion generation function of the vector map and provide users with high-quality map services.
[0178] In this embodiment, based on the above embodiments, on the one hand, a prediction model is obtained based on the powerful modeling capabilities of the Transformer architecture, ensuring that the generated vector map has high quality in terms of geometric shape, topological relationship and semantic information; on the other hand, the end-to-end architecture makes the maintenance and update of the model more convenient, and can quickly adapt to new data sources and application requirements, simplifying the entire generation process, reducing the data processing and conversion links, and reducing the complexity and error probability of the system.
[0179] Figure 8 A schematic diagram of the structure of a vector map generating device provided in an embodiment of the present application is shown as follows: Figure 8 As shown, the device includes:
[0180] The processing module 501 is used to obtain vector map data to be trained; and determine the predicted vector result of each vector line to be trained in the vector map data to be trained based on the initial model;
[0181] A training module 502 is configured to train the initial model based on the Hungarian matching algorithm according to the predicted vector results of each vector line to be trained to obtain a prediction model;
[0182] The generating module 503 is configured to process the vector map data to be processed based on the prediction model, and generate a vector map corresponding to the vector map data to be processed.
[0183] In one possible implementation, the training module 502 is specifically used to: process the predicted vector results of each vector line to be trained and the actual vector results of each vector line to be trained based on the Hungarian matching algorithm to obtain a matching index; wherein the matching index includes the predicted vector result of each vector line to be trained and the actual vector result that matches the predicted vector result; determine the multi-task loss function corresponding to the matching index based on the predicted vector result of each vector line to be trained and the corresponding matched actual vector result; wherein the multi-task loss function represents the gap between the predicted vector result of each vector line to be trained and the corresponding matched actual vector result; and train the initial model according to the multi-task loss function to obtain a prediction model.
[0184] In one possible implementation, the training module 502 is specifically configured to: determine a cost matrix corresponding to each vector line to be trained based on a predicted vector result of each vector line to be trained and a true vector result of each vector line to be trained; wherein the cost matrix represents the matching cost between each predicted vector result of each vector line to be trained and the true vector result of each vector line to be trained; and calculate and process the cost matrix based on a Hungarian matching algorithm to obtain a matching index.
[0185] In one possible embodiment, the training module 502 is further specifically used to: determine the classification loss corresponding to the matching index based on the predicted category information in the predicted vector result of each vector line to be trained in the matching index and the actual category information in the corresponding matched actual vector result; wherein the classification loss represents the difference between the predicted category distribution and the actual category distribution of all vector lines to be trained; determine the vector line loss and direction loss corresponding to the matching index based on the predicted coordinate information in the predicted vector result of each vector line to be trained in the matching index and the actual coordinate information in the corresponding matched actual vector result; wherein the vector line loss represents the difference between the predicted geometric shapes and the actual geometric shapes of all vector lines to be trained; the direction loss represents the difference between the predicted directions and the actual directions of all vector lines; and determine the multi-task loss function corresponding to the matching index based on the classification loss, vector line loss and direction loss corresponding to the matching index.
[0186] In one possible implementation, the processing module 501 is specifically configured to: encode the vector line to be trained based on the embedding layer of the initial model to obtain an initial embedding representation of the vector line to be trained; process the combined embedding representation and attribute embedding representation of the vector line to be trained based on the encoder and decoder of the initial model to obtain vector line features of the vector line to be trained; wherein the vector line features represent the geometric shape and semantic structure of the vector line to be trained; and process the vector line features of the vector line to be trained based on the output layer of the initial model to obtain a predicted vector result of the vector line to be trained.
[0187] In one possible implementation, the processing module 501 is specifically configured to: process the coordinate information of the vector line to be trained based on a multi-layer perceptron of the embedding layer to obtain a coordinate embedding representation of the vector line to be trained; process the category information of the vector line to be trained based on a category embedding layer of the embedding layer to obtain a category embedding representation of the vector line to be trained; process the category embedding representation of the vector line to be trained based on a conditional adjustment layer of the embedding layer to obtain a conditional embedding representation of the vector line to be trained; combine the coordinate embedding representation and the conditional embedding representation of the vector line to be trained based on the embedding layer to obtain a combined embedding representation of the vector line to be trained; perform position encoding processing on the attribute information of the vector line to be trained based on the embedding layer to obtain an attribute embedding representation of the vector line to be trained; wherein the combined embedding representation and the attribute embedding representation constitute an initial embedding representation.
[0188] In one possible implementation, the processing module 501 is further specifically used to: process the initial embedding representation of each vector line to be trained based on the multi-head self-attention mechanism of the encoder to obtain an attention feature set corresponding to each vector line to be trained; wherein the attention feature set includes at least one self-attention feature; each attention feature represents the global dependency relationship between each vector line to be trained; and process the attention feature set based on the multi-head attention mechanism of the decoder to obtain the vector line features of the vector line to be trained.
[0189] In a possible implementation, the device is further configured to: perform normalization processing on all vector points in the vector line to be trained, and at the same time, perform resampling processing on the vector line to be trained to obtain a processed vector line to be trained.
[0190] The device of this embodiment can execute the technical solution in the above method. Its specific implementation process and technical principles are the same and will not be repeated here.
[0191] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 9 As shown, the electronic device includes: a memory 601 and a processor 602; the memory 601 is a memory for storing instructions executable by the processor 602.
[0192] The processor 602 is configured to execute the method provided in the above embodiment.
[0193] The electronic device further includes a receiver 603 and a transmitter 604. The receiver 603 is used to receive instructions and data sent by other devices, and the transmitter 604 is used to send instructions and data to external devices.
[0194] The specific implementation process of the processor can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0195] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly implemented by a hardware processor or implemented by a combination of hardware and software modules in the processor.
[0196] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed on a computer, the computer executes the technical solution of the above embodiment.
[0197] The readable storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0198] An exemplary readable storage medium is coupled to a processor, such that the processor can read information from and write information to the readable storage medium. The readable storage medium may also be an integral part of the processor. The processor and the readable storage medium may reside in an application-specific integrated circuit. The processor and the readable storage medium may also reside in a device as discrete components.
[0199] An embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when at least one processor executes the computer program, the technical solutions in the above embodiments can be implemented.
[0200] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as a magnetic disk or optical disk.
[0201] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.
Claims
1. A vector map generation method, characterized in that: include: Get the vector map data to be trained; and determining a predicted vector result of each vector line to be trained in the vector map data to be trained based on the initial model; Based on the Hungarian matching algorithm, the initial model is trained according to the predicted vector result of each vector line to be trained to obtain a prediction model; The vector map data to be processed is processed based on the prediction model to generate a vector map corresponding to the vector map data to be processed.
2. The method according to claim 1, characterized in that The method of training the initial model based on the Hungarian matching algorithm according to the predicted vector result of each vector line to be trained to obtain a prediction model includes: Based on the Hungarian matching algorithm, the predicted vector results of each of the vector lines to be trained and the real vector results of each of the vector lines to be trained are processed to obtain a matching index; wherein the matching index includes the predicted vector result of each of the vector lines to be trained and the real vector result that matches the predicted vector result; Determining a multi-task loss function corresponding to the matching index based on the predicted vector result of each to-be-trained vector line and the corresponding matched real vector result in the matching index; wherein the multi-task loss function represents the difference between the predicted vector result of each to-be-trained vector line and the corresponding matched real vector result; The initial model is trained according to the multi-task loss function to obtain the prediction model.
3. The method according to claim 2, characterized in that The method of processing the predicted vector result of each of the vector lines to be trained and the real vector result of each of the vector lines to be trained based on the Hungarian matching algorithm to obtain a matching index includes: Determining a cost matrix corresponding to each of the vector lines to be trained based on the predicted vector results of each of the vector lines to be trained and the actual vector results of each of the vector lines to be trained; wherein the cost matrix represents a matching cost between each predicted vector result of each of the vector lines to be trained and the actual vector results of each of the vector lines to be trained; Based on the Hungarian matching algorithm, the cost matrix is calculated to obtain the matching index.
4. The method according to claim 2, characterized in that The processing according to the predicted vector result of each to-be-trained vector line in the matching index and the corresponding matched real vector result to determine the multi-task loss function corresponding to the matching index includes: Determining a classification loss corresponding to the matching index based on predicted category information in the predicted vector result of each to-be-trained vector line in the matching index and true category information in the corresponding matched true vector result; wherein the classification loss represents the difference between the predicted category distribution and the true category distribution of all to-be-trained vector lines; Determine, based on the predicted coordinate information in the predicted vector result of each vector line to be trained in the matching index and the actual coordinate information in the corresponding matched actual vector result, a vector line loss and a direction loss corresponding to the matching index; wherein the vector line loss represents the difference between the predicted geometric shapes and the actual geometric shapes of all vector lines to be trained; and the direction loss represents the difference between the predicted directions and the actual directions of all vector lines; Determine a multi-task loss function corresponding to the matching index according to the classification loss, vector line loss, and direction loss corresponding to the matching index.
5. The method according to claim 1, wherein The step of determining a predicted vector result for each vector line to be trained in the vector map data to be trained based on the initial model includes: Based on the embedding layer of the initial model, encoding the vector line to be trained is performed to obtain an initial embedding representation of the vector line to be trained; Based on the encoder and decoder of the initial model, the initial embedded representation of the vector line to be trained is processed to obtain vector line features of the vector line to be trained; wherein the vector line features represent the geometric shape and semantic structure of the vector line to be trained; Based on the output layer of the initial model, the vector line features of the vector line to be trained are processed to obtain a predicted vector result of the vector line to be trained.
6. The method according to claim 5, characterized in that The embedding layer based on the initial model encodes the vector line to be trained to obtain an initial embedding representation of the vector line to be trained, including: Processing the coordinate information of the vector line to be trained based on the multi-layer perceptron of the embedding layer to obtain a coordinate embedding representation of the vector line to be trained; Based on the category embedding layer of the embedding layer, the category information of the vector line to be trained is processed to obtain a category embedding representation of the vector line to be trained; Processing the category embedding representation of the vector line to be trained based on the conditional adjustment layer of the embedding layer to obtain a conditional embedding representation of the vector line to be trained; Based on the embedding layer, the coordinate embedding representation and the conditional embedding representation of the vector line to be trained are combined to obtain a combined embedding representation of the vector line to be trained; Based on the embedding layer, position encoding processing is performed on the attribute information of the vector line to be trained to obtain the attribute embedding representation of the vector line to be trained; wherein the combined embedding representation and the attribute embedding representation constitute the initial embedding representation.
7. The method according to claim 5, characterized in that The encoder and decoder based on the initial model process the initial embedding representation of the vector line to be trained to obtain the vector line features of the vector line to be trained, including: Based on the multi-head self-attention mechanism of the encoder, the initial embedded representation of each of the vector lines to be trained is processed to obtain a set of attention features corresponding to each of the vector lines to be trained; wherein the set of attention features includes at least one self-attention feature; each of the attention features represents a global dependency relationship between each of the vector lines to be trained; Based on the multi-head attention mechanism of the decoder, the attention feature set is processed to obtain the vector line features of the vector line to be trained.
8. The method according to any one of claims 1 to 7, characterized in that The method further comprises: Normalization processing is performed on all vector points in the vector line to be trained, and at the same time, resampling processing is performed on the vector line to be trained to obtain a processed vector line to be trained.
9. A vector map generating device, characterized in that: include: A processing module, used for obtaining vector map data to be trained; and determining a predicted vector result of each vector line to be trained in the vector map data to be trained based on the initial model; A training module, configured to train the initial model based on the Hungarian matching algorithm and according to the predicted vector result of each vector line to be trained, to obtain a prediction model; The generation module is used to process the vector map data to be processed based on the prediction model, and generate a vector map corresponding to the vector map data to be processed.
10. An electronic device / computer-readable storage medium / computer program product, characterized in that: The electronic device includes: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 8; and / or the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by the processor, the computer program product includes a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
System and method for map vectorization in advanced driving assistance systems
CN118552567A
Map generation method and device, training method and device, electronic equipment, storage medium and program product
CN119131177A
Model training method, map data fusion method and related equipment
CN119783045A
Online map generation model training method and device, online map generation method and device, electronic equipment and storage medium
CN120014106A