Model training method, map data fusion method and related devices

By extracting and fusing class and coordinate data from diverse local maps using a deep learning model with Transformer architectures, the method addresses the variability in crowd-sourced map data, achieving high-precision and efficient map fusion for autonomous driving.

CN119783045BActive Publication Date: 2025-07-15XIAOMI EV TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510272483.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-15
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

In the prior art, crowdsourcing map data has low flexibility, poor accuracy and low computing efficiency due to the uneven data source diversity and quality.

Method used

By obtaining the training data sets of multiple local vector maps, the category data and coordinate data of map elements are extracted, and feature extraction and rasterization are used to use the initial map data fusion model to perform feature fusion and feature refinement, high-precision global map data are generated, and the map data fusion model is trained to improve fusion efficiency and accuracy.

Benefits of technology

It realizes efficient and accurate integration of crowdsourcing map data with uneven diversity and quality, generates a high-precision and consistency global map, improves the flexibility and fusion efficiency of map data, and meets the needs of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783045B_ABST
    Figure CN119783045B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a model training method, a map data fusion method, and related devices, and pertains to the field of autonomous driving technology. The model training method includes: obtaining at least two local vector maps and global map data; extracting at least two local vector maps to obtain category data and coordinate data of multiple map elements; using an initial map data fusion model to perform feature extraction on the category data and coordinate data to obtain vector map features; performing rasterization processing based on the coordinate data to obtain raster map features; performing feature fusion and feature refinement on the vector map features and the raster map features to obtain refined map features; performing decoding and prediction processing on the refined map features to obtain global map prediction data; and training the initial map data fusion model based on the global map prediction data and the global map data to obtain a map data fusion model. The present disclosure can obtain accurate fusion results and has high fusion efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of autonomous driving technologies, and in particular, to a model training method, a map data fusion method, and related devices. Background Art

[0002] With the rapid development of autonomous driving technologies, high-precision maps have become an indispensable part of autonomous driving systems.

[0003] In related technologies, crowdsourced map data has become an important source of high-precision map data due to its advantages of fast update speed and wide coverage. However, due to the problems of diverse data sources and uneven data quality, the flexibility and accuracy of the fused crowdsourced map data are low, and the computational efficiency of the fusion is not high.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The present disclosure provides a model training method, a map data fusion method, and related devices.

[0006] According to a first aspect of an embodiment of the present disclosure, a model training method is provided, including:

[0007] Obtaining a training data set of a plurality of map data fusion training samples; the map data fusion training samples include at least two local vector maps and global map data; the global map data is obtained by fusing the at least two local vector maps;

[0008] Extracting the at least two local vector maps to obtain class data and coordinate data of a plurality of map elements in the at least two local vector maps; and

[0009] Using an initial map data fusion model to perform feature extraction on the class data and the coordinate data to obtain vector map features; performing rasterization processing on the coordinate data to obtain raster map features; performing feature fusion and feature refinement on the vector map features and the raster map features to obtain refined map features; performing decoding and prediction processing on the refined map features to obtain global map prediction data;

[0010] Training the initial map data fusion model based on the global map prediction data and the global map data to obtain a map data fusion model.

[0011] In some embodiments of the present disclosure, performing feature extraction on the class data and the coordinate data to obtain vector map features includes:

[0012] Extract features from the category data to obtain category features;

[0013] Extract features from the coordinate data to obtain location features;

[0014] Concatenate the category features and the location features to obtain the vector map features.

[0015] In some embodiments of the present disclosure, it further includes: calculating the central coordinates of each map element based on the coordinate data to obtain the 2D position information of the multiple map elements.

[0016] In some embodiments of the present disclosure, performing feature fusion and feature refinement on the vector map features and the raster map features to obtain the refined map features, including:

[0017] Performing feature fusion based on the vector map features and the raster map features to obtain initial map features;

[0018] Performing feature fusion based on the 2D position information and the raster map features to obtain the 2D position encoding corresponding to the initial map features;

[0019] Performing feature refinement on the initial map features using a Transformer encoder based on the 2D position encoding to obtain the refined map features;

[0020] Wherein, the initial map data fusion model includes the Transformer encoder.

[0021] In some embodiments of the present disclosure, performing rasterization processing based on the coordinate data to obtain raster map features, including:

[0022] Converting the coordinate data to obtain pixel coordinate data of a raster map to obtain a raster map image;

[0023] Extracting image features from the raster map image to obtain raster map image features, and generating spatial position coordinate data corresponding to the raster map image features;

[0024] Obtaining the raster map features according to the raster map image features and the spatial position coordinate data.

[0025] In some embodiments of the present disclosure, performing decoding and prediction processing on the refined map features to obtain global map prediction data, including:

[0026] Using a Transformer decoder to decode the refined map features to obtain prediction results of the multiple map elements;

[0027] Classify and regress the prediction results of the multiple map elements using a prediction head to obtain the global map prediction data;

[0028] Wherein, the initial map data fusion model includes the Transformer decoder and the prediction head.

[0029] In some embodiments of the present disclosure, decoding the refined map features using a Transformer decoder to obtain prediction results of the multiple map elements includes:

[0030] Obtain an initial content query vector; the initial content query vector is a vector of all zeros;

[0031] Generate core parameters of the attention mechanism of the cross-attention layer of the Transformer decoder based on the refined map features;

[0032] Refine the initial content query vector based on the core parameters of the attention mechanism to obtain a refined content query vector as the prediction results of the multiple map elements.

[0033] In some embodiments of the present disclosure, the global map prediction data includes:

[0034] Class prediction data, confidence score prediction data, and coordinate prediction data corresponding to the map elements of the global map.

[0035] In some embodiments of the present disclosure, extracting the at least two local vector maps to obtain class data and coordinate data of multiple map elements in the at least two local vector maps includes:

[0036] Extract the at least two local vector maps to obtain initial class data and initial coordinate data of multiple map elements in the at least two local vector maps;

[0037] Align the dimensions of the initial class data and the initial coordinate data to obtain class data and the coordinate data of multiple map elements in the at least two local vector maps.

[0038] According to a second aspect of the embodiments of the present disclosure, there is provided a map data fusion method, including:

[0039] Obtain multiple local vector maps to be fused;

[0040] Extract the multiple local vector maps to be fused to obtain class data and coordinate data of multiple map elements in the multiple local vector maps to be fused;

[0041] Input the category data and coordinate data of multiple map elements in the multiple local vector maps to be fused into a map data fusion model, and output the global map data obtained by fusing the multiple local vector maps to be fused;

[0042] Among them, the map data fusion model is trained according to the model training method described in the first aspect above.

[0043] According to a third aspect of the embodiments of the present disclosure, there is provided a map data fusion model training device, including:

[0044] A sample acquisition unit for acquiring a training data set of multiple map data fusion training samples; the map data fusion training samples include at least two local vector maps and global map data; the global map data is obtained by fusing the at least two local vector maps;

[0045] A sample data extraction unit for extracting the at least two local vector maps to obtain the category data and coordinate data of multiple map elements in the at least two local vector maps;

[0046] A sample training unit for using an initial map data fusion model to extract features from the category data and the coordinate data to obtain vector map features; performing rasterization processing based on the coordinate data to obtain raster map features; performing feature fusion and feature refinement on the vector map features and the raster map features to obtain refined map features; performing decoding and prediction processing on the refined map features to obtain global map prediction data; training the initial map data fusion model based on the global map prediction data and the global map data to obtain a map data fusion model.

[0047] According to a fourth aspect of the embodiments of the present disclosure, there is provided a map data fusion device, including:

[0048] A vector map acquisition unit for acquiring multiple local vector maps to be fused;

[0049] A map data extraction unit for extracting the multiple local vector maps to be fused to obtain the category data and coordinate data of multiple map elements in the multiple local vector maps to be fused;

[0050] A map fusion unit for inputting the category data and coordinate data of multiple map elements in the multiple local vector maps to be fused into a map data fusion model, and outputting the global map data obtained by fusing the multiple local vector maps to be fused;

[0051] Among them, the map data fusion model is trained according to the model training method described in the first aspect above.

[0052] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0053] A processor;

[0054] A memory for storing instructions executable by the processor;

[0055] Wherein, the processor is configured to: implement the model training method described in the first aspect above or implement the map data fusion method described in the second aspect above.

[0056] According to a sixth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enabling the mobile terminal to execute the model training method described in the first aspect above or the map data fusion method described in the second aspect above.

[0057] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0058] In the present disclosure, category data and coordinate data of multiple map elements are obtained by extracting at least two local vector maps of training samples. Feature extraction is performed on the category data and coordinate data through an initial map data fusion model to obtain vector map features; rasterization processing is performed based on the coordinate data to obtain raster map features; the vector map features and raster map features are subjected to feature fusion and feature refinement to obtain refined map features; decoding and prediction processing are performed on the refined map features to obtain global map prediction data; and the initial map data fusion model is trained based on the global map prediction data and the global map data in the training samples to obtain a map data fusion model, obtaining a map data fusion model that can effectively fuse vector map data. Just extract the category data and coordinate data of map elements from the vector map data to be fused collected through crowdsourcing and input them into the trained map data fusion model, then an accurate fusion result can be obtained, and the fusion efficiency is high. Relying on the trained map data fusion model for fusion, even if the data sources are diverse or the quality is uneven, accurate fusion can still be achieved, improving the flexibility of the fused map data.

[0059] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0061] Figure 1 It is a flowchart of a model training method shown according to some embodiments of the present disclosure.

[0062] Figure 2 It is a flowchart of the implementation process of step S104 shown according to some embodiments of the present disclosure.

[0063] Figure 3 It is the implementation process flow of obtaining vector map features shown according to some embodiments of the present disclosure Figure 1 .

[0064] Figure 4 It is the implementation process flow of obtaining vector map features shown according to some embodiments of the present disclosure Figure 2 .

[0065] Figure 5 It is a flowchart of the implementation process of obtaining raster map features shown according to some embodiments of the present disclosure.

[0066] Figure 6 It is a flowchart of the implementation process of obtaining refined map features shown according to some embodiments of the present disclosure.

[0067] Figure 7 It is a flowchart of the implementation process of obtaining global map prediction data shown according to some embodiments of the present disclosure.

[0068] Figure 8 It is a flowchart of the implementation process of step S702 shown according to some embodiments of the present disclosure.

[0069] Figure 9 It is a schematic diagram of the network structure of a map data fusion model provided by a specific example shown according to some embodiments of the present disclosure.

[0070] Figure 10 It is a flowchart of a map data fusion method shown according to some embodiments of the present disclosure.

[0071] Figure 11 It is a block diagram of a model training device shown according to some embodiments of the present disclosure.

[0072] Figure 12 It is a block diagram of a map data fusion device shown according to some embodiments of the present disclosure.

[0073] Figure 13 is a block diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0074] Here, some embodiments of the present disclosure will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those set forth herein, but may be changed as will be apparent after understanding the present disclosure, except for operations that must be performed in a specific order. Additionally, descriptions of features known in the art may be omitted for the sake of clarity and conciseness.

[0075] The implementation manners described in some embodiments of the present disclosure below do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0076] The following will describe in detail the specific implementation manners of the embodiments of the present disclosure with reference to the accompanying drawings.

[0077] Figure 1 is a flowchart of a model training method shown according to some embodiments of the present disclosure. As Figure 1 shown, the model training method can be applied to various electronic devices, including but not limited to terminals such as smart phones, computers, and smart tablets, vehicles, as well as stand-alone physical servers, server clusters or distributed systems composed of multiple physical servers, and cloud servers and other types of servers.

[0078] Figure 1 The model training method shown includes the following steps.

[0079] In step S102, a training data set of multiple map data fusion training samples is obtained.

[0080] In some embodiments of the present disclosure, the map data fusion training sample includes at least two local vector maps and global map data, and the global map data is obtained by fusing at least two local vector maps. Each local vector map may be from a crowdsourcing platform, collecting map data provided by different users or obtained by multiple acquisitions. Multiple data acquisitions are performed on the same geographical area to obtain richer data. Different data sources can be fully utilized to improve the quality and coverage of the map data.

[0081] It should be noted that crowd sourcing refers to a model in which tasks are distributed to a large number of individuals via the Internet. In the field of map data, crowd sourcing platforms allow users to upload and share map data of their regions.

[0082] It should be noted that a local vector map refers to a set of vectorized map elements within a specific geographical area. These elements can include: lane lines, which represent the boundaries of driving lanes on a road; lane centerlines, which represent the centerlines of lanes; stop lines, which represent the positions where vehicles must stop; road signs, which indicate directions or provide other traffic information; and zebra crossings, which are safe passageways for pedestrians to cross the road.

[0083] In step S104, at least two local vector maps are extracted to obtain the category data and coordinate data of multiple map elements in the at least two local vector maps.

[0084] It should be noted that the category data of map elements refers to the categories to which the map elements belong, which can include lane lines, lane centerlines, stop lines, road signs, and zebra crossings. Different categories can be identified by numerical numbers. For example, 0 - 4 can be used to represent lane lines, lane centerlines, stop lines, road signs, and zebra crossings respectively. The coordinate data of map elements refers to the specific position information of each map element, which can be a 2D coordinate array of each map element.

[0085] In step S106, the initial map data fusion model is used to extract features from the category data and coordinate data to obtain vector map features; rasterization processing is performed based on the coordinate data to obtain raster map features; the vector map features and raster map features are subjected to feature fusion and feature refinement to obtain the refined map features; and decoding and prediction processing are performed on the refined map features to obtain global map prediction data.

[0086] In some embodiments of the present disclosure, the initial map data fusion model is a deep learning architecture based on the attention mechanism that can understand and process map data and better fuse vector map data. It can effectively integrate local vector map data from crowd sourcing platforms and multiple acquisitions, generate a high - precision and consistent global map, and meet the requirements of autonomous driving systems and other applications. Among them, the attention mechanism is a computational model that simulates the process of human visual attention allocation.

[0087] In some embodiments of the present disclosure, an initially constructed initial map data fusion model is obtained. The category data and coordinate data of multiple map elements in at least two local vector maps are used as the input data of the initial map data fusion model. Feature extraction is performed on the category data and coordinate data to obtain vector map features. Rasterization processing is performed based on the coordinate data to obtain raster map features. The vector map features and raster map features are subjected to feature fusion and feature refinement to obtain refined map features. Then, decoding and prediction processing are performed on the refined map features to obtain global map prediction data that fuses at least two local vector maps.

[0088] In step S108, the initial map data fusion model is trained based on the global map prediction data and the global map data to obtain a map data fusion model.

[0089] In some embodiments of the present disclosure, after obtaining the global map prediction data using the initial map data fusion model, the initial map data fusion model is trained based on the global map data in the obtained training dataset, that is, the global map ground truth that fuses at least two local vector maps. After the training is completed, a map data fusion model is obtained.

[0090] As can be seen from the above steps, the model training method provided by the embodiments of the present disclosure extracts the category data and coordinate data of multiple map elements from at least two local vector maps of the training samples. Feature extraction is performed on the category data and coordinate data through the initial map data fusion model to obtain vector map features. Rasterization processing is performed based on the coordinate data to obtain raster map features. The vector map features and raster map features are subjected to feature fusion and feature refinement to obtain refined map features. Decoding and prediction processing are performed on the refined map features to obtain global map prediction data. And the initial map data fusion model is trained based on the global map prediction data and the global map data in the training samples to obtain a map data fusion model, and a map data fusion model that can effectively fuse vector map data is obtained. Just extract the category data and coordinate data of map elements from the vector map data to be fused collected through crowdsourcing and input them into the trained map data fusion model, and an accurate fusion result can be obtained, and the fusion efficiency is high. Relying on the trained map data fusion model for fusion, even if the data sources are diverse or the quality is uneven, accurate fusion can still be achieved, improving the flexibility of the fused map data.

[0091] In some exemplary embodiments of the present disclosure, as Figure 2 shown, it is a flowchart of the implementation process of step S104 provided by some exemplary embodiments of the present disclosure, including the following steps.

[0092] In step S202, at least two local vector maps are extracted to obtain the initial category data and initial coordinate data of multiple map elements in the at least two local vector maps.

[0093] In some embodiments of the present disclosure, before extracting at least two local vector maps, the at least two local vector maps can be preprocessed, which may include: converting vector map data from different sources into a unified coordinate system; removing noise points and outliers to improve data quality; and converting the data into a unified format for subsequent processing.

[0094] It should be noted that different vector map data may be stored in different formats, such as GeoJSON, KML, Shapefile, etc., and corresponding libraries are needed to parse these files. Each map element in the at least two local vector maps usually contains a category label indicating its type (such as lane line, lane center line, stop line, etc.), and these labels can be extracted from the parsed data structure to obtain the initial category data. The geometric information of the map element is usually stored in geometric objects (such as points, lines, polygons), and these geometric information can be extracted from the parsed data structure and their coordinates can be obtained to get the initial coordinate data.

[0095] In step S204, the initial category data and initial coordinate data are dimensionally aligned to obtain the category data and coordinate data of multiple map elements in the at least two local vector maps.

[0096] In some exemplary embodiments of the present disclosure, since the number of map elements contained in each local vector map may be different, it may cause the arrays to be misaligned during subsequent processing. Therefore, it is necessary to dimensionally align the initial category data and initial coordinate data. After obtaining the initial category data and initial coordinate data, zeros can be padded at the end of the array when forming a batch for dimensional alignment to facilitate subsequent processing.

[0097] In some exemplary embodiments of the present disclosure, as Figure 3 shown, the implementation process flow for obtaining vector map features provided by some exemplary embodiments of the present disclosure Figure 1 , includes the following steps.

[0098] In step S302, feature extraction is performed on the category data to obtain category features.

[0099] In some embodiments of the present disclosure, the category data of multiple map elements in at least two local vector maps is input into an Embedding layer to extract category features with a preset number of channels, and convert the category data into a vector form suitable for subsequent neural network processing, where the preset number of channels is related to the number of categories of map elements. The embodiments of the present disclosure can effectively convert the category information in the vector map data into rich feature representations through the Embedding layer, providing strong support for subsequent map fusion and feature extraction.

[0100] It should be noted that the Embedding layer is a technique for mapping discrete categorical data (such as words, category labels, etc.) into continuous vector representations to achieve category embedding and semantic representation.

[0101] In step S304, feature extraction is performed on the coordinate data to obtain position features.

[0102] In some embodiments of the present disclosure, the coordinate data of multiple map elements in at least two local vector maps is input into the Subgraph module of VectorNet to aggregate information of map elements within a local range, capture their relative positions and geometric relationships, and can also capture spatial dependency relationships in a larger range step by step by constructing hierarchical feature representations, and obtain position features. The embodiments of the present disclosure can effectively process complex road topologies through the use of the Subgraph module, improving the effect of the map fusion task.

[0103] It should be noted that VectorNet is a neural network architecture specifically designed for processing vectorized map data (such as lane lines, stop lines, etc.). A core component of VectorNet is its Subgraph module, which is used to aggregate information within a local range and capture the relative relationships between map elements.

[0104] In step S306, the category features and the position features are concatenated to obtain vector map features.

[0105] In some embodiments of the present disclosure, the category features and the position features are concatenated together along the channel dimension to obtain vector map features. Specifically, the channel dimension can include various different feature representations, such as category embeddings (generated through the Embedding layer), position coordinates (generated through linear transformation or more complex encoding methods), etc.

[0106] It should be noted that the initial map data fusion model includes a vector map feature extraction module, which is used to implement the above steps S302 to S306 to obtain vector map features. It can effectively integrate different types of map features and provide rich input data for further map fusion and feature extraction.

[0107] Those skilled in the art can understand that the Figure 3 shown order of steps is only an example and is not used to limit the protection scope of the embodiments of the present disclosure. For example, step S302 can be implemented simultaneously with step S304, or can be implemented after or before step S304.

[0108] In some exemplary embodiments of the present disclosure, as Figure 4 shown, it is the implementation process flow of obtaining vector map features provided by some exemplary embodiments of the present disclosure Figure 2 , Figure 4 shown, steps S402 to S406 in obtaining vector map features correspond to steps S302 to S306 in Figure 3 shown obtaining vector map features and will not be repeated here.

[0109] On the basis of Figure 3 , the following steps may further be included.

[0110] In step S408, based on the coordinate data, calculate the central coordinates of each map element to obtain the 2D position information of multiple map elements.

[0111] In some embodiments of the present disclosure, the geometric type can be determined, such as a point, a line segment or a polygon; calculate the central coordinates according to the geometric type. For example, the central coordinate of a point is the coordinate of the point itself, the central coordinate of a line segment is the average value of all endpoint coordinates, and the central coordinate of a polygon can be approximately obtained by calculating the average value of its vertex coordinates, or a more complex centroid formula can be used.

[0112] It should be noted that the vector map feature extraction module is used to implement the above steps S402 to S408, so that on the basis of obtaining vector map features, the 2D position information of multiple map elements can also be determined. It can effectively extract and simplify the position information of map elements and provide support for further map fusion and feature extraction.

[0113] Those skilled in the art can understand that the Figure 4 shown order of steps is only an example and is not used to limit the protection scope of the embodiments of the present disclosure. For example, step S408 can be implemented simultaneously with steps S402 to S406, or can be implemented after or before them.

[0114] In some exemplary embodiments of the present disclosure, as Figure 5 shown, it is a flowchart of an implementation process for obtaining grid map features provided by some exemplary embodiments of the present disclosure, including the following steps.

[0115] In step S502, pixel coordinate data of the grid map is obtained based on coordinate data conversion to obtain a grid map image.

[0116] In some embodiments of the present disclosure, first, a predefined grid map image is determined, including the geographical range (i.e., the minimum and maximum longitude and latitude) of the represented map area and the desired grid map resolution (the actual distance represented by each pixel). The width and height of the grid map image can be calculated based on the geographical range and resolution. Then, the geographical coordinates (such as longitude and latitude) are converted into a planar coordinate system. Since the Earth is a sphere, directly displaying geographical coordinates on a two-dimensional plane will cause distortion. Therefore, projection transformation (such as Web Mercator, etc.) is usually used to convert geographical coordinates into planar coordinates. And the pixel coordinates corresponding to each geographical point are calculated according to the geographical range and resolution. Finally, based on the pixel coordinates corresponding to each geographical point, points or other graphic elements are drawn at the corresponding positions on the grid map, and the converted grid map image can be obtained.

[0117] In step S504, image features are extracted based on the grid map image to obtain grid map image features, and spatial position coordinate data corresponding to the grid map image features is generated.

[0118] In some embodiments of the present disclosure, a suitable algorithm is selected to extract features in the image. For example, the grid map image can be input into a deep convolutional neural network for image feature extraction to adapt to feature extraction in complex scenarios. The extracted features can also be flattened and permuted in size to obtain grid map image features. At the same time, the coordinate data of each spatial position corresponding to the grid map image features is calculated.

[0119] It should be noted that before feature extraction, some preprocessing operations are usually required on the grid map image to improve the accuracy and efficiency of feature extraction. Common preprocessing steps include: resizing and resolution adjustment to ensure that the image size is suitable for subsequent processing. Denoising, using filters to remove noise in the image, such as Gaussian blur. Contrast enhancement, enhancing the image contrast through techniques such as histogram equalization.

[0120] In step S506, grid map features are obtained according to the grid map image features and the spatial position coordinate data.

[0121] It should be noted that the raster map features include raster map image features and spatial position coordinate data. The initial map data fusion model includes a raster map feature extraction module, which is used to implement steps S502 to S506 to obtain raster map features for subsequent modules to use.

[0122] It should be noted that the vector map features contain the overall information of more map elements, while the raster map features contain more local information of map elements. The two are highly complementary, so they are fused through a splicing operation. Accordingly, as Figure 6 shown, it is a flowchart of the implementation process for obtaining the refined map features provided by some exemplary embodiments of the present disclosure, including the following steps.

[0123] In step S602, based on the vector map features and the raster map features, feature fusion is performed to obtain the initial map features.

[0124] In some embodiments of the present disclosure, the raster map image features in the vector map features and the raster map features are fused through a splicing operation to obtain the fused initial map features.

[0125] In step S604, based on the 2D position information and the raster map features, feature fusion is performed to obtain the 2D position encoding corresponding to the initial map features.

[0126] In some embodiments of the present disclosure, the 2D position information obtained in step S408 and the spatial position coordinate data in the raster map features are fused through a splicing operation to obtain the fused position information, and the fused position information is subjected to 2D sine position encoding to obtain the 2D position encoding corresponding to the initial map features.

[0127] In step S606, based on the 2D position encoding, the Transformer encoder is used to refine the initial map features to obtain the refined map features.

[0128] It should be noted that the initial map data fusion model includes a Transformer Encoder, and the Transformer Encoder includes multiple Transformer Encoder Layers. Based on the 2D position encoding corresponding to the initial map features obtained in step S604, multiple Transformer Encoder Layers are used to perform feature refinement on the initial map features, which consists of self-attention processing, residual connection, and layer normalization processing, to obtain the refined map features.

[0129] As Figure 7As shown, the flowchart of the implementation process for obtaining global map prediction data provided by some exemplary embodiments of the present disclosure includes the following steps.

[0130] In step S702, the refined map features are decoded using a Transformer decoder to obtain prediction results of multiple map elements.

[0131] In step S704, the prediction results of multiple map elements are classified and regressed using a prediction head to obtain global map prediction data.

[0132] It should be noted that the initial map data fusion model includes a Transformer Decoder module and a prediction head.

[0133] In some embodiments of the present disclosure, after obtaining the refined map features using a Transformer Encoder, the refined map features are decoded using a Transformer Decoder to obtain prediction results of multiple map elements. Then, the prediction results of multiple map elements are classified and regressed using a prediction head to obtain global map prediction data. The initial map data fusion model may include multiple prediction heads, and each prediction head receives the outputs of multiple Transformer Decoder Layers to perform map element prediction, obtaining multiple sets of prediction results.

[0134] In some embodiments of the present disclosure, the global map prediction data includes: class prediction data, confidence score prediction data, and coordinate prediction data corresponding to the map elements of the global map. That is, the class prediction data, confidence score prediction data, and coordinate prediction data of each map element in the global map. The confidence score prediction data represents the credibility of each map element belonging to a certain class. For example, there are 5 class prediction data for a certain map element, which are label 0, 1, 2, 3, and 4 respectively. The confidence score prediction data includes 5 credibilities when the labels are 0, 1, 2, 3, and 4, and the label with the highest credibility represents the class of this map element.

[0135] It should be noted that each set of prediction heads includes two classification heads and regression heads both composed of MLP (Multi-Layer Perceptron). The class prediction data and confidence score prediction data are sequentially obtained from the output of each Transformer Decoder Layer through an MLP and a Softmax (activation function) layer, and the coordinate prediction data is obtained from the output of each Transformer Decoder Layer through an MLP and a Sigmoid (activation function) layer.

[0136] In some exemplary embodiments of the present disclosure, such as Figure 8As shown, it is a flowchart of the implementation process of step S702 provided by some exemplary embodiments of the present disclosure, including the following steps.

[0137] In step S802, an initial content query vector is obtained.

[0138] In some embodiments of the present disclosure, the initial content query vector is a zero vector. The Transformer decoder includes multiple Transformer Encoder Layers. The initial content query vector is input into the first Transformer Encoder Layer, and after self-attention calculation, cross-attention calculation, and feed-forward neural network calculation, the content query vector refined in the first layer is output. The content query vector refined in the first layer is input into the second Transformer Encoder Layer, and after the same process, the content query vector refined in the second layer is output. And so on until all the Transformer Encoder Layers are processed.

[0139] In step S804, based on the refined map features, the core parameters of the attention mechanism of the cross-attention layer of the Transformer decoder are generated.

[0140] In some embodiments of the present disclosure, based on the refined map features obtained in step S606, generating the core parameters of the attention mechanism of the cross-attention layer of the Transformer decoder may include: K (key), V (value).

[0141] In step S806, based on the core parameters of the attention mechanism, the initial content query vector is refined to obtain the refined content query vector, which is used as the prediction result of multiple map elements.

[0142] It should be noted that the refined content query vector includes the content query vectors processed by each Transformer Encoder Layer.

[0143] In some exemplary embodiments of the present disclosure, when specifically implementing step S108, the global map prediction data, including: the category prediction data, confidence score prediction data, and coordinate prediction data of each map element in the global map, can be matched with the global map ground truth, including: the category data and coordinate data of each real map element in the global map, using the Hungarian matching method to determine positive and negative samples, and then calculate the loss function. Among them, the design of the loss can adopt the loss function of Map TR, specifically including: classification loss, which is used to optimize the prediction of the model for the category of map elements. Regression loss, which is used to optimize the prediction of the model for the position coordinates of map elements. The total loss is usually the weighted sum of the classification loss and the regression loss to balance the influence of both on the final loss, which helps the model to improve the performance of both classification and regression tasks during training, so as to achieve more accurate map element detection and segmentation.

[0144] It should be noted that Map TR is a method for object detection and segmentation in a Bird's-Eye View (BEV) map.

[0145] It can be seen that the embodiments of the present disclosure design and train a map data fusion model based on a deep learning framework, and use a neural network model to automatically fuse vector map data from multiple crowdsourcing platforms or multiple acquisitions, so as to generate high-precision and consistent global map data with the trained map data fusion model.

[0146] To better illustrate the model training method of the map data fusion model provided by the embodiments of the present disclosure, a specific example is provided below for further explanation. The specific example provided is the specific model architecture and training process of applying the model training method provided by the examples of the present disclosure.

[0147] As Figure 9 shown, it is the network structure of the map data fusion model constructed in this specific example. The network input is multiple local vector maps collected and generated through crowdsourcing , where is the number of map elements included in each local vector map, is the category array of all map elements, is the 2D coordinate array of all map elements, represents the set of real numbers. Each map element is interpolated into points, and there are 5 types of map element categories, as shown in Table 1. Since the number of map elements in each input sample is different, when forming a batch, zero-padding is required at the end of the array for dimension alignment, and the final data input to the network is the map element category and coordinates ( is the batch size, ).

[0148]

[0149] Figure 9 The network output shown is the category of the fused map elements , confidence score and coordinates , where is the set maximum number of map element predictions.

[0150] The network includes five modules: a vector map feature extraction module; a raster map feature extraction module; a Transformer Encoder; a Transformer Decoder; and a prediction head.

[0151] Vector map feature extraction module: First, the map element category is fed into an Embedding layer to extract category features with a channel number of , where represents the preset number of channels. At the same time, the coordinates are input into the Subgraph module of VectorNet to extract location features , and then the two are concatenated along the channel dimension ( ), obtaining the vector map feature . At the same time, the center coordinates of each map element are calculated as the 2D position information of the map element , and the calculation formula of the entire module is shown in formula (1) below: (1)

[0152] (1)

[0153] where represents an operation specified in dimension 2 (the third dimension, that is, the channel dimension). represents taking the mean value.

[0154] Raster map feature extraction module: First, all map element coordinates are converted into pixel coordinates in the map image according to formula (2), where is the coordinate range of the map image , and are the pre-defined width and height of the raster map image, which are set values. Then, the map image with size is calculated according to formula (3), and the five map element categories respectively correspond to the map image The five image channels. Then the map image is input into the VGG19 network (a deep convolutional neural network) to extract image features with a size of . Finally, through size flattening and permutation operations, features are obtained, where . At the same time, a feature map is generated (where represents the height of the feature map, represents the width of the feature map), and the coordinates of each spatial position in the feature map are obtained by calculating the coordinates of each point in each grid of the feature map for use in subsequent modules.

[0155] (2)

[0156] (3)

[0157] Among them, represents the coordinates of each map element in the vector map . represents the map image in the coordinates of each map element, represents the class label of the map element, taking values from 0 to 4, represents the pixel coordinates of the map element.

[0158] Transformer Encoder: The vector map features contain more overall information of map elements, while the raster map features contain more local information of map elements. The two are highly complementary. Therefore, the two are fused through a concatenation operation to obtain the initial map features after fusion and position information . Then, through 2D sine position encoding ( ), the 2D position encoding corresponding to the initial map features is obtained. The calculation process is shown in formula (4).

[0159] (4)

[0160] Among them, represents specifying an operation in dimension 1 (the second dimension).

[0161] The fused map features will be input into the Encoder for feature refinement. The Transformer Encoder includes Transformer Encoder Layers. For the One Transformer Encoder Layer, with input features First, it passes through a Self-Attention module, where the features are added to the positional encoding to obtain Q1 (query) and K1 (key), which are directly used as (value). The output of the self-attention module goes through residual connection and LN (Layer Normalization) layer to obtain the updated features , and Equation (5) shows the calculation process of the self-attention module.

[0162] (5)

[0163] And the features then successively pass through the FFN (Feed-Forward Neural Network) layer, residual connection, and LN layer to obtain the refined map features , as shown in Equation (6).

[0164] (6)

[0165] Among them, takes values from 1 to .

[0166] After being processed by Transformer Encoder Layers, the refined map features are obtained.

[0167] Transformer Decoder: Input to the Transformer Decoder. At the same time, the Transformer Decoder also inputs a Content Query (content query) of all zeros and a learnable Position Query (position query) .

[0168] The Transformer Decoder includes Transformer Decoder Layers, and each layer contains three modules: Self-Attention, Cross-Attention, and FFN (feed-forward neural network).

[0169] For the One Transformer Decoder Layer, content query First, the self-attention is calculated through the Self-Attention module, where (query) and (key) are obtained by adding with the position query and is . The output of the self-attention module then passes through a residual connection and an LN layer to obtain the updated content query . The calculation process is shown in formula (7). The value of ranges from 1 to

[0170] (7)

[0171] The updated content query is then input into the Cross-Attention module to interact with the map features. First, the corresponding position encoding is calculated according to formula (8). If , is predicted sequentially through an MLP (Multi-Layer Perceptron), a Sigmoid (activation function), and a 2D sine position encoding from the position query . Otherwise, is obtained from the center point of the map element coordinates predicted by the output of the previous Transformer Decoder Layer through a 2D sine position encoding .

[0172] (8)

[0173] The (query) of the Cross-Attention module is obtained by adding the content query with the position encoding . The (key) is obtained by adding the refined map features with the 2D position encoding corresponding to the initial map features . The 3 (value) is the refined map features . Both contain content information and position information, which is beneficial to the accurate positioning of map elements and the rapid convergence of the model. Similar to the self-attention module, the output of the cross-attention module also passes through a residual connection and an LN layer to generate the updated content query , the calculation process is shown in formula (9).

[0174] (9)

[0175] Content query Then, it passes through the FFN layer, residual connection, and LN layer in sequence to obtain the refined content query , as shown in formula (10).

[0176] (10)

[0177] That is to say, for each Transformer Decoder Layer, the input After being processed by formulas (7), (8), (9), and (10), the output . The prediction results of multiple map elements are .

[0178] Prediction head: This network architecture contains groups of prediction heads, which respectively receive the outputs of TransformerDecoder Layers to perform map element prediction and obtain groups of prediction results.

[0179] Each group of prediction heads includes two classification heads and regression heads both composed of MLP. The prediction results are obtained by passing through the MLP and layers in sequence. The position prediction results are obtained by passing through the MLP and Sigmoid layers in sequence.

[0180] As shown in formula (11).

[0181] (11)

[0182] During the model training stage, the model prediction results will be matched with the map ground truth using the Hungarian algorithm to determine positive and negative samples, and then calculate the loss. The same loss function as Map TR is adopted.

[0183] During the model inference stage, that is, when applying the trained map data fusion model for vector map data fusion, only the prediction results of the last layer are output as the model output.

[0184] As can be seen from the above process, in this specific example, by constructing and training a map data fusion model, using a neural network and a deep learning architecture, the vector map data from multiple crowdsourcing platforms or multiple acquisitions is automatically fused to generate high-precision and consistent map data. Compared with the rule-based fusion method in the prior art, it has the following significant advantages:

[0185] Higher accuracy: By extracting and fusing the deep features of the data through a deep learning model, it can more accurately reflect the geometric shape, topological relationship and semantic information of the map, improving the accuracy of the fusion result.

[0186] Stronger robustness: The deep learning model can process data from different sources and of different qualities, and has strong robustness to data noise and anomalies, and can generate more reliable fusion results.

[0187] Better adaptability: The deep learning model can automatically adjust the fusion strategy according to the characteristics of the data, adapt to the data fusion requirements in different scenarios, and improve the versatility and flexibility of the method.

[0188] Higher efficiency: Through parallel computing and an optimized deep learning framework, it can significantly improve the computational efficiency of data fusion and meet the real-time requirements of the autonomous driving system.

[0189] It should be noted that in the technical solution of the present disclosure, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.

[0190] The present disclosure embodiment also provides a map data fusion method, Figure 10 which is a flowchart of a map data fusion shown according to some embodiments of the present disclosure. As Figure 10 shown, the map data fusion method can be applied to various electronic devices, including but not limited to terminal devices such as smart phones, computers, smart tablets, vehicles, and independent physical servers, server clusters or distributed systems composed of multiple physical servers, and cloud servers and other types of servers. Figure 10 The map data fusion method shown includes the following steps.

[0191] In step S1002, multiple local vector maps to be fused are obtained.

[0192] It should be noted that each local vector map to be fused can be from a crowdsourcing platform, collecting map data provided by different users or obtained through multiple acquisitions, and performing multiple data acquisitions on the same geographical area to obtain richer data.

[0193] In step S1004, multiple local vector maps to be fused are extracted to obtain the category data and coordinate data of multiple map elements in the multiple local vector maps to be fused.

[0194] In some embodiments of the present disclosure, the category data and coordinate data of multiple map elements in multiple local vector maps to be fused can be obtained by performing preprocessing, parsing and extraction, dimension alignment, etc. on the multiple local vector maps to be fused.

[0195] In step S1006, the category data and coordinate data of multiple map elements in multiple local vector maps to be fused are input into the map data fusion model, and the global map data obtained by fusing the multiple local vector maps to be fused is output.

[0196] It should be noted that the map data fusion model is trained according to the model training method described in the above embodiments.

[0197] In the embodiments of the present disclosure, the category data and coordinate data of multiple map elements in multiple local vector maps to be fused are fused by applying the map data fusion model constructed and trained based on the deep learning architecture in the above embodiments, and the global map data obtained by fusing the multiple local vector maps to be fused is obtained. It can achieve efficient and accurate fusion of crowdsourced vector map data, provide real-time and reliable map support for the autonomous driving system, and improve the safety and reliability of the autonomous driving system.

[0198] It can be understood that the same / similar parts among the various embodiments of the above methods in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments. For the relevant parts, refer to the descriptions of other method embodiments.

[0199] The following are the device embodiments of the present disclosure, which can be used to execute the method embodiments of the present disclosure. For the details not disclosed in the device embodiments of the present disclosure, refer to the method embodiments of the present disclosure.

[0200] Figure 11 It is a block diagram of a model training device shown according to some embodiments of the present disclosure. Referring to Figure 11 , the device includes: a sample acquisition unit 1101, a sample data extraction unit 1102, and a sample training unit 1103.

[0201] The sample acquisition unit 1101 is configured to acquire a training data set of multiple map data fusion training samples; the map data fusion training samples include at least two local vector maps and global map data; the global map data is obtained by fusing at least two local vector maps;

[0202] The sample data extraction unit 1102 is configured to extract at least two local vector maps, and obtain the category data and coordinate data of multiple map elements in the at least two local vector maps;

[0203] The sample training unit 1103 is configured to perform feature extraction on the category data and coordinate data by using an initial map data fusion model to obtain vector map features; perform rasterization processing based on the coordinate data to obtain raster map features; perform feature fusion and feature refinement on the vector map features and the raster map features to obtain refined map features; perform decoding and prediction processing on the refined map features to obtain global map prediction data; and train the initial map data fusion model based on the global map prediction data and the global map data to obtain a map data fusion model.

[0204] In some exemplary embodiments of the present disclosure, the sample data extraction unit 1102 is configured to:

[0205] Extract at least two local vector maps, and obtain the initial category data and initial coordinate data of multiple map elements in the at least two local vector maps;

[0206] Perform dimension alignment on the initial category data and the initial coordinate data to obtain the category data and coordinate data of multiple map elements in the at least two local vector maps.

[0207] In some exemplary embodiments of the present disclosure, the sample training unit 1103 is configured to:

[0208] Perform feature extraction on the category data to obtain category features;

[0209] Perform feature extraction on the coordinate data to obtain position features;

[0210] Concatenate the category features and the position features to obtain vector map features.

[0211] In some exemplary embodiments of the present disclosure, the sample training unit 1103 is further configured to:

[0212] Based on the coordinate data, calculate the center coordinates of each map element to obtain the 2D position information of multiple map elements.

[0213] Correspondingly, the sample training unit 1103 is configured to:

[0214] Perform feature fusion based on the vector map features and the raster map features to obtain initial map features;

[0215] Perform feature fusion based on the 2D position information and the raster map features to obtain a 2D position encoding corresponding to the initial map features;

[0216] Based on 2D position encoding, use a Transformer encoder to refine the initial map features to obtain the refined map features;

[0217] Among them, the initial map data fusion model includes a Transformer encoder.

[0218] In some exemplary embodiments of the present disclosure, the sample training unit 1103 is configured to:

[0219] Based on the coordinate data conversion, obtain the pixel coordinate data of the grid map to obtain the grid map image;

[0220] Extract image features from the grid map image to obtain grid map image features, and generate spatial position coordinate data corresponding to the grid map image features;

[0221] According to the grid map image features and the spatial position coordinate data, obtain the grid map features.

[0222] In some exemplary embodiments of the present disclosure, the sample training unit 1103 is configured to:

[0223] Use a Transformer decoder to decode the refined map features to obtain prediction results of multiple map elements;

[0224] Use a prediction head to classify and regress the prediction results of multiple map elements to obtain global map prediction data;

[0225] Among them, the initial map data fusion model includes a Transformer decoder and a prediction head.

[0226] In some embodiments of the present disclosure, the global map prediction data includes: class prediction data, confidence score prediction data, and coordinate prediction data corresponding to the map elements of the global map.

[0227] Further, in some exemplary embodiments of the present disclosure, the sample training unit 1103 is configured to:

[0228] Obtain an initial content query vector; the initial content query vector is a vector of all zeros;

[0229] Based on the refined map features, generate the core parameters of the attention mechanism of the cross-attention layer of the Transformer decoder;

[0230] Based on the core parameters of the attention mechanism, perform refinement processing on the initial content query vector to obtain the refined content query vector as the prediction results of multiple map elements.

[0231] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0232] Figure 12 is a block diagram of map data fusion shown according to some embodiments of the present disclosure. Referring to Figure 12 , the device includes: a vector map acquisition unit 1201, a map data extraction unit 1202, and a map fusion unit 1203.

[0233] The vector map acquisition unit 1201 is configured to acquire multiple local vector maps to be fused;

[0234] The map data extraction unit 1202 is configured to extract multiple local vector maps to be fused to obtain category data and coordinate data of multiple map elements in the multiple local vector maps to be fused;

[0235] The map fusion unit 1203 is configured to input the category data and coordinate data of multiple map elements in the multiple local vector maps to be fused into a map data fusion model, and output global map data that fuses the multiple local vector maps to be fused;

[0236] Among them, the map data fusion model is trained according to the model training method described in the above embodiments.

[0237] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0238] Figure 13 is a block diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. For example, the electronic device 1300 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc., or may also be a server.

[0239] Referring to Figure 13 , the electronic device 1300 may include one or more of the following components: a processing component 1302, a memory 1304, a power component 1306, a multimedia component 1308, an audio component 1310, an input / output (I / O) interface 1312, a sensor component 1314, and a communication component 1316.

[0240] The processing component 1302 generally controls the overall operation of the electronic device 1300, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1302 may include one or more processors 1320 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 1302 may include one or more modules to facilitate the interaction between the processing component 1302 and other components. For example, the processing component 1302 may include a multimedia module to facilitate the interaction between the multimedia component 1308 and the processing component 1302.

[0241] The memory 1304 is configured to store various types of data to support the operation of the device 1300. Examples of such data include instructions for any application or method operating on the electronic device 1300, contact data, phone book data, messages, pictures, videos, and the like. The memory 1304 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0242] The power component 1306 provides power to various components of the electronic device 1300. The power component 1306 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 1300.

[0243] The multimedia component 1308 includes a screen that provides an output interface between the electronic device 1300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 1308 includes a front camera and / or a rear camera. When the device 1300 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0244] The audio component 1310 is configured to output and / or input audio signals. For example, the audio component 1310 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 1300 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1304 or transmitted via the communication component 1316. In some embodiments, the audio component 1310 further includes a speaker for outputting audio signals.

[0245] The I / O interface 1312 provides an interface between the processing component 1302 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0246] The sensor component 1314 includes one or more sensors for providing an assessment of various aspects of the status of the electronic device 1300. For example, the sensor component 1314 can detect the on / off state of the device 1300, the relative positioning of components, such as the display and keypad of the electronic device 1300. The sensor component 1314 can also detect a change in the position of the electronic device 1300 or a component of the electronic device 1300, the presence or absence of user contact with the electronic device 1300, the orientation or acceleration / deceleration of the electronic device 1300, and a change in the temperature of the electronic device 1300. The sensor component 1314 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 1314 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1314 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0247] The communication component 1316 is configured to facilitate communication between the electronic device 1300 and other devices in a wired or wireless manner. The electronic device 1300 can access a wireless network based on communication standards, such as Wi-Fi, 3G, 4G, 5G, other communication standards, or a combination thereof. In some embodiments of the present disclosure, the communication component 1316 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of the present disclosure, the communication component 1316 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0248] In some embodiments of the present disclosure, the electronic device 1300 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to execute the above method.

[0249] In some embodiments of the present disclosure, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1304 including instructions, and the above instructions can be executed by a processor 1320 of the electronic device 1300 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0250] In some embodiments of the present disclosure, a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to execute the above model training method or map data fusion method.

[0251] In some embodiments of the present disclosure, a computer program product is also provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the above model training method or map data fusion method is implemented.

[0252] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0253] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A model training method, characterized in that, Comprising: Obtaining a training dataset of multiple map data fusion training samples; the map data fusion training samples include at least two local vector maps and global map data; The global map data is obtained by fusing the at least two local vector maps; Extracting the at least two local vector maps to obtain class data and coordinate data of multiple map elements in the at least two local vector maps; Using an initial map data fusion model to perform feature extraction on the class data and the coordinate data to obtain vector map features; Performing rasterization processing based on the coordinate data to obtain raster map features; Performing feature fusion and feature refinement on the vector map features and the raster map features to obtain refined map features; Performing decoding and prediction processing on the refined map features to obtain global map prediction data; Training the initial map data fusion model based on the global map prediction data and the global map data to obtain a map data fusion model; Wherein, performing feature fusion and feature refinement on the vector map features and the raster map features to obtain refined map features includes: Performing feature fusion based on the vector map features and the raster map features to obtain initial map features; Performing feature fusion based on the 2D position information of multiple map elements and the raster map features to obtain a 2D position encoding corresponding to the initial map features; Based on the 2D position encoding, using a Transformer encoder to perform feature refinement on the initial map features to obtain the refined map features; Wherein, the 2D position information of the multiple map elements is calculated based on the coordinate data.

2. The model training method according to claim 1, wherein Performing feature extraction on the class data and the coordinate data to obtain vector map features includes: Performing feature extraction on the class data to obtain class features; Performing feature extraction on the coordinate data to obtain position features; Concatenating the class features and the position features to obtain the vector map features.

3. The model training method according to claim 2, wherein Further comprising: Based on the coordinate data, calculating the central coordinates of each map element to obtain the 2D position information of the multiple map elements.

4. The model training method according to claim 1, wherein The initial map data fusion model includes the Transformer encoder.

5. The model training method according to claim 1, characterized in that Performing rasterization processing based on the coordinate data to obtain raster map features includes: Based on the coordinate data, converting to obtain pixel coordinate data of a raster map to obtain a raster map image; Extracting image features based on the raster map image to obtain raster map image features, and generating spatial position coordinate data corresponding to the raster map image features; According to the raster map image features and the spatial position coordinate data, obtaining the raster map features.

6. The model training method according to claim 1, wherein Performing decoding and prediction processing on the refined map features to obtain global map prediction data includes: Using a Transformer decoder to decode the refined map features to obtain prediction results of the multiple map elements; Using a prediction head to classify and regress the prediction results of the multiple map elements to obtain the global map prediction data; Among them, the initial map data fusion model includes the Transformer decoder and the prediction head.

7. The model training method according to claim 6, wherein The Transformer decoder is used to decode the refined map features to obtain prediction results of the multiple map elements, including: Obtain an initial content query vector; the initial content query vector is a vector of all zeros; Based on the refined map features, generate the core parameters of the attention mechanism of the mutual attention layer of the Transformer decoder; Based on the core parameters of the attention mechanism, perform refinement processing on the initial content query vector to obtain a refined content query vector as the prediction results of the multiple map elements.

8. The model training method according to claim 1, wherein The global map prediction data includes: The category prediction data, confidence score prediction data, and coordinate prediction data corresponding to the map elements of the global map.

9. The model training method according to claim 1, wherein Extract the at least two local vector maps to obtain the category data and coordinate data of multiple map elements in the at least two local vector maps, including: Extract the at least two local vector maps to obtain the initial category data and initial coordinate data of multiple map elements in the at least two local vector maps; Perform dimension alignment on the initial category data and the initial coordinate data to obtain the category data and coordinate data of multiple map elements in the at least two local vector maps.

10. A method for map data fusion, characterized in that, Include: Obtain multiple local vector maps to be fused; Extract the at least two local vector maps to be fused to obtain the category data and coordinate data of multiple map elements in the at least two local vector maps to be fused; Input the category data and coordinate data of multiple map elements in the at least two local vector maps to be fused into the map data fusion model, and output the global map data fusing the at least two local vector maps to be fused; Among them, the map data fusion model is trained according to the model training method described in any one of claims 1 to 9.

11. A model training device, characterized in that, Include: A sample acquisition unit for acquiring a training data set of multiple map data fusion training samples; the map data fusion training samples include at least two local vector maps and global map data; The global map data is obtained by fusing the at least two local vector maps; A sample data extraction unit for extracting the at least two local vector maps to obtain the category data and coordinate data of multiple map elements in the at least two local vector maps; A sample training unit for using the initial map data fusion model to extract features from the category data and the coordinate data to obtain vector map features; Perform rasterization processing based on the coordinate data to obtain raster map features; Perform feature fusion and feature refinement on the vector map features and the raster map features to obtain refined map features; Perform decoding and prediction processing on the refined map features to obtain global map prediction data; train the initial map data fusion model based on the global map prediction data and the global map data to obtain a map data fusion model; Among them, performing feature fusion and feature refinement on the vector map features and the raster map features to obtain refined map features includes: Performing feature fusion on the vector map features and the raster map features to obtain initial map features; Performing feature fusion on the 2D position information of multiple map elements and the raster map features to obtain the 2D position encoding corresponding to the initial map features; Based on the 2D position encoding, using a Transformer encoder to perform feature refinement on the initial map features to obtain the refined map features; Among them, the 2D position information of the multiple map elements is calculated based on the coordinate data.

12. A map data fusion device, characterized in that, Including: A vector map acquisition unit for acquiring multiple local vector maps to be fused; A map data extraction unit for extracting the multiple local vector maps to be fused to obtain the category data and coordinate data of multiple map elements in the multiple local vector maps to be fused; A map fusion unit for inputting the category data and coordinate data of multiple map elements in the multiple local vector maps to be fused into a map data fusion model and outputting global map data that fuses the multiple local vector maps to be fused; Among them, the map data fusion model is trained according to the model training method described in any one of claims 1 to 9.

13. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Among them, the processor is configured to: implement the model training method described in any one of claims 1 to 9, or implement the map data fusion method described in claim 10.

14. A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enabling the mobile terminal to execute the model training method described in any one of claims 1 to 9, or implement the map data fusion method described in claim 10.

Citation Information

Patent Citations

  • Trajectory prediction method and device

    CN118351342A

  • Local continuity vector map alignment method, system and device and storage medium

    CN118376225A