Feature extraction model training method and device, task model training method and device and medium
Through the feature extraction model training method based on high-precision map data, the problem of long research and development cycle and high operation and maintenance costs in the existing technology is solved, and more efficient feature extraction and task model training is achieved, reducing operation and maintenance costs.
Patent Information
- Application Number
- CN202311659804.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology builds a specific task model for each map processing task, resulting in a long R&D cycle and high operation and maintenance costs.
Through the feature extraction model training method based on high-precision map data, the map feature vector is generated and classified according to preset tasks, and used to train the map processing model until the model convergence condition is reached.
The context understanding and representation ability of the feature extraction model is improved, the traffic semantic information in different scenarios is identified, the information content and adaptability of the feature vector is enhanced, the construction and training time of the task model is reduced, and the operation and maintenance cost is reduced.
Smart Images

Figure CN120107923A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of high-precision map technology, and in particular to a method, device and medium for training a feature extraction model and a task model. Background Art
[0002] High-precision maps are electronic maps with higher location accuracy and richer geographic elements. Their high-precision road information and rich traffic semantic information play an important supporting role in the environmental perception and planning decisions of intelligent driving vehicles.
[0003] At present, when implementing specific map processing tasks (such as the generation of specific features and the identification of feature associations) based on basic high-precision map data, the main method is to build and train specific task models for each map processing task, and design refined manual intervention strategies for it. However, when there are many map processing tasks, the existing method of building specific tasks for each task will result in a long task development cycle and high operation and maintenance costs. Summary of the invention
[0004] In order to solve the above-mentioned technical problems of long R&D cycle and high operation and maintenance cost caused by building a specific task for each map processing task, the present disclosure provides a feature extraction model, a training method, a device and a medium for a task model.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for training a feature extraction model based on map data, comprising:
[0006] Generate map element samples of different types of map elements in the sample area based on high-precision map data;
[0007] The map feature samples are used as inputs of the feature extraction model, and the feature extraction model generates a map feature feature vector;
[0008] Classify the map element feature vectors according to the preset map processing tasks to obtain sample feature vectors of the corresponding map processing tasks;
[0009] Using the sample feature vector of the corresponding map processing task as the input of the map processing model of the map processing task, training the map processing model, and generating a model output result;
[0010] Based on the model output results, the feature extraction model and map processing model are iteratively trained until the model convergence conditions are met.
[0011] In a second aspect, the present disclosure also provides a method for training a task model based on map data, including:
[0012] Obtaining map feature samples in a sample area corresponding to a map processing task;
[0013] The map element sample is used as an input of the feature extraction model, and the feature extraction model generates a map element feature vector; wherein the feature extraction model is pre-trained by the training method of the feature extraction model based on map data as described above;
[0014] The map element feature vector is used as the input of the map processing model corresponding to the map processing task, and the map processing model generates the model output result;
[0015] Based on the model output results, the feature extraction model and map processing model are iteratively trained until the model convergence conditions are met.
[0016] In a third aspect, the present disclosure also provides a training device for a feature extraction model based on map data, comprising:
[0017] A sample generating unit, for generating map element samples of different types of map elements in a sample area based on high-precision map data;
[0018] A map element feature vector generating unit, used for taking a map element sample as an input of a feature extraction model, and generating a map element feature vector by the feature extraction model;
[0019] A sample feature vector determination unit is used to classify the map element feature vectors according to the preset map processing tasks to obtain the sample feature vectors of the corresponding map processing tasks;
[0020] A map processing model training unit, used to take the sample feature vector of the corresponding map processing task as the input of the map processing model of the map processing task, train the map processing model, and generate a model output result;
[0021] The feature extraction model training unit is used to iteratively train the feature extraction model and the map processing model based on the model output results until the model convergence conditions are met.
[0022] In a fourth aspect, the present disclosure also provides a training device for a task model based on map data, including:
[0023] A sample acquisition unit, used to acquire map element samples in a sample area corresponding to a map processing task;
[0024] An element feature vector generating unit, used to take a map element sample as an input of a feature extraction model, and generate a map element feature vector by the feature extraction model; wherein the feature extraction model is pre-trained by the above-mentioned training method of the feature extraction model based on map data;
[0025] A map processing model training unit, used for taking the map element feature vector as the input of the map processing model corresponding to the map processing task, and generating a model output result by the map processing model;
[0026] The task model training unit is used to iteratively train the feature extraction model and the map processing model based on the model output results until the model convergence conditions are met.
[0027] In a fifth aspect, an embodiment of the present disclosure further provides an electronic device, including:
[0028] A memory and a processor, the memory is used to store instructions executable by the processor;
[0029] A processor is used to read executable instructions from a memory and execute the executable instructions to implement a training method for a feature extraction model based on map data or a training method for a task model based on map data provided in any embodiment of the present disclosure.
[0030] In a sixth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the training method of a feature extraction model based on map data or the training method of a task model based on map data provided in any embodiment of the present disclosure is implemented.
[0031] In a seventh aspect, the embodiments of the present disclosure further provide a computer program product, which is used to execute the training method of a feature extraction model based on map data or the training method of a task model based on map data provided by any embodiment of the present disclosure.
[0032] The technical solution for training a feature extraction model based on map data provided by the disclosed embodiment uses basic high-precision map data as training samples to iteratively train the feature extraction model and the map processing model of the map processing task, so that the feature extraction model can learn the ability to extract more feature information related to the map processing task on the basis of the original feature extraction capability, so that the feature extraction model can have a stronger context understanding and representation capability, identify traffic semantic information in different scenarios, improve the information content and information richness of the feature vector output by the feature extraction model, and then improve the adaptability of the feature extraction model to different map processing tasks, providing a model basis for the subsequent construction and training of task models to reduce different map processing tasks.
[0033] The technical solution for training a task model based on map data provided by the embodiment of the present disclosure applies the feature extraction model obtained by pre-training to the training process of different map processing tasks. Since the feature extraction model already has the ability to provide more feature information, the feature vector output by it can provide the information required for the map processing task to a large extent, so that the task model corresponding to the map processing task can achieve model convergence after a small amount of training, thus avoiding manual strategy intervention in the process to a large extent, improving the task model training efficiency and model accuracy of downstream map processing tasks, thereby shortening the R&D cycle and reducing operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0035] Figure 1 A schematic diagram of a functional module based on map data provided by an embodiment of the present disclosure;
[0036] Figure 2 A flowchart of a method for training a feature extraction model based on map data provided by an embodiment of the present disclosure;
[0037] Figure 3 A schematic diagram of the structure of a high-precision map model provided in an embodiment of the present disclosure;
[0038] Figure 4 for Figure 2 A schematic diagram of a detailed process of S210 in a training method of a feature extraction model based on map data is shown;
[0039] Figure 5 for Figure 2 The diagram shows a detailed flow chart of S230 to S240 in the training method of the feature extraction model based on map data for the binary element prediction task;
[0040] Figure 6 for Figure 2 The diagram shows a detailed flow chart of S230 to S240 for the feature generation task in the training method of the feature extraction model based on map data;
[0041] Figure 7 A flowchart of a method for training a task model based on map data provided in an embodiment of the present disclosure;
[0042] Figure 8A schematic diagram of the structure of a training device for a feature extraction model based on map data provided by an embodiment of the present disclosure;
[0043] Fig. 9 A schematic diagram of the structure of a training device for a task model based on map data provided by an embodiment of the present disclosure;
[0044] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0045] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0046] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0047] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0048] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0049] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0050] Before describing in detail a method for training a feature extraction model based on map data provided by an embodiment of the present disclosure, the terms involved are described in detail, wherein:
[0051] Lane: An area on the road where vehicles can travel. A lane is represented by two lane lines on the left and right. Lanes have traffic attributes, such as: going straight, turning left, turning right, and U-turn.
[0052] Road: It is composed of lanes and can be expressed by lane groups, that is, a road consists of a group of lanes running side by side.
[0053] Stop line: Usually located at intersections, it is used to prompt drivers to stop and wait for passage. The stop lines in the map data include real stop lines in the real world and virtual stop lines. Virtual stop lines refer to stop lines that are not drawn on the actual road, but are specially added to the map data so that smart driving vehicles can stop at locations where there are no stop lines in the real world but need to stop.
[0054] Map elements: A general term for elements recorded in map data, including lane lines, roads, stop lines, traffic lights, etc.; the speed limit of a lane is an attribute of the lane element.
[0055] Map feature relationships: The relationships between features in the real world, including whether lane lines are in the same road group, whether roads are topologically connected, whether traffic signs are associated with roads, and the specific associated roads.
[0056] Map vector: All elements in map data are expressed and stored in the form of map vectors. The main types are: LineString expresses linear objects, such as lane lines; Polygon expresses surface objects, such as pedestrian safety islands. Both linestrings and polygons are essentially sequences of spatial points.
[0057] PointNet: A general point set function learner based on neural networks. It is usually used to process unordered point sets, such as point cloud data, and encode them into a feature vector for use in downstream tasks.
[0058] Transformer: A neural network module commonly used in sequence analysis tasks such as natural language processing. It includes an encoder and a decoder. Its core is the self-attention mechanism, which calculates the correlation between input elements through the characteristics of the input elements themselves.
[0059] Multilayer perceptron (MLP): A neural network module that is often used as a classifier. It essentially repeats linear transformation and nonlinear activation function operations.
[0060] In the related art, the production process of production tasks usually requires the construction of a new set of models and strategies for each new production task, that is, to build a unique model for each production task. Since the model needs to be retrained from scratch and the training task is single, the model development cycle is long and the model effect is poor, and manual strategies are required to intervene to improve the model effect. However, manual strategies need to be carefully designed for different production tasks, and the maintenance cost is high. In addition, the models and strategies in a single production task cannot be flexibly applied to other production tasks, that is, the model built for a single production task is not suitable for other production tasks.
[0061] Based on the above situation, the embodiment of the present disclosure provides a training method for a feature extraction model based on map data. In the pre-training stage, multiple training tasks are used to enable the model to learn traffic semantic information in multiple scenarios and establish associations between different map elements to improve the efficiency and effectiveness of downstream map processing tasks.
[0062] Figure 1 A schematic diagram of a functional module based on map data provided for an embodiment of the present disclosure includes a sample generation functional module, a high-precision data model architecture functional module, a pre-training task functional module and a task application functional module, wherein the sample generation functional module is used to divide the production task into multiple areas with a fixed-size pane, and then generate samples through feature vector difference and coordinate system transformation, one area corresponding to one sample; the high-precision data model architecture functional module includes a geometric encoder, a type encoder and an encoder based on a self-attention mechanism, which are used to extract feature information in the sample; the pre-training task functional module is used to customize pre-training task schemes, for example, univariate feature prediction tasks, binary feature prediction tasks and feature generation tasks; the task application functional module is used to provide various specific production task models that the pre-training model can quickly migrate, making full use of the scene semantic understanding ability of the pre-training model, which can greatly improve the development efficiency of the task model.
[0063] Figure 2 A flowchart of a method for training a feature extraction model based on map data provided in an embodiment of the present disclosure is provided, which can be applied to extracting general traffic semantic information in various production task scenarios based on map vector data. The method for training a feature extraction model based on map data can be performed by a pre-training device for a feature extraction model based on high-precision map data, which can be implemented in software and / or hardware and can be integrated on an electronic device with certain computing capabilities, such as a laptop computer, a desktop computer, a server, etc.
[0064] like Figure 2 As shown, a training method for a feature extraction model based on map data provided by an embodiment of the present disclosure may include:
[0065] S210. Generate map element samples of different types of map elements in the sample area based on high-precision map data.
[0066] The map element samples include at least one element sample corresponding to a preset map processing task.
[0067] Understandably, high-precision map data is obtained, and the high-precision map data includes map element information, which is presented in the form of images. By dividing the sample area of the high-precision map data, map element samples of different types of map elements in the sample area are obtained, where map elements include traffic lights, roads, stop lines, zebra crossings, and pedestrian safety islands, etc. Correspondingly, map element samples are road samples, zebra crossing samples, etc. Each type of map processing task contains an element sample, for example, the element sample corresponding to the intersection surface classification task is a road sample.
[0068] Among them, the preset map processing task types include at least one of a univariate element prediction task, a binary element prediction task and an element generation task, wherein the element refers to a map element.
[0069] Among them, the univariate feature prediction task is a task type that classifies and predicts the attributes of a single feature, the binary feature prediction task is a task type that classifies and predicts the feature relationship between two features, and the feature generation task is a task type that generates the geometric position information of a single feature.
[0070] It can be understood that a univariate feature prediction task refers to the classification prediction of a certain attribute of a single feature, which can be the prediction of the turning attribute of the lane line, the lane type prediction of the lane line, the road boundary attribute prediction, etc. The lane line is a map feature. A binary feature prediction task refers to the classification prediction of the relationship between two features, which can be whether the "lane line-lane line" is in the same road group, whether the "road-road" is topologically connected, the prediction of the junction point of the divergence, etc. "Lane line-lane line" is two map elements. For multiple elements involved in realizing the same map processing task, they can be the same or different. The feature generation task refers to predicting the geometric position information of the mask feature based on the feature information, which can be the generation of the geometric position information of the lane line and the geometric position information of the guide strip.
[0071] S220: taking the map element sample as the input of the feature extraction model, and generating the map element feature vector by the feature extraction model.
[0072] It is understandable that, based on the above S210, the feature extraction model is a pre-trained model, and the pre-trained model architecture (based on the pre-trained model architecture / pre-trained high-precision model architecture) refers to the upper-level model architecture suitable for subsequent map processing tasks, which is used to extract and fuse the features of each map element, and is also the model structure used and migrated by downstream specific generation tasks. The input of the feature extraction model is a map element sample, and the input map element sample can be composed of map element samples of at least one sample area (area), and the output is a map element feature vector of each map element sample, which can be used as the input of subsequent different map processing tasks. It is understandable that the features extracted by the feature extraction model include geometric features and type attribute features.
[0073] Among them, the feature extraction model includes a geometric encoding module, an attribute encoding module and an encoding module based on the self-attention mechanism.
[0074] For example, Figure 3 As shown, Figure 3A structural schematic diagram of a high-precision map model provided in an embodiment of the present disclosure, the high-precision map model includes a feature extraction model and a map processing model corresponding to the map processing tasks of various types of map elements, the feature extraction model includes a geometric encoding module, an attribute encoding module and an encoding module based on a self-attention mechanism, wherein the geometric encoding module can use PointNet as an encoder for transcoding a set of any size into a feature vector, PointNet is a type of encoder commonly used in computer vision for processing point cloud data, since the vectors of each element are converted into an order-independent spatial point set through vector differential, PointNet can be used for encoding and input as one of the feature vectors (geometric features) of the element into the lower-level encoder based on the self-attention mechanism. In the coding module; the attribute encoder can be understood as a type attribute encoder, which uses the type embedding method to convert each type attribute into an exclusive feature vector, and inputs it into the encoding module based on the self-attention mechanism in the lower layer as another feature vector (type attribute feature) of the element; the encoding module based on the self-attention mechanism is a neural network based on self-attention association, which specifically implements the Transformer layer neural network. The Transformer layer includes a self-attention layer and a residual connection layer, which is conducive to strengthening the information interaction between elements, allowing the neural network to understand the map traffic information of the sample context, and strengthening the representation of the feature vectors of each map element. The encoding module based on the self-attention mechanism outputs multiple feature feature vectors as the input of downstream tasks. The map processing model of each map processing task includes a map processing model for univariate feature prediction corresponding to the univariate feature prediction task type, a map processing model for binary feature prediction corresponding to the binary feature prediction task type, and a map processing model for feature generation corresponding to the feature generation task type, wherein the map processing model for univariate feature prediction can use an MLP neural network for prediction, and the input of the map processing model for univariate feature prediction is one feature feature vector among the above-mentioned multiple map feature vectors, the map processing model for binary feature prediction can also use an MLP neural network for prediction, and the input of the map processing model for binary feature prediction is two feature feature vectors among the above-mentioned multiple map feature vectors, and the map processing model for feature generation can also use a decoder (TransformerDecoder) for feature generation, and the input of the map processing model for feature generation is the above-mentioned multiple map feature vectors. It can be understood that the above-mentioned neural network structure can be determined by the user according to user needs, and will not be elaborated here.
[0075] Optionally, in the above S220, the map element sample is used as the input of the feature extraction model, and the feature extraction model generates the map element feature vector, which is specifically implemented by the following steps:
[0076] Each map feature sample is used as the input of the geometric encoding module, and the geometric encoding module generates the feature geometry vector of the corresponding map feature sample; each map feature sample is used as the input of the attribute encoding module, and the attribute encoding module generates the feature attribute vector of the corresponding map feature sample; the feature geometry vector and each feature attribute vector are used as the input of the encoding module based on the self-attention mechanism, and the encoding module based on the self-attention mechanism generates the map feature feature vector; wherein the map feature vector includes the geometric position information of the corresponding map feature sample, the type attribute information of the corresponding map feature sample, and the feature relationship information between the corresponding map feature sample and the remaining map feature samples.
[0077] It can be understood that each map element sample is input into the geometry encoding module and the attribute encoding module respectively. The map element sample is a spatial point set. The geometry encoding module encodes the spatial point set into a feature vector with geometric attributes to obtain the element geometry vector of the map element sample. The attribute encoding module converts the feature vector according to the element type of the map element sample into an element attribute vector of each element type, wherein the element attribute vector includes an entity type attribute vector and a subdivision type attribute vector. The entity type attribute facilitates the model to distinguish between elements, and the subdivision type attribute facilitates the model to distinguish within the element. The subdivision type attributes include the turning attributes of the lane line and the lane type attributes, etc. Subsequently, all feature geometry vectors and feature attribute vectors are input into the encoding module based on the self-attention mechanism. The encoding module based on the self-attention mechanism builds feature relationship information by strengthening the information interaction between features and extracting feature context semantic information, and generates a map feature feature vector. The map feature vector has a strong representation ability and can extract the map traffic information of the sample as a whole, so that the map feature vector after the feature extraction model can adapt to various downstream tasks more quickly and improve the training effect of the downstream task model. The map feature vector contains the geometric position information of the corresponding map feature sample, the type attribute information of the corresponding map feature sample, and the feature relationship information between the corresponding map feature sample and the remaining map feature samples.
[0078] S230: classify the map element feature vectors according to the preset map processing tasks to obtain sample feature vectors of the corresponding map processing tasks.
[0079] It is understandable that, based on the above S220, after obtaining the map element feature vector, a sample feature vector is determined for the map processing model corresponding to each map processing task, that is, each map processing task has a unique training sample, and the sample feature vector is composed of at least one map element feature vector. One possible classification method is to classify the map element feature vectors according to the task content (such as input and output) of the map processing task; another possible classification method is to label each map element sample with a task label, extract the feature vector, and classify based on the task label; other possible classification methods are not limited.
[0080] S240: Using the sample feature vector of the corresponding map processing task as the input of the map processing model of the map processing task, training the map processing model, and generating a model output result.
[0081] It is understandable that, based on the above S230, for each map processing task, the determined sample feature vector is input into the map processing model, the map processing model is trained, and the output result of the map processing model is generated. For example, the training sample of the lane line turning attribute prediction task is the lane line feature vector. The lane line feature vector includes the lane line related geometry information, type attribute information and element relationship information with other elements. The attribute information includes information of other attributes in addition to the turning attribute. The output result of the set map processing model is the lane line turning attribute result, for example, the turning result such as left turn or right turn. It is understandable that the execution order of the map processing models corresponding to each map processing task is not limited. Each time the sample feature vector of the map processing model of a map processing task is determined, the corresponding prediction task can be executed based on the sample feature vector.
[0082] S250. Based on the model output results, the feature extraction model and the map processing model are iteratively trained until the model convergence condition is reached.
[0083] It can be understood that on the basis of the above S240, the model output results of the map processing model of each map processing task and the sample standard results in each sample feature vector are obtained, and the model loss of the feature extraction model and the map processing model is calculated based on all model output results and the corresponding sample standard results. The network parameters of the feature extraction model and the map processing model are updated based on the model loss until the model convergence conditions are met, wherein the specific calculation method of the model loss is not limited.
[0084] The disclosed embodiment provides a method for training a feature extraction model based on map data. The feature extraction model and a variety of map processing models covering the map processing tasks of the high-precision map production process are trained in series. The variety of map processing models jointly train the feature extraction model to improve the feature extraction capability of the feature extraction model for high-precision map elements, thereby improving the universality of the feature extraction model.
[0085] In some embodiments, Figure 4 As shown, Figure 2 FIG. 1 is a schematic diagram of a detailed process of S210 in a method for training a feature extraction model based on map data. Figure 4 , S210 “Generate map element samples of different types of map elements in the sample area based on high-precision map data”, including:
[0086] S410. Generate map element vectors of different types of map elements in the sample area based on high-precision map data.
[0087] It can be understood that the high-precision map data is intercepted into a fixed-size grid to obtain the vectors of each map element contained in the sample area.
[0088] Optionally, the above-mentioned generation of map element vectors of different types of map elements in the sample area based on high-precision map data can be specifically achieved through the following steps:
[0089] The high-precision map data is segmented according to the area size of the sample area to determine at least one sample area; element vector processing is performed on the sample area to generate map element vectors of different types of map elements in the sample area.
[0090] It can be understood that the high-precision map is divided according to the area size of the sample area to obtain multiple sample areas. Specifically, the high-precision map is divided into multiple areas with a fixed-size pane. The area is the smallest processing unit in the model, and the area is also the sample area. Subsequently, based on the element vector processing method, at least one map element vector is determined for each area. At least one map element vector refers to the vector of different types of map elements in the area. Each area and the map element vectors of each map element contained in the area are taken as a sample to obtain multiple samples. Multiple samples can be used to train the model in batches.
[0091] S420: Perform mask processing on the target attribute of the target map element vector to generate a mask element vector corresponding to the target map element vector.
[0092] The target attribute is associated with the map processing task; the target map element vector is a map element vector having the target attribute.
[0093] It can be understood that, based on the above S410, the target attribute in the map element vector is masked, that is, the target attribute (mask element) is shielded to generate a mask element vector. Among them, the map element vector includes geometric information, and the geometric information includes multiple information, such as geometric position information, etc. The geometric position information of the map element vector is masked to generate a mask element vector. The mask element vector also includes vector data of other geometric information besides the geometric position information. Among them, the target attribute is associated with the map processing task. For example, the map processing task is a lane line task, the target attribute is a turning attribute, and the target map element vector refers to all map element vectors with target attributes. It is understandable that if the target attribute includes multiple information to be masked, multiple masking element vectors can be obtained by performing masking processing multiple times for each target map feature vector, that is, only one information to be masked is masked each time. For example, for two information to be masked, such as geometric position information and type attribute information, the map feature vector must be masked at least twice respectively, and the multiple masking element vectors include at least a first type of masking element vector that only masks out one information to be masked and a second type of masking element vector that masks out at least part of the information to be masked, that is, the geometric position information is masked to obtain a first type of masking element vector, and the type attribute information is masked to obtain another first type of masking element vector. Based on the two first type of masking element vectors, a second type of masking element vector that simultaneously masks out two information to be masked is obtained, or, the geometric position information and the type attribute information are simultaneously masked to obtain a second type of masking element vector.
[0094] S430 : Generate map element samples of different types of map elements in the sample area based on the mask element vector and the remaining map element vectors without target attributes.
[0095] It can be understood that in the above S420, the mask element vector and the remaining map element vectors without target attributes are subjected to element vector difference and coordinate system conversion to generate map element samples of different types of map elements.
[0096] Optionally, in the above S430, based on the mask element vector and the remaining map element vectors without target attributes, map element samples of various map elements in the sample area are generated, which can be specifically implemented by the following steps:
[0097] Coordinate difference processing is performed on each point included in the mask feature vector and the remaining map feature vector to generate an unordered spatial point set corresponding to the corresponding map feature vector; based on the unordered spatial point set corresponding to the corresponding map feature vector, map feature samples of various types of map elements in the sample area are generated.
[0098] It can be understood that for each sample area, coordinate difference processing is performed on the points contained in the mask feature vector and the remaining map feature vectors without target attributes to generate an unordered spatial point set corresponding to the corresponding map feature vector. Each sample area has at least one map feature vector, and at least one map feature vector includes the target map feature vector and the remaining map feature vectors. Among them, high-precision map elements are expressed and stored by map vectors. Map vectors are essentially an ordered set of spatial points. The order of each point in it usually expresses special traffic semantics. The arrangement order of shape points in line strings and vector surfaces is called the digitized direction of vector data, which expresses important information such as road driving direction. Therefore, in order to explicitly express the digitized direction, the shape points in the plane vector data can be differentiated front and back according to the order of vector connection, that is, the shape points Point<x,y> After vector differentiation, it is expanded into Point<x,y,dx,dy> , where x and y are coordinates,<dx,dy> is the unit vector after vector difference. After vector difference processing, the vector data is transformed from an ordered sequence of shape points into a set of shape points that are independent of the order, that is, an unordered spatial point set. Subsequently, the unordered spatial point set can be directly determined as a map element sample of various map elements.
[0099] Optionally, the above-mentioned generation of map element samples of different types of map elements in the sample area based on the unordered spatial point set corresponding to the corresponding map element vector can be specifically implemented by the following steps:
[0100] Construct a regional coordinate system with the center of the sample area as the coordinate origin; transform the unordered spatial point set corresponding to the corresponding map element vector from the world coordinate system to the regional coordinate system; based on the regional size of the sample area, perform coordinate normalization processing on the unordered spatial point set corresponding to the corresponding map element vector after the coordinate transformation, and generate map element samples of different types of map elements.
[0101] It is understandable that after obtaining the unordered spatial point set, all shape points in the unordered spatial point set are transformed into coordinate systems to generate map element samples, so as to reduce the difficulty of model learning, reduce computing power consumption, and further improve learning efficiency. Specifically, the samples obtained after completing the task segmentation still use the geodetic coordinate system (world coordinate system). In this case, a regional coordinate system with the center of the sample area as the coordinate origin is constructed, that is, the regional center of the sample area is set as the origin of the new coordinate system, and the coordinates of the unordered spatial point set are transformed from the world coordinate system to the new coordinate system. The coordinates of the unordered spatial point set after the coordinate transformation are normalized according to the size of the sample area, and can be normalized to the range of 0 to 1, that is, the sample is transformed into a coordinate system, and the coordinates are represented to the range of 0 to 1, so as to reduce the amount of data calculation and generate the final map element samples as the input of the model.
[0102] Optionally, based on S410 to S430, in S230, the map element feature vectors are classified according to preset map processing tasks, and the sample feature vectors of the corresponding map processing tasks are obtained, including:
[0103] Based on the number of feature vectors and the characteristics of feature vectors corresponding to the task content of the map processing task, the feature vectors of each map element are classified to obtain the sample feature vectors of the corresponding map processing task.
[0104] It can be understood that the task content includes univariate feature prediction task content, binary feature prediction task content and feature generation task content. Among them, the univariate feature prediction task content is to predict the attributes of a map element, the binary feature prediction task content is to predict the relationship between two map elements, and the feature generation task content is to generate geometric position information from multiple map elements.
[0105] Optionally, the map processing task includes a univariate feature prediction task, the task content of which is to classify and predict the type attributes of a single map feature. The number of feature vectors corresponding to the task content is 1, and the feature vector characteristics corresponding to the task content include mask features of target attributes associated with the map processing task; the target attributes include type attributes and / or geometric position attributes.
[0106] It can be understood that the unary feature prediction task is used to predict at least one attribute information of a map feature, and the corresponding feature vector number is 1, that is, it involves a map feature but the number of attributes to be predicted is not limited. For example, the target attributes to be predicted for the lane line map element include the lane line type attribute and the geometric position attribute. The lane line type attribute is solid line, dashed line, double solid line, etc.
[0107] Optionally, when the map processing task is a univariate feature prediction task, based on the number of feature vectors and feature vector characteristics corresponding to the task content of the map processing task, the feature vectors of each map element are classified to obtain a sample feature vector of the corresponding map processing task, which can be specifically achieved through the following steps:
[0108] The map element feature vectors having mask features of target attributes are selected from the map element feature vectors to obtain a first feature vector set; and one map element feature vector is selected from the first feature vector set as a sample feature vector for the unary feature prediction task.
[0109] It can be understood that each map element sample generated in S430 is input into the feature extraction model to generate each map element feature vector, and all map element feature vectors with mask features of target attributes are screened from each map element feature vector to obtain a first feature vector set. One map element feature vector is selected from the first feature vector set as a sample feature vector of the unary feature prediction task, and the selected map element feature vector is a vector obtained by the map element sample through the feature extraction model, which is recorded as a feature vector mask.
[0110] Optionally, based on the above embodiment, in S240, the sample feature vector of the corresponding map processing task is used as the input of the map processing model of the map processing task, the map processing model is trained, and the model output result is generated, which can be specifically implemented by the following steps:
[0111] The sample feature vector of the univariate factor prediction task is used as the input of the multi-layer perceptron model corresponding to the univariate factor prediction task, the multi-layer perceptron model corresponding to the univariate factor prediction task is trained, and the attribute attribution probability of the corresponding sample feature vector is generated.
[0112] It is understandable that the feature vector mask determined above is input into the multi-layer perceptron model corresponding to the univariate feature prediction task. The multi-layer perceptron model performs classification prediction based on the feature vectors of other elements in the sample feature vector that have a correlation with the target attribute, and predicts the target attribute. For example, for the map element of road, the target attribute is the turning attribute. According to the feature vectors of other elements, it is determined that the road is adjacent or connected to the road turning right, and it can be predicted that the turning attribute of the road is a right turn, that is, the target attribute is predicted based on the feature vectors of other elements. It is understandable that the univariate feature prediction task is a classification task, and the binary cross entropy loss function (BCE Loss, Binary CrossEntropy Loss) can be used to train the map processing model of the univariate feature prediction task based on the output results. The specific training process is not limited here.
[0113] The disclosed embodiment provides a method for training a feature extraction model based on map data. The feature extraction model is trained through a univariate element prediction task so that the feature extraction model has the ability to understand the classification scenario and can better adapt to the classification task.
[0114] In some embodiments, when the map processing task includes a binary feature prediction task, the feature vectors of any two map features are concatenated and then predicted, so that the feature extraction model can learn the relationship between the features during the model training process, which facilitates the subsequent migration of the feature extraction model to the binary feature prediction task. Figure 5 As shown, Figure 2The diagram shows a detailed flow chart of S230 to S240 for a binary element prediction task in a training method of a feature extraction model based on map data.
[0115] Among them, the map processing task includes a binary feature prediction task. The task content of the binary feature prediction task is to classify and predict the feature relationship between two map elements. The number of feature vectors corresponding to the task content is 2. The feature vector characteristics corresponding to the task content include mask features of target attributes that do not exist in association with the map processing task; target attributes include type attributes and / or geometric position attributes.
[0116] It can be understood that each map element has a corresponding feature vector.
[0117] See also Figure 5 , the above-mentioned “classifying the map element feature vectors according to the preset map processing tasks to obtain the sample feature vectors of the corresponding map processing tasks; using the sample feature vectors of the corresponding map processing tasks as the input of the map processing model of the map processing tasks, training the map processing model, and generating the model output results” includes:
[0118] S510 , filtering map element feature vectors having mask features without target attributes from the map element feature vectors to obtain a second feature vector set.
[0119] Exemplarily, a binary element prediction task may be to predict whether "lane line-lane line" are in the same road group, predict whether "road-road" are topologically connected, etc. Lane line-lane line are two map elements, that is, if lane line 1 and lane line 2 both belong to road 1, it means that lane line 1 and lane line 2 are in the same road group. Belonging to the same road group is the element relationship between the two lane lines.
[0120] It can be understood that the map element feature vectors without the mask feature of the target attribute are screened from each map element feature vector, and the map element feature vector without the mask feature of the target attribute refers to the feature vector of the remaining map element vectors that does not include the target attribute. The second feature vector set is constructed based on all the map element feature vectors obtained by screening, and the absence of the target attribute refers to the absence of the attribute of two lane lines on the same road.
[0121] S520 , select any two map element feature vectors from the second feature vector set and concatenate them to form an element pair feature vector as a sample feature vector for the binary element prediction task.
[0122] It can be understood that the binary feature prediction task refers to predicting the relationship between two features. Any two map feature vectors in the second feature vector set are concatenated to generate a feature pair feature vector. For example, the second feature vector set includes n important map pixel feature vectors. The map feature vector i and the map feature vector j in the n important map pixel feature vectors (denoted as map feature vector 1 to map feature vector n) are concatenated, denoted as concat (feature vector i, feature vector j), where i and j are less than or equal to n. The specific concatenation method is not limited here, and the feature pair feature vector is used as the sample feature vector of the map processing model corresponding to the binary feature prediction task.
[0123] S530, using the sample feature vector of the binary feature prediction task as the input of the multi-layer perceptron model corresponding to the binary feature prediction task, training the multi-layer perceptron model corresponding to the binary feature prediction task, and generating the feature relationship attribution probability corresponding to the corresponding sample feature vector.
[0124] It can be understood that on the basis of the above S520, the sample feature vector is input into the multi-layer perceptron model corresponding to the binary feature prediction task, and the multi-layer perceptron model is used for prediction to output the feature relationship attribution probability corresponding to the sample feature vector. For example, the feature relationship between roads is topologically connected, non-topologically connected, etc. Each feature relationship has an attribution probability, and the feature relationship with the maximum attribution probability can be determined as the output result.
[0125] It is understandable that the task type of the binary feature prediction task is a classification task, and the binary cross entropy loss function can also be used for training.
[0126] The disclosed embodiment provides a method for training a feature extraction model based on map data, by concatenating two map element feature vectors as a sample feature vector corresponding to a binary element prediction task, using a multilayer perceptron model to perform predictions and output results, and training the feature extraction model based on the output results, so that the feature extraction model can strengthen the connection between elements during feature extraction, so as to provide downstream tasks with semantic information that the original sample does not directly possess.
[0127] In some embodiments, the map processing task includes an element generation task, the task content of which is to generate geometric position information of a single map element, the number of feature vectors corresponding to the task content is multiple, and the feature vector characteristics corresponding to the task content at least include mask features with geometric position attributes. Figure 6 As shown, Figure 2 The diagram shows a detailed flow chart of S230 to S240 for the feature generation task in the training method of the feature extraction model based on map data.
[0128] It can be understood that when the map processing task includes a feature generation task, the geometric position information (target attribute) in the map feature vector is masked, and the model predicts the location of the masked feature (single map feature) based on other geometric information between the features. During the model training process, the feature extraction model can extract the feature vector of the masked feature, which facilitates the subsequent migration of the feature extraction model to the feature generation prediction task.
[0129] See also Figure 6 The above-mentioned “based on the number of feature vectors and the feature vector characteristics corresponding to the task content of the map processing task, classify the feature vectors of each map element to obtain the sample feature vectors of the corresponding map processing task; use the sample feature vectors of the corresponding map processing task as the input of the map processing model of the map processing task, train the map processing model, and generate the model output result” includes:
[0130] S610: Filter, from among the map element feature vectors, a map element feature vector having a mask feature with a geometric position attribute, and a plurality of map element feature vectors associated with the map element feature vector as sample feature vectors for the feature generation task.
[0131] It can be understood that after the map element samples are generated, the map element samples are input into the feature extraction model to generate feature vectors of each map element. Subsequently, the map element feature vectors with mask features of geometric location attributes and multiple map element feature vectors associated with the map element feature vector are selected from the map element feature vectors as sample feature vectors of the feature generation task, wherein the multiple map element feature vectors associated with the map element feature vector can be multiple map element feature vectors extracted from the local area of the map element sample, or can be selected from part of the map element feature vectors, that is, part of the input features are local mask features of the map elements. All map element feature vectors can also be determined as sample feature vectors corresponding to the feature generation task, that is, map element feature vector 1 to map element feature vector n are used as inputs of the map processing model corresponding to the feature generation task, that is, the feature generation task is for the geometric location attribute, and its input features are mask features of all map elements, including type attribute mask features and geometric location mask features.
[0132] S620: Using the sample feature vector of the feature generation task as a decoder based on the self-attention mechanism corresponding to the feature generation task, training the decoder to generate a geometric feature vector of the corresponding sample feature vector.
[0133] It can be understood that the feature vectors of each map element, including a single map element, are input into a decoder based on the self-attention mechanism. Specifically, the decoder based on the self-attention mechanism can be a Transformer Decoder, and then the decoder based on the self-attention mechanism outputs and extracts the feature vector of a single map element. The feature vector of the output single map element is calculated with the actual geometric feature vector to train the high-precision map model. The feature generation task is a generative task, and can be trained using a logistic regression loss function.
[0134] The disclosed embodiment provides a method for training a feature extraction model based on map data, predicting the location of mask elements according to the geometric information between elements, and pre-training the feature extraction model based on the prediction results to make it suitable for downstream generative tasks.
[0135] It should be noted that the map processing tasks may include any two of the unary feature prediction tasks, binary feature prediction tasks, and feature generation tasks. In this way, during the training process of the feature extraction model based on map data, the model training process corresponding to the corresponding preset task type can be executed in parallel, so that the feature extraction model can adapt to the downstream tasks of the two types of map processing tasks, further improving the applicability and transferability of the feature extraction model to downstream map processing tasks. The map processing tasks may include unary feature prediction tasks, binary feature prediction tasks, and feature generation tasks. In this way, during the training process of the feature extraction model based on map data, the model training process corresponding to the corresponding map processing tasks can be executed in parallel, so that the feature extraction model can adapt to the downstream one-classification tasks, two-classification tasks, and generation tasks, further improving the applicability and transferability of the feature extraction model to downstream map processing tasks.
[0136] Figure 7 A flowchart of a method for training a task model based on map data provided by an embodiment of the present disclosure is provided, which can be applicable to the case of migrating a pre-trained model to a downstream production task. The method for training a task model based on map data can be performed by a training device for a task model based on map data, which can be implemented in software and / or hardware and can be integrated on an electronic device with certain computing capabilities, such as a laptop, desktop computer, server, etc.
[0137] It should be noted that the training method of the task model based on map data is to perform training and fine-tuning of the downstream task model based on a pre-trained feature extraction model. Its implementation principle is the same as that of the feature extraction model based on high-precision map data. Therefore, the explanation of the same or similar terms and steps in this embodiment to those in the above embodiments can be found in the description of the above embodiments and will not be repeated here.
[0138] like Figure 7 As shown, the training method of the task model based on map data provided by the embodiment of the present disclosure may include:
[0139] S710: Obtain map element samples in a sample area corresponding to a map processing task.
[0140] It is understandable that a map processing task is determined, and the map processing task is a new downstream task to be trained. A plurality of map element samples in a sample area corresponding to the new downstream task are determined according to the method provided in the above embodiment.
[0141] S720: Use the map element sample as input to the feature extraction model, and generate a map element feature vector by the feature extraction model.
[0142] The feature extraction model is pre-trained by the above-mentioned training method of the feature extraction model based on map data.
[0143] It can be understood that the steps of generating the feature vectors of each map element by the feature extraction model refer to the above embodiment and will not be described in detail here.
[0144] S730: Using the map element feature vector as the input of a map processing model corresponding to the map processing task, and generating a model output result by the map processing model.
[0145] S740. Based on the model output results, iteratively train the feature extraction model and the map processing model until the model convergence condition is reached.
[0146] It can be understood that, based on the above S730, the feature extraction model and the map processing model are optimized and trained according to the model output results and the actual sample results until the convergence condition of the map processing model is reached.
[0147] It is understandable that in the generation process of high-precision maps, the downstream tasks based on map vector data are very rich, including classification tasks such as "traffic light-road" matching, intersection surface classification, intersection surface boundary prediction, as well as generative tasks such as virtual line generation, intersection surface generation, and interchange area generation. Downstream tasks such as classification tasks and generative tasks basically also conform to the three pre-training tasks involved in the feature extraction model training process (unary element prediction tasks, binary element prediction tasks, and element generation tasks). Therefore, the feature vector output by the feature extraction model obtained by pre-training using the above embodiments can largely meet the feature information requirements of downstream map processing tasks, thereby improving the effect of the task model, while reducing the training intensity of the task model of downstream tasks, greatly reducing the R&D cost of downstream task models, and shortening the development cycle.
[0148] Figure 8The structure diagram of a training device for a feature extraction model based on map data provided by an embodiment of the present disclosure is shown in FIG. Figure 8 As shown, the device 800 provided in the embodiment of the present disclosure may include:
[0149] A sample generating unit 810 is used to generate map element samples of different types of map elements in a sample area based on high-precision map data;
[0150] The feature feature vector generating unit 820 is used to use the map feature sample as the input of the feature extraction model, and generate the map feature feature vector by the feature extraction model;
[0151] The sample feature vector determination unit 830 is used to classify the map element feature vectors according to the preset map processing tasks to obtain the sample feature vectors of the corresponding map processing tasks;
[0152] A map processing model training unit 840 is used to use the sample feature vector of the corresponding map processing task as the input of the map processing model of the map processing task, train the map processing model, and generate a model output result;
[0153] The feature extraction model training unit 850 is used to iteratively train the feature extraction model and the map processing model based on the model output results until the model convergence condition is reached.
[0154] In some embodiments, the sample generation unit 810 is used to:
[0155] Generate map element vectors of different types of map elements in the sample area based on high-precision map data;
[0156] Masking the target attribute of the target map element vector to generate a mask element vector corresponding to the target map element vector; wherein the target attribute is associated with the map processing task; and the target map element vector is a map element vector having the target attribute;
[0157] Map feature samples of different types of map features in the sample area are generated based on the mask feature vectors and the remaining map feature vectors without the target attribute.
[0158] In some embodiments, the sample generation unit 810 is used to:
[0159] Perform coordinate difference processing on each point included in the mask element vector and the remaining map element vector to generate an unordered spatial point set corresponding to the corresponding map element vector;
[0160] Based on the unordered spatial point set corresponding to the corresponding map element vector, map element samples of various types of map elements in the sample area are generated.
[0161] In some embodiments, the sample generation unit 810 is used to:
[0162] Construct a regional coordinate system with the center of the sample area as the coordinate origin;
[0163] Convert the unordered spatial point set corresponding to the corresponding map element vector from the world coordinate system to the regional coordinate system;
[0164] Based on the area size of the sample area, coordinate normalization processing is performed on the unordered spatial point set corresponding to the corresponding map element vector after coordinate transformation to generate map element samples of different types of map elements.
[0165] In some embodiments, the sample feature vector obtaining unit 830 is used to:
[0166] Based on the number of feature vectors and the characteristics of feature vectors corresponding to the task content of the map processing task, the feature vectors of each map element are classified to obtain the sample feature vectors of the corresponding map processing task.
[0167] Among them, the map processing task includes a univariate feature prediction task. The task content of the univariate feature prediction task is to classify and predict the type attribute of a single map feature. The number of feature vectors corresponding to the task content is 1. The feature vector characteristics corresponding to the task content include mask features of target attributes associated with the map processing task; target attributes include type attributes and / or geometric position attributes.
[0168] In some embodiments, the sample feature vector obtaining unit 830 is used to:
[0169] Selecting map element feature vectors having mask features of target attributes from each map element feature vector to obtain a first feature vector set;
[0170] Select any map feature vector from the first feature vector set as a sample feature vector for the unary feature prediction task.
[0171] In some embodiments, the sample feature vector obtaining unit 830 is used to:
[0172] The sample feature vector of the univariate factor prediction task is used as the input of the multi-layer perceptron model corresponding to the univariate factor prediction task, the multi-layer perceptron model corresponding to the univariate factor prediction task is trained, and the attribute attribution probability of the corresponding sample feature vector is generated.
[0173] Among them, the map processing task includes a binary feature prediction task. The task content of the binary feature prediction task is to classify and predict the feature relationship between two map elements. The number of feature vectors corresponding to the task content is 2. The feature vector characteristics corresponding to the task content include mask features of target attributes that do not exist in association with the map processing task; target attributes include type attributes and / or geometric position attributes.
[0174] In some embodiments, the sample feature vector obtaining unit 830 is used to:
[0175] Selecting map element feature vectors of mask features without target attributes from each map element feature vector to obtain a second feature vector set;
[0176] Select any two map element feature vectors from the second feature vector set and concatenate them to form an element pair feature vector as a sample feature vector for the binary element prediction task.
[0177] In some embodiments, the sample feature vector obtaining unit 830 is used to:
[0178] The sample feature vector of the binary feature prediction task is used as the input of the multi-layer perceptron model corresponding to the binary feature prediction task, and the multi-layer perceptron model corresponding to the binary feature prediction task is trained to generate the feature relationship attribution probability corresponding to the corresponding sample feature vector.
[0179] Among them, the map processing task includes the feature generation task. The task content of the feature generation task is to generate the geometric position information of a single map element. The number of feature vectors corresponding to the task content is multiple, and the feature vector characteristics corresponding to the task content at least include mask features with geometric position attributes.
[0180] In some embodiments, the sample feature vector obtaining unit 830 is used to:
[0181] A map element feature vector having a mask feature with geometric position attributes and a plurality of map element feature vectors associated with the map element feature vector are selected from each map element feature vector as sample feature vectors for the feature generation task.
[0182] In some embodiments, the sample feature vector obtaining unit 830 is used to:
[0183] The sample feature vector of the feature generation task is used as a decoder based on the self-attention mechanism corresponding to the feature generation task, and the decoder is trained to generate a geometric feature vector of the corresponding sample feature vector.
[0184] Among them, the feature extraction model includes a geometric encoding module, an attribute encoding module and an encoding module based on the self-attention mechanism.
[0185] In some embodiments, the map element feature vector generation unit 820 is used to:
[0186] Each map element sample is used as an input of a geometric encoding module, and the geometric encoding module generates an element geometric vector of the corresponding map element sample;
[0187] Each map feature sample is used as an input of an attribute encoding module, and the attribute encoding module generates a feature attribute vector of the corresponding map feature sample;
[0188] The feature geometry vector and the feature attribute vector are used as inputs of an encoding module based on a self-attention mechanism, and a map feature feature vector is generated by the encoding module based on a self-attention mechanism; wherein the map feature feature vector includes the geometric position information of the corresponding map feature sample, the type attribute information of the corresponding map feature sample, and the feature relationship information between the corresponding map feature sample and the remaining map feature samples.
[0189] The training device for the feature extraction model based on map data provided in the embodiments of the present disclosure can execute the training method for the feature extraction model based on map data provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. For the contents not fully described in the embodiments of the device of the present disclosure, reference can be made to the description in any method embodiment of the present disclosure.
[0190] Fig. 9 The structure diagram of a training device for a task model based on map data provided by an embodiment of the present disclosure is shown in FIG. Fig. 9 As shown, the device 900 provided in the embodiment of the present disclosure may include:
[0191] The sample acquisition unit 910 is used to acquire the map element samples in the sample area corresponding to the map processing task;
[0192] The feature feature vector generating unit 920 is used to use the map feature sample as an input of the feature extraction model, and generate the map feature feature vector by the feature extraction model; wherein the feature extraction model is pre-trained by the above-mentioned training method of the feature extraction model based on map data;
[0193] A map processing model training unit 930 is used to use the map element feature vector as an input of a map processing model corresponding to the map processing task, and the map processing model generates a model output result;
[0194] The task model training unit 940 is used to iteratively train the feature extraction model and the map processing model based on the model output results until the model convergence conditions are met.
[0195] The training device for the task model based on map data provided in the embodiment of the present disclosure can execute the training method for the task model based on map data provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. For the contents not fully described in the embodiment of the device of the present disclosure, reference can be made to the description in any method embodiment of the present disclosure.
[0196] The disclosed embodiment also provides an electronic device, which may include at least a processor and a memory, wherein the memory may be used to store executable instructions. The processor may be used to read the executable instructions from the memory and execute the executable instructions to implement a method for training a feature extraction model based on map data or a method for training a task model based on map data in any of the above embodiments.
[0197] Fig.10 The structure diagram of an electronic device provided by the embodiment of the present disclosure is shown in FIG. Fig.10 As shown, the electronic device 1000 may include a processor 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are also stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0198] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange data.
[0199] It should be noted that Fig.10 The electronic device 1000 shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure. Fig.10 The electronic device 1000 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0200] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 1009, or installed from a storage device 1008, or installed from a ROM 1002. When the computer program is executed by the processor 1001, the functions defined in a training method for a feature extraction model based on map data or a training method for a task model based on map data provided in any embodiment of the present disclosure can be executed.
[0201] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0202] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0203] The computer-readable medium may be included in the electronic device, or may exist independently without being installed in the electronic device.
[0204] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes a training method for a feature extraction model based on map data or a training method for a task model based on map data provided in any embodiment of the present disclosure.
[0205] In embodiments of the present disclosure, computer program codes for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0206] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0207] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.
[0208] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0209] In the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a computer-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0210] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0211] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0212] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A training method for a feature extraction model based on map data, It is characterized in that include: Generate map element samples of different types of map elements in the sample area based on high-precision map data; Using the map element sample as input of a feature extraction model, and generating a map element feature vector by the feature extraction model; Classifying the map element feature vectors according to preset map processing tasks to obtain sample feature vectors of corresponding map processing tasks; Using the sample feature vector of the corresponding map processing task as the input of the map processing model of the map processing task, training the map processing model, and generating a model output result; Based on the model output results, the feature extraction model and the map processing model are iteratively trained until the model convergence condition is reached.
2. The method according to claim 1, It is characterized in that The generating of map element samples of different types of map elements in the sample area based on the high-precision map data includes: Based on the high-precision map data, generating map element vectors of different types of map elements in the sample area; Performing mask processing on the target attribute of the target map element vector to generate a mask element vector corresponding to the target map element vector; wherein the target attribute is associated with the map processing task; and the target map element vector is the map element vector having the target attribute; The map element samples of different types of map elements in the sample area are generated based on the mask element vectors and the remaining map element vectors without the target attribute.
3. The method according to claim 2, It is characterized in that The generating the map element samples of various types of map elements in the sample area based on the mask element vector and the remaining map element vectors without the target attribute comprises: Performing coordinate difference processing on each point included in the mask element vector and the remaining map element vector to generate an unordered spatial point set corresponding to the corresponding map element vector; The map element samples of each type of map element in the sample area are generated based on the unordered spatial point set corresponding to the corresponding map element vector.
4. The method according to claim 3, It is characterized in that The generating the map element samples of different types of map elements in the sample area based on the unordered spatial point set corresponding to the corresponding map element vector comprises: Constructing a regional coordinate system with the center of the sample area as the coordinate origin; Converting the unordered spatial point set corresponding to the corresponding map element vector from the world coordinate system to the regional coordinate system; Based on the area size of the sample area, coordinate normalization processing is performed on the unordered spatial point set corresponding to the corresponding map element vector after coordinate conversion to generate the map element samples of different types of map elements.
5. The method according to claim 1, It is characterized in that The step of classifying the map element feature vectors according to preset map processing tasks to obtain sample feature vectors of corresponding map processing tasks includes: Based on the number of feature vectors and the feature vector characteristics corresponding to the task content of the map processing task, the feature vectors of each map element are classified to obtain sample feature vectors of the corresponding map processing task.
6. The method according to claim 5, It is characterized in that The map processing task includes a unary element prediction task, the task content of the unary element prediction task is to classify and predict the type attribute of a single map element, the number of feature vectors corresponding to the task content is 1, and the feature vector characteristics corresponding to the task content include mask features of target attributes associated with the map processing task; the target attributes include type attributes and / or geometric position attributes; The number of feature vectors and feature vector characteristics corresponding to the task content of the map processing task are classified to obtain sample feature vectors of the corresponding map processing task, including: Selecting the map element feature vectors having the mask feature of the target attribute from the map element feature vectors to obtain a first feature vector set; Select any one of the map element feature vectors from the first feature vector set as a sample feature vector for the unary element prediction task.
7. The method according to claim 6, It is characterized in that The method of using the sample feature vector of the corresponding map processing task as the input of the map processing model of the map processing task, training the map processing model, and generating a model output result includes: The sample feature vector of the unary element prediction task is used as the input of the multi-layer perceptron model corresponding to the unary element prediction task, the multi-layer perceptron model corresponding to the unary element prediction task is trained, and the attribute attribution probability of the corresponding sample feature vector is generated.
8. The method according to claim 5, in, The map processing task includes a binary element prediction task, the task content of the binary element prediction task is to classify and predict the element relationship between two map elements, the number of feature vectors corresponding to the task content is 2, and the feature vector characteristics corresponding to the task content include mask features of target attributes associated with the map processing task; the target attributes include type attributes and / or geometric position attributes; The number of feature vectors and feature vector characteristics corresponding to the task content of the map processing task are classified to obtain sample feature vectors of the corresponding map processing task, including: Selecting the map element feature vectors without the mask feature of the target attribute from the map element feature vectors to obtain a second feature vector set; Select any two of the map element feature vectors from the second feature vector set and concatenate them to form an element pair feature vector as a sample feature vector for the binary element prediction task.
9. The method according to claim 8, It is characterized in that The method of using the sample feature vector of the corresponding map processing task as the input of the map processing model of the map processing task, training the map processing model, and generating a model output result includes: The sample feature vector of the binary element prediction task is used as the input of the multi-layer perceptron model corresponding to the binary element prediction task, the multi-layer perceptron model corresponding to the binary element prediction task is trained, and the element relationship attribution probability corresponding to the corresponding sample feature vector is generated.
10. The method according to claim 5, in, The map processing task includes an element generation task, the task content of which is to generate geometric position information of a single map element, the number of feature vectors corresponding to the task content is multiple, and the feature vector characteristics corresponding to the task content at least include mask features with geometric position attributes; The number of feature vectors and feature vector characteristics corresponding to the task content of the map processing task are classified to obtain sample feature vectors of the corresponding map processing task, including: The map element feature vector having the mask feature of the geometric position attribute and a plurality of the map element feature vectors associated with the map element feature vector are selected from the map element feature vectors as sample feature vectors for the element generation task.
11. The method according to claim 10, It is characterized in that The method of using the sample feature vector of the corresponding map processing task as the input of the map processing model of the map processing task, training the map processing model, and generating a model output result includes: The sample feature vector of the feature generation task is used as a decoder based on a self-attention mechanism corresponding to the feature generation task, and the decoder is trained to generate a geometric feature vector of the corresponding sample feature vector.
12. The method according to claim 1, in, The feature extraction model includes a geometric encoding module, an attribute encoding module and an encoding module based on a self-attention mechanism; The step of using the map element sample as an input of a feature extraction model and generating a map element feature vector by the feature extraction model includes: Taking each of the map element samples as input of the geometric encoding module, the geometric encoding module generates an element geometric vector of the corresponding map element sample; Taking each of the map element samples as input of the attribute encoding module, the attribute encoding module generates an element attribute vector of the corresponding map element sample; The feature geometry vector and the feature attribute vector are used as inputs of the encoding module based on the self-attention mechanism, and the encoding module based on the self-attention mechanism generates the map feature vector; wherein the map feature vector includes the geometric position information of the corresponding map feature sample, the type attribute information of the corresponding map feature sample, and the feature relationship information between the corresponding map feature sample and the remaining map feature samples.
13. A training method for a task model based on map data, It is characterized in that include: Obtaining map feature samples in a sample area corresponding to a map processing task; The map element sample is used as an input of a feature extraction model, and a map element feature vector is generated by the feature extraction model; wherein the feature extraction model is pre-trained by the training method of a feature extraction model based on map data according to any one of claims 1 to 12; Using the map element feature vector as input of a map processing model corresponding to the map processing task, and generating a model output result by the map processing model; Based on the model output results, the feature extraction model and the map processing model are iteratively trained until the model convergence condition is reached.
14. A training device for a feature extraction model based on map data, It is characterized in that include: A sample generating unit, for generating map element samples of different types of map elements in a sample area based on high-precision map data; A map element feature vector generating unit, used for taking the map element sample as an input of a feature extraction model, and generating a map element feature vector by the feature extraction model; A sample feature vector obtaining unit, used for classifying the map element feature vectors according to preset map processing tasks to obtain sample feature vectors of corresponding map processing tasks; A map processing model training unit, used to use the sample feature vector of the corresponding map processing task as the input of the map processing model of the map processing task, train the map processing model, and generate a model output result; The feature extraction model training unit is used to iteratively train the feature extraction model and the map processing model based on the model output results until the model convergence condition is reached.
15. A training device for a task model based on map data, It is characterized in that include: A sample acquisition unit, used to acquire map element samples in a sample area corresponding to a map processing task; an element feature vector generating unit, configured to use the map element sample as an input of a feature extraction model, and generate a map element feature vector by the feature extraction model; wherein the feature extraction model is pre-trained by the training method for a feature extraction model based on map data according to any one of claims 1 to 12; A map processing model training unit, used for taking the map element feature vector as an input of a map processing model corresponding to the map processing task, and generating a model output result by the map processing model; The task model training unit is used to iteratively train the feature extraction model and the map processing model based on the model output results until the model convergence conditions are met.
16. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a program or instruction, which enables a computer to execute the method for training a feature extraction model based on map data according to any one of claims 1 to 12 or the method for training a task model based on map data according to claim 13.