Map data reward model training method, related method and device

Through the proposed map data reward model training method, the reward model is trained using the map feature data set and importance matrix, and the problem of manual ranking results dependence in the existing technology is solved, and efficient model training and higher accuracy are achieved.

CN120198766APending Publication Date: 2025-06-24SHENYANG MXNAVI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311768032.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the prior art, when strengthening learning training map data generation model, the generation and training of reward models require a large amount of manual ranking results, resulting in large manual workload and subjectivity and inconsistency, affecting the model training effect.

Method used

A map data reward model training method is proposed. By obtaining the map feature data set, rating and ranking, a map geometric raster data is generated, the importance matrix of real map geometric raster data is extracted, and the pre-constructed reward model is trained based on these data to obtain the map data reward model.

Benefits of technology

It effectively reduces the workload of manual scoring and ranking, realizes small sample training, improves model training efficiency, reduces the subjectivity and inconsistency of manual labeling, and improves the accuracy of map data reward models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198766A_ABST
    Figure CN120198766A_ABST
Patent Text Reader

Abstract

The invention discloses a map data reward model training method, a related method and a related device. The map data reward model training method comprises the following steps: acquiring a map element data set; scoring the generated map geometric grating data in each sample of the map element data set to obtain scoring and sorting results of the generated map geometric grating data of each sample; extracting an importance degree matrix corresponding to the geometric grating data of the real map in each sample; and training a pre-constructed reward model based on the score and the sorting result, the map data pair in each sample and the importance degree matrix corresponding to each sample to obtain a map data reward model. In the training stage, the workload of manual scoring and ranking can be effectively reduced, small sample training is achieved, the model training efficiency is improved, and the model training time is saved. Moreover, the map data reward model obtained through training can better adapt to changes of different road layouts and features, so that the accuracy of the map data reward model is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of map processing, and in particular, to a method for training a map data reward model, related methods and devices. Background Art

[0002] With the booming development of the fields of intelligent transportation and autonomous driving, the demand for maps has exceeded the limitations of traditional navigation electronic maps, urgently requiring more accurate, detailed and highly complete map data. To meet the needs of these advanced applications, the rise of high-precision maps is rapidly becoming an industry consensus. In an autonomous driving system, a high-precision map is not just an additional function, but plays a crucial role, providing three core functions: map matching, environmental perception assistance, and path planning. These key functions have rigid requirements and are difficult to be replaced by other alternative solutions. The introduction of high-precision maps has significantly improved the safety and feasibility of autonomous driving systems, taking a crucial step towards the intelligence of driving.

[0003] In the generation algorithm of high-precision maps, reinforcement learning is an effective algorithm, and in reinforcement learning based on map data, the generation and training of a reward model is a very important step. Summary of the Invention

[0004] To obtain an accurate reward model, embodiments of the present invention provide a method for training a map data reward model, related methods and devices.

[0005] In a first aspect, embodiments of the present invention provide a method for training a map data reward model, the method comprising:

[0006] Obtaining a map feature data set; each sample of the map feature data set includes a map data pair and corresponding real map geometric raster data; the map data pair includes map feature raster data and generated map geometric raster data generated based on the map feature raster data;

[0007] Scoring and ranking the generated map geometric raster data in each of the samples to obtain the scoring and sorting results of the generated map geometric raster data of each sample;

[0008] Extracting the importance matrix corresponding to the real map geometric raster data in each sample;

[0009] Training a pre-constructed reward model based on the scoring and sorting results, the map data pair in each sample, and the importance matrix corresponding to each sample to obtain a map data reward model.

[0010] In one or some alternative embodiments of the embodiments of the present application, training a pre-constructed reward model based on the scoring and ranking results, the map data pairs in each sample, and the importance matrix corresponding to each sample to obtain a map data reward model includes:

[0011] According to the scores of the generated map geometric raster data in each sample, perturb the ranking results within a preset score threshold range to obtain the scores corresponding to the generated map geometric raster data in each sample;

[0012] Input the map data pairs in each sample into the pre-constructed reward model to obtain corresponding predicted scores and predicted importance matrices;

[0013] For each sample, calculate the total loss according to the predicted score, the predicted importance matrix, the score, and the importance matrix;

[0014] Update the pre-constructed reward model according to the total loss to obtain an updated reward model;

[0015] Repeat the above steps of model update until the updated reward model meets the preset conditions to obtain a map data reward model.

[0016] In one or some alternative embodiments of the embodiments of the present application, for each sample, calculating the total loss according to the predicted score, the predicted importance matrix, the score, and the importance matrix includes:

[0017] Calculate the regression loss based on the predicted score and the score;

[0018] Calculate the local weight loss based on the predicted importance matrix and the importance matrix;

[0019] Add the regression loss and the local weight loss after weighting to obtain the total loss.

[0020] In one or some alternative embodiments of the embodiments of the present application, according to the scores of the generated map geometric raster data in each sample, perturbing the ranking results within a preset score threshold range to obtain the scores corresponding to the generated map geometric raster data in each sample includes:

[0021] According to the scores of the generated map geometric raster data in each sample, perturb the ranking results within a preset score threshold range to obtain a new ranking result;

[0022] According to the new ranking result, obtain the scores corresponding to the generated map geometric raster data in each sample through normalization.

[0023] In one or some alternative embodiments of the embodiments of the present application, the extracting of the importance matrix corresponding to the real map geometric raster data in each of the samples includes:

[0024] Dividing the real map geometric raster data in each of the samples into a plurality of blocks according to a preset importance matrix size;

[0025] Counting the number of map elements in each block, and extracting a map element matrix corresponding to the real map geometric raster data in the sample;

[0026] For each sample, processing based on the map element matrix to obtain a corresponding importance matrix.

[0027] In a second aspect, an embodiment of the present invention provides a method for training a map data generation model, using a map data reward model obtained by a map data reward model training method; the method includes:

[0028] Obtaining a map data set; each map sample in the map data set includes map element raster data and map geometric raster data;

[0029] Based on the map data set, pre-training a pre-constructed initial map data generation model to obtain a plurality of intermediate map data generation models;

[0030] Based on the map data set, using the map data reward model to perform reinforcement learning training on the plurality of intermediate map data generation models in parallel to obtain a map data generation model.

[0031] In a third aspect, an embodiment of the present invention provides a map data reward model training device, and the device includes:

[0032] A first obtaining module, configured to obtain a map element data set; each sample in the map element data set includes a map data pair and a corresponding real map geometric raster data; the map data pair includes map element raster data and generated map geometric raster data generated based on the map element raster data;

[0033] A ranking module, configured to score and rank the generated map geometric raster data in each of the samples to obtain a scoring and sorting result of the generated map geometric raster data in each sample;

[0034] An extraction module, configured to extract an importance matrix corresponding to the real map geometric raster data in each of the samples;

[0035] A first training module, configured to train a pre-constructed reward model based on the scoring and sorting result, the map data pair in each of the samples, and the importance matrix corresponding to each sample to obtain a map data reward model.

[0036] In a fourth aspect, an embodiment of the present invention provides a device for training a map data generation model, the device comprising:

[0037] A second acquisition module, configured to acquire a map data set; each map sample in the map data set includes map feature raster data and map geometry raster data;

[0038] A generation module, configured to pre-train a pre-constructed initial map data generation model based on the map data set to obtain a plurality of intermediate map data generation models;

[0039] A third acquisition module, configured to acquire a map data reward model obtained by the above-mentioned map data reward model training method;

[0040] A second training module, configured to perform reinforcement learning training on the plurality of intermediate map data generation models in parallel based on the map data set using the map data reward model to obtain a map data generation model.

[0041] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned map data reward model training method, and / or, the above-mentioned map data generation model training method.

[0042] In a sixth aspect, an embodiment of the present invention provides a computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the above-mentioned map data reward model training method, and / or, the above-mentioned map data generation model training method.

[0043] In a seventh aspect, an embodiment of the present invention provides a computer program product containing instructions, and when the computer program product runs on a computer device, it causes the computer device to execute the above-mentioned map data reward model training method, and / or, the above-mentioned map data generation model training method.

[0044] In an eighth aspect, an embodiment of the present invention provides a chip, the chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run a computer program or instruction to implement the above-mentioned map data reward model training method, and / or, the above-mentioned map data generation model training method.

[0045] The beneficial effects of the above technical solutions provided by the embodiments of the present invention at least include:

[0046] The map data reward model training method provided by the embodiment of the present invention obtains a map element data set, scores and ranks the generated map geometric raster data to obtain a sorting result, calculates the importance matrix corresponding to the map geometric raster data, and finally jointly trains the reward model based on the data pairs, sorting results, and importance matrix in the map element data set. In this map data reward model training method, by scoring and ranking the generated map geometric raster data in each sample, the workload of manual scoring and ranking can be effectively reduced during the training phase, small-sample training can be achieved, the model training efficiency can be improved, and the model training time can be saved. Moreover, by extracting the importance matrix corresponding to the real map geometric raster data in the sample, different importance levels can be assigned to different positions in the real map geometric raster data, local information at different positions can be focused on, and the limitations on model training caused by the subjectivity and inconsistency of manual annotation can be effectively avoided or reduced, so that the trained map data reward model can better adapt to the changes in different road layouts and features, thereby effectively improving the accuracy of the map data reward model.

[0047] Other features and advantages of the present invention will be described in the following specification, and in part will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings.

[0048] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0049] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0050] Figure 1 It is a schematic flowchart of a map data reward model training method provided by an embodiment of the present invention;

[0051] Figure 2 It is a schematic structural diagram of an example map data reward model provided by an embodiment of the present invention;

[0052] Figure 3 It is a schematic flowchart of a map data generation model training method provided by an embodiment of the present invention;

[0053] Figure 4 It is a schematic structural diagram of a map data reward model training method device provided by an embodiment of the present application;

[0054] Figure 5It is a schematic structural diagram of the method and device for training a map data generation model provided by an embodiment of the present application. Detailed implementation manners

[0055] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0056] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0057] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0058] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined" or "in response to determining" or "once detecting [the described condition or event]" or "in response to detecting [the described condition or event]" according to the context.

[0059] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0060] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0061] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0062] In order to illustrate the technical solution of the present application, specific embodiments will be used for illustration below.

[0063] The inventor found that in the prior art, when using reinforcement learning to train a map data generation model, the generation and training of the reward model require a large amount of manual ranking results, resulting in a large amount of manual work. Moreover, due to the subjectivity and inconsistency of manual annotation, the results of model training are limited, resulting in problems with insufficient accuracy of the model. Based on this, after further research and development, the inventor made the present invention and provided a method for training a map data reward model, related methods and devices.

[0064] Embodiment 1

[0065] The embodiment of the present invention provides a method for training a map data reward model. Referring to Figure 1 as shown, the method includes:

[0066] S101: Obtain a map feature data set. Each sample in the map feature data set includes a map data pair and corresponding real map geometric raster data, and the map data pair includes map feature raster data and generated map geometric raster data generated based on the map feature raster data.

[0067] In the embodiment of the present application, the map feature data set in the above step S101 includes multiple samples, and each sample includes a map data pair and corresponding real map geometric raster data. The map data pair includes map feature raster data and generated map geometric raster data generated based on the map feature raster data. Among them, the map feature raster data is a raster image composed of pixel points, and each pixel point has a corresponding color or attribute value. The map geometric raster data is data that annotates the geographic information concerned by those skilled in the art on the basis of the map feature raster data. For example, buildings, roads, etc. are marked in the map feature raster data.

[0068] In the embodiment of the present application, the sources of the real map geometric raster data include manual annotation and existing high-precision maps. Among them, there may be too many elements marked in the existing high-precision maps, and it is necessary to preprocess and remove irrelevant information.

[0069] The sources of generating map geometric raster data include manual annotation and generation by a generation model. Among them, the generation model can be a generation model obtained by training using other methods, or a generation model during the training process based on reinforcement learning. In the specific implementation process, those skilled in the art can select multiple generation models to generate map geometric raster data with uneven quality.

[0070] S102: Score and rank the generated map geometric raster data in each sample to obtain the scoring and ranking results of the generated map geometric raster data in each sample.

[0071] In the embodiments of the present application, after preparing the samples, the generated map geometric raster data in the samples is scored based on an image quality metric, ranked based on the scoring results, and the ranking results of each map geometric raster data are obtained. The image quality metrics used can be, for example, FID (Fréchet Inception Distance), IS (Inception Score), etc. Among them, the image quality metrics FID and IS are both metrics for evaluating the image quality generated by a generation model. FID is a metric that quantifies the performance of a generation model by calculating the similarity between the distribution of the generated images and the distribution of real images, and IS is a metric that quantifies the image quality and diversity by calculating the image perplexity and the entropy of the image class distribution.

[0072] S103: Extract the importance matrix corresponding to the real map geometric raster data in each sample.

[0073] In the embodiments of the present application, the steps of extracting the importance matrix of the real map geometric raster data in each sample include:

[0074] According to the preset importance matrix size, divide the real map geometric raster data in each sample into multiple blocks;

[0075] Count the number of map elements in each block, and extract the map element matrix corresponding to the real map geometric raster data in the sample;

[0076] For each sample, process based on the map element matrix to obtain the corresponding importance matrix.

[0077] In the embodiments of the present application, the above-mentioned method for processing the map element matrix may be a normalization operation, and the normalization operation can be implemented using methods such as softmax. To make a clearer and more detailed description of the above-mentioned extraction of the importance matrix corresponding to the real map geometric raster data in each sample, it is exemplified that the real map geometric raster data in a sample is an image of 800 pixels × 800 pixels, and the preset size of the importance matrix is 25 × 25. According to the importance matrix with a size of 25 × 25, the image of 800 pixels × 800 pixels is divided into 625 blocks with a size of 32 pixels × 32 pixels, and the number of map elements in each block is counted to obtain the 25 × 25 map element matrix corresponding to the real map geometric raster data. The map element matrix is normalized to obtain the importance matrix.

[0078] In the embodiments of the present application, the operation of extracting the importance matrix is implemented by dividing the image into multiple blocks and counting the number of map elements in the blocks, which helps to capture the local features in the map geometric raster data, making the map data reward model tend to emphasize and focus on local specific areas, such as areas with more map elements, thereby improving the model accuracy.

[0079] S104: Train a pre-constructed reward model based on the scoring and sorting results, the map data pairs in each sample, and the importance matrix corresponding to each sample to obtain a map data reward model.

[0080] In the embodiments of the present application, the schematic diagram of the map data reward model in the above step S104 is as Figure 2 shown. The input is the map data pair in the sample, that is, the map element raster data (i.e., the map element raster map in the figure) and the generated map geometric raster data (i.e., the generated map element raster map in the figure), and the output is the predicted score and the predicted importance matrix, where the predicted score is Figure 2 "1,1,1" in Figure 2 The structure of the map data reward model exemplified in includes six consecutive convolutional layers, an average pooling layer, and a fully connected layer. The output of the average pooling layer is calculated through softmax to obtain the predicted importance matrix.

[0081] Figure 2 The training process of the map data reward model shown in specifically may include the following steps:

[0082] According to the scores of the generated map geometric raster data in each sample, perturb the sorting results within a preset score threshold range to obtain the scores corresponding to the generated map geometric raster data in each sample;

[0083] Input the map data pairs in each sample into the pre-constructed reward model to obtain the corresponding predicted scores and predicted importance matrices;

[0084] For each sample, a total loss is calculated based on the prediction score, the prediction importance matrix, the score, and the importance matrix.

[0085] Update the pre-constructed reward model according to the total loss to obtain an updated reward model; repeat the steps of the above first to fourth steps to update the reward model until the updated reward model meets the preset conditions to obtain a map data reward model.

[0086] In the embodiments of the present application, the preset conditions for the above model training may be that the number of iterations reaches a preset value, or the change range of the total loss value within the preset number of iterations does not exceed a preset range.

[0087] In a specific embodiment, the operation of disturbing the sorting result within a preset score threshold range according to the scores of the generated map geometric raster data in each sample to obtain the scores corresponding to the generated map geometric raster data in each sample is specifically as follows:

[0088] According to the scores of the generated map geometric raster data in each sample, disturb the sorting result within a preset score threshold range to obtain a new sorting result. According to the new sorting result, the scores corresponding to the generated map geometric raster data in each sample are obtained through normalization.

[0089] The above disturbing operation may specifically be that in the generated map geometric raster data after score sorting, if the difference between multiple consecutive generated map geometric raster data is less than the preset score threshold, it is determined that the difference between these consecutive map geometric raster data is small, and the scores of these consecutive map geometric raster data are shuffled to perform disturbance and re-sorting to obtain a new sorting result.

[0090] In the embodiments of the present application, according to the scores of the generated map geometric raster data in each sample, disturbing the sorting result within a preset score threshold range introduces some noise and variations, affecting the scores corresponding to the generated map geometric raster data in each sample, so that the model does not overly rely on a specific ranking order during training, thereby being able to reduce or avoid the overfitting problem in model training and improving the stability of map data reward model training.

[0091] In a specific embodiment, for each of the above samples, calculating the total loss according to the prediction score, the prediction importance matrix, the score, and the importance matrix is specifically as follows:

[0092] Calculate a regression loss based on the prediction score and the score;

[0093] Calculate a local weight loss based on the prediction importance matrix and the importance matrix;

[0094] Add the regression loss and the local weight loss after weighting to obtain the total loss.

[0095] In the embodiments of the present application, the loss function in the training of the map data reward model includes two parts: the regression loss and the local weight loss. The regression loss is calculated based on the predicted score and the true score corresponding to the generated map geometric raster data in the sample, and the local weight loss is calculated from the predicted importance matrix and the importance matrix corresponding to the true map geometric raster data in the sample.

[0096] In the embodiments of the present application, an importance matrix is introduced to form the local weight loss and provide local weights, enabling the map data reward model to learn the prior knowledge in the importance matrix, thereby improving the performance and adaptability of the model. During the training phase, the workload of manual scoring can be effectively reduced, small-sample training can be achieved, and thus the accuracy of the map data reward model can be effectively improved.

[0097] In a specific embodiment, it may be that based on the predicted importance matrix and the importance matrix, the local weight loss is calculated through the following formula 1:

[0098]

[0099] where LossB is the local weight loss, M*N is the size of the importance matrix, f represents the loss function, A i,j represents the value at the element (i, j) in the predicted importance matrix, and B i,j represents the value at the element (i, j) in the importance matrix. In this embodiment, the value at each element (i, j) in this importance matrix can be set manually.

[0100] In a specific embodiment, it may be that according to the regression loss and the local weight loss, they are weighted and added based on the following formula 2 to calculate the total loss:

[0101] Loss = αLossA + (1 - α)LossB Formula 2;

[0102] where Loss is the total loss, LossA and LossB respectively represent the regression loss and the local weight loss, and α represents the weight.

[0103] The map data reward model training method provided by the embodiments of the present application can effectively reduce the workload of manual scoring and ranking during the training phase by scoring and ranking the generated map geometric raster data in each sample, realize small-sample training, improve the model training efficiency, and save the model training time. Moreover, by extracting the importance matrix corresponding to the real map geometric raster data in the sample, different importance levels can be assigned to different positions in the real map geometric raster data, focusing on the local information of different positions, effectively avoiding or reducing the limitations on model training caused by the subjectivity and inconsistency of manual annotation, enabling the trained map data reward model to better adapt to the changes in different road layouts and features, thereby effectively improving the accuracy of the map data reward model.

[0104] Embodiment 2

[0105] Based on the same inventive concept, the embodiments of the present invention further provide a map data generation model training method. Using the map data reward model obtained by the map data reward model training method described in Embodiment 1, as shown in Figure 3 The method includes:

[0106] S301: Obtain a map data set; each map sample in the map data set includes map feature raster data and map geometric raster data;

[0107] S302: Based on the map data set, pre-train a pre-constructed initial map data generation model to obtain multiple intermediate map data generation models;

[0108] S303: Based on the map data set, use the map data reward model to perform reinforcement learning training on the multiple intermediate map data generation models in parallel to obtain a map data generation model.

[0109] In the embodiments of the present application, the map feature raster data in the map sample in the above step S301 can be obtained by converting pre-acquired map feature vector data. Specifically, the map feature vector data can be from crowdsourced map data. The map feature vector data is composed of geometric elements such as points, lines, and planes and related attribute data, and can be represented in vector form. The map feature raster data is a raster image composed of pixel points, and each pixel point has a corresponding color or attribute value. The process of converting map feature vector data into map feature raster data specifically involves converting elements such as points, lines, and planes on the vector data into pixels on the raster, and assigning appropriate colors or attribute values to each pixel. In the specific implementation process, those skilled in the art can choose to use GIS software or programming libraries to convert vector data into raster data, and use the rasterization function provided in the GIS software to convert multiple groups of pre-trained map feature vector data obtained into pre-trained map feature raster data.

[0110] In the embodiments of the present application, the map geometric raster data corresponding to the map element raster data in the map sample in step S301 above may be generated by pre-annotation based on manual work. The method of annotating the map element raster data may refer to the detailed description of the prior art and will not be elaborated here.

[0111] In step S302 above, based on the map data set, pre-train the initially constructed initial map data generation model to obtain multiple intermediate map data generation models, which can be specifically implemented in the following manner:

[0112] Divide the map data set into a training set and a test set;

[0113] Construct an initial map data generation model;

[0114] Input the training set into the initial map data generation model, minimize the loss function of the initial map data generation model, calculate the training set accuracy, and obtain the trained map data generation model;

[0115] Input the test set into the trained map data generation model and calculate the test set accuracy;

[0116] Repeat the model training step, and retain the trained map data generation model every first preset threshold number of iterations during the training process to obtain the multiple intermediate map data generation models.

[0117] In the embodiments of the present application, during the specific execution process, different models can be selected to construct the initial map data generation model. For example, a generative adversarial network or a Transformer can be selected.

[0118] In step S303 above, based on the map data set, use the map data reward model to perform reinforcement learning training on the multiple intermediate map data generation models in parallel. The specific process of obtaining the map data generation model may include:

[0119] For each intermediate map data generation model, input the map element raster data in the map data set into the intermediate map data generation model, and use the reinforcement learning algorithm to optimize the intermediate map data generation model based on the map data reward model to obtain the corresponding optimized map data generation model, and store the map geometric raster data output by the intermediate map data generation model;

[0120] Repeat the above steps of reinforcement learning optimization until each optimized map data generation model reaches the optimization termination condition to obtain multiple reinforced map data generation models;

[0121] The map data generation model is screened from multiple enhanced map data generation models.

[0122] In the embodiments of the present application, the reinforcement learning algorithm described can select algorithms such as Q-Learning, DQN (Deep Q Network), etc. according to requirements.

[0123] In the embodiments of the present application, after the map geometry raster data output by the above intermediate map data generation model is stored, it can be supplemented as a sample of the map data reward model to update the above map data reward model.

[0124] In a specific embodiment, for each of the above intermediate map data generation models, the map feature raster data in the map dataset is input into the intermediate map data generation model, and based on the map data reward model, the intermediate map data generation model is optimized using the reinforcement learning algorithm to obtain the corresponding optimized map data generation model. Specifically, it includes:

[0125] For each intermediate map data generation model, the map feature raster data is input into the intermediate map data generation model to obtain the target map geometry raster data;

[0126] Based on the map data reward model, evaluate the reward scores of the map feature raster data and the target map geometry raster data;

[0127] According to the reward scores, update the parameters of the intermediate map data generation model to obtain the corresponding optimized map data generation model.

[0128] In the embodiments of the present application, the above optimization termination condition can be that the number of reinforcement learning iterations reaches a preset threshold, or the change range of the reward score obtained by the reward model does not exceed a preset range.

[0129] The map data generation model training method provided by the embodiments of the present application uses the above map data reward model to evaluate the accuracy and quality of the generated map geometry raster data during the reinforcement learning optimization process. Based on the map data reward model trained by implementing the method of extracting the importance matrix, it can well adapt to the changes in different road layouts and features, effectively improve the accuracy of the map data reward model, and improve the accuracy of reinforcement learning. Therefore, the reward scores output by using this map data reward model can better guide the optimization of the intermediate map data generation model in reinforcement learning.

[0130] In a specific embodiment, the process of screening the map data generation model from the multiple enhanced map data generation models includes:

[0131] Use the raster data of the test map elements to test the multiple enhanced map data generation models respectively, and select the optimal map data generation model according to the test results.

[0132] In the embodiments of the present application, the above-mentioned raster data of the test map elements is the raster data of the map elements that do not belong to the map data set, and the raster data of the test map elements can be re-converted from the newly obtained vector data of the map elements.

[0133] In a specific embodiment, after obtaining the map data generation model, the method further includes: if the test result of testing the map data generation model with the raster data of the test map elements does not meet the preset requirements, then re-execute the above-mentioned steps of reinforcement learning optimization; or,

[0134] If the test result of testing the map data generation model with the raster data of the test map elements does not meet the preset requirements, and it is determined that the quality of the map data generation model is unqualified manually, then re-obtain a new map data set, re-execute the model training steps of S302 above, and obtain a new plurality of intermediate map data generation models.

[0135] The map data generation model trained by the embodiments of the present application can be used to process the vector data of the map elements to be processed to obtain the raster data of the map geometry. Specifically, the vector data of the map elements to be processed is converted into the raster data of the map elements and input into the map data generation model to obtain the corresponding raster data of the map geometry.

[0136] After obtaining the above-mentioned raster data of the map geometry, those skilled in the art can further convert the raster data of the map geometry into the vector data of the map geometry. The obtained vector data of the map geometry can be used to generate a high-precision map. The specific implementation process of converting the raster data of the map geometry into the vector data of the map geometry can refer to the detailed description of the prior art and will not be specifically limited here.

[0137] The map data generation model training method provided by the embodiments of the present application, based on the optimization method of reinforcement learning, can optimize the process of model training according to human feedback. Each intermediate map data generation model can be trained in a distributed and parallel manner, improving the model training efficiency, realizing rapid iterative optimization, selecting the best map data generation model from the obtained multiple optimized map data generation models, which is beneficial to improving the accuracy of the map data generation model, obtaining the raster data of the map geometry that meets the quality requirements, and further obtaining the vector data of the map geometry with higher quality for generating a high-precision map.

[0138] Embodiment III

[0139] Based on the same inventive concept, an embodiment of the present invention further provides a map data reward model training device. Referring to Figure 4 as shown, the device includes:

[0140] A first acquisition module 101, configured to acquire a map feature data set; each sample of the map feature data set includes a map data pair and corresponding real map geometric raster data; the map data pair includes map feature raster data and generated map geometric raster data generated based on the map feature raster data;

[0141] A ranking module 102, configured to score and rank the generated map geometric raster data in each of the samples to obtain the scoring and sorting results of the generated map geometric raster data of each sample;

[0142] An extraction module 103, configured to extract the importance matrix corresponding to the real map geometric raster data in each of the samples;

[0143] A first training module 104, configured to train a pre-constructed reward model based on the scoring and sorting results, the map data pairs in each of the samples, and the importance matrix corresponding to each of the samples to obtain a map data reward model.

[0144] Embodiment Four

[0145] Based on the same inventive concept, an embodiment of the present invention further provides a map data generation model training device. Referring to Figure 5 as shown, the device includes:

[0146] A second acquisition module 201, configured to acquire a map data set; each map sample in the map data set includes map feature raster data and map geometric raster data;

[0147] A generation module 202, configured to pre-train a pre-constructed initial map data generation model based on the map data set to obtain a plurality of intermediate map data generation models;

[0148] A third acquisition module 203, configured to acquire a map data reward model obtained by the map data reward model training method described in the first embodiment above;

[0149] A second training module 204, configured to perform reinforcement learning training on the plurality of intermediate map data generation models in parallel based on the map data set using the map data reward model to obtain a map data generation model.

[0150] Embodiment Five

[0151] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the map data reward model training method described in the first embodiment above, and / or the map data generation model training method described in the second embodiment above.

[0152] Embodiment Six

[0153] Based on the same inventive concept, an embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the map data reward model training method described in the first embodiment above, and / or the map data generation model training method described in the second embodiment above.

[0154] Embodiment Seven

[0155] Based on the same inventive concept, an embodiment of the present invention further provides a computer program product including instructions. When the computer program product runs on a computer device, it causes the computer device to execute the map data reward model training method described in the first embodiment above, and / or the map data generation model training method described in the second embodiment above.

[0156] Embodiment Eight

[0157] Based on the same inventive concept, an embodiment of the present invention further provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a computer program or instructions to implement the map data reward model training method described in the first embodiment above, and / or the map data generation model training method described in the second embodiment above.

[0158] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that include computer-usable program code.

[0159] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks or multiple blocks.

[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks or multiple blocks.

[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks or multiple blocks.

[0162] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for training a map data reward model, characterized in that Including: Obtain a map element data set; each sample of the map element data set includes a map data pair and corresponding real map geometry raster data; the map data pair includes map element raster data and generated map geometry raster data generated based on the map element raster data; Score the generated map geometry raster data in each of the samples to obtain the scores and sorting results of the generated map geometry raster data of each sample; Extract the importance matrix corresponding to the real map geometry raster data in each sample; Train a pre-constructed reward model based on the score and sorting results, the map data pair in each sample, and the importance matrix corresponding to each sample to obtain a map data reward model.

2. The method according to claim 1, wherein The training of the pre-constructed reward model based on the score and sorting results, the map data pair in each sample, and the importance matrix corresponding to each sample to obtain a map data reward model includes: According to the scores of the generated map geometry raster data in each sample, perturb the sorting results within a preset score threshold range to obtain the scores corresponding to the generated map geometry raster data in each sample; Input the map data pair in each sample into the pre-constructed reward model to obtain the corresponding predicted score and predicted importance matrix; For each sample, calculate the total loss according to the predicted score, the predicted importance matrix, the score, and the importance matrix; Update the pre-constructed reward model according to the total loss to obtain an updated reward model; Repeat the above steps of model update until the updated reward model meets the preset conditions to obtain a map data reward model.

3. The method according to claim 2, characterized in that The calculating the total loss according to the predicted score, the predicted importance matrix, the score, and the importance matrix for each sample includes: Calculate the regression loss based on the predicted score and the score; Calculate the local weight loss based on the predicted importance matrix and the importance matrix; Add the regression loss and the local weight loss after weighting to obtain the total loss.

4. The method according to claim 2, wherein The perturbing the sorting results within a preset score threshold range according to the scores of the generated map geometry raster data in each sample to obtain the scores corresponding to the generated map geometry raster data in each sample includes: According to the scores of the generated map geometry raster data in each sample, perturb the sorting results within a preset score threshold range to obtain a new sorting result; According to the new sorting result, obtain the scores corresponding to the generated map geometry raster data in each sample through normalization.

5. The method according to any one of claims 1-4, characterized in that The extracting the importance matrix corresponding to the real map geometry raster data in each sample includes: According to the preset importance matrix size, divide the real map geometry raster data in each sample into multiple blocks; Count the number of map elements in each block and extract the map element matrix corresponding to the real map geometry raster data in the sample; For each sample, process based on the map element matrix to obtain the corresponding importance matrix.

6. A method for training a map data generation model, characterized in that, A map data reward model obtained by using the map data reward model training method according to any one of claims 1-5; the method includes: Obtain a map data set; each map sample in the map data set includes map feature raster data and map geometry raster data; Based on the map data set, pre-train a pre-constructed initial map data generation model to obtain a plurality of intermediate map data generation models; Based on the map data set, use the map data reward model to perform reinforcement learning training on the plurality of intermediate map data generation models in parallel to obtain a map data generation model.

7. A training device for a map data reward model, characterized in that Includes: A first acquisition module for acquiring a map feature data set; each sample in the map feature data set includes a map data pair and corresponding true map geometry raster data; the map data pair includes map feature raster data and generated map geometry raster data generated based on the map feature raster data; A ranking module for scoring and ranking the generated map geometry raster data in each of the samples to obtain the scoring and sorting results of the generated map geometry raster data in each sample; An extraction module for extracting the importance matrix corresponding to the true map geometry raster data in each of the samples; A first training module for training a pre-constructed reward model based on the scoring and sorting results, the map data pairs in each of the samples, and the importance matrix corresponding to each of the samples to obtain a map data reward model.

8. A training device for a map data generation model, characterized in that Includes: A second acquisition module for acquiring a map data set; Each map sample in the map data set includes map feature raster data and map geometry raster data; A generation module for pre-training a pre-constructed initial map data generation model based on the map data set to obtain a plurality of intermediate map data generation models; A third acquisition module for acquiring a map data reward model obtained by using the map data reward model training method according to any one of claims 1-5; A second training module for performing reinforcement learning training on the plurality of intermediate map data generation models in parallel based on the map data set and using the map data reward model to obtain a map data generation model.

9. A computer-readable storage medium storing instructions that, when executed on a terminal, cause the terminal to execute the map data reward model training method according to any one of claims 1-5, and / or, the map data generation model training method according to claim 6.

10. A computer device, characterized in that, Includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the map data reward model training method according to any one of claims 1-5, and / or, the map data generation model training method according to claim 6.

11. A computer program product containing instructions that, when the computer program product runs on a computer device, causes the computer device to execute the map data reward model training method according to any one of claims 1-5, and / or, the map data generation model training method according to claim 6.

12. A chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run a computer program or instructions to implement the method for training a map data reward model according to any one of claims 1-5, and / or the method for training a map data generation model according to claim 6.