Online map generation model training method and device, online map generation method and device, electronic equipment and storage medium

Through the online map generation model training method, the map generation model is extracted and optimized, and the problems of incomplete and mismatch elements in high-precision map generation are solved, achieving more efficient map generation effect.

CN120014106APending Publication Date: 2025-05-16MUSHROOM CHELIAN INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510177353.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the process of real-time map generation of high-precision maps, problems of incomplete and mismatch of generated elements are often encountered, which affects the overall effect of the map.

Method used

The online map generation model training method is adopted. By inputting the data to be trained, it is passed through the network of the online map encoder and decoder respectively, extracting features and calculating the loss value, optimizing the network training to convergence, and obtaining the online map generation model.

Benefits of technology

Improve the integrity and matching accuracy of generated map elements, improve the overall effect of map generation, and solve the problems of incompleteness and mismatch of generated elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014106A_ABST
    Figure CN120014106A_ABST
Patent Text Reader

Abstract

The invention discloses an online map generation model training method and device, an online map generation method and device, electronic equipment and a storage medium. The training method comprises the steps of inputting to-be-trained data; enabling the to-be-trained data to respectively pass through a network at least comprising an online map encoder Map Encoder and an online map decoder Map Decoder to obtain a first feature and a second feature; determining an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder; and determining a final loss value LOSS according to the initial loss value LOSS, and training the network until convergence to obtain an online map generation model. According to the online map generation model training method, the training convergence time can be shortened, and meanwhile, the integrity of the generation elements and the accuracy of the generation elements can be guaranteed. Through the online map generation method provided by the invention, the overall effect of map generation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of high-precision map real-time map generation, and in particular to an online map generation model training method, an online map generation method, an apparatus and electronic equipment, and a storage medium. Background Art

[0002] In high-precision map processing, real-time map generation is usually required.

[0003] However, in real-time map generation, problems such as incomplete generated elements and mismatched generated elements are often encountered, which affects the overall effect of the generated map. Summary of the invention

[0004] The embodiments of the present application provide an online map generation model training method, an online map generation method, an apparatus and an electronic device, and a storage medium to achieve the integrity and accuracy of generated elements, thereby improving the overall effect of generated maps.

[0005] The present application embodiment adopts the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides an online map generation model training method, wherein the model training method comprises:

[0007] Input the data to be trained;

[0008] Passing the to-be-trained data through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder, respectively, to obtain a first feature and a second feature;

[0009] Determine an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder;

[0010] A final loss value LOSS is determined according to the initial loss value LOSS, and the network is trained until convergence to obtain an online map generation model.

[0011] In some embodiments, the data to be trained includes:

[0012] The image sensor is used to collect information of four images: front left, front right, rear left, and rear right;

[0013] The intrinsic parameters, extrinsic parameters of the image sensor and the four image information of the front left, front right, rear left and rear right are used as the data to be trained.

[0014] In some embodiments, the step of passing the to-be-trained data through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder to obtain the first feature and the second feature comprises:

[0015] Converting the front left, front right, rear left and rear right four image information into BEV data according to the online map encoder;

[0016] Using a preset attention mechanism, extracting common useful information from the BEV data according to the network of the online map decoder, and then outputting a BEV feature as the first feature;

[0017] After performing multi-layer convolution operations on the BEV feature, the preset attention mechanism is used to extract the feature information of the four images of the front left, front right, back left, and back right and use it as the second feature.

[0018] In some embodiments, determining an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder comprises:

[0019] A loss function is used to calculate an initial loss value LOSS between the first feature and the second feature output by the online map encoder MapEncoder and the online map decoder MapDecoder.

[0020] In some embodiments, determining a final loss value LOSS according to the initial loss value LOSS, training the network until convergence to obtain an online map generation model, comprises:

[0021] According to the initial loss value LOSS, the final loss value LOSS is calculated by back propagation and optimized to obtain an online map generation model.

[0022] In a second aspect, an embodiment of the present application further provides an online map generation method, wherein the online map generation method comprises:

[0023] Using the online map generation model training method described in the first aspect to obtain an online map generation model;

[0024] The input includes at least image information in four directions: front left, front right, back left, and back right;

[0025] The online map data is output according to the online map generation model.

[0026] In some embodiments, outputting online map data includes:

[0027] Online map data including at least lane lines, lane boundary lines, and lane indication signs.

[0028] In a third aspect, an embodiment of the present application further provides an online map generation model training device, wherein the device comprises:

[0029] Input module, used to input data to be trained;

[0030] A processing module is used to pass the to-be-trained data through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder to obtain a first feature and a second feature,

[0031] A loss determination module, configured to determine an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder;

[0032] The loss iteration module is used to determine the final loss value LOSS according to the initial loss value LOSS, and train the network until convergence to obtain an online map generation model.

[0033] In a fourth aspect, an embodiment of the present application further provides an electronic device, comprising: a processor; and a memory arranged to store computer executable instructions, wherein the executable instructions, when executed, cause the processor to perform the above method.

[0034] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device executes the above method.

[0035] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: input the data to be trained, and pass the data to be trained through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder, respectively, to obtain a first feature and a second feature. Further, determine the initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder, and finally, determine the final loss value LOSS based on the initial loss value LOSS, and train the network until convergence to obtain an online map generation model. Through the above method, the integrity of the generated map elements and the matching accuracy of the map generation elements are improved, thereby improving the overall effect of the generated map. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0037] Figure 1 This is a schematic diagram of the network structure of HDMapNet used in related technologies;

[0038] Figure 2 A schematic diagram of the process of an online map generation model training method in an embodiment of the present application;

[0039] Figure 3 Schematic diagram of the implementation principle of the online map generation model training method in the embodiment of the present application;

[0040] Figure 4 This is a schematic diagram of the structure of an online map generation model training device in an embodiment of the present application;

[0041] Figure 5 This is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0043] When using traditional online local map generation methods, most of them first use semantic segmentation based on multiple perspectives, then convert it to a bird's-eye view through IPM projection and splicing, and finally obtain the required map elements and topological structure through relatively complex post-processing. The maps generated in this way are very inaccurate and difficult to use in actual business.

[0044] In other solutions, based on the BEV paradigm, BEV features are first obtained through various methods, and then subsequent operations are performed on the BEV features, making the whole process much simpler. HDMapNet is a relatively early BEV online local map construction solution. The specific network structure is as follows Figure 1 shown.

[0045] When the inventor used HDMapNet to test relevant data sets (actual business related), it was found that its generation effect was very poor. Specifically, the poor effect was mainly reflected in: incomplete lane edge recognition, generation element errors and slow training convergence. Due to the above problems, this algorithm is difficult to use for online map generation.

[0046] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0047] The present application embodiment provides an online map generation model training method, such as Figure 2 As shown, a flow chart of the online map generation model training method in an embodiment of the present application is provided, and the method at least includes the following steps S210 to S240:

[0048] Step S210: input the data to be trained.

[0049] The data to be trained is relevant data that needs to be used for model training. Usually, the data to be trained includes but is not limited to the perception data collected by the vehicle's sensors.

[0050] Step S220 , the data to be trained are respectively passed through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder to obtain a first feature and a second feature.

[0051] like Figure 3 As shown, after the training data is processed by the online map encoder MapEncoder and the online map decoder MapDecoder, the first feature and the second feature can be obtained respectively. Figure 3 As shown, MapEncoder outputs the first feature and MapDecoder outputs the second feature.

[0052] It can be understood that the first feature and the second feature mainly refer to the extracted features of the image in different dimensions.

[0053] Step S230: determining an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder.

[0054] Based on the first feature and the second feature, a loss value LOSS between the features respectively output by the online map encoder and the online map decoder may be further determined.

[0055] It can be understood that the loss value is used to measure the difference between the model prediction result and the true label. The high or low loss value reflects the degree of fit of the model to the training data under the current parameters.

[0056] For example, a low loss value indicates that the difference between the model's prediction results and the true labels is small, that is, the model fits the training data well. On the contrary, a high loss value indicates that the difference between the model's prediction results and the true labels is large, that is, the model fits the training data poorly, which may indicate that the model has problems such as underfitting or overfitting, and further adjustment of the model structure or optimization algorithm is required. At the same time, the loss value is an important indicator for measuring model performance in deep learning. By continuously optimizing the loss value, the prediction accuracy and generalization ability of the model can be improved.

[0057] Step S240, determining a final loss value LOSS according to the initial loss value LOSS, and training the network until convergence to obtain an online map generation model.

[0058] The final generation effect is improved by calculating the loss between MapEncoder and MapDecoder. Therefore, the final loss value LOSS is obtained after continuous iteration based on the initial loss value LOSS. The process of determining the final loss value LOSS is the process of training the network until convergence to obtain the online map generation model. The online map generation model is obtained through training through the above steps.

[0059] Through the above method, the amount of training data is reduced by optimizing the training data. At the same time, the relevant mechanism is used in the network to optimize the map feature extraction process, and finally, the loss value between the first feature and the second feature output by the online map encoder and the online map decoder is continuously optimized to improve the overall generation effect of the online map.

[0060] Through the above method, the data to be trained is passed through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder respectively to obtain a first feature and a second feature. By continuously fusing feature information, the integrity of map generation elements and the matching accuracy of map generation elements are improved.

[0061] Different from the related technologies, when using the existing HDMapNet to process the actual data set, it is easy to have incomplete lane edge recognition, incorrect element generation and slow training convergence. The above method can not only overcome the problem of incomplete map element generation, but also improve the speed of training convergence.

[0062] In one embodiment of the present application, the data to be trained includes: information of four images of front left, front right, rear left and rear right acquired through an image sensor; and the intrinsic parameters, extrinsic parameters and the four images of front left, front right, rear left and rear right of the image sensor are used as the data to be trained.

[0063] The image sensor is used to collect information of four images, namely, the front left, the front right, the rear left, and the rear right, and the internal and external parameters (pre-calibrated) in the image sensor are used as input as the data to be trained.

[0064] Different from the related art, which uses six pictures of front left, front, front right, rear left, rear, and rear right, the embodiment of the present application uses four pictures of front left, front right, rear left, and rear right. Considering that the four directions of front left, front right, rear left, and rear right can cover the whole view, the four pictures of front left, front right, rear left, and rear right can be used directly by modifying the input channel and the corresponding internal and external parameters of the camera in the input part of the model.

[0065] In one embodiment of the present application, the data to be trained is passed through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder to obtain a first feature and a second feature, including: converting the four image information of the front left, front right, rear left, and rear right into BEV data according to the online map encoder; using a preset attention mechanism, extracting common useful information in the BEV data according to the network of the online map decoder, and then outputting BEV features as the first feature; after performing multi-layer convolution operations on the BEV features, using the preset attention mechanism to extract feature information of the four image information of the front left, front right, rear left, and rear right and use it as the second feature.

[0066] like Figure 3 As shown, the data to be trained is input into the network, and the camera internal and external parameters are configured during the network training. Specifically:

[0067] First, the network converts the front left, front right, rear left, and rear right image data into BEV data, that is, into a bird's-eye view.

[0068] The BEV data obtained above is used to extract common useful information using a three-layer transformer attention mechanism and then output as BEV features.

[0069] Finally, after performing multi-layer conv convolution operations on the BEV features obtained above, the transformer attention mechanism is used to extract information, that is, the feature information of the front left, front right, back left, and back right images is obtained.

[0070] It can be understood that the BEV (Bird's Eye View) feature, that is, the bird's-eye view feature, is a feature representation method commonly used in the fields of autonomous driving and computer vision.

[0071] It can be understood that the transformer attention mechanism is a technology widely used in the fields of natural language processing (NLP) and deep learning. It allows the model to dynamically assign weights when processing sequence data and focus on the most important parts, thereby achieving effective screening and integration of information. Using the transformer attention mechanism, by dynamically assigning weights to focus on important information, the efficiency and effect of the model when processing sequence data are greatly improved. Preferably, in the embodiments of the present application, the transformer attention mechanism adopts a three-layer attention mechanism.

[0072] In one embodiment of the present application, determining an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder includes: using a loss function to calculate the initial loss value LOSS between the first feature and the second feature output by the online map encoder MapEncoder and the online map decoder MapDecoder.

[0073] Preferably, in the embodiment of the present application, mse is used to calculate the loss between MapEncoder and MapDecoder.

[0074] It can be understood that the mse loss function is only an example, and other loss functions can be used in the embodiments of the present application, and are not specifically limited in the embodiments of the present application.

[0075] In one embodiment of the present application, the final loss value LOSS is determined according to the initial loss value LOSS, and the network is trained until convergence to obtain an online map generation model, including: according to the initial loss value LOSS, the final loss value LOSS is calculated by back propagation and optimized to obtain the online map generation model.

[0076] Based on the above structure, the network convergence speed can be accelerated by calculating the loss through back propagation and optimizing the final loss value LOSS.

[0077] Backpropagation is one of the core algorithms for training neural networks. It is used to calculate the gradients of each weight in the neural network and update the weights through optimization methods such as gradient descent to minimize the loss function of the neural network.

[0078] Specifically, the backpropagation process includes: (1) Forward propagation: The neural network receives input data and calculates the output value through forward propagation. The output value is usually the result of processing with a certain activation function. Then, the error between the predicted result and the true label is calculated through the loss function. (2) Calculate the gradient of the output layer: Calculate the gradient of the loss function with respect to the output layer activation value, and then calculate the gradient of the output layer activation value with respect to the output layer weight. (3) Backpropagate the error layer by layer: Starting from the output layer, use the chain rule to backpropagate the error layer by layer to the previous layer until it propagates to the input layer. The gradient of each layer can be calculated layer by layer using the chain rule.

[0079] An embodiment of the present application also provides an online map generation method, wherein the online map generation method comprises: adopting the online map generation model training method to obtain an online map generation model; inputting image information including at least four directions: front left, front right, rear left, and rear right; and outputting online map data according to the online map generation model.

[0080] The input includes at least image information in four directions: front left, front right, rear left, and rear right. It means that in the input image information, the orientation images are images in the four directions: front left, front right, rear left, and rear right. It may also include other road information images.

[0081] The online map generation model has the characteristics of completeness of generated map elements (such as lane edges) and matching accuracy of map generation elements (such as arrow element errors), thereby improving the overall effect of generated maps.

[0082] By adopting the online map generation model training method, an online map generation model is obtained, and according to the input image information in four directions of front left, front right, rear left and rear right, online map data can be output according to the online map generation model.

[0083] In one embodiment of the present application, the output online map data includes: online map data including at least lane lines, lane boundary lines, and lane indication marks.

[0084] The online map generation model is used to input image information in four directions: front left, front right, rear left, and rear right, and output online map data including at least lane lines, lane boundary lines, and lane indication marks.

[0085] It can be understood that the lane indication signs include but are not limited to road signs such as arrows and lane line information.

[0086] The present application embodiment also provides an online map generation model training device 400, such as Figure 4As shown, a schematic diagram of the structure of an online map generation model training device in an embodiment of the present application is provided, wherein the online map generation model training device 400 at least includes: an input module 410, a processing module 420, a loss determination module 430, and a loss iteration module 440, wherein:

[0087] In one embodiment of the present application, the input module 410 is specifically used to: input data to be trained.

[0088] The data to be trained is relevant data that needs to be used for model training. Usually, the data to be trained includes but is not limited to the perception data collected by the vehicle's sensors.

[0089] In one embodiment of the present application, the processing module 420 is specifically used to: pass the to-be-trained data through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder, respectively, to obtain the first feature and the second feature.

[0090] like Figure 3 As shown, after the training data is processed by the online map encoder MapEncoder and the online map decoder MapDecoder, the first feature and the second feature can be obtained respectively. Figure 3 As shown, MapEncoder outputs the first feature and MapDecoder outputs the second feature.

[0091] It can be understood that the first feature and the second feature mainly refer to the extracted features of the image in different dimensions.

[0092] In one embodiment of the present application, the loss determination module 430 is specifically used to determine an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder.

[0093] Based on the first feature and the second feature, a loss value LOSS between the features respectively output by the online map encoder and the online map decoder may be further determined.

[0094] It can be understood that the loss value is used to measure the difference between the model prediction result and the true label. The high or low loss value reflects the degree of fit of the model to the training data under the current parameters.

[0095] For example, a low loss value indicates that the difference between the model's prediction results and the true labels is small, that is, the model fits the training data well. On the contrary, a high loss value indicates that the difference between the model's prediction results and the true labels is large, that is, the model fits the training data poorly, which may indicate that the model has problems such as underfitting or overfitting, and further adjustment of the model structure or optimization algorithm is required. At the same time, the loss value is an important indicator for measuring model performance in deep learning. By continuously optimizing the loss value, the prediction accuracy and generalization ability of the model can be improved.

[0096] In one embodiment of the present application, the loss iteration module 440 is specifically used to: determine the final loss value LOSS according to the initial loss value LOSS, and train the network until convergence to obtain an online map generation model.

[0097] The final generation effect is improved by calculating the loss between MapEncoder and MapDecoder. Therefore, the final loss value LOSS is obtained after continuous iteration based on the initial loss value LOSS. The process of determining the final loss value LOSS is the process of training the network until convergence to obtain the online map generation model. The online map generation model is obtained through training through the above steps.

[0098] In one embodiment of the present application, the data to be trained includes:

[0099] The image sensor is used to collect information of four images: front left, front right, rear left, and rear right;

[0100] The intrinsic parameters, extrinsic parameters of the image sensor and the four image information of the front left, front right, rear left and rear right are used as the data to be trained.

[0101] In one embodiment of the present application, the processing module 420 is also used to

[0102] Converting the front left, front right, rear left and rear right four image information into BEV data according to the online map encoder;

[0103] Using a preset attention mechanism, extracting common useful information from the BEV data according to the network of the online map decoder, and then outputting a BEV feature as the first feature;

[0104] After performing multi-layer convolution operations on the BEV feature, the preset attention mechanism is used to extract the feature information of the four images of the front left, front right, back left, and back right and use it as the second feature.

[0105] In one embodiment of the present application, the loss determination module 430 is also used to

[0106] A loss function is used to calculate an initial loss value LOSS between the first feature and the second feature output by the online map encoder MapEncoder and the online map decoder MapDecoder.

[0107] In one embodiment of the present application, the loss iteration module 440 is also used to

[0108] According to the initial loss value LOSS, the final loss value LOSS is calculated by back propagation and optimized to obtain an online map generation model.

[0109] It can be understood that the above-mentioned online map generation model training device can implement the various steps of the online map generation model training method provided in the aforementioned embodiment, and the relevant explanations about the online map generation model training method are applicable to the online map generation model training device, which will not be repeated here.

[0110] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 5 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. Of course, the electronic device may also include hardware required for other services.

[0111] The processor, network interface and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0112] The memory is used to store the program. Specifically, the program may include a program code, and the program code includes a computer operation instruction. The memory may include a memory and a non-volatile memory, and provides instructions and data to the processor.

[0113] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming an online map generation model training device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0114] Input the data to be trained;

[0115] Passing the to-be-trained data through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder, respectively, to obtain a first feature and a second feature;

[0116] Determine an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder;

[0117] A final loss value LOSS is determined according to the initial loss value LOSS, and the network is trained until convergence to obtain an online map generation model.

[0118] The above application Figure 2 The method performed by the online map generation model training device disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or instructions in the form of software. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in a decoding processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0119] The electronic device may also perform Figure 2 The method for executing the online map generation model training device in the Figure 2 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0120] The present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, enable the electronic device to execute Figure 1 The method performed by the online map generation model training device in the illustrated embodiment is specifically used to perform:

[0121] Input the data to be trained;

[0122] Passing the to-be-trained data through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder, respectively, to obtain a first feature and a second feature;

[0123] Determine an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder;

[0124] A final loss value LOSS is determined according to the initial loss value LOSS, and the network is trained until convergence to obtain an online map generation model.

[0125] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0126] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0127] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0129] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0130] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0131] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0132] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0133] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0134] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A method for training an online map generation model, wherein: The model training method comprises: Input the data to be trained; Passing the to-be-trained data through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder, respectively, to obtain a first feature and a second feature; Determine an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder; A final loss value LOSS is determined according to the initial loss value LOSS, and the network is trained until convergence to obtain an online map generation model.

2. The method of claim 1, wherein: The data to be trained includes: The image sensor is used to collect information of four images: front left, front right, rear left, and rear right; The intrinsic parameters, extrinsic parameters of the image sensor and the four image information of the front left, front right, rear left and rear right are used as the data to be trained.

3. The method of claim 1, wherein: The step of passing the to-be-trained data through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder to obtain a first feature and a second feature comprises: Converting the front left, front right, rear left and rear right four image information into BEV data according to the online map encoder; Using a preset attention mechanism, extracting common useful information from the BEV data according to the network of the online map decoder, and then outputting a BEV feature as the first feature; After performing multi-layer convolution operations on the BEV feature, the preset attention mechanism is used to extract the feature information of the four images of the front left, front right, back left, and back right and use it as the second feature.

4. The method of claim 1, wherein: Determining an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder comprises: A loss function is used to calculate an initial loss value LOSS between the first feature and the second feature output by the online map encoder MapEncoder and the online map decoder MapDecoder.

5. The method of claim 4, wherein: Determining a final loss value LOSS according to the initial loss value LOSS, training the network until convergence to obtain an online map generation model, including: According to the initial loss value LOSS, the final loss value LOSS is calculated by back propagation and optimized to obtain an online map generation model.

6. An online map generation method, wherein: The online map generation method comprises: Using the online map generation model training method according to any one of claims 1 to 5 to obtain an online map generation model; The input includes at least image information in four directions: front left, front right, back left, and back right; According to the online map generation model, online map data is output.

7. The method according to any one of claims 1 to 5, wherein: The outputting online map data comprises: Online map data including at least lane lines, lane boundary lines, and lane indication signs.

8. An online map generation model training device, wherein: The device comprises: Input module, used to input data to be trained; A processing module is used to pass the to-be-trained data through a network including at least an online map encoder MapEncoder and an online map decoder MapDecoder to obtain a first feature and a second feature, A loss determination module, configured to determine an initial loss value LOSS between the first feature and the second feature output by both the online map encoder and the online map decoder; The loss iteration module is used to determine the final loss value LOSS according to the initial loss value LOSS, and train the network until convergence to obtain an online map generation model.

9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which, when executed, cause the processor to perform the method of any one of claims 1 to 5 and / or the method of any one of claims 6 to 7.

10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, enables the electronic device to execute any one of the methods of claims 1 to 7 and / or any one of the methods of claims 6 to 7.

Citation Information

Cited By

  • Vector map generation method and device, equipment, storage medium and program product

    CN120702447A