Training method and device of positioning and mapping model, electronic equipment and storage medium

CN117710456BActive Publication Date: 2026-09-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311676389.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2026-09-04
Estimated Expiration
2043-12-07

AI Technical Summary

Technical Problem

这种方法得到的定位和建图模型,主要依赖于训练采用的图像识别数据集的规模及标注结果,成本高,效率低

Benefits of technology

[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117710456B_ABST
    Figure CN117710456B_ABST
Patent Text Reader

Abstract

The present disclosure provides a positioning and mapping model training method and device, electronic equipment and storage medium, relating to the technical field of computers, especially to the technical field of artificial intelligence such as automatic driving, deep learning and computer vision, which can be applied to outdoor roads, automatic driving, intelligent auxiliary driving and other scenes. The specific scheme is: after obtaining the initial data set, first, based on the image data, map data and vehicle pose in each initial data, the corresponding prompt information is generated, then each image data and its corresponding prompt information is input into the preset large model, the first label of the road element contained in each image data output by the large model is obtained, and then the initial positioning and mapping model is trained by using each image data, each point cloud data and the first label of the road element contained in each image data, to obtain the main network in the positioning and mapping model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to the fields of artificial intelligence technology such as autonomous driving, deep learning, and computer vision. Specifically, it relates to a training method, apparatus, electronic device, and storage medium for localization and mapping models. Background Technology

[0002] Currently, localization and mapping models in autonomous driving scenarios typically use visual models as the backbone network, and these visual models are usually trained on large-scale image recognition datasets. This approach results in localization and mapping models that heavily rely on the size and annotation results of the image recognition datasets used for training, leading to high costs and low efficiency. Summary of the Invention

[0003] This disclosure aims to at least partially address one of the technical problems in the related art.

[0004] The first aspect of this disclosure proposes a method for training a localization and mapping model, comprising:

[0005] Obtain an initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data, and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element;

[0006] Based on the image data, map data, and vehicle pose in each initial data set, corresponding prompt information is generated;

[0007] Each image data and its corresponding prompt information are input into a preset large model, and the first label of the road element contained in each image data output by the large model is obtained;

[0008] Using the image data, the point cloud data, and the first label of the road element contained in the image data, an initial localization and mapping model is trained to obtain the main network in the localization and mapping model, wherein the main network is a network used to extract and fuse features from the input image data and point cloud data.

[0009] A second aspect of this disclosure provides a training apparatus for a localization and mapping model, comprising:

[0010] The first acquisition module is used to acquire an initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data, and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element;

[0011] The generation module is used to generate corresponding prompt information based on the image data, map data and vehicle pose in each initial data set;

[0012] The second acquisition module is used to input each image data and its corresponding prompt information into a preset large model, and to acquire the first label of the road element contained in each image data output by the large model.

[0013] The training module is used to train an initial localization and mapping model using the image data, the point cloud data, and the first labels of the road elements contained in the image data, so as to obtain the main network in the localization and mapping model, wherein the main network is a network used to extract and fuse features from the input image data and point cloud data.

[0014] A third aspect of this disclosure provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements a training method for a localization and mapping model as proposed in a first aspect of this disclosure.

[0015] A fourth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a training method for a localization and mapping model as proposed in a first aspect of this disclosure.

[0016] A fifth aspect of this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements a training method for a localization and mapping model as proposed in a first aspect of this disclosure.

[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0018] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0019] Figure 1 This is a schematic flowchart illustrating a training method for a localization and mapping model provided in an embodiment of the present disclosure.

[0020] Figure 2 This is a schematic flowchart illustrating a training method for a localization and mapping model provided in an embodiment of the present disclosure.

[0021] Figure 3 This is a schematic flowchart illustrating a training method for a localization and mapping model provided in an embodiment of the present disclosure.

[0022] Figure 4 A schematic diagram of a road element closed-loop data annotation engine provided in an embodiment of this disclosure;

[0023] Figure 5 This is a schematic flowchart illustrating a training method for a localization and mapping model provided in an embodiment of the present disclosure.

[0024] Figure 6 This is a schematic flowchart illustrating a training method for a localization and mapping model provided in an embodiment of the present disclosure.

[0025] Figure 7 This is a schematic flowchart illustrating a training method for a localization and mapping model provided in an embodiment of the present disclosure.

[0026] Figure 8 A schematic diagram of the main network and multi-task pre-training framework in the localization and mapping model provided in the embodiments of this disclosure;

[0027] Figure 9 This is a schematic diagram of the structure of the training device for the localization and mapping model provided in the embodiments of this disclosure;

[0028] Figure 10 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0030] This disclosure relates to the fields of artificial intelligence technologies such as autonomous driving, deep learning, and computer vision.

[0031] Artificial Intelligence (AI) is a new technological science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence.

[0032] Autonomous driving, also known as driverless driving, computer-controlled driving, or wheeled mobile robots, is a cutting-edge technology that relies on computer and artificial intelligence technologies to achieve complete, safe, and efficient driving without human intervention.

[0033] Deep learning learns the inherent patterns and hierarchical representations of sample data. The information gained during this learning process greatly aids in interpreting data such as text, images, and sound. The ultimate goal of deep learning is to enable machines to possess analytical and learning capabilities similar to humans, allowing them to recognize data such as text, images, and sound.

[0034] Computer vision refers to machine vision that uses cameras and computers to identify, track, and measure targets instead of human eyes, and further processes the images to make them more suitable for human observation or transmission to instruments for detection.

[0035] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information in this technical solution comply with relevant laws and regulations and do not violate public order and good morals.

[0036] The following description, with reference to the accompanying drawings, outlines a method, apparatus, electronic device, and storage medium for training localization and mapping models according to embodiments of the present disclosure.

[0037] Figure 1 This is a flowchart illustrating a training method for a localization and mapping model provided in an embodiment of this disclosure.

[0038] like Figure 1 As shown, the training method for this localization and mapping model may include the following steps:

[0039] Step 101: Obtain the initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data, and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element.

[0040] It should be noted that the initial dataset includes multiple initial datasets, each of which includes associated image data, point cloud data, map data, and vehicle pose. The image data can be surround-view image data collected in autonomous driving scenarios, and may include any data such as images and corresponding camera intrinsic parameters. The vehicle pose can be obtained by offline parsing of high-precision positioning data collected in driving scenarios; this disclosure does not impose any limitations on it.

[0041] Among them, road elements are elements used to represent the geometric features and geographical attributes of a geographical area. For example, road elements can be any elements such as landmarks, traffic signs, lane lines, etc., and this disclosure does not limit them.

[0042] Step 102: Generate corresponding prompt information based on the image data, map data, and vehicle pose in each initial data set.

[0043] It should be noted that the vehicle pose determines the road area contained in the collected point cloud data and image data. This means that the road area (road elements) collected from the same location on the road may differ depending on the vehicle pose. Furthermore, the conversion relationship between point cloud data and image data also differs under different vehicle poses. Therefore, in order to ensure that the large model can accurately output the road elements contained in the image data, this disclosure first requires generating accurate prompts based on the image data, map data, and vehicle pose.

[0044] In this disclosure, after obtaining the initial dataset, map data and vehicle pose can be obtained from the initial dataset. Then, based on the camera intrinsics, map data, and vehicle pose contained in the image data, corresponding prompt information can be generated for each image data.

[0045] Step 103: Input each image data and its corresponding prompt information into the preset large model, and obtain the first label of the road element contained in each image data output by the large model.

[0046] The preset large model is any pre-generated model capable of recognizing road elements. For example, the preset large model can be a segmentation large model, such as a Segment Anything Model (SAM), etc. This disclosure does not limit it.

[0047] The first tag can be a semantic tag corresponding to the road element, used to identify the type of the road element, such as a lane line or a traffic sign.

[0048] In this disclosure, by inputting the image contained in each image data and its corresponding prompt information into a preset large model, the large model can then identify and predict the image based on the prompt information to obtain the semantic label corresponding to the road element contained in the image, i.e., the first label.

[0049] Optionally, the output of the preset large model may also be the probability of the road elements contained in the image belonging to each type of road element. The type with the highest probability value can be used as the first label of the road elements contained in the image, etc. This disclosure does not limit this.

[0050] Step 104: Using the first labels of each image data, each point cloud data, and the road elements contained in each image data, train the initial localization and mapping model to obtain the main network in the localization and mapping model. The main network is a network used to extract and fuse features from the input image data and point cloud data.

[0051] In this disclosure, after obtaining the first labels of road elements contained in each image data output by the large model, the initial localization and mapping model can be trained based on each image data, each point cloud data, and the first labels of road elements contained in each image data to obtain the main network in the localization and mapping model. Since the data used for model training does not require manual annotation but is automatically annotated through the large model, it is not only low-cost but also highly efficient and accurate. Furthermore, when training the localization and mapping model, it is based not only on image data but also on point cloud data and the first labels of road elements, making the trained model more focused on the perception capabilities of road elements and three-dimensional (3D) perception capabilities, resulting in higher accuracy and reliability, and making it more suitable for use in autonomous driving scenarios.

[0052] In this embodiment, after obtaining the initial dataset, corresponding prompts are first generated based on the image data, map data, and vehicle poses in each initial dataset. Then, each image data and its corresponding prompt are input into a preset large model to obtain the first label of the road element contained in each image data output by the large model. Subsequently, the initial localization and mapping model is trained using each image data, each point cloud data, and the first label of the road element contained in each image data to obtain the main network in the localization and mapping model. Thus, by training the initial localization and mapping model based on each image data, each point cloud data, and the first label of the road element contained in each image data, the perception capability and 3D perception capability of the localization and mapping model are improved, thereby improving the accuracy and reliability of the localization and mapping model. Furthermore, the cost of data annotation is reduced by automatically labeling the training data through the large model.

[0053] Figure 2 This is a flowchart illustrating a training method for a localization and mapping model provided in one embodiment of the present disclosure.

[0054] like Figure 2 As shown, the training method for this localization and mapping model may include the following steps:

[0055] Step 201: Obtain the initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data, and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element.

[0056] The specific implementation of step 201 can be found in the detailed descriptions of other embodiments in this disclosure, and will not be repeated here.

[0057] Step 202: For initial data i in the initial dataset, determine the transformation matrix corresponding to the camera based on the vehicle pose in initial data i and the camera intrinsic parameters in the image data.

[0058] Where i is an integer less than or equal to N, and N is the number of initial data points contained in the initial dataset.

[0059] The transformation matrix is ​​used to transform the camera coordinate system into the vehicle coordinate system. It can be determined based on the camera's intrinsic parameters and the camera's extrinsic parameters determined by the vehicle's pose. This disclosure does not limit the specific parameters in this regard.

[0060] It should be noted that a transformation matrix has a one-to-one correspondence with the vehicle pose and the camera's intrinsic parameters.

[0061] Step 203: Collect a preset number of three-dimensional reference points from the vectors corresponding to each road element in the map data i in the initial data i.

[0062] In this disclosure, after obtaining map data i in the initial data i, three-dimensional reference points can be collected from the vectors corresponding to each road element in map data i.

[0063] Step 204: Based on the transformation matrix, map each three-dimensional reference point to obtain the corresponding two-dimensional reference point.

[0064] In this disclosure, after collecting a preset number of three-dimensional reference points from the vectors corresponding to each road element in the map data, the three-dimensional reference points can be projected onto the image plane based on the camera's intrinsic and extrinsic parameters to obtain the corresponding two-dimensional reference points.

[0065] It is understandable that the number of two-dimensional reference points obtained by mapping is the same as the number of three-dimensional reference points.

[0066] Step 205: Dilate the region containing all two-dimensional reference points to obtain auxiliary two-dimensional reference points in the dilated region, excluding the two-dimensional reference points.

[0067] Among them, the auxiliary two-dimensional reference point is the negative sample two-dimensional reference point in the expansion region other than the two-dimensional reference point.

[0068] In this disclosure, in order to alleviate the inaccuracy of the projection position of the two-dimensional reference point caused by the rolling shutter effect and the camera's intrinsic and extrinsic parameter errors, the region where all the two-dimensional reference points are located can be expanded to obtain auxiliary two-dimensional reference points other than the two-dimensional reference points in the expanded region.

[0069] It should be noted that the region containing all two-dimensional reference points can be expanded by increasing the range of the region containing all two-dimensional reference points. The expanded range is preset and is not limited in this disclosure.

[0070] Step 206: Based on the transformation matrix, map the contours corresponding to the road elements in map data i to obtain the two-dimensional bounding boxes corresponding to the road elements in map data i.

[0071] In this disclosure, after obtaining the transformation matrix, the contours corresponding to the road elements in map data i can be mapped to the image plane to obtain the two-dimensional bounding boxes corresponding to the road elements.

[0072] Step 207 generates prompt information corresponding to the initial data i based on the two-dimensional reference point, the auxiliary two-dimensional reference point, and the two-dimensional bounding box corresponding to the road element.

[0073] It should be noted that after processing the image data, map data, etc. in each initial data in the initial dataset in the above manner, a corresponding prompt message can be obtained.

[0074] In this disclosure, after obtaining the two-dimensional reference point, auxiliary two-dimensional reference point, and two-dimensional bounding box corresponding to the road element in the image data of the initial data i, prompt information corresponding to the initial data i can be generated based on the two-dimensional reference point, auxiliary two-dimensional reference point, and two-dimensional bounding box corresponding to the road element, which serves as the basis for interacting with the large model, thereby providing conditions for improving the accuracy of the large model.

[0075] Step 208: Input the image data in the initial data i and the prompt information corresponding to the initial data i into the preset large model, and obtain the first label of the road element contained in the image data in the initial data i output by the large model.

[0076] The specific implementation of step 208 can be found in the detailed descriptions of other embodiments in this disclosure, and will not be repeated here.

[0077] Step 209: When the first label and the second label of the same road element are different, the large model is corrected based on the difference between the second label and the first label. The second label is determined by the two-dimensional reference point obtained by mapping the three-dimensional reference point in the road element.

[0078] The second label can be a semantic label corresponding to the two-dimensional reference point, used to identify the type of road element corresponding to the two-dimensional reference point. It can be the same as the first label, or it can be different from the first label. This disclosure does not limit this.

[0079] It is understandable that 3D points collected from vectors corresponding to the same road element will have the same semantic labels for the mapped 2D reference points. Conversely, 3D points collected from vectors corresponding to different road elements will have different semantic labels for the mapped 2D reference points.

[0080] In this disclosure, since current large models are usually designed for general object segmentation, and due to the deviation in the projection results of road elements, the labels output by the large model may be inaccurate. Therefore, if the first label and the second label of the same road element are different, it can be assumed that the first label output by the large model is inaccurate. At this time, the inaccurate first label can be corrected based on the difference between the second label and the first label, and the corrected first label can be used to correct the large model.

[0081] In this embodiment of the disclosure, after obtaining the initial dataset, firstly, for the initial data i in the initial dataset, the transformation matrix corresponding to the camera is determined according to the vehicle pose in the initial data i and the camera intrinsic parameters in the image data. Then, a preset number of three-dimensional reference points are collected from the vectors corresponding to each road element in the map data i in the initial data i. Based on the transformation matrix, each three-dimensional reference point is mapped to obtain the corresponding two-dimensional reference point. The region where all two-dimensional reference points are located is expanded to obtain auxiliary two-dimensional reference points other than the two-dimensional reference points in the expanded region. Based on the transformation matrix, the contours corresponding to the road elements in the map data i are mapped to obtain the two-dimensional bounding boxes corresponding to the road elements in the map data i. Then, based on the two-dimensional reference points, auxiliary two-dimensional reference points, and two-dimensional bounding boxes, prompt information corresponding to the initial data i is generated. Finally, the image data in the initial data i and the prompt information corresponding to the initial data i are input into a preset large model to obtain the first label of the road element contained in the image data in the initial data i output by the large model. If the first label and the second label of the same road element are different, the large model is corrected based on the difference between the second label and the first label. Therefore, by correcting the large model based on the difference between the labels automatically labeled by the large model and the second label, the cost of data labeling is reduced and the accuracy and reliability of the large model are improved.

[0082] Figure 3 This is a flowchart illustrating a training method for a localization and mapping model provided in one embodiment of the present disclosure.

[0083] like Figure 3 As shown, the training method for this localization and mapping model may include the following steps:

[0084] Step 301: Obtain the initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data, and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element.

[0085] Step 302: For initial data i in the initial dataset, determine the transformation matrix corresponding to the camera based on the vehicle pose in initial data i and the camera intrinsic parameters in the image data.

[0086] Where i is an integer less than or equal to N, and N is the number of initial data points contained in the initial dataset.

[0087] Step 303: Collect a preset number of three-dimensional reference points from the vectors corresponding to each road element in the map data i in the initial data i.

[0088] Step 304: Based on the transformation matrix, map each three-dimensional reference point to obtain the corresponding two-dimensional reference point.

[0089] Step 305: Dilate the region containing all two-dimensional reference points to obtain auxiliary two-dimensional reference points in the dilated region, excluding the two-dimensional reference points.

[0090] Step 306: Based on the transformation matrix, map the contours corresponding to the road elements in map data i to obtain the two-dimensional bounding boxes corresponding to the road elements in map data i.

[0091] Step 307 generates prompt information corresponding to the initial data i based on the two-dimensional reference point, auxiliary two-dimensional reference point, and two-dimensional bounding box.

[0092] Step 308: Input the image data in the initial data i and the prompt information corresponding to the initial data i into the preset large model, and obtain the first label of the road element contained in the image data in the initial data i output by the large model.

[0093] The specific implementation of steps 301 to 308 can be found in the detailed descriptions of other embodiments in this disclosure, and will not be repeated here.

[0094] Step 309: If the first label of the same road element is different from the third label of its corresponding 2D bounding box, the large model is corrected based on the difference between the third label and the first label.

[0095] The third label can be a semantic label corresponding to the two-dimensional bounding box. It can be the same as the first label or it can be different from the first label. This disclosure does not limit this.

[0096] In this disclosure, if the first label is different from the third label corresponding to the two-dimensional bounding box, it can be considered that the first label output by the large model is inaccurate. In this case, the large model can be corrected based on the difference between the third label and the first label.

[0097] It is understood that in this disclosure, when the first label of the road element contained in each image data is different from the third label corresponding to the two-dimensional bounding box, the large model can be corrected based on the difference between the third label and the first label until a large model that meets the requirements is obtained.

[0098] The following is based on Figure 4 Taking this as an example, the process of training a large model based on image data, map data, and vehicle poses will be explained in detail. Figure 4 This is a schematic diagram of a road element closed-loop data annotation engine provided in an embodiment of this disclosure.

[0099] like Figure 4 As shown, the input data of the road element closed-loop data annotation engine provided in this disclosure is mainly divided into three parts: image data, high-precision map data, and vehicle pose. Among them, high-precision map data is a vectorized map representation.

[0100] The data labeling engine first generates visual cues using map data and vehicle poses. Then, it inputs these cues, along with the image data corresponding to the vehicle poses, into a large model to predict training labels. The training labels are the first labels corresponding to road elements in the image data. The large model is then further trained by combining the input data and the training labels. Through the iterative process of label generation and model training, high-quality road element segmentation results can be generated.

[0101] In this embodiment of the disclosure, after obtaining the initial dataset, firstly, for the initial data i in the initial dataset, the transformation matrix corresponding to the camera is determined according to the vehicle pose in the initial data i and the camera intrinsic parameters in the image data. Then, a preset number of three-dimensional reference points are collected from the vectors corresponding to each road element in the map data i in the initial data i. Based on the transformation matrix, each three-dimensional reference point is mapped to obtain the corresponding two-dimensional reference point. The region where all two-dimensional reference points are located is expanded to obtain auxiliary two-dimensional reference points other than the two-dimensional reference points in the expanded region. Based on the transformation matrix, the contours corresponding to the road elements in the map data i are mapped to obtain the two-dimensional bounding boxes corresponding to the road elements in the map data i. Then, based on the two-dimensional reference points, auxiliary two-dimensional reference points, and two-dimensional bounding boxes, the first prompt information corresponding to the initial data i is generated. Finally, the image data in the initial data i and the prompt information corresponding to the initial data i are input into a preset large model to obtain the first label of the road element contained in the image data in the initial data i output by the large model. If the first label of the same road element is different from the third label of its corresponding two-dimensional bounding box, the large model is corrected based on the difference between the third label and the first label. Therefore, by correcting the large model based on the differences between the labels automatically labeled by the large model and the third label, the efficiency of data labeling is improved, as well as the accuracy and reliability of the large model are enhanced.

[0102] Figure 5 This is a flowchart illustrating a training method for a localization and mapping model provided in one embodiment of the present disclosure.

[0103] like Figure 5As shown, the training method for this localization and mapping model may include the following steps:

[0104] Step 501: Obtain the initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data, and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element.

[0105] Step 502: For initial data i in the initial dataset, determine the transformation matrix corresponding to the camera based on the vehicle pose in initial data i and the camera intrinsic parameters in the image data.

[0106] Where i is an integer less than or equal to N, and N is the number of initial data points contained in the initial dataset.

[0107] The specific implementation of steps 501 to 502 can be found in the detailed descriptions of other embodiments in this disclosure, and will not be repeated here.

[0108] Step 503: Based on the vehicle pose in the initial data i, determine the target type of the road elements contained in the point cloud collected by the radar in the vehicle.

[0109] The target type refers to the type of road elements contained in the point cloud acquired by radar, which can be any type. For example, the target type of road elements can be any type such as vehicles, pedestrians, lane lines, landmarks, etc., and this disclosure does not limit this.

[0110] Step 504: Collect a preset number of three-dimensional reference points from the vectors corresponding to the road elements of the associated target type, wherein the vectors corresponding to the road elements of the associated target type are included in the map data i.

[0111] For example, based on the vehicle's current pose, if the road elements in the point cloud collected by the radar are determined to be lane lines, then three-dimensional reference points can be collected from the vectors corresponding to the lane lines.

[0112] In this disclosure, after determining the target type of a road element, three-dimensional reference points can be collected from the vector corresponding to the road element of the target type, thereby improving the efficiency of collecting three-dimensional reference points.

[0113] Step 505: Based on the transformation matrix, map each three-dimensional reference point to obtain the corresponding two-dimensional reference point.

[0114] Step 506: Dilate the region containing all two-dimensional reference points to obtain auxiliary two-dimensional reference points in the dilated region, excluding the two-dimensional reference points.

[0115] The specific implementation of steps 505 to 506 can be found in the detailed descriptions of other embodiments in this disclosure, and will not be repeated here.

[0116] Step 507: Based on the transformation matrix, map the contours corresponding to the road elements of the associated target type to obtain the two-dimensional bounding boxes corresponding to the road elements of the associated target type, wherein the contours corresponding to the road elements of the associated target type are included in the map data i.

[0117] In this disclosure, based on camera intrinsic and extrinsic parameters, the contours corresponding to road elements of associated target type in map data i can be mapped to the image plane to obtain the two-dimensional bounding boxes corresponding to road elements of associated target type, thereby providing conditions for improving the accuracy of large models.

[0118] Step 508: Based on the two-dimensional reference point, the auxiliary two-dimensional reference point, and the two-dimensional bounding box, generate the prompt information corresponding to the initial data i.

[0119] Step 509: Input the image data in the initial data i and the prompt information corresponding to the initial data i into the preset large model, and obtain the first label of the road element contained in the image data in the initial data i output by the large model.

[0120] Step 510: Using the first labels of each image data, each point cloud data, and the road elements contained in each image data, train the initial localization and mapping model to obtain the main network in the localization and mapping model. The main network is a network used to extract and fuse features from the input image data and point cloud data.

[0121] The specific implementation of steps 508 to 510 can be found in the detailed descriptions of other embodiments in this disclosure, and will not be repeated here.

[0122] In this embodiment, an initial dataset is obtained. For initial data i in the initial dataset, a transformation matrix corresponding to the camera is determined based on the vehicle pose in initial data i and the camera's intrinsic parameters in the image data. Based on the vehicle pose in initial data i, the target type of road elements contained in the point cloud collected by the radar in the vehicle is determined. A preset number of three-dimensional reference points are collected from the vectors corresponding to the road elements associated with the target type. Based on the transformation matrix, each three-dimensional reference point is mapped to obtain the corresponding two-dimensional reference point. The region where all two-dimensional reference points are located is dilated to obtain auxiliary two-dimensional reference points in the dilated region, excluding the two-dimensional reference points. Based on the transformation matrix, the contours corresponding to the road elements associated with the target type are mapped to obtain the two-dimensional bounding boxes corresponding to the road elements associated with the target type. Based on the two-dimensional reference points, auxiliary two-dimensional reference points, and two-dimensional bounding boxes, prompt information corresponding to initial data i is generated. The image data in initial data i and the prompt information corresponding to initial data i are input into a preset large model to obtain the first label of the road elements contained in the image data in initial data i output by the large model. Using each image data, each point cloud data, and the first label of the road elements contained in each image data, the initial localization and mapping model is trained to obtain the main network in the localization and mapping model. Therefore, by determining the target type of road elements contained in the point cloud collected by radar, the system obtains the prompt information of the road elements of the target type. The prompt information and image data are then input into the large model. Using the image data, point cloud data, and the labels of road elements output by the large model, the initial localization and mapping models are trained, thereby improving the accuracy and efficiency of the initial localization and mapping models and reducing the cost of data annotation.

[0123] Figure 6 This is a flowchart illustrating a training method for a localization and mapping model provided in one embodiment of the present disclosure.

[0124] like Figure 6 As shown, the training method for this localization and mapping model may include the following steps:

[0125] Step 601: Obtain the initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data, and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element.

[0126] The specific implementation of step 601 can be found in the detailed descriptions of other embodiments in this disclosure, and will not be repeated here.

[0127] Step 602: For image data i in the initial data i, input image data i into the image processing network in the initial localization and mapping model, and obtain the image features output by the image processing network corresponding to image data i.

[0128] Where i is an integer less than or equal to N, and N is the number of initial data points contained in the initial dataset.

[0129] The image processing network is a network that processes image data, and it can be any type of network. For example, the image processing network can be any visual network such as ResNet, ViT (Vision Transformer) network, or VoVNet real-time object detection backbone network, and this disclosure does not limit it.

[0130] Here, image features are the features used to characterize image data i.

[0131] Step 603: Input the image features into a preset semantic segmentation network and obtain the predicted labels output by the semantic segmentation network.

[0132] The preset semantic segmentation network is a pre-defined network for semantic segmentation of image features, and it can be any semantic segmentation network. For example, the preset semantic segmentation network can be an Atrous Spatial Pyramid Pooling Network (ASPP), etc., and this disclosure does not limit it.

[0133] The predicted label is used to represent the label that the image features are determined to belong to a certain road element after being processed by the semantic segmentation network.

[0134] In this disclosure, after processing image features, the semantic segmentation network can obtain different segmentation regions in the image features, as well as the probability of the target subject in each segmentation region belonging to each road element. By determining the road element with the highest probability of the target subject belonging to each road element in the segmentation region, the corresponding road element is determined as the predicted label. For example, if there are three types of road elements, after processing the image features, the semantic segmentation network obtains one segmentation region, and can obtain the probability of the target subject in this segmentation region belonging to each of the three road elements. If the probability of the target subject belonging to the lane line is the highest, then the predicted label can be the lane line. This disclosure does not limit this.

[0135] Step 604: Determine the first correction gradient based on the first difference between the predicted label and the first label of the road element contained in the image data i.

[0136] Wherein, the first correction gradient is the loss value between the predicted label and the first label.

[0137] For example, the predicted label is S j Where j represents the target category of the road element, and the first difference between the predicted label and the first label can be obtained by calculating the loss function. The loss function L...sem The formula is: L sem =-(1-S j ) γ log(S j ), where γ is the weight of different categories of road elements.

[0138] Step 605: Input the image features into a preset dense depth recognition network to obtain the predicted dense depth map corresponding to image data i.

[0139] Among them, the dense depth map is a depth map that contains precise depth values ​​for all regions of the image features.

[0140] In this disclosure, depth can be predicted in the form of probability distribution rather than metric. Specifically, the initial localization and mapping model first aggregates multi-scale features through the ASPP module, and then predicts the corresponding depth distribution for each pixel through a 1×1 convolutional layer. This can obtain the probability that the depth distribution of each pixel belongs to the preset depth interval of the corresponding pixel, thereby realizing the conversion from image to BEV space in the BEV model.

[0141] Step 606: Generate a reference dense depth map corresponding to image data i based on the point cloud data associated with image data i.

[0142] In this disclosure, when generating a reference dense depth map corresponding to image data i based on point cloud data associated with image data i, depth labels can first be constructed on the local map, and then the point cloud data can be mapped onto each frame of the local map to obtain the reference dense depth map, thereby providing conditions for improving the accuracy and reliability of image processing networks.

[0143] It should be noted that the local map can be obtained through a map building algorithm. This algorithm can provide the point cloud after removing motion distortion and the global pose corresponding to each frame of the point cloud. Then, the system can divide the point cloud into spatial parts based on the global pose and stitch the point cloud together to form a local map.

[0144] Step 607: Determine the second correction gradient based on the second difference between the predicted dense depth and the reference dense depth.

[0145] The second correction gradient is the loss value between the predicted dense depth and the reference dense depth.

[0146] Steps 603 to 604 and steps 605 to 607 can be executed in parallel, or steps 603 to 604 can be executed first, followed by steps 605 to 607, or steps 605 to 607 can be executed first, followed by steps 603 to 604, etc. This disclosure does not limit this.

[0147] Step 608: Based on the first correction gradient and the second correction gradient, the image processing network is corrected.

[0148] In this disclosure, after semantic segmentation and dense depth prediction of image features to obtain a first corrected gradient and a second corrected gradient, the image processing network can be corrected based on the first corrected gradient and the second corrected gradient.

[0149] In this embodiment, after obtaining the initial dataset, image data i in the initial data i is first input into the image processing network of the initial localization and mapping model to obtain the image features corresponding to image data i output by the image processing network. Then, the image features are input into a preset semantic segmentation network to obtain the predicted labels output by the semantic segmentation network. Based on the first difference between the predicted labels and the first labels of road elements contained in image data i, a first correction gradient is determined. The image features are then input into a preset dense depth recognition network to obtain the predicted dense depth map corresponding to image data i. A reference dense depth map corresponding to image data i is generated based on the point cloud data associated with image data i. A second correction gradient is determined based on the second difference between the predicted dense depth and the reference dense depth. Finally, the image processing network is corrected based on the first and second correction gradients. Therefore, by performing semantic segmentation and dense depth prediction on the image features and correcting the image processing network, the accuracy and reliability of the image processing network's output results are improved.

[0150] Figure 7 This is a flowchart illustrating a training method for a localization and mapping model provided in one embodiment of the present disclosure.

[0151] like Figure 7 As shown, the training method for this localization and mapping model may include the following steps:

[0152] Step 701: Obtain the initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data, and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element.

[0153] The specific implementation of step 701 can be found in the detailed descriptions of other embodiments in this disclosure, and will not be repeated here.

[0154] Step 702: For point cloud data i in the initial data i, input point cloud data i into the point cloud data processing network in the initial localization and mapping model to obtain point cloud features.

[0155] Where i is an integer less than or equal to N, and N is the number of initial data points contained in the initial dataset.

[0156] The point cloud data processing network is a network that processes point cloud data, and it can be any point cloud data processing network. For example, the point cloud data processing network can be a PointPillars network, or it can be a VoxelNet network, etc. This disclosure does not limit it.

[0157] Among them, point cloud features are the features used to characterize point cloud data i.

[0158] In this disclosure, point cloud data input from radar can be processed through a point cloud data processing network to extract point cloud features from a bird's-eye view.

[0159] Step 703: Input the point cloud features and associated image features into the fusion network in the initial localization and mapping model, respectively, to obtain the bird's-eye view BEV features after the point cloud and image are fused.

[0160] Among them, the associated image features are the output of the image processing network after processing the image data associated with point cloud data i.

[0161] It should be noted that a two-dimensional encoder can be used as the fusion network in this disclosure. For example, an inverted residual block can be used as the fusion network, etc. This disclosure does not limit this.

[0162] BEV stands for Bird's Eye View.

[0163] In this disclosure, in order to achieve a unified representation of image features and point cloud features, image features can be mapped from image space to BEV space by using transformation modules such as Transformer or LSS, combined with camera intrinsic and extrinsic parameters. When image features and point cloud features are in the same space, image features and point cloud features can be input into a fusion network to obtain BEV features after point cloud and image fusion.

[0164] It should be noted that different tasks can be achieved by fusing BEV features through the decoder unit. For example, in the application of high-precision positioning based on vectorized maps, a Transformer can be added to interact with BEV road elements, and the weights of different pose samples can be calculated through the cost volume, thereby achieving vehicle-side positioning in autonomous driving scenarios. This disclosure does not limit this.

[0165] Step 704: Input the BEV features into the preset 3D occupancy network and obtain the predicted 3D occupancy probability map output by the 3D occupancy network.

[0166] The preset 3D occupancy network is pre-set and can be any 3D occupancy network. For example, the 3D occupancy network can be an occupancy grid prediction network, etc., and this disclosure does not limit it.

[0167] Step 705: Generate a 3D occupancy probability map based on point cloud data i and associated image data.

[0168] In this disclosure, a corresponding 3D occupancy probability map is generated based on point cloud data i, associated image data, and a local map, which provides conditions for improving the accuracy and reliability of the fusion network.

[0169] It should be noted that this disclosure can add a 3D detection network (such as ResNet3D) to predict the probability of 3D occupied grids on the basis of BEV features, and use local maps to generate corresponding occupied grid labels for supervision, thereby realizing the training of 3D perception capabilities.

[0170] Step 706: Based on the third difference between the predicted 3D occupancy probability map and the reference 3D occupancy probability map, the fusion network is corrected.

[0171] In this disclosure, the fusion network can be corrected based on a third difference between the predicted 3D occupancy probability map and the reference 3D occupancy probability map.

[0172] The following is based on Figure 8 Taking this as an example, the pre-training process for image features and the fusion of BEV features will be explained in detail. Figure 8 This is a schematic diagram of the main network and multi-task pre-training framework in the localization and mapping model provided in the embodiments of this disclosure, wherein Transformer is a deep learning model based on attention mechanism, LSS is a three-dimensional bird's-eye view object detection edge perception LSS (Lift-Splat-Shoot) framework, and Occupancy is a grid.

[0173] like Figure 8 As shown in the dashed box, this disclosure allows for pre-training of image features and fused BEV features through sub-tasks related to localization and mapping. For image features, semantic segmentation and depth prediction of road elements can be used as sub-tasks to pre-train the image features, such as... Figure 8 The dashed box marked "①" indicates this. For fused BEV features, occupancy grid prediction can be used as a subtask to train the fused BEV features, such as... Figure 8 The dashed box labeled "②" shows the process. Through pre-training on these tasks, the backbone network used by the localization and mapping model can learn the semantic information and 3D location information of road elements before training for downstream tasks, thereby improving the road element perception and 3D perception capabilities of the localization and mapping model.

[0174] In this embodiment of the disclosure, after obtaining the initial dataset, point cloud data i is first input into the point cloud data processing network of the initial localization and mapping model to obtain point cloud features. Then, the point cloud features and associated image features are respectively input into the fusion network of the initial localization and mapping model to obtain the bird's-eye view BEV features after point cloud and image fusion. Next, the BEV features are input into a preset 3D occupancy network to obtain the predicted 3D occupancy probability map output by the 3D occupancy network. Based on point cloud data i and associated image data, a 3D occupancy probability map is generated. Finally, the fusion network is corrected based on the third difference between the predicted 3D occupancy probability map and the reference 3D occupancy probability map. Therefore, by performing 3D occupancy prediction based on the BEV features after point cloud and image fusion, and correcting the fusion network based on the difference between the predicted 3D occupancy probability map and the reference 3D occupancy probability map, the accuracy and reliability of the fusion network output results are improved.

[0175] To implement the above embodiments, this disclosure also proposes a training device for localization and mapping models.

[0176] Figure 9 This is a schematic diagram of the structure of the training device for the localization and mapping model provided in the embodiments of this disclosure.

[0177] like Figure 9 As shown, the training device 900 for the localization and mapping model includes: a first acquisition module 901, a generation module 902, a second acquisition module 903, and a training module 904.

[0178] The first acquisition module 901 is used to acquire an initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element.

[0179] The generation module 902 is used to generate corresponding prompt information based on the image data, map data and vehicle pose in each initial data set;

[0180] The second acquisition module 903 is used to input each image data and its corresponding prompt information into a preset large model, and to acquire the first label of the road element contained in each image data output by the large model.

[0181] Training module 904 is used to train the initial localization and mapping model using the first labels of each image data, each point cloud data, and the road elements contained in each image data, so as to obtain the main network in the localization and mapping model. The main network is a network used to extract and fuse features from the input image data and point cloud data.

[0182] In one possible implementation of this disclosure, the above-mentioned generation module 902 is specifically used for:

[0183] For initial data i in the initial dataset, the transformation matrix corresponding to the camera is determined based on the vehicle pose in the initial data i and the camera intrinsic parameters in the image data, where i is an integer less than or equal to N, and N is the number of initial data contained in the initial dataset.

[0184] From the vectors corresponding to each road element in the map data i in the initial data i, collect a preset number of three-dimensional reference points;

[0185] Based on the transformation matrix, each three-dimensional reference point is mapped to obtain the corresponding two-dimensional reference point;

[0186] Dilate the region containing all two-dimensional reference points to obtain auxiliary two-dimensional reference points in the dilated region, excluding the two-dimensional reference points themselves.

[0187] Based on the transformation matrix, the contours corresponding to the road elements in map data i are mapped to obtain the two-dimensional bounding boxes corresponding to the road elements in map data i.

[0188] Based on the two-dimensional reference point, the auxiliary two-dimensional reference point, and the two-dimensional bounding box, generate the prompt information corresponding to the initial data i.

[0189] In one possible implementation of this disclosure, the generation module 902 is further configured to:

[0190] When the first and second labels of the same road element are different, the large model is corrected based on the difference between the second and first labels, where the second label is determined by a two-dimensional reference point mapped from the three-dimensional reference points in the road element; and / or,

[0191] When the first label of the same road element is different from the third label of its corresponding 2D bounding box, the large model is corrected based on the difference between the third label and the first label.

[0192] In one possible implementation of this disclosure, the generation module 902 is further configured to:

[0193] Based on the vehicle pose in the initial data i, determine the target type of the road elements contained in the point cloud collected by the radar in the vehicle.

[0194] A preset number of 3D reference points are collected from the vectors corresponding to road elements of the associated target type, wherein the vectors corresponding to road elements of the associated target type are included in map data i.

[0195] In one possible implementation of this disclosure, the generation module 902 is further configured to:

[0196] Based on the transformation matrix, the contours corresponding to the road elements of the associated target type are mapped to obtain the two-dimensional bounding boxes corresponding to the road elements of the associated target type. The contours corresponding to the road elements of the associated target type are included in the map data i.

[0197] In one possible implementation of this disclosure, the training module 904 is specifically used for:

[0198] For image data i in the initial data i, input image data i into the image processing network in the initial localization and mapping model, and obtain the image features corresponding to image data i output by the image processing network, where i is an integer less than or equal to N, and N is the number of initial data contained in the initial dataset;

[0199] Input image features into a preset semantic segmentation network and obtain the predicted labels output by the semantic segmentation network;

[0200] A first correction gradient is determined based on the first difference between the predicted label and the first label of the road element contained in image data i;

[0201] Image features are input into a preset dense depth recognition network to obtain a predicted dense depth map corresponding to the second image data;

[0202] The second correction gradient is determined based on the second difference between the predicted density depth and the reference density depth;

[0203] The image processing network is modified based on the first and second correction gradients.

[0204] In one possible implementation of this disclosure, the training module 904 is further configured to:

[0205] Generate a reference dense depth map corresponding to image data i based on the point cloud data associated with image data i.

[0206] In one possible implementation of this disclosure, the training module 904 is further configured to:

[0207] For point cloud data i in the initial data i, input point cloud data i into the point cloud data processing network in the initial localization and mapping model to obtain point cloud features, where i is an integer less than or equal to N, and N is the number of initial data contained in the initial dataset.

[0208] The point cloud features and associated image features are respectively input into the fusion network in the initial localization and mapping model to obtain the bird's-eye view BEV features after the point cloud and image are fused. The associated image features are the output of the image processing network after processing the image data associated with the point cloud data i.

[0209] Input the BEV features into a preset 3D occupancy network and obtain the predicted 3D occupancy probability map output by the 3D occupancy network.

[0210] The fusion network is corrected based on the third difference between the predicted 3D occupancy probability map and the reference 3D occupancy probability map.

[0211] In one possible implementation of this disclosure, the training module 904 is further configured to:

[0212] A 3D occupancy probability map is generated based on point cloud data i and associated image data.

[0213] The functions and specific implementation principles of the modules described in this embodiment can be found in the above method embodiments, and will not be repeated here.

[0214] In this embodiment, after obtaining the initial dataset, corresponding prompts are first generated based on the image data, map data, and vehicle poses in each initial dataset. Then, each image data and its corresponding prompt are input into a preset large model to obtain the first label of the road element contained in each image data output by the large model. Subsequently, the initial localization and mapping model is trained using each image data, each point cloud data, and the first label of the road element contained in each image data to obtain the main network in the localization and mapping model. Thus, by training the initial localization and mapping model based on image data, point cloud data, and the first label of the road element contained in the image data, the perception capability and 3D perception capability of the localization and mapping model are improved, thereby improving the accuracy and reliability of the localization and mapping model. Furthermore, the cost of data annotation is reduced by automatically labeling the training data through the large model.

[0215] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0216] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0217] like Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0218] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0219] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as methods for training localization and mapping models. For example, in some embodiments, the methods for training localization and mapping models can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the methods for training localization and mapping models described above can be performed. Alternatively, in other embodiments, computing unit 1001 may be configured by any other suitable means (e.g., by means of firmware) to perform a method for training localization and mapping models.

[0220] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0221] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0222] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0223] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0224] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0225] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0226] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0227] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified. In the description of this disclosure, the words "if" and "suppose" as used may be interpreted as "when," "when," "in response to determination," or "in the circumstances."

[0228] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A training method for a localization and mapping model, comprising: Obtain an initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data, and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element; For initial data i in the initial dataset, the transformation matrix corresponding to the camera is determined based on the vehicle pose in the initial data i and the camera intrinsic parameters in the image data, where i is an integer less than or equal to N, and N is the number of initial data contained in the initial dataset. A preset number of three-dimensional reference points are collected from the vectors corresponding to each road element in the map data i in the initial data i; Based on the transformation matrix, each of the three-dimensional reference points is mapped to obtain the corresponding two-dimensional reference points; Expand the region containing all the two-dimensional reference points to obtain auxiliary two-dimensional reference points in the expanded region, excluding the two-dimensional reference points. Based on the transformation matrix, the contours corresponding to the road elements in the map data i are mapped to obtain the two-dimensional bounding boxes corresponding to the road elements in the map data i. Based on the two-dimensional reference point, the auxiliary two-dimensional reference point, and the two-dimensional bounding box, generate the prompt information corresponding to the initial data i; Each image data and its corresponding prompt information are input into a preset large model, and the first label of the road element contained in each image data output by the large model is obtained; Using the image data, the point cloud data, and the first label of the road element contained in the image data, an initial localization and mapping model is trained to obtain the main network in the localization and mapping model, wherein the main network is a network used to extract and fuse features from the input image data and point cloud data.

2. The method as described in claim 1, wherein, After inputting each image data and its corresponding prompt information into a preset large model, and obtaining the first label of the road element contained in each image data output by the large model, the method further includes: When the first and second labels of the same road element are different, the large model is corrected based on the difference between the second and first labels, wherein the second label is determined by a two-dimensional reference point mapped from the three-dimensional reference points in the road element; and / or, If the first label of the same road element is different from the third label of its corresponding two-dimensional bounding box, the large model is corrected based on the difference between the third label and the first label.

3. The method as described in claim 1, wherein, The step of collecting a preset number of three-dimensional reference points from the vectors corresponding to each road element in the map data i in the initial data i includes: Based on the vehicle pose in the initial data i, determine the target type of the road elements contained in the point cloud acquired by the radar in the vehicle. A preset number of three-dimensional reference points are collected from the vectors corresponding to the road elements associated with the target type, wherein the vectors corresponding to the road elements associated with the target type are included in the map data i.

4. The method of claim 3, wherein, The step of mapping the contours corresponding to road elements in map data i based on the transformation matrix to obtain the two-dimensional bounding boxes corresponding to road elements in map data i includes: Based on the transformation matrix, the contours corresponding to the road elements associated with the target type are mapped to obtain the two-dimensional bounding boxes corresponding to the road elements associated with the target type, wherein the contours corresponding to the road elements associated with the target type are included in the map data i.

5. The method as described in any one of claims 1-4, wherein, The training of the initial localization and mapping model includes: For image data i in the initial data i, the image data i is input into the image processing network in the initial localization and mapping model to obtain the image features output by the image processing network corresponding to the image data i, where i is an integer less than or equal to N, and N is the number of initial data contained in the initial dataset; The image features are input into a preset semantic segmentation network to obtain the predicted labels output by the semantic segmentation network; A first correction gradient is determined based on the first difference between the predicted label and the first label of the road element contained in the image data i; The image features are input into a preset dense depth recognition network to obtain the predicted dense depth map corresponding to the image data i; A second correction gradient is determined based on the second difference between the predicted density depth and the reference density depth; The image processing network is modified based on the first and second correction gradients.

6. The method of claim 5, wherein, Before determining the second correction gradient based on the second difference between the predicted density depth and the reference density depth, the method further includes: A reference dense depth map corresponding to image data i is generated based on the point cloud data associated with image data i.

7. The method of claim 5, wherein, The training of the initial localization and mapping model includes: For point cloud data i in the initial data i, the point cloud data i is input into the point cloud data processing network in the initial localization and mapping model to obtain point cloud features, wherein i is an integer less than or equal to N, and N is the number of initial data contained in the initial dataset; The point cloud features and associated image features are respectively input into the fusion network in the initial localization and mapping model to obtain the bird's-eye view BEV features after the point cloud and image are fused. The associated image features are the output of the image processing network after processing the image data associated with the point cloud data i. The BEV features are input into a preset 3D occupancy network to obtain the predicted 3D occupancy probability map output by the 3D occupancy network. The fusion network is corrected based on the third difference between the predicted 3D occupancy probability map and the reference 3D occupancy probability map.

8. The method of claim 7, wherein, Before correcting the fusion network based on the third difference between the predicted 3D occupancy probability map and the reference 3D occupancy probability map, the method further includes: The 3D occupancy probability map is generated based on the point cloud data i and the associated image data.

9. A training device for a localization and mapping model, comprising: The first acquisition module is used to acquire an initial dataset, wherein each initial data in the initial dataset includes image data, point cloud data, map data and vehicle poses corresponding to the image data, and the map data includes vectors corresponding to each road element; The generation module is used to: determine the transformation matrix corresponding to the camera for initial data i in the initial dataset, based on the vehicle pose in the initial data i and the camera's intrinsic parameters in the image data, where i is an integer less than or equal to N, and N is the number of initial data in the initial dataset; collect a preset number of 3D reference points from the vectors corresponding to each road element in the map data i in the initial data i; map each 3D reference point based on the transformation matrix to obtain a corresponding 2D reference point; dilate the region where all the 2D reference points are located to obtain auxiliary 2D reference points in the dilated region, excluding the 2D reference points; map the contours corresponding to the road elements in the map data i based on the transformation matrix to obtain 2D bounding boxes corresponding to the road elements in the map data i; and generate prompt information corresponding to the initial data i based on the 2D reference points, the auxiliary 2D reference points, and the 2D bounding boxes. The second acquisition module is used to input each image data and its corresponding prompt information into a preset large model, and to acquire the first label of the road element contained in each image data output by the large model. The training module is used to train an initial localization and mapping model using the image data, the point cloud data, and the first labels of the road elements contained in the image data, so as to obtain the main network in the localization and mapping model, wherein the main network is a network used to extract and fuse features from the input image data and point cloud data.

10. The apparatus of claim 9, wherein, The generation module is further configured to: When the first and second labels of the same road element are different, the large model is corrected based on the difference between the second and first labels, wherein the second label is determined by a two-dimensional reference point mapped from the three-dimensional reference points in the road element; and / or, If the first label of the same road element is different from the third label of its corresponding two-dimensional bounding box, the large model is corrected based on the difference between the third label and the first label.

11. The apparatus of claim 9, wherein, The generation module is further configured to: Based on the vehicle pose in the initial data i, determine the target type of the road elements contained in the point cloud acquired by the radar in the vehicle. A preset number of three-dimensional reference points are collected from the vectors corresponding to the road elements associated with the target type, wherein the vectors corresponding to the road elements associated with the target type are included in the map data i.

12. The apparatus of claim 11, wherein, The generation module is further configured to: Based on the transformation matrix, the contours corresponding to the road elements associated with the target type are mapped to obtain the two-dimensional bounding boxes corresponding to the road elements associated with the target type, wherein the contours corresponding to the road elements associated with the target type are included in the map data i.

13. The apparatus according to any one of claims 9-12, wherein, The training module is specifically used for: For image data i in the initial data i, the image data i is input into the image processing network in the initial localization and mapping model to obtain the image features output by the image processing network corresponding to the image data i, where i is an integer less than or equal to N, and N is the number of initial data contained in the initial dataset; The image features are input into a preset semantic segmentation network to obtain the predicted labels output by the semantic segmentation network; A first correction gradient is determined based on the first difference between the predicted label and the first label of the road element contained in the image data i; The image features are input into a preset dense depth recognition network to obtain the predicted dense depth map corresponding to the image data i; A second correction gradient is determined based on the second difference between the predicted density depth and the reference density depth; The image processing network is modified based on the first and second correction gradients.

14. The apparatus of claim 13, wherein, The training module is also used for: A reference dense depth map corresponding to image data i is generated based on the point cloud data associated with image data i.

15. The apparatus of claim 13, wherein, The training module is also used for: For point cloud data i in the initial data i, the point cloud data i is input into the point cloud data processing network in the initial localization and mapping model to obtain point cloud features, wherein i is an integer less than or equal to N, and N is the number of initial data contained in the initial dataset; The point cloud features and associated image features are respectively input into the fusion network in the initial localization and mapping model to obtain the bird's-eye view BEV features after the point cloud and image are fused. The associated image features are the output of the image processing network after processing the image data associated with the point cloud data i. The BEV features are input into a preset 3D occupancy network to obtain the predicted 3D occupancy probability map output by the 3D occupancy network. The fusion network is corrected based on the third difference between the predicted 3D occupancy probability map and the reference 3D occupancy probability map.

16. The apparatus of claim 15, wherein, The training module is also used for: The 3D occupancy probability map is generated based on the point cloud data i and the associated image data.

17. An electronic device, characterized in that, include: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that may be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.

19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN111339964A

  • Remote sensing image auxiliary processing method and device based on online learning

    CN113158855A