Element extraction and map generation method and device based on zero-shot deep learning

The zero-shot deep learning method enhances road element extraction in autonomous driving by using YOLOv8-OBB and PCA algorithms to project 3D point cloud data to 2D images, addressing precision, robustness, and real-time challenges, and reducing dataset creation costs.

CN119360051BActive Publication Date: 2025-07-15CHENGDU GREEN TURLON JIWU TECHNOLOGY CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411600138.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-07-15
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The existing road factor extraction technology has problems such as low accuracy, poor robustness, difficulty in projecting into 3D point cloud data, high labor cost for data set production, and poor real-time performance in autonomous driving.

Method used

Using a method based on zero-sample deep learning, a training model data set is generated through custom road feature identification templates, combined with the YOLOv8-OBB rotary box model and PCA algorithm, the 3D point cloud is mapped to the 2D image using intensity projection, and the road feature deep learning model of the 2D image is extracted, and the 2D key point information is restored back to the 3D point cloud, fusing the image texture information and the point cloud spatial information to improve extraction accuracy and robustness.

Benefits of technology

It improves the accuracy and robustness of road feature extraction, reduces the cost of data set production, improves real-time, is suitable for road feature extraction in different regions and scenarios, and improves the accuracy of map construction for autonomous driving, thereby enhancing safety and speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360051B_ABST
    Figure CN119360051B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and device for element extraction and map generation based on zero-shot deep learning. The present disclosure performs intensity projection on a three-dimensional point cloud image that requires road element extraction to obtain a two-dimensional intensity projection image; inputs the two-dimensional intensity projection image into a road element deep learning model to obtain the two-dimensional coordinate information of the key points of the road elements output by the road element deep learning model; the road element deep learning model sequentially includes a YOLOv8-OBB rotated box model, a calculation module including a PCA algorithm, and a multi-head centernet key point detection model; then, according to the reverse logic of the processing logic of the road element deep learning model, the two-dimensional coordinate information of the key points in the two-dimensional intensity projection image is determined; finally, according to the reverse logic of the processing logic of the intensity projection and the two-dimensional coordinate information of the key points in the two-dimensional intensity projection image, the three-dimensional coordinates of the road elements in the three-dimensional point cloud image are determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of autonomous driving technology, and particularly to the field of map construction technology, and discloses a method and device for feature extraction and map generation based on zero-shot deep learning. Background Art

[0002] With the development of autonomous driving technology, road features are not only identifiers for ensuring pedestrian and vehicle safety, but also indispensable basic elements in high-precision maps that meet the needs of autonomous driving. Existing road feature extraction technologies are mainly divided into image-based detection methods and point cloud-based detection methods.

[0003] The main advantage of the image-based detection method is that the image has high resolution and rich texture information, and the detection effect is relatively ideal. However, the image-based detection method has the following deficiencies:

[0004] Strong environmental dependence: Image detection is greatly affected by external conditions such as lighting and weather, resulting in poor robustness; Lack of spatial information: The two-dimensional image detection method cannot directly obtain three-dimensional geometric information, resulting in the inability to provide accurate three-dimensional positioning when constructing a high-precision map; High labor cost: There are certain differences in road features in each country, and the labor cost for making different data sets is high.

[0005] In the point cloud-based detection method, the three-dimensional point cloud data generated by lidar contains rich spatial information and can more accurately represent the physical location of road features. However, point cloud detection also has deficiencies:

[0006] High computational cost: Processing a large amount of point cloud data requires high computing resources, especially in real-time applications, it is difficult to complete the processing efficiently; Data sparsity: Due to the working mode of lidar, especially in the case of long distances or complex terrains, the point cloud data is often sparse and incomplete, affecting the accuracy of feature extraction. Summary of the Invention

[0007] The present disclosure provides at least one method and device for feature extraction and map generation based on zero-shot deep learning, which solves at least one of the problems such as low accuracy, poor robustness, difficulty in projecting into 3D point cloud data, high labor cost for making data sets, and poor real-time performance in actual production.

[0008] According to one aspect of the present disclosure, a method for feature extraction based on zero-shot deep learning is provided, including:

[0009] Obtain a pre-trained deep learning model for road features and a three-dimensional point cloud image that needs to perform road feature extraction;

[0010] Perform intensity projection on the three-dimensional point cloud image to obtain a two-dimensional intensity projection image;

[0011] Input the two-dimensional intensity projection image into the road element deep learning model to obtain the two-dimensional coordinate information of the key points of the road elements output by the road element deep learning model; the road element deep learning model sequentially includes a YOLOv8-OBB rotation box model, a calculation module including a PCA algorithm, and a multi-head centernet key point detection model;

[0012] According to the reverse logic of the processing logic of the road element deep learning model, determine the two-dimensional coordinate information of the key points on the two-dimensional intensity projection image;

[0013] According to the reverse logic of the processing logic of the intensity projection and the two-dimensional coordinate information of the key points on the two-dimensional intensity projection image, determine the three-dimensional coordinates of the road elements on the three-dimensional point cloud image;

[0014] Among them, the data set for training the multi-head centernet key point detection model is obtained according to the following steps:

[0015] Obtain multiple three-dimensional point cloud sample images, and perform intensity projection on each three-dimensional point cloud sample image respectively to obtain multiple two-dimensional intensity sample projection images;

[0016] Use each two-dimensional intensity sample projection image to determine multiple road element identification templates and the key points corresponding to each road element identification template;

[0017] Fill each road element identification template with the first target scene color, and rotate it at least once according to the first preset angle range to obtain multiple first road element template images, and use the data set formed by the first road element template images as the data set for training the multi-head centernet key point detection model;

[0018] Among them, the data set for training the YOLOv8-OBB rotation box model is obtained according to the following steps:

[0019] Fill each road element identification template with the second target scene color, and rotate it at least once according to the second preset angle range to obtain multiple second road element template images; paste the obtained second road element template images at at least one randomly determined position on at least one road background image to obtain multiple training sample images;

[0020] Use the data set formed by the training sample images as the data set for training the YOLOv8-OBB rotation box model;

[0021] Among them, the processing logic of the road element deep learning model is as follows:

[0022] Extract the rotation box position information of road elements from the two-dimensional intensity projection image using the YOLOv8-OBB rotation box model; the rotation box position information is two-dimensional coordinate information;

[0023] Use the PCA algorithm to perform principal component analysis on the rotation box position information of road elements to obtain the main direction of the rotation box;

[0024] According to the angle between the main direction of the rotation box and the vertical direction, perform rotation correction on the pixel points corresponding to the road elements on the two-dimensional intensity projection image;

[0025] Use the multi-head centernet keypoint detection model to extract keypoints from the rotation-corrected road elements to obtain the two-dimensional coordinate information of the keypoints.

[0026] In a possible implementation manner, the step of performing intensity projection on the three-dimensional point cloud image to obtain a two-dimensional intensity projection image includes:

[0027] According to the position information of each point in the three-dimensional point cloud image, determine the width and height of the point cloud;

[0028] According to the width and height of the point cloud, map each point in the three-dimensional point cloud image to the image grid to obtain the position index of each point in the two-dimensional intensity projection image;

[0029] For the image grid with only one point, fill the intensity of the point in the image grid into the corresponding image network; for the image grid with multiple point clouds, fill the maximum intensity into the corresponding image network.

[0030] In a possible implementation manner, the step of determining the width and height of the point cloud according to the position information of each point in the three-dimensional point cloud image includes:

[0031] According to the position information of each point in the three-dimensional point cloud image, determine the maximum value and minimum value of the point cloud in the x direction;

[0032] According to the maximum value of the point cloud in the x direction, the minimum value of the point cloud in the x direction, and the preset resolution, determine the width of the point cloud;

[0033] According to the position information of each point in the three-dimensional point cloud image, determine the maximum value and minimum value of the point cloud in the y direction;

[0034] According to the maximum value of the point cloud in the y direction, the minimum value of the point cloud in the y direction, and the preset resolution, determine the height of the point cloud;

[0035] Where the x direction and the y direction are perpendicular to each other.

[0036] In a possible implementation, mapping each point in the three-dimensional point cloud image to an image grid according to the width and height of the point cloud to obtain the position index of each point in the two-dimensional intensity projection image includes:

[0037] For each point, determine the first index of the point in the y direction according to the coordinate of the point in the y direction, the maximum value of the point cloud in the y direction, and the preset resolution; determine the second index of the point in the x direction according to the coordinate of the point in the x direction, the maximum value of the point cloud in the x direction, and the preset resolution; use the first index and the second index as the position index of the point in the two-dimensional intensity projection image.

[0038] In a possible implementation, performing principal component analysis on the rotated bounding box position information of the road element using the PCA algorithm to obtain the main direction of the rotated bounding box includes:

[0039] Determine the two-dimensional coordinates of the centroid of the rotated bounding box according to the rotated bounding box position information of the road element;

[0040] Perform a de-centralization process on each boundary point of the rotated bounding box using the two-dimensional coordinates of the centroid to obtain the de-centralized point corresponding to each boundary point;

[0041] Determine the covariance matrix using the coordinates of each de-centralized point;

[0042] Perform eigenvalue decomposition on the covariance matrix, and determine the main direction of the rotated bounding box according to the result of the eigenvalue decomposition.

[0043] In a possible implementation, performing a de-centralization process on each boundary point of the rotated bounding box using the two-dimensional coordinates of the centroid to obtain the de-centralized point corresponding to each boundary point includes:

[0044] For each boundary point, perform the de-centralization process using the following formula:

[0045] P′ i =(x i -C x ,y i -C y )

[0046] where (C x , C y ) represents the two-dimensional coordinates of the centroid, (x i , y i ) represents the coordinates of boundary point i, and P′ i represents the de-centralized point obtained by the de-centralization process.

[0047] In a possible implementation, the first target scene color and / or the second target scene color is white.

[0048] According to another aspect of the present disclosure, there is provided a feature extraction device based on zero-shot deep learning, including:

[0049] A data acquisition module, configured to acquire a pre-trained road feature deep learning model and a three-dimensional point cloud image for which road feature extraction is to be performed;

[0050] An intensity projection module, configured to perform intensity projection on the three-dimensional point cloud image to obtain a two-dimensional intensity projection image;

[0051] A key point extraction module, configured to input the two-dimensional intensity projection image into the road feature deep learning model to obtain two-dimensional coordinate information of key points of road features output by the road feature deep learning model; the road feature deep learning model sequentially includes a YOLOv8-OBB rotation box model, a calculation module including a PCA algorithm, and a multi-head centernet key point detection model;

[0052] A first reverse logic processing module, configured to determine two-dimensional coordinate information of the key point two-dimensional coordinate information on the two-dimensional intensity projection image according to the reverse logic of the processing logic of the road feature deep learning model;

[0053] A second reverse logic processing module, configured to determine three-dimensional coordinates of the road feature on the three-dimensional point cloud image according to the reverse logic of the intensity projection processing logic and the two-dimensional coordinate information of the key point two-dimensional coordinate information on the two-dimensional intensity projection image;

[0054] Wherein, the key point extraction module is further configured to determine a data set for training the multi-head centernet key point detection model:

[0055] Acquire multiple three-dimensional point cloud sample images, and perform intensity projection on each three-dimensional point cloud sample image respectively to obtain multiple two-dimensional intensity sample projection images;

[0056] Use each two-dimensional intensity sample projection image to determine multiple road feature identification templates and key points corresponding to each road feature identification template;

[0057] Fill each road feature identification template with a first target scene color, and rotate it at least once according to a first preset angle range to obtain multiple first road feature template images, and use the data set formed by the first road feature template images as the data set for training the multi-head centernet key point detection model;

[0058] Among them, the key point extraction module is further configured to determine a dataset for training the YOLOv8-OBB rotated bounding box model:

[0059] Fill each road element identification template with a second target scene color, and rotate it at least once according to a second preset angle range to obtain multiple second road element template images; paste the obtained second road element template images onto at least one road background image at at least one randomly determined position to obtain multiple training sample images;

[0060] Use the dataset formed by the training sample images as the dataset for training the YOLOv8-OBB rotated bounding box model;

[0061] Among them, the processing logic of the road element deep learning model is as follows:

[0062] Use the YOLOv8-OBB rotated bounding box model to extract the rotated bounding box position information of road elements from the two-dimensional intensity projection image; the rotated bounding box position information is two-dimensional coordinate information;

[0063] Use the PCA algorithm to perform principal component analysis on the rotated bounding box position information of road elements to obtain the main direction of the rotated bounding box;

[0064] According to the angle between the main direction of the rotated bounding box and the vertical direction, rotate the pixel points corresponding to the road elements on the two-dimensional intensity projection image;

[0065] Use the multi-head centernet key point detection model to extract key points from the rotated road elements to obtain the key point two-dimensional coordinate information.

[0066] According to another aspect of the present disclosure, there is provided a map generation method, including:

[0067] Use the element extraction method based on zero-shot deep learning described in any one of the above to determine the three-dimensional coordinates of each road element in the three-dimensional point cloud image on the three-dimensional point cloud image;

[0068] Generate a map including road elements according to the determined three-dimensional coordinates.

[0069] According to another aspect of the present disclosure, there is provided a map generation device, including:

[0070] A three-dimensional coordinate determination module, configured to use the element extraction method based on zero-shot deep learning described in any one of the above to determine the three-dimensional coordinates of each road element in the three-dimensional point cloud image on the three-dimensional point cloud image;

[0071] A map generation module, configured to generate a map including road elements according to the determined three-dimensional coordinates.

[0072] The element extraction and map generation method and device based on zero-shot deep learning of the present disclosure adopt the method of zero-shot custom deep learning, and automatically generate the dataset used for training the model through a custom road element identification template, avoiding the high cost of manual annotation. Especially in the case where different road elements are required in different countries and regions, the cost of making the dataset is significantly reduced. At the same time, the present disclosure proposes to map 3D point clouds to 2D images through intensity projection, and use the road element deep learning model corresponding to the 2D image to extract road elements, solving the problems that traditional image methods lack spatial information and cannot obtain three-dimensional precise positioning. Subsequently, the present disclosure restores the information of the extracted 2D key points back to the 3D point cloud, improving the extraction accuracy of road elements. The present disclosure effectively overcomes the deficiencies in a single mode, such as the problems that images are greatly affected by light and weather and point cloud data is sparse, by fusing the texture information of images and the spatial information of point clouds. This solution improves the extraction accuracy and robustness through multi-modal information. In addition, the present disclosure adopts a combination of the YOLOv8-OBB model and the PCA algorithm to quickly extract the rotation bounding boxes of road elements, and combines CenterNet for key point detection, improving the real-time performance of the solution in practical applications. And the present disclosure combines the detection of OBB rotation bounding boxes with the PCA algorithm to perform principal component analysis on the rotation bounding boxes of road elements and automatically rotate and align them, enabling the extracted road elements to handle situations with larger rotation angles, especially suitable for complex road environments. Furthermore, users can customize or revise the road element identification template according to actual needs, such as information on element types, key points, etc., making the road element deep learning model highly adaptable and capable of flexibly handling diverse road elements, suitable for different regions and scenarios.

[0073] In summary, the solution of the present disclosure solves the problems in actual production, such as low extraction accuracy of road elements, poor robustness, difficulty in projecting into 3D point cloud data, high labor cost for making datasets, and poor real-time performance. On this basis, the accuracy of the constructed map can be improved, which is beneficial to improving the safety and speed of autonomous driving.

[0074] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings

[0075] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0076] Figure 1 is a flowchart of the element extraction method based on zero-shot deep learning according to the present disclosure;

[0077] Figure 2AIt is a schematic diagram of a two-dimensional intensity projection image according to an embodiment of the present disclosure;

[0078] Figure 2B It is a schematic diagram of a road element identification template according to an embodiment of the present disclosure;

[0079] Figure 2C It is a schematic diagram of a road element identification template including key points according to an embodiment of the present disclosure;

[0080] Figure 3 It is a schematic diagram of a road element deep learning model according to an embodiment of the present disclosure;

[0081] Figure 4 It is a schematic structural diagram of an element extraction device based on zero-shot deep learning according to the present disclosure;

[0082] Figure 5 It is a schematic structural diagram of an electronic device according to the present disclosure. Detailed implementation manners

[0083] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0084] In view of the problems in actual production such as low accuracy, poor robustness, difficulty in projecting into 3D point cloud data, high manual cost for dataset production, and poor real-time performance in road element extraction, the present disclosure proposes a method and device for element extraction and map generation based on zero-shot deep learning. The solution of the present disclosure solves the above problems and can improve the accuracy of the constructed map on this basis, thereby facilitating the improvement of the safety and speed of autonomous driving.

[0085] The technical solution of the present disclosure will be described below through specific embodiments.

[0086] As Figure 1 shown, it is a flowchart of the element extraction method based on zero-shot deep learning in this embodiment. The execution subject of this embodiment is a computing device or component with data processing capabilities. Specifically, the method of this embodiment may include the following steps:

[0087] S110. Obtain a pre-trained road element deep learning model and a three-dimensional point cloud image for which road element extraction is required.

[0088] The road element deep learning model here is trained using an unlabeled sample dataset.

[0089] S120. Perform intensity projection on the three-dimensional point cloud image to obtain a two-dimensional intensity projection image.

[0090] Realize the projection of three-dimensional information onto a two-dimensional image to facilitate the subsequent extraction of key point information with three-dimensional features by the road element deep learning model. As Figure 2A shown, it is the two-dimensional intensity projection image obtained by performing intensity projection on the three-dimensional point cloud image.

[0091] S130. Input the two-dimensional intensity projection image into the road element deep learning model to obtain the two-dimensional coordinate information of the key points of the road elements output by the road element deep learning model; the road element deep learning model sequentially includes a YOLOv8-OBB rotated bounding box model, a calculation module containing the PCA algorithm, and a multi-head centernet key point detection model.

[0092] The YOLOv8-OBB rotated bounding box model is used to extract or determine the information of the road element rotated bounding box, the PCA algorithm is used to rectify the pixel points corresponding to the extracted road elements, and the multi-head centernet key point detection model is used to extract the two-dimensional information of the key points on the basis of the rectified pixel points.

[0093] The data set for training the YOLOv8-OBB rotated bounding box model and the data set for training the multi-head centernet key point detection model are both generated through a custom road element identification template, without the need for annotation, thus avoiding the manual annotation cost of data set generation. The above-mentioned custom road element identification template is a template manually identified or extracted by the user on the two-dimensional intensity projection image according to requirements, and the key points of the template are set. As Figure 2B shown is a schematic diagram of the road element identification template, and as Figure 2C shown is a schematic diagram of the key points on the road element identification template.

[0094] In some embodiments, the road element deep learning model is constructed according to the category data of the road element identification template and the number of key points of the corresponding category. In this embodiment, the YOLOv8-OBB rotated bounding box model and the multi-head centernet key point detection model are used together to construct a deep learning model for road element extraction.

[0095] S140. Determine the two-dimensional coordinate information of the key point two-dimensional coordinate information on the two-dimensional intensity projection image according to the reverse logic of the processing logic of the road element deep learning model.

[0096] Realize the mapping of the two-dimensional information of the extracted key points onto the two-dimensional intensity projection image.

[0097] S150. Determine the 3D coordinates of the road element on the 3D point cloud image according to the reverse logic of the processing logic of the intensity projection and the 2D coordinate information of the key point on the 2D intensity projection image.

[0098] The 2D information of the key points mapped to the 2D intensity projection image is mapped to the 3D point cloud image.

[0099] Among them, the data set for training the multi-head centernet key point detection model is obtained according to the following steps:

[0100] Step 1. Obtain multiple 3D point cloud sample images, and perform intensity projection on each 3D point cloud sample image respectively to obtain multiple 2D intensity sample projection images.

[0101] The algorithm for projecting the 3D point cloud sample image onto the 2D intensity sample projection image here is the same as the algorithm for performing intensity projection in S120.

[0102] Step 2. Use each 2D intensity sample projection image to determine multiple road element identification templates and the key points corresponding to each road element identification template.

[0103] Here, the road element identification template and the key points corresponding to each road element identification template can be determined manually from the 2D intensity sample projection image.

[0104] Step 3. Fill each road element identification template with the first target scene color, and perform at least one rotation respectively within the first preset angle range to obtain multiple first road element template images, and use the data set formed by the first road element template images as the data set for training the multi-head centernet key point detection model.

[0105] The above first target scene color can be white, or other colors matching the scene.

[0106] Among them, the data set for training the YOLOv8-OBB rotation box model is obtained according to the following steps:

[0107] Step 1. Fill each road element identification template with the second target scene color, and perform at least one rotation respectively within the second preset angle range to obtain multiple second road element template images; paste the obtained second road element template images at at least one randomly determined position onto at least one road background image to obtain multiple training sample images.

[0108] The above second target scene color can be white, or other colors matching the scene.

[0109] Although the above pasting positions are random, they are known quantities, so manual annotation is not required.

[0110] Step 2: Use the dataset formed by the training sample images as the dataset for training the YOLOv8-OBB rotated bounding box model.

[0111] The above dataset can include a training set and a test set, and the model training is completed jointly using the training set and the test set.

[0112] Among them, the deep learning model for road elements is as Figure 3 shown, and the processing logic of the deep learning model for road elements is as follows:

[0113] Step 1: Use the YOLOv8-OBB rotated bounding box model to extract the rotated bounding box position information of road elements from the two-dimensional intensity projection image; the rotated bounding box position information is two-dimensional coordinate information.

[0114] The rotated bounding box position information extracted by the above YOLOv8-OBB rotated bounding box model can be the coordinates of 4 two-dimensional points. In addition, the YOLOv8-OBB rotated bounding box model can also output the class information of the rotated bounding box.

[0115] Step 2: The calculation module uses the PCA algorithm to perform principal component analysis on the rotated bounding box position information of road elements to obtain the main direction of the rotated bounding box.

[0116] Specifically, first, according to the rotated bounding box position information of the road elements, determine the two-dimensional coordinates of the centroid of the rotated bounding box.

[0117] The two-dimensional coordinates of the centroid of the rotated bounding box can be determined using the following formula:

[0118]

[0119] Among them, (C x , C y ) represents the two-dimensional coordinates of the centroid, (x i , y i ) represents the coordinates of the boundary point i, and n represents the number of boundary points of the rotated bounding box.

[0120] After that, use the two-dimensional coordinates of the centroid to perform a centering process on each boundary point of the rotated bounding box to obtain the centered points corresponding to each boundary point.

[0121] For each boundary point, use the following formula to perform the centering process:

[0122] P′ i =(x i -C x ,y i -C y )

[0123] Among them, (C x , C y ) represents the two-dimensional coordinates of the centroid, (x i , y i ) represents the coordinates of the boundary point i, and P′ i represents the decentralized point obtained by the decentralization process.

[0124] After that, the covariance matrix is determined using the coordinates of each decentralized point.

[0125] The covariance matrix Cov is as follows:

[0126]

[0127] Finally, eigenvalue decomposition is performed on the covariance matrix, and the main direction of the rotated box is determined according to the result of the eigenvalue decomposition.

[0128] Performing eigenvalue decomposition on the covariance matrix, the eigenvectors give the main axis directions of the OBB (Oriented Bounding Box), the direction with the larger eigenvalue is the long axis direction of the OBB, and the direction with the smaller eigenvalue is the short axis direction. Among them, the eigenvector V represents the direction of the OBB, and the eigenvalue Λ represents the variance of the data in this direction, specifically as follows:

[0129] Cov = VΛV T

[0130] Step 3: According to the angle between the main direction of the rotated box and the vertical direction, the pixel points corresponding to the road elements on the two-dimensional intensity projection image are rotated to be upright.

[0131] Step 4: Use the multi-head centernet key point detection model to extract key points from the rotated road elements to obtain the two-dimensional coordinate information of the key points.

[0132] In some embodiments, the three-dimensional point cloud image is intensity-projected to obtain a two-dimensional intensity projection image, which can be specifically implemented using the following steps:

[0133] Step 1: Determine the width and height of the point cloud according to the position information of each point in the three-dimensional point cloud image.

[0134] Specifically, a point in the point cloud is represented as:

[0135]

[0136] In the formula, (x i , y i , Z i ) are the coordinates of point i, and I iis the intensity information of point i, and m is the number of points in the 3D point cloud image.

[0137] After that, according to the position information of each point in the 3D point cloud image, determine the maximum value max of the point cloud in the x direction x and the minimum value min x , that is

[0138]

[0139] According to the position information of each point in the 3D point cloud image, determine the maximum value max of the point cloud in the y direction y and the minimum value min y , that is

[0140]

[0141] After that, according to the maximum value of the point cloud in the x direction, the minimum value of the point cloud in the x direction, and the preset resolution resolution, determine the width width of the point cloud, that is

[0142]

[0143] In the formula, the preset resolution is set according to the actual situation, for example, set to 0.2.

[0144] According to the maximum value of the point cloud in the y direction, the minimum value of the point cloud in the y direction, and the preset resolution, determine the height height of the point cloud, that is

[0145]

[0146] The above x direction and y direction are perpendicular to each other. The x direction can be the horizontal direction, and the y direction can be the vertical direction.

[0147] Step 2: According to the width and height of the point cloud, map each point in the 3D point cloud image to the image grid to obtain the position index of each point in the two-dimensional intensity projection image.

[0148] For each point, determine the first index of the point in the y direction according to the coordinate of the point in the y direction, the maximum value of the point cloud in the y direction, and the preset resolution; determine the second index of the point in the x direction according to the coordinate of the point in the x direction, the maximum value of the point cloud in the x direction, and the preset resolution; use the first index and the second index as the position index of the point in the two-dimensional intensity projection image.

[0149] Specifically, the position index (A, B) of point i can be determined by the following formula:

[0150]

[0151] The above A and B are the indices of the column and row of the point in the two-dimensional intensity projection image.

[0152] This step realizes mapping the three-dimensional point cloud to the image grid.

[0153] Step 3: For the image grid with only one point, fill the intensity of the point in the image grid into the corresponding image network; for the image grid with multiple point clouds, fill the maximum intensity into the corresponding image network.

[0154] It can be expressed by the following formula:

[0155] grid intensity[A][B] = max(grid intensity[a][B], I i )

[0156] grid intensity is the intensity of the image network.

[0157] The above embodiments are proposed for the problems of low extraction accuracy, poor robustness, difficulty in projecting into 3D point cloud data, high artificial cost in dataset production, and poor real-time performance when automatically extracting road elements during the construction of a high-precision map, and can effectively solve the above problems.

[0158] The present disclosure provides a map generation method, including:

[0159] Step 1: Using the element extraction method based on zero-shot deep learning in any of the above embodiments, determine the three-dimensional coordinates of each road element in the three-dimensional point cloud image on the three-dimensional point cloud image.

[0160] Step 2: Generate a map including road elements according to the determined three-dimensional coordinates.

[0161] Using the element extraction method based on zero-shot deep learning in any of the above embodiments to construct a map can improve the accuracy of the constructed map, thereby facilitating improving the safety and speed of autonomous driving.

[0162] Based on the same inventive concept, the present disclosure provides an element extraction device based on zero-shot deep learning. The steps executed by the components of the device are the same as or similar to those of the above element extraction method based on zero-shot deep learning. Therefore, similar parts will not be elaborated. As Figure 4 shown, the element extraction device based on zero-shot deep learning in this embodiment includes:

[0163] A data acquisition module 410, configured to acquire a pre-trained road element deep learning model and a three-dimensional point cloud image that needs to perform road element extraction.

[0164] An intensity projection module 420, which is used to perform intensity projection on the three-dimensional point cloud image to obtain a two-dimensional intensity projection image.

[0165] A key point extraction module 430, which is used to input the two-dimensional intensity projection image into the road element deep learning model to obtain the two-dimensional coordinate information of the key points of the road elements output by the road element deep learning model; the road element deep learning model sequentially includes a YOLOv8-OBB rotation box model, a calculation module including a PCA algorithm, and a multi-head centernet key point detection model.

[0166] A first reverse logic processing module 440, which is used to determine the two-dimensional coordinate information of the key points on the two-dimensional intensity projection image according to the reverse logic of the processing logic of the road element deep learning model.

[0167] A second reverse logic processing module 450, which is used to determine the three-dimensional coordinates of the road elements on the three-dimensional point cloud image according to the reverse logic of the processing logic of the intensity projection and the two-dimensional coordinate information of the key points on the two-dimensional intensity projection image.

[0168] Wherein, the key point extraction module 430 is further used to determine the data set for training the multi-head centernet key point detection model:

[0169] Obtain multiple three-dimensional point cloud sample images, and perform intensity projection on each three-dimensional point cloud sample image respectively to obtain multiple two-dimensional intensity sample projection images;

[0170] Use each two-dimensional intensity sample projection image to determine multiple road element identification templates and the key points corresponding to each road element identification template;

[0171] Fill each road element identification template with a first target scene color, and rotate it at least once respectively according to a first preset angle range to obtain multiple first road element template images, and use the data set formed by the first road element template images as the data set for training the multi-head centernet key point detection model;

[0172] Wherein, the key point extraction module 430 is further used to determine the data set for training the YOLOv8-OBB rotation box model:

[0173] Fill each road element identification template with a second target scene color, and rotate it at least once respectively according to a second preset angle range to obtain multiple second road element template images; paste the obtained second road element template images at at least one randomly determined position on at least one road background image to obtain multiple training sample images;

[0174] Use the data set formed by the training sample images as the data set for training the YOLOv8-OBB rotated bounding box model;

[0175] Among them, the processing logic of the road element deep learning model is as follows:

[0176] Use the YOLOv8-OBB rotated bounding box model to extract the rotated bounding box position information of road elements from the two-dimensional intensity projection image; the rotated bounding box position information is two-dimensional coordinate information;

[0177] Use the PCA algorithm to perform principal component analysis on the rotated bounding box position information of road elements to obtain the main direction of the rotated bounding box;

[0178] According to the angle between the main direction of the rotated bounding box and the vertical direction, rotate the pixel points corresponding to the road elements on the two-dimensional intensity projection image;

[0179] Use the multi-head centernet keypoint detection model to extract keypoints from the rotated road elements to obtain the two-dimensional coordinate information of the keypoints.

[0180] Based on the same inventive concept, the present disclosure provides a map generation device. The steps executed by the components of this device are the same as or similar to those of the above map generation method. Therefore, similar parts will not be elaborated. The map generation device of this embodiment includes:

[0181] A three-dimensional coordinate determination module, configured to use the element extraction method based on zero-shot deep learning in any of the above embodiments to determine the three-dimensional coordinates of each road element in the three-dimensional point cloud image on the three-dimensional point cloud image.

[0182] A map generation module, configured to generate a map including road elements according to the determined three-dimensional coordinates.

[0183] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a computer-readable storage medium.

[0184] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0185] As Figure 5As shown, device 500 includes a computing unit 510, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 520 or a computer program loaded from a storage unit 580 into a random access memory (RAM) 530. In the RAM 530, various programs and data required for the operation of the device 500 can also be stored. The computing unit 510, the ROM 520, and the RAM 530 are connected to each other via a bus 540. An input / output (I / O) interface 550 is also connected to the bus 540.

[0186] Multiple components in the device 500 are connected to the I / O interface 550, including: an input unit 560, such as a keyboard, a mouse, etc.; an output unit 570, such as various types of displays, speakers, etc.; a storage unit 580, such as a magnetic disk, an optical disc, etc.; and a communication unit 590, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 590 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0187] The computing unit 510 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 510 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 510 executes the various methods and processes described above. For example, in some embodiments, any of the above methods can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 580. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 520 and / or the communication unit 590. When the computer program is loaded into the RAM 530 and executed by the computing unit 510, one or more steps of any of the above-described methods can be executed. Alternatively, in other embodiments, the computing unit 510 can be configured to execute any of the above-described methods in any other appropriate manner (e.g., by means of firmware).

[0188] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0189] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0190] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include electrical connections based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0191] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0192] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0193] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0194] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0195] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A feature extraction method based on zero-shot deep learning, characterized in that, Including: Obtain a pre-trained deep learning model for road elements and a 3D point cloud image for which road element extraction is required; Perform intensity projection on the 3D point cloud image to obtain a 2D intensity projection image; Input the 2D intensity projection image into the deep learning model for road elements to obtain the two-dimensional coordinate information of the key points of the road elements output by the deep learning model for road elements; the deep learning model for road elements sequentially includes a YOLOv8-OBB rotated bounding box model, a calculation module including a PCA algorithm, and a multi-head centernet key point detection model; According to the reverse logic of the processing logic of the deep learning model for road elements, determine the two-dimensional coordinate information of the key points on the 2D intensity projection image; According to the reverse logic of the processing logic of the intensity projection and the two-dimensional coordinate information of the key points on the 2D intensity projection image, determine the three-dimensional coordinates of the road elements on the 3D point cloud image; Among them, the data set for training the multi-head centernet key point detection model is obtained according to the following steps: Obtain multiple 3D point cloud sample images, and perform intensity projection on each 3D point cloud sample image to obtain multiple 2D intensity sample projection images; Use each 2D intensity sample projection image to determine multiple road element identification templates and the key points corresponding to each road element identification template; Fill each road element identification template with the first target scene color, and rotate it at least once within the first preset angle range to obtain multiple first road element template images, and use the data set formed by the first road element template images as the data set for training the multi-head centernet key point detection model; Among them, the data set for training the YOLOv8-OBB rotated bounding box model is obtained according to the following steps: Fill each road element identification template with the second target scene color, and rotate it at least once within the second preset angle range to obtain multiple second road element template images; paste the obtained second road element template images at at least one randomly determined position onto at least one road background image to obtain multiple training sample images; Use the data set formed by the training sample images as the data set for training the YOLOv8-OBB rotated bounding box model; Among them, the processing logic of the deep learning model for road elements is as follows: Use the YOLOv8-OBB rotated bounding box model to extract the rotated bounding box position information of the road elements from the 2D intensity projection image; the rotated bounding box position information is two-dimensional coordinate information; Use the PCA algorithm to perform principal component analysis on the rotated bounding box position information of the road elements to obtain the main direction of the rotated bounding box; According to the angle between the main direction of the rotated bounding box and the vertical direction, perform rotation correction on the pixel points corresponding to the road elements on the 2D intensity projection image; Use the multi-head centernet key point detection model to extract key points from the rotation-corrected road elements to obtain the two-dimensional coordinate information of the key points.

2. The method according to claim 1, wherein The performing intensity projection on the 3D point cloud image to obtain a 2D intensity projection image includes: Determine the width and height of the point cloud according to the position information of each point in the three-dimensional point cloud image; Map each point in the three-dimensional point cloud image to an image grid according to the width and height of the point cloud, and obtain the position index of each point in the two-dimensional intensity projection image; For an image grid with only one point, fill the intensity of the point in the image grid into the corresponding image network; for an image grid with multiple point clouds, fill the maximum intensity into the corresponding image network.

3. The method according to claim 2, characterized in that The determining the width and height of the point cloud according to the position information of each point in the three-dimensional point cloud image includes: Determine the maximum and minimum values of the point cloud in the x direction according to the position information of each point in the three-dimensional point cloud image; Determine the width of the point cloud according to the maximum value of the point cloud in the x direction, the minimum value of the point cloud in the x direction, and the preset resolution; Determine the maximum and minimum values of the point cloud in the y direction according to the position information of each point in the three-dimensional point cloud image; Determine the height of the point cloud according to the maximum value of the point cloud in the y direction, the minimum value of the point cloud in the y direction, and the preset resolution; Where the x direction and the y direction are perpendicular to each other.

4. The method according to claim 3, wherein The mapping each point in the three-dimensional point cloud image to an image grid according to the width and height of the point cloud, and obtaining the position index of each point in the two-dimensional intensity projection image includes: For each point, determine the first index of the point in the y direction according to the coordinate of the point in the y direction, the maximum value of the point cloud in the y direction, and the preset resolution; determine the second index of the point in the x direction according to the coordinate of the point in the x direction, the maximum value of the point cloud in the x direction, and the preset resolution; use the first index and the second index as the position index of the point in the two-dimensional intensity projection image.

5. The method according to claim 1, characterized in that, The using the PCA algorithm to perform principal component analysis on the rotation box position information of the road elements to obtain the main direction of the rotation box includes: Determine the two-dimensional coordinates of the centroid of the rotation box according to the rotation box position information of the road elements; Perform de-centralization processing on each boundary point of the rotation box by using the two-dimensional coordinates of the centroid to obtain the corresponding de-centralized point of each boundary point; Determine the covariance matrix by using the coordinates of each de-centralized point; Perform eigenvalue decomposition on the covariance matrix, and determine the main direction of the rotation box according to the result of the eigenvalue decomposition.

6. The method according to claim 5, wherein The performing de-centralization processing on each boundary point of the rotation box by using the two-dimensional coordinates of the centroid to obtain the corresponding de-centralized point of each boundary point includes: For each boundary point, perform the de-centralization processing by using the following formula: P′ i =(x i -C x ,y i -C y ) Among them, (C x , C y ) represents the two-dimensional coordinates of the centroid, (x i , y i ) represents the coordinates of the boundary point i, and P′ i represents the decentralized point obtained through the decentralization process.

7. The method according to claim 1, wherein The first target scene color and / or the second target scene color is white.

8. An element extraction device based on zero-shot deep learning, characterized in that, Including: A data acquisition module, configured to acquire a pre-trained deep learning model of road elements and a three-dimensional point cloud image that needs to extract road elements; An intensity projection module, configured to perform intensity projection on the three-dimensional point cloud image to obtain a two-dimensional intensity projection image; The key point extraction module is used to input the two-dimensional intensity projection image into the road element deep learning model to obtain the two-dimensional coordinate information of the key points of the road elements output by the road element deep learning model; the road element deep learning model sequentially includes a YOLOv8-OBB rotated box model, a calculation module including a PCA algorithm, and a multi-head centernet key point detection model; The first reverse logic processing module is used to determine the two-dimensional coordinate information of the key point two-dimensional coordinate information on the two-dimensional intensity projection image according to the reverse logic of the processing logic of the road element deep learning model; The second reverse logic processing module is used to determine the three-dimensional coordinates of the road elements on the three-dimensional point cloud image according to the reverse logic of the processing logic of the intensity projection and the two-dimensional coordinate information of the key point two-dimensional coordinate information on the two-dimensional intensity projection image; Wherein, the key point extraction module is further used to determine the data set for training the multi-head centernet key point detection model: Obtain multiple three-dimensional point cloud sample images, and perform intensity projection on each three-dimensional point cloud sample image to obtain multiple two-dimensional intensity sample projection images; Use each two-dimensional intensity sample projection image to determine multiple road element identification templates and the key points corresponding to each road element identification template; Fill each road element identification template with the first target scene color, and rotate it at least once according to the first preset angle range to obtain multiple first road element template images, and use the data set formed by the first road element template images as the data set for training the multi-head centernet key point detection model; Wherein, the key point extraction module is further used to determine the data set for training the YOLOv8-OBB rotated box model: Fill each road element identification template with the second target scene color, and rotate it at least once according to the second preset angle range to obtain multiple second road element template images; paste the obtained second road element template images onto at least one road background image at at least one randomly determined position to obtain multiple training sample images; Use the data set formed by the training sample images as the data set for training the YOLOv8-OBB rotated box model; Wherein, the processing logic of the road element deep learning model is as follows: Use the YOLOv8-OBB rotated box model to extract the rotated box position information of the road elements from the two-dimensional intensity projection image; the rotated box position information is two-dimensional coordinate information; Use the PCA algorithm to perform principal component analysis on the rotated box position information of the road elements to obtain the main direction of the rotated box; According to the angle between the main direction of the rotated box and the vertical direction, perform rotation correction on the pixel points corresponding to the road elements on the two-dimensional intensity projection image; Use the multi-head centernet key point detection model to extract key points from the rotation-corrected road elements to obtain the two-dimensional coordinate information of the key points.

9. A method for generating a map, characterized in that, Including: Use the method according to any one of claims 1 to 8 to determine the three-dimensional coordinates of each road element in the three-dimensional point cloud image; Generate a map including road elements based on the determined three-dimensional coordinates.

10. A map generation device, characterized in that, Including: A three-dimensional coordinate determination module, configured to use the method according to any one of claims 1 to 8 to determine the three-dimensional coordinates of each road element in the three-dimensional point cloud image on the three-dimensional point cloud image; A map generation module, configured to generate a map including road elements according to the determined three-dimensional coordinates.

Citation Information

Patent Citations

  • Map construction method and device based on laser point cloud

    CN110400363A

  • Multi-level semantic map construction method and device based on deep learning perception

    CN115655262A