Occupancy Grid-based Construction Method, Device and System for Mining Area Datasets

By constructing the occupancy raster dataset in the mining area environment, the problem of low accuracy of the mine's autonomous driving perception algorithm is solved, more refined environmental modeling is achieved, and the operation efficiency and safety of the mine are improved.

CN118968451BActive Publication Date: 2025-05-27WUXI QINGLIAN INTELLIGENT MINING TECHNOLOGY IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411029386.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-05-27
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

The existing technology is difficult to improve the accuracy of autonomous driving perception algorithms in mines, especially in unstructured road scenarios, where road areas and non-road areas are difficult to distinguish, and the lidar point cloud density is sparse, resulting in data set labeling challenges.

Method used

By obtaining the environmental image information and point cloud information of the roads and working areas of the mining area, a preliminary data set of the target object detection frame and semantic information is constructed, and the depth completion model is used for dense processing. Combined with dynamic and static point cloud separation and superposition, a grid three-dimensional reconstruction algorithm is used to generate an occupation raster, and the semantic information is assigned, and the occupation raster data set is finally constructed in the mining area environment.

Benefits of technology

It improves the perception accuracy of the mine’s autonomous driving perception algorithm, improves the operation efficiency, safety and economic benefits of the mine, and promotes the application of autonomous driving technology in mines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968451B_ABST
    Figure CN118968451B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of intelligent mines, and specifically discloses a method, device and system for constructing a mining area dataset based on occupancy grids, including: obtaining environmental image information and environmental point cloud information of the mining area road and the mining area operation area; constructing a preliminary dataset according to the environmental image information and the environmental point cloud information; constructing a depth completion model according to the preliminary dataset; performing dynamic and static point cloud separation and superposition processing on the initial dense point cloud data; performing three-dimensional reconstruction processing on the dense complete scene point cloud data according to the grid three-dimensional reconstruction algorithm; obtaining occupancy grid semantic information according to the dense complete scene point cloud data and the complete scene occupancy grid; and forming a dataset by representing all the semantic occupancy grids of the complete scenes to obtain an occupancy grid dataset in the mining area environment. The method for constructing a mining area dataset based on occupancy grids provided by the present invention can improve the accuracy of the autonomous driving perception algorithm in mines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart mine technology, and in particular to a method for constructing a mining area dataset based on an occupancy grid, a device for constructing a mining area dataset based on an occupancy grid, a system for constructing a mining area dataset based on an occupancy grid, and an autonomous driving perception system. Background Art

[0002] Mineral resources, as important raw materials for industrial production, include metal minerals, coal and other resources. They are of great significance to the national economic development and industrial manufacturing, and are an important foundation for supporting the national economic operation. Mines are important sources of mineral resource mining and production bases. Ensuring their efficient and safe operation is of great significance to mineral development. However, traditional mining methods have problems such as difficulty in recruiting workers, great safety hazards, and high operating costs. It is urgent to introduce more intelligent technical means to improve efficiency and safety.

[0003] By integrating a variety of intelligent algorithms, autonomous driving technology can significantly improve transportation safety and efficiency. The mining environment is a closed scene with fixed routes and slow vehicle speeds, making it one of the most suitable scenarios for the commercialization of unmanned driving. By introducing autonomous driving technology, intelligent control and unmanned operation of mining vehicles can be achieved, improving mining efficiency while reducing the risk of casualties. In addition, autonomous driving technology can make the operation of mining equipment more stable and reliable, reduce production and labor costs, and effectively improve economic benefits.

[0004] In the field of autonomous driving perception, the emerging occupancy grid technology has shown its superior accuracy and versatility. The autonomous driving perception dataset constructed in the traditional way only annotates the target object detection box information, while the dataset based on the occupancy grid technology describes the surrounding world more comprehensively through voxels and reflects the background other than the target object more finely, so that the corresponding autonomous driving algorithm can be trained to achieve higher perception accuracy. However, there are still few studies on occupancy grid datasets, and most datasets are still labeled in the traditional way. In addition, most datasets are still limited to structured road scenes such as urban roads, while mines are unstructured road scenes, where road areas and non-road areas are difficult to distinguish, and the scene is more open, resulting in sparse density of lidar point clouds, which further poses challenges to accurate dataset annotation.

[0005] Therefore, how to provide a method for constructing a mining area dataset to improve the perception accuracy of autonomous driving perception algorithms in mines has become a technical problem that needs to be urgently solved by technicians in this field. Summary of the invention

[0006] The present invention provides a method for constructing a mining area dataset based on an occupancy grid, a device for constructing a mining area dataset based on an occupancy grid, a system for constructing a mining area dataset based on an occupancy grid, and an autonomous driving perception system, so as to solve the problem existing in the related technology that the accuracy of the autonomous driving perception algorithm in mines cannot be improved.

[0007] As a first aspect of the present invention, a method for constructing a mining area dataset based on an occupancy grid is provided, which comprises:

[0008] Obtain environmental image information and environmental point cloud information of mining roads and mining operation areas;

[0009] Constructing a preliminary data set including a target object detection frame and semantic information according to the environmental image information and the environmental point cloud information;

[0010] Constructing a depth completion model according to the preliminary data set, and performing initial densification processing on the environmental point cloud information according to the depth completion model to obtain initial dense point cloud data;

[0011] Performing dynamic and static point cloud separation and superposition processing on the initial dense point cloud data to obtain dense complete scene point cloud data;

[0012] Performing three-dimensional reconstruction processing on the dense complete scene point cloud data according to a grid three-dimensional reconstruction algorithm to obtain a complete scene occupancy grid;

[0013] Acquiring occupancy grid semantic information according to the dense complete scene point cloud data and the complete scene occupancy grid, and obtaining a semantically represented occupancy grid representation of the complete scene;

[0014] The semantically represented occupancy grids of all complete scenes are combined into a dataset to obtain an occupancy grid dataset in a mining environment.

[0015] Furthermore, a preliminary data set including a target object detection frame and semantic information is constructed according to the environmental image information and the environmental point cloud information, including:

[0016] Preprocessing the environmental image information and the environmental point cloud information to remove abnormal data and obtain a preprocessed data set;

[0017] The preprocessed data set is annotated with semantic information, and a preliminary data set including a target object detection frame and semantic information is constructed based on the annotated preprocessed data set.

[0018] Further, a depth completion model is constructed according to the preliminary data set, and initial densification processing is performed on the environmental point cloud information according to the depth completion model to obtain initial dense point cloud data, including:

[0019] Train a neural network model according to the preliminary data set to obtain a depth completion model;

[0020] Input the environmental point cloud information into the depth completion model for data augmentation to obtain initial dense point cloud data.

[0021] Further, the depth completion model includes a data pre-fusion layer, a pre-trained backbone network module, an encoding network module, a decoding network module, and a dense point cloud output layer, which are sequentially arranged from input to output. Inputting the environmental point cloud information into the depth completion model for data augmentation to obtain initial dense point cloud data includes:

[0022] Input the environmental point cloud information and its corresponding environmental image information into the data pre-fusion layer for feature fusion processing to obtain a first feature map including image and depth information;

[0023] Input the first feature map into the pre-trained backbone network module for initialization processing to obtain a second feature map;

[0024] Input the second feature map into the encoding network module for encoding processing to obtain a fifth feature map, and the compactness of the fifth feature map is greater than that of the second feature map;

[0025] Input the fifth feature map into the decoding network module for decoding processing to obtain a dense depth map;

[0026] Input the dense depth map into the dense point cloud output layer for reduction processing to obtain initial dense point cloud data.

[0027] Further, perform dynamic and static point cloud separation and superposition processing on the initial dense point cloud data to obtain dense complete scene point cloud data, including:

[0028] Traverse the initial dense point cloud data of all frames within a preset time period;

[0029] Separate the dynamic and static parts of the initial dense point cloud data to obtain dynamic point clouds as dynamic objects and static point clouds as static scenes;

[0030] Unify the dynamic point clouds and the static point clouds to the world coordinate system, and superimpose the dynamic point clouds of different frames to obtain a dense dynamic object point cloud and superimpose the static point clouds of different frames to obtain a dense static scene point cloud;

[0031] Insert the dense dynamic object point cloud into the dense static scene point cloud and restore the complete scene to obtain dense complete scene point cloud data.

[0032] Further, perform three-dimensional reconstruction processing on the dense and complete scene point cloud data according to the grid three-dimensional reconstruction algorithm to obtain a complete scene occupancy grid, including:

[0033] Calculate the regional surface normal according to the distribution of the dense and complete scene point cloud data;

[0034] Render the regional surface normal to obtain a triangular facet grid map;

[0035] Perform rasterization processing on the triangular facet grid map to obtain a complete scene occupancy grid.

[0036] Further, obtain the occupancy grid semantic information according to the dense and complete scene point cloud data and the complete scene occupancy grid to obtain a semantic occupancy grid representation of the complete scene, including:

[0037] Find the grid closest to each point in the dense and complete scene point cloud data according to the nearest neighbor algorithm;

[0038] Assign semantic information to the closest grid to obtain a semantic occupancy grid representation of the complete scene.

[0039] As another aspect of the present invention, there is provided an occupancy grid-based mining area dataset construction device for implementing the occupancy grid-based mining area dataset construction method described above, where it includes:

[0040] An acquisition module for acquiring environmental image information and environmental point cloud information of the mining area road and the mining area operation area;

[0041] A preliminary dataset construction module for constructing a preliminary dataset including target object detection frames and semantic information according to the environmental image information and the environmental point cloud information;

[0042] An initial densification processing module for constructing a depth completion model according to the preliminary dataset and performing initial densification processing on the environmental point cloud information according to the depth completion model to obtain initial dense point cloud data;

[0043] A dynamic and static separation and superposition module for performing dynamic and static point cloud separation and superposition processing on the initial dense point cloud data to obtain dense and complete scene point cloud data;

[0044] A three-dimensional reconstruction module for performing three-dimensional reconstruction processing on the dense and complete scene point cloud data according to the grid three-dimensional reconstruction algorithm to obtain a complete scene occupancy grid;

[0045] A semantic assignment module for obtaining the occupancy grid semantic information according to the dense and complete scene point cloud data and the complete scene occupancy grid to obtain a semantic occupancy grid representation of the complete scene;

[0046] The module for obtaining the occupancy grid dataset is used to form a dataset of semantically represented occupancy grids of all complete scenes, and obtain the occupancy grid dataset in the mining environment.

[0047] As another aspect of the present invention, a system for constructing a mining area dataset based on an occupancy grid is provided, which comprises: a data acquisition vehicle, a multimodal sensor located on the data acquisition vehicle, and the above-mentioned mining area dataset construction device based on an occupancy grid, wherein the multimodal sensor is communicatively connected with the mining area dataset construction device based on the occupancy grid,

[0048] The multimodal sensor includes at least an image acquisition device and a radar device, and the image acquisition device and the radar device are capable of acquiring environmental image information and environmental point cloud information when the data acquisition vehicle travels on mining roads and working areas;

[0049] The occupancy grid-based mining area data set construction device can construct an occupancy grid data set in a mining area environment according to the environmental image information and the environmental point cloud information.

[0050] As another aspect of the present invention, an autonomous driving perception system is provided, which includes: an autonomous driving algorithm model and the above-mentioned mining area data set construction system based on occupancy grid,

[0051] The mining area data set construction system based on occupancy grid can construct an occupancy grid data set in a mining area environment;

[0052] The autonomous driving algorithm model can train a target autonomous driving algorithm according to the occupied grid data set in the mining environment to obtain a target autonomous driving algorithm model.

[0053] The method for constructing a mining area dataset based on an occupancy grid provided by the present invention collects environmental image information and environmental point cloud information of mining area roads and mining area working areas, and then constructs an occupancy grid dataset under a mining area environment based on the environmental image information and environmental point cloud information, which can make up for the lack of occupancy grid datasets in current mining area scenarios and promote the training and deployment of related algorithms; in addition, the embodiment of the present invention constructs a full process from raw data collection to mining area scene occupancy grid datasets in an end-to-end paradigm, opening up a complete chain of dataset construction steps; in view of the point cloud sparseness problem caused by the open mining area scene, a deep completion method is proposed to enhance data to ensure the accuracy of the constructed dataset; it can effectively promote the training and application of subsequent occupancy grid perception algorithms, so that the algorithms have more comprehensive and sophisticated environmental modeling capabilities, thereby further promoting the application of autonomous driving technology in mines, improving the perception accuracy of autonomous driving perception algorithms in mines, and thus improving the operating efficiency, safety and economic benefits of mines. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the following specific implementation manners, they are used to explain the present invention, but do not constitute a limitation to the present invention.

[0055] Figure 1 It is a flowchart of the method for constructing a mining area data set based on occupancy grids provided by the present invention.

[0056] Figure 2 It is a schematic structural diagram of the data acquisition vehicle provided by the present invention.

[0057] Figure 3 It is a flowchart of the construction of the preliminary data set provided by the present invention.

[0058] Figure 4 It is a flowchart of obtaining the initial dense point cloud data provided by the present invention.

[0059] Figure 5 It is a schematic structural diagram of the depth completion model provided by the present invention.

[0060] Figure 6 It is a flowchart of data enhancement through the depth completion model provided by the present invention.

[0061] Figure 7 It is a specific schematic structural diagram of the depth completion model provided by the present invention.

[0062] Figure 8 It is a flowchart of obtaining the dense and complete scene point cloud data provided by the present invention.

[0063] Figure 9 It is a flowchart of obtaining the complete scene occupancy grid provided by the present invention.

[0064] Figure 10 It is a flowchart of obtaining the semantic occupancy grid representation of the complete scene provided by the present invention.

[0065] Figure 11 It is an example of the result of the finally generated occupancy grid data set provided by the present invention.

[0066] Figure 12 It is a structural block diagram of the device for constructing a mining area data set based on occupancy grids provided by the present invention.

[0067] Figure 13 It is a structural block diagram of the system for constructing a mining area data set based on occupancy grids provided by the present invention.

[0068] Figure 14 It is a structural block diagram of the autonomous driving perception system provided by the present invention. Specific Implementation Manner

[0069] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0070] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0071] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present invention described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0072] In this embodiment, a method for constructing a mining area dataset based on occupancy grids is provided. Figure 1 is a flowchart of the method for constructing a mining area dataset based on occupancy grids according to the embodiment of the present invention, as Figure 1 shown, including:

[0073] S100. Obtain environmental image information and environmental point cloud information of the mining area road and the mining area operation area;

[0074] In the embodiment of the present invention, as Figure 2 shown, the data acquisition vehicle is equipped with multi-modal sensors and travels on the mining area road and the operation area to collect panoramic environmental images and point clouds. The multi-modal sensors in the embodiment of the present invention may specifically include lidar and multiple panoramic cameras, etc.

[0075] S200. Construct a preliminary dataset including target object detection frames and semantic information according to the environmental image information and the environmental point cloud information;

[0076] After collecting the environmental image information and the environmental point cloud information, preprocess the environmental image information and the environmental point cloud information and perform semantic information annotation, etc., to obtain a preliminary dataset including target detection frames and semantic information.

[0077] S300. Construct a depth completion model based on the preliminary data set, and perform initial densification processing on the environmental point cloud information according to the depth completion model to obtain initial dense point cloud data;

[0078] In the embodiment of the present invention, a depth completion model is constructed according to the preliminary data set, and the depth completion model can perform densification processing on the environmental point cloud information.

[0079] Specifically, the depth completion model in the embodiment of the present invention can be specifically obtained by training a neural network model with the preliminary data set.

[0080] S400. Perform dynamic and static point cloud separation and superposition processing on the initial dense point cloud data to obtain dense and complete scene point cloud data;

[0081] In the embodiment of the present invention, further densification processing is performed on the initial dense point cloud data to obtain dense and complete scene point cloud data.

[0082] S500. Perform three-dimensional reconstruction processing on the dense and complete scene point cloud data according to the grid three-dimensional reconstruction algorithm to obtain a complete scene occupancy grid;

[0083] In the embodiment of the present invention, three-dimensional reconstruction processing is performed on the dense and complete scene point cloud data through the three-dimensional reconstruction algorithm, and a complete scene occupancy grid representation can be obtained.

[0084] S600. Obtain semantic information of the occupancy grid according to the dense and complete scene point cloud data and the complete scene occupancy grid to obtain a semantic occupancy grid representation of the complete scene;

[0085] Semantic information is given to the obtained occupancy grid representation of the complete scene to obtain a semantic occupancy grid representation of the complete scene with semantics.

[0086] S700. Combine the semantic occupancy grid representations of all complete scenes into a data set to obtain an occupancy grid data set in the mining area environment.

[0087] After removing invalid data from the semantic occupancy grid representations of all complete scenes and combining them into a data set, an occupancy grid data set in the mining area environment is obtained.

[0088] The occupancy grid data set in the mining area environment can be input into an autonomous driving algorithm model for training to obtain a target autonomous driving algorithm model.

[0089] Therefore, the method for constructing a mining area dataset based on an occupancy grid provided in an embodiment of the present invention collects environmental image information and environmental point cloud information of mining area roads and mining operation areas, and then constructs an occupancy grid dataset in a mining environment based on the environmental image information and environmental point cloud information, which can make up for the lack of occupancy grid datasets in current mining scenarios and promote the training and deployment of related algorithms; in addition, the embodiment of the present invention uses an end-to-end paradigm to construct a full process from raw data collection to mining scene occupancy grid datasets, opening up a complete chain of dataset construction steps; in response to the point cloud sparseness problem caused by the open mining scene, a deep completion method is proposed to enhance data to ensure the accuracy of the constructed dataset; it can effectively promote the subsequent training and application of occupancy grid perception algorithms, so that the algorithms have more comprehensive and sophisticated environmental modeling capabilities, thereby further promoting the application of autonomous driving technology in mines, improving the perception accuracy of autonomous driving perception algorithms in mines, and thereby improving the operating efficiency, safety and economic benefits of mines.

[0090] In an embodiment of the present invention, a preliminary data set including a target object detection frame and semantic information is constructed based on the environmental image information and the environmental point cloud information, such as Figure 3 As shown, including:

[0091] S210, preprocessing the environmental image information and the environmental point cloud information to remove abnormal data and obtain a preprocessed data set;

[0092] In an embodiment of the present invention, each frame of image and point cloud is preprocessed to remove invalid and abnormal data, repair data noise such as image distortion caused by environmental interference, and coordinate system alignment and time synchronization are performed on different sensor data.

[0093] S220 , annotating the preprocessed data set with semantic information, and constructing a preliminary data set including a target object detection frame and semantic information based on the annotated preprocessed data set.

[0094] In the embodiment of the present invention, the data is specifically labeled and reviewed, and a data set including a target object detection frame and semantic information is preliminarily constructed.

[0095] After obtaining the preliminary data set, specifically, a depth completion model is constructed according to the preliminary data set, and the environment point cloud information is initially densified according to the depth completion model to obtain initial dense point cloud data, such as Figure 4 As shown, including:

[0096] S310, training a neural network model according to the preliminary data set to obtain a depth completion model;

[0097] It should be understood here that the depth completion model is obtained by training a preliminary data set based on a neural network model. Therefore, the structural composition of the depth completion model is actually the same as that of the neural network model.

[0098] The depth completion model can perform data augmentation on the input sparse point cloud to complete the preliminary densification process of the point cloud.

[0099] S320. Input the environmental point cloud information into the depth completion model for data augmentation to obtain initial dense point cloud data.

[0100] In the embodiment of the present invention, as Figure 5 shown, the depth completion model includes a data pre-fusion layer, a pre-trained backbone network module, an encoding network module, a decoding network module, and a dense point cloud output layer, which are arranged in sequence from input to output.

[0101] Specifically, inputting the environmental point cloud information into the depth completion model for data augmentation to obtain initial dense point cloud data, as Figure 6 shown, includes:

[0102] S321. Input the environmental point cloud information and its corresponding environmental image information into the data pre-fusion layer for feature fusion processing to obtain a first feature map including image and depth information;

[0103] In the embodiment of the present invention, the data pre-fusion layer can process the input data as follows: input sparse point cloud and the corresponding RGB image Using the distance of the point cloud as the pixel value, project the point cloud onto a two-dimensional plane to obtain a sparse depth map corresponding to the point cloud; then, after downsampling, align the size of the sparse depth map with the size of the RGB image; merge the RGB image and the sparse depth map in the channel dimension to obtain a feature map containing both image and depth information and output it. Denote this feature map as the first feature map

[0104] S322. Input the first feature map into the pre-trained backbone network module for initialization processing to obtain a second feature map;

[0105] In the embodiment of the present invention, the backbone network selects a pre-trained Resnet-18 network, which uses the pre-trained parameters of ImageNet as the initialization parameters, and the input is the first feature map F output by the data pre-fusion layer 1 , and outputs a second feature map F with a more compact scale 2 .

[0106] S323. Input the second feature map into the encoding network module for encoding processing to obtain a fifth feature map, where the compactness of the fifth feature map is greater than that of the second feature map.

[0107] In the embodiment of the present invention, the encoding network module is as Figure 7 shown, and is successively composed of a first convolutional layer, a normalization layer, and a second convolutional layer. The convolutional kernel size of the first convolutional layer is 3*3. Taking the second feature map F 2 as the input, and outputting a third feature map F 3 with a more compact scale. The activation function selects the Leaky ReLU function to alleviate the problem of zero gradient optimization for negative values during the network learning process. The function expression is as shown in the following formula:

[0108]

[0109] where x represents the input feature, and α represents the negative slope parameter, taking an empirical value of 0.01.

[0110] Further input the third feature map F 3 into the normalization layer. While retaining the horizontal and vertical scales of the feature map through 1*1 convolution, change the channel dimension, and perform normalization processing through batch normalization. Select appropriate mean and variance to adjust the feature map distribution, and output a fourth feature map F 4 ; The convolutional kernel size of the second convolutional layer is 3*3. Taking the fourth feature map F 4 as the input, and further outputting a feature with a more compact scale, that is, a fifth feature map F 5 , and the activation function still selects the Leaky ReLU function shown in the above formula.

[0111] S324. Input the fifth feature map into the decoding network module for decoding processing to obtain a dense depth map.

[0112] The decoding network module is successively composed of four transposed convolutional layers, a third convolutional layer, and a bilinear interpolation layer. The convolutional kernel sizes of the four transposed convolutional layers are all 3*3. Taking the fifth feature map F 5 as the input, through transposed convolution operation, perform weighted summation on each pixel of the input feature map and its surrounding pixels, so as to gradually upsample the feature map to obtain an upsampled sixth feature map F 6 , and the activation function of the transposed convolutional layer selects the ELU function. The function expression is as shown in the following formula:

[0113]

[0114] where x represents the input feature, and α is an adjustable parameter, taking an empirical value of 1.

[0115] The convolutional kernel size of the third - layer base is 3*3, using the sixth feature map F 6 as the input, and outputting a preliminary dense depth map representation; the bilinear interpolation layer takes the preliminary dense depth map as the input, restores the depth map size through interpolation, and outputs the final dense depth map representation P 2 ′ 。

[0116] The loss function selects the mean absolute error loss (MAE Loss), and its function expression is as shown in the following formula:

[0117]

[0118] where n is the number of training samples, P 2 and P 2 ′ respectively represent the ground - truth depth map and the predicted depth map.

[0119] S325. Input the dense depth map into the dense point cloud output layer for restoration processing to obtain initial dense point cloud data.

[0120] It should be understood that the dense point cloud output layer restores the two - dimensional depth map to a three - dimensional point cloud according to the coordinates and depth values at each pixel of the final dense depth map, and finally obtains a dense point cloud representation.

[0121] In the embodiment of the present invention, the initial dense point cloud data is subjected to dynamic and static point cloud separation and superposition processing to obtain dense complete scene point cloud data, as Figure 8 shown, including:

[0122] S410. Traverse the initial dense point cloud data of all frames within a preset time period;

[0123] S420. Separate the dynamic and static parts of the initial dense point cloud data to obtain dynamic point clouds as dynamic objects and static point clouds as static scenes;

[0124] It should be noted that by traversing all point cloud frames collected within a certain time, for each frame of point cloud data, dynamic and static are distinguished. The point cloud within the detection frame is regarded as a dynamic object, and the remaining point cloud is regarded as a static scene, and the dynamic and static point clouds are stripped and collected separately.

[0125] S430. Unify the dynamic point cloud and the static point cloud to the world coordinate system, and respectively superpose the dynamic point clouds of different frames to obtain a dense dynamic object point cloud and superpose the static point clouds of different frames to obtain a dense static scene point cloud;

[0126] Specifically, the dynamic and static point clouds are unified into the world coordinate system, and the point clouds of dynamic objects in different frames are superimposed to obtain a dense point cloud of dynamic objects; similarly, the static point cloud frames in different frames are superimposed to obtain a dense point cloud of the static scene.

[0127] S440. Insert the dense point cloud of dynamic objects into the dense point cloud of the static scene, and restore the complete scene to obtain dense complete scene point cloud data.

[0128] In the embodiment of the present invention, according to the position and orientation of the dynamic object point cloud in the vehicle coordinate system, the dense dynamic point cloud is inserted into the dense static scene point cloud according to the original position, and the complete scene is restored to obtain a dense complete scene point cloud.

[0129] In the embodiment of the present invention, three-dimensional reconstruction processing is performed on the dense complete scene point cloud data according to the grid three-dimensional reconstruction algorithm to obtain the occupied grid of the complete scene, as Figure 9 shown, including:

[0130] S510. Calculate the regional surface normal according to the distribution of the dense complete scene point cloud data;

[0131] S520. Render the regional surface normal to obtain a triangular facet mesh map;

[0132] S530. Perform rasterization processing on the triangular facet mesh map to obtain the occupied grid of the complete scene.

[0133] In the embodiment of the present invention, the Poisson reconstruction algorithm is specifically used. First, the regional surface normal is calculated from the point cloud distribution, then the triangular facet network diagram is rendered from the surface normal, further filling the point cloud holes to make the point cloud distribution uniform, and finally the network diagram is voxelized to obtain the occupied grid representation of the complete scene.

[0134] In the embodiment of the present invention, the semantic information of the occupied grid is obtained according to the dense complete scene point cloud data and the occupied grid of the complete scene, and the semantic occupied grid representation of the complete scene is obtained, as Figure 10 shown, including:

[0135] S610. Find the grid closest to each point in the dense complete scene point cloud data according to the nearest neighbor algorithm;

[0136] S620. Assign semantic information to the closest grid to obtain the semantic occupied grid representation of the complete scene.

[0137] In the embodiment of the present invention, the KNN nearest neighbor algorithm can be specifically used. For each point in the point cloud, the closest grid is found, and semantic information is assigned to the grid, thereby obtaining the semantic occupied grid representation of the complete scene.

[0138] After obtaining the semantic occupancy grid representation of the complete scene, review and fine-tune the occupancy grid, remove invalid and incorrect data, and obtain the final complete scene semantic occupancy grid representation, as Figure 11 shown. The final multi-frame semantic occupancy grids can then be used as ground truth to form a dataset and provide supervision for the algorithm.

[0139] In summary, the method for constructing a mining area dataset based on occupancy grids provided by the present invention obtains a dataset in the form of occupancy grids, performs data augmentation for the problem of sparse point cloud density in the mining area scene, ensures the accuracy of the constructed dataset, promotes the training and deployment of the corresponding autonomous driving algorithm model in the mining area scene, and thus further promotes the application of autonomous driving technology in mines, improving the operation efficiency, safety and economic benefits of mines.

[0140] As another aspect of the present invention, there is provided an apparatus 100 for constructing a mining area dataset based on occupancy grids, which is used to implement the method for constructing a mining area dataset based on occupancy grids described above. Among them, as Figure 12 shown, it includes:

[0141] An acquisition module 110, configured to acquire environmental image information and environmental point cloud information of a mining area road and a mining area operation area;

[0142] A preliminary dataset construction module 120, configured to construct a preliminary dataset including target object detection frames and semantic information according to the environmental image information and the environmental point cloud information;

[0143] An initial densification processing module 130, configured to construct a depth completion model according to the preliminary dataset, and perform initial densification processing on the environmental point cloud information according to the depth completion model to obtain initial dense point cloud data;

[0144] A dynamic and static separation and superposition module 140, configured to perform dynamic and static point cloud separation and superposition processing on the initial dense point cloud data to obtain dense complete scene point cloud data;

[0145] A three-dimensional reconstruction module 150, configured to perform three-dimensional reconstruction processing on the dense complete scene point cloud data according to a grid three-dimensional reconstruction algorithm to obtain a complete scene occupancy grid;

[0146] A semantic assignment module 160, configured to obtain occupancy grid semantic information according to the dense complete scene point cloud data and the complete scene occupancy grid to obtain a semantic occupancy grid representation of the complete scene;

[0147] An occupancy grid dataset acquisition module 170, configured to form a dataset with all semantic occupancy grid representations of the complete scene to obtain an occupancy grid dataset in the mining area environment.

[0148] In summary, the device for constructing a mining area dataset based on an occupancy grid provided by the present invention collects environmental image information and environmental point cloud information of mining area roads and mining area working areas, and then constructs an occupancy grid dataset under a mining environment based on the environmental image information and environmental point cloud information, which can make up for the lack of occupancy grid datasets in current mining scenarios and promote the training and deployment of related algorithms; in addition, the embodiment of the present invention adopts an end-to-end paradigm to construct a full process from raw data collection to mining scene occupancy grid datasets, and opens up a complete chain of dataset construction steps; in view of the point cloud sparseness problem caused by the open mining scene, a deep completion method is proposed to enhance the data to ensure the accuracy of the constructed dataset; it can effectively promote the training and application of subsequent occupancy grid perception algorithms, so that the algorithms have more comprehensive and sophisticated environmental modeling capabilities, thereby further promoting the application of autonomous driving technology in mines, improving the perception accuracy of autonomous driving perception algorithms in mines, and thus improving the operating efficiency, safety and economic benefits of mines.

[0149] The specific working principle of the device for constructing a mining area dataset based on an occupancy grid provided by the present invention can be referred to the description of the method for constructing a mining area dataset based on an occupancy grid in the foregoing text, which will not be repeated here.

[0150] As another aspect of the present invention, a mining area dataset construction system 10 based on occupancy grid is provided, wherein: Figure 13 As shown, it includes: a data collection vehicle 200, a multimodal sensor 210 located on the data collection vehicle 200 and the mining area data set construction device 100 based on the occupancy grid as described above, wherein the multimodal sensor 210 is in communication connection with the mining area data set construction device 100 based on the occupancy grid.

[0151] The multimodal sensor 210 includes at least an image acquisition device and a radar device, and the image acquisition device and the radar device can collect environmental image information and environmental point cloud information when the data acquisition vehicle travels on the mine road and the working area;

[0152] The occupancy grid-based mining area dataset construction device 100 can construct an occupancy grid dataset in a mining area environment according to the environmental image information and the environmental point cloud information.

[0153] In the embodiment of the present invention, the structural diagram of the data collection vehicle 200 is as follows: Figure 2 As shown, the multimodal sensor is carried on the mining road and the working area to collect the surrounding environment image and point cloud. The multimodal sensor 210 in the embodiment of the present invention may specifically include an image acquisition device 211 and a radar device 212. The image acquisition device may specifically be a camera, and the radar device may specifically include a laser radar and a millimeter wave radar.

[0154] The mining area data set construction system based on occupancy grid provided by the present invention adopts the mining area data set construction device based on occupancy grid, collects environmental image information and environmental point cloud information of mining area roads and mining area operation areas, and then constructs the occupancy grid data set under the mining area environment based on the environmental image information and environmental point cloud information, which can make up for the lack of occupancy grid data set in the current mining area scene and promote the training and deployment of related algorithms; in addition, the embodiment of the present invention constructs the whole process from raw data collection to mining area scene occupancy grid data set in an end-to-end paradigm, and opens up a complete chain of data set construction steps; in view of the point cloud sparseness problem caused by the open mining area scene, a deep completion method is proposed to enhance the data to ensure the accuracy of the constructed data set; it can effectively promote the training and application of the subsequent occupancy grid perception algorithm, so that the algorithm has a more comprehensive and detailed environmental modeling capability, thereby further promoting the application of autonomous driving technology in mines, improving the perception accuracy of the autonomous driving perception algorithm in mines, and thus improving the operation efficiency, safety and economic benefits of mines.

[0155] The specific working principle of the mining area dataset construction system based on the occupancy grid provided by the present invention can be referred to the description of the mining area dataset construction method based on the occupancy grid in the previous text, which will not be repeated here.

[0156] As another aspect of the present invention, an automatic driving perception system 1 is provided, wherein: Figure 14 As shown, it includes: an autonomous driving algorithm model 20 and the mining area data set construction system 10 based on the occupancy grid as described above,

[0157] The mining area data set construction system 10 based on occupancy grid can construct an occupancy grid data set in a mining area environment;

[0158] The autonomous driving algorithm model 20 can train a target autonomous driving algorithm according to the grid data set occupied in the mining environment to obtain a target autonomous driving algorithm model.

[0159] It should be understood that the autonomous driving algorithm model can train the target autonomous driving algorithm according to the occupancy grid dataset constructed by the occupancy grid-based mining dataset construction system 10 in the above-mentioned mining environment to obtain the target autonomous driving algorithm model.

[0160] This autonomous driving perception system of the present invention constructs a system based on the occupancy grid-based mining area dataset described above. By collecting environmental image information and environmental point cloud information of the mining area road and the mining area operation area, and then constructing an occupancy grid dataset in the mining area environment based on this environmental image information and environmental point cloud information, it can make up for the lack of the occupancy grid dataset in the current mining area scenario, promote the training and deployment of related algorithms, effectively promote the training and application of subsequent occupancy grid perception algorithms, enable the algorithm to have a more comprehensive and refined environmental modeling ability, thereby further promoting the application of autonomous driving technology in mines, improving the perception accuracy of the autonomous driving perception algorithm in mines, and further enhancing the operation efficiency, safety and economic benefits of mines.

[0161] Regarding the specific principle of improving the autonomous driving perception accuracy provided by the autonomous driving perception system of the present invention, reference can be made to the description of the occupancy grid-based mining area dataset construction device described above, and details will not be repeated here.

[0162] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principle of the present invention. However, the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also regarded as the protection scope of the present invention.

Claims

1. A method for constructing a mining area dataset based on an occupancy grid, characterized in that: include: Obtain environmental image information and environmental point cloud information of mining roads and mining operation areas; Constructing a preliminary data set including a target object detection frame and semantic information according to the environmental image information and the environmental point cloud information; Constructing a depth completion model according to the preliminary data set, and performing initial densification processing on the environmental point cloud information according to the depth completion model to obtain initial dense point cloud data; Performing dynamic and static point cloud separation and superposition processing on the initial dense point cloud data to obtain dense complete scene point cloud data; Performing three-dimensional reconstruction processing on the dense complete scene point cloud data according to a grid three-dimensional reconstruction algorithm to obtain a complete scene occupancy grid; Acquiring occupancy grid semantic information according to the dense complete scene point cloud data and the complete scene occupancy grid, and obtaining a semantically represented occupancy grid representation of the complete scene; The semantically represented occupancy grids of all complete scenes are combined into a dataset to obtain an occupancy grid dataset in a mining environment. The initial dense point cloud data is subjected to dynamic and static point cloud separation and superposition processing to obtain dense complete scene point cloud data, including: Traversing the initial dense point cloud data of all frames within a preset time period; Separating the dynamic and static points of the initial dense point cloud data to obtain a dynamic point cloud as a dynamic object and a static point cloud as a static scene; Unifying the dynamic point cloud and the static point cloud into a world coordinate system, and respectively superimposing the dynamic point clouds of different frames to obtain a dense dynamic object point cloud and superimposing the static point clouds of different frames to obtain a dense static scene point cloud; Inserting the dense dynamic object point cloud into the dense static scene point cloud and restoring the complete scene to obtain dense complete scene point cloud data; Performing three-dimensional reconstruction processing on the dense complete scene point cloud data according to a grid three-dimensional reconstruction algorithm to obtain a complete scene occupancy grid includes: Calculate the regional surface normal according to the distribution of the dense complete scene point cloud data; Rendering the surface normals of the region to obtain a triangular facet mesh; Rasterizing the triangular face grid to obtain a complete scene occupancy grid; Acquiring occupancy grid semantic information according to the dense complete scene point cloud data and the complete scene occupancy grid to obtain a semantically represented occupancy grid representation of the complete scene, including: Find the grid closest to each point in the dense complete scene point cloud data according to the nearest neighbor algorithm; The semantic information is assigned to the grid with the closest distance to obtain the semantically occupied grid representation of the complete scene.

2. The method for constructing a mining area dataset based on an occupancy grid according to claim 1, characterized in that: A preliminary data set including a target object detection frame and semantic information is constructed based on the environmental image information and the environmental point cloud information, including: Preprocessing the environmental image information and the environmental point cloud information to remove abnormal data and obtain a preprocessed data set; The preprocessed data set is annotated with semantic information, and a preliminary data set including a target object detection frame and semantic information is constructed based on the annotated preprocessed data set.

3. The method for constructing a mining area dataset based on an occupancy grid according to claim 1, characterized in that: Constructing a depth completion model according to the preliminary data set, and performing initial densification processing on the environment point cloud information according to the depth completion model to obtain initial dense point cloud data, including: Training a neural network model according to the preliminary data set to obtain a depth completion model; The environmental point cloud information is input into the depth completion model for data enhancement to obtain initial dense point cloud data.

4. The method for constructing a mining area dataset based on an occupancy grid according to claim 3, characterized in that: The depth completion model includes a data pre-fusion layer, a pre-trained backbone network module, an encoding network module, a decoding network module and a dense point cloud output layer, which are arranged in sequence from input to output. The environmental point cloud information is input into the depth completion model for data enhancement to obtain initial dense point cloud data, including: Inputting the environmental point cloud information and its corresponding environmental image information into a data front fusion layer for feature fusion processing to obtain a first feature map including image and depth information; Inputting the first feature map into a pre-trained backbone network module for initialization processing to obtain a second feature map; Inputting the second feature map into the encoding network module for encoding processing to obtain a fifth feature map, wherein the compactness of the fifth feature map is greater than the compactness of the second feature map; Inputting the fifth feature map into the decoding network module for decoding processing to obtain a dense depth map; The dense depth map is input into the dense point cloud output layer for restoration processing to obtain initial dense point cloud data.

5. A mining area data set construction device based on an occupancy grid, used to implement the mining area data set construction method based on an occupancy grid as claimed in any one of claims 1 to 4, characterized in that: include: An acquisition module is used to acquire environmental image information and environmental point cloud information of mining roads and mining operation areas; A preliminary data set construction module, used to construct a preliminary data set including a target object detection frame and semantic information according to the environmental image information and the environmental point cloud information; An initial densification processing module, used to construct a depth completion model according to the preliminary data set, and perform initial densification processing on the environmental point cloud information according to the depth completion model to obtain initial dense point cloud data; A dynamic and static separation and superposition module is used to perform dynamic and static point cloud separation and superposition processing on the initial dense point cloud data to obtain dense complete scene point cloud data; A three-dimensional reconstruction module, used to perform three-dimensional reconstruction processing on the dense complete scene point cloud data according to a grid three-dimensional reconstruction algorithm to obtain a complete scene occupancy grid; A semantic assignment module, used for acquiring occupancy grid semantic information according to the dense complete scene point cloud data and the complete scene occupancy grid, to obtain a semantically represented occupancy grid representation of the complete scene; The module for obtaining the occupancy grid dataset is used to form a dataset of semantically represented occupancy grids of all complete scenes, and obtain the occupancy grid dataset in the mining environment.

6. A mining area data set construction system based on occupancy grid, characterized in that: include: A data collection vehicle, a multimodal sensor located on the data collection vehicle, and the mining area data set construction device based on the occupancy grid according to claim 5, wherein the multimodal sensor is communicatively connected to the mining area data set construction device based on the occupancy grid, The multimodal sensor includes at least an image acquisition device and a radar device, and the image acquisition device and the radar device are capable of acquiring environmental image information and environmental point cloud information when the data acquisition vehicle travels on mining roads and working areas; The occupancy grid-based mining area data set construction device can construct an occupancy grid data set in a mining area environment according to the environmental image information and the environmental point cloud information.

7. An autonomous driving perception system, characterized in that: include: The autonomous driving algorithm model and the mining area data set construction system based on the occupancy grid as described in claim 6, The mining area data set construction system based on occupancy grid can construct an occupancy grid data set in a mining area environment; The autonomous driving algorithm model can train a target autonomous driving algorithm according to the occupied grid data set in the mining environment to obtain a target autonomous driving algorithm model.

Citation Information

Patent Citations

  • Driving scene simulation method, system and equipment based on three-dimensional occupation grid and medium

    CN116452766A

  • Semantic scene completion method based on image and point cloud fusion in automatic driving scene

    CN116503825A