Mining area 3D occupation prediction method and device, vehicle, equipment, medium and chip
By integrating the 2D images and 3D point cloud features of the mining area, and introducing 2D semantics and deep supervision using 3D Gaussian splashing technology, the problem of low training accuracy of the 3D occupation prediction scheme in the mining area is solved, and higher prediction accuracy and robustness are achieved, and adapting to complex mining area environments.
Patent Information
- Application Number
- CN202510787770.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
In the prior art, the 3D occupancy prediction scheme in the mining area only uses the 3D occupancy truth value for supervision, resulting in low training accuracy and it is difficult to provide efficient environmental modeling and path planning in complex and changing mining area environments.
Combining the 2D image and 3D point cloud features in the mining area, through feature fusion and 3D Gaussian splashing technology, 2D semantics and deep supervision are introduced to improve the model's perception of sparse targets and realize multimodal and multidimensional signal supervision.
It enhances the adaptability of the mining area's autonomous driving system in complex environments, improves the accuracy and robustness of 3D occupancy prediction, ensures driving safety and improves operating efficiency.
Smart Images

Figure CN120299008A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving and modeling in mining areas, and particularly to a 3D occupancy prediction method, device, computer device, readable storage medium, and chip for mining areas. Background Art
[0002] The mining area environment is more complex and changeable compared to ordinary urban roads and highways, such as large terrain undulations, dusty, complex lighting conditions, and frequent appearance of dynamic obstacles. Therefore, ensuring the safety and operation efficiency of the autonomous driving system in the mining area environment is the core issue of current research.
[0003] As an important environmental modeling method, the 3D Occupancy Grid can accurately model the surrounding environment in a rasterized manner, providing the occupancy status and probability information of each grid, thereby providing reliable data support for path planning and obstacle avoidance decision-making. In the mining area environment, the 3D Occupancy Grid technology can more effectively model irregular terrains, dynamic obstacles (such as mining trucks, pedestrians), and unpredictable environmental changes (such as collapses, falling rocks), improving the adaptability of the autonomous driving system.
[0004] However, in the 3D occupancy prediction scheme in the related art, only 3D occupancy ground truth is used for supervision, and the 3D occupancy ground truth data is often sparse in the mining area scenario, so it cannot fully guide the training of the model. Summary of the Invention
[0005] In view of this, this application provides a 3D occupancy prediction method, device, vehicle, computer device, readable storage medium, and chip for mining areas, which solves the problem of low training accuracy caused by only using 3D occupancy ground truth for supervision in the 3D occupancy prediction scheme in the related art. In a first aspect, an embodiment of this application provides a 3D occupancy prediction method for mining areas, including: Obtaining 3D image features based on 2D images of the mining area, and obtaining 3D point cloud features based on 3D point clouds of the mining area; Performing feature fusion on the 3D image features and the 3D point cloud features to obtain fused features; Performing occupancy prediction on the fused features through a 3D occupancy prediction head, and using 3D occupancy ground truth for supervision to obtain 3D occupancy distribution prediction values; Converting the 3D occupancy distribution prediction values into 2D semantic segmentation prediction values and depth prediction values through 3D Gaussian splashing technology, and using 2D semantic segmentation ground truth to supervise the 2D semantic segmentation prediction values, and using depth ground truth to supervise the depth prediction values.
[0006] Second aspect, an embodiment of the present application provides a 3D occupancy prediction device for a mining area, including: An image feature acquisition module, configured to obtain 3D image features based on 2D images of the mining area; A point cloud feature acquisition module, configured to obtain 3D point cloud features based on 3D point clouds of the mining area; A feature fusion module, configured to perform feature fusion on the 3D image features and the 3D point cloud features to obtain fused features; A prediction module, configured to perform occupancy prediction on the fused features through a 3D occupancy prediction head and use 3D occupancy ground truth for supervision to obtain 3D occupancy distribution prediction values; A first supervision module, configured to convert the 3D occupancy distribution prediction values into 2D semantic segmentation prediction values and depth prediction values through 3D Gaussian splashing technology, and use 2D semantic segmentation ground truth to supervise the 2D semantic segmentation prediction values, and use depth ground truth to supervise the depth prediction values.
[0007] Third aspect, an embodiment of the present application provides a vehicle, and the vehicle includes the device as in the second aspect.
[0008] Fourth aspect, an embodiment of the present application provides a computer device, including a first processor and a first memory. The first memory stores a program or instruction that runs on the first processor. When the program or instruction is executed by the first processor, the steps of the method as in the first aspect are implemented.
[0009] Fifth aspect, an embodiment of the present application provides a computer-readable storage medium, and a program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps of the method as in the first aspect are implemented.
[0010] Sixth aspect, an embodiment of the present application provides a chip, which includes at least one second processor and a communication interface. The communication interface is coupled to the at least one second processor, and the at least one second processor is configured to run a program or instruction to implement the method as in the first aspect.
[0011] The 3D occupancy prediction method, device, vehicle, computer device, readable storage medium, and chip provided by the embodiments of the present application obtain 3D image features based on 2D images of the mining area, and obtain 3D point cloud features based on 3D point clouds of the mining area. The 3D image features and the 3D point cloud features are fused to obtain fused features. Occupancy prediction processing is performed on the fused features, and 3D occupancy ground truth is used to supervise the output results to obtain 3D occupancy distribution prediction values. The 3D occupancy distribution prediction values are converted into 2D semantic segmentation prediction values and depth prediction values through 3D Gaussian splashing technology, and 2D semantic segmentation ground truth is used to supervise the 2D semantic segmentation prediction values, and depth ground truth is used to supervise the depth prediction values.
[0012] In the embodiments of the present application, 2D semantics and depth supervision are introduced through 3D Gaussian splash on the basis of 3D occupancy ground truth supervision, realizing multi-modal and multi-dimensional signal supervision for 3D occupancy prediction. Based on the distribution advantage of sparse targets in 2D semantic segmentation, the perception ability of the model for sparse targets is enhanced. At the same time, since the 2D semantic segmentation signal is denser than the 3D occupancy ground truth signal, it can be used as an effective supplement to the 3D occupancy ground truth to improve the model prediction effect.
[0013] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 A flowchart showing the process of the 3D occupancy prediction method for a mining area in an embodiment of the present application; Figure 2 A logic diagram showing the training stage of the 3D occupancy model for a mining area in an embodiment of the present application; Figure 3 A logic diagram showing the inference stage of the 3D occupancy for a mining area in an embodiment of the present application; Figure 4 A structural block diagram showing the 3D occupancy prediction device for a mining area in an embodiment of the present application; Figure 5 A structural schematic diagram showing a vehicle in an embodiment of the present application; Figure 6 A structural schematic diagram showing a computer device in an embodiment of the present application; Figure 7 A structural schematic diagram showing a computer-readable storage medium in an embodiment of the present application; Figure 8 A structural schematic diagram showing a chip in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.
[0016] In the description and claims of this application, terms such as "first" and "second" are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally means that the related objects before and after are in an "or" relationship.
[0017] Autopilot in mining areas relies on high-precision 3D occupancy prediction to provide fine-grained environmental perception to ensure driving safety and improve operation efficiency. However, simply relying on image input in an unstructured and complex terrain environment like a mining area is likely to lead to incorrect occupancy prediction, affecting the stability and reliability of the system. In addition, due to the sparse and uneven distribution of occupancy data in mining area scenes, the methods in related technologies are difficult to provide a complete and accurate environmental model. The 3D occupancy prediction method, device, computer device, readable storage medium, and chip provided in the embodiments of this application fuse multi-modal inputs, can make up for the limitations of a single sensor, and combine multi-dimensional supervision signals. Since the semantic proportion of sparse or small targets in the 2D space is larger than that in the 3D space, introducing the supervision of 2D semantic signals can alleviate the sparsity of occupancy data of sparse or small targets, effectively improving the accuracy and robustness of 3D occupancy prediction, and enabling the autopilot system to have stronger adaptability in the complex environment of mining areas. The following will combine the accompanying drawings to detail the 3D occupancy prediction method, device, vehicle, computer device, readable storage medium, and chip provided in the embodiments of this application through specific embodiments and their application scenarios. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0018] The embodiments of this application provide a 3D occupancy prediction method for mining areas, as Figure 1 and Figure 2 shown. The method includes: Step 101, obtaining 3D image features based on a mining area 2D image, and obtaining 3D point cloud features based on a mining area 3D point cloud.
[0019] In this step, a mining area 2D image is obtained through an image acquisition device (such as a monocular camera) of a vehicle, and the mining area 2D image is converted into 3D image features. In addition, a mining area 3D point cloud is collected by using the lidar of the vehicle, and 3D point cloud features are obtained based on the mining area 3D point cloud.
[0020] The vehicles include pickup trucks, micro trucks, light trucks, minivans, dump trucks, cargo trucks, tractors, trailers, special vehicles, mining trucks, wide-body trucks, articulated trucks, excavators, electric shovels, bulldozers, etc. within the mining area.
[0021] In one embodiment of the present application, obtaining 3D image features based on the 2D image of the mining area includes: performing 2D feature extraction on the 2D image of the mining area through a 2D image feature extraction network to obtain 2D image features; predicting the depth distribution vector of the 2D image of the mining area through a depth prediction network, and using the depth distribution vector to convert the 2D image features into 3D image features.
[0022] In this embodiment, the 2D image feature extraction network includes ResNet (Residual Network) and FPN (Feature Pyramid Network). ResNet can be ResNet50, ResNet-101, ResNet-152, etc. The 2D image of the mining area is sequentially subjected to feature extraction through ResNet and FPN to obtain 2D image features. The obtained 2D image features are input into the depth prediction network to predict the depth distribution vector for each pixel, and the depth ground truth is used to supervise the depth distribution vector to ensure the accuracy of the predicted depth distribution vector. Then, using the predicted depth distribution vector, based on the LSS (Lift-Splat-Shoot) method, the 2D image features are converted into 3D image features. Among them, LSS (Lift-Splat-Shoot) is a method for 3D object detection, especially suitable for scenarios such as autonomous driving that require perceiving 3D spatial information from 2D images. Specifically, based on the predicted depth distribution vector, the 2D image features are mapped into voxels or pillars in the three-dimensional space, and the features within each voxel or pillar in the three-dimensional space are aggregated and fused to obtain 3D image features.
[0023] In one embodiment of the present application, obtaining 3D point cloud features based on the 3D point cloud of the mining area includes: grouping the 3D point cloud of the mining area according to voxels of a preset size to obtain a plurality of voxel groups; performing coordinate mean operation on the 3D point cloud of the mining area included in each voxel group to obtain the point cloud data corresponding to the voxel group; Performing sparse convolution on the point cloud data of multiple voxel groups to obtain 3D point cloud features.
[0024] In this embodiment, the 3D point cloud of the mining area is grouped according to voxels of a preset size. Within each voxel group, coordinate mean operation is performed on the included point cloud to obtain the point cloud data corresponding to the voxel group, and then the point cloud data is subjected to sparse convolution to obtain 3D point cloud features.
[0025] In the embodiment of the present application, the input is multimodal information including images and point clouds, which can make up for the limitations of a single sensor and improve the accuracy of 3D occupancy prediction.
[0026] Step 102: Perform feature fusion on the 3D image features and 3D point cloud features to obtain fused features.
[0027] In this step, the 3D image features and 3D point cloud features are fused through a feature fusion network. For example, the 3D image features and 3D point cloud features can be concatenated in the feature channels and then reduced to the dimension before fusion through an MLP (Multilayer Perceptron) network.
[0028] Step 103: Perform occupancy prediction on the fused features through a 3D occupancy prediction head and use the 3D occupancy ground truth for supervision to obtain the 3D occupancy distribution prediction value.
[0029] In this step, occupancy prediction processing is performed on the fused features, and the output result is supervised using the 3D occupancy ground truth to obtain the 3D occupancy distribution prediction value.
[0030] In an embodiment of the present application, performing occupancy prediction on the fused features through a 3D occupancy prediction head and using the 3D occupancy ground truth for supervision to obtain the 3D occupancy distribution prediction value includes: Enhance the fused features through a 3D feature pyramid network; Input the enhanced fused features into the 3D occupancy prediction head to output the initial 3D occupancy distribution prediction value; And use the 3D occupancy ground truth to supervise the initial 3D occupancy distribution prediction value to obtain the 3D occupancy distribution prediction value.
[0031] In this embodiment, after the fused features are enhanced through a 3D feature pyramid network, the initial 3D occupancy distribution prediction value is output through the 3D occupancy prediction head. The initial 3D occupancy distribution prediction value is supervised based on the 3D occupancy ground truth, and the 3D occupancy ground truth is pre-obtained training data. Specifically, the loss is calculated according to the 3D occupancy ground truth and the initial 3D occupancy distribution prediction value to verify the prediction effect.
[0032] In the embodiment of the present application, performing 3D occupancy prediction after enhancing the fused features can improve the accuracy of 3D occupancy prediction.
[0033] Step 104: Convert the 3D occupancy distribution prediction value into a 2D semantic segmentation prediction value and a depth prediction value through 3D Gaussian splashing technology, and use the 2D semantic segmentation ground truth to supervise the 2D semantic segmentation prediction value and use the depth ground truth to supervise the depth prediction value.
[0034] While this application uses 3D occupancy ground truth for 3D occupancy prediction supervision, it also adds 2D semantic and depth supervision through 3D Gaussian splatting, achieving multi-dimensional signal supervision for 3D occupancy prediction.
[0035] In this step, after obtaining the 3D occupancy distribution prediction value through 3D occupancy prediction, the 3D occupancy distribution prediction value is converted into a 2D semantic segmentation prediction value and a depth prediction value through the 3D Gaussian splatting technique. Among them, 3D Gaussian Splatting is a technique that diffuses discrete 3D data into a continuous space through a Gaussian kernel function. Its core idea is to project the features of each 3D point (such as position, color, depth, semantics, etc.) onto the 2D image plane in the form of a Gaussian distribution to generate a smooth and continuous distribution.
[0036] In the embodiment of this application, the 3D point cloud feature of each 3D occupancy distribution prediction value corresponds to a voxel in space. The center of this voxel is used as the Gaussian kernel center of the Gaussian splatting. The rotation matrix and the scale matrix are both set to fixed values. For example, the rotation matrix can be set to 0, that is, no rotation, the scale matrix is set to a fixed value and no stretching, and the opacity is calculated by the prediction value of the empty class according to the following formula:
[0037] Among them, is the opacity required for 3D Gaussian rendering, is the prediction value of the empty class.
[0038] Through the rotation matrix and the scale matrix, geometric transformation of 3D data is achieved to ensure that the projection matches the observation perspective and scale. By controlling the superposition effect of the projection through the opacity, the realism and readability of the scene are improved.
[0039] Furthermore, the 2D semantic segmentation prediction value and the depth prediction value are supervised using the 2D semantic segmentation ground truth and the depth ground truth respectively. The 2D semantic segmentation ground truth and the depth ground truth are pre-acquired training data. Among them, the depth ground truth can be obtained by projecting the stitching of single-frame or multi-frame mining area 3D point clouds, and the 2D semantic segmentation ground truth can be obtained from the input image through a large semantic segmentation model. For example, the loss is calculated through the 2D semantic segmentation ground truth and the 2D semantic segmentation prediction value, and the loss is calculated through the depth ground truth and the depth prediction value to verify the 2D semantic and depth prediction effects.
[0040] The embodiment of the present application provides a 3D occupancy prediction method for mining areas based on 3D Gaussian splash, multi-modal fusion, and multi-dimensional signal supervision. This method introduces 2D semantic and depth supervision through 3D Gaussian splash on the basis of 3D occupancy ground truth supervision, realizing multi-modal and multi-dimensional signal supervision for 3D occupancy prediction. Based on the distribution advantage of sparse targets in 2D semantic segmentation, the perception ability of the model for sparse targets is enhanced. At the same time, since the 2D semantic segmentation signal is denser than the 3D occupancy ground truth signal, it can be used as an effective supplement to the 3D occupancy ground truth to improve the model prediction effect.
[0041] When using a 3D occupancy grid to perform 3D occupancy prediction on the mining area environment and then realizing the modeling of the mining area environment, the inference process is as Figure 3 shown. 3D image features are obtained based on the image, 3D point cloud features are obtained based on the point cloud, and then the 3D image features and 3D point cloud features are fused to obtain 3D fusion features. After performing feature enhancement processing on the 3D fusion features, 3D occupancy prediction is carried out.
[0042] Since the prediction target of 3D occupancy is limited to a local cube near the vehicle body, which is a small area relative to the visible mine scene, the projection of the 3D occupancy ground truth can only reflect the occupancy projection of the elements within this cube. While the 2D semantic segmentation directly corresponding to the input image corresponds to the projection of all occupied voxels within the camera frustum, which means that the projections of those voxels far from the vehicle body are also included, creating a difference from the projection of the 3D occupancy ground truth. Therefore, the 2D semantic segmentation ground truth obtained from the training data is not suitable for directly supervising the 2D semantic segmentation prediction value. Considering this problem, it is necessary to first calculate a mask for the 2D segmentation result according to the distance between the object and the image acquisition device. In one embodiment of the present application, the method further includes: generating a mask for the near-view area involved in 3D occupancy based on the depth ground truth obtained from the 3D point cloud of the mining area, and obtaining the masked 2D semantic segmentation ground truth through this mask; Supervising the 2D semantic segmentation prediction value using the 2D semantic segmentation ground truth includes: supervising the 2D semantic segmentation prediction value using the masked 2D semantic segmentation ground truth.
[0043] In this embodiment, the depth ground truth obtained from the 3D point cloud of the mining area, that is, the sparse depth map, extracts the near-view area from the sparse depth map, generates a mask for the near-view area, and then obtains the 2D semantic segmentation ground truth aligned with this mask, excluding the far-view pixels, to obtain the masked 2D semantic segmentation ground truth, reducing the difference between the projection of the 2D semantic segmentation and the 3D occupancy ground truth, avoiding misleading the prediction by objects at a relatively long distance, and then supervising the 2D semantic segmentation prediction value using the masked 2D semantic segmentation ground truth, thereby improving the adaptability of the data and the reliability of supervision.
[0044] In one embodiment of the present application, masking the regional depth map to obtain the masked 2D semantic segmentation ground truth includes: Removing target points that do not conform to the preset 3D occupancy grid depth interval from the regional depth map; After removing the target points, obtaining the connected regions in the regional depth map and calculating the connected area of the connected regions; If the connected area of the connected region is greater than the preset threshold, the connected region is retained and the mask value of the connected region is set to "1"; if the connected area of the connected region is less than or equal to the preset threshold, the mask value of the connected region is set to "0"; Performing morphological dilation on the connected regions with a mask value of "1" to fill the gaps and obtaining the masked 2D semantic segmentation ground truth.
[0045] In this embodiment, the regional depth map obtained by lidar projection is often sparse and cannot be directly used as a mask. First, a mask is calculated for the 2D segmentation result based on the distance. First, target points that do not conform to the preset 3D occupancy grid depth interval are removed from the regional depth map. These target points not in the preset 3D occupancy grid depth interval, that is, points without depth information, belong to points at a relatively far distance. For example, points with a distance greater than 80 meters are excluded. The 3D occupancy grid depth interval is preset and related to the training dataset. Different training datasets can correspond to different 3D occupancy grid depth intervals.
[0046] Furthermore, all connected regions in the regional depth map are obtained, realizing connecting the sparse points into regions and calculating the connected area. A threshold is set. When the connected area of a certain connected region is greater than the threshold, it is retained and its mask value is set to "1", that is, True. When the connected area of a certain connected region is less than or equal to the threshold, it is not retained and its mask value is set to "0". Then, morphological dilation is performed on the retained connected regions to fill the small gaps. Finally, the masked depth-based 2D semantic segmentation ground truth is obtained. The masked depth-based 2D semantic segmentation ground truth has the same size as the image. Furthermore, the masked 2D semantic segmentation ground truth is used to supervise the 2D semantic segmentation prediction value. For the pixels with a mask value of "1" in the masked 2D semantic segmentation ground truth, a loss calculation is performed with the 2D semantic segmentation prediction value. For the pixels with a mask value of "0", no loss calculation is performed with the 2D semantic segmentation prediction value, improving the supervision reliability.
[0047] In the image branch, when converting 2D image features to 3D image features, depth information obtained from lidar input is considered for feature enhancement. However, in related technologies, during the process of feature extraction in the lidar branch, there is no participation of semantic information, resulting in the features extracted by the lidar branch lacking the enhancement of semantic information. In an embodiment of the present application, before fusing the 3D image features and 3D point cloud features to obtain the fused features, the method further includes: Converting the 3D point cloud features into 2D semantic segmentation prediction values through 3D Gaussian splash technology, and supervising the 2D semantic segmentation prediction values using the masked 2D semantic segmentation ground truth.
[0048] In this embodiment, before fusing the 3D image features and 3D point cloud features, the 3D point cloud features are converted into 2D semantic segmentation prediction values through 3D Gaussian splash technology, and then the 2D semantic segmentation prediction values are supervised using the 2D semantic segmentation ground truth. In particular, the masked 2D semantic segmentation ground truth can be used to supervise the 2D semantic segmentation prediction values.
[0049] The feature fusion scheme in related technologies does not provide semantic information guidance for the lidar branch before fusion. In the embodiment of the present application, in the lidar branch before feature fusion, rendering 2D semantic segmentation from point cloud features and supervising with the masked 2D ground truth can improve the semantic consistency of lidar features, align the lidar features with image features, enhance the feature quality before fusion, and make it better adapt to multi-modal perception tasks.
[0050] In an embodiment of the present application, converting the 3D point cloud features into 2D semantic segmentation prediction values through 3D Gaussian splash technology includes: Extracting semantic information from the 3D point cloud features through a 3D convolutional network to obtain 3D semantic distribution prediction values; Projecting the 3D semantic distribution prediction values into the 2D space through 3D Gaussian splash technology to obtain 2D semantic segmentation prediction values.
[0051] In this embodiment, the structural information in the 3D point cloud features can be used to generate semantic predictions. First, a two-layer 3×3 3D convolutional network is used to extract semantic information from the 3D point cloud features to obtain the 3D semantic distribution prediction values based on the sparse radar features. Then, through the 3D Gaussian splash rendering method, the 3D semantic distribution prediction values are projected into the 2D space to obtain the 2D semantic segmentation prediction values. Then, the masked 2D semantic segmentation ground truth is used for supervision to help the radar feature branch obtain semantic information. Among them, the 3D point cloud features of each 3D occupancy distribution prediction value correspond to a voxel in space. The center of the voxel is used as the center of the Gaussian kernel of the Gaussian splash. The rotation matrix and the scale matrix are both set to fixed values. For example, the rotation matrix can be set to 0, that is, no rotation, the scale matrix is set to a fixed value and not stretched, and the opacity is calculated from the prediction value of the empty class according to the following formula:
[0052] Wherein, is the opacity required for 3D Gaussian rendering, is the prediction value of the empty class.
[0053] In the related art, the fusion effect is not supervised when performing multi-modal feature fusion. Therefore, the fusion effect cannot be effectively guaranteed. In one embodiment of the present application, before the occupancy prediction of the fusion features is performed by the 3D occupancy prediction head, the method further includes: Performing dimensionality reduction on the fusion features and outputting RGB prediction values and opacity prediction values; Converting the RGB prediction values and the opacity prediction values into 2D RGB prediction values through the 3D Gaussian splash technology, and using the RGB ground truth to supervise the 2D RGB prediction values.
[0054] In this embodiment, the fused 3D fusion features should have both semantic context information and structural information at the same time. The fusion features are dimensionally reduced through an MLP network or a 1×1 3D convolutional network, and RGB prediction values and opacity prediction values are output. Then, the RGB prediction values and the opacity prediction values are converted into 2D RGB prediction values through 3D Gaussian splash, and the RGB ground truth is used to supervise the 2D RGB prediction values to ensure the fusion effect of the fusion features. It should be noted that the RGB ground truth here can also be the masked RGB ground truth.
[0055] In the present application, after the fusion features, 2D RGB prediction values are output through 3D Gaussian splash and supervised, enhancing the geometric consistency and semantic consistency of the fusion features, and helping the model learn more consistent cross-modal representations.
[0056] As a specific implementation of the above 3D occupancy prediction method for mining areas, an embodiment of the present application provides a 3D occupancy prediction device for mining areas. AsFigure 4 As shown in Figure 4 , the 3D occupancy prediction device 400 for the mining area includes: an image feature acquisition module 401, a point cloud feature acquisition module 402, a feature fusion module 403, a prediction module 404, and a first supervision module 405.
[0057] Among them, the image feature acquisition module 401 is used to obtain 3D image features based on the 2D image of the mining area; The point cloud feature acquisition module 402 is used to obtain 3D point cloud features based on the 3D point cloud of the mining area; The feature fusion module 403 is used to perform feature fusion on the 3D image features and the 3D point cloud features to obtain fused features; The prediction module 404 is used to perform occupancy prediction on the fused features through a 3D occupancy prediction head and use 3D occupancy ground truth for supervision to obtain a 3D occupancy distribution prediction value; The first supervision module 405 is used to convert the 3D occupancy distribution prediction value into a 2D semantic segmentation prediction value and a depth prediction value through 3D Gaussian splashing technology, and use the 2D semantic segmentation ground truth to supervise the 2D semantic segmentation prediction value, and use the depth ground truth to supervise the depth prediction value.
[0058] Furthermore, the image feature acquisition module 401 is specifically used for: Performing 2D feature extraction on the 2D image of the mining area through a 2D image feature extraction network to obtain 2D image features; Predicting the depth distribution vector of the 2D image of the mining area through a depth prediction network, and using the depth distribution vector to convert the 2D image features into 3D image features.
[0059] Furthermore, the point cloud feature acquisition module 402 is specifically used for: Grouping the 3D point cloud of the mining area according to a preset voxel size to obtain a plurality of voxel groups; Performing coordinate mean operation on the 3D point cloud of the mining area included in each voxel group to obtain point cloud data corresponding to the voxel group; Performing sparse convolution on the point cloud data of multiple voxel groups to obtain 3D point cloud features.
[0060] Furthermore, the prediction module 404 is specifically used for: Performing enhancement processing on the fused features through a 3D feature pyramid network; Inputting the enhanced fused features into the 3D occupancy prediction head to output an initial 3D occupancy distribution prediction value; And using the 3D occupancy ground truth to supervise the initial 3D occupancy distribution prediction value to obtain a 3D occupancy distribution prediction value.
[0061] Further, the device further includes: a mask processing module, configured to: generate a mask for the near-view region involved in 3D occupancy based on the depth ground truth obtained from the mining area 3D point cloud, and obtain the masked 2D semantic segmentation ground truth through this mask; The first supervision module 405 is specifically configured to: supervise the 2D semantic segmentation prediction value using the masked 2D semantic segmentation ground truth.
[0062] Further, the mask processing module is specifically configured to: Remove target points that do not conform to the preset 3D occupancy grid depth interval from the regional depth map; After removing the target points, obtain the connected regions in the regional depth map and calculate the connected area of the connected regions; If the connected area of the connected region is greater than the preset threshold, retain the connected region and set the mask value of the connected region to "1"; if the connected area of the connected region is less than or equal to the preset threshold, set the mask value of the connected region to "0"; Perform morphological dilation processing on the connected regions with a mask value of "1" to fill the gaps and obtain the masked 2D semantic segmentation ground truth.
[0063] Further, the device further includes: a second supervision module, configured to: Before fusing the 3D image features and the 3D point cloud features to obtain the fused features, convert the 3D point cloud features into 2D semantic segmentation prediction values through the 3D Gaussian splash technique, and supervise the 2D semantic segmentation prediction values using the masked 2D semantic segmentation ground truth.
[0064] Further, the second supervision module is specifically configured to: Extract semantic information from the 3D point cloud features through a 3D convolutional network to obtain a 3D semantic distribution prediction value; Project the 3D semantic distribution prediction value into the 2D space through the 3D Gaussian splash technique to obtain a 2D semantic segmentation prediction value.
[0065] Further, the device further includes: a third supervision module, configured to: Before performing occupancy prediction on the fused features through the 3D occupancy prediction head, perform dimensionality reduction processing on the fused features and output an RGB prediction value and an opacity prediction value; Convert the RGB prediction value and the opacity prediction value into 2D RGB prediction values through the 3D Gaussian splash technique, and supervise the 2D RGB prediction values using the RGB ground truth.
[0066] The mining area 3D occupancy prediction device 400 in the embodiments of the present application can be a computer device or a component in a computer device, such as an integrated circuit or a chip. The mining area 3D occupancy prediction device 400 provided in the embodiments of the present application can implement Figure 1 each process implemented by the mining area 3D occupancy prediction method embodiment. To avoid repetition, it will not be elaborated here.
[0067] The embodiments of the present application also provide a vehicle, as Figure 5 shown. The vehicle 500 includes the above-mentioned mining area 3D occupancy prediction device 400.
[0068] The above vehicle 500 can execute the mining area 3D occupancy prediction method described in the above embodiments through the mining area 3D occupancy prediction device 400. It can be understood that the implementation manner of the vehicle 500 controlling the mining area 3D occupancy prediction device 400 can be set according to the actual application scenario, and the embodiments of the present application do not make specific limitations.
[0069] The vehicle 500 is provided with a terminal, and the terminal includes but is not limited to: in-vehicle terminal, in-vehicle controller, in-vehicle module, in-vehicle component, in-vehicle chip, in-vehicle unit, in-vehicle radar, or other sensors such as an in-vehicle camera. The vehicle can implement the method provided in the present application through the in-vehicle terminal, in-vehicle controller, in-vehicle module, in-vehicle component, in-vehicle chip, in-vehicle unit, in-vehicle radar, or camera. The vehicles in the present application include passenger vehicles and commercial vehicles. Common models of commercial vehicles include but are not limited to: pickup trucks, micro trucks, light trucks, minivans, dump trucks, trucks, tractors, trailers, special vehicles, and mining vehicles, etc. Mining vehicles include but are not limited to mining trucks, wide-body vehicles, articulated vehicles, excavators, electric shovels, bulldozers, etc. The present application does not further limit the type of intelligent vehicles, and any type of vehicle is within the protection scope of the present application.
[0070] The embodiments of the present application also provide a computer device, as Figure 6 shown. The computer device 600 includes a first processor 601 and a first memory 602. A program or instruction that can run on the first processor 601 is stored on the first memory 602. When the program or instruction is executed by the first processor 601, it implements each step of the mining area 3D occupancy prediction method embodiment described above and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0071] The first memory 602 can be used to store software programs and various data. The first memory 602 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the first memory 602 may include a volatile memory or a non-volatile memory, or the first memory 602 may include both a volatile and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The first memory 602 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.
[0072] The first processor 601 may include one or more processing units; optionally, the first processor 601 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the first processor 601 either.
[0073] The embodiments of the present application also provide a readable storage medium, such as Figure 7 shown, a program or instruction 701 is stored on the readable storage medium 700. When the program or instruction 701 is executed by a processor, it implements each process of the above-mentioned embodiment of the 3D occupancy prediction method for the mining area, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0074] The method described in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. The computer-readable storage medium 700 may include computer storage media and communication media, and may also include any medium that can transfer a computer program from one place to another. The storage medium may be any target medium accessible by a computer.
[0075] As a possible design, the computer-readable storage medium 700 may include a Compact Disc Read-Only Memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; the computer-readable storage medium may include a magnetic disk storage or other magnetic disk storage devices. Moreover, any connecting line may also be appropriately referred to as a computer-readable storage medium. For example, if software is transmitted using coaxial cable, fiber optic cable, twisted pair, DSL (Digital Subscriber Line), or wireless technologies such as infrared, radio, and microwave from a website, server, or other remote source, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, magnetic disks and optical discs include optical discs (CDs), laser discs, optical discs, Digital Versatile Discs (DVDs), floppy disks, and Blu-ray discs, where magnetic disks typically reproduce data magnetically, while optical discs use lasers to optically reproduce data.
[0076] The embodiment of the present application also provides a chip, as Figure 8 shown. The chip 800 includes at least one processor (such as the second processor 801) and a communication interface 802. The communication interface 802 is coupled to the second processor 801. The second processor 801 is used to run programs or instructions to implement each process of the above embodiment of the 3D occupancy prediction method for the mining area, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0077] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip.
[0078] Preferably, the chip 800 further includes a memory, such as the second memory 803. The second memory 803 stores the following elements: executable modules or data structures, or subsets thereof, or extended sets thereof.
[0079] In the embodiments of the present application, the second memory 803 may include a read-only memory and a random access memory, and provide instructions and data to the second processor 801. A part of the second memory 803 may further include a non-volatile random access memory (NVRAM).
[0080] In the embodiments of the present application, the second processor 801, the communication interface 802, and the second memory 803 are coupled together through a bus system 804. Among them, in addition to the data bus, the bus system 804 may further include a power bus, a control bus, a status signal bus, etc. For the sake of description, in Figure 8 all kinds of buses are labeled as the bus system 804.
[0081] The above-described 3D occupancy prediction method for the mining area in the embodiments of the present application can be applied to or implemented by the second processor 801. The second processor 801 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the second processor 801 or the instructions in the form of software. The above-mentioned second processor 801 may be a general-purpose processor (e.g., a microprocessor or a conventional processor), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. The second processor 801 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention.
[0082] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising such element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0083] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. A 3D occupancy prediction method for a mining area, characterized in that, Including: Obtaining 3D image features based on the 2D image of the mining area, and obtaining 3D point cloud features based on the 3D point cloud of the mining area; Performing feature fusion on the 3D image features and the 3D point cloud features to obtain fused features; Performing occupancy prediction on the fused features through a 3D occupancy prediction head and using 3D occupancy ground truth for supervision to obtain 3D occupancy distribution prediction values; Converting the 3D occupancy distribution prediction values into 2D semantic segmentation prediction values and depth prediction values through 3D Gaussian splashing technology, and using 2D semantic segmentation ground truth to supervise the 2D semantic segmentation prediction values, and using depth ground truth to supervise the depth prediction values.
2. The method according to claim 1, characterized in that, The obtaining 3D image features based on the 2D image of the mining area includes: Performing 2D feature extraction on the 2D image of the mining area through a 2D image feature extraction network to obtain 2D image features; Predicting the depth distribution vector of the 2D image of the mining area through a depth prediction network, and using the depth distribution vector to convert the 2D image features into 3D image features.
3. The method according to claim 1, characterized in that, The obtaining 3D point cloud features based on the 3D point cloud of the mining area includes: Grouping the 3D point cloud of the mining area according to preset size voxels to obtain a plurality of voxel groups; Performing coordinate mean operation on the 3D point cloud of the mining area included in each voxel group to obtain the point cloud data corresponding to the voxel group; Performing sparse convolution on the point cloud data of multiple voxel groups to obtain the 3D point cloud features.
4. The method according to claim 1, characterized in that The performing occupancy prediction on the fused features through a 3D occupancy prediction head and using 3D occupancy ground truth for supervision to obtain 3D occupancy distribution prediction values includes: Performing enhancement processing on the fused features through a 3D feature pyramid network; Inputting the enhanced fused features into the 3D occupancy prediction head to output initial 3D occupancy distribution prediction values; And using 3D occupancy ground truth to supervise the initial 3D occupancy distribution prediction values to obtain 3D occupancy distribution prediction values.
5. The method according to claim 1, characterized in that, The method further includes: Generating a mask for the near-view region involved in 3D occupancy based on the depth ground truth obtained from the 3D point cloud of the mining area, and obtaining the masked 2D semantic segmentation ground truth through this mask; The using 2D semantic segmentation ground truth to supervise the 2D semantic segmentation prediction values includes: Using the masked 2D semantic segmentation ground truth to supervise the 2D semantic segmentation prediction values.
6. The method according to claim 5, characterized in that, The performing masking processing on the regional depth map to obtain the masked 2D semantic segmentation ground truth includes: Removing target points that do not conform to the preset 3D occupancy grid depth interval from the regional depth map; After removing the target points, obtaining the connected regions in the regional depth map and calculating the connected area of the connected regions; If the connected area of the connected region is greater than a preset threshold, retaining the connected region and setting the mask value of the connected region to "1"; if the connected area of the connected region is less than or equal to the preset threshold, setting the mask value of the connected region to "0"; Performing morphological dilation processing on the connected regions with a mask value of "1" to fill in the gaps to obtain the masked 2D semantic segmentation ground truth.
7. The method according to claim 5, characterized in that, Before the feature fusion of the 3D image feature and the 3D point cloud feature to obtain a fused feature, the method further includes: Converting the 3D point cloud feature into a 2D semantic segmentation prediction value by 3D Gaussian splashing technology, and supervising the 2D semantic segmentation prediction value with a masked 2D semantic segmentation ground truth.
8. The method according to claim 7, characterized in that, The converting the 3D point cloud feature into a 2D semantic segmentation prediction value by 3D Gaussian splashing technology includes: Extracting semantic information from the 3D point cloud feature through a 3D convolutional network to obtain a 3D semantic distribution prediction value; Projecting the 3D semantic distribution prediction value into a 2D space by 3D Gaussian splashing technology to obtain the 2D semantic segmentation prediction value.
9. The method according to claim 1, characterized in that, Before the occupancy prediction of the fused feature by the 3D occupancy prediction head, the method further includes: Performing dimensionality reduction processing on the fused feature and outputting an RGB prediction value and an opacity prediction value; Converting the RGB prediction value and the opacity prediction value into the 2D RGB prediction value by 3D Gaussian splashing technology, and supervising the 2D RGB prediction value with an RGB ground truth.
10. A 3D occupancy prediction device for a mining area, characterized in that, including: An image feature acquisition module, configured to obtain a 3D image feature based on a 2D image of a mining area; A point cloud feature acquisition module, configured to obtain a 3D point cloud feature based on a 3D point cloud of a mining area; A feature fusion module, configured to perform feature fusion on the 3D image feature and the 3D point cloud feature to obtain a fused feature; A prediction module, configured to perform occupancy prediction on the fused feature through a 3D occupancy prediction head and supervise with a 3D occupancy ground truth to obtain a 3D occupancy distribution prediction value; A first supervision module, configured to convert the 3D occupancy distribution prediction value into a 2D semantic segmentation prediction value and a depth prediction value by 3D Gaussian splashing technology, and supervise the 2D semantic segmentation prediction value with a 2D semantic segmentation ground truth, and supervise the depth prediction value with a depth ground truth.
11. A vehicle, characterized in that, The vehicle includes the 3D occupancy prediction device for a mining area according to claim 10.
12. A computer device, characterized in that, Including a first processor and a first memory, the first memory stores a program or instruction running on the first processor, and when the program or instruction is executed by the first processor, the steps of the 3D occupancy prediction method for a mining area according to any one of claims 1 to 9 are implemented.
13. A computer-readable storage medium having a program or instructions stored thereon, characterized in that, When the program or instruction is executed by a processor, the steps of the 3D occupancy prediction method for a mining area according to any one of claims 1 to 9 are implemented.
14. A chip, characterized in that, The chip includes at least one second processor and a communication interface, the communication interface is coupled to the at least one second processor, and the at least one second processor is configured to run a program or instruction to implement the steps of the 3D occupancy prediction method for a mining area according to any one of claims 1 to 9.
Citation Information
Patent Citations
3D voxel occupation prediction method based on multi-view representation and multi-modal fusion
CN118865159A
Panoramic 3D occupancy prediction method based on 3D Gaussian sputtering
CN119091412A
System and method for predicting a map from an image
WO2021175434A1