Occupied grid labeling method and device and storage medium
Through the purely visual occupancy grid truth value generation scheme, the generalization ability and self-supervisation optimization of the large model are used to solve the problems of high cost and low accuracy of radar point cloud annotation, and high-precision and efficient occupancy grid annotation are achieved.
Patent Information
- Application Number
- CN202510134501.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, radar point clouds are used to occupy grid labels with high cost, difficult acquisition, sparseness and alignment difficulties, which affects the accuracy of occupying grids.
A purely visual occupancy raster truth generation scheme is adopted, and the generalization ability of the large model is used to self-supervise and optimize the image sequence, combining feature level and image level consistency to generate accurate occupancy raster results.
The occupancy grid results generated by pure visual methods have higher accuracy and feasibility, avoiding the high cost and complex operations of radar technology and improving the efficiency of dynamic things processing.
Smart Images

Figure CN120220146A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image processing technologies, and in particular, to an occupancy grid annotation method, apparatus, and storage medium. Background Art
[0002] Currently, most of the automatic occupancy grid ground truth annotation technologies rely on lidar point clouds. However, the construction and acquisition costs of lidar ground truth vehicles are relatively high, making it difficult to collect and use them in large quantities. In addition, the sparsity of the collected point clouds also requires multi-frame fusion and densification operations to solve, and the processing of dynamic objects is even more complicated. Moreover, point cloud holes caused by reasons such as long distances, occlusion, or sensor errors also need to be filled. And there are always some gaps in the alignment of lidar point clouds and images, and the inconsistency in the visibility range also needs to be adjusted and corrected. Therefore, providing ground truth for visual occupancy grids through lidar involves many complex post-processing fusion operations and numerous factors affecting accuracy. Summary of the Invention
[0003] The purpose of the embodiments of the present disclosure is to provide an occupancy grid annotation method, apparatus, and storage medium, which adopt a pure vision-based occupancy grid ground truth generation scheme, and utilize the generalization ability of large models to perform self-supervised optimization on image sequences in a set scenario, and fully utilize feature-level and image-level consistency to solve the accuracy problem of prediction results.
[0004] To achieve the above purpose, a first aspect of the embodiments of the present disclosure provides an occupancy grid annotation method, the method including: obtaining 2D image features of an image sequence based on a vision foundation large model; obtaining a 2D image segmentation mask of the image sequence based on a segmentation large model; obtaining 3D spatial points of the image sequence by sparse reconstruction; obtaining an occupancy feature field of the image sequence based on a voxel encoder, and through geometric texture mapping and segmentation mapping of the occupancy feature field, obtaining a 2D image space of occupancy grid output and rendering output; in a self-supervised learning manner, strengthening the consistency between the 2D image space of the geometric texture mapping rendering output and the original pixels corresponding to the image sequence, and the consistency between the 2D image space of the segmentation mapping rendering output and the 2D image segmentation mask, and the consistency between the voxelized 3D spatial points and the voxel positions of the occupancy grid output, to obtain a 3D occupancy grid result corresponding to the self-supervised learning image sequence.
[0005] In some embodiments of the present disclosure, the obtaining 2D image features of an image sequence based on a vision foundation large model includes: extracting 2D image features in the image sequence by using a DINO v2 model.
[0006] In some embodiments of the present disclosure, the obtaining of the 2D image segmentation mask of the image sequence based on the segmentation large model includes: segmenting the image sequence using the Grounded SAM v2 model to obtain the 2D image segmentation mask.
[0007] In some embodiments of the present disclosure, the obtaining of the 3D spatial points of the image sequence by using sparse reconstruction includes: based on the sfm algorithm, extracting sparse feature points from the image sequence, obtaining the matching relationships of different visual feature points through the lightglue depth matcher, and using triangulation to obtain the 3D spatial points of the image sequence.
[0008] In some embodiments of the present disclosure, the obtaining of the occupancy feature field of the image sequence based on the voxel encoder includes: converting each image in the image sequence from 2D image features to 3D grid space and obtaining the voxel features of each image through 3D deformable convolution; fusing the voxel features of each image through the 3D deformable attention mechanism to obtain the occupancy feature field of the image sequence.
[0009] In some embodiments of the present disclosure, the obtaining of the 2D image space of the occupancy grid output and the rendering output by subjecting the occupancy feature field to geometric texture mapping and segmentation mapping includes: obtaining the 3D occupancy grid corresponding to the occupancy feature field through the geometric decoder; obtaining the semantic label of each occupancy grid voxel in the 3D occupancy grid corresponding to the occupancy feature field through the semantic decoder; determining the occupancy grid output based on the 3D occupancy grid and its semantic label; obtaining the texture mapping of each occupancy grid in the 3D occupancy grid according to the projection matrix from the spatial position of the 3D occupancy grid to the original image of the image sequence, and obtaining the 2D image space from the rendering output of the texture mapping; obtaining the 2D image space of the rendering output of the segmentation mapping through the semantic label of each occupancy grid voxel in the 3D occupancy grid.
[0010] In some embodiments of the present disclosure, the image sequence is a multi-view image sequence in a set scene.
[0011] In some embodiments of the present disclosure, the method further includes: using the 2D image features as a supervision signal to supervise the occupancy feature field.
[0012] According to a second aspect of the present disclosure, there is provided an occupancy grid annotation device, the device comprising: a 2D image feature extraction module for obtaining 2D image features of an image sequence based on a vision foundation large model; a mask acquisition module for obtaining a 2D image segmentation mask of the image sequence based on a segmentation large model; a voxelization module for obtaining 3D spatial points of the image sequence by means of sparse reconstruction; an encoding module for obtaining an occupancy feature field of the image sequence based on a voxel encoder, and subjecting the occupancy feature field to geometric texture mapping and segmentation mapping to obtain a 2D image space of occupancy grid output and rendering output; a self-supervised learning module for strengthening, by means of self-supervised learning, the consistency between the 2D image space of the geometric texture mapping rendering output and the original pixels corresponding to the image sequence, the consistency between the 2D image space of the segmentation mapping rendering output and the 2D image segmentation mask, and the consistency between the voxel positions of the 3D spatial points after voxelization and the occupancy grid output, so as to obtain a 3D occupancy grid result corresponding to the image sequence after self-supervised learning.
[0013] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium, characterized in that an occupancy grid annotation program is stored thereon, and when the occupancy grid annotation program is executed by a processor, the occupancy grid annotation method described in the first aspect of the present disclosure is implemented.
[0014] Other features and advantages of the embodiments of the present disclosure will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification, and are used to explain the embodiments of the present disclosure together with the following specific implementation, but do not constitute a limitation to the embodiments of the present disclosure. In the drawings:
[0016] Figure 1 is a flowchart of an occupancy grid annotation method provided by an embodiment of the present disclosure;
[0017] Figure 2 is a schematic architecture diagram of an occupancy grid annotation device provided by an embodiment of the present disclosure;
[0018] Figure 3 is a schematic architecture diagram of an occupancy grid annotation method provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of the present disclosure without creative efforts also fall within the scope of protection of the present disclosure.
[0020] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the subject matter of the present disclosure belongs. Further, it will be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the specification and the relevant art, and will not be interpreted in an idealized or overly formal form unless otherwise clearly defined herein.
[0021] Figure 1 It is a schematic flowchart of an occupancy grid annotation method provided by an embodiment of the present disclosure. As Figure 1 shown, the method includes the following steps:
[0022] Step S110, obtaining 2D image features of the image sequence based on a vision foundation large model.
[0023] Among them, in the embodiments of the present disclosure, the image sequence is a multi-view image sequence in a set scene.
[0024] Among them, the DINO v2 model is used to extract the 2D image features in the image sequence.
[0025] Step S120, obtaining a 2D image segmentation mask of the image sequence based on a segmentation large model.
[0026] Among them, the Grounded SAM v2 model is used to segment the image sequence to obtain a 2D image segmentation mask. Specifically, category prompts to be segmented are preset, and segmentation labels corresponding to each category prompt are obtained in each image in the image sequence, so as to obtain the 2D image segmentation mask corresponding to the image sequence.
[0027] Step S130, obtaining 3D spatial points of the image sequence by using sparse reconstruction.
[0028] Among them, based on the sfm algorithm, sparse feature points are extracted from the image sequence, the matching relationship of different visual feature points is obtained through a lightglue depth matcher, and 3D spatial points of the image sequence are obtained by using triangulation.
[0029] Step S140: Obtain the occupancy feature field of the image sequence based on the voxel encoder, and through geometric texture mapping and segmentation mapping of the occupancy feature field, obtain the 2D image space of the occupancy grid output and the rendering output.
[0030] Among them, each image in the image sequence is converted from 2D image features to 3D grid space, and through 3D deformable convolution, the voxel features of each image are obtained. Then, the voxel features of each image are fused through a 3D deformable attention mechanism to obtain the occupancy feature field of the image sequence.
[0031] Obtain the 3D occupancy grid corresponding to the occupancy feature field through a geometric decoder, and obtain the semantic label of each occupancy grid voxel in the 3D occupancy grid corresponding to the occupancy feature field through a semantic decoder. Finally, determine the occupancy grid output through the 3D occupancy grid and its semantic label.
[0032] In addition, the texture mapping of each occupancy grid in the 3D occupancy grid can also be obtained according to the projection matrix from the spatial position of the 3D occupancy grid to the original image of the image sequence, that is, through geometric texture mapping, obtain the scene geometric expression (such as the density of probability expression or SDF (Sign Distance Function)) and texture expression (such as the color of probability expression or spherical harmonic basis elements, etc.), and obtain the 2D image space from the rendering output of the texture mapping, and obtain the 2D image space of the rendering output of the segmentation mapping through the semantic label of each occupancy grid voxel in the 3D occupancy grid, that is, the 2D image space of the rendering output of the segmentation mapping.
[0033] Step S150: Through self-supervised learning, strengthen the consistency between the 2D image space of the geometric texture mapping rendering output and the original pixels corresponding to the image sequence, the consistency between the 2D image space of the segmentation mapping rendering output and the 2D image segmentation mask, and the consistency between the voxelization of 3D space points and the voxel positions of the occupancy grid output, to obtain the 3D occupancy grid result corresponding to the image sequence after self-supervised learning.
[0034] Through the embodiments of the present disclosure, a pure-vision occupancy grid ground truth generation scheme is adopted, and the generalization ability of the large model is utilized to perform self-supervised optimization on the image sequence in a set scene, and the feature-level and image-level consistencies are fully utilized to solve the accuracy problem of the prediction result.
[0035] In an implementation manner of the embodiments of the present disclosure, the 2D image features are used as a supervision signal to supervise the occupancy feature field, so that when the occupancy feature field is projected onto the 2D image features, it is consistent with the 2D image features.
[0036] Through the true value pre-training in the embodiments of the present disclosure, and then through self-supervised scene fitting, a large model is used to avoid the collapse of self-supervised training. Through multiple self-supervised optimizations, the generation of pure visual occupancy grid true values is obtained. The original image self-supervision is dense self-supervision, ensuring overall consistency in the 2D space, and the 3D space point self-supervision is sparse self-supervision, ensuring the consistency of 3D key points in the 3D space. The large model obtains weak label auxiliary supervision, the vision-based large model obtains generalization feature supervision, and the segmentation large model obtains generalization semantic supervision.
[0037] Correspondingly, Figure 2 is a schematic architecture diagram of an occupancy grid annotation device 20 provided by an embodiment of the present disclosure. As Figure 2 shown, the device 20 includes: a 2D image feature extraction module 21, a mask acquisition module 22, a voxelization module 23, an encoding module 24, and a self-supervised learning module 25.
[0038] Among them, the 2D image feature extraction module 21 is used to obtain 2D image features of an image sequence based on a vision-based large model;
[0039] The mask acquisition module 22 is used to obtain a 2D image segmentation mask of the image sequence based on a segmentation large model;
[0040] The voxelization module 23 is used to obtain 3D space points of the image sequence by using sparse reconstruction;
[0041] The encoding module 24 is used to obtain an occupancy feature field of the image sequence based on a voxel encoder, and through geometric texture mapping and segmentation mapping of the occupancy feature field, obtain a 2D image space of occupancy grid output and rendering output;
[0042] The self-supervised learning module 25 is used to strengthen the consistency between the 2D image space of the geometric texture mapping rendering output and the original pixels corresponding to the image sequence, the consistency between the 2D image space of the segmentation mapping rendering output and the 2D image segmentation mask, and the consistency between the voxel positions of the 3D space points after voxelization and the occupancy grid output through self-supervised learning, so as to obtain a 3D occupancy grid result corresponding to the image sequence after self-supervised learning.
[0043] Among them, the image sequence is a multi-view image sequence under a set scene.
[0044] In some embodiments of the present disclosure, the 2D image feature extraction module 21 is further used to extract 2D image features in the image sequence by using the DINO v2 model.
[0045] In some embodiments of the present disclosure, the mask acquisition module 22 is further configured to segment the image sequence using the Grounded SAM v2 model to obtain 2D image segmentation masks.
[0046] In some embodiments of the present disclosure, the voxelization module 23 is further configured to extract sparse feature points from the image sequence based on the sfm algorithm, obtain the matching relationships of different visual feature points through the lightglue depth matcher, and use triangulation to obtain the 3D spatial points of the image sequence.
[0047] In some embodiments of the present disclosure, the encoding module 24 is further configured to convert each image in the image sequence from 2D image features to 3D mesh space, and obtain voxel features of each image through 3D deformable convolution; fuse the voxel features of each image through a 3D deformable attention mechanism to obtain the occupancy feature field of the image sequence.
[0048] In some embodiments of the present disclosure, the encoding module 24 is further configured to obtain a 3D occupancy grid corresponding to the occupancy feature field through a geometric decoder; obtain the semantic label of each occupancy grid voxel in the 3D occupancy grid corresponding to the occupancy feature field through a semantic decoder; determine the occupancy grid output through the 3D occupancy grid and its semantic label; obtain the texture mapping of each occupancy grid in the 3D occupancy grid according to the projection matrix from the spatial position of the 3D occupancy grid to the original image of the image sequence, and obtain the 2D image space from the rendering output of the texture mapping; obtain the rendering output of the segmentation mapping of the 2D image space through the semantic label of each occupancy grid voxel in the 3D occupancy grid.
[0049] In some embodiments of the present disclosure, the self-supervised learning module 25 is further configured to use the 2D image features as a supervision signal to supervise the occupancy feature field.
[0050] To facilitate understanding of the embodiments of the present disclosure, Figure 3 is a schematic architecture diagram of an occupancy grid annotation method provided by an embodiment of the present disclosure. As Figure 3 shown, multi-view image sequences in a set scenario respectively pass through a vision foundation large model, a segmentation large model, sparse reconstruction, and a voxel encoder to obtain 2D image features, 2D image segmentation masks, 3D spatial points, and an occupancy feature field.
[0051] Among them, using the 2D image features as a supervision signal to supervise the occupancy feature field specifically means mapping the occupancy feature field to a 2D image to ensure that it is consistent with the 2D image features after mapping.
[0052] Among them, the occupied feature field undergoes geometric texture mapping to obtain the scene geometric expression and texture expression rendering output 2D image space of each grid in the 3D occupancy grid. Additionally, the occupied feature field undergoes segmentation mapping to obtain the spatial semantic label rendering output 2D image space of each grid in the 3D occupancy grid. The 3D occupancy grid obtained by the above geometric texture mapping and the spatial semantic label of each grid obtained by the segmentation mapping are combined to form the occupancy grid output.
[0053] In addition, through self-supervised learning, the 2D image space rendered and output by the geometric texture mapping should be consistent with the original pixels corresponding to the image sequence, and the 2D image space rendered and output by the segmentation mapping should be consistent with the 2D image segmentation mask, and the 3D space points should be voxelized to be consistent with the voxel positions of the occupancy grid output. Self-supervised optimization runs through the entire occupancy grid generation process. A predetermined upper limit of the number of iterations is set for each self-supervised learning. There will be an iteration difference for each iteration. When the iteration difference between the previous and current iterations is less than the set threshold or reaches the predetermined upper limit of the number of iterations, the 3D occupancy grid result corresponding to the image sequence after self-supervised learning is finally obtained.
[0054] Furthermore, on the other hand, an embodiment of the present disclosure also provides a machine-readable storage medium, on which an occupancy grid annotation program is stored. When the occupancy grid annotation program is executed by a processor, it implements the occupancy grid annotation method described in the above embodiments.
[0055] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a system or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0056] The present disclosure is described with reference to the flowcharts and / or block diagrams of systems (devices) and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0057] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0058] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0059] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0060] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0061] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0062] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0063] The above are only embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the scope of the claims of the present disclosure.
Claims
1. A method for marking an occupied grid, characterized in that: The method comprises: Obtain 2D image features of image sequences based on the visual basic model; Obtaining a 2D image segmentation mask of the image sequence based on the segmentation model; Obtaining 3D spatial points of the image sequence using sparse reconstruction; Obtaining an occupation feature field of the image sequence based on a voxel encoder, and subjecting the occupation feature field to geometric texture mapping and segmentation mapping to obtain a 2D image space of an occupation grid output and a rendering output; By means of self-supervised learning, the consistency between the 2D image space output by the geometric texture mapping rendering and the original pixels corresponding to the image sequence, the consistency between the 2D image space output by the segmentation mapping rendering and the 2D image segmentation mask, and the consistency between the voxel positions of the 3D space points after voxelization and the occupancy grid output are strengthened, and a 3D occupancy grid result corresponding to the image sequence after self-supervised learning is obtained.
2. The method for marking an occupied grid according to claim 1, characterized in that: The 2D image features of the image sequence obtained based on the visual basic model include: The DINO v2 model is used to extract 2D image features in the image sequence.
3. The occupancy grid annotation method according to claim 1, characterized in that: The 2D image segmentation mask of the image sequence obtained based on the segmentation large model includes: The Grounded SAM v2 model is used to segment the image sequence to obtain a 2D image segmentation mask.
4. The occupancy grid annotation method according to claim 1, characterized in that: The step of obtaining the 3D spatial points of the image sequence by sparse reconstruction comprises: Based on the SFM algorithm, sparse feature points are extracted from the image sequence, the matching relationship between different visual feature points is obtained through the Lightglue deep matcher, and the 3D spatial points of the image sequence are obtained by using the triangulation method.
5. The occupancy grid annotation method according to claim 1, characterized in that: The obtaining of the occupation feature field of the image sequence based on the voxel encoder comprises: Convert each image in the image sequence from 2D image features to 3D grid space, and obtain voxel features of each image through 3D deformable convolution; The voxel features of each image are fused through a 3D deformable attention mechanism to obtain an occupancy feature field of the image sequence.
6. The method for marking an occupied grid according to claim 1, characterized in that: The step of subjecting the occupied feature field to geometric texture mapping and segmentation mapping to obtain a 2D image space of occupied grid output and rendered output includes: Obtaining a 3D occupancy grid corresponding to the occupancy feature field through a geometric decoder; Obtaining a semantic label of each occupied grid voxel in the 3D occupied grid corresponding to the occupied feature field through a semantic decoder; Determining the occupancy grid output by using the 3D occupancy grid and its semantic label; Obtaining a texture map of each occupied grid in the 3D occupied grid according to a projection matrix from the spatial position of the 3D occupied grid to the original image of the image sequence, and obtaining a 2D image space from a rendering output of the texture map; The 2D image space of the rendered output of the segmentation map is obtained by the semantic label of each occupied grid voxel in the 3D occupied grid.
7. The occupancy grid annotation method according to claim 1, characterized in that: The image sequence is a multi-view image sequence under a set scene.
8. The occupancy grid annotation method according to claim 1, characterized in that: The method further comprises: The 2D image features are used as supervisory signals to supervise the occupancy feature field.
9. An occupancy grid labeling device, characterized in that: The device comprises: 2D image feature extraction module, used to obtain 2D image features of image sequences based on the visual basic model; A mask acquisition module, used for obtaining a 2D image segmentation mask of the image sequence based on the segmentation model; A voxelization module, used for obtaining 3D spatial points of the image sequence by sparse reconstruction; An encoding module, used for obtaining an occupation feature field of the image sequence based on a voxel encoder, and subjecting the occupation feature field to geometric texture mapping and segmentation mapping to obtain a 2D image space of an occupation grid output and a rendering output; The self-supervised learning module is used to enhance the consistency between the 2D image space output by the geometric texture mapping rendering and the original pixels corresponding to the image sequence, the consistency between the 2D image space output by the segmentation mapping rendering and the 2D image segmentation mask, and the consistency between the voxel positions of the 3D space points after voxelization and the occupancy grid output by the self-supervised learning, so as to obtain the 3D occupancy grid result corresponding to the image sequence after self-supervised learning.
10. A computer-readable storage medium, characterized in that: An occupancy grid labeling program is stored thereon, and when the occupancy grid labeling program is executed by a processor, an occupancy grid labeling method according to any one of claims 1-8 is implemented.