Semantic occupancy prediction model training method and device based on multi-modal feature fusion
A multi-modal feature fusion approach for semantic occupancy prediction models addresses the limitations of laser radar by integrating depth and semantic information from RGB images, improving scene reconstruction accuracy and detail in complex environments.
Patent Information
- Application Number
- CN202510437097.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-15
AI Technical Summary
The existing three-dimensional scene reconstruction method relies on lidar to cause high cost, large size, lack of semantic priors and low long-distance accuracy, making it difficult to achieve accurate three-dimensional scene reconstruction.
The semantic occupancy prediction model with multimodal feature fusion is adopted. By combining RGB images and pre-trained depth estimation model and large visual model, the fusion training of depth information and semantic information is carried out, and feature conversion and refinement is used to generate detailed semantic voxel maps.
More accurate and detailed image reconstruction is achieved, improving prediction accuracy of long-distance targets, especially in complex environments, reducing costs and improving the semantic understanding of the model.
Smart Images

Figure CN120318633A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the fields of computer vision and natural language processing technologies, and more specifically, to a training method and apparatus for a semantic occupancy prediction model based on multi-modal feature fusion. Background Art
[0002] In related fields, as one of the key tasks in the field of computer vision, the application scope of comprehensive three-dimensional scene understanding can range from autonomous driving and path planning to augmented reality and virtual reality. How to accurately and comprehensively reconstruct a three-dimensional scene will directly affect the implementation of downstream tasks (such as planning, navigation, and environmental mapping). In the prior art, due to the inherent limitations of sensing resolution, occlusion, and incomplete observations of available sensors, it has become a challenge to achieve accurate three-dimensional scene reconstruction. Semantic scene completion (abbreviated as SSC) has become one of the relatively effective methods to solve such challenges, which is achieved by jointly inferring the complete scene geometry and semantics from limited and usually fragmented sensor data.
[0003] However, in many actual application scenarios, the existing methods still have great limitations compared with human perception. These existing methods usually rely heavily on lidar as the main modality (because of its high accuracy in capturing 3D geometric measurements), but lidar sensors have disadvantages such as high cost and large volume. Summary of the Invention
[0004] Embodiments of the present disclosure provide a training method and apparatus for a semantic occupancy prediction model based on multi-modal feature fusion. One of its purposes is to solve the problems of lack of semantic prior and low accuracy at long distances. By generating queries rich in depth and semantic information through a coarse-to-fine semantic occupancy prediction model based on multi-modal representation fusion and refining the queries through multi-modal feature fusion processing, more accurate and detailed image reconstruction can be achieved.
[0005] In one general aspect, a training method for a semantic occupancy prediction model based on multi-modal feature fusion is provided. The training method includes: obtaining input sample data, where the input sample data includes sample RGB image data; inputting the input sample data into a pre-trained depth estimation model and a pre-trained large visual model respectively to perform depth estimation processing and semantic feature extraction processing, and obtaining corresponding depth information data and semantic information data respectively. Among them, the depth information data includes depth feature data, and the semantic information data includes first semantic feature data of a two-dimensional image and segmentation mask data; using the depth feature data and the first semantic feature data to perform the first stage of training on the semantic occupancy prediction model based on multi-modal feature fusion to obtain a trained semantic occupancy prediction model; inputting the input sample data into a preset backbone network for feature extraction to obtain corresponding two-dimensional image feature data; using the two-dimensional image feature data to perform the second stage of training on the trained semantic occupancy prediction model to obtain a final semantic occupancy prediction model.
[0006] Optionally, the step of using the depth feature data and the first semantic feature data to perform the first stage of training on the semantic occupancy prediction model based on multi-modal feature fusion to obtain a trained semantic occupancy prediction model may include: inputting the depth feature data and the first semantic feature data into the semantic occupancy prediction model based on multi-modal feature fusion to perform feature fusion processing to obtain corresponding initial voxel space data; obtaining semantic feature map data based on the initial voxel space data; calculating a first target loss in the first stage based on the semantic feature map data, and training the semantic occupancy prediction model based on multi-modal feature fusion based on the first target loss in the first stage, so as to obtain a trained semantic occupancy prediction model.
[0007] Optionally, the step of using the two-dimensional image feature data to perform a second stage of training on the trained semantic occupancy prediction model to obtain a final semantic occupancy prediction model may include: distilling the knowledge in the pre-trained large visual model into the trained semantic occupancy prediction model using a preset semantic decoder to obtain second semantic feature data corresponding to the two-dimensional image feature data; performing three-dimensional spatial transformation processing on the binary classification query data determined based on the initial voxel space data as the training result of the first stage and the two-dimensional image feature data using a deformable cross-attention mechanism to obtain corresponding three-dimensional query feature data; performing feature fusion processing on the three-dimensional query feature data, the mask token corresponding to the two-dimensional image feature data, and the initial voxel space data using a deformable self-attention mechanism to obtain corresponding three-dimensional voxel space features; performing upsampling and linear mapping on the three-dimensional voxel space features for the voxel space through a preset occupancy prediction head structure to obtain semantic voxel map data corresponding to the input sample data; calculating a second-stage second objective loss based on the second semantic feature data and the semantic voxel map data, and training the semantic occupancy prediction model based on multi-modal feature fusion based on the second-stage second objective loss, so as to obtain a final semantic occupancy prediction model.
[0008] Optionally, the first objective loss in the first stage may be a semantic scene completion loss associated with semantic feature map data determined based on the prediction result of the trained semantic occupancy prediction model.
[0009] Optionally, the objective loss function in the second stage of training is calculated by the following formula: , where, represents the objective loss function, represents the binary cross-entropy loss function for the preset semantic decoder, represents the geometric scale loss determined based on the semantic voxel map data, represents the semantic scale loss determined based on the semantic voxel map data, represents the semantic scene completion loss determined based on the semantic voxel map data, , , and each represent preset hyperparameters.
[0010] In another general aspect, a semantic occupancy prediction method based on multi-modal feature fusion is provided. The semantic occupancy prediction method includes: obtaining RGB image data to be predicted; inputting the RGB image data to be predicted into a semantic occupancy prediction model based on multi-modal feature fusion to obtain corresponding semantic voxel map data, where the semantic occupancy prediction model based on multi-modal feature fusion is trained by using the training method described above.
[0011] In another general aspect, a training device for a semantic occupancy prediction model based on multi-modal feature fusion is provided. The training device includes: a data acquisition module configured to: obtain input sample data, where the input sample data includes sample RGB image data; a first feature extraction module configured to: input the input sample data into a pre-trained depth estimation model and a pre-trained large visual model respectively to perform depth estimation processing and semantic feature extraction processing, and obtain corresponding depth information data and semantic information data respectively, where the depth information data includes depth feature data, and the semantic information data includes first semantic feature data of a two-dimensional image and segmentation mask data; a first model training module configured to: perform a first stage of training on a semantic occupancy prediction model based on multi-modal feature fusion by using the depth feature data and the first semantic feature data to obtain a trained semantic occupancy prediction model; a second feature extraction module configured to: input the input sample data into a preset backbone network for feature extraction to obtain corresponding two-dimensional image feature data; a second model training module configured to: perform a second stage of training on the trained semantic occupancy prediction model by using the two-dimensional image feature data to obtain a final semantic occupancy prediction model, where in the training process of each stage, the target loss for adjusting the prediction model parameters is: minimizing the difference between the model prediction results calculated in each stage and the corresponding ground truth data.
[0012] Optionally, the operation of the first model training module performing a first stage of training on a semantic occupancy prediction model based on multi-modal feature fusion by using the depth feature data and the first semantic feature data to obtain a trained semantic occupancy prediction model may include: inputting the depth feature data and the first semantic feature data into the semantic occupancy prediction model based on multi-modal feature fusion to perform feature fusion processing to obtain corresponding initial voxel space data; obtaining semantic feature map data based on the initial voxel space data; calculating a first target loss in the first stage based on the semantic feature map data, and training the semantic occupancy prediction model based on multi-modal feature fusion based on the first target loss in the first stage, thereby obtaining a trained semantic occupancy prediction model.
[0013] Optionally, the operation of the second model training module using the two-dimensional image feature data to perform a second stage of training on the trained semantic occupancy prediction model to obtain a final semantic occupancy prediction model may include: distilling the knowledge in the pre-trained large visual model into the trained semantic occupancy prediction model using a preset semantic decoder to obtain second semantic feature data corresponding to the two-dimensional image feature data; performing three-dimensional space conversion processing on the binary classification query data determined based on the initial voxel space data as the training result of the first stage and the two-dimensional image feature data using a deformable cross-attention mechanism to obtain corresponding three-dimensional query feature data; performing feature fusion processing on the three-dimensional query feature data, the mask token corresponding to the two-dimensional image feature data, and the initial voxel space data using a deformable self-attention mechanism to obtain corresponding three-dimensional voxel space features; performing upsampling and linear mapping on the three-dimensional voxel space features for the voxel space through a preset occupancy prediction head structure to obtain semantic voxel map data corresponding to the input sample data; calculating a second-stage second objective loss based on the second semantic feature data and the semantic voxel map data, and training the semantic occupancy prediction model based on multi-modal feature fusion based on the second-stage second objective loss, thereby obtaining a final semantic occupancy prediction model.
[0014] Optionally, the first objective loss in the first stage may be a semantic scene completion loss associated with semantic feature map data determined based on the prediction result of the trained semantic occupancy prediction model.
[0015] Optionally, the objective loss function in the second stage of training is calculated by the following formula: , where, represents the objective loss function, represents the binary cross-entropy loss function for the preset semantic decoder, represents the geometric scale loss determined based on the semantic voxel map data, represents the semantic scale loss determined based on the semantic voxel map data, represents the semantic scene completion loss determined based on the semantic voxel map data, , , and each represent preset hyperparameters.
[0016] In another general aspect, a computer program product is provided, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the training method of the semantic occupancy prediction model based on multi-modal feature fusion and the semantic occupancy prediction method based on multi-modal feature fusion as described above are implemented.
[0017] In another general aspect, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device / server, the electronic device / server is enabled to execute the training method of the semantic occupancy prediction model based on multi-modal feature fusion and the semantic occupancy prediction method based on multi-modal feature fusion as described above.
[0018] In another general aspect, a computing device is provided, which includes: at least one processor; at least one memory storing computer-executable instructions. Wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the training method of the semantic occupancy prediction model based on multi-modal feature fusion and the semantic occupancy prediction method based on multi-modal feature fusion as described above.
[0019] According to the training method of the semantic occupancy prediction model based on multi-modal feature fusion and the semantic occupancy prediction method based on multi-modal feature fusion of the embodiments of the present disclosure, by generating queries rich in depth and semantic information through a coarse-to-fine semantic occupancy prediction model based on multi-modal representation fusion and refining the queries through multi-modal feature fusion processing, more accurate and detailed image reconstruction can be achieved. In addition, through the coarse-to-fine semantic occupancy prediction model and prediction method based on multi-modal representation fusion, there is a prominent improvement effect in improving the prediction accuracy of distant targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Through the description below with reference to the drawings showing embodiments, the above and other objects and features of the embodiments of the present disclosure will become clearer, where: Figure 1 is a flowchart showing the training method of the semantic occupancy prediction model based on multi-modal feature fusion according to an embodiment of the present disclosure; Figure 2 is a flowchart showing an example of the training method of the semantic occupancy prediction model based on multi-modal feature fusion according to an embodiment of the present disclosure; Figure 3 is a schematic diagram showing an exemplary deformable multi-modal fusion process according to an embodiment of the present disclosure; Figure 4 is a schematic diagram showing an exemplary coarse-to-fine voxel generation process according to an embodiment of the present disclosure; Figure 5A is a schematic diagram showing the semantic segmentation effect according to an embodiment of the present disclosure; Figure 5B is a schematic diagram showing a performance comparison table of the semantic occupancy prediction method according to an embodiment of the present disclosure and other methods; Figure 6 is a flowchart showing the semantic occupancy prediction method based on multi-modal feature fusion according to an embodiment of the present disclosure; Figure 7 is a block diagram showing a training device of a semantic occupancy prediction model based on multi-modal feature fusion according to an embodiment of the present disclosure; Figure 8 is a block diagram showing a computing device according to an embodiment of the present disclosure. Detailed Description of Embodiments
[0021] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent after understanding the disclosure of this application. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of this application, except for operations that must occur in a specific order. In addition, descriptions of features known in the art may be omitted for greater clarity and conciseness.
[0022] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings, wherein like reference numerals always refer to like elements. The following embodiments will be described with reference to the accompanying drawings to explain the present disclosure.
[0023] In the related art, although progress has been made in semantic scene completion, there are still many problems with existing methods. For example, existing methods may produce different outputs depending on the model, which may hinder the development of general algorithms.
[0024] To solve the problems in the related art, the present disclosure proposes a training method and device for a semantic occupancy prediction model based on multi-modal feature fusion. By adopting a camera-based SSC solution and utilizing the rich visual information captured by the camera, a more cost-effective solution is provided, and a depth and semantic information-rich query is generated through a coarse-to-fine semantic occupancy prediction model based on multi-modal representation fusion and refined through multi-modal feature fusion processing, which can accurately reconstruct occluded regions and maintain cross-camera geometric consistency, thereby achieving excellent performance when applied to complex environments (e.g., complex driving environments).
[0025] The following refers toFigures 1 to 8 Describe in detail a training method for a semantic occupancy prediction model based on multi-modal feature fusion and a semantic occupancy prediction method based on multi-modal feature fusion according to an embodiment of the present disclosure.
[0026] First, refer to Figures 1 to 5B to describe in detail a training method for a semantic occupancy prediction model based on multi-modal feature fusion according to an embodiment of the present disclosure. Figure 1 is a flowchart showing a training method 100 for a semantic occupancy prediction model based on multi-modal feature fusion according to an embodiment of the present disclosure. Figure 2 is a flowchart showing an example of a training method for a semantic occupancy prediction model based on multi-modal feature fusion according to an embodiment of the present disclosure. Figure 3 is a schematic diagram showing an exemplary deformable multi-modal fusion process according to an embodiment of the present disclosure. Figure 4 is a schematic diagram showing an exemplary voxel generation process from coarse to fine according to an embodiment of the present disclosure. Figure 5A is a schematic diagram showing a semantic segmentation effect according to an embodiment of the present disclosure. Figure 5B is a schematic diagram showing a performance comparison table of a semantic occupancy prediction method and other methods according to an embodiment of the present disclosure.
[0027] Refer to Figure 1 , according to an embodiment of the present disclosure, in step S101, obtain input sample data.
[0028] Here, the input sample data includes sample RGB image data.
[0029] , according to an embodiment of the present disclosure, in step S102, input the input sample data into a pre-trained depth estimation model and a pre-trained large visual model respectively to perform depth estimation processing and semantic feature extraction processing, and obtain corresponding depth information data and semantic information data respectively.
[0030] Here, the depth information data includes depth feature data, and the semantic information data includes first semantic feature data of a two-dimensional image (for example, a two-dimensional image corresponding to the input image data) and segmentation mask data.
[0031] , according to an embodiment of the present disclosure, in step S103, use the depth feature data and the first semantic feature data to perform the first stage of training on the semantic occupancy prediction model based on multi-modal feature fusion, and obtain a trained semantic occupancy prediction model.
[0032] As an example, step S103 may further include the following steps S1031 to S1033: In step S1031, the depth feature data and the first semantic feature data are input into a semantic occupancy prediction model based on multi-modal feature fusion to perform feature fusion processing, obtaining corresponding initial voxel space data.
[0033] In step S1032, based on the initial voxel space data, semantic feature map data and corresponding binary classification query data are obtained.
[0034] For example, the semantic feature map data is used to calculate the objective loss in the first stage, and the binary classification query data is used to calculate the objective loss in the second stage.
[0035] In step S1033, the first objective loss in the first stage is calculated based on the semantic feature map data, and the semantic occupancy prediction model based on multi-modal feature fusion is trained based on the first objective loss in the first stage, thereby obtaining a trained semantic occupancy prediction model.
[0036] According to the present disclosure, by using the semantic features in the two-dimensional image and the semantic features in the segmentation mask for the first-stage training process, it is possible to fuse them with the depth features to correct the depth and improve the quality of the voxel space, and by using the segmentation mask for the distillation process in the second stage, the semantic understanding ability of the model is enhanced.
[0037] In addition, according to the present disclosure, a method for realizing the fusion of depth features and image features in the first stage is provided, which can integrate multi-modal features, and by using a pre-trained large-scale visual model to enhance feature representation, the training cost can also be reduced.
[0038] For example, an example of the first-stage training process of the semantic occupancy prediction model based on multi-modal feature fusion can be as Figure 2 shown in steps S201 to S203 and its example process can be as Figure 3 shown, and these steps and processes will be elaborated in detail later.
[0039] According to an embodiment of the present disclosure, in step S104, the input sample data is input into a preset backbone network for feature extraction, obtaining corresponding two-dimensional image feature data.
[0040] According to an embodiment of the present disclosure, in step S105, the trained semantic occupancy prediction model is trained in the second stage using the two-dimensional image feature data, obtaining a final semantic occupancy prediction model. Here, the semantic occupancy prediction model in the present disclosure can be applied to the field of semantic occupancy prediction. For example, the semantic occupancy prediction model in the present disclosure can be applied to 3D perception in the field of autonomous driving, and can infer the complete 3D geometry and semantic information of the scene only through a two-dimensional image.
[0041] As an example, step S105 may further include the following steps S1051 to S1055: In step S1051, a preset semantic decoder is used to distill the knowledge in the pre-trained large visual model into the trained semantic occupancy prediction model, obtaining second semantic feature data corresponding to the two-dimensional image feature data. Here, the second semantic feature data is used to calculate the objective loss in the second stage.
[0042] In step S1052, a deformable cross-attention mechanism is used to perform three-dimensional space transformation processing on the binary classification query data and the two-dimensional image feature data determined based on the initial voxel space data as the training result of the first stage, obtaining corresponding three-dimensional query feature data.
[0043] In step S1053, a deformable self-attention mechanism is used to perform feature fusion processing on the three-dimensional query feature data, the mask token corresponding to the two-dimensional image feature data, and the initial voxel space data, obtaining corresponding three-dimensional voxel space features.
[0044] In step S1054, a preset occupancy prediction head structure is used to perform upsampling and linear mapping on the three-dimensional voxel space features for the voxel space, obtaining semantic voxel map data corresponding to the input sample data.
[0045] In step S1055, the second objective loss in the second stage is calculated based on the second semantic feature data and the semantic voxel map data, and the semantic occupancy prediction model based on multi-modal feature fusion is trained based on the second objective loss in the second stage, thereby obtaining the final semantic occupancy prediction model.
[0046] According to the present disclosure, by distilling the knowledge of the large-scale visual model into the model of the present disclosure, a method of applying this knowledge to the occupancy network prediction task is realized, improving the accuracy of image segmentation and fully enhancing the visual understanding ability of the model of the present disclosure.
[0047] In addition, according to the present disclosure, by adopting a deformable attention mechanism to construct the network, the problem of high computational complexity of the traditional attention mechanism in processing high-resolution images and long sequences can be solved.
[0048] In addition, according to the present disclosure, by adopting a lightweight deformable attention method in the second-stage training process, using the initial voxels obtained in the first stage to enhance the three-dimensional query features, and extracting knowledge from the pre-trained large-scale visual model at the same time, the semantic understanding ability of the model can be improved, ensuring the optimization of the model performance without further increasing the model size.
[0049] For example, an example of the second-stage training process of the semantic occupancy prediction model based on multi-modal feature fusion can be as shown in steps S204 to S208 in Figure 2 and its example process can be as shown in Figure 4 shown, and these steps and processes will be elaborated in detail later.
[0050] In the present disclosure, during the training process of each stage, the target loss for adjusting the prediction model parameters is: minimizing the difference between the model prediction results calculated in each stage and the corresponding ground truth data.
[0051] Specifically, as an example, in the first stage, the first target loss in the first stage is the semantic scene completion loss associated with the semantic feature map data determined based on the prediction result of the trained semantic occupancy prediction model.
[0052] By using the first target loss, the depth can be corrected and the quality of the voxel space can be improved. For example, in the first stage, by combining the depth data of the depth estimation model with the semantic interpretation extracted from the visual model, queries rich in both depth and semantic information are generated, and these information are initially fused through, for example, the UNet architecture, thus creating an initial rough representation.
[0053] In addition, as an example, in the second stage, the target loss function in the training of the second stage can be calculated by the following formula (1): (1) where, represents the target loss function in the training of the second stage, represents the binary cross-entropy loss function for the preset semantic decoder, represents the geometric scale loss determined based on the semantic voxel map data, represents the semantic scale loss determined based on the semantic voxel map data, represents the semantic scene completion loss determined based on the semantic voxel map data, , , and each represent preset hyperparameters.
[0054] By using the target loss function in the training of the second stage, the alignment accuracy of the model in terms of geometric scale and semantic scale can be improved, and the semantic scene completion effect can be enhanced. For example, in the second stage, the above queries are refined by the multi-modal feature fusion module, that is, the image features are combined with the initial queries, and this multi-modal fusion can achieve more accurate and detailed reconstruction, especially with a significant improvement in the prediction accuracy of distant targets, successfully solving one of the common problems in the camera-based SSC method.
[0055] According to the present disclosure, the voxel map obtained by the above training method of the present disclosure shows a clearer segmentation, with less overlap between voxels of different categories, and has achieved remarkable breakthrough results in the image segmentation of some small objects and long-tailed objects (such as trucks and bicycles, etc.).
[0056] Next, with reference to Figure 2 to illustrate the flowchart of the training method 200 of the semantic occupancy prediction model based on multi-modal feature fusion. Specifically, an example of the first-stage training process of the semantic occupancy prediction model based on multi-modal feature fusion can be as shown in Figure 2 steps S201 to S203, and an example of the second-stage training process of the semantic occupancy prediction model based on multi-modal feature fusion can be as shown in Figure 2 steps S204 to S208.
[0057] In step S201, depth information can be obtained by using a pre-trained depth estimation network.
[0058] For example, the input data for this step includes the left and right view RGB image pairs from a binocular camera.
[0059] In addition, the pre-trained depth estimation network can be, for example, a pre-trained binocular depth estimation network, which adopts an encoder-decoder architecture. Here, the encoder is responsible for extracting multi-scale features of the image, and the decoder generates a pixel-level depth map based on these features. As an example, this depth estimation network can be pre-trained on a public dataset (such as, but not limited to, the KITTI dataset or the SceneFlow dataset, etc.) to ensure the robustness and accuracy of depth prediction.
[0060] In this step, point cloud information can be generated. Specifically, by combining the above depth map and relevant camera parameters, the two-dimensional pixel points are back-projected into the three-dimensional point cloud space to obtain the initial point cloud information.
[0061] In addition, in this step S201, semantic information can also be obtained by using a pre-trained large visual model.
[0062] For example, the input data for this step includes: the same RGB image data as the input data described above.
[0063] In addition, the pre-trained large visual model (which can also be called a large-scale visual model) can extract rich semantic features from the image, including segmentation masks, target regions, and boundary information, etc. In addition, as an example, an example of this large visual model can be Grounded-SAM, etc., but not limited to this.
[0064] Specifically, a multi-scale feature extraction can be performed on the input image using the image encoder of a large visual model. For example, semantic embedding features can be generated by calculating the semantic boundaries of the target region and their corresponding class labels. For example, a pixel-level segmentation mask can be generated through the segmentation head of the model to describe the semantic category of each pixel in the scene.
[0065] For example, the semantic information obtained through the above steps may include the semantic features and segmentation masks of a two-dimensional image (e.g., a two-dimensional image corresponding to the input image data). Here, the semantic features are used for the first-stage training, fused with the depth features to correct the depth and improve the quality of the voxel space, and the segmentation masks are used for the distillation process in the second-stage training (such as the "distillation module" shown Figure 3 ), to enhance the semantic understanding ability of the model (e.g., the ability in accurately segmenting the boundaries and shapes of objects, etc.).
[0066] In step S202, an initial voxel space is generated by fusing the above depth features and semantic features .
[0067] Specifically, first, the initial point cloud information is fused with the image features extracted by the large visual model.
[0068] Secondly, the two-dimensional information is transferred to the three-dimensional space through a lightweight U-Net, thereby realizing the extraction and fusion of multi-modal features.
[0069] Finally, an initial voxel space is preliminarily constructed through a three-dimensional (3D) convolutional layer , and its expression is shown in Equation (2) below: (2) Where, represents the above-extracted image features, represents the above depth features, , respectively represent the channels, height, and width. In addition, the subscript raw indicates that this feature is the result directly generated after being processed in the preliminary stage of the network and has not undergone further refinement or optimization.
[0070] In step S203, a classification segmentation head is applied to the obtained initial voxel space , and a semantic feature map is obtained. Each channel here corresponds to a class occupancy prediction, and the specific expression is shown in Equation (3) below: (3) Where, the above , , and respectively represent the channel, height, width, and depth, represents class channels, represents the classification segmentation head.
[0071] Here, in order to retain more rich and complete abstract feature information, is retained during the second-stage training process, while is only used for the calculation of the loss function in the first stage.
[0072] In addition, during the training process of the first stage, the present disclosure also uses, for example, LMSCNet to obtain a total of binary classification queries . Here, each voxel is labeled as 1 if it contains at least one point. The subscript d represents the feature dimension. will be used as the mask index of the deformable attention mechanism for the second-stage training.
[0073] In addition, during the training process of the first stage, the above semantic features of the first stage are used for the calculation of the loss function. For example, the loss function can be the semantic scene completion loss .
[0074] In the exemplary first-stage processing of the present disclosure (for example, the processing shown in each box of Figure 3 ), through the above steps based on the fusion of depth and image features, multi-modal features can be integrated, and by using a pre-trained large-scale visual model to enhance the feature representation, the training cost is effectively reduced at the same time.
[0075] In step S204, a preset backbone network (for example, Resnet 50) is used to extract image features . Here, the subscript of represents the features obtained by processing the two-dimensional image through the network, represents the feature map dimension range, represents the height of the feature map, represents the width of the feature map, represents the number of channels of the feature map.
[0076] In addition, it should be noted that although Figure 2 shows that step S204 is executed after step S203, the extraction step of the above image features can also be executed in parallel with any step in S201 to S203, and the present disclosure is not limited thereto.
[0077] In step S205, knowledge distillation is performed to enhance the semantic understanding ability of the model. For example, the distillation module shown in Figure 4 can be used to distill the knowledge in Grounded-SAM into the model and enhance the semantic understanding ability of the model.
[0078] Specifically, by introducing a semantic decoder , whose input is the two-dimensional image features extracted above , and using the segmentation mask labels generated in the first stage as the ground truth, the specific expression is shown in Equation (4) below: (4) where represents the semantic features obtained through the semantic decoder, and the superscript indicates that this feature is a feature representation in the two-dimensional space.
[0079] Here, the binary cross-entropy loss is used to calculate the difference between the prediction result and the segmentation mask label to optimize the network parameters (which can also refer to optimizing the model parameters in this disclosure).
[0080] In step S206, the deformable cross-attention mechanism is used to guide the two-dimensional features to be embedded into the three-dimensional space.
[0081] For example, the input data for this step includes: two-dimensional image features and binary classification queries .
[0082] In addition, regarding the specific application of the deformable cross-attention mechanism (DCA): using the binary classification query as the guiding index, through the use of the deformable cross-attention mechanism (DCA), the two-dimensional image features are embedded into the three-dimensional space, that is, the three-dimensional query features are obtained, thus effectively guiding the conversion of the 2D feature map to a structured 3D representation. The specific expression is shown in Equation (5) below: (5) where represents the finally generated three-dimensional query features, which are used to guide the subsequent operations in the 3D voxel space, represents the deformable cross-attention mechanism, represents the two-dimensional image features, represents the binary classification query. In addition, the subscript indicates that the query features are generated in the current stage, and the superscript indicates that these query features have been projected into the three-dimensional space and have a three-dimensional structured representation.
[0083] Here, different from traditional fixed attention, DCA can dynamically adjust the spatial position of attention according to the input query to improve the flexibility and accuracy of feature mapping.
[0084] In step S207, a deformable self-attention mechanism is used to refine the voxel features and enhance the representation ability. The specific steps are as follows: First, the initial voxel space obtained in the first stage is fused with the three-dimensional query features obtained in the second stage In addition, mask tokens based on can be added to the voxel space to complete the scene .
[0085] Then, by using the deformable self-attention mechanism (DSA), the completed voxel space is updated, which will be used for the prediction operation. The specific expression is shown in Equation (6) below: (6) where represents the optimized three-dimensional query features and is input into the self-attention mechanism, represents the deformable self-attention mechanism, represents the three-dimensional voxel space features refined by the deformable self-attention module.
[0086] By using DSA, the features can be refined in the three-dimensional voxel space, that is, by introducing a dynamic offset different from the traditional self-attention mechanism, the attention is focused on the regions with important features in the three-dimensional space.
[0087] In step S208, the refined voxel space features pass through the occupancy prediction head, and the voxel space is upsampled and linearly mapped to obtain the final semantic voxel map .
[0088] Here, the subscript represents the current time step, represents the number of categories, , , represent the 3D volume dimensions.
[0089] In addition, in the second stage, multiple losses can be used for joint training to optimize the model parameters. Specifically, for the semantic decoder , the binary cross-entropy loss function can be used, for example. In addition, for the finally output semantic voxel map, the geometric scale loss , the semantic scale loss , the semantic scene completion loss The total loss function in the second stage here can be expressed as the above formula (1), which will not be elaborated here.
[0090] In the exemplary second stage processing of the present disclosure (e.g., the processing shown in each box of Figure 4 ), by adopting a lightweight deformable attention method and using the initial voxel space obtained in the first stage to enhance the 3D query features, and extracting knowledge from a pre-trained large-scale vision model, the semantic understanding of the model can be improved, and the optimization of the model performance can be ensured without further increasing the model scale.
[0091] According to the present disclosure, in terms of the application effect of the above training method according to the embodiments of the present disclosure, the effectiveness of the coarse-to-fine semantic occupancy prediction method based on multi-modal representation fusion of the present disclosure has been verified on a preset dataset. For example, the prediction method of the present disclosure has achieved state-of-the-art performance (i.e., the best performance) in the camera-based SemanticKITTI (which is a large-scale autonomous driving dataset) benchmark test.
[0092] For example, it can be as shown in the semantic segmentation effect of the prediction method of the present disclosure in Figure 5A and Figure 5B . Specifically, in Figure 5A , the "Camera View" in Diagram (1) represents the "camera view" diagram, the "Ground Truth" in Diagram (2) represents the corresponding "ground truth" diagram, the "CFOcc" in Diagram (3) represents the result diagram of the prediction method of the present disclosure, and Diagrams (4) to (6) respectively represent the result diagrams of other methods as comparative examples.
[0093] In Figure 5B , "DMFNet" in the leftmost column of the table represents the network model in the first stage of the present disclosure, "CFOcc" represents the prediction method of the present disclosure, and the methods other than the above two methods are other methods as comparative examples. In addition, "Camera" in the second column of the table represents the (monocular) camera image / RGB image, that is, pure RGB image data (two-dimensional data); "Camera and Depth" represents the RGB-D image (color + depth) / multi-modal camera image, that is, multi-modal image data (multi-modal data, strictly pixel-aligned).
[0094] In addition, in Figure 5A , the colored squares represent the class labels represented by their corresponding colors, and in Figure 5BThe corresponding various types of tags are also shown in the table. For example, the meanings of the respective tags in these two figures are as follows: "road: standard road, sidewalk: sidewalk, parking: parking, other-grnd (as shown in Figure 5A ) / other-ground (as shown in Figure 5B ): non-standard road, building: building, vegetation: vegetation, trunk: tree trunk, terrain. (as shown in Figure 5A ) / terrain (as shown in Figure 5B ): grassland, traf.-sign (as shown in Figure 5A ) / traffic-sign (as shown in Figure 5B ): traffic sign, pole: pole, car: ordinary passenger car, person: pedestrian, bicyclist: bicyclist, motorcyclist: motorcyclist, fence: fence, truck: freight truck, bicycle: bicycle, motorcycle: motorcycle, other-veh. (as shown in Figure 5A ) / other-vehicle (as shown in Figure 5B ): other vehicles".
[0095] From Figure 5A and Figure 5B , it can be seen that the prediction method of the present disclosure is optimal in terms of prediction results and algorithm performance compared with other methods. For example, Figure 5B the value of the algorithm evaluation parameter mIoU (i.e., mean intersection over union) shown is optimal. It should be noted that in Figure 5B , the bold black numerical values are the optimal values, and the underlined numerical values are the sub-optimal values.
[0096] For example, the model proposed by the present disclosure shows greater advantages than the existing models in the close-range scenario, which enables the method of the present disclosure to be effectively applied to practical scenarios, such as the autonomous driving scenario. In addition, since the accurate perception of the close range by the model can improve its judgment of a farther distance, there is a prominent improvement effect in improving the prediction accuracy of distant targets.
[0097] In addition, by performing coarse-to-fine enhancement on each category in the second stage, the model obtains very significant beneficial effects in the segmentation of small objects and long-tail objects (such as trucks and bicycles, etc.), thus proving the effectiveness and robustness of the model of the present disclosure in complex real-world scenarios.
[0098] In addition, the model proposed by the present disclosure can also show clearer segmentation, resulting in less overlap between voxels of different categories.
[0099] Here, it should be noted that the above descriptions of Figure 2 each step are merely exemplary, and the steps in the method according to the present disclosure are not limited thereto.
[0100] Next, with reference to Figure 6 the specific steps of the semantic occupancy prediction method 600 based on multi-modal feature fusion according to an embodiment of the present disclosure will be described.
[0101] Figure 6 FIG. is a flowchart showing the semantic occupancy prediction method 600 based on multi-modal feature fusion according to an embodiment of the present disclosure.
[0102] With reference to Figure 6 , in step S601, RGB image data to be predicted is obtained.
[0103] In step S602, the RGB image data to be predicted is input into the semantic occupancy prediction model based on multi-modal feature fusion to obtain corresponding semantic voxel map data.
[0104] Here, the semantic occupancy prediction model based on multi-modal feature fusion is trained using the training method 100 as described above.
[0105] Next, with reference to Figure 7 the specific operations of the training device 700 for the semantic occupancy prediction model based on multi-modal feature fusion according to an embodiment of the present disclosure will be described.
[0106] Figure 7 FIG. is a block diagram showing the training device 700 for the semantic occupancy prediction model based on multi-modal feature fusion according to an embodiment of the present disclosure.
[0107] With reference to Figure 7 , the training device 700 for the semantic occupancy prediction model based on multi-modal feature fusion according to an embodiment of the present disclosure may include: a data acquisition module 710, a first feature extraction module 720, a first model training module 730, a second feature extraction module 740, and a second model training module 750.
[0108] According to an embodiment of the present disclosure, the data acquisition module 710 may execute: obtaining input sample data. Here, the input sample data includes sample RGB image data.
[0109] According to an embodiment of the present disclosure, the first feature extraction module 720 may perform: inputting the input sample data into a pre-trained depth estimation model and a pre-trained large visual model respectively to perform depth estimation processing and semantic feature extraction processing, and obtaining corresponding depth information data and semantic information data respectively.
[0110] Here, by way of example, the depth information data includes depth feature data, and the semantic information data includes first semantic feature data of a two-dimensional image and segmentation mask data.
[0111] According to an embodiment of the present disclosure, the first model training module 730 may perform: using the depth feature data and the first semantic feature data to perform the first stage of training on a semantic occupancy prediction model based on multi-modal feature fusion, and obtaining a trained semantic occupancy prediction model.
[0112] By way of example, the first model training module 730 may further perform the following operations 731) to 733): In operation 731), input the depth feature data and the first semantic feature data into a semantic occupancy prediction model based on multi-modal feature fusion to perform feature fusion processing, and obtain corresponding initial voxel space data.
[0113] In operation 732), based on the initial voxel space data, obtain semantic feature map data and corresponding binary classification query data.
[0114] Here, the semantic feature map data is used to calculate the objective loss of the first stage, and the binary classification query data is used to calculate the objective loss of the second stage.
[0115] In operation 733), calculate the first objective loss of the first stage based on the semantic feature map data, and train the semantic occupancy prediction model based on multi-modal feature fusion based on the first objective loss of the first stage, thereby obtaining a trained semantic occupancy prediction model.
[0116] According to an embodiment of the present disclosure, the second feature extraction module 740 may perform: inputting the input sample data into a preset backbone network for feature extraction, and obtaining corresponding two-dimensional image feature data.
[0117] According to an embodiment of the present disclosure, the second model training module 750 may perform: using the two-dimensional image feature data to perform the second stage of training on the trained semantic occupancy prediction model, and obtaining a final semantic occupancy prediction model.
[0118] By way of example, the second model training module 750 may further perform the following operations 751) to 755): In operation 751), based on the two-dimensional image feature data, the knowledge in the pre-trained large visual model is distilled into the trained semantic occupancy prediction model by using a preset semantic decoder, and the second semantic feature data corresponding to the two-dimensional image feature data is obtained. Here, the second semantic feature data is used to calculate the objective loss in the second stage.
[0119] In operation 752), a three-dimensional spatial transformation process is performed on the binary classification query data and the two-dimensional image feature data determined based on the initial voxel space data as the training result of the first stage by using a deformable cross-attention mechanism, and the corresponding three-dimensional query feature data is obtained.
[0120] In operation 753), a feature fusion process is performed on the three-dimensional query feature data, the mask token corresponding to the two-dimensional image feature data, and the initial voxel space data by using a deformable self-attention mechanism, and the corresponding three-dimensional voxel space feature is obtained.
[0121] In operation 754), upsampling and linear mapping for the voxel space are performed on the three-dimensional voxel space feature through a preset occupancy prediction head structure, and the semantic voxel map data corresponding to the input sample data is obtained.
[0122] In operation 755), the second objective loss in the second stage is calculated based on the second semantic feature data and the semantic voxel map data, and the semantic occupancy prediction model based on multi-modal feature fusion is trained based on the second objective loss in the second stage, so as to obtain the final semantic occupancy prediction model.
[0123] According to the present disclosure, in the training process of each stage, the objective loss for adjusting the prediction model parameters is: minimizing the difference between the model prediction result calculated in each stage and the corresponding ground truth data.
[0124] For example, in the first stage, the first objective loss in the first stage is the semantic scene completion loss associated with the semantic feature map data determined based on the prediction result of the trained semantic occupancy prediction model.
[0125] In addition, in the second stage, the objective loss function in the training of the second stage can be calculated through the above formula (1).
[0126] It should be noted that the operations performed on the above respective structural blocks may be similar to the related content described with reference to Figure 1 and will not be elaborated here.
[0127] Figure 8 is a block diagram showing a computing device 800 according to an embodiment of the present disclosure.
[0128] Referring to Figure 8, a computing device 800 according to an embodiment of the present disclosure may include a processor 810 and a memory 820. The processor 810 may include (but is not limited to) a central processing unit (CPU), a digital signal processor (DSP), a microcomputer, a field programmable gate array (FPGA), a system on chip (SoC), a microprocessor, an application specific integrated circuit (ASIC), etc. The memory 820 may store computer-executable instructions to be executed by the processor 810. The memory 820 includes high-speed random access memory and / or non-volatile computer-readable storage media. When the processor 810 executes the computer-executable instructions stored in the memory 820, the training method of the semantic occupancy prediction model based on multi-modal feature fusion and the semantic occupancy prediction method based on multi-modal feature fusion as described above can be implemented.
[0129] The training method of the semantic occupancy prediction model based on multi-modal feature fusion and the semantic occupancy prediction method based on multi-modal feature fusion according to the embodiments of the present disclosure can be written as computer programs / instructions to form a computer program product and stored on a computer-readable storage medium. When the computer programs / instructions are executed by a processor, the training method of the semantic occupancy prediction model based on multi-modal feature fusion and the semantic occupancy prediction method based on multi-modal feature fusion as described above can be realized. When the instructions in the computer-readable storage medium are executed by the processor of an electronic device / server, the electronic device / server can be enabled to execute the training method of the semantic occupancy prediction model based on multi-modal feature fusion and the semantic occupancy prediction method based on multi-modal feature fusion as described above. Examples of computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device configured to store a computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer such that the processor or computer can execute the computer program. In one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system such that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0130] The training method of the semantic occupancy prediction model based on multi-modal feature fusion and the semantic occupancy prediction method based on multi-modal feature fusion according to the embodiments of the present disclosure can generate queries rich in depth and semantic information through a coarse-to-fine semantic occupancy prediction model based on multi-modal representation fusion and refine the queries through multi-modal feature fusion processing, enabling more accurate and detailed image reconstruction.
[0131] On the other hand, through a coarse-to-fine semantic occupancy prediction model and prediction method based on multi-modal representation fusion, there is a prominent improvement effect in improving the prediction accuracy of distant targets.
[0132] On the other hand, according to the training method of the semantic occupancy prediction model based on multi-modal feature fusion and the semantic occupancy prediction method according to the embodiments of the present disclosure, by implementing a deformable multi-modal fusion network, a coarse-to-fine voxel generation network, and a semantic distillation module, the method of the present disclosure has achieved state-of-the-art performance in camera-based semantic scene completion, and has proven its effectiveness and robustness in complex real-world scenarios through experiments.
[0133] On the other hand, according to the two-stage coarse-to-fine semantic occupancy prediction method based on multi-modal representation fusion of the embodiments of the present disclosure, it is possible to improve the existing semantic occupancy prediction method by introducing a large visual model into the semantic occupancy task and through a semantic-assisted loss and a coarse-to-fine voxel generation framework. By combining a large visual model, more comprehensive knowledge distillation is carried out into the semantic occupancy task, improving the performance of the overall system framework while also being able to maintain a balance in terms of efficiency.
[0134] Although some embodiments of the present disclosure have been disclosed and described, those skilled in the art should understand that these embodiments can be modified and varied without departing from the concept and spirit of the present disclosure as defined by the claims and their equivalents.
Claims
1. A training method for a semantic occupancy prediction model based on multi-modal feature fusion, characterized in that The training method includes: Obtain input sample data, where the input sample data includes sample RGB image data; Input the input sample data into a pre-trained depth estimation model and a pre-trained large-scale vision model respectively to perform depth estimation processing and semantic feature extraction processing respectively, and obtain corresponding depth information data and semantic information data. Among them, the depth information data includes depth feature data, and the semantic information data includes first semantic feature data of a two-dimensional image and segmentation mask data; Use the depth feature data and the first semantic feature data to perform the first-stage training on a semantic occupancy prediction model based on multi-modal feature fusion to obtain a trained semantic occupancy prediction model; Input the input sample data into a preset backbone network for feature extraction to obtain corresponding two-dimensional image feature data; Use the two-dimensional image feature data to perform the second-stage training on the trained semantic occupancy prediction model to obtain a final semantic occupancy prediction model.
2. The training method according to claim 1, wherein The step of using the depth feature data and the first semantic feature data to perform the first-stage training on a semantic occupancy prediction model based on multi-modal feature fusion to obtain a trained semantic occupancy prediction model includes: Input the depth feature data and the first semantic feature data into the semantic occupancy prediction model based on multi-modal feature fusion to perform feature fusion processing to obtain corresponding initial voxel space data; Based on the initial voxel space data, obtain semantic feature map data; Calculate the first objective loss of the first stage based on the semantic feature map data, and train the semantic occupancy prediction model based on multi-modal feature fusion based on the first objective loss of the first stage, so as to obtain a trained semantic occupancy prediction model.
3. The training method according to claim 1, wherein The step of using the two-dimensional image feature data to perform the second-stage training on the trained semantic occupancy prediction model to obtain a final semantic occupancy prediction model includes: Use a preset semantic decoder to distill the knowledge in the pre-trained large-scale vision model into the trained semantic occupancy prediction model to obtain second semantic feature data corresponding to the two-dimensional image feature data; Use a deformable cross-attention mechanism to perform three-dimensional space transformation processing on the binary classification query data determined based on the initial voxel space data as the result of the first-stage training and the two-dimensional image feature data to obtain corresponding three-dimensional query feature data; Use a deformable self-attention mechanism to perform feature fusion processing on the three-dimensional query feature data, the mask token corresponding to the two-dimensional image feature data, and the initial voxel space data to obtain corresponding three-dimensional voxel space features; Perform upsampling and linear mapping on the three-dimensional voxel space features for the voxel space through a preset occupancy prediction head structure processing to obtain semantic voxel map data corresponding to the input sample data; Calculate the second objective loss in the second stage based on the second semantic feature data and the semantic voxel map data, and train the semantic occupancy prediction model based on multi-modal feature fusion based on the second objective loss in the second stage, so as to obtain the final semantic occupancy prediction model.
4. The training method according to claim 2, wherein The first objective loss in the first stage is the semantic scene completion loss associated with the semantic feature map data determined based on the prediction result of the trained semantic occupancy prediction model.
5. The training method according to claim 3, wherein The objective loss function in the training of the second stage is calculated by the following formula: , Among them, represents the target loss function, represents the binary cross-entropy loss function for the preset semantic decoder, represents the geometric scale loss determined based on the semantic voxel map data, represents the semantic scale loss determined based on the semantic voxel map data, represents the semantic scene completion loss determined based on the semantic voxel map data, , , and each represents a preset hyperparameter.
6. A semantic occupancy prediction method based on multi-modal feature fusion, characterized in that, The semantic occupancy prediction method includes: Obtain the RGB image data to be predicted; Input the RGB image data to be predicted into the semantic occupancy prediction model based on multi-modal feature fusion to obtain the corresponding semantic voxel map data. The semantic occupancy prediction model based on multi-modal feature fusion is trained by using the training method described in any one of claims 1 to 5.
7. A training device for a semantic occupancy prediction model based on multi-modal feature fusion, characterized in that, The training device includes: A data acquisition module configured to: acquire input sample data, where the input sample data includes sample RGB image data; A first feature extraction module configured to: input the input sample data into a pre-trained depth estimation model and a pre-trained large visual model respectively to perform depth estimation processing and semantic feature extraction processing respectively, and obtain corresponding depth information data and semantic information data. Among them, the depth information data includes depth feature data, and the semantic information data includes the first semantic feature data of the two-dimensional image and segmentation mask data; A first model training module configured to: use the depth feature data and the first semantic feature data to perform the first stage of training on the semantic occupancy prediction model based on multi-modal feature fusion to obtain the trained semantic occupancy prediction model; A second feature extraction module configured to: input the input sample data into a preset backbone network for feature extraction to obtain corresponding two-dimensional image feature data; A second model training module configured to: use the two-dimensional image feature data to perform the second stage of training on the trained semantic occupancy prediction model to obtain the final semantic occupancy prediction model.
8. A computer program product, characterized in that, The computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, they implement the training method of the semantic occupancy prediction model based on multi-modal feature fusion described in any one of claims 1 to 5 and the semantic occupancy prediction method based on multi-modal feature fusion described in claim 6.
9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the training method of the semantic occupancy prediction model based on multi-modal feature fusion described in any one of claims 1 to 5 and the semantic occupancy prediction method based on multi-modal feature fusion described in claim 6.
10. A computing device, characterized in that, The computing device includes: at least one processor; at least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the training method of the semantic occupancy prediction model based on multi-modal feature fusion as described in any one of claims 1 to 5 and the semantic occupancy prediction method based on multi-modal feature fusion as described in claim 6.
Citation Information
Cited By
Logistics vehicle obstacle sensing method and system based on sparse semantic occupancy network
CN122392027A