Method, device, equipment and storage medium for generating point set occupancy network based on large model constraints

By acquiring segmented images and deep features and utilizing feature splicing and cross-attention mechanisms to enhance feature fusion, the problems of low computational efficiency and insufficient prediction accuracy in existing technologies are solved, efficient and accurate scene voxel generation and occupancy semantic label prediction are achieved, and the adaptability and accuracy of the model are improved.

CN119963574BActive Publication Date: 2025-10-03VOYAH AUTOMOBILE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510028322.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-10-03
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Existing technologies have low computational efficiency in voxel prediction tasks and cannot meet the needs of real-time applications. The generated occupancy prediction performance has poor responsiveness to different scenarios and cannot effectively adapt to changing traffic environments and complex scenarios, resulting in insufficient prediction accuracy and reliability, and an inability to fine-tune voxel reconstruction in different areas.

Method used

Acquire segmented image features and image depth features, adapt the occupied voxel task through feature enhancement, determine the application area and target feature adaptation output information, reconstruct the scene to locate the scene mapping points, generate a pseudo-lidar dataset, use feature splicing and cross-attention mechanism to enhance feature fusion, and improve the accuracy of scene voxel generation and occupied semantic labels.

Benefits of technology

It improves the accuracy and details of scene voxel reconstruction, enhances the generalization and adaptability of the model, enables it to flexibly respond to different scene changes, and significantly improves computational efficiency and prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963574B_ABST
    Figure CN119963574B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device and storage medium for generating a point set occupancy network based on large model constraints, which relates to the field of intelligent driving technology, including: obtaining segmented image features and image depth features; performing feature enhancement and adaptation of occupied voxel tasks based on the segmented image features and the image depth features, determining the application area and target feature adaptation output information; reconstructing the scene based on the application area and the target feature adaptation output information to locate the scene mapping points, determine the simulated LiDAR dataset, and complete scene voxel generation and prediction of occupied semantic labels based on the simulated LiDAR dataset. The present application extracts image segmentation and depth features, adapts to the occupied voxel task and enhances features using feature splicing and cross-attention mechanisms, reconstructs scene mapping points to generate a simulated LiDAR dataset, thereby efficiently and accurately completing scene voxel generation and occupancy semantic label prediction, significantly improving the accuracy and details of scene voxel reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent driving technology, and in particular to a method, apparatus, device, and storage medium for generating a point set occupancy network based on large model constraints. Background Art

[0002] With the continuous development of intelligent driving, more refined scene recognition has become an urgent need. To facilitate path planning, obstacle avoidance, and decision adjustment, it is necessary to accurately identify detailed information on the occupancy status of various objects in the scene, including the object's position, shape, and semantic category information, so as to more accurately predict dense scene voxels. Therefore, in order to construct a detailed three-dimensional scene model, it is urgent to efficiently and accurately generate point sets from image data and predict occupancy networks to achieve more functions.

[0003] Currently, existing approaches rely on the fusion of data from multiple sensors, such as using a cylindrical three-view representation of point clouds to predict 3D semantic occupancy, building a 3D spatial model based on three-plane features, and using pure vision for 3D perception, achieving 3D occupancy prediction through image feature fusion and spatial transformation. For example, point clouds are encoded using pre-trained 2D neural networks and feature aggregation networks to obtain a three-view feature representation, and 3D semantic occupancy prediction is performed using a lightweight prediction head. Alternatively, multi-view camera image features are converted into three-plane features, followed by feature interaction and decoding to achieve 3D semantic occupancy prediction.

[0004] However, existing approaches require the prediction of dense scene voxels in voxel prediction tasks, which leads to low computational efficiency. This is especially true when processing large datasets, as the computational resources consumed are enormous and cannot meet the requirements of real-time or near-real-time applications. Furthermore, the generated occupancy prediction performance is poorly responsive to different scenarios and cannot effectively adapt to changing traffic environments and complex scenes. This results in insufficient prediction accuracy and reliability in practical applications such as autonomous driving and robot navigation. Furthermore, the methods for scene voxel generation and occupancy feature voxel generation are limited, making it impossible to fine-tune voxel reconstruction of different regions, such as the ground, environment, and object areas. This results in insufficient accuracy and detail in the reconstruction results. Therefore, how to more efficiently and accurately generate point set occupancy networks based on large model constraints has become an urgent problem to be solved.

[0005] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of this application is to provide a method, device, equipment and storage medium for generating a point set occupancy network based on a large model constraint, aiming to solve the technical problem of how to generate a point set occupancy network based on a large model constraint more efficiently and accurately.

[0007] To achieve the above objectives, this application proposes a method for generating a point set occupancy network based on large model constraints, the method comprising:

[0008] Obtain segmentation image features and image depth features;

[0009] Performing a feature enhancement and adaptation occupied voxel task based on the segmented image features and the image depth features, and determining application areas and target feature adaptation output information;

[0010] Based on the application area and the target feature adaptation output information, the scene positioning scene mapping point is reconstructed, the pseudo-lidar data set is determined, and scene voxel generation and occupancy semantic label prediction are completed based on the pseudo-lidar data set.

[0011] In one embodiment, the step of obtaining the segmented image features and the image depth features includes:

[0012] Obtain surround view camera data, segmentation model, and depth model;

[0013] Inputting the surround view camera data into the segmentation model to identify the target segmentation subject, extract the segmentation image, and determine the segmentation image features;

[0014] The surround view camera data is input into the depth cutting model to identify the depth of the target body, extract the image depth, and determine the image depth feature.

[0015] In one embodiment, the step of performing the feature lifting and adaptation occupied voxel task based on the segmented image features and the image depth features, and determining the application area and target feature adaptation output information includes:

[0016] Acquire a feature interaction mode, wherein the feature interaction mode includes a cross-attention interaction mode and a feature splicing interaction mode;

[0017] Performing feature splicing on the image feature and the image depth feature based on a feature splicing interaction mode in the feature interaction mode to determine a mixed feature code;

[0018] Based on the cross-attention interaction mode in the feature interaction mode, the image feature, the image depth feature and the mixed feature encoding are subjected to a feature enhancement adaptation occupied voxel task to obtain target feature adaptation output information, wherein the target feature adaptation output information includes segmentation feature adaptation output information and image depth feature adaptation output information;

[0019] The target area is segmented based on the segmented image features, the image depth features and the mixed features to obtain an application area, where the application area includes a ground area, an environment area and an object area.

[0020] In one embodiment, the step of reconstructing the scene positioning scene mapping points based on the application area and the target feature adaptation output information and determining the pseudo-lidar data set includes:

[0021] Reconstructing the scene based on the application area and the target feature adaptation output information to locate the scene mapping point, and determining the occupied voxel mapping point set and the occupied semantic label;

[0022] A scene representation is performed based on the occupied voxel mapping point set and the occupied semantic label to construct pseudo-lidar data.

[0023] In one embodiment, the steps of reconstructing the scene based on the application area and the target feature adaptation output information to locate the scene mapping point and determining the occupied voxel mapping point set and the occupied semantic label include:

[0024] Obtaining pixel data, the pixel data including pixel position, pixel center area, and camera hardware focal length;

[0025] generating an occupancy semantic label based on the application area, and selecting a scene reconstruction method, wherein the scene reconstruction method includes a ground area reconstruction method, an environment area reconstruction method, and an object area reconstruction method;

[0026] The scene mapping points are located based on the scene reconstruction method, the target feature adaptation output information and the pixel data reconstructed scene, and an occupied voxel mapping point set is calculated.

[0027] In one embodiment, the steps of locating scene mapping points based on the scene reconstruction method, the target feature adaptation output information, and the pixel data reconstructed scene, and calculating an occupied voxel mapping point set include:

[0028] When the scene reconstruction method is a ground area reconstruction method or the scene reconstruction method is an environment area reconstruction method, a mapping point set is calculated based on the image depth feature adaptation output information in the target feature adaptation output information and the pixel data, the mapping point set including a ground area mapping point set and an environment area mapping point set;

[0029] When the scene reconstruction method is an object region reconstruction method, the potential occluded objects are supplemented based on the pixel data, the segmentation feature adaptation output information, and the image depth feature adaptation output information in the target feature adaptation output information, and a supplemented mapping point set is calculated;

[0030] An occupied voxel mapping point set is obtained based on the mapping point set and the completed mapping point set.

[0031] In one embodiment, the step of constructing simulated LiDAR data by performing scene representation based on the occupied voxel mapping point set and the occupied semantic label includes:

[0032] Performing voxelization processing on the occupied voxel mapping point set to construct a scene and determine scene data;

[0033] Semantic categories are bound based on the scene data and the occupancy semantic labels to obtain simulated lidar data.

[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a point set occupancy network generation device based on a large model constraint, the point set occupancy network generation device based on a large model constraint comprising:

[0035] An acquisition module, used to obtain segmentation image features and image depth features;

[0036] A processing module, configured to perform a feature enhancement and adaptation occupied voxel task based on the segmented image features and the image depth features, and determine application areas and target feature adaptation output information;

[0037] An execution module is used to reconstruct the scene positioning scene mapping points based on the application area and the target feature adaptation output information, determine the simulated lidar data set, and complete scene voxel generation and predicted occupancy semantic labels based on the simulated lidar data set.

[0038] In addition, to achieve the above-mentioned purpose, the present application also proposes a point set occupancy network generation device based on large model constraints, the device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the point set occupancy network generation method based on large model constraints as described above.

[0039] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the point set occupancy network generation method based on large model constraints as described above are implemented.

[0040] One or more technical solutions proposed in this application have at least the following technical effects:

[0041] This embodiment proposes a point set occupancy network generation method based on large model constraints, which obtains segmented image features and image depth features; performs feature enhancement and adaptation of occupied voxel tasks based on the segmented image features and the image depth features, determines the application area and target feature adaptation output information; reconstructs the scene positioning scene mapping points based on the application area and the target feature adaptation output information, determines the simulated LiDAR dataset, and completes scene voxel generation and prediction of occupied semantic labels based on the simulated LiDAR dataset. This application improves the effectiveness of feature fusion by obtaining segmented image features and image depth features, adapting to the occupied voxel task, and enhancing features using feature splicing and cross-attention mechanisms, reconstructing scene mapping points to generate a simulated LiDAR dataset, thereby efficiently and accurately completing scene voxel generation and occupied semantic label prediction, improving the accuracy and details of scene voxel reconstruction, and at the same time, can flexibly respond to different scene changes, significantly improving the generalization ability and adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0044] Figure 1 A flowchart of the first embodiment of the method for generating a point set occupancy network based on large model constraints provided by this application;

[0045] Figure 2 A flowchart diagram of the second embodiment of the method for generating a point set occupancy network based on large model constraints provided by this application;

[0046] Figure 3 This is a schematic diagram of object occlusion in the point set occupancy network generation method based on large model constraints in this application;

[0047] Figure 4 A schematic diagram of a simplified flow chart of a method for generating a point set occupancy network based on a large model constraint provided in an embodiment of the present application;

[0048] Figure 5 This is a schematic diagram of the module structure of a point set occupancy network generation device based on large model constraints according to an embodiment of the present application;

[0049] Figure 6Schematic diagram of the device structure of the hardware operating environment involved in the point set occupancy network generation method based on large model constraints in an embodiment of the present application.

[0050] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0051] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0052] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0053] The main solutions of the embodiments of the present application are: obtaining segmented image features and image depth features; performing feature enhancement and adaptation of occupied voxels based on the segmented image features and the image depth features, and determining the application area and target feature adaptation output information; reconstructing the scene positioning scene mapping points based on the application area and the target feature adaptation output information, determining the simulated lidar data set, and completing scene voxel generation and predicting occupancy semantic labels based on the simulated lidar data set.

[0054] In this embodiment, for ease of description, the following description is made by taking the identification of a point set occupancy network generation device based on a large model constraint as the execution subject.

[0055] Since existing technologies need to predict dense scene voxels in voxel prediction tasks, this leads to low computational efficiency. Especially when processing large-scale data sets, the computing resources consumed are huge and cannot meet the real-time or near-real-time application requirements. In addition, the generated occupancy prediction performance has poor responsiveness to different scenarios and cannot effectively adapt to changing traffic environments and complex scenarios. As a result, in practical applications such as autonomous driving and robot navigation, the prediction accuracy and reliability are insufficient. At the same time, the methods for scene voxel generation and occupancy feature voxel generation are single and cannot finely handle voxel reconstruction in different areas, such as the ground, environment and object areas, resulting in insufficient accuracy and details of the reconstruction results.

[0056] The present application provides a solution to obtain segmented image features and image depth features; perform feature enhancement and adaptation of occupied voxels based on the segmented image features and the image depth features, and determine the application area and target feature adaptation output information; reconstruct the scene positioning scene mapping points based on the application area and the target feature adaptation output information, determine the simulated LiDAR data set, and complete scene voxel generation and predict occupancy semantic labels based on the simulated LiDAR data set.

[0057] It can be seen from the above embodiments that the present application improves the effectiveness of feature fusion by acquiring segmentation image features and image depth features, adapting to the occupied voxel task, and enhancing features using feature splicing and cross-attention mechanisms, reconstructing scene mapping points to generate a pseudo-lidar dataset, thereby efficiently and accurately completing scene voxel generation and occupancy semantic label prediction, improving the accuracy and details of scene voxel reconstruction, and at the same time, can flexibly respond to different scene changes, significantly improving the generalization ability and adaptability of the model.

[0058] Based on this, the embodiment of the present application provides a method for generating a point set occupancy network based on a large model constraint, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the method for generating a point set occupancy network based on large model constraints of this application.

[0059] In this embodiment, the method for generating a point set occupancy network based on large model constraints includes steps S10 to S30:

[0060] Step S10, obtaining segmented image features and image depth features;

[0061] It should be noted that the segmented image features reflect the features of semantic information of different objects and scenes segmented in the image, and the image depth features reflect the features of depth information of each pixel in the image relative to the camera.

[0062] It can be understood that based on the segmented image features and image depth features, objects in the scene can be identified and located more accurately, the accuracy of scene parsing can be significantly improved, the three-dimensional effect of scene reconstruction can be enhanced, and the features that integrate segmentation and depth information can adapt to more complex and changeable scenes, thereby improving the applicability of the model in different environments.

[0063] For ease of understanding, the following description is given by taking the acquisition of segmented image features and image depth features as an example, wherein the information acquisition device is an information acquisition module, and the storage device is a memory.

[0064] The information acquisition module obtains surround-view camera data, segmentation cutting model and depth cutting model, that is, uses surround-view cameras to capture images. For example, 6 surround-view cameras are set to collect surround-view camera data, obtain SAM segmentation cutting model and DepthAnything depth cutting model, input the surround-view camera data into the segmentation cutting model to identify the target segmentation subject, extract the segmented image, determine the segmented image features, input the surround-view camera data into the depth cutting model to identify the target subject depth, extract the image depth, and determine the image depth features. That is, the SAM segmentation cutting model extracts segmentation image features from the images captured by the surround-view camera and encodes them through the feature encoding network. At this time, the segmented image features are represented as:

[0065] FS =E S (I)

[0066] The DepthAnything depth cutting model extracts depth features from the images captured by the surround camera and encodes them through the feature encoding network. At this time, the image depth features are represented as:

[0067] F D =E D (I)

[0068] Among them, E S and E D They represent the feature encoding networks of SAM and Depth models respectively, and I is the input image.

[0069] Subsequent processing is performed based on the segmented image features and the image depth features.

[0070] In a feasible implementation, step S10 may include steps A11 to A13:

[0071] Step A11, obtaining surround view camera data, segmentation model and depth model;

[0072] It should be noted that the surround view camera data reflects the characteristics of the all-round visual information of the vehicle's surrounding environment, the segmentation model reflects the characteristics of the data set for identifying and segmenting the target subject in the image, and the depth cutting model reflects the characteristics of the data set for identifying the distance between the target subject and the camera.

[0073] Step A12: inputting the surround view camera data into the segmentation model to identify the target segmentation subject, extract the segmentation image, and determine the segmentation image features;

[0074] It can be understood that the segmented image features may include segmented object image contours, segmented object shapes, and segmented object categories, which are extracted from images captured by the surround view camera through the segmentation model and used to identify and distinguish different targets and backgrounds in the image.

[0075] Step A13: input the surround view camera data into the depth cutting model to identify the depth of the target body, extract the image depth, and determine the image depth feature.

[0076] It can be understood that the image depth feature can represent the distance between the object and the camera. By extracting it from the surround camera data through the depth cutting model, the scene reconstruction is no longer limited to the two-dimensional plane, but can construct a more realistic three-dimensional space model.

[0077] Step S20, performing a feature enhancement and adaptation occupied voxel task based on the segmented image features and the image depth features, and determining an application area and target feature adaptation output information;

[0078] It should be noted that the application area reflects the characteristics of the areas divided by different subjects in the identified scene, and the target feature adaptation output information reflects the characteristics of the data adapted to different environments after feature enhancement.

[0079] It is understandable that through refined area division and targeted feature adaptation output information, scene reconstruction can be made more accurate, and the accuracy of scene parsing can be significantly improved. It can also be applied to different application areas, enabling the model to have a deeper understanding of the various components of the scene, strengthening the model's understanding of the scene, and at the same time, reducing unnecessary calculations, making the entire processing flow more efficient.

[0080] For ease of understanding, the following description is made by taking the determination of application area and target feature adaptation output information as an example, wherein the information acquisition device is an information acquisition module, the storage device is a memory, and the processing device is a processing module.

[0081] The information acquisition module obtains the segmented image features F S and image depth feature F D , obtaining a feature interaction mode, the feature interaction mode including a cross-attention interaction mode and a feature splicing interaction mode, performing feature splicing on the image features and the image depth features based on the feature splicing interaction mode in the feature interaction mode, and determining a hybrid feature code, that is, forming a comprehensive feature representation by feature splicing the extracted segmented image features and image depth features, and obtaining a hybrid feature code, which is expressed as:

[0082] F concat =Concat(F S ,F D )

[0083] Among them, Concat represents the feature concatenation operation.

[0084] Based on the cross attention interaction mode in the feature interaction mode, the image feature, the image depth feature and the hybrid feature encoding are subjected to a feature enhancement adaptation voxel occupation task to obtain target feature adaptation output information, which includes segmentation feature adaptation output information and image depth feature adaptation output information. That is, a multi-head cross attention mechanism is introduced to calculate the adaptation feature output, which is expressed as:

[0085] F adapter =MultiHeadCrossAttention(F S / F D ,F concat )

[0086] Among them, MultiHeadCrossAttention represents the cross-attention operation, which can dynamically adjust the attention weight according to the importance of the feature.

[0087] At this time, F adapter Using multi-head cross attention, the features of DepthAnything with good SAM features are cross-fused, so F adapter After obtaining the common properties of the two, Sigmoid can be used to activate the areas that both of them pay more attention to, so as to retain them.

[0088] Sigmoid represents the Sigmoid calculation function. The formula of sigmoid is expressed as:

[0089]

[0090] From this, we can use sigmoid activation to normalize feature inputs to the range [0, 1]. The resulting data can be adapted for both SAM and DepthAnything. Sigmoid preserves the features required for autonomous driving, unlike SAM and DepthAnything, which are derived by training networks on a wide variety of global data. This ensures that the resulting adapted features are better suited for autonomous driving scenarios.

[0091] Then calculate the target feature adaptation output information. In order to better apply it to scene segmentation and depth estimation, design an adaptive feature matching and output mechanism, that is, the adaptation network performs sigmoid calculation on the adaptation feature output, and then multiplies the attention features obtained by sigmoid with the SAM and DepthAnything intermediate features respectively to obtain attention enhancement features. Then, this part of the features is calculated using depth-wise separable convolution and activation function to obtain the final feature adaptation output, that is, the target feature adaptation output.

[0092] Calculate the segmentation feature adaptation output information in the target feature adaptation output information, expressed as:

[0093]

[0094] Calculate the image depth feature adaptation output information in the target feature adaptation output information, expressed as:

[0095]

[0096] in, represents the final output of the SAM network, Represents the final output of the DepthAnything network. GELU represents the GELU activation function. DWConv stands for depthwise separable convolution. Similar to sigmoid, GELU is used to make the adapted features more suitable for autonomous driving.

[0097] The target area is segmented based on the segmented image features, the image depth features and the mixed feature encoding to obtain an application area, which includes a ground area, an environment area and an object area. That is, a common general segmentation and depth estimation method is used to obtain a segmentation result and a depth estimation result, and the segmentation result and the depth estimation result are applied to different areas, such as a ground area, an environment area and an object area, to obtain an application area.

[0098] The output information is adapted for subsequent processing based on the application area and target features.

[0099] In a feasible implementation, step S20 may include steps B11 to B14:

[0100] Step B11, obtaining a feature interaction mode, wherein the feature interaction mode includes a cross-attention interaction mode and a feature splicing interaction mode;

[0101] It should be noted that the feature interaction method reflects the features of processing and fusing the extracted different feature data.

[0102] It is understandable that the cross-attention interaction method focuses on dynamically adjusting the attention weight according to the importance of the feature, while the feature splicing interaction method focuses on directly merging features of different modalities to form a hybrid feature encoding. Through the combination of feature splicing and cross-attention mechanism, it can more effectively integrate features from different modes, enhance the expressive power of features, and better adapt to different scenarios and conditions, thereby improving the prediction accuracy and reliability of the model.

[0103] Step B12, performing feature splicing on the image feature and the image depth feature based on the feature splicing interaction mode in the feature interaction mode, and determining a mixed feature code;

[0104] It should be noted that the hybrid feature encoding reflects the characteristics of the rich integrated information obtained by performing feature splicing operations on features.

[0105] It can be understood that by integrating the segmentation image features and image depth features to obtain the hybrid feature encoding, the features can be represented more comprehensively, the model's ability to understand the scene can be enhanced, it can better adapt to different environments and conditions, and the model's prediction accuracy and reliability in changing traffic environments and complex scenarios can be improved.

[0106] Step B13, performing a feature enhancement adaptation occupied voxel task on the image features, the image depth features, and the hybrid feature encoding based on the cross-attention interaction mode in the feature interaction mode, to obtain target feature adaptation output information, wherein the target feature adaptation output information includes segmentation feature adaptation output information and image depth feature adaptation output information;

[0107] It can be understood that by determining the target feature adaptation output information to improve the accuracy of segmentation and depth estimation, after feature fusion, the model can further integrate features from different models through feature splicing, and the adaptation feature output module dynamically adjusts the feature matching strategy according to the specific content and scene of the input image.

[0108] Step B14: encoding and segmenting a target area based on the segmented image features, the image depth features, and the mixed features to obtain an application area, wherein the application area includes a ground area, an environment area, and an object area.

[0109] It's clear that the model is able to identify key areas in the image and optimize the feature representation accordingly to improve the accuracy of the segmentation results. The depth estimation output module also uses a similar approach to ensure the accuracy of depth information. This adaptive mechanism enables the model to flexibly respond to different image features, improving its generalization and adaptability, thereby more accurately segmenting the application area.

[0110] Step S30: reconstructing the scene positioning scene mapping points based on the application area and the target feature adaptation output information, determining a pseudo-lidar data set, and completing scene voxel generation and predicting occupancy semantic labels based on the pseudo-lidar data set.

[0111] It should be noted that the pseudo lidar dataset reflects the features of the scene localization scene mapping point set reconstructed based on the application area and target feature adaptation output information.

[0112] It is understandable that the pseudo lidar dataset may include occupied voxel mapping point sets and occupied semantic labels, which can provide detailed spatial information for three-dimensional scene reconstruction and understanding, while providing accurate environmental perception capabilities for applications such as autonomous driving and robot navigation, providing high-density and accurate spatial information, making three-dimensional scene reconstruction and understanding more accurate, and enhancing the model's ability to parse complex scenes.

[0113] For ease of understanding, the following is explained by taking the determination of the simulated lidar data set as an example, wherein the information acquisition device is the information acquisition module, the storage device is the memory, and the execution device is the execution module.

[0114] The information acquisition module obtains the application area, such as the ground area, the environment area and the object area, obtains the target feature adaptation output information, reconstructs the scene positioning scene mapping point based on the application area and the target feature adaptation output information, determines the occupied voxel mapping point set and the occupied semantic label, and performs scene representation based on the occupied voxel mapping point set and the occupied semantic label to construct the pseudo-lidar data, that is, based on the application area, 3D mapping points of three parts, such as the ground mapping point, the environment mapping point and the object mapping point, so as to more efficiently represent the occupancy estimation result, directly voxelize the mapped 3D points, and at the same time, for the semantic category of each voxel area, use the semantic result of the segmented area above to bind it, which can more quickly and efficiently predict the semantic label of the occupancy semantics.

[0115] This embodiment proposes a point set occupancy network generation method based on large model constraints, which obtains segmented image features and image depth features; performs feature enhancement and adaptation of occupied voxels based on the segmented image features and the image depth features, and determines the application area and target feature adaptation output information; reconstructs the scene positioning scene mapping points based on the application area and the target feature adaptation output information, determines a simulated lidar data set, and completes scene voxel generation and predicts occupancy semantic labels based on the simulated lidar data set. The technical problem of how to more efficiently generate point set occupancy networks based on large model constraints is solved. Compared with the existing technology, this application obtains segmented image features and image depth features, uses features to enhance the adaptive occupancy voxel task, and reconstructs scene positioning scene mapping points based on the output information of application area and target feature adaptation, which significantly improves the effectiveness of feature fusion. By optimizing the multimodal feature fusion and enhancement system, the amount of calculation required to predict dense scene voxels is reduced, and the computational efficiency is significantly improved. By introducing the cross-attention enhancement mechanism, the key features are strengthened and the secondary features are suppressed, making feature fusion more effective and improving the accuracy and robustness of feature expression. By fine-tuning the reconstruction of the ground, environment and object areas, the accuracy and details of the scene voxel reconstruction are improved, providing richer information for three-dimensional scene reconstruction and understanding, and being able to flexibly respond to different image features and scene changes, improving the generalization ability and adaptability of the model.

[0116] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated later.

[0117] In this embodiment, refer to Figure 2 , Figure 2 This is a flow chart of Example 2 of the method for generating a point set occupancy network based on large model constraints of this application. Step S30 specifically includes steps S31 to S32:

[0118] Step S31, reconstructing the scene based on the application area and the target feature adaptation output information to locate the scene mapping point, and determining the occupied voxel mapping point set and the occupied semantic label;

[0119] It should be noted that the occupied voxel mapping point set reflects the characteristics of the positioning points in the scene accurately reconstructed by applying the region and target feature adaptation output information, and the occupied semantic label reflects the characteristics of the semantic category information of each labeled voxel.

[0120] For ease of understanding, the following description is made by taking the determination of occupied voxel mapping point sets and occupied semantic labels as an example, wherein the information acquisition device is the information acquisition module, the storage device is the memory, and the execution device is the execution module.

[0121] The information acquisition module obtains pixel data. Since the surround view camera is composed of 6 different cameras, it is necessary to use the calibration parameters of different cameras separately to obtain pixel data. The pixel data includes pixel position (u, v), pixel center area c x and c y and the camera hardware focal length f x and f y , obtain an application area, and generate a possession semantic label based on the application area.

[0122] A scene reconstruction method is selected, wherein the scene reconstruction method includes a ground area reconstruction method, an environment area reconstruction method, and an object area reconstruction method. When the scene reconstruction method is the ground area reconstruction method or the scene reconstruction method is the environment area reconstruction method, a mapping point set is calculated based on the image depth feature adaptation output information in the target feature adaptation output information and the pixel data. The mapping point set includes a ground area mapping point set and an environment area mapping point set. That is, when the scene reconstruction method is the ground area reconstruction method or the scene reconstruction method is the environment area reconstruction method, pixels and depth can be used to estimate data to convert pixel point data into point data in 3D space.

[0123] The image depth feature adaptation output information in the target feature adaptation output information is expressed as:

[0124] D I ∈R H×W

[0125] Where R is a set of real numbers, H is the depth, and W is the width.

[0126] By converting pixel point data into point data in 3D space, we can obtain a ground area mapping point set or an environment area mapping point set, thereby obtaining a mapping point set, which is expressed as:

[0127]

[0128] z=D(x,y) (u,v)

[0129] Among them, (u, v) represents the pixel position, c x and c y Represents the pixel center area of ​​the camera, this data is determined by the camera's hardware sensor. x and f y Represents the focal length of the camera hardware.

[0130] This allows for quick point mapping of the ground and the environment, such as trees, lawns, and sidewalks. However, there are certain occlusions and problems with objects, such as Figure 3 As shown, Figure 3 This is a diagram of object occlusion in the point set occupancy network generation method based on large model constraints in this application. The gray parallelogram represents the cross-section of the foreground and background of the object as seen by the ego vehicle's camera. In other words, the left side of the cross-section represents the background area of ​​the object, and the right side of the cross-section represents the foreground area of ​​the object. We cannot directly see the background area because this part of the area is in the blind spot of the ego vehicle's camera, and only the foreground area can be seen.

[0131] Therefore, when the scene reconstruction method is an object region reconstruction method, the potential occluded objects are complemented based on the pixel data, the segmentation feature adaptation output information, and the image depth feature adaptation output information in the target feature adaptation output information, and the complement mapping point set is calculated, that is, the image depth feature adaptation output information in the target feature adaptation output information is still expressed as:

[0132] D I ∈R H×W

[0133] Where R is a set of real numbers, H is the depth, and W is the width.

[0134] Based on the segmentation feature adaptation output information in the target feature adaptation output information, calculate the object segmentation area D S The maximum and minimum values ​​of , and then calculate the uniform sampling between the minimum and maximum depths, and set the uniform sampling interval to Δ.

[0135] Through calculation, we can obtain richer depth data of the object segmentation area. Expressed as:

[0136]

[0137] Among them, Max and Min represent the calculation methods of maximum and minimum values.

[0138] After obtaining more depth data, the 3D points of the object area can be recalculated according to the ground area reconstruction method and the environment area reconstruction method to enrich the point information of the object area and obtain a complete mapping point set.

[0139] An occupied voxel mapping point set is obtained based on the mapping point set and the completed mapping point set, and subsequent processing is performed based on the occupied voxel mapping point set and the occupied semantic label.

[0140] In a feasible implementation, step S31 may include steps C11 to C13:

[0141] Step C11, obtaining pixel data, wherein the pixel data includes pixel position, pixel center area, and camera hardware focal length;

[0142] It should be noted that the pixel data reflects the characteristics of a data set consisting of the parameters of the surround view camera itself and the information of the collected objects.

[0143] It can be understood that by accurately obtaining pixel data, two-dimensional image data can be more accurately converted into three-dimensional spatial information, thereby improving the accuracy of scene reconstruction, and more accurately identifying and classifying objects in the image, improving the depth and accuracy of semantic understanding, and at the same time, providing more reliable environmental perception and reducing the risks caused by inaccurate data.

[0144] Step C12: generating an occupancy semantic label based on the application area and selecting a scene reconstruction method, wherein the scene reconstruction method includes a ground area reconstruction method, an environment area reconstruction method, and an object area reconstruction method;

[0145] It should be noted that the scene reconstruction method reflects the refinement of features of each area in the scene by selecting and applying different scene reconstruction methods.

[0146] It is understandable that using specialized reconstruction methods for different areas can more accurately reconstruct the ground, environment and object areas in the scene, improve the accuracy and details of scene reconstruction, reduce unnecessary calculations, increase processing speed, adapt to a variety of environments and conditions, and significantly improve the accuracy of scene reconstruction.

[0147] Step C13: reconstructing the scene based on the scene reconstruction method, the target feature adaptation output information and the pixel data to locate scene mapping points and calculate an occupied voxel mapping point set.

[0148] It can be understood that the occupancy voxel mapping point set can represent the position and spatial distribution of objects in three-dimensional space, making scene reconstruction more accurate and enhancing the model's ability to parse complex scenes. The occupancy semantic label can include object geometry information and object category information, such as vehicle, pedestrian and road information, providing rich contextual information for the semantic understanding of the scene.

[0149] In a feasible implementation, step C13 may include steps D11 to D13:

[0150] Step D11, when the scene reconstruction method is a ground area reconstruction method or the scene reconstruction method is an environment area reconstruction method, calculating a mapping point set based on the image depth feature adaptation output information in the target feature adaptation output information and the pixel data, the mapping point set including a ground area mapping point set and an environment area mapping point set;

[0151] It should be noted that the mapping point set reflects the characteristics of the three-dimensional coordinates of the target identification terrain in the ground or environmental area reconstruction scene.

[0152] It is understandable that the use of the mapping point set can finely reconstruct the scene, provide accurate three-dimensional spatial information, make the scene reconstruction more accurate, and enhance the model's ability to analyze complex scenes.

[0153] Step D12, when the scene reconstruction method is an object region reconstruction method, complementing potential occluded objects based on the pixel data, the segmentation feature adaptation output information, and the image depth feature adaptation output information in the target feature adaptation output information, and calculating a complement mapping point set;

[0154] It should be noted that the completion mapping point set reflects the characteristics of the three-dimensional coordinates constructed by completing the target object in the object region reconstruction scene.

[0155] It is understandable that the use of the mapping point set can better fill the background area and construct the obstructed object part, making the scene reconstruction more accurate.

[0156] Step D13: obtaining an occupied voxel mapping point set based on the mapping point set and the completed mapping point set.

[0157] It can be understood that the use of the mapping point set and the supplementary mapping point set can adapt to various scenes, integrate information from different areas to form a complete three-dimensional scene, improve the integrity and accuracy of scene reconstruction, and enhance environmental perception capabilities when processing occluded and invisible areas, providing more accurate and reliable three-dimensional spatial information.

[0158] Step S32: constructing simulated LiDAR data by performing scene representation based on the occupied voxel mapping point set and the occupied semantic label.

[0159] As you can see, in each region, the model uses pixel values ​​and depth estimation information to construct simulated LiDAR data for continuous depth estimation, generating occupancy feature voxels. Occupancy feature voxels are voxels in three-dimensional space that represent the occupancy status of objects in the scene, providing more reliable environmental perception capabilities for applications such as autonomous driving and robotic navigation.

[0160] For ease of understanding, the construction of simulated lidar data is taken as an example for explanation, where the information collection device is the information collection module, the storage device is the memory, and the execution device is the execution module.

[0161] The information acquisition module obtains the occupied voxel mapping point set, that is, obtains the 3D mapping points of the ground area, the 3D mapping points of the environment area and the 3D mapping points of the object area, obtains the occupied semantic label, performs voxel processing on the occupied voxel mapping point set to construct a scene, determines the scene data, binds the semantic category based on the scene data and the occupied semantic label, obtains the simulated lidar data, that is, the 3D mapping points of the ground area, the 3D mapping points of the environment area and the 3D mapping points of the object area, in order to more efficiently represent the occupancy estimation result, directly voxelizes the 3D mapping points of the ground area, the 3D mapping points of the environment area and the 3D mapping points of the object area, and for the semantic category of each voxel area, we use the segmentation area F S The occupancy semantic labels are bound to the image, and the obtained image dense semantic output can predict the semantic labels of occupancy semantics more quickly and efficiently, thereby constructing pseudo-LiDAR data, and completing scene voxel generation and prediction of occupancy semantic labels based on the pseudo-LiDAR data set.

[0162] In a feasible implementation, step S32 may include steps E11 to E12:

[0163] Step E11, performing voxelization processing on the occupied voxel mapping point set to construct a scene and determine scene data;

[0164] It should be noted that the scene data reflects the characteristics of a data set that converts an occupied voxel mapping point set into a detailed three-dimensional scene representation for practical application.

[0165] It is understandable that voxelizing point sets to construct scenes improves the accuracy and completeness of scene representation, making the reconstruction of three-dimensional scenes more detailed and accurate, and providing richer and more structured scene information for various applications.

[0166] Step E12: Binding semantic categories based on the scene data and the occupancy semantic labels to obtain simulated LiDAR data.

[0167] It is understandable that after the three-dimensional scene is reconstructed, specific semantic categories are assigned to each voxel in the scene, thereby converting the dataset into scene understanding data with clear physical and semantic meanings, improving the accuracy of scene understanding, and thus identifying the location, nature and category of objects, and enhancing the application value of the data, and optimizing resource utilization and processing efficiency through precise semantic labeling.

[0168] This embodiment proposes a method for generating a point set occupancy network based on large model constraints. Based on the application area and the target feature adaptation output information, the method reconstructs the scene, locates the scene mapping points, determines the occupied voxel mapping point set and the occupied semantic labels, and constructs simulated LiDAR data based on the occupied voxel mapping point set and the occupied semantic labels. This method solves the technical problem of how to more accurately generate a point set occupancy network based on large model constraints. Compared to the existing technology, the technical solution of this application significantly improves the accuracy and detail of scene reconstruction by reconstructing the scene to locate the scene mapping points, and determining the occupied voxel mapping point set and the occupied semantic labels, and then constructing simulated LiDAR data. This enhances the model's adaptability and generalization capabilities for different scenarios, and improves the accuracy and reliability of predictions, meeting the needs of real-time or near-real-time applications and providing more accurate environmental perception capabilities.

[0169] For example, in order to help understand the implementation process of the point set occupancy network generation method based on large model constraints obtained by combining this embodiment with the above embodiment 1, please refer to Figure 4 , Figure 4 This paper provides a brief flowchart of a method for generating a point set occupancy network based on large model constraints. Specifically:

[0170] Referring to Example 1, segmented image features and image depth features are obtained; based on the segmented image features and the image depth features, a feature enhancement and adaptation occupancy voxel task is performed to determine the application area and target feature adaptation output information; based on the application area and the target feature adaptation output information, a scene positioning scene mapping point is reconstructed, a pseudo-LiDAR dataset is determined, and scene voxel generation and occupancy semantic label prediction are completed based on the pseudo-LiDAR dataset. Referring to Example 2, a scene positioning scene mapping point is reconstructed based on the application area and the target feature adaptation output information to determine the occupancy voxel mapping point set and occupancy semantic label; based on the occupancy voxel mapping point set and the occupancy semantic label, a scene representation is performed to construct pseudo-LiDAR data. After the surround-view camera collects data, it inputs the SAM segmentation-cutting model and the DepthAnything depth-cutting model to extract image features S and image features D. Feature splicing is performed to obtain hybrid feature encoding, and then cross-attention enhancement and feature splicing are performed to obtain adapted feature output, thereby calculating the segmentation result output and depth estimation result output. The ground area, environmental area and object area are divided. Pixel values ​​and depth estimates are selected in the ground area and environmental area to construct pseudo-lidar data. The object area first performs continuous depth estimation of the segmented area, and then pixel values ​​and depth estimation are performed to construct pseudo-lidar data, thereby completing scene voxel generation and occupancy prediction.

[0171] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the point set occupancy network generation method based on large model constraints of the present application. More forms of simple transformations based on this technical concept are all within the scope of protection of the present application.

[0172] This application also provides a point set occupancy network generation device based on large model constraints, please refer to Figure 5 The point set occupancy network generation device based on large model constraints includes:

[0173] An acquisition module 10 is used to acquire segmented image features and image depth features;

[0174] A processing module 20 is configured to perform a feature enhancement and adaptation occupied voxel task based on the segmented image features and the image depth features, and determine an application area and target feature adaptation output information;

[0175] The execution module 30 is used to reconstruct the scene positioning scene mapping points based on the application area and the target feature adaptation output information, determine the pseudo-lidar data set, and complete scene voxel generation and predicted occupancy semantic labels based on the pseudo-lidar data set.

[0176] The acquisition module 10 is further used to acquire surround view camera data, segmentation cut model and depth cut model;

[0177] Inputting the surround view camera data into the segmentation model to identify the target segmentation subject, extract the segmentation image, and determine the segmentation image features;

[0178] The surround view camera data is input into the depth cutting model to identify the depth of the target body, extract the image depth, and determine the image depth feature.

[0179] The processing module 20 is further configured to obtain a feature interaction mode, wherein the feature interaction mode includes a cross-attention interaction mode and a feature splicing interaction mode;

[0180] Performing feature splicing on the image feature and the image depth feature based on a feature splicing interaction mode in the feature interaction mode to determine a mixed feature code;

[0181] Based on the cross-attention interaction mode in the feature interaction mode, the image feature, the image depth feature and the mixed feature encoding are subjected to a feature enhancement adaptation occupied voxel task to obtain target feature adaptation output information, wherein the target feature adaptation output information includes segmentation feature adaptation output information and image depth feature adaptation output information;

[0182] The target area is segmented based on the segmented image features, the image depth features and the mixed features to obtain an application area, where the application area includes a ground area, an environment area and an object area.

[0183] The execution module 30 is further configured to reconstruct the scene positioning scene mapping points based on the application area and the target feature adaptation output information, and determine the occupied voxel mapping point set and occupied semantic labels;

[0184] A scene representation is performed based on the occupied voxel mapping point set and the occupied semantic label to construct pseudo-lidar data.

[0185] The execution module 30 is further configured to obtain pixel data, wherein the pixel data includes pixel position, pixel center area, and camera hardware focal length;

[0186] generating an occupancy semantic label based on the application area, and selecting a scene reconstruction method, wherein the scene reconstruction method includes a ground area reconstruction method, an environment area reconstruction method, and an object area reconstruction method;

[0187] The scene mapping points are located based on the scene reconstruction method, the target feature adaptation output information and the pixel data reconstructed scene, and an occupied voxel mapping point set is calculated.

[0188] The execution module 30 is further configured to calculate a mapping point set based on the image depth feature adaptation output information in the target feature adaptation output information and the pixel data when the scene reconstruction method is a ground area reconstruction method or the scene reconstruction method is an environment area reconstruction method, wherein the mapping point set includes a ground area mapping point set and an environment area mapping point set;

[0189] When the scene reconstruction method is an object region reconstruction method, the potential occluded objects are supplemented based on the pixel data, the segmentation feature adaptation output information, and the image depth feature adaptation output information in the target feature adaptation output information, and a supplemented mapping point set is calculated;

[0190] An occupied voxel mapping point set is obtained based on the mapping point set and the completed mapping point set.

[0191] The execution module 30 is further configured to perform voxelization processing on the occupied voxel mapping point set to construct a scene and determine scene data;

[0192] Semantic categories are bound based on the scene data and the occupancy semantic labels to obtain simulated lidar data.

[0193] The device for generating a point set occupancy network based on large model constraints provided in this application utilizes the method for generating a point set occupancy network based on large model constraints described in the aforementioned embodiments, thereby solving the technical problem of more efficiently and accurately generating a point set occupancy network based on large model constraints. Compared to the prior art, the device for generating a point set occupancy network based on large model constraints provided in this application has the same beneficial effects as the method for generating a point set occupancy network based on large model constraints described in the aforementioned embodiments. Other technical features of the device for generating a point set occupancy network based on large model constraints are the same as those disclosed in the aforementioned embodiments and are not further elaborated upon here.

[0194] The present application provides a point set occupancy network generation device based on a large model constraint, the point set occupancy network generation device based on a large model constraint comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the point set occupancy network generation method based on the large model constraint in the above-mentioned embodiment 1.

[0195] Reference below Figure 6, which shows a schematic diagram of the structure of a device for generating a point set occupancy network based on large model constraints suitable for implementing embodiments of the present application. The device for generating a point set occupancy network based on large model constraints in embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The point set occupancy network generation device based on large model constraints shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0196] like Figure 6 As shown, the device for generating a point set occupancy network based on large model constraints may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for the operation of the device for generating a point set occupancy network based on large model constraints are also stored in RAM 1004. The processing device 1001, ROM 1002, and RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the large model constraint-based point set occupancy network generation device to communicate wirelessly or wired with other devices to exchange data. Although the diagram shows a large model constraint-based point set occupancy network generation device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have instead.

[0197] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0198] The device for generating a point set occupancy network based on large model constraints provided by this application utilizes the method for generating a point set occupancy network based on large model constraints described in the aforementioned embodiment, thereby solving the technical problem of more efficiently and accurately generating a point set occupancy network based on large model constraints. Compared to the prior art, the device for generating a point set occupancy network based on large model constraints provided by this application has the same beneficial effects as the method for generating a point set occupancy network based on large model constraints described in the aforementioned embodiment. Other technical features of the device for generating a point set occupancy network based on large model constraints are the same as those disclosed in the aforementioned embodiment and are not further elaborated upon here.

[0199] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0200] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0201] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, and the computer-readable program instructions are used to execute the method for generating a point set occupancy network based on large model constraints in the above-mentioned embodiment.

[0202] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0203] The computer-readable storage medium may be included in the device for generating a point set occupancy network based on large model constraints; or may exist independently without being assembled into the device for generating a point set occupancy network based on large model constraints.

[0204] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by a point set occupancy network generation device based on a large model constraint, the point set occupancy network generation device based on a large model constraint: obtains segmented image features and image depth features; performs feature enhancement and adaptation occupancy voxel tasks based on the segmented image features and the image depth features, determines the application area and target feature adaptation output information; reconstructs the scene positioning scene mapping points based on the application area and the target feature adaptation output information, determines a simulated lidar data set, and completes scene voxel generation and predicts occupancy semantic labels based on the simulated lidar data set.

[0205] The computer program code for performing the operations of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).

[0206] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0207] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0208] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned method for generating a point set occupancy network based on large model constraints. This computer-readable storage medium addresses the technical problem of more efficiently and accurately generating a point set occupancy network based on large model constraints. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the method for generating a point set occupancy network based on large model constraints provided in the aforementioned embodiments, and are not further elaborated here.

[0209] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for generating a point set occupancy network based on large model constraints, characterized in that: The method includes: Obtain segmentation image features and image depth features; Performing a feature enhancement and adaptation occupied voxel task based on the segmented image features and the image depth features, and determining application areas and target feature adaptation output information; Reconstructing the scene positioning scene mapping points based on the application area and the target feature adaptation output information, determining a pseudo-lidar data set, and completing scene voxel generation and predicting occupancy semantic labels based on the pseudo-lidar data set; The step of performing a feature enhancement adaptation occupied voxel task based on the segmented image features and the image depth features and determining target feature adaptation output information comprises: Acquire a feature interaction mode, wherein the feature interaction mode includes a cross-attention interaction mode and a feature splicing interaction mode; Performing feature splicing on the image feature and the image depth feature based on a feature splicing interaction mode in the feature interaction mode to determine a mixed feature code; Based on the cross-attention interaction mode in the feature interaction mode, the image features, the image depth features and the mixed feature encoding are subjected to a feature enhancement adaptation occupied voxel task to obtain target feature adaptation output information.

2. The method according to claim 1, wherein The step of obtaining the segmentation image features and the image depth features comprises: Obtain surround view camera data, segmentation model, and depth model; Inputting the surround view camera data into the segmentation model to identify the target segmentation subject, extract the segmentation image, and determine the segmentation image features; The surround view camera data is input into the depth cutting model to identify the depth of the target body, extract the image depth, and determine the image depth feature.

3. The method according to claim 1, wherein The target feature adaptation output information includes segmentation feature adaptation output information and image depth feature adaptation output information; The step of determining the application area based on the segmented image features, the image depth features and the mixed feature coding includes: The target area is segmented based on the segmented image features, the image depth features and the mixed features to obtain an application area, where the application area includes a ground area, an environment area and an object area.

4. The method according to claim 1, wherein The steps of reconstructing the scene positioning scene mapping points based on the application area and the target feature adaptation output information and determining the simulated laser radar data set include: Reconstructing the scene based on the application area and the target feature adaptation output information to locate the scene mapping point, and determining the occupied voxel mapping point set and the occupied semantic label; A scene representation is performed based on the occupied voxel mapping point set and the occupied semantic label to construct pseudo-lidar data.

5. The method according to claim 4, wherein: The steps of reconstructing the scene positioning scene mapping points based on the application area and the target feature adaptation output information and determining the occupied voxel mapping point set and occupied semantic labels include: Obtaining pixel data, the pixel data including pixel position, pixel center area, and camera hardware focal length; generating an occupancy semantic label based on the application area, and selecting a scene reconstruction method, wherein the scene reconstruction method includes a ground area reconstruction method, an environment area reconstruction method, and an object area reconstruction method; The scene mapping points are located based on the scene reconstruction method, the target feature adaptation output information and the pixel data reconstructed scene, and an occupied voxel mapping point set is calculated.

6. The method according to claim 5, wherein The steps of reconstructing the scene based on the scene reconstruction method, the target feature adaptation output information and the pixel data to locate the scene mapping points and calculate the occupied voxel mapping point set include: When the scene reconstruction method is a ground area reconstruction method or the scene reconstruction method is an environment area reconstruction method, a mapping point set is calculated based on the image depth feature adaptation output information in the target feature adaptation output information and the pixel data, the mapping point set including a ground area mapping point set and an environment area mapping point set; When the scene reconstruction method is an object region reconstruction method, the potential occluded objects are supplemented based on the pixel data, the segmentation feature adaptation output information, and the image depth feature adaptation output information in the target feature adaptation output information, and a supplemented mapping point set is calculated; An occupied voxel mapping point set is obtained based on the mapping point set and the completed mapping point set.

7. The method according to claim 4, wherein The step of constructing simulated LiDAR data by performing scene representation based on the occupied voxel mapping point set and the occupied semantic label comprises: Performing voxelization processing on the occupied voxel mapping point set to construct a scene and determine scene data; Semantic categories are bound based on the scene data and the occupancy semantic labels to obtain simulated lidar data.

8. A point set occupancy network generation device based on large model constraints, characterized in that: The device comprises: An acquisition module, used to obtain segmentation image features and image depth features; A processing module, configured to perform a feature enhancement and adaptation occupied voxel task based on the segmented image features and the image depth features, and determine application areas and target feature adaptation output information; An execution module is configured to reconstruct a scene based on the application area and the target feature adaptation output information to locate a scene mapping point, determine a pseudo-lidar data set, and complete scene voxel generation and predict occupancy semantic labels based on the pseudo-lidar data set; The processing module is further configured to obtain a feature interaction mode, wherein the feature interaction mode includes a cross-attention interaction mode and a feature splicing interaction mode; Performing feature splicing on the image feature and the image depth feature based on a feature splicing interaction mode in the feature interaction mode to determine a mixed feature code; Based on the cross-attention interaction mode in the feature interaction mode, the image features, the image depth features and the mixed feature encoding are subjected to a feature enhancement adaptation occupied voxel task to obtain target feature adaptation output information.

9. A point set occupancy network generation device based on large model constraints, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for generating a point set occupancy network based on large model constraints according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method for generating a point set occupancy network based on large model constraints as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method, device and equipment for sound barrier engineering and storage medium

    CN117576311A

  • Network occupancy prediction method and device, equipment, storage medium and product

    CN118864873A