Open vocabulary occupation prediction method and system, terminal equipment and medium

By constructing an open vocabulary occupancy prediction model, and using feature guidance and feature distillation loss to optimize the alignment of the semantic features and visual language features of RGB images, the accuracy of open vocabulary occupancy prediction is solved and the accuracy of the perception of the autonomous driving environment is improved.

CN120236257APending Publication Date: 2025-07-01NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510379293.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing open vocabulary occupancy prediction methods have low accuracy in autonomous driving, especially in real scenarios, and it is difficult to identify unseen target categories, and the labeling of real data is expensive. At the same time, the ignorance of visual-language feature encoding similarity during distillation leads to increased ambiguity.

Method used

An open vocabulary occupancy prediction model is built, and a coding module that acquires the semantic features of RGB images based on feature guidance is combined with feature distillation losses to achieve spatial alignment of RGB images and visual language features. It uses vehicle-mounted camera data and environmental category labels for training, and uses feature distillation losses and grouped cosine similarity losses to enhance feature alignment and differential capture.

Benefits of technology

It improves the accuracy of open vocabulary occupancy prediction, reduces distillation ambiguity, enhances the spatial alignment of semantic features and visual language features, and improves the accuracy of perception of autonomous driving environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236257A_ABST
    Figure CN120236257A_ABST
Patent Text Reader

Abstract

The invention provides an open vocabulary occupation prediction method and system, terminal equipment and a medium. The method comprises the following steps: acquiring training data; training a pre-constructed open vocabulary occupancy prediction model by using the training data to obtain a trained open vocabulary occupancy prediction model; wherein the open vocabulary occupancy prediction model comprises a coding module for acquiring semantic features of an RGB image based on feature guidance, an updating module for realizing spatial alignment between the semantic features of the RGB image and visual language features based on feature distillation loss, and a reasoning module for predicting open vocabulary occupancy according to the semantic features of the RGB image; and inputting to-be-predicted RGB image data into the trained open vocabulary occupation prediction model to obtain an open vocabulary occupation prediction result corresponding to the to-be-predicted RGB image data. According to the invention, the accuracy of open vocabulary occupation prediction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and specifically relates to an open-vocabulary occupancy prediction method, system, terminal device, and medium. Background Art

[0002] Comprehensive 3D scene understanding is crucial for autonomous driving. For autonomous vehicles, the perception of the surrounding environment is the basis for downstream tasks such as path planning and navigation. Open-vocabulary occupancy prediction, i.e., 3D occupancy prediction (3DOccupancy Prediction), reconstructs the scene using voxels with semantic information, providing a complete 3D perception for autonomous vehicles. However, most of the current research work is trained on closed category sets and relies on 3D voxel labels for supervised training, which brings two major challenges. First, applying these models to real-world scenarios is challenging because they may have difficulty recognizing unseen target categories and may misclassify targets not included in the training dataset. Second, preparing real data is both challenging and costly because it requires fine-grained voxel annotation of hundreds of scenes. Although some technicians have proposed using rendering techniques to project 3D representations into 2D, enabling the model to be trained using 2D labels or pseudo-labels without voxel labels, their semantic understanding ability is still limited by the limited training categories.

[0003] In some other studies, technicians achieved zero-shot open-vocabulary occupancy prediction by distilling knowledge from large vision-language models. They backprojected 3D voxel features into 2D and then used the 2D vision-language features extracted from the vision-language model as supervision labels. Thus, the 3D voxel features are aligned with the vision-language feature space within the vision-language model, enabling voxels to be segmented into any category according to text descriptions without semantic labels. However, during the distillation process, traditional methods ignore the inherent similarity between vision-language feature encodings when calculating the similarity loss. As Figure 1 shown, for some categories, such as "car" and "truck", their vision-language feature encodings show a high degree of similarity, which causes confusion for the model when learning the correct feature encoding distribution, resulting in ambiguity during the distillation process and reducing the accuracy of open-vocabulary occupancy prediction, leading to a deviation in the perception of the surrounding environment by autonomous driving. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an open-vocabulary occupancy prediction method, system, terminal device, and medium to improve the accuracy of open-vocabulary occupancy prediction.

[0005] In a first aspect, the present invention provides an open-vocabulary occupancy prediction method, which includes the following steps:

[0006] Obtain training data; the training data includes RGB image data captured by an on-vehicle camera for describing the surrounding environment of the vehicle, on-vehicle camera parameters, and environmental class labels corresponding to the RGB image data;

[0007] Use the training data to train a pre-constructed open-vocabulary occupancy prediction model to obtain a trained open-vocabulary occupancy prediction model; wherein, the open-vocabulary occupancy prediction model includes an encoding module for obtaining semantic features of the RGB image based on feature guidance, an update module for achieving spatial alignment between the semantic features of the RGB image and visual language features based on feature distillation loss, and an inference module for performing open-vocabulary occupancy prediction according to the semantic features of the RGB image;

[0008] Input the RGB image data to be predicted into the trained open-vocabulary occupancy prediction model to obtain an open-vocabulary occupancy prediction result corresponding to the RGB image data to be predicted.

[0009] Optionally, the encoding module is composed of an encoder, a decoder, a dimension converter, a feature guidance branch, a first multi-layer perceptron network, and a second multi-layer perceptron network;

[0010] Obtaining semantic features of the RGB image based on feature guidance includes:

[0011] The encoder extracts 2D image features of the RGB image data based on an image backbone network and inputs the 2D image features into a pre-trained deep network to obtain a depth map corresponding to the RGB image data;

[0012] The feature guidance branch extracts 2D visual language features of the RGB image data based on a visual language model;

[0013] The dimension converter aggregates the 2D image features and the depth map to obtain 3D sparse voxel features of the RGB image data, and performs feature dimension elevation on the 2D visual language features to obtain 3D visual language features;

[0014] The decoder decodes the 3D sparse voxel features to obtain 3D dense volume features of the RGB image data;

[0015] The first multi-layer perceptron network fuses the 3D dense volume features and the 3D visual language features to obtain a semantic body of the RGB image data;

[0016] The second multi-layer perceptron network obtains the density of the volume according to the 3D dense volume features.

[0017] Optionally, to achieve spatial alignment between the semantic features of the RGB image and the vision-language features based on the feature distillation loss, it includes:

[0018] Generate multiple rays according to the in-vehicle camera parameters;

[0019] Calculate the 2D rendered semantic features corresponding to each ray;

[0020] Extract the 2D vision-language features corresponding to each ray from the 2D vision-language features according to the pixel positions corresponding to each ray;

[0021] Construct a feature distillation loss based on the 2D rendered semantic features and the 2D vision-language features corresponding to each ray to achieve spatial alignment between the semantic features of the RGB image and the vision-language features.

[0022] Optionally, calculating the 2D rendered semantic features corresponding to each ray includes:

[0023] For each ray respectively, sample the ray to obtain multiple sampling points;

[0024] Calculate the termination probability and cumulative transmittance corresponding to each sampling point, and calculate the 2D rendered semantic features corresponding to each ray according to the termination probability and cumulative transmittance.

[0025] Optionally, calculating the 2D rendered semantic features corresponding to each ray according to the termination probability and cumulative transmittance includes:

[0026] By using the calculation formula

[0027]

[0028] α(p k )=1-exp(-σ(p k )d)

[0029]

[0030] Obtain the 2D rendered semantic feature S 2D (r) of the r-th ray; where p k represents the sampling point on the r-th ray, k = 1, 2,..., n, n represents the number of sampling points, α(p k ) represents the termination probability of the sampling point p k , T(p k ) represents the cumulative transmittance of the sampling point p k , S(p k ) represents the semantic feature of the semantic body S at the sampling point p k , σ(p k ) represents the sampling point p kThe termination probability, d represents the sampling interval, t represents the t-th sampling point on the r-th ray. When k≠1, t = 1,..., k - 1; when k = 1, t = 1.

[0031] Optionally, the expression of the feature distillation loss is as follows:

[0032] L FD (V1, V2) = λ1L SCS (V1, V2) + λ2L MSE (V1, V2)

[0033]

[0034] Wherein, L FD (V1, V2) represents the feature distillation loss of V1 and V2. V1 and V2 represent 2D vision-language features and 2D image features respectively. λ1 and λ2 are hyperparameters. L SCS (V1, V2) represents the grouped cosine loss of V1 and V2. L MsE (V1, V2) represents the mean square error loss of V1 and V2. n represents the number of groups, and i represents the i-th group.

[0035] Optionally, open-vocabulary occupancy prediction based on RGB image semantic features includes:

[0036] Determine the state of each voxel in the semantic body according to the density; the state is occupied or unoccupied;

[0037] For the voxels with the state of occupied, through the calculation formula

[0038]

[0039] Obtain the open-vocabulary occupancy prediction result O(x, y, z) of the voxel (x, y, z); wherein, S T (x, y, z) represents the similarity score between the voxel S(x, y, z) with the state of occupied and the text feature T vl ; represents the tensor product operation, τ represents the preset density threshold, σ(x, y, z) represents the density of the voxel (x, y, z), and the open-vocabulary occupancy prediction result represents the environmental category label corresponding to the voxel.

[0040] In a second aspect, the present invention discloses an open-vocabulary occupancy prediction system, including:

[0041] A data acquisition module, configured to acquire training data; the training data includes RGB image data captured by an on-vehicle camera for describing the environment around the vehicle, on-vehicle camera parameters, and the environmental category label corresponding to the RGB image data;

[0042] A model training module for training a pre-constructed open-vocabulary occupancy prediction model using training data to obtain a trained open-vocabulary occupancy prediction model; wherein, the open-vocabulary occupancy prediction model includes an encoding module for obtaining semantic features of RGB images based on feature guidance, an updating module for achieving spatial alignment between semantic features of RGB images and vision-language features based on a feature distillation loss, and an inference module for performing open-vocabulary occupancy prediction according to the semantic features of RGB images.

[0043] A prediction module for inputting RGB image data to be predicted into the trained open-vocabulary occupancy prediction model to obtain an open-vocabulary occupancy prediction result corresponding to the RGB image data to be predicted.

[0044] In a third aspect, the present invention discloses a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above method is implemented.

[0045] In a fourth aspect, the present invention discloses a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0046] The beneficial effects of the present invention are:

[0047] The open-vocabulary occupancy prediction method disclosed by the present invention constructs an open-vocabulary occupancy prediction model. The encoding module of the open-vocabulary occupancy prediction model obtains semantic features of RGB images based on feature guidance, can more effectively learn the distribution of target feature encoding, reduce distillation ambiguity, and thus improve the accuracy of open-vocabulary occupancy prediction; its updating module achieves spatial alignment between semantic features of RGB images and vision-language features based on a feature distillation loss, can enhance the alignment between semantic features and the vision-language feature space, accurately capture the differences between vision-language model feature embeddings, improve the effectiveness of distillation, further reduce distillation ambiguity, and improve the accuracy of open-vocabulary occupancy prediction. Description of the Drawings

[0048] Figure 1 A similarity comparison diagram for visual-language encodings of different categories;

[0049] Figure 2 A flowchart of the open-vocabulary occupancy prediction method in one implementation manner of the present application;

[0050] Figure 3 A structural diagram of the open-vocabulary occupancy prediction model in one implementation manner of the present application;

[0051] Figure 4 A structural diagram of the encoding module in one implementation manner of the present application;

[0052] Figure 5 It is a similarity comparison diagram of different types of visual language encodings calculated using grouped cosine similarity loss in one implementation manner of this application;

[0053] Figure 6 It is a schematic structural diagram of an open vocabulary occupancy prediction system in one implementation manner of this application;

[0054] Figure 7 It is a schematic structural diagram of a terminal device in one implementation manner of this application. Detailed implementation manners

[0055] Aiming at the problem of low accuracy in open vocabulary occupancy prediction by traditional methods, the present invention discloses an open vocabulary occupancy prediction method, system, terminal device and medium. This method constructs an open vocabulary occupancy prediction model. The encoding module of the open vocabulary occupancy prediction model obtains RGB image semantic features based on feature guidance, can more effectively learn the distribution of target feature encodings, reduce distillation ambiguity, and thus improve the accuracy of open vocabulary occupancy prediction; its update module realizes the spatial alignment between RGB image semantic features and visual language features based on feature distillation loss, can enhance the alignment between semantic features and visual language feature spaces, accurately capture the differences between visual language model feature embeddings, improve the effectiveness of distillation, further reduce distillation ambiguity, and improve the accuracy of open vocabulary occupancy prediction.

[0056] The open vocabulary occupancy prediction method disclosed by the present invention will be described below.

[0057] As Figure 2 shown, the open vocabulary occupancy prediction method includes the following steps:

[0058] Step 21, obtain training data.

[0059] In an embodiment of the present invention, the training data includes RGB image data captured by an in-vehicle camera for describing the vehicle surrounding environment, in-vehicle camera parameters, and environmental category labels corresponding to the RGB image data. The above training data can be obtained through a public database, and the acquisition path thereof will not be elaborated herein.

[0060] Specifically, the RGB image data is a set of multi-view RGB images, denoted as I = {I1, I2,..., I N}, where N represents the number of in-vehicle cameras.

[0061] The parameters of the vehicle-mounted camera include the internal parameters K and the external parameters [R|t] of the camera. Among them, the internal parameters specifically include the focal length, the principal point, the distortion coefficient, the camera matrix, and the pose. The principal point is the origin of the image coordinate system, which is usually located at the center of the image and represents the offset between the camera coordinate system and the image coordinate system. The distortion coefficient is used to describe and correct the barrel or pillow distortion that appears in the RGB image. The camera matrix is a matrix that converts the three-dimensional object coordinates into two-dimensional image coordinates. The external parameters specifically include the rotation matrix R and the translation vector T. The rotation matrix describes the rotation relationship between the camera coordinate system and the world coordinate system. The translation vector describes the translation of the origin of the camera coordinate system relative to the origin of the world coordinate system.

[0062] The environmental category labels corresponding to the RGB image data are used to describe the environmental categories. In the field of autonomous driving, the environmental categories may include roads, street lights, pedestrians, other vehicles, etc. In the embodiments of the present invention, the environmental category labels are presented in text form and can be specifically expressed as C = {C1, C2,..., C L}, where L represents the total number of categories, and each category supports multiple text descriptions. For example: {"driveable surface": "road", "surface on which a car can drive"}.

[0063] Step 22: Use the training data to train the pre-constructed open vocabulary occupancy prediction model to obtain the trained open vocabulary occupancy prediction model.

[0064] As Figure 3 shown, in the embodiments of the present invention, the open vocabulary occupancy prediction model includes an encoding module 301 for obtaining the semantic features of the RGB image based on feature guidance, an updating module 302 for realizing the spatial alignment between the semantic features of the RGB image and the vision-language features based on the feature distillation loss, and an inference module 303 for performing open vocabulary occupancy prediction according to the semantic features of the RGB image.

[0065] The encoding module is described below.

[0066] In the embodiments of the present invention, the encoding module is composed of an encoder 401, a decoder 402, a dimension converter 403, a feature guiding branch 404, a first multi-layer perceptron network 405, and a second multi-layer perceptron network 406, as specifically Figure 4 shown.

[0067] In the process of the encoding module obtaining the semantic features of the RGB image based on feature guidance, the functions of its respective sub-modules are described as follows:

[0068] The encoder 401 extracts 2D image features of RGB image data based on an image backbone network, and inputs the 2D image features into a pre-trained deep network to obtain a depth map corresponding to the RGB image data. Exemplarily, the image backbone network can extract feature information from the original image. The selectable image backbone networks include at least VGGNet, ResNet, etc. The deep network is used to convert the RGB image into a depth map. Common deep networks include Monodepth2 and DeepLabV3.

[0069] The feature guidance branch 404 extracts 2D visual language features of RGB image data based on a vision-language model. Due to the ambiguity problem in the feature distillation process, it may be difficult for the network to accurately learn the distribution of the vision-language model feature encoding space. In the embodiments of the present invention, relevant visual language features are extracted from the input image and provided to the network as supplementary guidance, which enables the network to more effectively learn the distribution of v-l features and improve the accuracy of open vocabulary occupancy prediction.

[0070] Specifically, in the feature guidance branch 404, through the calculation formula

[0071]

[0072] the 2D visual language features are obtained where f CLIP (·) represents the process of feature extraction by the image encoder of CLIP, and I represents the RGB image.

[0073] The dimension converter 403 aggregates the 2D image features and the depth map to obtain 3D sparse voxel features of the RGB image data, and performs feature dimension elevation on the 2D visual language features to obtain 3D visual language features.

[0074] The process of performing feature dimension elevation on the 2D visual language features to obtain 3D visual language features is specifically as follows:

[0075] Through the calculation formula

[0076]

[0077] the 3D visual language features are obtained where Trans(·) represents dimension conversion. In the embodiments of the present invention, the 2D to 3D conversion can adopt the LSS method (a prior art).

[0078] The decoder 402 decodes the 3D sparse voxel features to obtain 3D dense volume features of the RGB image data.

[0079] The first multi-layer perceptron network 405 fuses the 3D dense volume features and 3D vision-language features to obtain the semantic body of the RGB image data.

[0080] Specifically, through the calculation formula

[0081]

[0082] the semantic body S is obtained; where V 3D represents the 3D dense volume features, concat(·) represents the concatenation operation, and MLP(·) represents the multi-layer perceptron network.

[0083] The second multi-layer perceptron network 406 obtains the density of the volume according to the 3D dense volume features.

[0084] It should be noted that in the embodiments of the present invention, the above volume actually refers to the three-dimensional features generated by the neural network. The two perceptron networks (the first multi-layer perceptron network and the second multi-layer perceptron network) respectively process this three-dimensional feature. The first multi-layer perceptron network predicts the density of the three-dimensional space based on the features, and this density represents the probability that each voxel block in the three-dimensional space is not "empty". The second multi-layer perceptron network then predicts the semantic information of these voxel blocks based on the three-dimensional features.

[0085] Next, the process of realizing the spatial alignment between the RGB image semantic features and the vision-language features based on the feature distillation loss in the update module will be described.

[0086] The update module uses the volume rendering technology to back-project the semantic density field (SDF, Semantic Density Field) into the 2D space to generate the predicted vision-language features. Then, the training loss between these rendered features and the vision-language features extracted from the image encoder of the vision-language model is calculated, and the model parameters of the encoding module are updated according to the training loss until the training loss is less than the preset loss threshold.

[0087] Next, the process of realizing the spatial alignment between the RGB image semantic features and the vision-language features based on the feature distillation loss in the update module will be described.

[0088] Specifically, it includes steps I to IV.

[0089] Step I, generate multiple rays according to the vehicle-mounted camera parameters.

[0090] Specifically, given a set of camera poses and their corresponding images, n rays are generated according to the internal and external parameters of the camera Each ray r iThey are all calculated based on specific camera poses. It starts from the origin of the camera and shoots into the 3D space through the pixel (x, y) of the corresponding image. In addition, in some other embodiments of the present invention, in order to improve the geometric perception ability of the network, the present invention adopts the 2D depth supervision and auxiliary ray generation strategy in RenderOcc.

[0091] Step II, calculate the 2D rendered semantic features corresponding to each ray.

[0092] To render the semantic features at the pixel (x, y), the present invention samples K points at intervals of d on the ray r For each point p k , its termination probability α(p k ) and cumulative transmittance T(p k ) are respectively expressed as:

[0093] α(p k ) = 1 - exp(-σ(p k )d)

[0094]

[0095] The 2D rendered semantic features of the ray r can be obtained through the calculation formula

[0096]

[0097] ; where S 2D (r) represents the 2D rendered semantic features corresponding to the r-th ray, p k represents the sampling point on the r-th ray, k = 1, 2,..., n, n represents the number of sampling points, α(p k ) represents the termination probability of the sampling point p k , T(p k ) represents the cumulative transmittance of the sampling point p k , S(p k ) represents the semantic features of the semantic body S at the sampling point p k , σ(p k ) represents the termination probability of the sampling point p k , d represents the sampling interval, t represents the t-th sampling point on the r-th ray, when k ≠ 1, t = 1,..., k - 1, when k = 1, t = 1.

[0098] It should be noted that assuming a ray of light enters space and multiple points are sampled on this line, each point has a density probability, which can be understood as the passing probability of the light ray passing through this point. Therefore, by performing cumulative calculations along the propagation direction of the light ray, the termination probability of the light ray reaching each point can be obtained. For example, the density probability of point A is 0.6, and the density probability of point B is 0.3. If the light ray passes through A and B in sequence, the termination probability of the light ray at point A is 0.6, and the termination probability at point B is (1 - 0.6) * 0.3. The cumulative transmittance is the probability of the light ray passing through each point, which is the termination probability.

[0099] Step III: According to the pixel positions corresponding to each ray, extract the 2D visual language features corresponding to each ray from the 2D visual language features.

[0100] Specifically, by sampling the visual language feature map extracted by the encoder of the visual language model at the pixel position (x, y), the corresponding feature ground truth is obtained.

[0101] Step IV: According to the 2D rendering semantic features and the 2D visual language features corresponding to each ray, construct a feature distillation loss to achieve spatial alignment between the semantic features of the RGB image and the visual language features.

[0102] To save memory, the present invention reduces the feature channel dimension of the encoder-decoder network. Therefore, after rendering, the present invention uses a mapping layer to align the feature channel dimension with the feature ground truth.

[0103] It should be noted that label-based supervision has clear boundaries between categories, while feature-based supervision does not. Therefore, under label-based supervised training, incorrect pixel semantic classification will significantly increase the loss value. However, under feature-based supervision, even if the network classification is incorrect, the loss may remain unchanged. In an embodiment of the present invention, five different category feature encodings generated by the visual language image encoder are extracted, and the cosine similarity function is used to calculate the similarity between them. As Figure 1 shown, the similarity scores between these categories are extremely high. Therefore, this leads to training ambiguity and impairs the effectiveness of training. For example, from Figure 1It is observed that the cosine similarity between the sampled visual-linguistic feature encodings of 'car' and 'pedestrian' is as high as 0.9, indicating that in the visual-linguistic feature space, the feature encodings of these two categories may be highly similar. During the training process, if the network wrongly generates a feature encoding representing 'pedestrian' instead of 'car' on the voxels within the car region, the loss feedback will be too small to effectively correct the network. Because the cosine similarity loss between the predicted feature encoding and the visual-linguistic feature encoding of 'car' is similar to the loss value between the predicted feature encoding and the visual-linguistic feature encoding of 'pedestrian'.

[0104] To amplify the differences between feature encodings, the present invention adopts a calculation method of grouped cosine similarity loss. By splitting the entire encoding vector, it highlights the differences between different regions and ignores the influence of feature vectors in the common regions. Specifically, as Figure 4 shown, in the embodiments of the present invention, the similar parts in the visual-linguistic feature encoding are called the common zone, and the different parts are called the distinct zone. The features belonging to the distinct zone are the key to identifying the target semantic information. Therefore, to amplify the differences between visual-linguistic encodings, especially the feature encodings in the distinct zone, the present invention splits the encoding into multiple sub-encodings along the feature dimension to calculate the cosine similarity loss, named Subgroup Cosine Similarity Loss (SCS Loss), and its calculation formula is:

[0105]

[0106] wherein. L SCS (V1, V2) represents the grouped cosine loss of V1 and V2, and V1 and V2 represent 2D visual-linguistic features and 2D image features respectively.

[0107] The present invention recalculates the similarity between feature encodings, and the results are as Figure 5 shown. Compared with Figure 1 , the similarity between visual-linguistic feature encodings is significantly reduced, indicating that the present invention is effective in amplifying the differences between embeddings.

[0108] The cosine similarity loss mainly emphasizes minimizing the angular distance between two feature encodings in the feature space. Mathematically, two vectors are considered equivalent if and only if their directions and lengths are equal. Therefore, the present invention adopts the mean squared error (MSE) loss to guide the network to learn the length features of the target feature embeddings, which ensures that the network can more accurately capture the distribution of the feature space in the visual-linguistic model. We call this loss the feature distillation loss, and its expression is as follows:

[0109] L FD (V1,V2)=λ1L SCS (V1,V2)+λ2L MSE (V1,V2)

[0110] Among them, L FD (V1, V2) represents the feature distillation loss of V1 and V2, λ1 and λ2 are hyperparameters, L MSE (V1, V2) represents the mean square error loss of V1 and V2, n represents the number of groups, and i represents the i-th group.

[0111] In some other embodiments of the present invention, in order to take into account both deep supervision and SDF regularization, SILogLoss Ldepth, distortion loss and TV loss

[46] can also be introduced to construct the loss function in the update module, which is specifically expressed as L = L FD +L depth +L reg .

[0112] The following is an explanation of the training process of the open word occupancy prediction model.

[0113] During the training process, the model parameters are first initialized. After the training data is input into the model, the feature distillation loss is calculated, and the model parameters of the encoding module are back-propagated according to the feature distillation loss until the feature distillation loss value is less than the preset loss threshold, thus obtaining a trained open word occupancy prediction model.

[0114] The reasoning module is described below.

[0115] The process of the inference module predicting open word occupancy based on the semantic features of the RGB image includes steps A to B:

[0116] Step A, according to the density, determines the state of each voxel in the semantic volume.

[0117] The above status is occupied or unoccupied;

[0118] Step B: for the voxels in the occupied state, calculate the formula

[0119]

[0120] Get the open word occupancy prediction result O(x,y,z) of voxel (x,y,z).

[0121] Among them, S T (x, y, z) represents the voxel S(x, y, z) with the status of occupied and text feature T vl The similarity score between represents the tensor product operation, τ represents a preset density threshold, σ(x, y, z) represents the density of the voxel (x, y, z), and the open vocabulary occupancy prediction result represents the environmental category label corresponding to the voxel.

[0122] Step 23: Input the RGB image data to be predicted into the trained open vocabulary occupancy prediction model to obtain the open vocabulary occupancy prediction result corresponding to the RGB image data to be predicted.

[0123] To verify the effectiveness of the present invention, in an embodiment of the present invention, the open vocabulary occupancy prediction method provided by the present invention was evaluated on the Occ3d-nuScenes dataset and compared with the state-of-the-art LangOcc method. The results show that the open vocabulary occupancy prediction method provided by the present invention achieved performance comparable to that of the LangOcc method. It should be noted that LangOcc reduced the embedding space of the vision-language model by pre-training an autoencoder, making a concession between maintaining the open vocabulary expression ability and the low-dimensional subspace. The open vocabulary occupancy prediction method provided by the present invention directly distills knowledge from the original vision-language model, retaining as much as possible the generalization ability of the vision-language model. Without compromising the generalization ability of the vision-language model, it achieved performance comparable to that of the state-of-the-art LangOcc method in the zero-shot open vocabulary occupancy prediction task, verifying the accuracy of the open vocabulary occupancy prediction of the present invention.

[0124] In summary, the open vocabulary occupancy prediction method disclosed in the present invention constructs an open vocabulary occupancy prediction model. The encoding module of the open vocabulary occupancy prediction model obtains the RGB image semantic features based on feature guidance, which can more effectively learn the distribution of the target feature encoding, reduce the distillation ambiguity, and thus improve the accuracy of the open vocabulary occupancy prediction. Its update module realizes the spatial alignment between the RGB image semantic features and the vision-language features based on the feature distillation loss, which can enhance the alignment between the semantic features and the vision-language feature space, accurately capture the differences between the vision-language model feature embeddings, improve the effectiveness of the distillation, further reduce the distillation ambiguity, and improve the accuracy of the open vocabulary occupancy prediction.

[0125] As Figure 6 shown, the present invention discloses an open vocabulary occupancy prediction system, which includes:

[0126] A data acquisition module 601, configured to acquire training data; the training data includes RGB image data captured by an in-vehicle camera for describing the environment around the vehicle, in-vehicle camera parameters, and the environmental category label corresponding to the RGB image data;

[0127] A model training module 602 is configured to train a pre-constructed open-vocabulary occupancy prediction model using training data to obtain a trained open-vocabulary occupancy prediction model. The open-vocabulary occupancy prediction model includes an encoding module for obtaining RGB image semantic features based on feature guidance, an update module for achieving spatial alignment between RGB image semantic features and vision-language features based on a feature distillation loss, and an inference module for performing open-vocabulary occupancy prediction based on RGB image semantic features.

[0128] A prediction module 603 is configured to input RGB image data to be predicted into the trained open-vocabulary occupancy prediction model to obtain an open-vocabulary occupancy prediction result corresponding to the RGB image data to be predicted.

[0129] It should be noted that for the information interaction, execution process, etc. between the above-mentioned device / units, since they are based on the same concept as the method embodiments of the present application, their specific functions and the technical effects brought can be specifically referred to the method embodiment part, and will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments and will not be elaborated here.

[0130] As Figure 7 shown, an embodiment of the present invention provides a terminal device. As Figure 7 shown, the terminal device D10 in this embodiment includes at least one processor D100 ( Figure 7 only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, the steps in any of the above method embodiments are implemented.

[0131] Specifically, when the processor D100 executes the computer program D102, it obtains training data; trains a pre-constructed open-vocabulary occupancy prediction model using the training data to obtain a trained open-vocabulary occupancy prediction model; and inputs the RGB image data to be predicted into the trained open-vocabulary occupancy prediction model to obtain the open-vocabulary occupancy prediction result corresponding to the RGB image data to be predicted. Among them, this method constructs an open-vocabulary occupancy prediction model. The encoding module of this open-vocabulary occupancy prediction model obtains the RGB image semantic features based on feature guidance, can more effectively learn the distribution of target feature encoding, reduce the distillation ambiguity, and thus improve the accuracy of open-vocabulary occupancy prediction; its update module realizes the spatial alignment between the RGB image semantic features and the vision-language features based on the feature distillation loss, can enhance the alignment between the semantic features and the vision-language feature space, accurately capture the differences between the vision-language model feature embeddings, improve the effectiveness of distillation, further reduce the distillation ambiguity, and improve the accuracy of open-vocabulary occupancy prediction.

[0132] The so-called processor D100 may be a central processing unit (CPU, Central Processing Unit), and this processor D100 may also be other general-purpose processors, digital signal processors (DSP, Digital Signal Processor), application specific integrated circuits (ASIC, Application Specific Integrated Circuit), field-programmable gate arrays (FPGA, Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0133] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as the hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart media card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. equipped on the terminal device D10. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store the operating system, application programs, boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store the data that has been output or will be output.

[0134] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0135] An embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device can implement the steps in the above-mentioned various method embodiments when executed.

[0136] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is exemplary only and is not intended to imply that the scope of protection of the present application is limited to these examples; under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.

[0137] One or more embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the present application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for predicting open vocabulary occupancy, characterized in that: include: Get training data; The training data includes RGB image data taken by a vehicle-mounted camera for describing the vehicle's surrounding environment, vehicle-mounted camera parameters, and an environment category label corresponding to the RGB image data; The pre-built open vocabulary occupancy prediction model is trained using the training data to obtain a trained open vocabulary occupancy prediction model; wherein the open vocabulary occupancy prediction model includes an encoding module for acquiring RGB image semantic features based on feature guidance, an updating module for achieving spatial alignment between the RGB image semantic features and visual language features based on feature distillation loss, and an inference module for predicting open vocabulary occupancy based on the RGB image semantic features; The RGB image data to be predicted is input into the trained open word occupancy prediction model to obtain the open word occupancy prediction result corresponding to the RGB image data to be predicted.

2. The open vocabulary occupancy prediction method according to claim 1, characterized in that: The encoding module is composed of an encoder, a decoder, a dimension converter, a feature guide branch, a first multi-layer perceptron network, and a second multi-layer perceptron network; The obtaining of RGB image semantic features based on feature guidance includes: The encoder extracts 2D image features of the RGB image data based on an image backbone network, and inputs the 2D image features into a pre-trained deep network to obtain a depth map corresponding to the RGB image data; The feature guiding branch extracts 2D visual language features of the RGB image data based on a visual language model; The dimension converter aggregates the 2D image features and the depth map to obtain 3D sparse voxel features of the RGB image data, and performs feature dimension upgrading on the 2D visual language features to obtain 3D visual language features; The decoder decodes the 3D sparse voxel features to obtain 3D dense volume features of the RGB image data; The first multi-layer perceptron network fuses the 3D dense volume feature and the 3D visual language feature to obtain a semantic body of the RGB image data; The second multilayer perceptron network obtains the density of the volume according to the 3D dense volume feature.

3. The open vocabulary occupancy prediction method according to claim 2, characterized in that: The spatial alignment between the semantic features of the RGB image and the visual language features is achieved based on the feature distillation loss, including: Generate multiple rays according to the vehicle-mounted camera parameters; Calculate the 2D rendering semantic features corresponding to each ray; Extracting the 2D visual language feature corresponding to each ray from the 2D visual language feature according to the pixel position corresponding to each ray; According to the 2D rendering semantic features and the 2D visual language features corresponding to each ray, a feature distillation loss is constructed to achieve spatial alignment between the RGB image semantic features and the visual language features.

4. The open vocabulary occupancy prediction method according to claim 3, characterized in that: The calculating of the 2D rendering semantic features corresponding to each ray includes: For each ray, sampling is performed on the ray to obtain a plurality of sampling points; The termination probability and the cumulative transmittance corresponding to each of the sampling points are calculated, and the 2D rendering semantic features corresponding to each of the rays are calculated according to the termination probability and the cumulative transmittance.

5. The open vocabulary occupancy prediction method according to claim 4, characterized in that: The calculating, according to the termination probability and the cumulative transmittance, the 2D rendering semantic feature corresponding to each ray includes: By calculating the formula α(p k )=1-exp(-σ(p k )d) Get the 2D rendering semantic feature S corresponding to the rth ray 2D (r); where p k represents the sampling point on the rth ray, k = 1, 2, ..., n, n represents the number of sampling points, α(p k ) represents the sampling point p k The termination probability, T(p k ) represents the sampling point p k The cumulative transmittance, S(p k ) represents the semantic body S at sampling point p k The semantic features of σ(p k ) represents the sampling point p k The termination probability of , d represents the sampling interval, t represents the t-th sampling point on the r-th ray, when k≠1, t=1,...,k-1, when k=1, t=1.

6. The open vocabulary occupancy prediction method according to claim 5, characterized in that: The expression of the characteristic distillation loss is as follows: <h2 style=";text-align:left;direction:ltr">L<h2 style=";text-align:left;direction:ltr"> FD <h2 style=";text-align:left;direction:ltr"> (V1,V2)=λ1L<h2 style=";text-align:left;direction:ltr"> SCS <h2 style=";text-align:left;direction:ltr"> (V1,V2)+λ2L<h2 style=";text-align:left;direction:ltr"> MSE <h2 style=";text-align:left;direction:ltr"> (V1,V2) Among them, L FD (V1, V2) represents the feature distillation loss of V1 and V2, V1 and V2 represent 2D visual language features and 2D image features respectively, λ1 and λ2 are hyperparameters, L SCS (V1, V2) represents the group cosine loss of V1, V2, L MSE (V1, V2) represents the mean square error loss of V1 and V2, n represents the number of groups, and i represents the i-th group.

7. The open vocabulary occupancy prediction method according to claim 6, characterized in that: The performing open word occupancy prediction according to the semantic features of the RGB image includes: Determining the state of each voxel in the semantic volume according to the density; the state is occupied or unoccupied; For voxels in the occupied state, the formula is used Get the open word occupancy prediction result O(x,y,z) of voxel (x,y,z); where S T (x, y, z) represents the voxel S(x, y, z) with the status of occupied and text feature T vl The similarity score between represents a tensor product operation, τ represents a preset density threshold, σ(x, y, z) represents the density of the voxel (x, y, z), and the open vocabulary occupancy prediction result represents the environment category label corresponding to the voxel.

8. An open word occupancy prediction system, characterized in that include: A data acquisition module, used to acquire training data; The training data includes RGB image data taken by a vehicle-mounted camera for describing the vehicle's surrounding environment, vehicle-mounted camera parameters, and an environment category label corresponding to the RGB image data; A model training module, used to train a pre-built open vocabulary occupancy prediction model using the training data to obtain a trained open vocabulary occupancy prediction model; wherein the open vocabulary occupancy prediction model includes an encoding module for acquiring RGB image semantic features based on feature guidance, an updating module for achieving spatial alignment between the RGB image semantic features and visual language features based on feature distillation loss, and an inference module for predicting open vocabulary occupancy based on the RGB image semantic features; The prediction module is used to input the RGB image data to be predicted into the trained open word occupancy prediction model to obtain the open word occupancy prediction result corresponding to the RGB image data to be predicted.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.