Gaussian coding model training method, indoor occupancy prediction method, device and medium

By training a Gaussian coding model and combining it with an opacity-guided autoencoder and a geometric perception cross encoder, the problem of insufficient accuracy in existing occupancy prediction models is solved, enabling the robot to achieve high-precision perception of indoor scene occupancy.

CN120635680BActive Publication Date: 2025-11-04BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511128838.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-04
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing occupancy prediction models lack accuracy in robot environmental perception, affecting the execution of downstream tasks.

Method used

By training a Gaussian coding model, opacity-guided autoencoders and geometry-aware cross-encoders are used to extract features from indoor scene images and fuse geometric and semantic data to generate Gaussian features. Gaussian voxel mapping is then used to obtain predicted spatial occupancy information.

Benefits of technology

This improves the robot's accuracy in perceiving indoor scene occupancy, supporting more effective execution of downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635680B_ABST
    Figure CN120635680B_ABST
Patent Text Reader

Abstract

The application provides a Gaussian coding model training method, an indoor occupancy prediction method, equipment and a medium. The model training method comprises: performing feature extraction on a sample indoor scene image to obtain a sample image feature map corresponding to the sample indoor scene image; updating an initially randomly generated Gaussian feature by using a preset initial Gaussian coding model according to the sample image feature map to obtain a target Gaussian feature corresponding to the sample indoor scene image; obtaining predicted spatial occupancy information corresponding to the sample indoor scene image according to the target Gaussian feature; and performing parameter training on the preset initial Gaussian coding model according to the sample spatial occupancy information and the predicted spatial occupancy information to obtain a target Gaussian coding model. The target Gaussian coding model is used to predict a current Gaussian feature of a current indoor scene image collected by a robot, and then predicted spatial occupancy information is obtained. The accuracy is high, and the robot can accurately perceive the occupancy of the indoor scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a Gaussian coding model training method, an indoor occupancy prediction method, equipment and a medium. BACKGROUND

[0002] Occupancy prediction is a key task for robots to realize embodied perception, which can help robots obtain spatial geometry and semantic information of the environment and support downstream tasks such as exploration and decision-making.

[0003] In related technologies, based on the scene image collected by the robot, an occupancy prediction model is established to predict the occupancy of the surrounding environment scene of the robot, so as to indicate whether the surrounding environment scene of the robot is occupied and by what object. However, the existing occupancy prediction model is used to predict the occupancy, which is insufficient in accuracy and affects the execution of downstream tasks. SUMMARY

[0004] Therefore, the embodiments of the present application provide a Gaussian coding model training method, an indoor occupancy prediction method, equipment and a medium to solve the problem that the existing occupancy prediction model is used to predict the occupancy, which is insufficient in accuracy and affects the execution of downstream tasks.

[0005] In a first aspect, the embodiments of the present application provide a Gaussian coding model training method, comprising:

[0006] An indoor scene data set is obtained, which includes a sample indoor scene image, wherein the sample indoor scene image is labeled with sample space occupancy information, and the sample space occupancy information includes actual semantic categories and actual occupancy voxel information of occupied objects in the sample indoor scene image;

[0007] Feature extraction is performed on the sample indoor scene image to obtain a sample image feature map corresponding to the sample indoor scene image;

[0008] According to the sample image feature map, an initial Gaussian feature randomly generated is updated by using a preset initial Gaussian coding model to obtain a target Gaussian feature corresponding to the sample indoor scene image;

[0009] According to the target Gaussian feature, predicted space occupancy information corresponding to the sample indoor scene image is obtained, and the predicted space occupancy information includes predicted semantic categories and predicted occupancy voxel information of occupied objects in the sample indoor scene image;

[0010] According to the sample space occupancy information and the predicted space occupancy information, the preset initial Gaussian coding model is trained to obtain a target Gaussian coding model.

[0011] In an optional implementation, the preset initial Gaussian coding model comprises: an opacity-guided self-encoder and a geometry-aware cross-encoder.

[0012] The preset initial Gaussian coding model comprises: an opacity-guided self-encoder and a geometry-aware cross-encoder.

[0013] The preset initial Gaussian coding model comprises: an opacity-guided self-encoder and a geometry-aware cross-encoder.

[0014] The preset initial Gaussian coding model comprises: an opacity-guided self-encoder and a geometry-aware cross-encoder.

[0015] In an optional implementation, the preset initial Gaussian coding model further comprises: an optimization module.

[0016] The preset initial Gaussian coding model comprises: an opacity-guided self-encoder and a geometry-aware cross-encoder.

[0017] The preset initial Gaussian coding model comprises: an opacity-guided self-encoder and a geometry-aware cross-encoder.

[0018] The preset initial Gaussian coding model comprises: an opacity-guided self-encoder and a geometry-aware cross-encoder.

[0019] In an optional implementation, the opacity-guided self-encoder comprises: a voxelization module and a plurality of sparse convolutional layers.

[0020] The preset initial Gaussian coding model comprises: an opacity-guided self-encoder and a geometry-aware cross-encoder.

[0021] The preset initial Gaussian coding model comprises: an opacity-guided self-encoder and a geometry-aware cross-encoder.

[0022] The preset initial Gaussian coding model comprises: an opacity-guided self-encoder and a geometry-aware cross-encoder.

[0023] In an optional implementation, the geometry-aware cross-encoder comprises: a sampling module, a semantic mixing module and a geometry mixing module.

[0024] The geometric cross-encoder is used to perform geometric cross-perception on the sample image feature map according to the encoded high-level feature, to obtain a cross-perception high-level feature.

[0025] The sampling module is used to sample the sample image feature map according to the encoded high-level feature, to obtain an image feature of a sampling Gaussian point corresponding to the encoded high-level feature.

[0026] The semantic mixing module is used to perform semantic fusion on the image feature of the sampling Gaussian point and semantic information in the encoded high-level feature, to obtain a semantic fusion feature of the sampling Gaussian point.

[0027] The geometric mixing module is used to perform geometric fusion on the semantic fusion feature of the sampling Gaussian point and geometric information in the encoded high-level feature, to obtain a geometric fusion feature of the sampling Gaussian point as the cross-perception high-level feature.

[0028] In an optional implementation, the sampling module is used to sample the sample image feature map according to the encoded high-level feature, to obtain an image feature of a sampling Gaussian point corresponding to the encoded high-level feature, including:

[0029] The sampling module is used to sample the encoded high-level feature, to obtain the sampling Gaussian point, and project the sampling Gaussian point into the sample image feature map, to obtain the image feature of the sampling Gaussian point.

[0030] In an optional implementation, the target high-level feature is used to obtain the predicted space occupation information corresponding to the sample indoor scene image, including:

[0031] A preset Gaussian voxel mapping module is used to perform voxel mapping on the target high-level feature, to obtain the predicted space occupation information.

[0032] In a second aspect, the embodiments of the present application further provide an indoor occupation prediction method, including:

[0033] A current indoor scene image collected by a target robot in a target indoor scene is obtained.

[0034] A feature of the current indoor scene image is extracted, to obtain a current image feature map corresponding to the current indoor scene image.

[0035] The initial high-level feature is updated by using a target Gaussian encoding model according to the current image feature map, to obtain a current high-level feature corresponding to the current indoor scene image, the target Gaussian encoding model being a model trained by using the method in any one of the first aspect.

[0036] Based on the current Gaussian features, the current space occupancy information corresponding to the current indoor scene image is obtained. The current space occupancy information includes the semantic category of the occupied object in the current indoor scene image and the voxel information of the occupied object.

[0037] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the method described in any of the first aspects.

[0038] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the method described in any of the first aspects.

[0039] This application provides a Gaussian coding model training method, an indoor occupancy prediction method, a device, and a medium. The model training method includes: extracting features from a sample indoor scene image to obtain a sample image feature map; updating randomly generated initial Gaussian features using a preset initial Gaussian coding model based on the sample image feature map to obtain target Gaussian features corresponding to the sample indoor scene image; obtaining predicted spatial occupancy information corresponding to the sample indoor scene image based on the target Gaussian features; and training the preset initial Gaussian coding model with parameter tuning based on the sample spatial occupancy information and the predicted spatial occupancy information to obtain the target Gaussian coding model. By using the target Gaussian coding model to predict the current Gaussian features of the current indoor scene image collected by the robot, and thus obtaining predicted spatial occupancy information, the accuracy is high, facilitating the robot's accurate perception of indoor scene occupancy. Attached Figure Description

[0040] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 A flowchart illustrating the Gaussian coding model training method provided in this application embodiment. Figure 1 ;

[0042] Figure 2 A flowchart illustrating the Gaussian coding model training method provided in this application embodiment. Figure 2 ;

[0043] Figure 3 A flowchart of the Gaussian coding model training method provided by the embodiment of the present application is shown in FIG. 1. Figure 3

[0044] Figure 4 A flowchart of the Gaussian coding model training method provided by the embodiment of the present application is shown in FIG. 1. Figure 4

[0045] Figure 5 A flowchart of the indoor occupancy prediction method provided by the embodiment of the present application is shown in FIG. 2.

[0046] Figure 6 A structural diagram of the Gaussian coding model training device provided by the embodiment of the present application is shown in FIG. 3.

[0047] Figure 7 A structural diagram of the indoor occupancy prediction device provided by the embodiment of the present application is shown in FIG. 4.

[0048] Figure 8 A structural diagram of the electronic device provided by the embodiment of the present application is shown in FIG. 5. DETAILED DESCRIPTION

[0049] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0050] In order to solve the problem that the existing occupancy prediction model is not accurate enough to predict the occupancy, and affects the execution of downstream tasks, the target Gaussian coding model is trained to predict the current Gaussian feature of the current indoor scene image collected by the robot, that is, to describe the space scene by using the Gaussian feature, and then to obtain the predicted space occupancy information, which is more accurate, and is convenient for the robot to accurately perceive the indoor scene occupancy, and is beneficial to the robot to execute downstream tasks.

[0051] Figure 1 A flowchart of the Gaussian coding model training method provided by the embodiment of the present application is shown in FIG. 1. Figure 1 The execution subject of the embodiment can be an electronic device, such as a notebook computer, a desktop computer, etc. ​​

[0052] As Figure 1 shown, the method can include:

[0053] S101, acquiring an indoor scene dataset.

[0054] The indoor scene dataset includes: a plurality of sample indoor scene images, wherein the sample indoor scene images are labeled with sample space occupancy information, and the sample space occupancy information includes: actual semantic categories of occupied objects in the sample indoor scene images and actual occupancy voxel information.

[0055] The indoor scene may, for example, be a scene containing objects such as a ceiling, a floor, a wall, a table, a sofa, etc., and different indoor scenes contain different object objects.

[0056] The occupied objects in the sample indoor scene images are objects in the occupied voxel grid in the three-dimensional field of view space corresponding to the sample indoor scene images, wherein the three-dimensional field of view space is a field of view space captured by a camera for the sample indoor scene images, that is, the sample indoor scene images are obtained by image acquisition of the three-dimensional field of view space of the indoor scene, and the three-dimensional field of view space is divided into a plurality of voxel grids, and the occupied objects are objects in the occupied voxel grid in the plurality of voxel grids.

[0057] The actual semantic categories of the occupied objects in the sample indoor scene images are used to indicate the categories of the occupied objects in the sample indoor scene images, such as a ceiling, a floor, a wall, a table, a sofa, etc., and the actual occupancy voxel information is used to indicate the voxel grid in which the occupied objects in the sample indoor scene images are located in the three-dimensional field of view space corresponding to the sample indoor scene images.

[0058] S102, performing feature extraction on the sample indoor scene images to obtain sample image feature maps corresponding to the sample indoor scene images.

[0059] A preset image encoder is used to perform feature extraction on the sample indoor scene images to obtain image features of the sample indoor scene images and generate sample image feature maps corresponding to the sample indoor scene images.

[0060] In some embodiments, multi-scale feature extraction is performed on the sample indoor scene images to obtain image features of the sample indoor scene images at different scales (different resolutions), and multi-scale sample image feature maps are generated based on the multi-scale image features, wherein the multi-scale features can capture features at different scales in the image and can provide a more comprehensive image representation.

[0061] It should be noted that the preset image encoder can be a pre-trained image encoder.

[0062] S103. Based on the feature map of the sample image, the randomly generated initial Gaussian features are updated using a preset initial Gaussian coding model to obtain the target Gaussian features corresponding to the sample indoor scene image.

[0063] Initial Gaussian features are randomly generated and used to provide an initial description of objects in the indoor scene from dimensions such as geometric information, opacity information, and semantic information (semantic category).

[0064] The sample image feature map and initial Gaussian features are used as input. A pre-defined initial Gaussian coding model is used to update the initial Gaussian features to obtain target Gaussian features. These target Gaussian features are used to describe objects within the indoor scene corresponding to the sample indoor scene image, considering dimensions such as geometric information, opacity, and semantic category. Both the initial and target Gaussian features can be 3D Gaussian features.

[0065] S104. Based on the target Gaussian features, obtain the predicted space occupancy information corresponding to the sample indoor scene image.

[0066] The predicted space occupancy information includes: the predicted semantic category of occupied objects in the sample indoor scene image and the predicted occupancy voxel information.

[0067] Among them, the predicted semantic category of occupied objects in the sample indoor scene image is used to indicate the category of occupied objects in the predicted sample indoor scene image, and the predicted voxel information is used to indicate the voxel grid where the occupied objects in the predicted sample indoor scene image are located in the corresponding three-dimensional field of view space.

[0068] In an optional implementation, a preset Gaussian voxel mapping module is used to perform voxel mapping on the target Gaussian features to obtain predicted space occupancy information.

[0069] The preset Gaussian voxel mapping module can be a pre-trained Gaussian voxel mapping module used to convert sparse 3D Gaussian representations into dense voxel grids. In this embodiment, the Gaussian voxel mapping module is used to map the target Gaussian features onto the voxel grid to obtain preset space occupancy information.

[0070] S105. Based on the sample space occupancy information and the predicted space occupancy information, the preset initial Gaussian coding model is trained by adjusting the parameters to obtain the target Gaussian coding model.

[0071] Based on the sample space occupancy information and the predicted space occupancy information, the model loss is calculated, and the model loss is used to tune the parameters of the preset initial Gaussian coding model until the model loss does not exceed the preset loss threshold. The Gaussian coding model obtained when the model loss does not exceed the preset loss threshold is taken as the target Gaussian coding model.

[0072] The specific selection of the preset loss threshold can be selected according to actual selection, and the embodiment does not particularly limit this.

[0073] In some embodiments, the semantic loss (semantic affinity loss) is calculated according to the actual semantic category and the predicted semantic category, the occupancy loss (geometric affinity loss) is calculated according to the actual occupancy voxel information and the predicted occupancy voxel information, and the model loss is calculated according to the semantic loss and the occupancy loss, wherein the model loss can be a weighted sum of the semantic loss and the occupancy loss.

[0074] It should be noted that, in some embodiments, the sample space occupancy information and the predicted space occupancy information can also be used to train the preset image encoder, the preset Gaussian voxel mapping module and the preset initial Gaussian encoding model, to obtain a target image encoder, a target Gaussian voxel mapping module and a target Gaussian encoding model.

[0075] In the Gaussian encoding model training method provided in the embodiment, the target Gaussian encoding model is obtained by training, so that the target Gaussian encoding model is used to predict the current Gaussian feature of the current indoor scene image collected by the robot, that is, the Gaussian feature is used to describe the space scene, and then the predicted space occupancy information is obtained, which has high accuracy and is convenient for the robot to accurately perceive the occupancy of the indoor scene.

[0076] Figure 2 Flowchart of the Gaussian encoding model training method provided in the embodiment of the present application Figure 2 As shown in the optional embodiment, the preset initial Gaussian encoding model includes an opacity-guided autoencoder and a geometric perception cross-encoder. Figure 3 The step S103 of updating the randomly generated initial Gaussian feature according to the sample image feature map by using the preset initial Gaussian encoding model can include:

[0077] S201, using an opacity-guided autoencoder to perform autoencoding on the initial Gaussian feature according to the opacity information in the initial Gaussian feature to obtain an encoded Gaussian feature.

[0078] The number of initial Gaussian features is multiple, that is, multiple initial Gaussian features are used to describe the preliminary description of the indoor scene. The initial Gaussian feature includes opacity information, and the opacity information is used to represent the visibility of the object in the indoor scene described by the initial Gaussian feature. The higher the opacity of the object, the stronger the visibility, and the more semantic information needs to be retained. The lower the opacity of the object, the lower the visibility, for example, the opacity of the foreground object (such as a chair) is higher than that of the background (such as a wall), and the semantic information of the chair should be retained in priority when overlapping.

[0079] The opacity-guided autoencoder is used to perform autoencoding on the initial high-frequency features according to the opacity information, to enhance the initial high-frequency features with high opacity (non-empty Gaussian) among the plurality of initial high-frequency features and weaken the initial high-frequency features with low opacity (empty Gaussian) among the plurality of initial high-frequency features, to obtain encoded high-frequency features.

[0080] In S202, the geometric perception cross encoder is used to perform geometric cross perception on the sample image feature map according to the encoded high-frequency features, to obtain target high-frequency features.

[0081] The encoded high-frequency features are input into the geometric perception cross encoder, and the image features of the sampled Gaussian points in the encoded high-frequency features are determined from the sample image feature map, and the image features of the sampled Gaussian points are geometrically fused and semantically fused according to the semantic information and the geometric information in the encoded high-frequency features, to obtain the target high-frequency features.

[0082] In an optional embodiment, the preset initial Gaussian encoding model further includes an optimization module.

[0083] The step S202 of using the geometric perception cross encoder to perform geometric cross perception on the sample image feature map according to the encoded high-frequency features to obtain the target high-frequency features can include:

[0084] The geometric perception cross encoder is used to perform geometric cross perception on the sample image feature map according to the encoded high-frequency features, to obtain cross-perception high-frequency features.

[0085] The optimization module is used to obtain the target high-frequency features according to the cross-perception high-frequency features.

[0086] The encoded high-frequency features are input into the geometric perception cross encoder, and the image features of the sampled Gaussian points in the encoded high-frequency features are determined from the sample image feature map, and the image features of the sampled Gaussian points are geometrically fused and semantically fused according to the semantic information and the geometric information in the encoded high-frequency features, to obtain the cross-perception high-frequency features, and then the optimization module is used to perform feature optimization on the cross-perception high-frequency features, to obtain the target high-frequency features.

[0087] It should be noted that the optimization module can optimize the geometric information, the opacity information and the semantic information in the cross-perception high-frequency features, so as to more accurately match the real structure and semantic information of the indoor scene corresponding to the sample indoor scene image.

[0088] In the Gaussian encoding model training method provided in the embodiment, the opacity-guided autoencoder is used to enhance the non-empty Gaussian and weaken the empty Gaussian, and the geometric perception cross encoder is used for semantic fusion and geometric fusion for model training, thereby improving the model training accuracy.

[0089] Figure 3A flowchart of a Gaussian coding model training method provided by an embodiment of the present application Figure 4 As shown in FIG. 2, in an optional embodiment, the opacity-guided autoencoder comprises a voxelization module and a plurality of sparse convolutional layers. Figure 4 The step S201 of using the opacity-guided autoencoder to perform autoencoding on the initial Gaussian features according to the opacity information in the initial Gaussian features to obtain the encoded Gaussian features can include:

[0090] S301, using the voxelization module to perform voxelization processing on the initial Gaussian features to obtain Gaussian features corresponding to a plurality of voxels.

[0091] The number of the initial Gaussian features is a plurality, and the Gaussian mean in the geometric information of one initial Gaussian feature corresponds to one Gaussian point in the three-dimensional space. All Gaussian points are aggregated to form a point cloud of the three-dimensional space, which is used to indicate the spatial distribution of the object in the three-dimensional space.

[0092] The voxelization module discretizes the point cloud into a plurality of voxels (sparse voxel grid) according to a preset voxel size, wherein the plurality of voxels are voxels containing Gaussian points, and empty voxels are not retained, and then the corresponding initial Gaussian features are assigned to the corresponding voxels to obtain Gaussian features corresponding to the corresponding voxels.

[0093] It should be noted that the Gaussian mean is used as the coordinate point of the voxel, and the semantic information and the opacity information in the initial Gaussian features are used as the attribute value of the voxel, and the sparse tensor of the voxel is generated by combining the two, wherein the sparse tensor of the voxel contains geometric information and semantic information and opacity information. Thus, the storage and calculation overheads are reduced compared with the dense tensor by using the spatial sparsity of the indoor scene (most areas are empty).

[0094] S302, sequentially using a plurality of sparse convolutional layers to add the spatial opacity information in the initial Gaussian features to the Gaussian features corresponding to the plurality of voxels layer by layer to generate the encoded Gaussian features.

[0095] The number of the plurality of sparse convolutional layers can be 2, 3, etc., which is not particularly limited in the embodiment. The input of the first sparse convolutional layer in the plurality of sparse convolutional layers is the Gaussian features corresponding to the plurality of voxels, the input of the second sparse convolutional layer is the output of the first sparse convolutional layer, and the input of the last sparse convolutional layer is the output of the last sparse convolutional layer. The encoded Gaussian features include a plurality of encoded Gaussian features corresponding to the plurality of voxels.

[0096] In some embodiments, for a sparse convolution layer, the opacity information is added to the high-dimensional features corresponding to the voxels by using the sparse convolution layer to perform convolution operation on the non-empty regions, thereby reducing the amount of calculation and memory occupation. In addition, the convolution operation can also enhance the high-dimensional features of voxels with high opacity (non-empty Gaussian) and weaken the high-dimensional features of voxels with low opacity (empty Gaussian).

[0097] The processing of the sparse convolution layer is represented as follows:

[0098] OGSPConv(Q, o) = SPConv3x3(Q) o (Sigmoid(o) + Sigmoid(SPConv1x1(Q)))

[0099] wherein Q is the high-dimensional feature corresponding to the voxel (feature matrix of the sparse tensor), o is the opacity information, SPConvkxk represents the sparse convolution of a (k x k) kernel, o represents the dot product operation, Sigmoid(o) represents the application of Sigmoid to the opacity o, which is mapped to (0, 1), and the higher the opacity, the closer the weight to 1, indicating that the corresponding Gaussian is more likely to be an object region, and the feature thereof needs to be strengthened, wherein Sigmoid is an activation function.

[0100] In some embodiments, the opacity-guided autoencoder includes a voxelization module, a first sparse convolution layer, and a second sparse convolution layer.

[0101] The multi-scale module is further included between the first sparse convolution layer and the second sparse convolution layer, wherein the multi-scale module is used to adjust the spatial receptive field of the first sparse convolution layer to capture information of different scales in the scene and more comprehensively describe the scene. For example, the receptive field of a single-layer 3 x 3 convolution is 3 x 3 pixels, and after multi-layer convolution stacking by using the multi-scale module, the spatial receptive field is 7 x 7 pixels.

[0102] In the Gaussian coding model training method provided in the embodiment, when multiple object objects overlap in space, they will jointly affect the semantic prediction of the same voxel. For example, a voxel may be misclassified as the semantic category of any overlapping Gaussian (such as a voxel containing both “chair” and “wall” Gaussians may be misjudged as a wall). Based on this, the opacity-guided autoencoder is used to strengthen the semantic contribution of the foreground object and suppress the background interference by using the opacity information, thereby achieving accurate classification of the voxel.

[0103] Figure 5 Flowchart of the Gaussian coding model training method provided in the embodiment Figure 6 As shown in FIG. 4, the Gaussian coding model training method provided in the embodiment includes the following steps. Figure 6As shown, in an optional embodiment, the geometry-aware cross-encoder includes a sampling module, a semantic mixing module, and a geometry mixing module.

[0104] The step S202 of obtaining the cross-aware high-frequency feature by performing geometry-aware cross-encoding on the sample image feature map according to the encoding high-frequency feature can include:

[0105] S401, using the sampling module, sampling the sample image feature map according to the encoding high-frequency feature to obtain the image feature of the sampling Gaussian point corresponding to the encoding high-frequency feature.

[0106] Using the sampling module, the encoding high-frequency feature is sampled to obtain a sampling Gaussian point, and the sampling Gaussian point is projected into the sample image feature map to obtain the image feature of the sampling Gaussian point.

[0107] Wherein, the encoding high-frequency feature is visualized as an ellipsoid, and the offset Δm0 is initialized, and the Gaussian geometry offset Δm is obtained by using matrix multiplication (MatMul) according to the encoding high-frequency feature and the scaling information s and rotation information r in the geometry information of the encoding high-frequency feature. It is expressed as follows:

[0108] Δm=MatMul(r,Δm0s)

[0109] Then, according to the Gaussian geometry offset Δm and the Gaussian mean in the geometry information of the encoding high-frequency feature, the sampling Gaussian point is determined. It is expressed as follows:

[0110] P=m+Δm

[0111] Wherein, m is the mean coordinate of the Gaussian mean, corresponding to the ellipsoid center point, and P is the position coordinate of the sampling Gaussian point. Wherein, the sampling Gaussian point is used to describe the geometric contour of the corresponding object, and the sampling Gaussian point is a 3D point, for example, for an encoding high-frequency feature representing a table, P can be an edge point of the table.

[0112] The camera intrinsic matrix and the camera extrinsic matrix when the indoor scene image is collected are obtained, and the sampling Gaussian point is projected into the sample image feature map according to the camera intrinsic matrix and the camera extrinsic matrix to obtain the image feature of the sampling Gaussian point.

[0113] It is expressed as follows:

[0114] Qp=Sampling(Q,π(P,E,K),F)

[0115] Wherein, Qp is the image feature of the sampling Gaussian point, π(P, E, K) is a projection function for converting the sampling Gaussian point into a 2D image coordinate, E is a camera extrinsic matrix, K is a camera intrinsic matrix, Sampling(Q, …, F) represents a sampling function for extracting the image feature corresponding to the 2D image coordinate from the sample image feature map as the image feature of the sampling Gaussian point according to the 2D image coordinate.

[0116] It should be noted that the process of Gaussian projection to the image is essentially to establish a cross-modal mapping between the 3D scene representation and the 2D visual perception, to realize geometric alignment, and to constrain the image feature. In the model training process, the representation gradually converges to match the object geometry and semantics, thereby realizing high-precision indoor scene occupancy prediction.

[0117] S402, using a semantic mixing module, performing semantic fusion according to the image feature of the sampling Gaussian point and the semantic information in the encoded Gaussian feature to obtain a semantic fusion feature of the sampling Gaussian point.

[0118] Wherein, the semantic information in the encoded Gaussian feature is used to strengthen the key semantic information of the image feature of the sampling Gaussian point.

[0119] Using the semantic mixing module, the semantic weight is obtained according to the semantic information in the encoded Gaussian feature, wherein the semantic weight is used to indicate the importance of the semantic information in the Gaussian encoded feature, and then the semantic fusion is performed according to the semantic weight and the image feature of the sampling Gaussian point to obtain the semantic fusion feature of the sampling Gaussian point, wherein the semantic fusion feature is the image feature of the sampling Gaussian point fused with the semantic weight.

[0120] is expressed as follows:

[0121] Qs = ReLU (LayerNorm (MatMul (Qp, Ws)))

[0122] Wherein, Ws is the semantic weight, MatMul represents matrix multiplication operation, LayerNorm represents layer normalization processing, ReLU represents activation function, and Qs is the semantic fusion feature.

[0123] S403, using a geometric mixing module, performing geometric fusion according to the semantic fusion feature of the sampling Gaussian point and the geometric information in the encoded Gaussian feature to obtain the geometric fusion feature of the sampling Gaussian point as the cross-perception Gaussian feature.

[0124] Wherein, the geometric information in the encoded Gaussian feature is used to strengthen the geometric information of the semantic fusion feature of the sampling Gaussian point.

[0125] The geometric weight is obtained according to geometric information in the encoded Gaussian feature by using a geometric mixing module, the geometric weight is used to indicate the importance of the geometric information in the Gaussian encoding feature, then the geometric fusion is performed according to the geometric weight and the semantic fusion feature of the sampled Gaussian point, to obtain the geometric fusion feature of the sampled Gaussian point, and the geometric fusion feature is the cross perception Gaussian feature, wherein the geometric fusion feature is the semantic fusion feature fused with the geometric weight.

[0126] is represented as follows:

[0127] Qg = ReLU (LayerNorm(MatMul(Wg, Qs)))

[0128] Wherein, Qg is the geometric fusion feature, and Wg is the geometric weight.

[0129] In the Gaussian encoding model training method provided in the embodiment, the image features of the sampled Gaussian points of semantic perception and geometric perception are obtained through the semantic weight and the geometric weight, the strengthening effect of the semantic information and the geometric information on the image features is highlighted. The model can more accurately grasp the semantic and geometric features of the objects in the scene, thereby realizing more accurate indoor scene occupancy prediction.

[0130] Figure 1 The flowchart of the indoor occupancy prediction method provided in the embodiment of the application, the execution subject of the embodiment can be an electronic device, such as a target robot.

[0131] As shown in Figure 6 , the method can include:

[0132] S501, obtaining a current indoor scene image collected by a target robot in a target indoor scene.

[0133] Wherein, the target robot can be a humanoid robot, and the target robot is in the target indoor scene, and the target indoor scene is any one of the above indoor scenes.

[0134] The current indoor image frame can be an image frame obtained by a camera of the target robot collecting the target indoor scene, and the camera of the target robot can be a monocular camera, and the current indoor image frame can be a monocular RGB image.

[0135] S502, performing feature extraction on the current indoor scene image to obtain a current image feature map corresponding to the current indoor scene image.

[0136] A preset image encoder is used to perform feature extraction on the current indoor scene image to obtain an image feature of the current indoor scene image, and a current image feature map corresponding to the current indoor scene image is generated.

[0137] In some embodiments, the target image encoder trained above is used to extract features from the current indoor scene image to obtain a current image feature map.

[0138] S503, according to the current image feature map, the initial Gaussian feature is updated by using the target Gaussian encoding model to obtain the current Gaussian feature corresponding to the current indoor scene image.

[0139] The opacity information in the randomly generated initial Gaussian feature is encoded by using the opacity-guided autoencoder in the target Gaussian encoding model to obtain an encoded Gaussian feature, and then the geometric cross perception of the current image feature map is performed according to the encoded Gaussian feature by using the geometric perception cross encoder in the target Gaussian encoding model to obtain the current Gaussian feature corresponding to the current indoor scene image. The target Gaussian encoding model is a model trained by using the above method.

[0140] For the implementation process of the current Gaussian feature, refer to the implementation process of the target Gaussian feature in the above occupancy world model training method, which will not be described here.

[0141] S504, according to the current Gaussian feature, the current spatial occupancy information corresponding to the current indoor scene image is obtained.

[0142] The current spatial occupancy information includes the semantic category of the occupied object in the current indoor scene image and the occupancy voxel information.

[0143] The semantic category of the occupied object in the current indoor scene image is used to indicate the category of the occupied object in the current indoor scene image, and the occupancy voxel information is used to indicate the voxel grid where the occupied object in the current indoor scene image is located in the three-dimensional field of view space corresponding to the current indoor scene image.

[0144] The current Gaussian feature is voxel mapped by using a preset Gaussian voxel mapping module to obtain the current spatial occupancy information corresponding to the current indoor scene image.

[0145] In some embodiments, the current Gaussian feature is voxel mapped by using the target Gaussian voxel mapping module trained above to obtain the current spatial occupancy information.

[0146] In this embodiment, the target Gaussian encoding model is used to predict the current Gaussian feature of the current indoor scene image collected by the robot, and then the predicted spatial occupancy information is obtained, which has high accuracy and is convenient for the robot to accurately perceive the occupancy of the indoor scene.

[0147] Figure 7 Structure of the Gaussian encoding model training device provided in the embodiment of the present application Figure 2The device can be integrated in an electronic device, such as a notebook computer, a desktop computer, etc. As shown in Figure 8 The device can include:

[0148] The acquisition module 601 is configured to acquire an indoor scene dataset, the indoor scene dataset including: a sample indoor scene image, wherein the sample indoor scene image is labeled with sample space occupation information, and the sample space occupation information includes actual semantic categories and actual occupied voxel information of an occupied object in the sample indoor scene image.

[0149] The processing module 602 is configured to perform feature extraction on the sample indoor scene image to obtain a sample image feature map corresponding to the sample indoor scene image.

[0150] The processing module 602 is further configured to update an initial Gaussian feature randomly generated according to the sample image feature map by using a preset initial Gaussian coding model to obtain a target Gaussian feature corresponding to the sample indoor scene image.

[0151] The acquisition module 601 is further configured to acquire predicted space occupation information corresponding to the sample indoor scene image according to the target Gaussian feature, wherein the predicted space occupation information includes predicted semantic categories and predicted occupied voxel information of the occupied object in the sample indoor scene image.

[0152] The processing module 602 is further configured to perform parameter training on the preset initial Gaussian coding model according to the sample space occupation information and the predicted space occupation information to obtain a target Gaussian coding model.

[0153] In an optional implementation, the preset initial Gaussian coding model includes an opacity-guided autoencoder and a geometry-aware cross-encoder.

[0154] The processing module 602 is specifically configured to:

[0155] The opacity-guided autoencoder is used to perform autoencoding on the initial Gaussian feature according to opacity information in the initial Gaussian feature to obtain an encoded Gaussian feature.

[0156] The geometry-aware cross-encoder is used to perform geometry-aware cross perception on the sample image feature map according to the encoded Gaussian feature to obtain the target Gaussian feature.

[0157] In an optional implementation, the preset initial Gaussian coding model further includes an optimization module.

[0158] The processing module 602 is specifically configured to:

[0159] The geometry-aware cross-encoder is used to perform geometry-aware cross perception on the sample image feature map according to the encoded Gaussian feature to obtain the cross-perception Gaussian feature.

[0160] The optimization module is used to obtain a target high-dimensional feature according to the cross-perception high-dimensional feature.

[0161] In an optional implementation, the opacity-guided autoencoder comprises a voxelization module and a plurality of sparse convolution layers.

[0162] The processing module 602 is specifically configured to:

[0163] The voxelization module is used to voxelize the initial high-dimensional feature to obtain high-dimensional features corresponding to a plurality of voxels.

[0164] The plurality of sparse convolution layers are sequentially used to add spatial opacity information in the initial high-dimensional feature to the high-dimensional features corresponding to the plurality of voxels layer by layer to generate an encoded high-dimensional feature.

[0165] In an optional implementation, the geometry-perception cross-encoder comprises a sampling module, a semantic mixing module and a geometry mixing module.

[0166] The processing module 602 is specifically configured to:

[0167] The sampling module is used to sample a sample image feature map according to the encoded high-dimensional feature to obtain an image feature of a sampled Gaussian point corresponding to the encoded high-dimensional feature.

[0168] The semantic mixing module is used to perform semantic fusion according to the image feature of the sampled Gaussian point and semantic information in the encoded high-dimensional feature to obtain a semantic fusion feature of the sampled Gaussian point.

[0169] The geometry mixing module is used to perform geometry fusion according to the semantic fusion feature of the sampled Gaussian point and geometry information in the encoded high-dimensional feature to obtain a geometry fusion feature of the sampled Gaussian point as the cross-perception high-dimensional feature.

[0170] In an optional implementation, the processing module 602 is specifically configured to:

[0171] The sampling module is used to sample the encoded high-dimensional feature to obtain a sampled Gaussian point, and project the sampled Gaussian point into a sample image feature map to obtain an image feature of the sampled Gaussian point.

[0172] In an optional implementation, the obtaining module 601 is specifically configured to:

[0173] The preset Gaussian voxel mapping module is used to voxel map the target high-dimensional feature to obtain predicted spatial occupancy information.

[0174] The description of the processing procedure of each module in the apparatus and the interaction procedure between the modules can refer to the related description in the above method embodiments, which will not be described in detail here.

[0175] Figure 8 The structure of the Gaussian coding model training device provided in the embodiment of the present application is shown in the figure ​ The device can be integrated in an electronic device, such as a target robot.

[0176] The acquisition module 701 is configured to acquire a current indoor scene image collected by a target robot in a target indoor scene.

[0177] The processing module 702 is configured to perform feature extraction on the current indoor scene image to obtain a current image feature map corresponding to the current indoor scene image.

[0178] The processing module 702 is further configured to update an initial Gaussian feature randomly generated according to the current image feature map by using a target Gaussian coding model to obtain a current Gaussian feature corresponding to the current indoor scene image, the target Gaussian coding model being a model trained by the method.

[0179] The acquisition module 701 is further configured to acquire current space occupancy information corresponding to the current indoor scene image according to the current Gaussian feature, the current space occupancy information including a semantic category of an occupied object in the current indoor scene image and occupancy voxel information.

[0180] The processing flow of each module in the device and the interaction flow between the modules will be described in the above method embodiments, and will not be described in detail here.

[0181] ​ The structure of the electronic device provided in the embodiment of the present application is shown in the figure ​ As shown in the figure, the device can include a processor 801, a memory 802, and a bus 803, the memory 802 storing machine-readable instructions executable by the processor 801, when the electronic device is running, the processor 801 and the memory 802 communicate through the bus 803, and the processor 801 executes the machine-readable instructions to perform the above method.

[0182] The embodiment of the present application further provides a computer-readable storage medium, the computer-readable storage medium storing a computer program, the computer program being executed by a processor to perform the above method.

[0183] In the embodiment of the present application, the computer program executed by the processor can also execute other machine-readable instructions to perform other methods described in the embodiments, for specific method steps and principles, refer to the description of the embodiments, and will not be described in detail here.

[0184] In the embodiments of the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. The embodiments described above are merely exemplary, for example, the division of the units is only a logical function division, and there can be another division manner in actual implementation; for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, and electrical, mechanical or other forms.

[0185] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.

[0186] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.

[0187] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0188] It should be noted that: similar reference numerals and letters in the following drawings represent similar items, and therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third" and the like are only used to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0189] Finally, it should be noted that the above-described embodiments are merely specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, but not to limit the same. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that any skilled person in the art can still modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some of the technical features, within the technical range disclosed by the present application. The modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. All should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for training a Gaussian coding model, characterized in that, include: An indoor scene dataset is obtained, comprising: sample indoor scene images, wherein the sample indoor scene images are labeled with sample space occupancy information, the sample space occupancy information comprising: the actual semantic category of the occupied objects in the sample indoor scene images and the actual voxel information of the occupied objects, the actual semantic category being used to indicate the category of the occupied objects in the sample indoor scene images, and the actual voxel information of the occupied objects being used to indicate the voxel grid in the corresponding three-dimensional field of view of the sample indoor scene images; Feature extraction is performed on the sample indoor scene image to obtain the sample image feature map corresponding to the sample indoor scene image; Based on the feature map of the sample image, the randomly generated initial Gaussian features are updated using a preset initial Gaussian coding model to obtain the target Gaussian features corresponding to the sample indoor scene image; Based on the target Gaussian features, the predicted space occupancy information corresponding to the sample indoor scene image is obtained. The predicted space occupancy information includes: the predicted semantic category of the occupied object in the sample indoor scene image and the predicted occupancy voxel information. The predicted semantic category is used to indicate the category of the occupied object in the predicted sample indoor scene image. The predicted occupancy voxel information is used to indicate the voxel grid where the occupied object in the predicted sample indoor scene image is located in the three-dimensional field space corresponding to the sample indoor scene image. Based on the sample space occupancy information and the prediction space occupancy information, the preset initial Gaussian coding model is trained by parameter tuning to obtain the target Gaussian coding model.

2. The method according to claim 1, characterized in that, The preset initial Gaussian coding model includes: an opacity-guided autoencoder and a geometry-aware cross encoder; The step of updating the randomly generated initial Gaussian features using a preset initial Gaussian coding model based on the feature map of the sample image to obtain the target Gaussian features corresponding to the sample indoor scene image includes: Using the opacity-guided autoencoder, the initial Gaussian features are autoencoded based on the opacity information in the initial Gaussian features to obtain encoded Gaussian features; The geometrically perceptive cross encoder is used to perform geometric cross-perception on the feature map of the sample image based on the encoded Gaussian features to obtain the target Gaussian features.

3. The method according to claim 2, characterized in that, The preset initial Gaussian coding model also includes: an optimization module; The step of employing the geometrically perceptive cross-encoder to perform geometric cross-perception on the feature map of the sample image based on the encoded Gaussian features to obtain the target Gaussian features includes: Using the geometrically perceptive cross encoder, geometric cross-perception is performed on the feature map of the sample image based on the encoded Gaussian features to obtain cross-perceptive Gaussian features; The optimization module is used to obtain the target Gaussian features based on the cross-perception Gaussian features.

4. The method according to claim 2, characterized in that, The opacity-guided autoencoder includes: a voxelization module, and multiple sparse convolutional layers; The process of employing the opacity-guided autoencoder to autoencode the initial Gaussian features based on the opacity information in the initial Gaussian features to obtain encoded Gaussian features includes: The voxelization module is used to voxelize the initial Gaussian features to obtain Gaussian features corresponding to multiple voxels. The spatial opacity information in the initial Gaussian features is added layer by layer to the Gaussian features corresponding to the multiple voxels to generate the encoded Gaussian features.

5. The method according to claim 3, characterized in that, The geometry-aware cross encoder includes: a sampling module, a semantic mixing module, and a geometric mixing module; The step of employing the geometrically perceptive cross encoder to perform geometric cross-perception on the feature map of the sample image based on the encoded Gaussian features, and obtaining cross-perceptive Gaussian features, includes: Using the sampling module, the sample image feature map is sampled according to the encoded Gaussian features to obtain the image features of the sampled Gaussian points corresponding to the encoded Gaussian features; Using the semantic fusion module, semantic fusion is performed based on the image features of the sampled Gaussian points and the semantic information in the encoded Gaussian features to obtain the semantic fusion features of the sampled Gaussian points; The geometric fusion module is used to perform geometric fusion based on the semantic fusion features of the sampled Gaussian points and the geometric information in the encoded Gaussian features, and the resulting geometric fusion features of the sampled Gaussian points are used as the cross-perceptual Gaussian features.

6. The method according to claim 5, characterized in that, The step of using the sampling module to sample the feature map of the sample image according to the encoded Gaussian features to obtain the image features of the sampled Gaussian points corresponding to the encoded Gaussian features includes: The sampling module is used to sample the encoded Gaussian features to obtain the sampled Gaussian points, and the sampled Gaussian points are projected onto the sample image feature map to obtain the image features of the sampled Gaussian points.

7. The method according to claim 1, characterized in that, The step of obtaining the predicted spatial occupancy information corresponding to the sample indoor scene image based on the target Gaussian features includes: A preset Gaussian voxel mapping module is used to perform voxel mapping on the target Gaussian features to obtain the predicted space occupancy information.

8. A method for predicting indoor occupancy, characterized in that, include: Acquire the current indoor scene image collected by the target robot within the target indoor scene; Feature extraction is performed on the current indoor scene image to obtain the current image feature map corresponding to the current indoor scene image; Based on the current image feature map, the randomly generated initial Gaussian features are updated using a target Gaussian coding model to obtain the current Gaussian features corresponding to the current indoor scene image. The target Gaussian coding model is a model trained using the method described in any one of claims 1-7. Based on the current Gaussian features, the current space occupancy information corresponding to the current indoor scene image is obtained. The current space occupancy information includes: the semantic category of the occupied object in the current indoor scene image and the voxel information of the occupied object. The semantic category of the occupied object is used to indicate the category of the occupied object in the current indoor scene image, and the voxel information of the occupied object is used to indicate the voxel grid where the occupied object in the current indoor scene image is located in the three-dimensional field space corresponding to the current indoor scene image.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Scene semantic occupancy prediction method based on probabilistic three-dimensional Gaussian

    CN120126116A