Scene understanding method, device, electronic device and computer-readable storage medium

Through the three-dimensional Gaussian points based on the three-dimensional Gaussian points, the problem of domain differences between the two-dimensional feature space and the three-dimensional feature space in the existing technology is solved, and efficient three-dimensional scene understanding is achieved.

CN119850992BActive Publication Date: 2025-06-03FALCON INNOVATIONS TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510318425.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-03
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing three-dimensional scene understanding methods are mainly focused on the two-dimensional semantic segmentation task, resulting in significant domain differences between the two-dimensional feature space and the three-dimensional feature space, resulting in limitations in language feature learning.

Method used

By directly implementing three-dimensional scene understanding based on three-dimensional Gaussian points, obtaining multi-frame depth images of the scene, constructing multiple three-dimensional Gaussian points, including high-dimensional language features and geometric features, clustering to obtain three-dimensional superprimitives, and fusing them to obtain a three-dimensional instance segmentation representation.

Benefits of technology

It effectively reduces the amount of parameters required for scene expression and semantic modeling, improves the accuracy and efficiency of three-dimensional scene understanding, and solves the problem of domain differences between two-dimensional feature space and three-dimensional feature space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850992B_ABST
    Figure CN119850992B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method, apparatus, electronic device, and computer-readable storage medium for scene understanding, relating to the technical field of three-dimensional scene understanding, aiming to solve the technical problem that three-dimensional scene understanding based on two-dimensional features is not accurate enough. The method includes: obtaining multiple depth images of a scene; constructing multiple three-dimensional Gaussian points of the scene according to the multiple depth images; the three-dimensional Gaussian points include: high-dimensional language features and geometric features; clustering the multiple three-dimensional Gaussian points according to the similarity of the high-dimensional language features and the similarity of the geometric features of the multiple three-dimensional Gaussian points to obtain multiple three-dimensional super primitives; fusing the multiple three-dimensional super primitives to obtain a three-dimensional instance segmentation representation of the scene. In this way, the present solution can directly implement three-dimensional scene understanding based on three-dimensional Gaussian points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of three-dimensional scene understanding, and specifically to a scene understanding method, device, electronic device, and computer-readable storage medium. Background Art

[0002] Three-dimensional scene understanding is a key task in the field of computer vision and is widely applied to scenarios such as intelligent robots, Augmented Reality (AR), and autonomous driving. Three-dimensional scene understanding is that a computer processes and analyzes scene data to understand the geometric structure and semantic information in the scene.

[0003] The three-dimensional scene understanding methods in the related art mainly focus on two-dimensional semantic segmentation tasks, and perform feature extraction and learning after rendering a three-dimensional scene onto a two-dimensional image plane. However, the domain difference between the two-dimensional feature space and the three-dimensional feature space is significant, resulting in limitations in language feature learning. Therefore, the three-dimensional scene understanding methods in the related art still need to be improved. Summary of the Invention

[0004] The embodiments of the present application provide a scene understanding method, device, electronic device, and computer-readable storage medium, which can directly achieve three-dimensional scene understanding based on three-dimensional Gaussian points.

[0005] In a first aspect, the embodiments of the present application provide a scene understanding method, including:

[0006] Obtain multiple frames of depth images of a scene;

[0007] Construct multiple three-dimensional Gaussian points of the scene according to the multiple frames of depth images; the three-dimensional Gaussian points include: high-dimensional language features and geometric features;

[0008] Cluster the multiple three-dimensional Gaussian points according to the similarity of the high-dimensional language features and the similarity of the geometric features of the multiple three-dimensional Gaussian points to obtain multiple three-dimensional super primitives;

[0009] Fuse the multiple three-dimensional super primitives to obtain a three-dimensional instance segmentation representation of the scene.

[0010] In one embodiment, the constructing multiple three-dimensional Gaussian points of the scene according to the multiple frames of depth images includes:

[0011] Construct multiple initial three-dimensional Gaussian points of the scene according to the pixel points in the multiple frames of depth images; the initial three-dimensional Gaussian points include: the geometric features and semantic encodings; the semantic encodings of the initial three-dimensional Gaussian points are learned from the latent space pyramid three-plane features;

[0012] Input the semantic encoding of the initial 3D Gaussian points into a 3D decoder to obtain the high-dimensional language features;

[0013] Combine the high-dimensional language features and the initial 3D Gaussian points to obtain the 3D Gaussian points of the scene.

[0014] In one embodiment, the training steps of the 3D decoder at least include:

[0015] Obtain the semantic encoding samples of the initial 3D Gaussian point samples;

[0016] According to the semantic encoding samples of the initial 3D Gaussian point samples, determine the high-dimensional language feature samples corresponding to the initial 3D Gaussian point samples and the confidence levels of the high-dimensional language feature samples;

[0017] Input the initial 3D Gaussian point samples into an initial 3D decoder to obtain the predicted high-dimensional language features of the initial 3D Gaussian point samples;

[0018] According to the confidence levels of the high-dimensional language feature samples, and the cosine similarity between the predicted high-dimensional language features corresponding to the initial 3D Gaussian point samples and the high-dimensional language feature samples, construct a loss function;

[0019] Based on the loss function, train the initial 3D decoder to obtain the trained 3D decoder.

[0020] In one embodiment, the determining the high-dimensional language feature samples corresponding to the initial 3D Gaussian point samples according to the semantic encoding samples of the initial 3D Gaussian point samples includes:

[0021] Obtain multi-frame depth image samples of the scene samples corresponding to the initial 3D Gaussian point samples;

[0022] Extract the language feature maps of the multi-frame depth image samples;

[0023] Perform multi-angle projection on the initial 3D Gaussian point samples to obtain the two-dimensional pixel position information corresponding to the initial 3D Gaussian point samples in the multi-frame depth image samples;

[0024] Obtain the visual relationships of the initial 3D Gaussian point samples in the multi-frame depth image samples;

[0025] According to the visual relationships, determine the target depth image samples corresponding to the initial 3D Gaussian point samples; the initial 3D Gaussian point samples are visible in the target depth image samples;

[0026] Determine the two-dimensional language features corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample according to the language feature maps of multiple frames of the depth image samples and the two-dimensional pixel position information corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample;

[0027] Perform average pooling on the two-dimensional language features corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample to obtain the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample.

[0028] In one embodiment, determining the confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample according to the semantic coding sample of the initial three-dimensional Gaussian point sample includes:

[0029] Obtain the variance of the two-dimensional language features corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample;

[0030] Determine the confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample according to the visual relationship and the variance of the two-dimensional language features corresponding to the initial three-dimensional Gaussian point sample.

[0031] In one embodiment, the fusing of multiple three-dimensional super-primitives to obtain the three-dimensional instance segmentation representation of the scene includes:

[0032] Determine the affine coefficients between multiple three-dimensional super-primitives according to the projections of the three-dimensional Gaussian points of multiple three-dimensional super-primitives and the instance segmentation masks of multiple frames of the depth images;

[0033] Fuse multiple three-dimensional super-primitives with affine coefficients less than the affine coefficient threshold according to the affine coefficients between multiple three-dimensional super-primitives to obtain multiple preliminary three-dimensional instance segmentation representations;

[0034] Perform iterative fusion on multiple preliminary three-dimensional instance segmentation representations to obtain the three-dimensional instance segmentation representation of the scene.

[0035] In one embodiment, the determining of the affine coefficients between multiple three-dimensional super-primitives according to the three-dimensional Gaussian projections of multiple three-dimensional super-primitives and the instance segmentation masks of multiple frames of the depth images includes:

[0036] Project the three-dimensional super-primitives to obtain multiple two-dimensional images corresponding to multiple frames of the depth images for the three-dimensional super-primitives;

[0037] Determine the distribution of the instance segmentation masks in multiple two-dimensional images according to the two-dimensional images and the instance segmentation masks of multiple frames of the depth images;

[0038] Determine the affine coefficients among multiple three-dimensional super-primitives under the depth image of the same frame according to the distribution of the instance segmentation masks of the two-dimensional images corresponding to the multiple three-dimensional super-primitives in the depth image of the same frame;

[0039] Obtain the visible ratio of the two-dimensional image in the corresponding depth image;

[0040] Determine the affine coefficients among multiple three-dimensional super-primitives according to the affine coefficients among multiple three-dimensional super-primitives under the depth images of multiple same frames and the visible ratios of the two-dimensional image in the depth images of multiple same frames.

[0041] In a second aspect, an embodiment of the present application provides a scene understanding device, including:

[0042] An acquisition module, configured to acquire multiple frames of depth images of a scene;

[0043] A construction module, configured to construct multiple three-dimensional Gaussian points of the scene according to multiple frames of the depth images; the three-dimensional Gaussian points include: high-dimensional language features and geometric features;

[0044] A clustering module, configured to cluster multiple three-dimensional Gaussian points according to the similarity of high-dimensional language features and the similarity of geometric features of multiple three-dimensional Gaussian points to obtain multiple three-dimensional super-primitives;

[0045] A fusion module, configured to fuse multiple three-dimensional super-primitives to obtain a three-dimensional instance segmentation representation of the scene.

[0046] In one embodiment, the construction module includes:

[0047] A construction unit, configured to construct multiple initial three-dimensional Gaussian points of the scene according to the pixel points in multiple frames of the depth images; the initial three-dimensional Gaussian points include: the geometric features and semantic encodings; the semantic encodings of the initial three-dimensional Gaussian points are learned from the implicit space pyramid three-plane features;

[0048] An input unit, configured to input the semantic encodings of the initial three-dimensional Gaussian points into a three-dimensional decoder to obtain the high-dimensional language features;

[0049] A combination unit, configured to combine the high-dimensional language features and the initial three-dimensional Gaussian points to obtain the three-dimensional Gaussian points of the scene.

[0050] In one embodiment, the training steps of the three-dimensional decoder at least include:

[0051] Obtain semantic encoding samples of initial three-dimensional Gaussian point samples;

[0052] Determine the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample and the confidence of the high-dimensional language feature sample according to the semantic coding sample of the initial three-dimensional Gaussian point sample;

[0053] Input the initial three-dimensional Gaussian point sample into the initial three-dimensional decoder to obtain the predicted high-dimensional language feature of the initial three-dimensional Gaussian point sample;

[0054] Construct a loss function according to the confidence of the high-dimensional language feature sample, and the cosine similarity between the predicted high-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample and the high-dimensional language feature sample;

[0055] Train the initial three-dimensional decoder based on the loss function to obtain the trained three-dimensional decoder.

[0056] In one embodiment, the determining the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample according to the semantic coding sample of the initial three-dimensional Gaussian point sample includes:

[0057] Obtain multiple-frame depth image samples of the scene sample corresponding to the initial three-dimensional Gaussian point sample;

[0058] Extract the language feature maps of multiple frames of the depth image samples;

[0059] Perform multi-angle projection on the initial three-dimensional Gaussian point sample to obtain the two-dimensional pixel position information corresponding to the initial three-dimensional Gaussian point sample in multiple frames of the depth image samples;

[0060] Obtain the visible relationship of the initial three-dimensional Gaussian point sample in multiple frames of the depth image samples;

[0061] Determine the target depth image sample corresponding to the initial three-dimensional Gaussian point sample according to the visible relationship; the initial three-dimensional Gaussian point sample is visible in the target depth image sample;

[0062] Determine the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample according to the language feature maps of multiple frames of the depth image samples and the two-dimensional pixel position information corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample;

[0063] Perform average pooling on the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample to obtain the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample.

[0064] In one embodiment, the determining the confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample according to the semantic coding sample of the initial three-dimensional Gaussian point sample includes:

[0065] Obtain the variance of the two-dimensional language features corresponding to the initial three-dimensional Gaussian point samples in the target depth image samples;

[0066] Determine the confidence of the high-dimensional language feature samples corresponding to the initial three-dimensional Gaussian point samples according to the visual relationship and the variance of the two-dimensional language features corresponding to the initial three-dimensional Gaussian point samples.

[0067] In one embodiment, the fusion module includes:

[0068] A determination unit, configured to determine the affine coefficients between multiple three-dimensional super-primitives according to the projections of the three-dimensional Gaussian points of the multiple three-dimensional super-primitives and the instance segmentation masks of multiple frames of the depth images;

[0069] A fusion unit, configured to fuse multiple three-dimensional super-primitives with affine coefficients less than an affine coefficient threshold according to the affine coefficients between multiple three-dimensional super-primitives, to obtain multiple preliminary three-dimensional instance segmentation representations;

[0070] An iteration unit, configured to iteratively fuse multiple preliminary three-dimensional instance segmentation representations to obtain a three-dimensional instance segmentation representation of the scene.

[0071] In one embodiment, the determination unit includes:

[0072] A projection sub-unit, configured to project the three-dimensional super-primitive to obtain multiple two-dimensional images of the three-dimensional super-primitive corresponding to multiple frames of the depth images;

[0073] A distribution determination sub-unit, configured to determine the distribution of the instance segmentation masks in multiple two-dimensional images according to the two-dimensional images and the instance segmentation masks of multiple frames of the depth images;

[0074] A coefficient determination sub-unit, configured to determine the affine coefficients between multiple three-dimensional super-primitives under the depth images of the same frame according to the distribution of the instance segmentation masks of the two-dimensional images corresponding to the multiple three-dimensional super-primitives in the depth images of the same frame;

[0075] A ratio acquisition sub-unit, configured to acquire the visual ratio of the two-dimensional image in the corresponding depth image;

[0076] An affine coefficient determination sub-unit, configured to determine the affine coefficients between multiple three-dimensional super-primitives according to the affine coefficients between multiple three-dimensional super-primitives under multiple depth images of the same frame and the visual ratios of the two-dimensional images in multiple depth images of the same frame.

[0077] In a third aspect, an embodiment of the present application further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps in the above-mentioned scene understanding method are implemented.

[0078] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned scene understanding method are implemented.

[0079] In a fifth aspect, an embodiment of the present application further provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners described in the embodiments of the present application.

[0080] In summary, in the embodiments of the present application, multiple three-dimensional Gaussian points of a scene can be directly constructed based on multi-frame depth images of the scene, and the three-dimensional Gaussian points include high-dimensional language features and geometric features. Furthermore, multiple three-dimensional Gaussian points can be clustered according to the similarity of the high-dimensional language features and the similarity of the geometric features of the three-dimensional Gaussian points to obtain multiple three-dimensional super primitives, and the multiple three-dimensional super primitives are fused to obtain a three-dimensional instance segmentation representation of the scene. In this way, three-dimensional scene understanding is directly achieved based on three-dimensional Gaussian points, and the problem of domain differences between the two-dimensional feature space and the three-dimensional feature space is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] In order to more clearly illustrate the technical solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0082] Figure 1 is a schematic diagram of the steps of the scene understanding method provided by an embodiment of the present application;

[0083] Figure 2 is a schematic diagram of the structure of the scene understanding device provided by an embodiment of the present application;

[0084] Figure 3 is a schematic diagram of the structure of the electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0085] Next, the technical solutions in the present application will be clearly and completely described in conjunction with the accompanying drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0086] 3D scene understanding can support downstream tasks such as object detection, path planning, and virtual scene interaction, providing comprehensive environmental perception capabilities for agents. Especially in the open-vocabulary scene understanding task, the system needs to be able to understand diverse semantic information, which poses higher requirements for modeling capabilities and semantic expression.

[0087] In recent years, neural rendering techniques, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), have shown remarkable performance in the novel view synthesis task. 3DGS has received particular attention for its fast training and rendering capabilities, as well as its explicit scene representation based on point clouds. Related technologies usually combine Vision Language Models (VLM) and utilize two-dimensional pixel-level semantic features to achieve open-vocabulary scene understanding. However, most of the methods in related technologies still mainly focus on two-dimensional semantic segmentation tasks, extracting and learning features after rendering to a two-dimensional image plane. Although these methods have achieved certain success in two-dimensional semantic segmentation tasks, in 3D scene understanding tasks, there are still problems such as inaccurate 3D language feature learning and lack of instance-level information modeling.

[0088] In related technologies, the feature representation of discrete Gaussian basis elements may destroy the inherent smoothness of semantically similar objects. At the same time, due to the feature accumulation method based on alpha-blending, the domain difference between the two-dimensional feature space and the three-dimensional feature space is significant, resulting in limitations in language feature learning. In addition, the feature compression and distillation process in the two-dimensional space will further weaken the discriminative ability of features. Thus, it will cause the problem of inaccurate 3D language feature learning.

[0089] In related technologies, object detection is usually performed through the similarity heatmap of language features and text queries. However, this method cannot distinguish multi-instance objects with the same semantics, resulting in inconsistent instance-level segmentation results and causing the problem of lack of instance-level information modeling, while instance-level information is crucial for general semantic 3D scene understanding.

[0090] The reasons for the inaccurate three-dimensional language feature learning and the lack of instance-level information modeling are as follows: Most of the related technologies stay at the feature distillation in the two-dimensional space, ignoring the interaction relationship between semantic and geometric features in the three-dimensional space. At the same time, the lack of instance information is mainly due to the lack of an effective three-dimensional graph structure modeling mechanism, which cannot use semantic consistency and geometric relationships to perform instance-level partitioning of the scene.

[0091] Based on the above problems, an embodiment of the present application proposes a scene understanding method, which is based on three-dimensional Gaussian projection and aims to fundamentally improve the ability of three-dimensional scene semantic and instance-level understanding.

[0092] The three-dimensional panoramic segmentation technology of the related technology has the following problems: In multi-view scene reconstruction, a large number of three-dimensional points or voxels need to be explicitly stored, resulting in huge storage overhead and low running efficiency; poor adaptability to open vocabulary, the methods of the related technology rely on a fixed label set, and the adaptability to unseen categories or open vocabulary is insufficient, making it difficult to be generalized to complex actual scenes; insufficient performance of joint segmentation of semantics and geometry, lacking an effective semantic, geometric and multi-view feature fusion mechanism, resulting in inaccurate segmentation results.

[0093] The goal of an embodiment of the present application is to achieve efficient and accurate three-dimensional panoramic segmentation in an open vocabulary environment, and solve the problems of storage overhead, open vocabulary adaptation and semantic geometry joint optimization by improving the scene representation and segmentation algorithm.

[0094] To achieve the above goal, an embodiment of the present application provides the following technical solutions:

[0095] Use three-dimensional Gaussian points for scene expression, and reconstruct the geometric and appearance information of the scene. Instead of directly encoding independent high-frequency features for each three-dimensional Gaussian, use low-dimensional multi-resolution tri-planes to model the smooth semantic characteristics of the three-dimensional scene, and use a three-dimensional feature decoder to lift the low-dimensional semantic encoding to a high-dimensional language feature space, so as to be able to perform scene understanding and segmentation of open vocabulary. In this way, the number of parameters required for scene expression and semantic modeling can be effectively reduced.

[0096] To achieve three-dimensional instance segmentation, use the geometric normal vectors and high-dimensional language features reconstructed by three-dimensional Gaussian points as the basis, and use millions of three-dimensional Gaussians in the scene to construct the vertices of graph clustering, thereby reducing the number of parameters of the clustering algorithm and accelerating the clustering speed. And use the segmentation model to obtain the segmentation masks (masks) from different perspectives, so as to construct the affine relationship between the graph vertices. Based on the relationship between the constructed graph vertices and affine edges, perform progressive clustering to obtain a three-dimensional consistent instance segmentation result.

[0097] Therefore, the technical solution proposed in an embodiment of the application can effectively solve the problems of storage overhead, open vocabulary adaptability, and segmentation accuracy in the current three-dimensional panoramic segmentation field, and has broad practical application value.

[0098] In one embodiment, as Figure 1 shown, a scene understanding method is provided. Although the logical order is shown in the step schematic diagram, in some cases, the steps shown or described can be executed in an order different from that shown in the drawings. Specifically, the scene understanding method can be applied to an electronic device, which can be a terminal or a server. Among them, the terminal can include, but is not limited to, one or more of a smart phone, a tablet computer, a portable computer, a desktop computer, and a vehicle-mounted computer. The server can be a physical server or a cloud server providing various cloud services. It should be noted that the application does not limit the number of terminals or servers. According to the implementation needs, there can be any number of terminals or servers. For example, the server can be a single server or a server cluster composed of multiple servers, etc.

[0099] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments.

[0100] According to Figure 1 the scene understanding method shown, the method at least includes steps S110 to S140, which are introduced in detail as follows:

[0101] In step S110, multiple frames of depth images of the scene are obtained.

[0102] The scene can be any three-dimensional scene. A depth image is an image with depth information. Depth information refers to the distance information of the object being photographed from the photographing device. The depth image can be an image taken by a color depth (Red Green BlueDepth, RGB-D) camera. Multiple frames of depth images can be multiple frames of depth images continuously taken by a color depth camera.

[0103] In step S120, multiple three-dimensional Gaussian points of the scene are constructed according to the multiple frames of depth images; the three-dimensional Gaussian points include: high-dimensional language features and geometric features.

[0104] Three-dimensional Gaussian points are points in three-dimensional space that follow a Gaussian distribution. Three-dimensional Gaussian points include high-dimensional language features and geometric features. High-dimensional language features are obtained by decoding low-dimensional semantic encodings and have a higher dimension than low-dimensional semantic encodings. Among them, the semantic encoding is learned from the implicit space pyramid three-plane features. Geometric features include position features, rotation features, and / or scale features, etc. In an embodiment of the application, the scene is characterized by three-dimensional Gaussian points.

[0105] Compared with low-dimensional semantic encoding, high-dimensional language features have stronger feature expression capabilities. A higher dimension means a larger feature space, which can encode more information. Therefore, based on high-dimensional language features, richer and more detailed semantic information can be captured, such as more complex textures, shapes, or context relationships. High-dimensional language features have a finer-grained discrimination ability and can make more refined distinctions between similar but different semantics; for example, based on high-dimensional language features, the specific breed of a cat can be distinguished, while based on low-dimensional language features, it may only be possible to distinguish whether an animal is a cat. High-dimensional language features may contain more abstract information. High-dimensional language features can encode higher-level abstract information, such as the function of an object, the semantic relationship of a scene, etc., which is very important for complex tasks (such as image generation, semantic segmentation).

[0106] The three-dimensional Gaussian points are constructed based on multiple frames of depth images of the captured scene. From the multiple frames of depth images, the image color and depth sequences can be obtained. Optionally, a depth sequence can be a sequence composed of the depth information of an object in multiple frames of depth images. When there are multiple objects, there can be multiple depth sequences. Optionally, a depth sequence can be a sequence composed of the depth information of pixel points representing the same physical point in multiple frames of depth images.

[0107] Based on the image color and depth sequences, using the structure from motion algorithm, by analyzing the depth images captured by the color-depth camera at different positions, the six-degree-of-freedom pose of the color-depth camera can be calculated. The six-degree-of-freedom pose of the color-depth camera refers to the position and orientation of the color-depth camera in three-dimensional space, specifically including three translational degrees of freedom and three rotational degrees of freedom. The three translational degrees of freedom refer to the translational degrees of freedom along the three rectangular coordinate axes x, y, and z, and the three rotational degrees of freedom refer to the rotational degrees of freedom around the three coordinate axes x, y, and z.

[0108] Optionally, to save storage overhead and improve running efficiency, each depth image can be sampled to obtain multiple sampled pixel points. Based on the six-degree-of-freedom pose of the camera and the depth information of the sampled pixel points, projecting back into three-dimensional space, the sampled point cloud corresponding to the sampled pixel points can be obtained. For example, 20% of the pixel points can be uniformly sampled from each depth image to obtain multiple sampled pixel points, and then the sampled sparse point cloud can be obtained based on the multiple sampled pixel points. The information of the point cloud corresponding to the pixel point is determined according to the pixel point.

[0109] Initialize the 3D Gaussian points based on the sampled point cloud to obtain multiple original 3D Gaussian points, with each original 3D Gaussian point corresponding to a sampled point cloud. The color feature, opacity feature, position feature, rotation feature, and / or scale feature, etc. of each original 3D Gaussian point can be determined. The color feature and position feature of the original 3D Gaussian point can be determined according to the color feature and position feature of the corresponding point cloud. The opacity of the original 3D Gaussian point can be initialized to 0.5, and the rotation feature can be initialized to the identity matrix. The scale feature of the original 3D Gaussian point can be determined according to the distance between the original 3D Gaussian point and its surrounding adjacent 3D Gaussian points.

[0110] Assign a semantic encoding to the original 3D Gaussian point to obtain the initial 3D Gaussian point. The dimension of the semantic encoding of the initial 3D Gaussian point can be 6, and this semantic encoding is learned from the latent space pyramid tri-planar features. The latent space pyramid tri-planar contains fine-grained encoding features and coarse-grained encoding features; the fine-grained encoding feature is an optimization variable with a shape of [3, Hf , Wf , 6]; the coarse-grained encoding feature is an optimization variable with a shape of [3, Hc , Wc , 6]. Among them, 3 represents the three decomposition planes of xy, xz, and yz, 6 represents the feature dimension of the semantic encoding, Hf , Wf , Hc , Wc respectively represent the plane resolutions of the fine-grained and coarse-grained. The optimization variable refers to the parameter or variable that needs to be adjusted by the optimization algorithm during the model training process.

[0111] The initial 3D Gaussian points can be used to represent the 3D scene compactly and efficiently, which can reduce the storage requirements. At the same time, encoding the latent feature field using the pyramid tri-planar structure can enhance the semantic expression ability of the 3D scene.

[0112] Input the semantic encoding of the initial 3D Gaussian point into the trained 3D decoder, and the high-dimensional language feature corresponding to the initial 3D Gaussian point output by the 3D decoder can be obtained. Combining the high-dimensional language feature corresponding to the initial 3D Gaussian point with the initial 3D Gaussian point can obtain the 3D Gaussian points of the scene. The original 3D Gaussian points do not have semantic encoding or high-dimensional language features; the initial 3D Gaussian points have semantic encoding; the 3D Gaussian points have high-dimensional language features.

[0113] In step S130, cluster the multiple 3D Gaussian points according to the similarity of the high-dimensional language features and the similarity of the geometric features of the multiple 3D Gaussian points to obtain multiple 3D super primitives.

[0114] The similarity of the high-dimensional language features between multiple 3D Gaussian points can be determined by calculating the cosine similarity or Euclidean distance of the high-dimensional language features between the multiple 3D Gaussian points. The similarity of the geometric features between the multiple 3D Gaussian points can be calculated by a Gaussian kernel function or other distance metric methods. According to the similarity of the high-dimensional language features and the similarity of the geometric features of the multiple 3D Gaussian points, the multiple 3D Gaussian points are clustered to obtain multiple 3D super primitives.

[0115] Optionally, a language-guided Graph Cuts algorithm can be used to cluster the multiple 3D Gaussian points to obtain multiple 3D super primitives. Optionally, spectral clustering, hierarchical clustering, or a K-means clustering algorithm can be adopted to cluster the multiple 3D Gaussian points to obtain multiple 3D super primitives.

[0116] In one embodiment, two 3D Gaussian points with the smallest sum of the similarity of the high-dimensional language features and the similarity of the geometric features can be clustered into one class to obtain a cluster. The mean of the high-dimensional language features of the 3D Gaussian points in the cluster is determined as the high-dimensional language feature of the clustering center, and the mean of the geometric features of the 3D Gaussian points in the cluster is determined as the geometric feature of the clustering center. A high-dimensional language feature similarity threshold and a geometric feature similarity threshold are obtained. It is sequentially determined whether each of the remaining 3D Gaussian points belongs to the existing clusters. When the similarity of the high-dimensional language feature of the 3D Gaussian point and the high-dimensional language feature of the clustering center is greater than the high-dimensional language feature similarity threshold, and the similarity of the geometric feature of the 3D Gaussian point and the geometric feature of the clustering center is greater than the geometric feature similarity threshold, the 3D Gaussian point is clustered into the cluster corresponding to the clustering center, and the high-dimensional language feature and the geometric feature of the 3D Gaussian point are used to update the high-dimensional language feature and the geometric feature of the clustering center. When the 3D Gaussian point cannot be clustered into the cluster corresponding to the clustering center, another cluster can be generated according to the 3D Gaussian point, and it is continued to determine whether each of the remaining 3D Gaussian points belongs to the existing clusters until the clustering of all 3D Gaussian points is completed, and each cluster is determined as a 3D super primitive.

[0117] The 3D Gaussian points in the 3D super primitive have geometric and semantic consistency, and the 3D Gaussian points in the 3D super primitive have similar geometric features and high-dimensional language features. The 3D super primitive can be used as a vertex unit for subsequent graph clustering.

[0118] In step S140, the multiple 3D super primitives are fused to obtain a 3D instance segmentation representation of the scene.

[0119] Using the Segment Anything Model (SAM), obtain the instance segmentation mask of the depth image. The instance segmentation mask can distinguish different objects in the depth image and determine the types of each object. According to the three-dimensional Gaussian projections of multiple three-dimensional hyper-primitives and the instance segmentation masks of multiple frames of depth images, determine the affine coefficients between the multiple three-dimensional hyper-primitives. The affine coefficients can determine the specific form of the transformation.

[0120] Take the three-dimensional hyper-primitives as the vertices in the graph structure, and take the affine coefficients between the three-dimensional hyper-primitives as the edges in the graph structure. According to the affine coefficients between the multiple three-dimensional hyper-primitives, fuse the multiple three-dimensional hyper-primitives with affine coefficients less than the affine coefficient threshold to obtain multiple preliminary three-dimensional instance segmentation representations.

[0121] Take the preliminary three-dimensional instance segmentation representations as the vertices in the new graph structure, and take the affine coefficients between the preliminary three-dimensional instance segmentation representations as the edges in the new graph structure, and fuse the multiple preliminary three-dimensional instance segmentation representations. Repeat the fusion multiple times to achieve step-by-step iterative fusion from local to global, and finally obtain the three-dimensional instance segmentation representation of the scene, complete the three-dimensional instance segmentation, and achieve three-dimensional scene understanding.

[0122] Adopting the technical solution of the embodiment of the present application, multiple three-dimensional Gaussian points of the scene can be directly constructed according to multiple frames of depth images of the scene, and the three-dimensional Gaussian points include high-dimensional language features and geometric features. Furthermore, according to the similarity of the high-dimensional language features and the similarity of the geometric features of the three-dimensional Gaussian points, cluster the multiple three-dimensional Gaussian points to obtain multiple three-dimensional hyper-primitives, and fuse the multiple three-dimensional hyper-primitives to obtain the three-dimensional instance segmentation representation of the scene. In this way, three-dimensional scene understanding is directly realized based on three-dimensional Gaussian points, and the problem of domain difference between the two-dimensional feature space and the three-dimensional feature space is solved.

[0123] On the basis of the above technical solution, as an embodiment, the constructing multiple three-dimensional Gaussian points of the scene according to multiple frames of the depth image may include: constructing multiple initial three-dimensional Gaussian points of the scene according to the pixel points in multiple frames of the depth image; the initial three-dimensional Gaussian points include: the geometric features and semantic encodings; the semantic encodings of the initial three-dimensional Gaussian points are learned from the implicit space pyramid three-plane features; input the semantic encodings of the initial three-dimensional Gaussian points into a three-dimensional decoder to obtain the high-dimensional language features; combine the high-dimensional language features and the initial three-dimensional Gaussian points to obtain the three-dimensional Gaussian points of the scene.

[0124] Uniformly sample each depth image to obtain multiple sampled pixel points. According to the six-degree-of-freedom pose of the camera and the depth information of the sampled pixel points, the depth image can be projected back into three-dimensional space to obtain the sampled point cloud corresponding to the sampled pixel points. Determine the information of the point cloud corresponding to the pixel point according to the pixel point. Initialize the three-dimensional Gaussian points according to the sampled point cloud to obtain multiple original three-dimensional Gaussian points, and each original three-dimensional Gaussian point corresponds to a sampled point cloud. The color feature, opacity feature, position feature, rotation feature, and / or scale feature, etc. of each original three-dimensional Gaussian point can be determined. Determine the geometric feature of the original three-dimensional Gaussian point according to the position feature, rotation feature, and / or scale feature of the original three-dimensional Gaussian point.

[0125] Assign a semantic encoding to the original three-dimensional Gaussian point to obtain the initial three-dimensional Gaussian point. The semantic encoding is a representation of semantics and contains semantic information. The dimension of the semantic encoding of the initial three-dimensional Gaussian point can be 6, and this semantic encoding is learned from the latent space pyramid triplane features. The latent space pyramid triplane contains fine-grained encoding features and coarse-grained encoding features; the fine-grained encoding feature is an optimization variable with a shape of [3, Hf , Wf , 6]; the coarse-grained encoding feature is an optimization variable with a shape of [3, Hc , Wc , 6]. Among them, 3 represents the three decomposition planes of xy, xz, and yz, 6 represents the feature dimension of the semantic encoding, Hf , Wf , Hc , Wc respectively represent the plane resolutions of the fine-grained and coarse-grained.

[0126] The three planes of the latent space pyramid include the xy, xz, and yz decomposition planes, and data sampling is performed on these three planes. For each position in the three-dimensional scene, information is obtained from different perspectives of the three planes to comprehensively capture the semantic features of each initial three-dimensional Gaussian point in three-dimensional space. A pyramid-shaped feature structure is constructed on each decomposition plane. By downsampling the original data to different degrees, feature layers with different resolutions are obtained, forming a pyramid shape. The low-level coarse-grained encoded features have a larger receptive field and can capture overall and macroscopic semantic information; the high-level fine-grained encoded features have a smaller receptive field and can retain more detailed information. The features extracted from the three decomposition planes contain spatial information in different directions. These features are fused, and the features from the three decomposition planes are combined together to form a feature representation that includes the three decomposition planes. The fusion method can be element-wise addition or concatenation. In the pyramid structure, features at different levels have different semantic granularities. To make full use of this information, cross-level feature fusion can be performed, fusing the low-level coarse-grained features with the high-level fine-grained features. A neural network model can be used to learn the mapping relationship between the features of the latent space pyramid three planes and the semantic encoding of the initial three-dimensional Gaussian points. The neural network model can take the fused pyramid three-plane features as input and, through multi-layer operations such as convolution, pooling, and attention, further abstract and refine the features, gradually learning the semantic encoding. For example, the convolution layer in a convolutional neural network can automatically extract local patterns in the features, the pooling layer can reduce the dimension of the features and extract the main features, and the attention mechanism can dynamically focus on the importance of different features, thereby learning semantic information. By training the neural network model, the semantic encoding can have better discriminability and representativeness.

[0127] Inputting the semantic encoding of the initial three-dimensional Gaussian point into a pre-trained three-dimensional decoder can obtain the high-dimensional language features corresponding to the initial three-dimensional Gaussian point output by the three-dimensional decoder. Combining the high-dimensional language features corresponding to the initial three-dimensional Gaussian point with the initial three-dimensional Gaussian point can obtain the three-dimensional Gaussian points of the scene. The original three-dimensional Gaussian points do not have semantic encoding or high-dimensional language features; the initial three-dimensional Gaussian points have semantic encoding; the three-dimensional Gaussian points have high-dimensional language features.

[0128] High-dimensional language features have a higher dimension than low-dimensional semantic encodings. Compared with low-dimensional semantic encodings, high-dimensional language features have stronger feature expression capabilities. A higher dimension means a larger feature space, which can encode more information. Therefore, based on high-dimensional language features, richer and more detailed semantic information can be captured, such as more complex textures, shapes, or context relationships. High-dimensional language features have a finer-grained discrimination ability and can make more refined distinctions between similar but different semantics. High-dimensional language features may contain more abstract information. High-dimensional language features can encode higher-level abstract information, such as the functions of objects, semantic relationships in scenes, etc., which are very important for complex tasks (such as image generation, semantic segmentation).

[0129] Adopting the technical solution of the embodiment of the present application, the initial three-dimensional Gaussian points can be used to represent the three-dimensional scene compactly and efficiently, thereby reducing the storage requirements. At the same time, the pyramid three-plane structure is used to encode the latent feature field, which can enhance the semantic expression ability of the three-dimensional scene; the high-dimensional language features can be obtained through the three-dimensional decoder, and richer and more detailed semantic information can be obtained, which is beneficial to subsequent three-dimensional scene understanding based on the three-dimensional Gaussian points containing high-dimensional language features.

[0130] In another embodiment, a deep learning network specialized in processing point cloud data can be used to obtain the semantic encoding of the initial three-dimensional Gaussian points; the point cloud corresponding to the initial three-dimensional Gaussian points is input into the deep learning network. The deep learning network automatically extracts the local and global features of the point cloud through operations such as multi-layer convolution and pooling, learns the semantic information of the point cloud, and determines the semantic encoding of the initial three-dimensional Gaussian points corresponding to the point cloud according to the semantic information of the point cloud.

[0131] In another embodiment, the high-dimensional language features can be obtained by referring to the method for obtaining the high-dimensional language feature samples of the initial three-dimensional Gaussian point samples described later.

[0132] On the basis of the above technical solution, as an embodiment, the training steps of the three-dimensional decoder at least include: obtaining the semantic encoding samples of the initial three-dimensional Gaussian point samples; determining the high-dimensional language feature samples corresponding to the initial three-dimensional Gaussian point samples and the confidence of the high-dimensional language feature samples according to the semantic encoding samples of the initial three-dimensional Gaussian point samples; inputting the initial three-dimensional Gaussian point samples into the initial three-dimensional decoder to obtain the predicted high-dimensional language features of the initial three-dimensional Gaussian point samples; constructing a loss function according to the confidence of the high-dimensional language feature samples, and the cosine similarity between the predicted high-dimensional language features corresponding to the initial three-dimensional Gaussian point samples and the high-dimensional language feature samples; training the initial three-dimensional decoder based on the loss function to obtain the trained three-dimensional decoder.

[0133] The method for obtaining the initial three-dimensional Gaussian point samples can refer to the method for obtaining the initial three-dimensional Gaussian points described above. The method for obtaining the semantic coding samples of the initial three-dimensional Gaussian point samples can refer to the method for obtaining the semantic coding of the initial three-dimensional Gaussian points described above.

[0134] Based on the semantic coding samples of the initial three-dimensional Gaussian point samples, the high-dimensional language feature samples corresponding to the initial three-dimensional Gaussian point samples can be determined. Based on the semantic coding samples of the initial three-dimensional Gaussian point samples, the confidence of the high-dimensional language feature samples corresponding to the initial three-dimensional Gaussian point samples can be determined. The specific method for determining the high-dimensional language feature samples corresponding to the initial three-dimensional Gaussian point samples and the confidence of the high-dimensional language feature samples will be described in detail later. The confidence of the high-dimensional language feature samples can characterize the accuracy of the high-dimensional language feature samples.

[0135] The initial three-dimensional decoder is a three-dimensional decoder to be trained. The initial three-dimensional decoder can predict high-dimensional language features, but the initial three-dimensional decoder has not been trained well yet, so the predicted high-dimensional language features are not accurate enough. Inputting the initial three-dimensional Gaussian point samples into the initial three-dimensional decoder can obtain the predicted high-dimensional language features corresponding to the initial three-dimensional Gaussian point samples output by the initial three-dimensional decoder.

[0136] Construct a loss function based on the high-dimensional language feature samples, and use the loss function to perform supervised training on the initial three-dimensional decoder so that the predicted high-dimensional language features output by the initial three-dimensional decoder can approach the high-dimensional language feature samples. Optionally, a loss function can be constructed according to the confidence of the high-dimensional language feature samples, as well as the cosine similarity between the predicted high-dimensional language features corresponding to the initial three-dimensional Gaussian point samples and the high-dimensional language feature samples. With the goal of minimizing this loss function, the initial three-dimensional decoder is trained to obtain a trained three-dimensional decoder.

[0137] Optionally, the loss function can be determined by the following formula:

[0138] ;

[0139] where, represents the loss function, represents the number of initial three-dimensional Gaussian point samples, represents the th confidence of the high-dimensional language feature samples corresponding to the initial three-dimensional Gaussian point samples, represents the predicted high-dimensional language feature samples, represents the high-dimensional language feature samples corresponding to the initial three-dimensional Gaussian point samples, represents the cosine distance.

[0140] By adopting the technical solution of the embodiment of the present application, a three-dimensional decoder can be pre-trained. Thus, when obtaining the high-dimensional language features of the three-dimensional Gaussian points, the three-dimensional decoder can be used to quickly obtain the high-dimensional language features of the three-dimensional Gaussian points, thereby improving the efficiency of obtaining high-dimensional language features.

[0141] Based on the above technical solution, as an embodiment, determining the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample according to the semantic coding sample of the initial three-dimensional Gaussian point sample may include: obtaining multi-frame depth image samples of the scene sample corresponding to the initial three-dimensional Gaussian point sample; extracting the language feature maps of the multi-frame depth image samples; performing multi-angle projection on the initial three-dimensional Gaussian point sample to obtain the two-dimensional pixel position information corresponding to the initial three-dimensional Gaussian point sample in the multi-frame depth image samples; obtaining the visible relationship of the initial three-dimensional Gaussian point sample in the multi-frame depth image samples; determining the target depth image sample corresponding to the initial three-dimensional Gaussian point sample according to the visible relationship; the initial three-dimensional Gaussian point sample is visible in the target depth image sample; determining the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample according to the language feature maps of the multi-frame depth image samples and the two-dimensional pixel position information corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample; performing average pooling on the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample to obtain the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample.

[0142] The scene sample can be any three-dimensional scene. By using a color depth camera to photograph the scene sample, multi-frame depth image samples are obtained. The initial three-dimensional Gaussian point sample can be obtained from the multi-frame depth image samples. The semantic segmentation model can be used to extract features from the depth image samples to obtain a language feature map with a shape of H×W×D corresponding to each frame of the depth image sample, where D represents the dimension of the extracted language features, and H×W represents the size. Optionally, the semantic segmentation model can be a Language-driven semantic segmentation (LSeg) model.

[0143] Based on multiple-frame depth image samples, an initial three-dimensional Gaussian point sample is obtained. According to the angles of the depth image samples, the initial three-dimensional Gaussian point sample is projected at multiple angles to obtain the two-dimensional pixel position information corresponding to each initial three-dimensional Gaussian point sample in the multiple-frame depth image samples. According to the two-dimensional pixel position information corresponding to the initial three-dimensional Gaussian point sample in the multiple-frame depth image samples, the distance information of the two-dimensional pixels corresponding to the initial three-dimensional Gaussian point sample in the multiple-frame depth image samples can be determined. According to the distance information of the two-dimensional pixels corresponding to the initial three-dimensional Gaussian point sample in the depth image sample, and the depth distance of the initial three-dimensional Gaussian point sample under the viewing angle corresponding to the depth image sample, the visible relationship of the initial three-dimensional Gaussian point sample under the viewing angle corresponding to the depth image sample can be determined. When the distance of the two-dimensional pixel corresponding to the initial three-dimensional Gaussian point sample in the depth image sample is less than the depth distance of the initial three-dimensional Gaussian point sample under the viewing angle corresponding to the depth image sample, and the difference between the depth distance of the initial three-dimensional Gaussian point sample under the viewing angle corresponding to the depth image sample and the distance of the two-dimensional pixel corresponding to the initial three-dimensional Gaussian point sample in the depth image sample is less than the visibility threshold, it indicates that the initial three-dimensional Gaussian point sample is occluded, and the visible relationship under the viewing angle corresponding to the depth image sample is invisible.

[0144] According to the visible relationships of the respective initial three-dimensional Gaussian point samples in the multiple-frame depth image samples, the target depth images corresponding to the respective initial three-dimensional Gaussian point samples are determined. The visible relationship of the initial three-dimensional Gaussian point sample in the corresponding target depth image is visible.

[0145] According to the two-dimensional pixel position information corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample, and the language feature map of the target depth image sample corresponding to the initial three-dimensional Gaussian point sample, the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the corresponding target depth image sample is determined. Average pooling is performed on the two-dimensional language features corresponding to the initial three-dimensional Gaussian point sample in multiple target depth image samples to obtain the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample.

[0146] By adopting the technical solution of the embodiment of the present application, the high-dimensional language feature sample of the initial three-dimensional Gaussian point sample can be determined according to the visible relationship of the initial three-dimensional Gaussian point sample at multiple viewing angles and the two-dimensional language features of the initial three-dimensional Gaussian point sample at multiple viewing angles, so that the calculated high-dimensional language feature sample of the initial three-dimensional Gaussian point sample can be used as the learning target of the three-dimensional decoder subsequently to quickly obtain the high-dimensional language features.

[0147] Based on the above technical solution, as an embodiment, determining the confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample according to the semantic coding sample of the initial three-dimensional Gaussian point sample may include: obtaining the variance of the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample; determining the confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample according to the visual relationship and the variance of the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample.

[0148] Obtain the variance of the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the visible target depth image sample. Since an initial three-dimensional Gaussian point sample corresponds to a physical point, the two-dimensional language features of this initial three-dimensional Gaussian point sample in multiple target depth image samples should be the same or close. Therefore, if the variance of the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the visible target depth image sample is large, the confidence of the high-dimensional language feature sample obtained based on the two-dimensional language feature of this initial three-dimensional Gaussian point sample is not high.

[0149] Based on the visual relationship of the initial three-dimensional Gaussian point sample and the variance of the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample, the confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample can be calculated. Optionally, the confidence can be calculated according to the following formula:

[0150] ;

[0151] where, represents the confidence, represents the number of visible times, represents the variance of the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample; the number of visible times represents the number of target depth image samples corresponding to the initial three-dimensional Gaussian point sample.

[0152] By adopting the technical solution of the embodiment of the present application, the confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample can be calculated according to the variance of the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample and the visual relationship, so that the three-dimensional decoder can learn more high-confidence and more accurate high-dimensional language feature samples during learning, thereby improving the performance of the three-dimensional decoder.

[0153] Based on the above technical solutions, as an embodiment, the fusing of the multiple three-dimensional super primitives to obtain the three-dimensional instance segmentation representation of the scene may include: determining the affine coefficients between the multiple three-dimensional super primitives according to the projections of the three-dimensional Gaussian points of the multiple three-dimensional super primitives and the instance segmentation masks of multiple frames of the depth images; fusing the multiple three-dimensional super primitives with affine coefficients less than the affine coefficient threshold according to the affine coefficients between the multiple three-dimensional super primitives to obtain multiple preliminary three-dimensional instance segmentation representations; and performing iterative fusion on the multiple preliminary three-dimensional instance segmentation representations to obtain the three-dimensional instance segmentation representation of the scene.

[0154] Multiple frames of depth images can be respectively input into the segmentation model to obtain the instance segmentation masks of each frame of depth image. The instance segmentation masks can distinguish different objects in the depth image and determine the types of each object. The labels of the instance segmentation masks corresponding to different types of objects are different.

[0155] Project the three-dimensional Gaussian points in the multiple three-dimensional super primitives into multiple frames of depth images. According to the projections of the three-dimensional Gaussian points in the multiple three-dimensional super primitives in the multiple frames of depth images and the instance segmentation masks of the multiple frames of depth images, the distribution of the labels of the instance segmentation masks of the projections of the three-dimensional super primitives under the multiple frames of depth images can be determined.

[0156] According to the distribution situations corresponding to the multiple three-dimensional super primitives, the affine coefficients between the multiple three-dimensional super primitives can be obtained by calculating the JS divergence (Jensen-Shannon divergence) of different distributions. The affine coefficients can determine the specific form of the transformation.

[0157] Taking the three-dimensional super primitives as the vertices in the graph structure and the affine coefficients between the three-dimensional super primitives as the edges in the graph structure, fusing the multiple three-dimensional super primitives with affine coefficients less than the affine coefficient threshold according to the affine coefficients between the multiple three-dimensional super primitives to obtain multiple preliminary three-dimensional instance segmentation representations. The affine coefficient threshold can be set according to actual requirements. The initial three-dimensional instance segmentation representation includes multiple three-dimensional super primitives with close distances.

[0158] Taking the preliminary three-dimensional instance segmentation representations as the vertices in the new graph structure and referring to the method of calculating the affine coefficients between the three-dimensional super primitives, calculate the affine coefficients between the initial three-dimensional instance segmentation representations. Taking the affine coefficients between the preliminary three-dimensional instance segmentation representations as the edges in the new graph structure, fuse the multiple preliminary three-dimensional instance segmentation representations with affine coefficients less than the affine coefficient threshold.

[0159] Repeat the fusion process multiple times to achieve step-by-step iterative fusion from local to global, and finally obtain the three-dimensional instance segmentation representation of the scene, complete the three-dimensional instance segmentation, and achieve three-dimensional scene understanding.

[0160] Adopting the technical solution of the embodiment of the present application, based on the graph structure modeling mechanism, the three-dimensional super-primitives are gradually fused, and the process of gradual fusion is based on the affine coefficients, and the affine coefficients are determined according to the distribution of the labels of the instance segmentation masks of the projections of the three-dimensional super-primitives in multiple frames of depth images. Therefore, the affine coefficients utilize semantic information. Generally speaking, the fusion process utilizes both graph structure information and semantic information, synthesizes the interaction relationship between semantics and geometric features in three-dimensional space, realizes accurate fusion, and realizes the instance-level division of the scene by using semantic consistency and geometric relationships.

[0161] On the basis of the above technical solution, as an embodiment, the determining the affine coefficients between multiple three-dimensional super-primitives according to the three-dimensional Gaussian projections of the multiple three-dimensional super-primitives and the instance segmentation masks of multiple frames of the depth images may include: projecting the three-dimensional super-primitives to obtain multiple two-dimensional images corresponding to the multiple frames of the depth images of the three-dimensional super-primitives; determining the distribution of the instance segmentation masks in the multiple two-dimensional images according to the two-dimensional images and the instance segmentation masks of multiple frames of the depth images; determining the affine coefficients between the multiple three-dimensional super-primitives under the depth images of the same frame according to the distribution of the instance segmentation masks of the two-dimensional images corresponding to the multiple three-dimensional super-primitives in the depth images of the same frame; obtaining the visible ratio of the two-dimensional images in the corresponding depth images; and determining the affine coefficients between the multiple three-dimensional super-primitives according to the affine coefficients between the multiple three-dimensional super-primitives under the depth images of multiple same frames and the visible ratios of the two-dimensional images in the depth images of multiple same frames.

[0162] When calculating the affine coefficients between multiple three-dimensional super-primitives, the three-dimensional Gaussian points in the multiple three-dimensional super-primitives are projected into multiple frames of depth images. According to the projections of the three-dimensional Gaussian points in the three-dimensional super-primitives in the multiple frames of depth images and the instance segmentation masks of the multiple frames of depth images, the distribution of the labels of the instance segmentation masks of the projections of the three-dimensional super-primitives in the multiple frames of depth images can be determined.

[0163] According to the distribution of the labels of the instance segmentation masks of multiple three-dimensional super-primitives in the k-th frame, the affine coefficients between the multiple three-dimensional super-primitives in the k-th frame can be obtained by calculating the JS divergence of different distributions. Optionally, the affine coefficients between the three-dimensional super-primitives in the k-th frame can be calculated by the following formula:

[0164] ;

[0165] wherein, represents the affine coefficient between the i-th three-dimensional super-primitive and the j-th three-dimensional super-primitive in the k-th frame, Characterize the distribution of the i-th three-dimensional super primitive at the k-th frame, Characterize the distribution of the j-th three-dimensional super primitive at the k-th frame, = ( + ) / 2, Characterize and average distribution, Characterize the label.

[0166] Since the instance segmentation of a single frame may not be accurate, resulting in noise in the calculation of the affine coefficient, the information of multiple frames is fused according to the visible ratio of the three-dimensional super primitive in different frames to obtain the final multi-frame fusion affine coefficient. Among them, the visible ratio can be used to determine the number of three-dimensional Gaussian points visible in the depth image of the k-th frame through the method for determining the visible relationship described above, and the visible ratio of the two-dimensional image projected by the three-dimensional super primitive in the depth image is obtained by dividing the number of three-dimensional Gaussian points visible in the depth image of the k-th frame by the total number of three-dimensional Gaussian points in the three-dimensional super primitive.

[0167] According to the affine coefficients between multiple three-dimensional super primitives in the depth images of multiple identical frames, and the visible ratios of the two-dimensional images in the depth images of multiple identical frames, the affine coefficients between multiple three-dimensional super primitives are determined according to the following formula:

[0168] ;

[0169] Among them, Characterize the affine coefficient between the i-th three-dimensional super primitive and the j-th three-dimensional super primitive, Characterize the affine coefficient between the i-th three-dimensional super primitive and the j-th three-dimensional super primitive at the k-th frame, where k represents the number of frames, Characterize the visible ratio of the i-th three-dimensional super primitive in the depth image of the k-th frame, Characterize the visible ratio of the j-th three-dimensional super primitive in the depth image of the k-th frame.

[0170] Adopting the technical solution of the embodiment of the present application, the final multi-frame fusion affine coefficient is obtained according to the affine coefficients of the three-dimensional super primitive in multiple frames, and whether to fuse the three-dimensional super primitive is determined based on the multi-frame fusion affine coefficient, and a more accurate fusion result can be obtained.

[0171] Adopting the technical solution of the embodiment of the present application, there is no need to explicitly store the semantic features of the three-dimensional space. By using multi-resolution three planes to model the language feature information of the scene, the number of parameters of the model and segmentation can be effectively reduced. In addition, three-dimensional hyper-primitives are constructed through the reconstructed geometric features and high-dimensional language features, and an affine relationship between the three-dimensional hyper-primitives is constructed using a segmentation model, thereby performing clustering to obtain a three-dimensional instance segmentation result. Through mathematical derivation and experimental comparison, it is proved that the scene representation scheme based on three-dimensional hyper-primitives has better progressive segmentation convergence, and the calculation method of the affine coefficient as the edge effectively reduces the influence of noise and improves the stability and accuracy of segmentation.

[0172] By introducing the joint optimization of semantics and geometry in the embodiment of the present application, in multiple open-vocabulary segmentation tasks, both three-dimensional semantic segmentation and instance segmentation of the embodiment of the present application have achieved effective improvement, significantly superior to the method with a fixed label set, especially excellent in terms of unseen categories. In addition, the embodiment of the present application also supports open-vocabulary query and instance segmentation. Therefore, the technical solution proposed by the embodiment of the present application can effectively solve the problems of storage overhead, open-vocabulary adaptability, and segmentation accuracy in the current three-dimensional panoramic segmentation field, and has broad practical application value.

[0173] To facilitate better implementation of the scene understanding method of the present application, the present application also provides a scene understanding device based on the above scene understanding method. The meanings of the nouns are the same as those in the above scene understanding method, and the specific implementation details can refer to the description in the method embodiment.

[0174] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of the scene understanding device provided by the embodiment of the present application. The scene understanding device includes:

[0175] An acquisition module 201, configured to acquire multiple frames of depth images of a scene;

[0176] A construction module 202, configured to construct multiple three-dimensional Gaussian points of the scene according to the multiple frames of depth images; the three-dimensional Gaussian points include: high-dimensional language features and geometric features;

[0177] A clustering module 203, configured to cluster the multiple three-dimensional Gaussian points according to the similarity of the high-dimensional language features and the similarity of the geometric features of the multiple three-dimensional Gaussian points to obtain multiple three-dimensional hyper-primitives;

[0178] A fusion module 204, configured to fuse the multiple three-dimensional hyper-primitives to obtain a three-dimensional instance segmentation representation of the scene.

[0179] In one embodiment, the construction module 202 includes:

[0180] A construction unit, configured to construct multiple initial 3D Gaussian points of the scene according to pixel points in multiple frames of the depth image; the initial 3D Gaussian points include: the geometric features and semantic encodings; the semantic encodings of the initial 3D Gaussian points are learned from the implicit space pyramid tri-plane features.

[0181] An input unit, configured to input the semantic encodings of the initial 3D Gaussian points into a 3D decoder to obtain the high-dimensional language features.

[0182] A combination unit, configured to combine the high-dimensional language features and the initial 3D Gaussian points to obtain the 3D Gaussian points of the scene.

[0183] In one embodiment, the training steps of the 3D decoder at least include:

[0184] Obtain semantic encoding samples of initial 3D Gaussian point samples.

[0185] According to the semantic encoding samples of the initial 3D Gaussian point samples, determine the high-dimensional language feature samples corresponding to the initial 3D Gaussian point samples and the confidence levels of the high-dimensional language feature samples.

[0186] Input the initial 3D Gaussian point samples into an initial 3D decoder to obtain the predicted high-dimensional language features of the initial 3D Gaussian point samples.

[0187] Construct a loss function according to the confidence levels of the high-dimensional language feature samples, and the cosine similarity between the predicted high-dimensional language features corresponding to the initial 3D Gaussian point samples and the high-dimensional language feature samples.

[0188] Based on the loss function, train the initial 3D decoder to obtain the trained 3D decoder.

[0189] In one embodiment, the determining the high-dimensional language feature samples corresponding to the initial 3D Gaussian point samples according to the semantic encoding samples of the initial 3D Gaussian point samples includes:

[0190] Obtain multiple frames of depth image samples of the scene samples corresponding to the initial 3D Gaussian point samples.

[0191] Extract the language feature maps of multiple frames of the depth image samples.

[0192] Perform multi-angle projection on the initial 3D Gaussian point samples to obtain the two-dimensional pixel position information corresponding to the initial 3D Gaussian point samples in multiple frames of the depth image samples.

[0193] Obtain the visual relationships of the initial 3D Gaussian point samples in multiple frames of the depth image samples.

[0194] Determine the target depth image sample corresponding to the initial three-dimensional Gaussian point sample according to the visual relationship; the initial three-dimensional Gaussian point sample is visible in the target depth image sample.

[0195] Determine the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample according to the language feature maps of multiple frames of the depth image samples and the two-dimensional pixel position information corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample.

[0196] Perform average pooling on the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample to obtain the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample.

[0197] In one embodiment, determining the confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample according to the semantic coding sample of the initial three-dimensional Gaussian point sample includes:

[0198] Obtain the variance of the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample in the target depth image sample.

[0199] Determine the confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample according to the visual relationship and the variance of the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample.

[0200] In one embodiment, the fusion module 204 includes:

[0201] A determination unit, configured to determine the affine coefficients between multiple three-dimensional super-primitives according to the projections of the three-dimensional Gaussian points of the multiple three-dimensional super-primitives and the instance segmentation masks of multiple frames of the depth images.

[0202] A fusion unit, configured to fuse multiple three-dimensional super-primitives with affine coefficients less than the affine coefficient threshold according to the affine coefficients between the multiple three-dimensional super-primitives to obtain multiple preliminary three-dimensional instance segmentation representations.

[0203] An iteration unit, configured to iteratively fuse the multiple preliminary three-dimensional instance segmentation representations to obtain the three-dimensional instance segmentation representation of the scene.

[0204] In one embodiment, the determination unit includes:

[0205] A projection sub-unit, configured to project the three-dimensional super-primitive to obtain multiple two-dimensional images corresponding to multiple frames of the depth image of the three-dimensional super-primitive.

[0206] A distribution determination subunit, configured to determine the distribution of instance segmentation masks in multiple two-dimensional images according to the two-dimensional image and instance segmentation masks of multiple frames of the depth images;

[0207] A coefficient determination subunit, configured to determine the affine coefficients between multiple three-dimensional super primitives under the depth images of the same frame according to the distribution of instance segmentation masks of two-dimensional images of the depth images corresponding to multiple three-dimensional super primitives respectively in the same frame;

[0208] A ratio acquisition subunit, configured to acquire the visible ratio of the two-dimensional image in the corresponding depth image;

[0209] An affine coefficient determination subunit, configured to determine the affine coefficients between multiple three-dimensional super primitives according to the affine coefficients between multiple three-dimensional super primitives under multiple depth images of the same frame and the visible ratio of the two-dimensional image in multiple depth images of the same frame.

[0210] By adopting the technical solution of the embodiment of the present application, multiple three-dimensional Gaussian points of a scene can be directly constructed according to multiple frames of depth images of the scene, and the three-dimensional Gaussian points include high-dimensional language features and geometric features. Furthermore, multiple three-dimensional Gaussian points can be clustered according to the similarity of high-dimensional language features and the similarity of geometric features of the three-dimensional Gaussian points to obtain multiple three-dimensional super primitives, and multiple three-dimensional super primitives are fused to obtain a three-dimensional instance segmentation representation of the scene. In this way, three-dimensional scene understanding is directly realized based on three-dimensional Gaussian points, and the problem of domain difference between the two-dimensional feature space and the three-dimensional feature space is solved. For the specific limitation of the scene understanding device, reference can be made to the limitation of the scene understanding method in the above text, which will not be elaborated here. Each module in the above scene understanding device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or independent of the processor, or stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0211] In addition, the present application further provides an electronic device, as Figure 3 shown, which shows a schematic structural diagram of the electronic device involved in the present application. Specifically:

[0212] The electronic device may include a processor 301 with one or more processing cores and a memory 302 with one or more computer-readable storage media and other components. Those skilled in the art can understand that Figure 3 the structural diagram of the electronic device shown in does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine some components, or different component arrangements. Among them:

[0213] The processor 301 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 302, and by invoking the data stored in the memory 302, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 301 either.

[0214] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and data processing by running the software programs and modules stored in the memory 302. The memory 302 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 302 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 302 can also include a memory controller to provide the processor 301 with access to the memory 302.

[0215] In one embodiment, the electronic device further includes a power supply 303 for supplying power to each component. Preferably, the power supply 303 can be logically connected to the processor 301 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 303 can also include any components such as one or more DC or AC power supplies, a recharge system, a power device debugging circuit, a power converter or inverter, and a power status indicator.

[0216] In one embodiment, the electronic device may further include an input unit 304. The input unit 304 can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0217] Although not shown, the electronic device may further include a display unit and the like, which will not be elaborated here. Specifically, in this embodiment, the processor 301 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 302 according to the following instructions, and the processor 301 will run the application programs stored in the memory 302, so as to implement the steps in any of the scenario understanding methods provided in the embodiments of the present application.

[0218] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0219] In one embodiment, an electronic device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the method described in any embodiment of the present application is implemented.

[0220] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment of the present application is implemented.

[0221] In some embodiments, a computer program product is further proposed, including a computer program or instruction. When the computer program or instruction is executed by a processor, the method described in any embodiment of the present application is implemented.

[0222] For the specific implementation of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0223] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0224] Therefore, the present application provides a computer-readable storage medium, on which a computer program is stored. The computer program can be loaded by a processor to execute the steps in any of the scenario understanding methods provided by the present application.

[0225] For the specific implementation of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0226] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0227] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the scene understanding methods provided by this application, the beneficial effects achievable by any of the scene understanding methods provided by this application can be realized. For details, see the previous embodiments and will not be elaborated here.

[0228] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the element.

[0229] The above has introduced in detail a scene understanding method, device, electronic device and computer-readable storage medium provided by this application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A scene understanding method, characterized in that: include: Acquire multiple frames of depth images of the scene; Constructing a plurality of three-dimensional Gaussian points of the scene according to the plurality of frames of the depth image; The three-dimensional Gaussian points include: high-dimensional language features and geometric features; Clustering the plurality of three-dimensional Gaussian points according to the similarity of the high-dimensional linguistic features and the similarity of the geometric features of the plurality of three-dimensional Gaussian points to obtain a plurality of three-dimensional super-primitives; fusing a plurality of the three-dimensional superprimitives to obtain a three-dimensional instance segmentation representation of the scene; The step of constructing a plurality of three-dimensional Gaussian points of the scene according to the plurality of frames of the depth image comprises: Constructing a plurality of initial three-dimensional Gaussian points of the scene according to the pixel points in the multiple frames of the depth image; the initial three-dimensional Gaussian points include: the geometric features and the semantic coding; the semantic coding of the initial three-dimensional Gaussian points is learned from the three-plane features of the latent space pyramid; Inputting the semantic coding of the initial three-dimensional Gaussian points into a three-dimensional decoder to obtain the high-dimensional language features; Combining the high-dimensional language features with the initial three-dimensional Gaussian points to obtain three-dimensional Gaussian points of the scene; The step of fusing the plurality of three-dimensional superprimitives to obtain a three-dimensional instance segmentation representation of the scene includes: Determine affine coefficients between the plurality of the three-dimensional superprimitives according to projections of the three-dimensional Gaussian points of the plurality of the three-dimensional superprimitives and instance segmentation masks of the plurality of frames of the depth images; wherein the affine coefficients are determined according to the distribution of labels of the instance segmentation masks of the projections of the three-dimensional superprimitives under the plurality of frames of the depth images; According to the affine coefficients between the plurality of the three-dimensional superprimitives, a plurality of the three-dimensional superprimitives whose affine coefficients are less than an affine coefficient threshold are merged to obtain a plurality of preliminary three-dimensional instance segmentation representations; Iteratively fuse the multiple preliminary 3D instance segmentation representations to obtain a 3D instance segmentation representation of the scene.

2. A scene understanding method according to claim 1, characterized in that: The training step of the three-dimensional decoder at least includes: Obtain semantic coding samples of initial three-dimensional Gaussian point samples; Determining, according to the semantic coding samples of the initial three-dimensional Gaussian point samples, high-dimensional language feature samples corresponding to the initial three-dimensional Gaussian point samples and the confidence of the high-dimensional language feature samples; Inputting the initial three-dimensional Gaussian point samples into an initial three-dimensional decoder to obtain predicted high-dimensional language features of the initial three-dimensional Gaussian point samples; Constructing a loss function according to the confidence of the high-dimensional language feature sample and the cosine similarity between the predicted high-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample and the high-dimensional language feature sample; Based on the loss function, the initial 3D decoder is trained to obtain the trained 3D decoder.

3. A scene understanding method according to claim 2, characterized in that: The step of determining the high-dimensional language feature samples corresponding to the initial three-dimensional Gaussian point samples according to the semantically encoded samples of the initial three-dimensional Gaussian point samples comprises: Acquire multiple frames of depth image samples of scene samples corresponding to the initial three-dimensional Gaussian point samples; Extracting language feature maps of multiple frames of depth image samples; Performing multi-angle projection on the initial three-dimensional Gaussian point samples to obtain two-dimensional pixel position information corresponding to the initial three-dimensional Gaussian point samples in multiple frames of the depth image samples; Acquire the visual relationship between the initial three-dimensional Gaussian point samples in multiple frames of depth image samples; Determining, according to the visible relationship, a target depth image sample corresponding to the initial three-dimensional Gaussian point sample; the initial three-dimensional Gaussian point sample is visible in the target depth image sample; Determine the two-dimensional language features corresponding to the initial three-dimensional Gaussian point samples in the target depth image samples according to the language feature graphs of the multiple frames of the depth image samples and the two-dimensional pixel position information corresponding to the initial three-dimensional Gaussian point samples in the target depth image samples; Average pooling is performed on the two-dimensional language features corresponding to the initial three-dimensional Gaussian point samples in the target depth image samples to obtain high-dimensional language feature samples corresponding to the initial three-dimensional Gaussian point samples.

4. A scene understanding method according to claim 3, characterized in that: Determining the confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample according to the semantic coding sample of the initial three-dimensional Gaussian point sample includes: Obtaining the variance of the two-dimensional language features corresponding to the initial three-dimensional Gaussian point samples in the target depth image samples; The confidence of the high-dimensional language feature sample corresponding to the initial three-dimensional Gaussian point sample is determined according to the visual relationship and the variance of the two-dimensional language feature corresponding to the initial three-dimensional Gaussian point sample.

5. A scene understanding method according to claim 1, characterized in that: The step of determining affine coefficients between the plurality of the three-dimensional superprimitives according to projections of the three-dimensional Gaussian points of the plurality of the three-dimensional superprimitives and instance segmentation masks of the plurality of frames of the depth images comprises: Projecting the three-dimensional superprimitive to obtain a plurality of two-dimensional images corresponding to a plurality of frames of the depth image of the three-dimensional superprimitive; Determine, according to the instance segmentation masks of the two-dimensional image and the multiple frames of the depth image, a distribution of the instance segmentation masks in the multiple two-dimensional images; Determine affine coefficients between the plurality of three-dimensional superprimitives in the depth image of the same frame according to distribution of instance segmentation masks of the two-dimensional images of the depth image of the same frame corresponding to the plurality of three-dimensional superprimitives; Obtaining a visible ratio of the two-dimensional image in the corresponding depth image; Affine coefficients between the plurality of three-dimensional superprimitives are determined according to affine coefficients between the plurality of three-dimensional superprimitives in the depth images of the same frames and visible ratios of the two-dimensional images in the depth images of the same frames.

6. A scene understanding device, characterized in that: include: An acquisition module, used to acquire multiple frames of depth images of a scene; A construction module, used to construct a plurality of three-dimensional Gaussian points of the scene according to a plurality of frames of the depth image; The three-dimensional Gaussian points include: high-dimensional language features and geometric features; A clustering module, used for clustering the plurality of three-dimensional Gaussian points according to the similarity of the high-dimensional linguistic features and the similarity of the geometric features of the plurality of three-dimensional Gaussian points to obtain a plurality of three-dimensional super-primitives; A fusion module, used for fusing a plurality of the three-dimensional superprimitives to obtain a three-dimensional instance segmentation representation of the scene; The building blocks include: A construction unit, configured to construct a plurality of initial three-dimensional Gaussian points of the scene according to pixel points in the multiple frames of the depth image; the initial three-dimensional Gaussian points include: the geometric features and semantic coding; the semantic coding of the initial three-dimensional Gaussian points is learned from the three-plane features of the latent space pyramid; An input unit, used for inputting the semantic coding of the initial three-dimensional Gaussian point into a three-dimensional decoder to obtain the high-dimensional language feature; A combining unit, used for combining the high-dimensional language features and the initial three-dimensional Gaussian points to obtain the three-dimensional Gaussian points of the scene; The fusion module includes: A determination unit, configured to determine affine coefficients between the plurality of the three-dimensional superprimitives according to projections of the three-dimensional Gaussian points of the plurality of the three-dimensional superprimitives and instance segmentation masks of the plurality of frames of the depth images; wherein the affine coefficients are determined according to the distribution of labels of the instance segmentation masks of the projections of the three-dimensional superprimitives under the plurality of frames of the depth images; a fusion unit, configured to fuse, according to the affine coefficients between the plurality of the three-dimensional superprimitives, the plurality of three-dimensional superprimitives whose affine coefficients are less than an affine coefficient threshold, to obtain a plurality of preliminary three-dimensional instance segmentation representations; An iterative unit is used to iteratively fuse the multiple preliminary 3D instance segmentation representations to obtain a 3D instance segmentation representation of the scene.

7. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in a scene understanding method as described in any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the scene understanding method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Three-dimensional Gaussian representation scene segmentation method based on group coding

    CN118429363A

  • 3D Object Reconstruction Method, Computer Apparatus and Storage Medium

    US20210327126A1