Semantic three-dimensional scene understanding method and device based on large language model, equipment and storage medium
Through the semantic three-dimensional scene understanding method based on large language model, semantic information of three-dimensional Gaussian ellipsoids is processed and embedded, and the 3D model parameters are optimized, the problem of insufficient semantic understanding and interaction capabilities of the robot system in diverse scenarios is solved, and higher positioning accuracy and scene recognition accuracy are achieved, and natural human-computer interaction is promoted.
Patent Information
- Application Number
- CN202510009870.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-03
AI Technical Summary
In the prior art, the semantic understanding and interaction capabilities of robot systems in diverse scenarios are insufficient, resulting in low positioning accuracy and scene recognition accuracy in different environments.
The semantic three-dimensional scene understanding method based on the large language model is adopted. By collecting multi-angle pictures of indoor scenes, processing them into three-dimensional Gaussian ellipsoids and image semantic text, inputting the large language model for common sense training, predicting the scene type, compressing the semantic information into the three-dimensional Gaussian ellipsoid, and optimizing the 3D model parameters through differentiable rendering end-to-end training to form a 3D scene representation with embedded semantic information.
It effectively improves the scene understanding and interaction capabilities of the robot system in complex environments, enhances the positioning accuracy and scene recognition accuracy in different environments, and achieves more natural human-computer interaction.
Smart Images

Figure CN119941989A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional scene reconstruction, and in particular to a semantic three-dimensional scene understanding method, device, equipment and storage medium based on a large language model. Background Art
[0002] A key challenge in robotics is scene understanding. To achieve widespread deployment of robotic systems, robots must be able to locate themselves in different environments and have a semantic understanding of the environment and the entities in it. For example, when given the instruction “the robot goes to the table in the living room to get the remote control,” the robot needs to: (1) understand the meaning of “living room” and have a basic understanding of typical objects in the living room; (2) use this understanding to identify whether it is in the living room; and (3) identify the target object, the remote control, through object segmentation.
[0003] VSLAM (Visual Simultaneous Localization and Mapping) can build a map of an unknown environment while calculating the position and posture of the visual sensor. Although VSLAM algorithms have made significant progress in recent years, in the field of 3D reconstruction, compared with traditional methods based on voxels, grids, and point clouds, the latest methods based on deep learning and radiation fields have higher reconstruction accuracy, but the demand for computing resources has increased significantly. In addition, these methods focus on visual authenticity and ignore scene interactivity. That is, robots have not yet reached the common sense level of humans in scene, object recognition, and position understanding in terms of environmental perception and scene understanding, and it is still challenging to achieve natural human-computer interaction.
[0004] Therefore, there is an urgent need for a semantic three-dimensional scene understanding method based on a large language model, which can improve the semantic understanding and interaction capabilities of the robot system in diverse scenarios, and further enhance the positioning accuracy and scene recognition accuracy of the robot system in different environments, so as to achieve natural human-computer interaction. Summary of the invention
[0005] The main purpose of the present invention is to provide a semantic three-dimensional scene understanding method, device, equipment and storage medium based on a large language model, aiming to solve the technical problem in the prior art that the robot system has insufficient semantic understanding and interaction capabilities in diversified scenarios, resulting in low positioning accuracy and scene recognition accuracy of the robot system in different environments.
[0006] To achieve the above object, the present invention provides a semantic three-dimensional scene understanding method based on a large language model, the method comprising the following steps:
[0007] Collect multi-angle pictures of indoor scenes, and process the multi-angle pictures to obtain three-dimensional Gaussian ellipsoids and image semantic texts corresponding to the multi-angle pictures;
[0008] Inputting the image semantic text into a preset large language model for common sense training, and predicting the indoor scene type based on the training results to obtain corresponding high-level semantics;
[0009] Compressing the image semantic text and the high-level semantics and embedding them into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid;
[0010] Based on the target three-dimensional Gaussian ellipsoid, the parameters of the 3D model corresponding to the indoor scene are optimized through differentiable rendering end-to-end training to form a 3D scene representation embedded with semantic information, so as to build a deep scene understanding from local objects to global scenes.
[0011] Optionally, the step of processing the multi-angle image to obtain a three-dimensional Gaussian ellipsoid and image semantic text corresponding to the multi-angle image includes:
[0012] Processing the multi-angle images using structure-from-motion technology to obtain a sparse point cloud;
[0013] Generate a corresponding three-dimensional Gaussian ellipsoid based on the sparse point cloud;
[0014] Using an image segmentation model to segment objects in the multi-angle images to obtain multiple masks;
[0015] Each of the masks is input into a contrastive language-image pre-training model for semantic text alignment to obtain image semantic text.
[0016] Optionally, the step of processing the multi-angle images using structure from motion technology to obtain a sparse point cloud includes:
[0017] Using structure-from-motion technology to perform feature extraction and feature matching on the multi-angle images;
[0018] Estimate camera extrinsic parameters according to the feature extraction results and feature matching results, and perform camera calibration by presetting camera intrinsic parameters;
[0019] A sparse point cloud corresponding to the multi-angle image is generated based on the camera extrinsic parameters and the motion recovery structure technology.
[0020] Optionally, the step of collecting multi-angle pictures of indoor scenes includes:
[0021] Collect multi-angle pictures of indoor scenes through visual sensors according to preset picture collection rules;
[0022] The multi-angle pictures are screened according to the structural similarity index between adjacent pictures in the multi-angle pictures to obtain the screened multi-angle pictures.
[0023] Optionally, the step of compressing the image semantic text and the high-level semantics and embedding them into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid comprises:
[0024] Concatenating the image semantic text and the high-level semantics to obtain a nested feature aggregation tensor;
[0025] Using a fully connected layer of a multilayer perceptron to reduce the dimension of the nested feature aggregation tensor to obtain a first code and a second code;
[0026] The first code and the second code are embedded in the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid.
[0027] Optionally, the step of optimizing parameters of the 3D model corresponding to the indoor scene through differentiable rendering end-to-end training based on the target three-dimensional Gaussian ellipsoid to form a 3D scene representation embedded with semantic information includes:
[0028] Performing differentiable rendering on the target three-dimensional Gaussian ellipsoid to obtain a semantic 2D image;
[0029] Decoding the semantic 2D picture to obtain decoded semantic information;
[0030] Comparing the decoded semantic information with the nested feature aggregation tensor to obtain a comparison result;
[0031] Based on the comparison result, the parameters of the 3D model corresponding to the indoor scene are iteratively optimized through a preset loss function to form a 3D scene representation embedded with semantic information.
[0032] Optionally, the preset loss function is:
[0033]
[0034] In the formula, λ is the weight value, L1 is the absolute error loss, E is encoding, D is decoding, S N is the nested feature aggregation tensor, L SSIM represents the structural similarity loss, and N represents the total number of samples.
[0035] In addition, to achieve the above-mentioned purpose, the present invention also proposes a semantic three-dimensional scene understanding device based on a large language model, the device comprising:
[0036] An image processing module is used to collect multi-angle images of indoor scenes and process the multi-angle images to obtain a three-dimensional Gaussian ellipsoid and image semantic text corresponding to the multi-angle images;
[0037] A semantic training module, used for inputting the image semantic text into a preset large language model for common sense training, and predicting the indoor scene type based on the training results to obtain corresponding high-level semantics;
[0038] A semantic embedding module, used for compressing the image semantic text and the high-level semantics and embedding them into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid;
[0039] A differential rendering module is used to optimize the parameters of the 3D model corresponding to the indoor scene based on the target three-dimensional Gaussian ellipsoid through end-to-end differentiable rendering training to form a 3D scene representation embedded with semantic information, so as to build a deep scene understanding from local objects to global scenes.
[0040] In addition, to achieve the above-mentioned purpose, the present invention also proposes a semantic three-dimensional scene understanding device based on a large language model, the device comprising: a memory, a processor, and a semantic three-dimensional scene understanding program based on a large language model stored in the memory and executable on the processor, the semantic three-dimensional scene understanding program based on a large language model being configured to implement the steps of the semantic three-dimensional scene understanding method based on a large language model as described above.
[0041] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, on which a semantic three-dimensional scene understanding program based on a large language model is stored. When the semantic three-dimensional scene understanding program based on a large language model is executed by a processor, the steps of the semantic three-dimensional scene understanding method based on a large language model as described above are implemented.
[0042] The present invention discloses the method of collecting multi-angle pictures of indoor scenes, processing the multi-angle pictures, obtaining a three-dimensional Gaussian ellipsoid and image semantic text corresponding to the multi-angle pictures; inputting the image semantic text into a preset large language model for common sense training, and predicting the type of indoor scenes based on the training results to obtain corresponding high-level semantics; compressing the image semantic text and the high-level semantics and embedding them into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid; based on the target three-dimensional Gaussian ellipsoid, performing parameter optimization on a 3D model corresponding to the indoor scene through end-to-end training of differentiable rendering, forming a 3D scene representation embedded with semantic information, so as to construct a deep scene understanding from local objects to global scenes. Since the present invention inputs the image semantic text into a preset large language model for common sense training, predicts the indoor scene type based on the training results to obtain the corresponding high-level semantics, then compresses the image semantic text and the high-level semantics and embeds them into the three-dimensional Gaussian ellipsoid, and finally trains the 3D model parameters corresponding to the indoor scene end-to-end through differentiable rendering, compared with the prior art, the present invention effectively improves the scene understanding and interaction capabilities of the robot system in complex environments, thereby enhancing the positioning accuracy and scene recognition accuracy of the robot system in different environments, and achieving more natural human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a flowchart of the first embodiment of the semantic three-dimensional scene understanding method based on a large language model of the present invention;
[0044] Figure 2 It is a system flow chart of the semantic three-dimensional scene understanding method based on the large language model of the present invention;
[0045] Figure 3 It is a system block diagram of the semantic three-dimensional scene understanding method based on a large language model of the present invention;
[0046] Figure 4 This is a flowchart of large language model prediction in the semantic three-dimensional scene understanding method based on a large language model of the present invention;
[0047] Figure 5 It is a block diagram of a large language model prediction in a semantic three-dimensional scene understanding method based on a large language model of the present invention;
[0048] Figure 6 It is a flow chart of a second embodiment of a semantic three-dimensional scene understanding method based on a large language model of the present invention;
[0049] Figure 7 It is a structural block diagram of the first embodiment of the semantic three-dimensional scene understanding device based on a large language model of the present invention;
[0050] Figure 8It is a structural diagram of a semantic three-dimensional scene understanding device based on a large language model in a hardware operating environment involved in an embodiment of the present invention.
[0051] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0052] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0053] The embodiment of the present invention provides a semantic three-dimensional scene understanding method based on a large language model, referring to Figure 1 , Figure 1 It is a flowchart of the first embodiment of the semantic three-dimensional scene understanding method based on a large language model of the present invention.
[0054] In this embodiment, the semantic three-dimensional scene understanding method based on a large language model includes steps S10 to S40:
[0055] Step S10: Collect multi-angle pictures of the indoor scene, and process the multi-angle pictures to obtain the three-dimensional Gaussian ellipsoid and image semantic text corresponding to the multi-angle pictures.
[0056] It should be noted that the execution subject of this embodiment can be a computer server device with data processing, network communication and program running functions used in robot application scenarios, such as a server, a tablet computer, a smart phone, a vehicle-mounted mobile device, etc., or an electronic device, a robot, etc. that can realize the above functions. The following takes a robot system (hereinafter referred to as the system) as an example to illustrate this embodiment and the following embodiments.
[0057] For example, refer to Figure 2 and Figure 3 , Figure 2 It is a system flow chart of the semantic three-dimensional scene understanding method based on the large language model of the present invention; Figure 3 This is a system block diagram of the semantic three-dimensional scene understanding method based on a large language model of the present invention.
[0058] In a specific implementation, the system can use the visual sensor to collect indoor scene pictures at a uniform speed for 30 seconds (the specific interval time can be customized, for example, 40 seconds, 50 seconds, etc., and this embodiment and the following embodiments are described with 30 seconds) per frame. The collection of multi-angle pictures can be expressed as {I n ∈R H×W n=1,2,…,N}, where H and W represent the height and width of the pictures in the set respectively, and N represents the number of pictures in the set.
[0059] In order to obtain the three-dimensional Gaussian ellipsoid and image semantic text corresponding to the multi-angle pictures, the motion recovery structure technology (SfM, Structure from Motion) can be used to process the multi-angle pictures to obtain a sparse point cloud; the corresponding three-dimensional Gaussian ellipsoid is generated based on the sparse point cloud; the image segmentation model (SAM, Segment Anything Model) is used to segment the objects in the multi-angle pictures to obtain multiple masks; each of the masks is input into the contrastive language-image pretraining model (CLIP, Contrastive Language-Image Pretraining) for semantic text alignment to obtain the image semantic text.
[0060] In order to obtain a sparse point cloud, the step of processing the multi-angle image using the motion recovery structure technology to obtain a sparse point cloud may include: performing feature extraction and feature matching on the multi-angle image using the motion recovery structure technology; estimating camera extrinsic parameters based on the feature extraction results and feature matching results, and performing camera calibration by presetting camera intrinsic parameters; and generating a sparse point cloud corresponding to the multi-angle image based on the camera extrinsic parameters and the motion recovery structure technology.
[0061] In the specific implementation, SfM performs n Feature extraction and feature matching are performed to estimate the camera external parameters, and the camera calibration is completed through the known internal parameters. Through external parameter estimation, SfM can determine the relative position and orientation of the camera in each image. Subsequently, based on the multi-view geometric relationship, SfM uses a triangulation method to infer the position of the feature points in three-dimensional space, thereby generating a sparse point cloud. 3DGS (three-dimensional Gaussian distribution projection technology) defines a sparse point cloud based on position (mean), covariance matrix ∑ and opacity α′ to initialize a three-dimensional Gaussian ellipsoid. Specifically, the initial covariance matrix is estimated as an isotropic Gaussian distribution, in which the variance of each axis is equal to the mean of the distances of the three nearest points. This estimation method can provide a reasonable initial geometric structure representation and lay the foundation for further optimization. Subsequently, based on Tile rasterization, the process allows α synthesis of various anisotropic three-dimensional Gaussian ellipsoids, and α synthesis is shown in formula (1).
[0062]
[0063] Among them, c i is the color of the i-th sampling point, that is, the color of the i-th three-dimensional Gaussian ellipsoid, α i It is determined by the projection function ∑′ of the i-th three-dimensional Gaussian ellipsoid onto the 2D plane multiplied by the learning transparency rate α′ of the i-th point, as shown in formula (2).
[0064] α i =α′Σ′ (2)
[0065] Anisotropic three-dimensional Gaussian ellipsoids can compactly represent subtle structures, thereby achieving accurate representation of the scene.
[0066] It should be noted that Picture I n Input the SAM model, which can accurately segment objects based on pixels. It automatically generates multiple segmentation masks through its pre-trained segmentation algorithm. These masks correspond to different objects in the image. The objects have clear boundaries, and each mask is a binary image. It is passed as input to the CLIP model for semantic text alignment. The CLIP model finds the corresponding semantic label for each segmented object and aligns it with the mask to generate a semantically consistent representation. Mathematically, pixel-level semantic alignment is expressed as:
[0067] S n =V(I n ⊙M(v)) (3)
[0068] Among them, M(v) is the mask corresponding to each pixel v, ⊙ represents the element-by-element multiplication operation, and V is the mask extracted by the CLIP model to achieve semantic alignment.
[0069] It should be added that the three-dimensional Gaussian ellipsoid refers to a Gaussian distribution centered on a point in the point cloud, where the initial estimation of the covariance matrix is based on the average distance between the point and the three nearest points. Subsequently, by optimizing the 3D model parameters corresponding to the indoor scene, the covariance matrix is gradually adjusted and optimized, which can more accurately represent the local distribution characteristics of the point in the 3D space.
[0070] Step S20: inputting the image semantic text into a preset large language model for common sense training, and predicting the indoor scene type based on the training results to obtain corresponding high-level semantics.
[0071] It should be understood that common sense training refers to the education and training of individuals in common or general knowledge so that they have the ability to act wisely in their daily lives.
[0072] It should be noted that the types of objects contained in indoor scenes are complex, and some individual objects are highly repetitive. In the pre-training stage, in order to meet universality, all object semantic labels (i.e., the basic semantics in the image semantic text) are input into the large language model for common sense training. In the prediction stage, it is unreasonable to input all object semantic labels into the large language model: first, because long sentences are redundant, it is difficult to extract important information from them. Second, objects that do not have "significant" information may mask unique information related to specific indoor scenes, making the query inaccurate and ultimately predicting incorrectly. Therefore, in the prediction stage, it is necessary to preferentially select K keywords.
[0073] For example, refer to Figure 4 and Figure 5 , Figure 4 This is a flowchart of large language model prediction in the semantic three-dimensional scene understanding method based on a large language model of the present invention; Figure 5 This is a block diagram of a large language model prediction in a semantic three-dimensional scene understanding method based on a large language model of the present invention. In a specific implementation, step S20 may include: steps S201 to 206:
[0074] Step S201: In the pre-training stage, the probability of objects appearing in each indoor scene (room) is calculated. First, the true co-occurrence frequency is used, that is, the number of times each object label appears in different types of rooms C ( k ,r m ). The room type is determined by {r m ∈L R |m=1,2,…R} means, where r m represents the mth room, L R Represents a set of rooms. Item tags are represented by {o k ∈L O |k=1,2,…O} means, where o k represents the kth item, L O Represents a set of objects. In order to ensure that the frequencies of occurrence of different items in room types are comparable, the true co-occurrence frequencies need to be normalized to calculate the conditional probability. The conditional probability of an item in each room type satisfies the following formula:
[0075]
[0076] Step S202: Calculate the probability distribution of room types under the condition that a given item appears. Derived from Bayesian theorem (5):
[0077]
[0078] Among them, room type r m The prior probability a is r m Total number of rooms by type. Indicates that in all room types, item o k The weighted average probability of occurrence, R represents the total number of rooms, and r represents the room.
[0079] Step S203: Known conditional probability p(r m |o k ), the metric value is calculated by entropy. Entropy measures the distribution uncertainty of items appearing in different room types. When the entropy value is the largest, it means that item o kThe distribution in all room types is very uniform, with no particular preference for any one type of room. When the entropy value is minimum, it means that the item o k Mainly concentrated in a specific room type, indicating that the item is more "prominent" because its location is more clear. The measurement formula is as follows:
[0080]
[0081] Step S204: Select K keywords to participate in the large language model prediction. In order to improve the prediction accuracy, select K items with the smallest entropy. The formula is as follows:
[0082]
[0083] Step S205: In the prediction stage, the preset large language model selected in this embodiment can be Next-token Prediction Models, and the room type is predicted by Next-token Prediction Models. The latest Next-token Prediction Models are based on the self-attention mechanism of Transformer for reasoning and generation, and meet parallel processing. In short, the model calculates the corresponding Q (query) for each token (word): representing the query vector of the current token, K (key): representing the key vector of all previous tokens, V (value): representing the value vector of all tokens. Then calculate its probability through formula (8), and select the one with the highest probability as the next token.
[0084] P(W)=softmax(Q i ·K 1:i-1 )V (8)
[0085] Where W is the sentence, Q i ·K 1:i-1 The query vector Q representing the i-th token and the key vector K representing tokens 1 to i-1 are set together. The softmax function converts the scores of all tokens into a probability distribution, indicating the likelihood of each token being the next word.
[0086] A priori construct a room query list L R ,When the robot is in a room, it is given a “choose word to fill in the blank” sentence, the content is shown in formula (9), which is the high-level semantics.
[0087]
[0088] The "selection word" word library is L o and L R, the “fill in the blanks” are underlined.
[0089] Step S206: Score the predicted room type using formula (10), and the room type with the highest score is the predicted room type, i.e., formula (11).
[0090] S(W m )=P(W m ) (10)
[0091]
[0092] Step S30: compressing the image semantic text and the high-level semantics and embedding them into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid.
[0093] It should be noted that in order to compress the image semantic text and high-level semantics and embed them into the three-dimensional Gaussian ellipsoid, the image semantic text and the high-level semantics can be spliced to obtain a nested feature aggregation tensor; the nested feature aggregation tensor can be reduced in dimension using the fully connected layer of the multilayer perceptron to obtain a first code and a second code; the first code and the second code are embedded in the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid.
[0094] Step S40: Based on the target three-dimensional Gaussian ellipsoid, the parameters of the 3D model corresponding to the indoor scene are optimized through differentiable rendering end-to-end training to form a 3D scene representation embedded with semantic information, so as to construct a deep scene understanding from local objects to global scenes.
[0095] In order to ensure the consistency of the semantic scene, the target three-dimensional Gaussian ellipsoid can be differentiably rendered to obtain a semantic 2D picture; the semantic 2D picture can be decoded to obtain decoded semantic information; the decoded semantic information can be compared with the nested feature aggregation tensor to obtain a comparison result; based on the comparison result, the parameters of the 3D model corresponding to the indoor scene can be iteratively optimized through a preset loss function to form a 3D scene representation with embedded semantic information.
[0096] In the specific implementation, after obtaining the high-level semantics corresponding to the room type, the high-level semantics is combined with S through the CONCAT function. n (i.e. image semantic text) is concatenated into S N(i.e., nested feature aggregation tensor), forming a nested feature aggregation tensor of segmented objects (2D images) - basic semantics - advanced semantics. Taking into account the memory consumption, the need for semantic text embedding in the three-dimensional Gaussian ellipsoid to meet the requirements of simplicity, encoding format consistency, and real-time optimization of 3D model parameters, MLP (Multilayer Perceptron) is used to process the basic semantics-advanced semantic tensor. Specifically, it is reduced in dimension through a fully connected layer and encoded into the first code f1 and the second code f2, and then f1 and f2 are embedded in the three-dimensional Gaussian ellipsoid. The semantic 2D picture generated by the three-dimensional Gaussian ellipsoid can be rendered differently. After decoding, the generated semantic information is compared with the input semantic information (i.e., the nested feature aggregation tensor), and the preset loss function L is used. c The 3D model corresponding to the indoor scene is optimized. This process is continuously iterated to ensure the semantic scene consistency. It should be noted that the number of iterations in this example can be 3000 times (it should be noted that the number of iterations can be custom set, for example, 1000 times, 5000 times, etc., and this embodiment and the following embodiments are described with 3000 times).
[0097]
[0098] In the formula, λ is the weight value, L1 is the absolute error loss, E is encoding, D is decoding, S N is the nested feature aggregation tensor, L SSIM represents the structural similarity loss, and N represents the total number of samples.
[0099] The present embodiment discloses collecting multi-angle pictures of indoor scenes, processing the multi-angle pictures, obtaining a three-dimensional Gaussian ellipsoid and image semantic text corresponding to the multi-angle pictures; inputting the image semantic text into a preset large language model for common sense training, and predicting the type of indoor scenes based on the training results to obtain corresponding high-level semantics; compressing the image semantic text and the high-level semantics and embedding them into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid; based on the target three-dimensional Gaussian ellipsoid, optimizing the parameters of the 3D model corresponding to the indoor scene through end-to-end training of differentiable rendering to form a 3D scene representation embedded with semantic information, so as to construct a deep scene understanding from local objects to global scenes. Since this embodiment inputs the image semantic text into the preset large language model for common sense training, predicts the indoor scene type based on the training results to obtain the corresponding high-level semantics, then compresses the image semantic text and the high-level semantics and embeds them into the three-dimensional Gaussian ellipsoid, and finally trains the 3D model parameters corresponding to the indoor scene end-to-end through differentiable rendering, compared with the existing technology, this embodiment effectively improves the scene understanding and interaction capabilities of the robot system in complex environments, thereby enhancing the positioning accuracy and scene recognition accuracy of the robot system in different environments, and achieving more natural human-computer interaction.
[0100] refer to Figure 6 , Figure 6 It is a flowchart diagram of the second embodiment of the semantic three-dimensional scene understanding method based on a large language model of the present invention.
[0101] Based on the above first embodiment, in this embodiment, step S10 includes steps S101 to S102:
[0102] Step S101: Collect multi-angle pictures of indoor scenes using a visual sensor according to preset picture collection rules.
[0103] Step S102: screening the multi-angle pictures according to the structural similarity index between adjacent pictures in the multi-angle pictures to obtain screened multi-angle pictures.
[0104] It should be noted that in order to ensure the validity of the multi-angle image data volume and the efficiency of subsequent calculation and processing of data, the multi-angle images can be screened, which can not only improve the efficiency of data processing, but also ensure the validity of data processing.
[0105] In a specific implementation, the multi-angle pictures can be screened by SSIM (Structural Similarity Index), and the pictures with similarity between adjacent frames greater than a preset threshold will be discarded to obtain the screened multi-angle pictures.
[0106] It should be understood that the above-mentioned preset threshold can be a user-defined setting, such as 70%, 80%, etc. This embodiment takes 80% as an example for explanation. The multi-angle images with a similarity of more than 80% for each adjacent frame will be discarded to obtain the filtered multi-angle images.
[0107] This embodiment discloses collecting multi-angle pictures of indoor scenes by using a visual sensor according to a preset picture collection rule; filtering the multi-angle pictures according to the structural similarity index between adjacent pictures in the multi-angle pictures to obtain the filtered multi-angle pictures. Compared with the prior art, since the present invention filters the multi-angle pictures by the structural similarity index, it can not only improve the efficiency of subsequent data processing, but also ensure the effectiveness of subsequent data processing.
[0108] In addition, an embodiment of the present invention also proposes a storage medium, on which a semantic three-dimensional scene understanding program based on a large language model is stored. When the semantic three-dimensional scene understanding program based on a large language model is executed by a processor, the steps of the semantic three-dimensional scene understanding method based on a large language model as described above are implemented.
[0109] Reference Figure 7 , Figure 7 This is a structural block diagram of the first embodiment of the semantic three-dimensional scene understanding device based on a large language model of the present invention.
[0110] like Figure 7 As shown, the semantic three-dimensional scene understanding device based on the large language model proposed in the embodiment of the present invention includes: an image processing module 701, a semantic training module 702, a semantic embedding module 703 and a differential rendering module 704.
[0111] The image processing module 701 is used to collect multi-angle images of indoor scenes and process the multi-angle images to obtain the three-dimensional Gaussian ellipsoid and image semantic text corresponding to the multi-angle images.
[0112] The semantic training module 702 is used to input the image semantic text into a preset large language model for common sense training, and predict the indoor scene type based on the training results to obtain corresponding high-level semantics.
[0113] The semantic embedding module 703 is used to compress the image semantic text and the high-level semantics and embed them into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid.
[0114] The differential rendering module 704 is used to optimize the parameters of the 3D model corresponding to the indoor scene based on the target three-dimensional Gaussian ellipsoid through differentiable rendering end-to-end training to form a 3D scene representation embedded with semantic information, so as to build a deep scene understanding from local objects to global scenes.
[0115] The image processing module 701 is also used to process the multi-angle image using motion recovery structure technology to obtain a sparse point cloud; generate a corresponding three-dimensional Gaussian ellipsoid based on the sparse point cloud; use an image segmentation model to segment objects in the multi-angle image to obtain multiple masks; input each of the masks into a contrast language-image pre-training model to perform semantic text alignment to obtain image semantic text.
[0116] The image processing module 701 is also used to perform feature extraction and feature matching on the multi-angle image using the motion recovery structure technology; estimate the camera extrinsic parameters based on the feature extraction results and feature matching results, and perform camera calibration by presetting the camera intrinsic parameters; and generate a sparse point cloud corresponding to the multi-angle image based on the camera extrinsic parameters and the motion recovery structure technology.
[0117] The semantic embedding module 703 is also used to concatenate the image semantic text and the high-level semantics to obtain a nested feature aggregation tensor; use the fully connected layer of the multi-layer perceptron to reduce the dimension of the nested feature aggregation tensor to obtain a first code and a second code; embed the first code and the second code into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid.
[0118] The differential rendering module 704 is also used to perform differentiable rendering on the target three-dimensional Gaussian ellipsoid to obtain a semantic 2D image; decode the semantic 2D image to obtain decoded semantic information; compare the decoded semantic information with the nested feature aggregation tensor to obtain a comparison result; based on the comparison result, iteratively optimize the parameters of the 3D model corresponding to the indoor scene through a preset loss function to form a 3D scene representation embedded with semantic information.
[0119] The embodiment of the device discloses collecting multi-angle pictures of indoor scenes, processing the multi-angle pictures, obtaining a three-dimensional Gaussian ellipsoid and image semantic text corresponding to the multi-angle pictures; inputting the image semantic text into a preset large language model for common sense training, and predicting the type of indoor scenes based on the training results to obtain corresponding high-level semantics; compressing the image semantic text and the high-level semantics and embedding them into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid; based on the target three-dimensional Gaussian ellipsoid, optimizing the parameters of the 3D model corresponding to the indoor scene through end-to-end training of differentiable rendering to form a 3D scene representation embedded with semantic information, so as to construct a deep scene understanding from local objects to global scenes. Since the embodiment of the device inputs the image semantic text into the preset large language model for common sense training, predicts the indoor scene type based on the training results to obtain the corresponding high-level semantics, and then compresses the image semantic text and the high-level semantics and embeds them into the three-dimensional Gaussian ellipsoid, and finally trains the 3D model parameters corresponding to the indoor scene through end-to-end differentiable rendering, compared with the prior art, the embodiment of the device effectively improves the scene understanding and interaction capabilities of the robot system in complex environments, thereby enhancing the positioning accuracy and scene recognition accuracy of the robot system in different environments, and achieving more natural human-computer interaction.
[0120] Based on the first embodiment of the semantic three-dimensional scene understanding device based on a large language model of the present invention, a second embodiment of the semantic three-dimensional scene understanding device based on a large language model of the present invention is proposed.
[0121] In this embodiment, the image processing module 701 is also used to collect multi-angle images of indoor scenes through visual sensors according to preset image collection rules; the multi-angle images are filtered according to the structural similarity index between adjacent images in the multi-angle images to obtain filtered multi-angle images.
[0122] Other embodiments or specific implementations of the semantic three-dimensional scene understanding device based on the large language model of the present invention can refer to the above-mentioned method embodiments and will not be repeated here.
[0123] The present application provides a semantic three-dimensional scene understanding device based on a large language model, and the semantic three-dimensional scene understanding device based on a large language model includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the semantic three-dimensional scene understanding method based on the large language model in the above-mentioned embodiment one.
[0124] Reference below Figure 8, which shows a schematic diagram of the structure of a semantic three-dimensional scene understanding device based on a large language model suitable for implementing the embodiment of the present application. The semantic three-dimensional scene understanding device based on a large language model in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The semantic three-dimensional scene understanding device based on a large language model shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0125] like Figure 8 As shown, the semantic three-dimensional scene understanding device based on the large language model may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 to the random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the semantic three-dimensional scene understanding device based on the large language model are also stored. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the semantic three-dimensional scene understanding device based on a large language model to communicate wirelessly or wired with other devices to exchange data. Although the semantic three-dimensional scene understanding device based on a large language model with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have alternatively.
[0126] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0127] The semantic three-dimensional scene understanding device based on a large language model provided by the present application adopts the semantic three-dimensional scene understanding method based on a large language model in the above-mentioned embodiment, which can solve the technical problem that the robot system in the prior art has insufficient semantic understanding and interaction capabilities in diversified scenes, resulting in low positioning accuracy and scene recognition accuracy of the robot system in different environments. Compared with the prior art, the beneficial effects of the semantic three-dimensional scene understanding device based on a large language model provided by the present application are the same as the beneficial effects of the semantic three-dimensional scene understanding method based on a large language model provided by the above-mentioned embodiment, and the other technical features in the semantic three-dimensional scene understanding device based on a large language model are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0128] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0129] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0130] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0131] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0132] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory / random access memory, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0133] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A semantic three-dimensional scene understanding method based on a large language model, characterized in that: The method comprises: Collect multi-angle pictures of indoor scenes, and process the multi-angle pictures to obtain three-dimensional Gaussian ellipsoids and image semantic texts corresponding to the multi-angle pictures; Inputting the image semantic text into a preset large language model for common sense training, and predicting the indoor scene type based on the training results to obtain corresponding high-level semantics; Compressing the image semantic text and the high-level semantics and embedding them into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid; Based on the target three-dimensional Gaussian ellipsoid, the parameters of the 3D model corresponding to the indoor scene are optimized through differentiable rendering end-to-end training to form a 3D scene representation embedded with semantic information, so as to build a deep scene understanding from local objects to global scenes.
2. The semantic three-dimensional scene understanding method based on a large language model as claimed in claim 1, characterized in that: The step of processing the multi-angle image to obtain a three-dimensional Gaussian ellipsoid and image semantic text corresponding to the multi-angle image includes: Processing the multi-angle images using structure-from-motion technology to obtain a sparse point cloud; Generate a corresponding three-dimensional Gaussian ellipsoid based on the sparse point cloud; Using an image segmentation model to segment objects in the multi-angle images to obtain multiple masks; Each of the masks is input into a contrastive language-image pre-training model for semantic text alignment to obtain image semantic text.
3. The semantic three-dimensional scene understanding method based on a large language model as claimed in claim 2, characterized in that: The step of processing the multi-angle images using the motion recovery structure technology to obtain a sparse point cloud includes: Using structure-from-motion technology to perform feature extraction and feature matching on the multi-angle images; Estimate camera extrinsic parameters according to the feature extraction results and feature matching results, and perform camera calibration by presetting camera intrinsic parameters; A sparse point cloud corresponding to the multi-angle image is generated based on the camera extrinsic parameters and the motion recovery structure technology.
4. The semantic three-dimensional scene understanding method based on a large language model as claimed in claim 1, characterized in that: The step of collecting multi-angle pictures of the indoor scene includes: Collect multi-angle pictures of indoor scenes through visual sensors according to preset picture collection rules; The multi-angle pictures are screened according to the structural similarity index between adjacent pictures in the multi-angle pictures to obtain the screened multi-angle pictures.
5. The semantic three-dimensional scene understanding method based on a large language model as claimed in claim 1, characterized in that: The step of compressing the image semantic text and the high-level semantics and embedding them into the three-dimensional Gaussian ellipsoid to obtain the target three-dimensional Gaussian ellipsoid comprises: Concatenating the image semantic text and the high-level semantics to obtain a nested feature aggregation tensor; Using a fully connected layer of a multilayer perceptron to reduce the dimension of the nested feature aggregation tensor to obtain a first code and a second code; The first code and the second code are embedded in the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid.
6. The semantic three-dimensional scene understanding method based on a large language model as claimed in claim 5, characterized in that: The step of optimizing parameters of the 3D model corresponding to the indoor scene through end-to-end training of differentiable rendering based on the target three-dimensional Gaussian ellipsoid to form a 3D scene representation embedded with semantic information includes: Performing differentiable rendering on the target three-dimensional Gaussian ellipsoid to obtain a semantic 2D image; Decoding the semantic 2D picture to obtain decoded semantic information; Comparing the decoded semantic information with the nested feature aggregation tensor to obtain a comparison result; Based on the comparison result, the parameters of the 3D model corresponding to the indoor scene are iteratively optimized through a preset loss function to form a 3D scene representation embedded with semantic information.
7. The semantic three-dimensional scene understanding method based on a large language model as claimed in claim 6, characterized in that: The preset loss function is: In the formula, λ is the weight value, L1 is the absolute error loss, E is encoding, D is decoding, S N is the nested feature aggregation tensor, L SSIM represents the structural similarity loss, and N represents the total number of samples.
8. A semantic three-dimensional scene understanding device based on a large language model, characterized in that: The device comprises: An image processing module is used to collect multi-angle images of indoor scenes and process the multi-angle images to obtain a three-dimensional Gaussian ellipsoid and image semantic text corresponding to the multi-angle images; A semantic training module, used for inputting the image semantic text into a preset large language model for common sense training, and predicting the indoor scene type based on the training results to obtain corresponding high-level semantics; A semantic embedding module, used for compressing the image semantic text and the high-level semantics and embedding them into the three-dimensional Gaussian ellipsoid to obtain a target three-dimensional Gaussian ellipsoid; The differential rendering module is used to optimize the parameters of the 3D model corresponding to the indoor scene through end-to-end training of differentiable rendering based on the target three-dimensional Gaussian ellipsoid, so as to form a 3D scene representation embedded with semantic information, so as to build a deep scene understanding from local objects to global scenes.
9. A semantic three-dimensional scene understanding device based on a large language model, characterized in that: The device includes: a memory, a processor, and a semantic three-dimensional scene understanding program based on a large language model stored in the memory and executable on the processor, wherein the semantic three-dimensional scene understanding program based on a large language model is configured to implement the steps of the semantic three-dimensional scene understanding method based on a large language model as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores a semantic three-dimensional scene understanding program based on a large language model. When the semantic three-dimensional scene understanding program based on a large language model is executed by a processor, the steps of the semantic three-dimensional scene understanding method based on a large language model as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Three-dimensional Gaussian scene stylization method based on text driving
CN119006760A
Real-scene three-dimensional scene reconstruction method, system and device and storage medium
CN119068123A
Probabilistic approach to unifying representations for robotic mapping
US20240091950A1
Cited By
Nerve radiation field plant rendering method and device fused with large language model
CN120411325A
Video generation method and device, electronic equipment, storage medium and program product
CN120434373A
Three-dimensional digital human flaw repairing method, related device and storage medium
CN120726242A
A three-dimensional digital human flaw repairing method and related device and storage medium
CN120726242B
Three-dimensional scene construction method and device, electronic equipment, storage medium and product
CN121236316A