A three-dimensional scene perception interaction method and system based on multimodal collaborative representation

By using the combination of multimodal collaborative representation and large language model in indoor three-dimensional scene understanding, the problem of restricted expert models and lack of object positioning methods in the prior art is solved, and efficient understanding and multi-task processing of three-dimensional scenes are achieved.

CN118658154BActive Publication Date: 2025-05-13ZHEJIANG UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410818080.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-05-13
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

In the understanding of indoor three-dimensional scenarios, the problem that expert models are limited by fixed input and output forms and the three-dimensional multimodal large language model lacks effective object positioning methods.

Method used

A three-dimensional scene-aware interaction method based on multimodal collaborative representation is adopted. By obtaining three-dimensional scene point cloud data, multi-view depth image data and text query annotations, object features are extracted using object detectors and pre-trained encoders to form multimodal collaborative representations of objects, and single-stage joint training is performed using large language models.

Benefits of technology

It realizes a general understanding of three-dimensional scenes, improves the performance of three-dimensional scene positioning, description and question-answer, can handle complex scenes and tasks, and has flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118658154B_ABST
    Figure CN118658154B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional scene perception and interaction method and system based on multimodal collaborative representation, which belongs to the field of indoor three-dimensional scene understanding. Three-dimensional scene point cloud data, multi-view depth image data and text query annotations are obtained, object instances are detected from the point cloud, and multi-view projections of the object instances are obtained; object point cloud features and image features are extracted using a pre-trained encoder, and then projected to the embedded feature space of the language model respectively; object point cloud embedding features and image embedding features are connected using object identifiers to form a multimodal collaborative representation of the object, thereby expressing three-dimensional scene information in the language model input, and finally using the reasoning and general dialogue capabilities of the large language model to achieve general three-dimensional scene perception and interaction. The present invention realizes general three-dimensional scene perception and interaction by introducing object-level multimodal collaborative representation into the large language model, and simultaneously improves performance in multiple indoor three-dimensional scene downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of indoor three-dimensional scene understanding, and in particular to a three-dimensional scene perception interaction method and system based on multimodal collaborative representation. Background Art

[0002] At present, the field of indoor three-dimensional scene understanding has attracted widespread attention and has become an important field of artificial intelligence. This field aims to combine indoor three-dimensional scene point clouds and human text instructions to achieve various understanding tasks in three-dimensional scenes, including visual positioning, dense description generation, and scene question and answering, which plays an important role in the development of embodied intelligence and indoor robots.

[0003] Existing technologies can be divided into three categories: 1) traditional expert models, 2) multi-task expert models, and 3) three-dimensional multimodal large language models.

[0004] The traditional expert model designs a special structure for a specific task and achieves good performance, but is limited by its model structure and is difficult to expand to other 3D scene tasks; the multi-task expert model adds different task heads on the basis of the backbone network to adapt to the output required by different tasks, so that one model can solve multiple tasks. However, these task heads usually require special design and have a fixed output form. They lack flexibility in practical applications and cannot handle complex scenes and tasks in practical applications. The 3D multimodal large language model integrates 3D point cloud modal information into the large language model and uses the large language model to achieve general reasoning and dialogue. However, the current 3D multimodal large language model still has a large gap with the expert model in tasks such as 3D visual positioning due to the lack of effective object positioning methods. Summary of the invention

[0005] In order to overcome the problem that the existing expert model is limited to fixed input and output forms and to improve the problem that the existing three-dimensional multimodal large language model lacks an effective object positioning method, the present invention provides a three-dimensional scene perception interaction method and system based on multimodal collaborative representation to achieve general indoor three-dimensional scene understanding.

[0006] The specific technical solution adopted by the present invention is:

[0007] In a first aspect, the present invention proposes a three-dimensional scene perception interaction method based on multimodal collaborative representation, comprising:

[0008] Obtain 3D scene point cloud data, multi-view depth image data and text query annotations as training sets, use object detectors to detect object instances from the original 3D scene, extract object point clouds, and combine multi-view depth images to obtain multi-view projections of object instances;

[0009] The object point cloud features are extracted using a pre-trained point cloud encoder, and the object image features are extracted from the multi-view depth image and multi-view projection using a pre-trained image encoder. Then, two learnable linear projection layers are used to project the object point cloud features and the object image features into the embedding feature space of the language model to obtain the object point cloud embedding features and the object image embedding features.

[0010] The object identifier and text query are segmented and embedded features are obtained using a learnable text embedding layer. The embedded features of the object identifier are connected to the object point cloud embedding features and the object image embedding features to form a multimodal collaborative representation of the object. The embedded features of the text query are concatenated with the multimodal collaborative representation as the input of the language model, and the language model is used to generate the answer to the text query.

[0011] Using the training set, the learnable linear projection layer, text embedding layer and language model are jointly trained in a single stage, and the trained model is used to complete perceptual interaction of various three-dimensional scene tasks.

[0012] Furthermore, the detecting of object instances from the original three-dimensional scene and extracting object point clouds includes:

[0013] Use the pre-trained 3D object instance segmentation model as the object detector, input the original 3D scene point cloud, extract the n object instances with the highest confidence from the scene, obtain the segmentation mask of each object instance, and use the segmentation mask to extract the point cloud corresponding to each object from the original 3D scene point cloud. Each point in the point cloud contains the 3D coordinates and RGB color of the point.

[0014] Furthermore, the method of acquiring multi-view projections of object instances by combining multi-view depth images includes:

[0015] The multi-view depth image and the camera's internal and external parameter information are used to obtain the mapping relationship between the pixels in each depth image and the three-dimensional point cloud. For each point cloud extracted from the object, the pixel position in the corresponding multi-view depth image is obtained to obtain the multi-view projection.

[0016] Furthermore, the method for extracting object point cloud features and object image features includes:

[0017] Input the point cloud corresponding to the object instance into the pre-trained point cloud feature encoder to extract the object point cloud features;

[0018] Input the multi-view depth image into the pre-trained image encoder, extract the feature map of each depth image, extract the pixel feature vector of the object instance from the feature map of each depth image using multi-view projection, take the mean as the image feature of the object instance corresponding to the depth image, and make a weighted average according to the projection area of ​​the object instance in each depth image to obtain the final object image feature;

[0019] Traverse all object instances to obtain all object point cloud features and object image features.

[0020] Furthermore, the object identifier and the text query are segmented and an embedding feature is obtained using a learnable text embedding layer, and the embedding feature of the object identifier is connected with the object point cloud embedding feature and the object image embedding feature to form a multimodal collaborative representation of the object, including:

[0021] Introducing object identifiers with the same number as object instances, which are used to refer to objects in text interaction, and transforming different 3D scene tasks into a unified form;

[0022] Tokenize object identifiers and text queries and obtain embedded features using a learnable text embedding layer;

[0023] Arrange the object identifier embedding features, point cloud embedding features and image embedding features in order to obtain the multimodal collaborative representation of the object Among them, i represents the embedding feature of the i-th object identifier, are the point cloud embedding features and image embedding features of the i-th object instance respectively.

[0024] Furthermore, object identifiers are used to transform different three-dimensional scene tasks into a unified question-answering format, wherein the three-dimensional scene tasks include a visual positioning task, a dense description generation task, and a scene question-answering task.

[0025] Furthermore, in the single-stage joint training, the loss function is calculated as follows:

[0026]

[0027] in, is the loss function, s prefix is a preamble sequence consisting of the multimodal co-representation of the object and the embedded features of the text query, k is the length of the target answer sequence of the text query, is the sequence of the first i-1 words in the answer sequence, θ is a trainable parameter, is the probability of generating the i-th word in the answer sequence given the first i-1 words in the preamble sequence and the answer sequence.

[0028] In a second aspect, the present invention proposes a three-dimensional scene perception interaction system based on multimodal collaborative representation, which is used to implement the above-mentioned three-dimensional scene perception interaction method based on multimodal collaborative representation.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] The present invention is a three-dimensional scene perception interaction method and system based on multimodal collaborative representation. During implementation, the present invention uses multimodal collaborative representation and a large language model.

[0031] (1) By using multimodal collaborative representation, the model of the present invention not only obtains rich shape and spatial information from the 3D point cloud modality, but also obtains rich semantic information from the 2D image modality, which enhances the overall understanding of the scene and thus realizes advanced 3D scene positioning, description, and question-answering.

[0032] (2) By using a large language model, the present invention leverages the powerful understanding and conversation capabilities of the large language model to extract rich scene and object instance information from multimodal collaborative representations, unify the question-answering format of multiple downstream tasks, and thus realize a method for jointly training data from multiple different tasks, and effectively improve the performance of the model in various data sets and tasks.

[0033] In summary, by jointly using multimodal collaborative representation and a large language model, the present invention can fully learn the shape, space, and semantic information of three-dimensional scenes, and perform joint training on multiple different three-dimensional scene datasets, thereby achieving advanced three-dimensional scene localization, description, and question-answering performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0035] The present invention will be further described and illustrated below in conjunction with the accompanying drawings and specific implementation methods.

[0036] The present invention proposes a three-dimensional scene perception interaction method based on multimodal collaborative representation, which mainly includes the following steps:

[0037] Step 1, obtaining three-dimensional scene point cloud data, multi-view depth image data and text query annotations as training sets, wherein the text query annotations consist of questions and answers; using an object detector to detect object instances from the original three-dimensional scene and extract object point clouds;

[0038] Step 2, combining the correspondence between the pixel points in the multi-view depth image and the object point cloud to obtain the multi-view projection of the object instance;

[0039] Step 3, using the pre-trained point cloud encoder to extract object point cloud features, and using the pre-trained image encoder to extract object image features from multi-view depth images and multi-view projections;

[0040] Step 4: Use two learnable linear projection layers to project the object point cloud features and the object image features into the embedding feature space of the language model to obtain the object point cloud embedding features and the object image embedding features;

[0041] Step 5: Use the object identifier to connect the object point cloud embedding features and the object image embedding features to form a multimodal collaborative representation of the object, so as to express the three-dimensional scene information in the language model input; splice the embedded features of the text query after the multimodal collaborative representation as the input of the language model;

[0042] Step 6: Use object identifiers to convert different 3D scene tasks into a unified form, use the reasoning and dialogue capabilities of the pre-trained large language model to complete different tasks, and generate answers to text queries;

[0043] Step 7: Acquire and integrate multiple 3D scene downstream task data sets to form a unified single-round dialogue data format, and use the cross-entropy loss of the language model prediction results to perform single-stage joint training on the multimodal linear projection layer and language model, thereby realizing a general 3D scene perception interaction method and system that can simultaneously complete various downstream tasks.

[0044] In the above step 1, an optional method for obtaining an object instance in a three-dimensional scene is as follows: and a method for obtaining a multi-view projection thereof is as follows:

[0045] Object instance: Use the pre-trained 3D object instance segmentation model as an object detector, input the original 3D scene point cloud, and extract n object instances from the scene; object detectors are well-known in the art and will not be introduced here. The object detector will generate a detection box with confidence for the input point cloud or image, and only retain the n object instances corresponding to the top n detection boxes with confidence according to a preset number threshold. Based on the n object instances detected, obtain the segmentation mask of each object instance, use the segmentation mask to extract the point cloud corresponding to each object from the original 3D scene point cloud, and record the point cloud of the i-th object as Where m i Represents the point cloud size of the i-th object. Each point contains 6-dimensional information, namely 3D coordinates and RGB colors.

[0046] In step 2 above, an optional method for obtaining multi-view projections of an object instance is as follows:

[0047] Multi-view projection of objects: Utilize the multi-view depth image and the camera’s internal and external parameter information to calculate the mapping relationship between the pixels in each depth image and the three-dimensional point cloud. For each extracted point cloud of an object, obtain the corresponding pixel position in the multi-view depth image, i.e., multi-view projection.

[0048] In the above step 3, an optional way to obtain the object point cloud features and object image features is as follows:

[0049] 3.1) The point cloud P corresponding to the i-th object instance i Input the pre-trained point cloud encoder Uni3D to obtain its point cloud features Uni3D can also be replaced by other encoders available in the art.

[0050] 3.2) For multi-view depth images, the pre-trained image encoder DINOv2 is first used to extract a feature map of each depth image with a size of 16×16×1024, where 16×16 is the feature map size; DINOv2 can also be replaced by other existing encoders in the field.

[0051] 3.3) For the i-th object instance, use its multi-view projection to extract its corresponding pixel position in each depth image, and extract the feature vector corresponding to the area where the object exists from the feature map of each depth image according to the corresponding pixel position. The feature vectors extracted from each depth image are averaged as the image feature of the object instance corresponding to the depth image. Finally, the final multi-view image feature of the object is obtained by weighted average according to the projection area of ​​the object instance in each depth image.

[0052] 3.4) Traverse all n object instances and obtain the point cloud features of all objects and image features

[0053] In the above step 4, an optional way to obtain the object point cloud embedding features and the object image embedding features is as follows:

[0054] 4.1) For the i-th object instance, use the Fourier position code pos(· to obtain the position code vector corresponding to its point cloud

[0055] 4.2) Point cloud features for the i-th object instance Use a point cloud-text projection layer f p (·, project it into the embedding feature space of the language model, and add the corresponding position encoding vector to obtain the point cloud embedding feature pos(P i); the point cloud-text projection layer is a linear layer structure.

[0056] 4.3) Image features for the i-th object instance Use an image-text projection layer f v (·), project it into the embedding feature space of the language model, and add the corresponding position encoding vector to obtain the image embedding feature pos(P i ); the image-text projection layer is a linear layer structure.

[0057] 4.4) Traverse all n object instances and obtain the point cloud embedding feature sequence of all objects and image embedding feature sequence

[0058] In the above step 5, an optional way of obtaining the multimodal collaborative representation of the object is as follows:

[0059] 5.1) For n object instances in a three-dimensional scene, n object identifiers <OBJ001>, <OBJ002>, …, <OOJn> are introduced to refer to specific objects in text dialogues. In step 1, the object instances are detected from the original three-dimensional scene using an object detector, and then the detected n object instances are displayed in a visual manner. It is preferred that the n learnable object identifiers correspond to the n object instances one by one. When performing a three-dimensional scene task, the object identifiers can be input or output instead of the object instances; the object identifiers are segmented using a text segmenter, and then the corresponding embedding features are obtained through a text embedding layer, which are recorded as {O i} i=1…n ; Text queries also need to use the text segmenter and text embedding layer to perform the same processing to generate embedded features for text queries.

[0060] 5.2) Each object instance is assigned an object identifier. In order to bind the identifier and its corresponding object in the language model, the object identifier embedding features, object point cloud embedding features and object image embedding features are arranged in order. For the i-th object, its multimodal collaborative representation is

[0061] 5.3) For all n object instances, arrange their multimodal collaborative representations in order to obtain a multimodal collaborative representation sequence By leveraging the understanding and reasoning capabilities of the large language model itself, we can achieve understanding of object information, object identifiers, and the relationships between objects. This sequence ultimately constitutes a multimodal collaborative representation of the entire scene, strengthening the model’s understanding of the entire scene.

[0062] In step 6 above, an optional way to convert different tasks into a unified form is as follows:

[0063] The 3D scene tasks involved in this invention are mainly divided into three types: visual positioning, dense description generation and scene question answering. Among them, visual positioning requires the model to output the 3D bounding box of the queried object in the 3D scene, dense description generation requires the model to generate descriptions of the specified objects in the 3D scene, and scene question answering requires the model to give a brief answer to the questions asked; for each task, the corresponding Chinese question answering template is listed here:

[0064] Visual positioning:

[0065] User: <object description>. Please find the object in the scene that best fits this description and give its identifier.

[0066] Model: <objxxx>。

[0067] Dense description generation:

[0068] User: Please describe <OBJXXX> in the scene in detail, including its appearance and its spatial relationship with surrounding objects.

[0069] Model: <Object description>.

[0070] Scene Q&A:

[0071] User: Please answer this question in short text: <Question>.

[0072] Model: <Answer>.

[0073] Among them, for the visual localization task, the output object identifier can uniquely correspond to an object instance in the scene. Since the corresponding point cloud of the object has been obtained in the previous steps, the corresponding 3D bounding box can be directly calculated as the final answer output. For the dense description generation task, only the object identifier needs to be directly used in the user input text to refer to the object instance. For the scene Q&A task, only through simple prompt engineering, the ability of the large language model itself can be used to control the form of the output answer.

[0074] In the above step 7, an optional way of the single-stage joint training is as follows:

[0075] After converting all 3D scene tasks into a unified form through the previous steps, the training data of different 3D scene downstream datasets can be mixed together for single-stage joint training. In this embodiment, five datasets are jointly trained, and the cross-entropy loss function of the language model is directly used to supervise the model training; the training objective is to optimize all trainable parameters (denoted as θ here), so as to minimize the negative log-likelihood of the target answer sequence s res ;

[0076] Given the previous sequence s prefix (including the multi-modal collaborative representation sequence of the 3D scene and the text query input by the user), the calculation formula of the loss function is:

[0077]

[0078] where k is the length of the target answer sequence of the text query, is the sequence composed of the first i - 1 tokens in the answer sequence, and the trainable parameters θ include the point cloud-text projection layer, the image-text projection layer, the text embedding layer, and the language model itself.

[0079] Using this loss function and the gradient descent learning method, the trainable parameters θ in the model are trained to achieve joint training on multiple data sets, so that the model can complete multiple three-dimensional scene tasks at the same time.

[0080] The above method is applied to the following embodiments to demonstrate the technical effects of the present invention, and the specific steps in the embodiments are not repeated here.

[0081] The present invention conducts experiments on five 3D scene datasets ScanRefer, Multi3DRefer, Scan2Cap, ScanQA and SQA3D, among which ScanRefer and Multi3DRef are used to test the 3D scene visual positioning task, Scan2Cap is used to test the 3D scene dense description generation task, and ScanQA and SQA3D are used to test the 3D scene question answering task. In order to objectively evaluate the performance of the present invention, the present invention uses the Acc and F1 evaluation criteria to evaluate the visual positioning effect in the selected test set, and uses the EM, BLEU, CIDEr, METEOR, ROUGE and SPICE evaluation criteria to evaluate the description generation and question answering effects, and compares with the following existing models:

[0082] Comparison 1. Traditional expert models, including ScanRefer, InstanceRefer, 3DVG-Transformer, ConcreteNet, Scan2Cap, ScanQA, Vote2Cap-DETR and SQA3D. These expert models are designed for a specific task and have achieved good performance, but are limited by their model structure and are difficult to be extended to other 3D scene tasks.

[0083] Comparison 2. Multi-task expert models, including 3DJCG, 3D-VLP, M3DRef-CLIP, D3Net and 3D-VisTA, add different task heads on the basis of the backbone network to adapt to the output required by different tasks, so as to solve multiple tasks with one model. However, these task heads usually require special design and have fixed output forms. They lack flexibility in practical applications and cannot handle complex scenarios and tasks in practical applications.

[0084] Comparison 3. Multimodal large language models, including 3D-LLM, LL3DA and Scene-LLM. These models integrate 3D point cloud modal information into the large language model and use the large language model to achieve general reasoning and dialogue. However, the current 3D multimodal large language model still has a large gap with the expert model in tasks such as 3D visual positioning due to the lack of effective object positioning methods.

[0085] According to the steps described in the specific implementation manner, the experimental results obtained are shown in Tables 1 to 5, and the model of the present invention is represented by SceneChat.

[0086] Table 1: Test results of the present invention for the three-dimensional scene single target visual positioning obtained by ScanRefer dataset

[0087]

[0088]

[0089] Table 2: Test results of multi-target visual positioning in three-dimensional scenes obtained by the present invention for the Multi3DRefer dataset

[0090]

[0091] Table 3: Test results of the present invention for the generation of dense descriptions of 3D scenes obtained from the Scan2Cap dataset

[0092]

[0093] Table 4: Test results of the three-dimensional scene question answering obtained by the present invention on the ScanQA dataset

[0094] Model EM B-1 B-4 ROUGE METEOR CIDEr SPICE ScanQA 21.05 30.24 10.08 33.33 13.14 64.86 13.43 3D-VLP 21.65 30.53 11.15 34.51 13.53 66.97 14.18 3D-LLM 20.50 39.30 12.00 35.70 14.50 69.40 - LL3D - - 13.53 37.31 15.88 76.79 - SceneChat (the present invention) 21.47 43.20 14.26 41.67 18.33 88.49 20.76

[0095] Table 5: Test results of the three-dimensional scene question answering obtained by the present invention for the SQA3D dataset

[0096] Model What Is How Can Which Others Average SQA3D 31.6 63.8 46.0 69.5 43.8 45.3 46.5 3D-VisTA 34.8 63.3 45.4 69.8 47.2 48.1 48.5 Scene-LLM 40.9 69.1 45.0 70.8 47.2 52.3 54.2 SceneChat (the present invention) 48.9 66.0 55.9 67.2 47.5 53.7 55.5

[0097] It can be found from Table 1 and Table 2 that the 3D visual positioning performance of the method of the present invention is significantly better than the original advanced method. For the single-target visual positioning task ScanRefer, the Acc@0.25 score is improved from 51.90 to 55.52 (6.97%), and the Acc@0.5 score is improved from 46.53 to 50.23 (7.95%); for the multi-target visual positioning task Multi3DRefer, the F1@0.25 score is improved from 42.8 to 57.1 (33.41%), and the F1@0.5 score is improved from 38.4 to 52.4 (36.45%). Benefiting from the multimodal collaborative representation of the scene by the present invention, object identifier information is added to clearly refer to objects in the scene. The model of the present invention can introduce object identifiers in question-answering interaction, realizing efficient reference and positioning of scene objects.

[0098] It can be seen from Tables 3, 4, and 5 that the method of the present invention also achieves the most advanced performance for description generation and text question answering tasks. Benefiting from the unified multimodal collaborative representation of the scene proposed by the present invention, the model of the present invention not only obtains rich shape and spatial information from the 3D point cloud modality, but also obtains rich semantic information from the 2D image modality, which enhances the overall understanding of the scene and thus achieves advanced description and question answering effects.

[0099] It is worth noting that all experimental results in Tables 1 to 5 are based on the same trained model of the present invention, which demonstrates the powerful multi-task processing capability of the model of the present invention.

[0100] In this embodiment, a three-dimensional scene perception interaction system based on multimodal collaborative representation is also provided, which is used to implement the above embodiment. The terms "module", "unit", etc. used below can implement a combination of software and / or hardware for predetermined functions. Although the system described in the following embodiments is preferably implemented in software, it is also possible to implement hardware, or a combination of software and hardware.

[0101] The object instance segmentation module is used to detect object instances from the original 3D scene using an object detector, extract object point clouds, and obtain multi-view projections of object instances in combination with multi-view depth images;

[0102] A point cloud feature encoder, which is used to extract point cloud features of an object;

[0103] An image feature encoder, which is used to extract object image features;

[0104] Point cloud-text projection layer, which is used to project the object point cloud features into the embedding feature space of the language model;

[0105] The image-text projection layer is used to project the object image features into the embedding feature space of the language model;

[0106] A text embedding layer, which is used to obtain embedded features of object identifiers and text queries;

[0107] The object joint representation module is used to connect the embedded features of the introduced object identifiers with the object point cloud embedded features and the object image embedded features one by one to form a multimodal collaborative representation of the object; the embedded features of the text query are spliced ​​after the multimodal collaborative representation as the input of the language model;

[0108] A language model that generates answers to text queries based on a preamble consisting of a multimodal co-representation of objects and an embedded feature of the text query.

[0109] The training module is used to combine downstream task datasets corresponding to multiple 3D scenes to perform single-stage joint training on the learnable point cloud-text projection layer, image-text projection layer, text embedding layer and language model.

[0110] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be repeated here. The system embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.

[0111] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, the corresponding computer program instructions in the non-volatile memory are read into the memory by the processor of any device with data processing capabilities and run.

[0112] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.< / objxxx>

Claims

1. A three-dimensional scene perception interaction method based on multimodal collaborative representation, characterized in that: include: Obtain 3D scene point cloud data, multi-view depth image data and text query annotations as training sets, use object detectors to detect object instances from the original 3D scene, extract object point clouds, and combine multi-view depth images to obtain multi-view projections of object instances; The object point cloud features are extracted using a pre-trained point cloud encoder, and the object image features are extracted from the multi-view depth image and multi-view projection using a pre-trained image encoder. Then, two learnable linear projection layers are used to project the object point cloud features and the object image features into the embedding feature space of the language model to obtain the object point cloud embedding features and the object image embedding features. The method for extracting object point cloud features and object image features comprises: inputting the point cloud corresponding to the object instance into a pre-trained point cloud feature encoder to extract the object point cloud features; inputting the multi-view depth image into a pre-trained image encoder to extract the feature map of each depth image, extracting the feature vector of the pixel points where the object instance exists from the feature map of each depth image by multi-view projection, taking the mean as the image feature of the object instance corresponding to the depth image, taking the weighted average according to the projection area of ​​the object instance in each depth image, and obtaining the final object image feature; traversing all object instances to obtain all object point cloud features and object image features; The object identifier and text query are segmented and embedded features are obtained using a learnable text embedding layer. The embedded features of the object identifier are connected to the object point cloud embedding features and the object image embedding features to form a multimodal collaborative representation of the object. The embedded features of the text query are concatenated with the multimodal collaborative representation as the input of the language model, and the language model is used to generate the answer to the text query. Using the training set, the learnable linear projection layer, text embedding layer and language model are jointly trained in a single stage, and the trained model is used to complete perceptual interaction of various three-dimensional scene tasks.

2. The three-dimensional scene perception interaction method based on multimodal collaborative representation according to claim 1 is characterized in that: The detecting of object instances from the original three-dimensional scene and extracting the object point cloud comprises: Use the pre-trained 3D object instance segmentation model as the object detector, input the original 3D scene point cloud, extract the n object instances with the highest confidence from the scene, obtain the segmentation mask of each object instance, and use the segmentation mask to extract the point cloud corresponding to each object from the original 3D scene point cloud. Each point in the point cloud contains the 3D coordinates and RGB color of the point.

3. The three-dimensional scene perception interaction method based on multimodal collaborative representation according to claim 1 is characterized in that: The method of obtaining a multi-view projection of an object instance by combining a multi-view depth image includes: The multi-view depth image and the camera's internal and external parameter information are used to obtain the mapping relationship between the pixels in each depth image and the three-dimensional point cloud. For each point cloud extracted from the object, the pixel position in the corresponding multi-view depth image is obtained to obtain the multi-view projection.

4. The three-dimensional scene perception interaction method based on multimodal collaborative representation according to claim 1 is characterized in that: The object identifier and the text query are segmented and the embedding features are obtained by using a learnable text embedding layer, and the embedding features of the object identifier are connected with the object point cloud embedding features and the object image embedding features to form a multimodal collaborative representation of the object, including: Introducing object identifiers with the same number as object instances, which are used to refer to objects in text interaction, and transforming different 3D scene tasks into a unified form; Tokenize object identifiers and text queries and obtain embedded features using a learnable text embedding layer; Arrange the object identifier embedding features, point cloud embedding features and image embedding features in order to obtain the multimodal collaborative representation of the object Among them, i represents the embedding feature of the i-th object identifier, are the point cloud embedding features and image embedding features of the i-th object instance respectively.

5. The three-dimensional scene perception interaction method based on multimodal collaborative representation according to claim 1 is characterized in that: Different three-dimensional scene tasks are converted into a unified question-answering format using object identifiers, wherein the three-dimensional scene tasks include a visual positioning task, a dense description generation task, and a scene question-answering task.

6. The three-dimensional scene perception interaction method based on multimodal collaborative representation according to claim 1 is characterized in that: In the single-stage joint training, the loss function is calculated as: in, is the loss function, s prefix is a preamble sequence consisting of the multimodal co-representation of the object and the embedded features of the text query, k is the length of the target answer sequence of the text query, is the sequence of the first i-1 words in the answer sequence, θ is a trainable parameter, is the probability of generating the i-th word in the answer sequence given the first i-1 words in the preamble sequence and the answer sequence.

7. A three-dimensional scene perception interaction system based on multimodal collaborative representation, used to implement the method of claim 1, characterized in that: include: The object instance segmentation module is used to detect object instances from the original 3D scene using an object detector, extract object point clouds, and obtain multi-view projections of object instances in combination with multi-view depth images; A point cloud feature encoder, which is used to extract point cloud features of an object; The object point cloud feature extraction method comprises: inputting the point cloud corresponding to the object instance into a pre-trained point cloud feature encoder to extract the object point cloud features; traversing all object instances to obtain all object point cloud features; An image feature encoder is used to extract object image features; the object image feature extraction method comprises: inputting a multi-view depth image into a pre-trained image encoder, extracting a feature map of each depth image, extracting a pixel feature vector of an object instance from the feature map of each depth image using multi-view projection, taking the mean as the image feature of the object instance corresponding to the depth image, taking a weighted average according to the projection area of ​​the object instance in each depth image, and obtaining the final object image feature; traversing all object instances to obtain all object image features; Point cloud-text projection layer, which is used to project the object point cloud features into the embedding feature space of the language model; The image-text projection layer is used to project the object image features into the embedding feature space of the language model; A text embedding layer, which is used to obtain embedded features of object identifiers and text queries; The object joint representation module is used to connect the embedded features of the introduced object identifiers with the object point cloud embedded features and the object image embedded features one by one to form a multimodal collaborative representation of the object; the embedded features of the text query are spliced ​​after the multimodal collaborative representation as the input of the language model; A language model that generates answers to text queries based on a preamble consisting of a multimodal co-representation of objects and an embedded feature of the text query; The training module is used to combine downstream task datasets corresponding to multiple 3D scenes to perform single-stage joint training on the learnable point cloud-text projection layer, image-text projection layer, text embedding layer and language model.

8. The three-dimensional scene perception interaction system based on multimodal collaborative representation according to claim 7 is characterized in that: The training module adopts the cross entropy loss function of the language model.

9. The three-dimensional scene perception interaction system based on multimodal collaborative representation according to claim 7, characterized in that: The downstream task data sets corresponding to the multiple three-dimensional scenes include a visual positioning task data set, a dense description generation task data set and a scene question answering task data set.

Citation Information

Patent Citations

  • 3D visual question and answer method based on three-mode knowledge distillation

    CN117216225A