Self-adaptive three-dimensional large language model system based on query guidance
By using query-guided adaptive pruning and multimodal vector representation enhancement modules, the problems of redundant object information and semantic lack in 3D large language models are solved, achieving more efficient 3D scene understanding and reasoning capabilities, and improving the accuracy of 3D question answering and description generation.
Patent Information
- Application Number
- CN202511369889.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-24
AI Technical Summary
Existing 3D large language models suffer from problems such as redundant object information interfering with reasoning and a lack of 2D image semantic information in 3D scene understanding tasks, which limits the accuracy of reasoning and semantic understanding capabilities.
The system employs a query-guided adaptive cropping module (QGAP) and a multimodal object-level vector representation enhancement module (MOFE). By calculating the relevance of objects to the task and fusing semantic information from two-dimensional images, it adaptively selects key object-level vector representations and enhances the three-dimensional scene feature matrix.
It improves the accuracy and robustness of 3D scene understanding, reduces noise interference, enhances semantic understanding and reasoning capabilities for complex scenes, and improves the performance of 3D question answering, description generation, and interactive reasoning tasks.
Smart Images

Figure CN120849595A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of 3D scene understanding and multimodal artificial intelligence, and in particular relates to a 3D large language model system that integrates query-guided adaptive pruning and multimodal semantic enhancement mechanisms. It can be widely used in tasks such as 3D scene question answering, description generation and multimodal reasoning. Background Art
[0002] With the development of artificial intelligence technology, 3D scene understanding has become a core research direction in fields such as autonomous driving, augmented reality and virtual reality, robot navigation, and intelligent manufacturing. Efficiently and accurately analyzing the spatial structure and object relationships in a 3D environment is of great significance for enhancing the perception, scene understanding, and environmental interaction capabilities of AI systems.
[0003] In recent years, the rapid development of large language models (LLMs) has driven their application in multimodal tasks beyond natural language processing, especially demonstrating strong reasoning and generalization capabilities in 3D semantic understanding tasks. Many studies have fused LLMs with 3D point cloud data to construct 3D large language models (3D-LLMs) for tasks such as scene description, question answering reasoning, and visual localization, significantly improving the cognitive level of AI systems in complex 3D environments. Despite the progress made by existing methods, the current mainstream 3D-LLM architecture still faces the following two key challenges:
[0004] Redundant object information interferes with reasoning: Existing methods generally use 3D encoders such as PointNet and PointNet++ to encode the entire scene into a large number of object-level tokens. Only some of these objects are relevant to the current task, while the rest of the irrelevant information introduces noise, increases the risk of model illusion, and affects the accuracy of reasoning.
[0005] Lack of semantically rich 2D visual vector representations: While 3D point cloud data has advantages in spatial structure, it lacks semantic information such as color, texture, material properties, and high-level contextual relationships, limiting the model's comprehensive understanding of complex scenes. Furthermore, most existing methods fail to effectively integrate multimodal semantic vector representations from 2D images, resulting in limited semantic representation capabilities.
[0006] Furthermore, current state-of-the-art (SOTA) 3D-LLM methods such as LEO and Point-LLM often simply concatenate object-level vector representations with text tokens and input them directly into a large language model, neglecting the systematic modeling of the two issues mentioned above. This hinders the further development of models in scene semantic understanding, spatial reasoning, and task execution. Therefore, there is an urgent need for a novel 3D language model architecture that integrates task-related guidance and multimodal vector representation enhancement to improve the model's ability to select key objects, enhance the semantic expression of 3D tokens, and ultimately achieve accurate perception, semantic understanding, and intelligent response to complex 3D environments. Summary of the Invention
[0007] This invention aims to address two key problems existing in current 3D Large Language Models (3D-LLMs) for 3D scene understanding tasks: First, existing models often input a large number of task-irrelevant object-level vector representations, leading to noise and illusions during inference and reducing understanding accuracy; second, relying solely on point cloud data while ignoring the rich semantic information in 2D images makes it difficult to achieve deep multimodal semantic fusion and scene understanding. Therefore, this invention proposes a novel 3D language model system that integrates a task-relevance-guided object selection mechanism with a multimodal vector representation enhancement strategy, significantly improving the accuracy and robustness of 3D scene question answering, generation, and inference tasks.
[0008] The technical solution of this invention:
[0009] A query-guided adaptive three-dimensional large language model system includes the following steps:
[0010] Step 1: 3D vision-language alignment construction;
[0011] The SentencePiece word segmenter is used to encode the system prompt text and user questions. Simultaneously, a pre-trained PointNet++ point cloud encoder is used to extract point-level vector representations of each object in the input 3D scene, obtaining the geometric vector representation of each object. Subsequently, the geometric vector representations of each object are converted into object-level vector representations at a uniform scale to obtain the 3D feature matrix. Simultaneously, predict the category label for each object in the 3D scene. ,in, Refers to the first in a three-dimensional scene One object, The number of objects in the 3D scene. The vector dimension is matched to the input dimension of the large language model; after the three-dimensional feature matrix undergoes query-guided adaptive pruning in step 2 and multimodal object-level vector representation enhancement in step 3, it is fed into the large language model along with the system prompt text and the user question code to obtain the language response sequence. ;
[0012] Step 2: Query-guided adaptive pruning;
[0013] (2.1) Question representation generation: The user question is encoded using a frozen BERT encoder to obtain the query semantic vector. Meanwhile, the BERT encoder converts the category label of each object into a category semantic vector. ;
[0014] (2.2) Semantic relevance calculation: Calculate the category semantic vector and the query semantic vector. The cosine similarity is used to form a list of semantic relevance:
[0015]
[0016] (2.3) Global task-guided modeling: Constructing a learnable global task query vector Through its relationship with the query semantic vector Cross-attention, two layers of self-attention, and pooling operations are used to obtain the pruning scaling factor. Calculate the number of objects that need to be retained. :
[0017]
[0018]
[0019] in, This represents the total number of objects in the current 3D scene. This indicates a linear layer, used to perform linear transformations on the input vector; This represents the pooling operation, used to aggregate vector representations; This indicates that two self-attention layers are used to capture the correlation between elements within a sequence; Indicates a cross-attention layer; Indicates the rounding operation;
[0020] (2.4) Based on the similarity ranking results, find the top-ranked features in the three-dimensional feature matrix. The object-level vector representation most relevant to the task;
[0021] Step 3: Enhancement of multimodal object-level vector representation;
[0022] (3.1) Point-level two-dimensional vector representation extraction: First, the pixel-level vector representation of each multi-view image is extracted from the multi-view images corresponding to the three-dimensional scene using a pre-trained 2D encoder. ,in, and These represent the height and width of the multi-view image, respectively. The number of channels is represented by a pixel-level vector; the point cloud data corresponding to the 3D scene is obtained by using camera intrinsic and extrinsic parameters. Projected onto image pixels The corresponding pixel vector representations are averaged from multiple perspectives to form a point-level two-dimensional vector representation. , The number of points in the point cloud data corresponding to the 3D scene:
[0023]
[0024] in, The number of multi-view images in a 3D scene;
[0025] (3.2) Object-level 2D vector representation fusion and alignment: Using the object mask corresponding to the 3D scene, the point-level 2D vector representation is fused into an object-level 2D vector representation. And through a linear mapping matrix Mapping to dimensions that match the large language model:
[0026]
[0027] (3.3) Cross-modal fusion operation: using a three-dimensional feature matrix As a query vector, and the mapped object-level two-dimensional vector representation As keys, they are fused into a fusion vector representation through a cross-attention mechanism:
[0028]
[0029] in, It is the mapping matrix in cross-attention. These represent the query vector, key vector, and value vector in the cross-attention calculation, respectively. This represents the normalization function; subsequently, through a linear layer, the normalized random inactivation and residual connection are normalized, followed by another normalization operation, and finally, the result is calculated based on the value obtained in step (2.3). Values are selected from the three-dimensional feature matrix. The enhanced 3D feature matrix is obtained by using the object-level vector representations most relevant to the task. ;
[0030] Step 4: Language modeling and training optimization;
[0031] (4.1) Input Construction: The system prompt text and user question encoded by the SentencePiece tokenizer and the three-dimensional feature matrix. The sequences are concatenated to form a unified multimodal input sequence, which is then fed into a pre-trained large language model to generate language responses. ;
[0032] (4.2) Definition of training objective function: for language response sequences The conditional language modeling loss function is defined as follows:
[0033]
[0034] in, Indicates a batch size. Indicates the length of the target response sequence. Indicates the current time step of the forecast. This indicates the position of the currently processed data within the batch. This includes system prompt text, user questions, and the final enhanced 3D feature matrix. This indicates the language response result for the current batch and the current prediction time step. This indicates the language response results for the current batch and the time step prior to the current time step. Indicates by parameters Parameterized conditional probability distribution.
[0035] The beneficial effects of this invention are:
[0036] (1) The query-guided adaptive pruning module aims to address the problem of excessive objects and insufficient task relevance in existing 3D large language models. Based on the user's task instructions or question content, QGAP automatically evaluates the relevance between the vector representations of each object in the scene and the user's question. The system introduces a learnable global task query vector to interact with the user's question and adaptively determine the number of objects to retain. Ultimately, the model only retains the vector representations of objects closely related to the task, eliminating redundant or irrelevant information, thereby reducing inference noise and illusions and improving the model's understanding and inference accuracy in complex scenes.
[0037] (2) Multimodal Object-Level Vector Enhancement: The enhancement module aims to address the problem that existing methods rely solely on point cloud data and lack rich semantic information. MOFE projects the semantic vector representations from multi-view images into three-dimensional space by fusing two-dimensional images and three-dimensional point cloud information. Based on this, the two-dimensional vector representations are aggregated to the object level and then fused with the object-level vector representations filtered by the QGAP module through cross-attention calculation, so that the final enhanced three-dimensional feature matrix possesses both geometric details and semantic information. Through this multimodal vector representation enhancement mechanism, the model can better understand the material, color, semantic attributes, and contextual relationships of objects, achieving a richer and more accurate understanding of three-dimensional scenes.
[0038] (3) Query-guided adaptive 3D large language model system (3D-SceneQ): The overall system 3D-SceneQ of this invention is based on the two key modules mentioned above, constructing a novel query-guided adaptive 3D large language model system. First, the QGAP module adaptively filters object-level vector representations according to task requirements, ensuring that the input vector representation data is more compact and task-relevant. Subsequently, MOFE further integrates two-dimensional semantic information with three-dimensional geometric vector representations to generate a semantically rich three-dimensional feature matrix. Finally, the enhanced three-dimensional feature matrix, system prompt text, and user questions are input into the large language model to complete the question-answering, description, and reasoning tasks of the three-dimensional scene. Through this design, 3D-SceneQ not only effectively reduces noise and illusion problems, but also enhances the model's semantic understanding and reasoning ability for complex scenes, demonstrating excellent accuracy and robustness in tasks such as 3D question answering, 3D description generation, and interactive reasoning. Attached Figure Description
[0039] Figure 1 This is an overview diagram of the 3D-SceneQ framework;
[0040] Figure 2 This is a flowchart of the Query Guided Adaptive Clipping (QGAP) module;
[0041] Figure 3 This is a flowchart of the Multimodal Object-Level Vector Representation Enhancement Module (MOFE). Detailed Implementation
[0042] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.
[0043] Example
[0044] A query-guided adaptive 3D large language model system, 3D-SceneQ, integrates a query-guided adaptive pruning mechanism and a multimodal semantic enhancement strategy to improve semantic understanding and language reasoning capabilities in complex 3D scenes. Specifically, it includes the following steps:
[0045] Step 1: 3D vision-language alignment construction;
[0046] like Figure 1 As shown, the SentencePiece word segmenter is used to encode the system prompt text and user questions. Simultaneously, a pre-trained PointNet++ point cloud encoder is used to extract point-level vector representations of each object in the input 3D scene, obtaining the geometric vector representation of each object. Subsequently, the geometric vector representations of each object are converted into object-level vector representations at a uniform scale, resulting in the 3D feature matrix. Simultaneously, predict the category label for each object in the 3D scene. ,in, Refers to the first in a three-dimensional scene One object, The number of objects in the 3D scene. The vector dimension is matched to the input dimension of the large language model; after the three-dimensional feature matrix undergoes query-guided adaptive pruning in step 2 and multimodal object-level vector representation enhancement in step 3, it is fed into the large language model along with the system prompt text and the user question code to obtain the language response sequence. ;
[0047] Step 2: Query-guided adaptive pruning;
[0048] (2.1) As Figure 2 As shown, the question representation is generated by encoding the user question using a frozen BERT encoder to obtain a query semantic vector. Meanwhile, the BERT encoder converts the category label of each object into a category semantic vector. ;
[0049] (2.2) Semantic relevance calculation: Calculate the category semantic vector and the query semantic vector. The cosine similarity is used to form a list of semantic relevance:
[0050]
[0051] (2.3) Global task-guided modeling: Constructing a learnable global task query vector Through its relationship with the query semantic vector Cross-attention, two layers of self-attention, and pooling operations are used to obtain the pruning scaling factor. Calculate the number of objects that need to be retained. :
[0052]
[0053]
[0054] in, This represents the total number of objects in the current 3D scene. This indicates a linear layer, used to perform linear transformations on the input vector; This represents the pooling operation, used to aggregate vector representations; This indicates that two self-attention layers are used to capture the correlation between elements within a sequence; Indicates a cross-attention layer; Indicates the rounding operation;
[0055] (2.4) Based on the similarity ranking results, find the top-ranked features in the three-dimensional feature matrix. The object-level vector representation most relevant to the task;
[0056] Step 3: Enhancement of multimodal object-level vector representation;
[0057] (3.1) Extraction of point-level two-dimensional vector representation: such as Figure 3 As shown, firstly, a pre-trained 2D encoder is used to extract pixel-level vector representations of each multi-view image from the multi-view images corresponding to the 3D scene. ,in, and These represent the height and width of the multi-view image, respectively. The number of channels is represented by a pixel-level vector; the point cloud data corresponding to the 3D scene is obtained by using camera intrinsic and extrinsic parameters. Projected onto image pixels The corresponding pixel vector representations are averaged from multiple perspectives to form a point-level two-dimensional vector representation. , The number of points in the point cloud data corresponding to the 3D scene:
[0058]
[0059] in, This represents the number of multi-view images in a 3D scene.
[0060] (3.2) Object-level 2D vector representation fusion and alignment: Using the object mask corresponding to the 3D scene, the point-level 2D vector representation is fused into an object-level 2D vector representation. And through a linear mapping matrix Mapping to dimensions that match the large language model:
[0061]
[0062] (3.3) Cross-modal fusion operation: using a three-dimensional feature matrix As a query vector, and the mapped object-level two-dimensional vector representation As keys, they are fused into a fusion vector representation through a cross-attention mechanism:
[0063]
[0064] in, It is the mapping matrix in cross-attention. These represent the query vector, key vector, and value vector in the cross-attention calculation, respectively. This represents the normalization function; subsequently, through a linear layer, the normalized random inactivation and residual connection are normalized, followed by another normalization operation, and finally, the result is calculated based on the value obtained in step (2.3). Values are selected from the three-dimensional feature matrix. The enhanced 3D feature matrix is obtained by using the object-level vector representations most relevant to the task. ;
[0065] Step 4: Language modeling and training optimization;
[0066] (4.1) Input Construction: The system prompt text and user question encoded by the SentencePiece tokenizer and the three-dimensional feature matrix. The sequences are concatenated to form a unified multimodal input sequence, which is then fed into a pre-trained large language model to generate language responses. ;
[0067] (4.2) Definition of training objective function: for language response sequences The conditional language modeling loss function is defined as follows:
[0068]
[0069] in, Indicates a batch size. Indicates the length of the target response sequence. Indicates the current time step of the forecast. This indicates the position of the currently processed data within the batch. This includes system prompt text, user questions, and the final enhanced 3D feature matrix. This indicates the language response result for the current batch and the current prediction time step. This indicates the language response results for the current batch and the time step prior to the current time step. Indicates by parameters Parameterized conditional probability distribution.
[0070] Experimental setup and effect verification:
[0071] To evaluate the proposed 3D-SceneQ model, multiple 3D vision-language datasets were used in the experimental datasets, covering the alignment and instruction fine-tuning processes for different tasks, as detailed below:
[0072] (1) 3D Vision-Language Alignment Stage: The model is trained to bridge the gap between 3D scene representation and natural language using three types of descriptive data: Object-level descriptions: Data from Cap3D provides fine-grained textual descriptions from the Objaverse, detailing different 3D objects. Object referential representations in context: Data from ScanScribe and ReferIt3D provides referential information about objects in a 3D scene based on user questions, and uses LLM-generated annotations to enhance contextual diversity. Scene-level summaries: Data from 3RScan captures global information such as object layout, spatial relationships, and scene semantics.
[0073] (2) Instruction Fine-tuning Stage: In this stage, the model's ability to understand, reason, and respond to natural language instructions in the 3D environment is improved through instruction fine-tuning. Datasets based on ScanNet and 3RScan were used, covering open-ended 3D scene description and question answering tasks. The model needs to generate scene descriptions or answer questions about object attributes, locations, and relationships.
[0074] (3) Evaluation metrics: To evaluate the model's performance in different tasks, a series of standard evaluation metrics were used, as follows:
[0075] 1) Language generation task: For the language generation task, CIDEr, BLEU, METEOR and ROUGE are reported to measure the relevance, fluency and overlap with the reference text of the generated text, respectively.
[0076] 2) Open-ended generation tasks: In open-ended generation tasks, since multiple correct answers may exist, sentence similarity metrics are also used to better capture semantic alignment. Question answering tasks: For question answering tasks that require factual accuracy, EM@1 (Top-1 accuracy) is reported, which reflects the proportion of times the model's predictions perfectly match the real answers.
[0077] Implementation Details: In terms of implementation, Vicuna-7B was used as the backbone network of the language model, directly using a pre-trained version of the LEO model for initialization. The model components are as follows: the 3D encoder, the 2D semantic segmentation model used in MOFE, the language segmenter and language backbone network, and the BERT encoder in QGAP, all of which were kept frozen during training. The LoRA method was used for efficient parameter fine-tuning of Vicuna-7B, with the low-rank set to r=16. Other trainable parameters, including QGAP, MOFE, and the object-to-text mapping layer, were trained from scratch. The training process used the AdamW optimizer with a learning rate of... A total of 10 rounds of training were conducted. All experiments were performed on two NVIDIA A8000 GPUs.
[0078] Table 1 compares the performance of 3D-SceneQ with current mainstream models on Scan2Cap and SQA3D benchmark tests.
[0079]
[0080] Here, "EM@1" represents the Top-1 exact match accuracy; the n-gram evaluation metric in Scan2Cap uses an IoU threshold of ≥0.5.
[0081] Table 2 compares the performance of 3D-SceneQ with current mainstream models on the ScanQA benchmark.
[0082]
[0083] “EM@1” indicates Top-1 exact match accuracy.
[0084] Table 3 compares the performance of 3D-SceneQ with current mainstream models on Scan2Cap and SQA3D benchmark tests.
[0085]
[0086] “EM@1” indicates Top-1 exact match accuracy; the n-gram evaluation metric in Scan2Cap uses an IoU threshold of ≥0.5.
[0087] Table 4 compares the performance of 3D-SceneQ with current mainstream models on the ScanQA benchmark.
[0088]
[0089] “EM@1” indicates Top-1 exact match accuracy.
[0090] As shown in Tables 1 and 2, 3D-SceneQ achieved state-of-the-art (SOTA) performance on the Scan2Cap, ScanQA, and SQA3D tasks, surpassing powerful task-specific benchmark models and fine-tuning-specific models. For example, in the Scan2Cap task, 3D-SceneQ improved METEOR to 30.5 and ROUGE to 65.3, exceeding the previous best models by 2.6 and 7.2 points, respectively. In the ScanQA task, CIDEr improved to 104.7 and BLEU-4 improved to 13.6, reflecting that the model generated more relevant and fluent answers. On the SQA3D test set, EM@1 reached 67.0%, which is 17 percentage points higher than the strongest pre-test method, demonstrating its ability to generate accurate and context-rich outputs.
[0091] Key Module Effectiveness Verification: By comparing with several state-of-the-art (SOTA) methods, the superior performance of the proposed method in 3D scene understanding tasks was verified. Specifically, a series of ablation experiments were conducted to deeply analyze the individual and combined effects of the two core modules, QGAP and MOFE. Tables 3 and 4 show the performance comparison of QGAP and MOFE. Using the removal of components related to the agent representation in the LEO architecture as a baseline model for comparison, the following four main conclusions were drawn:
[0092] Random pruning helps reduce redundancy but may lose important information. The RandFilter variant randomly discards 40% of objects, but this indiscriminate compression causes Scan2Cap's CIDEr to drop from 72.4 to 52.3. This result suggests that reducing redundant information helps inference effectiveness, but pruning must be based on task relevance.
[0093] QGAP improves description quality and inference accuracy. By learning the filtering ratio based on the user's question, QGAP retains objects that are highly relevant to the task, improving Scan2Cap CIDEr to 78.4 and SQA3D EM@1 to 70.1, outperforming the baseline model and random pruning.
[0094] MOFE improves overall performance, especially on the ScanQA task. By fusing geometric, visual, and speech signals, MOFE improves Scan2Cap CIDEr to 77.8, showing a significant performance boost, particularly in the ScanQA task.
[0095] The synergy between QGAP and MOFE delivers the most balanced performance improvements. 3D-SceneQ combines QGAP’s task-related clipping with MOFE’s enhancements, achieving significant performance improvements across all benchmarks, such as a 5.7-point improvement in Scan2Cap CIDEr and a 20.0-point improvement in SQA3D EM@1.
Claims
1. A query-guided adaptive three-dimensional large language model system, characterized in that, Includes the following steps: Step 1: 3D vision-language alignment construction; The SentencePiece word segmenter is used to encode the system prompt text and user questions. At the same time, the pre-trained point cloud encoder PointNet++ is used to extract the point-level vector representation of each object in the input 3D scene to obtain the geometric vector representation of each object. The geometric vector representation of each object is then converted into a uniform-scale object-level vector representation to obtain the three-dimensional feature matrix. Simultaneously, predict the category label for each object in the 3D scene. ,in, Refers to the first in a three-dimensional scene One object, The number of objects in the 3D scene. The vector dimension that matches the input dimension of the large language model; Step 2: Query-guided adaptive pruning; Step 3: Enhancement of multimodal object-level vector representation; Step 4: Language modeling and training optimization.
2. The query-guided adaptive three-dimensional large language model system according to claim 1, characterized in that, The specific implementation process of step 2 is as follows: (2.1) Question representation generation: The user question is encoded using a frozen BERT encoder to obtain the query semantic vector. Meanwhile, the BERT encoder converts the category label of each object into a category semantic vector. ; (2.2) Semantic relevance calculation: Calculate the category semantic vector and the query semantic vector. The cosine similarity is used to form a list of semantic relevance: (2.3) Global task-guided modeling: Constructing a learnable global task query vector Through its relationship with the query semantic vector Cross-attention, two layers of self-attention, and pooling operations are used to obtain the pruning scaling factor. Calculate the number of objects that need to be retained. : in, This represents the total number of objects in the current 3D scene. This indicates a linear layer, used to perform linear transformations on the input vector; This represents the pooling operation, used to aggregate vector representations; This indicates that two self-attention layers are used to capture the correlation between elements within a sequence; Indicates a cross-attention layer; Indicates the rounding operation; (2.4) Based on the similarity ranking results, find the top-ranked features in the three-dimensional feature matrix. The object-level vector representation that is most relevant to the task.
3. The query-guided adaptive three-dimensional large language model system according to claim 1, characterized in that, The specific implementation process of step 3 is as follows: (3.1) Point-level two-dimensional vector representation extraction: First, the pixel-level vector representation of each multi-view image is extracted from the multi-view images corresponding to the three-dimensional scene using a pre-trained 2D encoder. ,in, and These represent the height and width of the multi-view image, respectively. The number of channels is represented by a pixel-level vector; the point cloud data corresponding to the 3D scene is obtained by using camera intrinsic and extrinsic parameters. Projected onto image pixels The corresponding pixel vector representations are averaged from multiple perspectives to form a point-level two-dimensional vector representation. , The number of points in the point cloud data corresponding to the 3D scene: in, The number of multi-view images in a 3D scene; (3.2) Object-level 2D vector representation fusion and alignment: Using the object mask corresponding to the 3D scene, the point-level 2D vector representation is fused into an object-level 2D vector representation. And through a linear mapping matrix Mapping to dimensions that match the large language model: (3.3) Cross-modal fusion operation: using a three-dimensional feature matrix As a query vector, and the mapped object-level two-dimensional vector representation As keys, they are fused into a fusion vector representation through a cross-attention mechanism: in, It is the mapping matrix in cross-attention. These represent the query vector, key vector, and value vector in the cross-attention calculation, respectively. This represents the normalization function; subsequently, through a linear layer, the normalized random inactivation and residual connection are normalized, followed by another normalization operation, and finally, the result is calculated based on the value obtained in step (2.3). Values are selected from the three-dimensional feature matrix. The enhanced 3D feature matrix is obtained by representing the object-level vectors most relevant to the task. .
4. The query-guided adaptive three-dimensional large language model system according to claim 1, characterized in that, The specific implementation process of step 4 is as follows: (4.1) Input Construction: The system prompt text and user question encoded by the SentencePiece tokenizer and the three-dimensional feature matrix. The sequences are concatenated to form a unified multimodal input sequence, which is then fed into a pre-trained large language model to generate language responses. ; (4.2) Definition of training objective function: for language response sequences The conditional language modeling loss function is defined as follows: in, Indicates a batch size. Indicates the length of the target response sequence. Indicates the current time step of the forecast. This indicates the position of the currently processed data within the batch. This includes system prompt text, user questions, and the final enhanced 3D feature matrix. This indicates the language response result for the current batch and the current prediction time step. This indicates the language response results for the current batch and the time step prior to the current time step. Indicates by parameters Parameterized conditional probability distribution.
Citation Information
Patent Citations
Music video question answering method based on music feature guidance in few-sample scene
CN120179876A
Blind guiding scene identification method based on multi-modal visual large model
CN120472387A
Intelligent geometric reasoning and semantic understanding method based on three-dimensional large language model
CN120542438A
Visual language multi-modal fusion method based on parameter-free cross attention
CN120654176A
Three-dimensional point cloud data semantic category classification method and system based on multi-modal data
CN120656007A
Cited By
Three-dimensional language understanding method and system based on explicit double-reference cognitive map
CN122265551A
A transformer-based multi-modal indoor three-dimensional scene understanding method
CN122391836A
A Transformer-based multimodal indoor 3D scene understanding method
CN122391836B