Data enhancement method, device and system based on combination of three-dimensional scene and language data

By jointly preprocessing, enhancing and filtering 3D scenes and language data, a high-quality 3D-language joint dataset was constructed, which solved the problems of insufficient scale, diversity and annotation quality of existing datasets, achieved multimodal fusion of data and improved generalization capabilities, and supported 3D visual understanding and robotic tasks in complex scenarios.

CN120671074AActive Publication Date: 2025-09-19BEIJING UNIV OF POSTS & TELECOMM
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510751059.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-19
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Existing 3D scene understanding datasets have deficiencies in data scale, sample diversity, annotation quality, and multimodal fusion, which limit the generalization ability and performance of models in complex scenarios.

Method used

By preprocessing 3D scene data and text annotation data, multimodal data enhancement, and semantic quality filtering, a high-quality 3D-language joint dataset is constructed, including rotation and translation of RGB-D images, synonym replacement and logical inversion of text, and semantic filtering using BERT and RoBERTa models to achieve full integration of multimodal information.

Benefits of technology

A large-scale, high-quality 3D-language joint dataset was generated, which improved the diversity and generalization ability of the data, supported applications such as 3D visual understanding and robotic task planning, and reduced the cost of data collection and annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671074A_ABST
    Figure CN120671074A_ABST
Patent Text Reader

Abstract

The invention provides a data enhancement method, device and system based on combination of a three-dimensional scene and language data. The method comprises the following steps: acquiring 3D scene data and corresponding text annotation data; preprocessing the scene data and the text labeling data to obtain preprocessed 3D-language joint data; and performing multi-modal data enhancement and semantic quality filtering processing on the preprocessed 3D-language joint data in sequence to obtain a target 3D-language joint data set. According to the method, various data sources such as 3D point cloud data, RGB-D images, question and answer pairs and intensive description are integrated, automatic construction of a high-quality large-scale data set is realized by utilizing data preprocessing, multi-modal data enhancement and semantic quality filtering, and the data quality of 3D scene understanding and visual question and answer tasks can be improved while the data quality of the 3D scene understanding and visual question and answer tasks is improved. The diversity and generalization ability of data are enhanced, and powerful support is provided for applications such as 3D visual understanding and robot task planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data enhancement method, device and system based on the combination of three-dimensional scene and language data. Background Art

[0002] With the rapid development of artificial intelligence technology and deep learning algorithms, computer vision and natural language processing technologies have made great progress in image recognition, scene understanding, and language generation. In recent years, visual tasks based on 3D scene data, such as 3D question answering, 3D dense description, and robot interaction, have gradually become a research hotspot. Traditional 3D scene understanding methods mainly rely on devices such as structured light, laser scanning, and RGB-D cameras to collect 3D point cloud data, and then process it in combination with manually annotated text information. Common datasets such as ScanNet, ScanQA, and ScanRefer have provided certain data support for related tasks. However, these datasets still have shortcomings in terms of data scale, sample diversity, and annotation accuracy, which are mainly manifested in the following aspects:

[0003] 1. Limited data scale: Existing datasets are usually based on manual collection and annotation, and the number of samples is relatively limited. This makes it difficult to meet the deep learning model's demand for large-scale training data, limiting the model's generalization ability in complex scenarios.

[0004] 2. Uneven annotation quality: Due to subjective differences and high annotation costs in manual annotation, the semantic information in question-answer pairs and scene descriptions often contains inconsistencies, duplications, ambiguities, and missing information, which may cause the model to learn incorrect semantic mappings during training.

[0005] 3. Insufficient multimodal fusion: Current data construction methods mostly focus on acquiring data from a single modality. However, in the joint modeling of 3D scenes and language data, how to fully integrate 3D point clouds, images, and text information to achieve cross-modal information complementarity and enhancement remains a technical challenge that needs to be solved urgently.

[0006] 4. Limited data augmentation methods: Existing data augmentation methods mainly rely on traditional data augmentation techniques such as rotation, translation, and scaling. However, these methods rarely process text data and cannot effectively improve the data diversity of question-answer pairs or dense descriptions, thus affecting the performance of the model in practical applications. Summary of the Invention

[0007] The purpose of this application is to provide a data enhancement method, device and system based on the combination of three-dimensional scene and language data, which can improve the data quality of 3D scene understanding and visual question answering tasks while enhancing the diversity and generalization ability of data, providing strong support for applications such as 3D visual understanding and robot task planning.

[0008] In a first aspect, the present application provides a data enhancement method based on the combination of three-dimensional scene and language data, the method comprising: acquiring 3D scene data and corresponding text annotation data; the 3D scene data comprising: scene point cloud data, RGB-D images and camera pose information; the text annotation data comprising: question-answer pairs and dense description information related to the location, attributes and relationships of scene objects; preprocessing the 3D scene data and text annotation data respectively to obtain preprocessed 3D-language joint data; the preprocessing comprising: normalization processing of the scene point cloud data, consistency processing of the RGB-D images, and data cleaning processing and grammatical structure analysis processing of the text annotation data; performing multimodal data enhancement on the preprocessed 3D-language joint data respectively to obtain enhanced joint data; the multimodal data enhancement comprising: multiple image expansion operations of rotation, translation and scale transformation of the RGB-D images, as well as data expansion operations of three strategies of synonym replacement, logical inversion and template generation for question-answer pairs and dense description information; performing semantic quality filtering processing on the enhanced joint data to obtain the target 3D-language joint dataset.

[0009] Furthermore, the above-mentioned steps of preprocessing the 3D scene data and text annotation data respectively include: normalizing the scene point cloud data in the 3D scene data, and performing resolution and color consistency processing on the RGB-D image; performing data cleaning on the text annotation data to eliminate duplicates, grammatical errors, and samples with incomplete information; using natural language processing technology to segment the text annotation data after data cleaning, remove stop words, and analyze the grammatical structure to obtain text annotation data with semantic consistency and accuracy that meet the requirements.

[0010] Furthermore, the above-mentioned steps of performing multimodal data enhancement on the preprocessed 3D-language joint data to obtain enhanced joint data include: performing multiple operations such as rotation, translation and scale transformation on the RGB-D image in the preprocessed 3D scene data to achieve image expansion; based on the pre-trained language model, using three strategies of synonym replacement, logical inversion, and template generation to perform data expansion on question-answer pairs and dense description information; combining the image-expanded 3D scene data and the text annotation data after data expansion to obtain enhanced joint data.

[0011] Furthermore, the above-mentioned step of performing semantic quality filtering on the enhanced joint data to obtain the target 3D-language joint data includes: using the pre-trained BERT model to calculate the semantic similarity between the generated question and answer text and the original question and answer text, and filtering the text with a semantic similarity less than a first threshold; using the RoBERTa model to perform natural language inference, calculate the consistency score between the generated dense description text and the original dense description text, and filter the description text with a consistency score less than a second threshold to obtain the target 3D-language joint data.

[0012] Furthermore, the above method also includes: encoding the text data, visual cue data and three-dimensional point cloud data in the target 3D-language joint dataset respectively to obtain the encoding features corresponding to the three types of data; fusing the encoding features corresponding to the three types of data to obtain fused features; inputting the fused features into a pre-trained large-scale language model for natural language decoding to obtain the target natural language.

[0013] Furthermore, the above-mentioned steps of respectively encoding the text data, visual cue data and three-dimensional point cloud data in the target 3D-language joint data to obtain encoding features corresponding to the three types of data include: encoding the text data based on the Transformer architecture to generate high-dimensional representation features; flattening and linearly mapping the visual cue data through a multi-layer perceptron to obtain spatial prior embedding features of the visual cue data; and using the Vote2Cap-DETR++ network to extract features from the three-dimensional point cloud data to obtain deep geometric features of the three-dimensional point cloud data.

[0014] Furthermore, the above-mentioned step of fusing the encoding features corresponding to the three types of data to obtain fused features includes: preliminarily fusing the spatial prior embedding features corresponding to the visual cue data and the deep geometric features corresponding to the three-dimensional point cloud data in a splicing manner to obtain joint features; deeply fusing the high-dimensional representation features corresponding to the text data and the joint features through a cross-modal attention mechanism to obtain interactive features; performing self-attention processing on the interactive features and the high-dimensional representation features corresponding to the text data, and superimposing the features through residual connections to form fused features.

[0015] Secondly, the present application also provides a data enhancement device based on the combination of three-dimensional scene and language data, the device includes: a data acquisition module for acquiring 3D scene data and corresponding text annotation data; 3D scene data includes: scene point cloud data, RGB-D image and camera pose information; text annotation data includes: question-answer pairs and dense description information related to the location, attributes and relationships of scene objects; a preprocessing module for preprocessing the 3D scene data and text annotation data respectively to obtain preprocessed 3D-language joint data; preprocessing includes: normalization of scene point cloud data, RGB-D image and camera pose information; text annotation data includes: question-answer pairs and dense description information related to the location, attributes and relationships of scene objects; preprocessing module for preprocessing the 3D scene data and text annotation data respectively to obtain preprocessed 3D-language joint data; preprocessing includes: normalization of scene point cloud data, RGB-D image and camera pose information The consistency processing of B-D images, as well as data cleaning and grammatical structure analysis of text annotation data; the data enhancement module is used to perform multimodal data enhancement on the preprocessed 3D-language joint data to obtain enhanced joint data; multimodal data enhancement includes: multiple image expansion operations such as rotation, translation and scale transformation of RGB-D images, as well as data expansion operations of three strategies: synonym replacement, logical inversion, and template generation for question-answer pairs and dense description information; the filtering processing module is used to perform semantic quality filtering on the enhanced joint data to obtain the target 3D-language joint dataset.

[0016] In a third aspect, the present application also provides a data enhancement system based on the combination of three-dimensional scenes and language data, including a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, and the processor executing the computer-executable instructions to implement the method described in the first aspect.

[0017] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method described in the first aspect.

[0018] The data enhancement method, device and system based on the combination of three-dimensional scene and language data provided by the present application first obtains 3D scene data and corresponding text annotation data; wherein the 3D scene data includes: scene point cloud data, RGB-D image and camera pose information; the text annotation data includes: question-answer pairs and dense description information related to the location, attributes and relationship of scene objects; then, the scene point cloud data is normalized and the RGB-D image is processed for consistency to ensure the consistency of the coordinate range; the text annotation data is cleaned and the grammatical structure is analyzed to remove duplicate or grammatically incorrect samples; then, the RG-D image is processed to ensure the consistency of the coordinate range; then, the RGB-D image is processed to remove duplicate or grammatically incorrect samples; then, the RGB ... Multiple image augmentation operations such as rotation, translation, and scale transformation of B-D images, as well as data augmentation operations using three strategies: synonym replacement, logical inversion, and template generation for question-answer pairs and dense description information, are performed to improve the diversity of enhanced joint data. Finally, semantic quality filtering is performed on the enhanced joint data to ensure that it remains logically consistent with the data. Ultimately, a large-scale, high-quality 3D-language joint dataset is generated, which can enhance the data diversity and generalization ability while improving the data quality of 3D scene understanding and visual question answering tasks, providing strong support for applications such as 3D visual understanding and robotic task planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 A flowchart of a data enhancement method based on the combination of three-dimensional scene and language data provided in an embodiment of the present application;

[0021] Figure 2 A flowchart of constructing a 3D-language joint dataset provided in an embodiment of the present application;

[0022] Figure 3 A flowchart of a text data enhancement process provided in an embodiment of the present application;

[0023] Figure 4 A flowchart of a semantic quality filtering process provided in an embodiment of the present application;

[0024] Figure 5 A flowchart of a multimodal embedding encoding process provided in an embodiment of the present application;

[0025] Figure 6A structural block diagram of a data enhancement device based on the combination of three-dimensional scene and language data provided in an embodiment of the present application;

[0026] Figure 7 A structural diagram of a data enhancement system based on the combination of three-dimensional scenes and language data provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions of this application in conjunction with the embodiments. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0028] In existing technologies, question-answering and description of three-dimensional scenes usually rely on relatively scarce data resources, and traditional data collection methods require a large amount of annotation by professionals, which is both time-consuming and costly. To alleviate the problem of data scarcity, this embodiment proposes an overall solution based on multimodal data generation and enhancement. By combining ScanNet three-dimensional scene data, ScanQA question-answer annotation, and ScanRefer dense description, it achieves the diversification and large-scale expansion of data sources. At the same time, this embodiment uses advanced semantic quality control technology to filter and optimize the generated data to ensure high accuracy and semantic consistency of the data.

[0029] Based on this, the embodiments of the present application provide a data enhancement method, device and system based on the combination of three-dimensional scenes and language data, which can improve the data quality of 3D scene understanding and visual question answering tasks while enhancing the diversity and generalization ability of data, providing strong support for applications such as 3D visual understanding and robot task planning.

[0030] To facilitate understanding of this embodiment, a method for constructing a 3D-language joint dataset disclosed in an embodiment of the present application is first introduced in detail.

[0031] This application provides a data augmentation method based on the combination of 3D scene and language data. This method aims to construct a high-quality, large-scale joint 3D-language dataset through deep fusion and intelligent augmentation of multimodal data. This method not only expands the dataset size, improving its diversity and generalization capabilities, but also effectively eliminates low-quality samples, ensuring the semantic consistency and accuracy of the data. Figure 1 A flowchart of a data enhancement method based on the combination of three-dimensional scene and language data provided in an embodiment of the present application, the method specifically includes the following steps:

[0032] Step S102: Acquire 3D scene data and corresponding text annotation data; the 3D scene data includes scene point cloud data, RGB-D images, camera pose information, and corresponding semantic annotation information; the text annotation data includes question-answer pairs and dense description information related to the location, attributes, and relationships of scene objects;

[0033] First, we acquire 3D scene data and text annotation data from multiple standard datasets (such as ScanNet, ScanQA, and ScanRefer). 3D scene data includes indoor scene point clouds, RGB-D images, and camera pose information; text annotation data includes question-answer pairs and dense descriptions related to the location, attributes, and relationships of scene objects. By integrating these multi-source data, we achieve complementarity and fusion between different data sources, providing a sample foundation for subsequent data preprocessing and enhancement.

[0034] In specific implementation, scene point cloud data and its corresponding RGB-D image and camera pose information can be extracted from the indoor scene dataset, and the scene point cloud data can be stored in a standard format (such as .ply format) to facilitate subsequent data processing.

[0035] On the other hand, question-answer pairs containing the location, attributes, and relationship information of scene objects can be obtained from the question-answering dataset; dense description information can be extracted from the description dataset and combined with the question-answering data to form multi-source text annotation data.

[0036] Step S104: Preprocessing the 3D scene data and the text annotation data to obtain preprocessed 3D-language joint data. The preprocessing includes normalization of the scene point cloud data, consistency processing of the RGB-D image, and data cleaning and grammatical structure analysis of the text annotation data.

[0037] By normalizing the point cloud data of each scene, the range of each coordinate point is limited to a preset interval; by performing data cleaning on the text annotation data and eliminating duplicate, grammatically incorrect or semantically incomplete samples, the accuracy and completeness of the text data can be ensured.

[0038] Step S106: Multimodal data enhancement is performed on the pre-processed 3D-language joint data to obtain enhanced joint data. Multimodal data enhancement includes: multiple image augmentation operations such as rotation, translation, and scaling of RGB-D images, as well as data augmentation operations such as synonym replacement, logical inversion, and template generation for question-answer pairs and dense description information.

[0039] The specific contents of the three strategies mentioned above, namely synonym replacement, logic inversion, and template generation, are as follows:

[0040] (1) Synonym replacement: The pre-trained BERT model is used to replace the keywords in the question-answer pair with synonyms, and the replacement probability is set to a preset value to maintain semantic stability;

[0041] (2) Logical reversal: Generate the negation sentence in the question-answer pair through the predefined rule template. The generated reverse question must be verified for semantic consistency;

[0042] (3) Template generation: Based on the pre-trained T5 model, the description text is converted into structured question-answer pairs, and a variety of pre-designed question templates are used for generation.

[0043] Data augmentation through the above three strategies can improve sample diversity.

[0044] Step S108 : performing semantic quality filtering on the enhanced joint data to obtain a target 3D-language joint dataset.

[0045] In specific implementations, deep learning models can be used to calculate semantic similarity and semantic consistency, filter the enhanced joint data, and obtain a high-quality 3D-language joint dataset that meets preset threshold requirements. For example, the BERT model can be used to extract text features from the original and generated question-answer pairs and calculate cosine similarity to select question-answer pairs with similarity at least below a preset threshold. The RoBERTa model can also be used to perform natural language inference on the generated and original descriptions, filtering out logically inconsistent text samples based on the consistency score.

[0046] The purpose of the above-mentioned multimodal data enhancement and semantic quality filtering is to improve the diversity and generalization ability of the data while maintaining the original semantic information of the data, so as to support various application needs such as 3D scene question answering, semantic description, and robot task planning.

[0047] The data enhancement method based on the combination of three-dimensional scene and language data provided in the embodiments of this application integrates multiple data sources such as scene point cloud data, RGB-D images, camera pose information, and text annotations (including question-answer pairs and dense descriptions), and utilizes data preprocessing, multimodal data enhancement, and semantic quality filtering to automatically construct high-quality large-scale datasets. This method has broad application prospects in fields such as 3D scene question answering, dense description generation, robot navigation, and task planning. It provides solid data support for multimodal understanding of complex scenes, reduces data annotation costs and acquisition difficulties, and improves the model's ability to extract and robustly analyze semantic information in complex scenes.

[0048] The embodiment of the present application also provides another data enhancement method based on the combination of three-dimensional scenes and language data, which is implemented on the basis of the above embodiment; this embodiment focuses on describing the specific implementation process of preprocessing, data enhancement, and semantic quality filtering.

[0049] The above steps of preprocessing the 3D scene data and text annotation data respectively include:

[0050] (1) Normalize the scene point cloud data in the 3D scene data and process the resolution and color consistency of the RGB-D image;

[0051] Normalizing the scene point cloud data can ensure the uniformity of the coordinate range of the point cloud in each scene and eliminate data distribution differences caused by factors such as acquisition equipment and scene size. At the same time, preprocessing the RGB-D image can ensure the consistency of image resolution and color information.

[0052] (2) Clean the text annotation data to remove duplicates, grammatical errors, and incomplete samples; use natural language processing technology to segment the cleaned text annotation data, remove stop words, and analyze the grammatical structure to obtain text annotation data that meets the requirements of semantic consistency and accuracy.

[0053] In order to improve the diversity and coverage of the dataset, the embodiments of the present application introduce multiple data augmentation techniques to respectively expand the text annotation data and the 3D scene data. That is, the steps of performing multimodal data augmentation on the pre-processed 3D-language joint data to obtain enhanced joint data include:

[0054] (1) Performing multiple operations such as rotation, translation, and scale transformation on the RGB-D image in the preprocessed 3D scene data to achieve image expansion;

[0055] Performing slight rotations, translations, and scale transformations on 3D scene data to simulate possible angle and position changes in real environments can increase the model's robustness to different scene perspectives.

[0056] (2) Based on the pre-trained language model, three strategies, namely synonym replacement, logical inversion, and template generation, are used to expand the data of question-answer pairs and dense description information; the image-expanded 3D scene data and the text annotation data after data expansion are combined to obtain enhanced joint data.

[0057] In practical applications, pre-trained language models (such as BERT and T5) can be used to expand question-answer pairs and dense descriptions using strategies such as synonym replacement, logical inversion, and template generation. For example, synonym replacement can be used to transform "the sofa is on the left side of the living room" into "the sofa is on the west side of the living room," generating more variants while maintaining semantic consistency. Logical inversion and template generation methods can also be used to convert descriptive text into question-answer pairs, enriching the data's expression.

[0058] To ensure the high quality of the enhanced joint data, the embodiment of the present application uses a deep learning model to perform semantic similarity and consistency detection on the generated question-answer pairs and descriptions. That is, the above-mentioned step of performing semantic quality filtering on the enhanced joint data to obtain the target 3D-language joint data includes:

[0059] (1) Using the pre-trained BERT model to calculate the semantic similarity between the generated question-answer text and the original question-answer text, the text with a semantic similarity less than a first threshold is filtered out;

[0060] In specific implementation, the pre-trained BERT model can be used to extract text features and calculate the cosine similarity between the generated question-answer pairs and the original question-answer pairs; a similarity threshold can be set to filter out samples with higher semantic similarity to ensure that the generated question-answer pairs are highly consistent with the original data in semantics.

[0061] (2) Using the RoBERTa model for natural language inference, the consistency score between the generated dense description text and the original dense description text is calculated, and the description text with a consistency score less than the second threshold is filtered to obtain the target 3D-language joint data.

[0062] In specific implementation, the RoBERTa model can be used for natural language reasoning, and the generated description text and the original annotations can be evaluated for logical consistency, and samples with semantic contradictions or logical errors can be eliminated to ensure that the final dataset has high accuracy at the semantic level.

[0063] After the aforementioned preprocessing, data augmentation, and semantic filtering, the selected high-quality data is integrated to form the final 3D-language joint dataset. This dataset includes tens of thousands of question-answer pairs and dense descriptions, covering multiple 3D scenes. With rich semantic information and diverse expression forms, it can effectively support applications such as 3D scene question answering, dense description generation, and robotic task planning.

[0064] Compared with traditional datasets, the data enhancement method based on the combination of three-dimensional scenes and language data provided in this embodiment generates a dataset with wide coverage, high data quality and good semantic consistency through intelligent data preprocessing, semantic-preserving data enhancement strategies and strict quality filtering methods. It not only significantly expands the data scale, but also achieves significant improvements in data diversity, semantic consistency and quality control. It provides strong data support for complex applications such as 3D visual question answering, dense description generation and robot task planning, and greatly reduces the cost and cycle of data collection and labeling.

[0065] In summary, the data enhancement method based on the combination of 3D scene and language data provided in the embodiments of the present application has the following advantages:

[0066] 1. Large data scale and wide coverage: Through multi-source data integration and intelligent data enhancement methods, it is possible to generate high-quality datasets containing tens of thousands of question-answer pairs and dense descriptions, significantly improving the scale and scenario coverage of the original dataset.

[0067] 2. High data quality and good semantic consistency: Advanced natural language processing technology and deep learning models are used to perform semantic filtering on the generated data, effectively eliminating low-quality and inconsistent samples, and ensuring the accuracy and consistency of the data in semantic expression.

[0068] 3. Full fusion of multimodal information: The embodiment of the present application processes 3D scene data and text data simultaneously. Through preprocessing, data enhancement and filtering strategies, it achieves full fusion of multimodal information, providing strong data support for visual question answering and dense description generation in complex scenes.

[0069] 4. Reduce data collection and annotation costs: The use of automated data enhancement and semantic filtering technology significantly reduces the workload and cost of manual collection and annotation, improves the efficiency and accuracy of data construction, and has high practical value and economic benefits.

[0070] The method provided in the embodiments of this application not only constructs a high-quality 3D-language joint dataset by preprocessing, enhancing, filtering, and integrating raw data; in a preferred embodiment, it also achieves cross-modal upgrade from sparse data to dense data, and utilizes advanced deep learning models to efficiently encode text, visual cues, and 3D point cloud data, ultimately generating high-quality natural language output through decoding using a large-scale pre-trained language model. Specifically, the above method may further include the following steps:

[0071] (1) Encode the text data, visual cue data, and 3D point cloud data in the target 3D-language joint dataset separately to obtain the encoding features corresponding to the three types of data;

[0072] In specific implementation, text data can be encoded based on the Transformer architecture to generate high-dimensional representation features; visual cue data can be flattened and linearly mapped through a multi-layer perceptron to obtain spatial prior embedding features of the visual cue data; and the Vote2Cap-DETR++ network can be used to extract features from three-dimensional point cloud data to obtain deep geometric features of the three-dimensional point cloud data.

[0073] (2) Fusing the coding features corresponding to the three types of data to obtain fused features;

[0074] In specific implementation, a splicing method can be used to preliminarily fuse the spatial prior embedding features corresponding to the visual cue data and the deep geometric features corresponding to the three-dimensional point cloud data to obtain joint features; the high-dimensional representation features corresponding to the text data and the joint features are deeply fused through the cross-modal attention mechanism to obtain interactive features; self-attention processing is performed on the interactive features and the high-dimensional representation features corresponding to the text data, and the features are superimposed through residual connections to form fused features.

[0075] (3) The fused features are input into the pre-trained large-scale language model for natural language decoding to obtain the target natural language.

[0076] Based on the above method embodiments, the embodiments of the present application also provide a data enhancement system based on the combination of three-dimensional scenes and language data. That is, this embodiment proposes an end-to-end system based on multimodal data generation, semantic quality filtering, and large-scale pre-trained language model decoding for question-answering and dense description tasks in three-dimensional scenes. This system can not only automatically construct a large-scale, high-quality 3D-language joint data set, but also achieve innovative technological breakthroughs in key links such as multimodal information fusion, data enhancement, and text generation, thereby demonstrating excellent semantic reasoning and spatial understanding capabilities when processing complex three-dimensional scenes. By preprocessing, enhancing, filtering, and integrating the original data, this embodiment achieves a cross-modal upgrade from sparse data to dense data, and uses advanced deep learning models to efficiently encode text, visual cues, and three-dimensional point cloud data, and ultimately generates high-quality natural language output through large-scale pre-trained language model decoding.

[0077] like Figure 2 As shown in the figure, the overall system process begins with raw data acquisition, including 3D point cloud data from ScanNet, question-and-answer annotations from ScanQA, and dense descriptions from ScanRefer. After preprocessing, the raw data enters the multimodal data augmentation module. This module uses a pretrained language model to perform synonym replacement, logical inversion, and template generation on the text data, thereby expanding the data diversity while maintaining semantic consistency. After the augmented data undergoes preliminary synthesis, it is then subjected to semantic quality filtering using the pretrained model to ensure the high accuracy and consistency of the output dataset, ultimately completing the integrated output of the dataset. Figure 2 The overall process of dataset construction is intuitively demonstrated, and each step from raw data acquisition, preprocessing, multimodal data enhancement, semantic quality filtering to final dataset integration output, as well as the data flow and processing relationships between them, is described in detail, fully reflecting the systematic and automated features of this embodiment in data construction.

[0078] In the text data enhancement link, such as Figure 3As shown, the embodiment of the present application uses a pre-trained language model to expand the original text data from multiple angles and in all directions. Specifically, for the questions and answers in the original question-answer pair, the system first uses synonym replacement technology to perform semantic expansion, and generates multiple versions of question-answer text by introducing different but semantically consistent vocabulary. At the same time, in response to the logical relationship in the original text, this embodiment uses logical inversion technology to generate question-answer pairs with reverse semantics, and then combines it with a predefined template generation method to use the T5 model to convert dense descriptions into question-answer pairs. The conversion process can be described by the following formula:

[0079]

[0080] Among them, C represents the original description text, P template Indicates a preset text template, symbol Represents a text concatenation operation. This series of text data enhancement strategies not only effectively expands the data scale, but also significantly improves data diversity while maintaining text semantic consistency. Figure 2 The specific process of the text data enhancement module is demonstrated in detail. The calling relationship and data flow between each processing module are fully reflected, providing sufficient candidate data for subsequent semantic quality filtering.

[0081] After data enhancement is completed, this system further uses the semantic quality filtering module to strictly screen the generated text. Figure 4 The overall process of the semantic quality filtering module is demonstrated. The core idea is to use the pre-trained BERT model and RoBERTa model to calculate the semantic similarity and consistency between the generated text and the original text, and filter the data according to the preset threshold. Specifically, for the generation process of question-answer pairs, the system uses the BERT model to calculate the cosine similarity between the original question and the generated question. The calculation formula is:

[0082] SQ(Q orig ,Q gen )=cos(f BERT (Q orig ),f BERT (Q gen ));

[0083] Among them, f BERTrepresents the semantic vector obtained based on the BERT model, and cos(·,·) represents the cosine similarity function. By setting a task-specific threshold (for example, the threshold is set to 0.82 for question-answering tasks), the system can eliminate samples that deviate significantly from the original question and answer in semantics. In addition, to further ensure the semantic consistency of the generated data, the system uses the RoBERTa model for natural language inference (NLI) to calculate the consistency score between the description texts. The calculation formula is:

[0084]

[0085] Among them, C orig Represents the original description text, represents the i-th generated description, and n is the number of generated samples. After this series of rigorous semantic filtering steps, the system ultimately selects high-quality, highly consistent text data, which is then combined with other modality data to form a large-scale 3D-language joint dataset. Figure 4 It illustrates in detail the process of using the BERT and RoBERTa models to perform semantic similarity and consistency detection on enhanced text, indicates the data information transmission and threshold screening mechanism between each link, and intuitively demonstrates how to accurately eliminate low-quality samples from massive data to ensure that the overall quality of the dataset meets the expected requirements.

[0086] In terms of multimodal data encoding, this embodiment adopts a unified embedded coding architecture to map text, visual cues and 3D point cloud data into the same high-dimensional feature space. Specifically, for text data, this system encodes the input text sequence based on the Transformer architecture. After embedding lookup and position encoding, a high-dimensional representation E is generated. t , and its calculation process can be described as:

[0087] E t =Transformer enc (Embed(T)+P t );

[0088] Among them, Embed(T) represents the embedding operation of converting text into word vectors, P t Represents position encoding; for visual cue data, the system first uses the target detection algorithm or user annotation to obtain 3D bounding box information It is then flattened and linearly mapped through a multi-layer perceptron (MLP) to obtain the spatial prior embedding E of the visual cue v :E v =MLP(Flatten(V)) and for 3D point cloud data This system uses the Vote2Cap-DETR++ network for feature extraction and generates deep geometric features E of point cloud data. s :

[0089]

[0090] Figure 4 A schematic diagram of the multimodal embedding coding structure is intuitively presented, which provides detailed explanations of the specific processes of text encoding, visual cue encoding, and 3D point cloud encoding, respectively. It shows how each module uses the Transformer, MLP, and Vote2Cap-DETR++ network to extract and map high-dimensional features, thereby constructing a unified multimodal feature space.

[0091] When realizing multimodal fusion, the system first performs a preliminary fusion of visual cues and 3D point cloud data, and then uses a splicing method to combine the encoding results E of the two. v With E s Merge to get the joint feature E vs :

[0092] E vs =Concat(E v ,E s );

[0093] Then, the system encodes the text E through a cross-modal attention mechanism t With the joint feature E vs Perform deep fusion and calculate the interaction feature E f1 :

[0094]

[0095] Among them, W q 、W k and W v are the projection matrices of query, key, and value respectively, and d is the embedding dimension. To ensure the coherence of text information and the complete expression of context, the system further performs self-attention processing on the interaction features and the original text encoding, and superimposes them through residual connections to form the final fusion representation E f :

[0096] E f =LayerNorm(TransformerLayer(E f1 ||E t )+E t );

[0097] Among them, “||” represents the splicing operation of features. Through the above multimodal fusion and encoding process, the system achieves efficient fusion of text, visual cues and 3D point cloud information, thereby providing comprehensive and accurate semantic and spatial information support for subsequent natural language generation. After completing the multimodal feature fusion, the system will fused feature E f The input is fed into a pre-trained large-scale language model for natural language decoding. Specifically, the OPT-1.3B model is used as a decoder, and dynamic prefix projection technology is used to map the fused features to the prefix information required by the language model, thereby guiding the generation process. The decoding process can be expressed as:

[0098] h t =LLM(p 1:t ;LinearProj(E f ));

[0099] Among them, LinearProj(E f ) means that the fusion feature is converted into prefix information suitable for language model input through linear mapping, p 1:t represents the generated token sequence, and the probability distribution of generating the next token is:

[0100] p(w t+1 )=softmax(h t W vocab );

[0101] Through the above-mentioned generation mechanism, the system can fully utilize multimodal information while generating natural, fluent and semantically accurate question and answer or description texts, thereby significantly improving the model's performance in semantic understanding and reasoning of three-dimensional scenes.

[0102] This embodiment adopts an end-to-end joint training strategy, which integrates multimodal data generation, semantic quality filtering, multimodal encoding and large-scale language model decoding. During the training process, the system minimizes the question-answer generation loss L QA and description generation loss L cap , the total loss function is formed, which is expressed as:

[0103] L total =λ QA ·L QA +λ cap ·L cap ;

[0104] Among them, λ QA and λ capare the weight coefficients for each task, and the optimal value is obtained through cross-validation. The question-answering generation loss is mainly based on the cross-entropy error between the generated question and the real question, while the description generation loss combines indicators such as ROUGE and CIDEr to comprehensively evaluate the degree of match between the generated text and the real text. To ensure stability in large-scale data training, the system uses the AdamW optimizer and combines it with the cosine annealing learning rate scheduling strategy to effectively avoid the model from falling into local optimality and overfitting problems. Table 1 below shows the question-answering task-comparison method and detailed information, and Table 2 shows the dense subtitle task-comparison method and detailed information;

[0105] Table 1

[0106]

[0107] Table 2

[0108]

[0109]

[0110] The experimental results shown in Tables 1 and 2 demonstrate that the system proposed in this embodiment achieves significant performance improvements over existing methods in multiple 3D scene question answering and dense description tasks. For example, on the ScanQA dataset, the CIDEr metric improved by approximately 2.15%, while on the ScanRefer dataset, the CIDEr@0.5 metric improved by approximately 1.84%. Furthermore, ablation experiments verified the positive contributions of the data augmentation strategy, semantic filtering mechanism, and multimodal fusion module to the overall system performance, further demonstrating the technical advantages of this embodiment in large-scale data construction and end-to-end joint training.

[0111] This embodiment performs comprehensive preprocessing and multimodal data enhancement on raw 3D point cloud data, question-answer annotation data, and dense description data, uses advanced pre-trained language models to generate and semantically filter text, and combines a cross-modal attention mechanism to achieve multimodal feature fusion, thereby generating high-quality natural language output under end-to-end joint training. Figure 2 、 Figure 3 、 Figure 4 and Figure 5The core modules and their internal working principles of this embodiment are intuitively demonstrated from four perspectives: overall process, text data enhancement, semantic quality filtering, and multimodal embedding coding. This demonstrates the entire process from data generation and quality control to multimodal feature fusion and natural language decoding. This system not only achieves a dual improvement in data scale and quality through a variety of advanced technologies in the data construction and preprocessing stages, but also fully leverages the advantages of deep learning models in the multimodal information fusion and language generation stages, achieving efficient expression and precise generation of semantic information in complex three-dimensional scenes. This provides an efficient, accurate, and highly practical solution for question-answering and dense description tasks in complex 3D scenes.

[0112] The method and system provided in this embodiment have significant technical advantages. On the one hand, through multimodal data generation and enhancement, it effectively overcomes the problems of scarce data volume and insufficient semantic information in traditional 3D scene description. On the other hand, by utilizing pre-trained language models and multimodal fusion technology, the system can capture the deep correlation between text, visual cues, and 3D geometric information in a high-dimensional feature space, significantly improving the accuracy and naturalness of question answering and description generation. Through the careful design of each module of the system and end-to-end joint optimization, this embodiment not only constructs a unified 3D-language data construction and generation framework in theory, but also realizes the automated construction and efficient utilization of large-scale high-quality data in practice, providing solid technical support for the further development of 3D vision and language tasks. Therefore, this embodiment proposes a new 3D scene question answering and dense description method through the systematic integration of data generation, semantic filtering, multimodal encoding, and natural language generation. While ensuring the high semantic consistency and accuracy of the generated text, this method achieves efficient understanding and expression of complex spatial information through a large-scale pre-trained model and cross-modal attention mechanism, thus significantly outperforming existing technologies in multiple evaluation indicators. This technology not only has high academic research value, but also shows great potential in practical applications, providing a new solution for the future development of intelligent three-dimensional scene analysis and human-computer interaction technology.

[0113] Based on the above method embodiment, the present application embodiment also provides a data enhancement device based on the combination of three-dimensional scene and language data, see Figure 6As shown, the device includes: a data acquisition module 62, which is used to acquire 3D scene data and corresponding text annotation data; the 3D scene data includes: scene point cloud data, RGB-D images and camera pose information; the text annotation data includes: question-answer pairs and dense description information related to the location, attributes and relationships of scene objects; a preprocessing module 64, which is used to preprocess the 3D scene data and text annotation data respectively to obtain preprocessed 3D-language joint data; the preprocessing includes: normalization processing of scene point cloud data, consistency processing of RGB-D images, and data cleaning processing and grammatical structure analysis processing of text annotation data; a data enhancement module 66, which is used to perform multimodal data enhancement on the preprocessed 3D-language joint data respectively to obtain enhanced joint data; the multimodal data enhancement includes: multiple image expansion operations such as rotation, translation and scale transformation of RGB-D images, and data expansion operations of three strategies such as synonym replacement, logical inversion and template generation for question-answer pairs and dense description information; a filtering processing module 68, which is used to perform semantic quality filtering processing on the enhanced joint data to obtain the target 3D-language joint data set.

[0114] Furthermore, the above-mentioned preprocessing module 64 is used to normalize the scene point cloud data in the 3D scene data, and perform resolution and color consistency processing on the RGB-D image; perform data cleaning on the text annotation data to eliminate duplicates, grammatical errors and samples with incomplete information; use natural language processing technology to segment the text annotation data after data cleaning, remove stop words and analyze the grammatical structure to obtain text annotation data with semantic consistency and accuracy that meet the requirements.

[0115] Furthermore, the data enhancement module 66 is used to perform various operations such as rotation, translation, and scale transformation on the RGB-D image in the preprocessed 3D scene data to achieve image expansion; based on the pre-trained language model, three strategies of synonym replacement, logical inversion, and template generation are used to expand the question-answer pairs and dense description information; the image-expanded 3D scene data and the text annotation data after data expansion are combined to obtain enhanced joint data.

[0116] Furthermore, the above-mentioned filtering processing module 68 is used to use the pre-trained BERT model to calculate the semantic similarity between the generated question and answer text and the original question and answer text, and filter the text with a semantic similarity less than a first threshold; use the RoBERTa model to perform natural language inference, calculate the consistency score between the generated dense description text and the original dense description text, and filter the description text with a consistency score less than a second threshold to obtain the target 3D-language joint data.

[0117] Furthermore, the above-mentioned device also includes: a language generation module, which is used to encode the text data, visual prompt data and three-dimensional point cloud data in the target 3D-language joint dataset respectively to obtain the encoding features corresponding to the three types of data; fuse the encoding features corresponding to the three types of data to obtain fused features; input the fused features into a pre-trained large-scale language model for natural language decoding to obtain the target natural language.

[0118] Furthermore, the above-mentioned language generation module is used to encode text data based on the Transformer architecture to generate high-dimensional representation features; flatten and linearly map the visual cue data through a multi-layer perceptron to obtain the spatial prior embedding features of the visual cue data; and use the Vote2Cap-DETR++ network to extract features from three-dimensional point cloud data to obtain the deep geometric features of the three-dimensional point cloud data.

[0119] Furthermore, the above-mentioned language generation module is used to use a splicing method to preliminarily fuse the spatial prior embedding features corresponding to the visual cue data and the deep geometric features corresponding to the three-dimensional point cloud data to obtain joint features; deeply fuse the high-dimensional representation features corresponding to the text data and the joint features through a cross-modal attention mechanism to obtain interactive features; perform self-attention processing on the interactive features and the high-dimensional representation features corresponding to the text data, and superimpose features through residual connections to form fused features.

[0120] The device provided in the embodiment of the present application has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the embodiment of the device, reference can be made to the corresponding content in the aforementioned method embodiment.

[0121] Based on the above method embodiments, the present application also provides a data enhancement system based on the combination of three-dimensional scene and language data. This system is a three-dimensional scene question-answering and dense description system based on multimodal data generation, fusion, and decoding of large-scale pre-trained language models (hereinafter referred to as the "3D-MoRe system"). The system aims to automatically generate large amounts of high-quality three-dimensional language data, and utilize advanced data enhancement, semantic filtering, and cross-modal interaction technologies to effectively improve the understanding and reasoning ability of semantic information in complex three-dimensional scenes, thereby achieving results close to manual annotation in 3D question-answering and dense description tasks.

[0122] like Figure 7 As shown, it is a structural diagram of the system, wherein the system includes a processor 71 and a memory 70, the memory 70 stores computer executable instructions that can be executed by the processor 71, and the processor 71 executes the computer executable instructions to implement the above method.

[0123] exist Figure 7In the illustrated embodiment, the system further includes a bus 72 and a communication interface 73 , wherein the processor 71 , the communication interface 73 and the memory 70 are connected via the bus 72 .

[0124] Among them, the memory 70 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 73 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 72 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 72 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0125] The processor 71 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 71 or by software instructions. The above processor 71 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor 71 reads the information in the memory and completes the steps of the method of the above embodiment in combination with its hardware.

[0126] An embodiment of the present application also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the above-mentioned method. The specific implementation can be found in the above-mentioned method embodiment, which will not be repeated here.

[0127] The computer program products of the methods, devices, and systems provided in the embodiments of the present application include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the previous method embodiments. For specific implementation, please refer to the method embodiments and will not be repeated here.

[0128] Unless otherwise specifically stated, the relative steps, numerical expressions and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present application.

[0129] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0130] In the description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate the description of this application and simplify the description. They do not indicate or imply that the devices or components referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0131] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in the present application, or perform equivalent replacements for some of the technical features thereof. These modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims.

Claims

1. A data enhancement method based on the combination of three-dimensional scene and language data, characterized in that: The method comprises: Acquire 3D scene data and corresponding text annotation data; the 3D scene data includes scene point cloud data, RGB-D images, and camera pose information; the text annotation data includes question-answer pairs and dense description information related to the location, attributes, and relationships of scene objects; Preprocessing the 3D scene data and the text annotation data separately to obtain preprocessed 3D-language joint data; the preprocessing includes: normalization of scene point cloud data, consistency processing of the RGB-D image, and data cleaning and grammatical structure analysis of the text annotation data; Performing multimodal data augmentation on the preprocessed 3D-language joint data to obtain enhanced joint data; the multimodal data augmentation includes: multiple image augmentation operations such as rotation, translation, and scaling of the RGB-D image; and data augmentation operations using three strategies: synonym replacement, logical inversion, and template generation for the question-answer pairs and dense description information; Semantic quality filtering is performed on the enhanced joint data to obtain a target 3D-language joint dataset.

2. The method according to claim 1, characterized in that The steps of preprocessing the 3D scene data and the text annotation data respectively include: Normalizing the scene point cloud data in the 3D scene data, and performing resolution and color consistency processing on the RGB-D image; The text annotation data is cleaned to eliminate duplicates, grammatical errors, and incomplete samples; natural language processing technology is used to segment the text annotation data after data cleaning, remove stop words, and analyze the grammatical structure to obtain text annotation data with semantic consistency and accuracy that meet the requirements.

3. The method according to claim 1, characterized in that The step of performing multimodal data enhancement on the pre-processed 3D-language joint data to obtain enhanced joint data includes: Performing rotation, translation, and scale transformation operations on the RGB-D image in the pre-processed 3D scene data to achieve image expansion; Based on the pre-trained language model, three strategies, namely synonym replacement, logical inversion, and template generation, are used to perform data expansion on the question-answer pairs and dense description information; The image-augmented 3D scene data and the data-augmented text annotation data are combined to obtain enhanced joint data.

4. The method according to claim 1, wherein The step of performing semantic quality filtering on the enhanced joint data to obtain target 3D-language joint data comprises: Use the pre-trained BERT model to calculate the semantic similarity between the generated question and answer text and the original question and answer text, and filter out texts with semantic similarity less than the first threshold; The RoBERTa model is used for natural language inference to calculate the consistency score between the generated dense description text and the original dense description text. The description texts with a consistency score less than the second threshold are filtered out to obtain the target 3D-language joint data.

5. The method according to claim 1, wherein The method further comprises: Encoding the text data, visual cue data, and three-dimensional point cloud data in the target 3D-language joint dataset respectively to obtain encoding features corresponding to the three types of data; The coding features corresponding to the three types of data are fused to obtain the fusion features; The fused features are input into a pre-trained large-scale language model for natural language decoding to obtain the target natural language.

6. The method according to claim 5, characterized in that The steps of respectively encoding the text data, visual cue data, and 3D point cloud data in the target 3D-language joint data to obtain encoding features corresponding to the three types of data include: Encode text data based on the Transformer architecture to generate high-dimensional representation features; The visual cue data is flattened and linearly mapped through a multi-layer perceptron to obtain the spatial prior embedding features of the visual cue data; The Vote2Cap-DETR++ network is used to extract features from 3D point cloud data to obtain the deep geometric features of the 3D point cloud data.

7. The method according to claim 6, characterized in that The steps of fusing the coding features corresponding to the three types of data to obtain fused features include: The spatial prior embedding features corresponding to the visual cue data and the deep geometric features corresponding to the 3D point cloud data are preliminarily fused in a splicing manner to obtain the joint features. Deeply fusing the high-dimensional representation features corresponding to the text data with the joint features through a cross-modal attention mechanism to obtain interactive features; Self-attention processing is performed on the interaction features and the high-dimensional representation features corresponding to the text data, and features are superimposed through residual connections to form fusion features.

8. A data enhancement device based on the combination of three-dimensional scene and language data, characterized in that: The device comprises: A data acquisition module is configured to acquire 3D scene data and corresponding text annotation data; the 3D scene data includes scene point cloud data, RGB-D images, and camera pose information; the text annotation data includes question-answer pairs and dense description information related to the location, attributes, and relationships of scene objects; a preprocessing module for preprocessing the 3D scene data and the text annotation data respectively to obtain preprocessed 3D-language joint data; the preprocessing includes: normalization of the scene point cloud data, consistency processing of the RGB-D image, and data cleaning and grammatical structure analysis of the text annotation data; a data augmentation module for performing multimodal data augmentation on the preprocessed 3D-language joint data to obtain enhanced joint data; the multimodal data augmentation includes: multiple image augmentation operations such as rotation, translation, and scaling of the RGB-D image; and data augmentation operations using three strategies: synonym replacement, logical inversion, and template generation for the question-answer pairs and dense description information; The filtering processing module is used to perform semantic quality filtering on the enhanced joint data to obtain a target 3D-language joint data set.

9. A data enhancement system based on the combination of three-dimensional scenes and language data, characterized in that: The system includes a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text content auditing method and system based on sensitive word

    CN106445998A

  • Deep web resource detection system based on long short-term memory neural network

    CN108874943A

  • Dialogue generation and corpus expansion method and device, computer equipment and storage medium

    CN111061847A

  • Characteristic fusion and clustering-based turning singing identification method and device, and storage medium

    CN115472181A

  • Visual knowledge reasoning question and answer method for multi-source heterogeneous knowledge joint enhancement

    CN117010500A