Three-dimensional vision-text positioning method and system based on multi-view and multi-text

By extracting multi-view features in three-dimensional vision and text modes and performing multi-modal fusion, the scoring mechanism and attention module of perspective guidance are used to solve the problem of perspective neglect in the existing methods, and a more accurate and robust three-dimensional vision-text positioning is achieved.

CN117197439BActive Publication Date: 2025-08-12SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311215321.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-20
Publication Date
2025-08-12
Estimated Expiration
2043-09-20

AI Technical Summary

Technical Problem

The existing three-dimensional vision-text positioning methods only focus on three-dimensional modes, ignoring the importance of perspective cues and multi-view angles in text modes, resulting in the impact of model performance.

Method used

By simultaneously extracting multi-view features from three-dimensional vision and text modes, performing multi-modal fusion, and using the scoring mechanism and attention module guided by perspective, the scene-independent knowledge is memorized to achieve three-dimensional vision-text positioning.

Benefits of technology

It improves the accuracy and robustness of three-dimensional vision-text positioning, and enhances the ability to learn and extract perspective knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117197439B_ABST
    Figure CN117197439B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer vision technology, and in particular to a three-dimensional vision-text positioning method and system based on multi-perspective and multi-text, the method comprising: extracting three-dimensional scene features under multiple perspectives from an input three-dimensional scene to obtain multi-perspective three-dimensional features; expanding the perspective and expression of the text to obtain text with multi-perspective information, and performing feature extraction to obtain multi-perspective text features; performing multi-modal fusion of the multi-perspective three-dimensional features and the multi-perspective text features to obtain multi-perspective fusion features; scoring the multiple perspectives of the multi-perspective fusion features based on a perspective-guided scoring mechanism, using multi-perspective representative features to memorize scene-independent knowledge, and performing three-dimensional vision-text positioning. The present application simultaneously performs multi-perspective processing and understanding of three-dimensional vision and text, and uses multi-perspective representative features to memorize scene-independent knowledge, so as to achieve more accurate and robust three-dimensional vision-text positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer vision technology, and in particular to a three-dimensional vision-text positioning method and system based on multiple perspectives and multiple texts. Background Art

[0002] The 3D visual grounding task involves aligning a given natural language description with a target object in a 3D scene to find the target object described in the text. Mainstream approaches model the scene based on language and 3D point cloud data, using various multimodal learning methods to align the two modalities for effective localization.

[0003] Existing 3D vision-text localization methods aim to address the perspective ambiguity between text and 3D scenes. Specifically, they focus on the correspondence between the perspective of the text description and the 3D scene angle. They use multi-perspective approaches to improve localization accuracy by enhancing the model's robustness to perspective, starting from the 3D modality. However, these methods focus solely on the 3D modality, ignoring the perspective cues embedded in the text modality and neglecting to measure the importance of multiple perspectives, which can impact model performance. Summary of the Invention

[0004] The embodiments of the present application provide a three-dimensional vision-text positioning method and system based on multiple perspectives and multiple texts. The method simultaneously performs multi-perspective processing and understanding of three-dimensional vision and text, and uses multi-perspective representative features to memorize scene-independent knowledge to achieve more accurate and robust three-dimensional vision-text positioning.

[0005] To solve the above technical problems, in the first aspect, an embodiment of the present application provides a three-dimensional vision-text positioning method based on multi-perspective and multi-text, which includes the following steps: first, extracting three-dimensional scene features under multiple perspectives from the input three-dimensional scene to obtain multi-perspective three-dimensional features; next, expanding the perspective and expression of the text to obtain text with multi-perspective information, and performing feature extraction on the text with multi-perspective information to obtain multi-perspective text features; then performing multi-modal fusion of the multi-perspective three-dimensional features and the multi-perspective text features to obtain multi-perspective fusion features; finally, based on a perspective-guided scoring mechanism, scoring multiple perspectives of the multi-perspective fusion features, using multi-perspective representative features to memorize scene-independent knowledge, and performing three-dimensional vision-text positioning.

[0006] In some exemplary embodiments, a scoring mechanism based on perspective guidance is used to score multi-perspective fusion features, and multi-perspective representative features are used to memorize scene-independent knowledge, including the following steps: first, the cosine similarity between three-dimensional features and perspective representative features under multiple perspectives is calculated; then, based on the cosine similarity, the multiple perspectives of the multi-perspective fusion features are scored to obtain scores for the multiple perspectives; finally, the scores are used as weights for final prediction.

[0007] In some exemplary embodiments, after multi-view three-dimensional features and multi-view text features are multimodally fused, and before multiple viewpoints of the multi-view fusion features are scored, it also includes: a view-guided attention module that extracts view information from view representative features; and a view-guided attention module that is used to enhance the learning and extraction of view knowledge by the text modality.

[0008] In some exemplary embodiments, three-dimensional scene features under multiple perspectives are extracted from an input three-dimensional scene to obtain multi-perspective three-dimensional features, including the following steps: inputting the point cloud data of each object in the three-dimensional scene into a point cloud encoder separately to extract the three-dimensional features of each object; performing three-dimensional rotation on the center point of each object to obtain the position information of all objects at each perspective; and adding the position information directly to the object features to obtain the three-dimensional scene features of the object corresponding to each perspective.

[0009] In some exemplary embodiments, the input three-dimensional scene includes point cloud data of each object and the three-dimensional coordinates of the center point of the object.

[0010] In some exemplary embodiments, a pre-trained large language model is used to expand the perspective and expression of the text; and a text feature extractor is used to extract features from the text with multi-perspective information to obtain multi-perspective text features.

[0011] In some exemplary embodiments, before extracting three-dimensional scene features from multiple perspectives from an input three-dimensional scene, it also includes: obtaining a data set and dividing the data set into a training data set and a test data set; the data set is a data set of object positioning and language expression in the three-dimensional scene.

[0012] In some exemplary embodiments, after scoring the multiple perspectives of the multi-perspective fusion features, using the multi-perspective representative features to memorize scene-independent knowledge, and performing three-dimensional vision-text positioning, it also includes: verifying the three-dimensional vision-text positioning results, and applying the three-dimensional vision-text positioning method.

[0013] In a second aspect, an embodiment of the present application also provides a three-dimensional vision-text positioning system based on multi-perspective and multi-text, comprising a connected data processing module and a model training module; wherein the data processing module is used to obtain three-dimensional scene data; the model training module includes a multi-perspective three-dimensional feature acquisition unit, a multi-perspective text feature acquisition unit, a multimodal fusion unit and a perspective-guided scoring unit connected in sequence; wherein; the multi-perspective three-dimensional feature acquisition unit is used to extract three-dimensional scene features under multiple perspectives from the input three-dimensional scene to obtain multi-perspective three-dimensional features; the multi-perspective text feature acquisition unit is used to expand the perspective and expression of the text to obtain a text with multi-perspective information, and perform feature extraction on the text with multi-perspective information to obtain multi-perspective text features; the multimodal fusion unit performs multimodal fusion on the multi-perspective three-dimensional features and the multi-perspective text features to obtain multi-perspective fusion features; the perspective-guided scoring unit is used to score multiple perspectives of the multi-perspective fusion features according to the perspective-guided scoring mechanism, and use the multi-perspective representative features to memorize scene-independent knowledge to perform three-dimensional vision-text positioning.

[0014] In some exemplary embodiments, the above-mentioned three-dimensional vision-text positioning system based on multiple perspectives and multiple texts further includes a verification and application module connected to the model training module; the verification and application module is used to verify the three-dimensional vision-text positioning results and apply the three-dimensional vision-text positioning method; the data processing module includes a data set module and a data set division module; the data set module is used to obtain a data set; the data set is a data set of object positioning and language expression in a three-dimensional scene; the data set division module is used to divide the data set into a training data set and a test data set.

[0015] The technical solution provided by the embodiments of the present application has at least the following advantages:

[0016] The embodiment of the present application provides a three-dimensional vision-text positioning method and system based on multiple perspectives and multiple texts, the method comprising the following steps: first, extracting three-dimensional scene features under multiple perspectives from an input three-dimensional scene to obtain multi-perspective three-dimensional features; next, expanding the perspective and expression of the text to obtain a text with multi-perspective information, and performing feature extraction on the text with multi-perspective information to obtain multi-perspective text features; then performing multi-modal fusion on the multi-perspective three-dimensional features and the multi-perspective text features to obtain multi-perspective fusion features; finally, based on a perspective-guided scoring mechanism, scoring multiple perspectives of the multi-perspective fusion features, using multi-perspective representative features to memorize scene-independent knowledge, and performing three-dimensional vision-text positioning.

[0017] This application proposes a 3D vision-text localization method based on multiple perspectives and multiple texts. It simultaneously processes and understands the two modalities of 3D vision and text from multiple perspectives, and uses multi-perspective representative features to memorize scene-independent knowledge, which can achieve more accurate and robust 3D vision-text localization. At the same time, the method of this application also designs a multimodal fusion converter with inter-perspective interaction and an attention module for perspective guidance of the text modality, which can promote the framework's fusion of multimodal knowledge and enhance the learning and extraction of perspective knowledge by the text modality. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] One or more embodiments are exemplarily described by the pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Unless otherwise stated, the pictures in the drawings do not constitute proportional limitations.

[0019] Figure 1 A flowchart of a multi-view and multi-text 3D vision-text localization method provided in one embodiment of the present application;

[0020] Figure 2 A flowchart of the architecture of a multi-view and multi-text 3D vision-text localization method provided in one embodiment of the present application;

[0021] Figure 3 A schematic diagram of a training framework for a multi-view and multi-text 3D vision-text localization method provided in one embodiment of the present application;

[0022] Figure 4 A schematic diagram of a scoring mechanism framework for perspective guidance provided in one embodiment of the present application;

[0023] Figure 5 An architecture diagram of a multi-view and multi-text 3D vision-text localization system provided in one embodiment of the present application;

[0024] Figure 6 A comparison chart of experimental results between the multi-view and multi-text 3D vision-text localization method provided in one embodiment of the present application and other models;

[0025] Figure 7 A comparison chart of experimental results of the multi-view and multi-text 3D vision-text localization method provided in one embodiment of the present application and the existing best model. DETAILED DESCRIPTION

[0026] As can be seen from the background technology, existing methods only focus on the three-dimensional modality, ignore the perspective cues embedded in the text modality, and ignore the measurement of the importance of multiple perspectives, which will affect the performance of the model.

[0027] In the field of computer vision, the goal of 3D text localization is to achieve 3D scene understanding and reasoning by associating natural language descriptions with objects in the 3D environment. This technology can be widely used in autonomous driving, robot navigation, augmented reality, virtual reality and other fields.

[0028] At present, the methods of 3D text visual positioning can be divided into two categories: multimodal fusion assisted by 2D images and multimodal fusion directly based on 3D images: The first is multimodal fusion assisted by 2D images. Using the semantic information of 2D images to assist the feature representation of 3D scenes and the matching of visual languages is a relatively classic method. It uses the target detection results of 2D images to generate candidate regions of 3D scenes, and calculates the similarity between 2D images and 3D point clouds through a bidirectional attention mechanism, and then matches them with language features. This method can utilize the rich semantic information of 2D images to improve the representation ability of 3D scenes. In recent years, the work on 3D text visual positioning assisted by 2D images has also overcome the limitations of a single perspective, using multiple perspectives of 2D images to extract features of 3D scenes, and using attention mechanisms to fuse information from different perspectives.

[0029] The second is multimodal fusion directly based on 3D images. Direct 3D image-based methods are the mainstream approach for 3D text visual localization, and their model structure can be further divided into single-stage and two-stage approaches. Single-stage approaches utilize the semantic features of key points and language descriptions in the 3D scene to generate candidate regions. They then use self-attention and bidirectional attention mechanisms to calculate the similarity between the key points and language descriptions. Finally, a progressive strategy is used to gradually narrow the candidate regions or directly select the most matching region as the target object. This approach focuses on leveraging local information in the 3D scene to improve the accuracy and efficiency of candidate regions. Two-stage approaches first use a 3D object detector to obtain object proposal regions, then calculate similarity based on the fusion of the detected object features and the semantic features of the text, and select the object with the highest similarity as the output. Although existing two-stage approaches attempt to construct a view-robust multimodal representation and alleviate the view difference problem starting from the 3D modality, they ignore the view cues embedded in the text modality and apply the same weight to different viewpoints. This application proposes a multi-view method for 3D multimodality that improves the ability to capture view knowledge from both text and 3D modalities.

[0030] Currently, the existing two-stage methods based directly on 3D images have two options to alleviate this potential visual-text misalignment problem: (1) manually aligning the 3D scene with the paired text; (2) simultaneously inputting multiple views into the network in the 3D modality to obtain better viewpoint robustness. However, these methods have two major limitations. First, they only focus on solving the viewpoint dependency problem from the 3D modality, while ignoring the lack of viewpoint clues in the text input. Second, for multi-view input, they do not introduce a specially designed module to capture viewpoint knowledge, which is very important for distinguishing the relative importance of each viewpoint.

[0031] In order to solve the above technical problems, the present application provides a three-dimensional visual-text localization method based on multiple perspectives and multiple texts, comprising the following steps: first, extracting three-dimensional scene features from multiple perspectives from the input three-dimensional scene to obtain multi-perspective three-dimensional features; next, expanding the perspective and expression of the text to obtain text with multi-perspective information, and extracting features from the text with multi-perspective information to obtain multi-perspective text features; then performing multi-modal fusion of the multi-perspective three-dimensional features and the multi-perspective text features to obtain multi-perspective fusion features; finally, based on a perspective-guided scoring mechanism, scoring the multiple perspectives of the multi-perspective fusion features, using multi-perspective representative features to memorize scene-independent knowledge, and performing three-dimensional visual-text localization. The method proposed in this application simultaneously grasps perspective knowledge from text and three-dimensional modalities, designs a perspective-guided attention module to give different weights to perspectives, and further improves the three-dimensional localization performance.

[0032] The following detailed description of the various embodiments of the present application is provided in conjunction with the accompanying drawings. However, those skilled in the art will appreciate that many technical details are provided in the various embodiments of the present application to facilitate a better understanding of the present application. However, even without these technical details and the various variations and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.

[0033] See Figure 1 The embodiment of the present application provides a three-dimensional vision-text positioning method based on multiple perspectives and multiple texts, including:

[0034] Step S1: extracting three-dimensional scene features from multiple perspectives from an input three-dimensional scene to obtain multi-perspective three-dimensional features.

[0035] Step S2: Expand the perspective and expression of the text to obtain a text with multi-perspective information, and perform feature extraction on the text with multi-perspective information to obtain multi-perspective text features.

[0036] Step S3: Perform multimodal fusion on the multi-view 3D features and the multi-view text features to obtain multi-view fusion features.

[0037] Step S4: Based on the perspective-guided scoring mechanism, multiple perspectives of the multi-perspective fusion feature are scored, and the multi-perspective representative features are used to memorize scene-independent knowledge and perform three-dimensional vision-text positioning.

[0038] This application addresses the technical issues in the existing technology of "focusing only on the 3D modality, ignoring the perspective cues embedded in the text modality, and ignoring the measurement of the importance of multiple perspectives." This application explores how to simultaneously mine perspective knowledge from text and 3D modalities to solve the bottleneck of accuracy, which is of great value. Based on this, this application proposes a 3D vision-text localization method based on multiple perspectives and multiple texts. This method simultaneously processes and understands 3D vision and text from multiple perspectives, and uses multi-perspective representative features to memorize scene-independent knowledge to achieve more accurate and robust 3D vision-text localization.

[0039] This application provides a multi-view method for three-dimensional text visual positioning, exploring how to master perspective knowledge from text and three-dimensional modalities. For the text branch, this application uses the diverse language knowledge of large-scale language models to expand a single positioning text into multiple geometrically consistent descriptions. At the same time, in the three-dimensional modality, this application introduces a fusion module with inter-perspective attention to enhance the interaction between objects from different perspectives. In addition, this application also proposes a set of learnable multi-view representative features (Multi-view Prototype) to memorize scene-independent knowledge from different perspectives, and enhances the enhancement framework model from two aspects, specifically: a perspective-guided attention module for more robust text features, and a perspective-guided scoring strategy for final prediction.

[0040] In some embodiments, before extracting 3D scene features from multiple perspectives from the input 3D scene in step S1, the method further includes: obtaining a dataset and dividing the dataset into a training dataset and a test dataset; the dataset is a dataset of object positioning and language expression in the 3D scene. The training dataset is used to train the model in the model training phase; the test dataset is used to verify the prediction results in the verification and application phase.

[0041] The following is a detailed introduction to the three-dimensional vision-text positioning method based on multiple perspectives and multiple texts provided by this application.

[0042] See Figure 2 The process of the multi-view and multi-text based three-dimensional vision-text positioning method of this application is roughly divided into three stages: data processing, model training, verification and application.

[0043] In the data processing stage, this application selected a dataset for object localization and language expression in 3D scenes. This application used the original training set and test set division of the dataset for the vision-text localization task.

[0044] During the model training stage, this application proposes the core part of this application: a three-dimensional vision-text positioning method based on multi-view and multi-text.

[0045] The model framework consists of four parts: three-dimensional feature extraction, text expansion and feature extraction, multimodal fusion converter and perspective-guided scoring mechanism.

[0046] In some embodiments, three-dimensional scene features under multiple perspectives are extracted from an input three-dimensional scene to obtain multi-perspective three-dimensional features, including the following steps: inputting the point cloud data of each object in the three-dimensional scene into a point cloud encoder separately to extract the three-dimensional features of each object; performing three-dimensional rotation on the center point of each object to obtain the position information of all objects at each perspective; and adding the position information directly to the object features to obtain the three-dimensional scene features of the object corresponding to each perspective.

[0047] like Figure 3 As shown, in the process of 3D feature extraction, multi-view 3D features are extracted from the multi-view 3D scene. The 3D feature extraction part is used to extract 3D features from multiple viewpoints from the 3D point cloud for subsequent feature fusion. Specifically, the input 3D scene consists of two parts, the point cloud data of each object and the 3D coordinates of the center point of the object. First, the point cloud data of each object is input separately into the point cloud encoder to extract the 3D features of each object, and then the center point of each object is rotated in three dimensions to obtain the position information of all objects at each viewpoint. By adding the position information directly to the object features, the 3D scene features corresponding to each viewpoint can be obtained.

[0048] In the text expansion and feature extraction process, this application first expands and locates the text through a pre-trained large language model, expanding the text's perspective and expression. Then, the text containing multi-perspective information is input into the text feature extractor to extract features.

[0049] In some embodiments, after multi-view 3D features and multi-view text features are multi-modally fused, and before multiple views of the multi-view fusion features are scored, the system further includes: a view-guided attention module to extract view information from view representative features; and a view-guided attention module to enhance the learning and extraction of view knowledge by the text modality. Figure 3As shown, after multimodal fusion of multi-view 3D features and multi-view text features, a visually guided prediction value is obtained through a perspective-guided attention module. This application enhances the learning and extraction of perspective knowledge in the text modality by designing a perspective-guided attention module and extracting perspective information from the perspective representative features introduced later.

[0050] Please continue to see Figure 3 In the multimodal fusion stage, a multimodal fusion converter is used for multimodal fusion. The multimodal fusion converter consists of four converters, each of which includes a self-attention layer within the view, a multimodal interactive attention layer, and a self-attention layer between views, which are connected in sequence; the input is multi-view three-dimensional features and multi-view text features. The fusion converter module includes a self-attention layer within the view, a multimodal interactive attention layer, and a self-attention layer between views. Among them, the self-attention layer between views proposed in this application is used to fuse features between multiple views, aiming at deeper mining of view information. Through this module, multi-view three-dimensional features and multi-view text features can be effectively fused and aligned.

[0051] See Figure 4 In some embodiments, a scoring mechanism based on perspective guidance is used to score multi-perspective fusion features, and multi-perspective representative features are used to memorize scene-independent knowledge, including the following steps: first, the cosine similarity between the three-dimensional features and the perspective representative features under multiple perspectives is calculated; then, based on the cosine similarity, the multiple perspectives of the multi-perspective fusion features are scored to obtain scores for the multiple perspectives; finally, the scores are used as weights for final prediction.

[0052] In the perspective-guided scoring mechanism, this application first proposes perspective-representative features to memorize scene-independent knowledge. Specifically, the cosine similarity between the 3D features of multiple perspectives and the perspective-representative features is calculated to assign scores to the multiple perspectives. This score represents the relative importance of each perspective. This score is then used as a weight in the weighted summation during the subsequent prediction process.

[0053] In some embodiments, after scoring multiple perspectives of the multi-perspective fusion features, using the multi-perspective representative features to memorize scene-independent knowledge, and performing three-dimensional vision-text positioning, it also includes: verifying the three-dimensional vision-text positioning results, and applying the three-dimensional vision-text positioning method.

[0054] like Figure 2As shown, verification and application are the final stages of the technical solution process of this application, and the feasibility and effectiveness of the proposed method will be verified through actual experiments. In the experimental verification stage, this application selected other visual-text positioning methods for quantitative experiments and qualitative experiments. In the quantitative experiment part, the model of this application achieved the best results in most indicators; in the qualitative experiment, through the visualization of the results, this application found that the method of this application achieved the same level of results as other methods, and in some data, the method of this application showed better positioning effects.

[0055] See Figure 5 The embodiment of the present application also provides a three-dimensional vision-text positioning system based on multi-view and multi-text, including a data processing module 101 and a model training module 102 connected to each other; wherein the data processing module 101 is used to obtain three-dimensional scene data; the model training module 102 includes a multi-view three-dimensional feature acquisition unit 1021, a multi-view text feature acquisition unit 1022, a multimodal fusion unit 1023 and a scoring unit based on view guidance 1024 connected in sequence; wherein; the multi-view three-dimensional feature acquisition unit 1021 is used to extract three-dimensional scene features under multiple viewpoints from the input three-dimensional scene, and obtain multi-view Three-dimensional features; the multi-perspective text feature acquisition unit 1022 is used to expand the perspective and expression of the text to obtain a text with multi-perspective information, and perform feature extraction on the text with multi-perspective information to obtain multi-perspective text features; the multimodal fusion unit 1023 performs multimodal fusion on the multi-perspective three-dimensional features and the multi-perspective text features to obtain multi-perspective fusion features; the perspective-guided scoring unit 1024 is used to score multiple perspectives of the multi-perspective fusion features according to the perspective-guided scoring mechanism, use the multi-perspective representative features to memorize scene-independent knowledge, and perform three-dimensional vision-text positioning.

[0056] In some embodiments, the above-mentioned three-dimensional vision-text positioning system based on multiple perspectives and multiple texts also includes a verification and application module 103 connected to the model training module 102; the verification and application module 103 is used to verify the three-dimensional vision-text positioning results and apply the three-dimensional vision-text positioning method; the data processing module 101 includes a data set module 1011 and a data set division module 1012; the data set module 1011 is used to obtain a data set; the data set is a data set for object positioning and language expression in a three-dimensional scene; the data set division module 1012 is used to divide the data set into a training data set and a test data set.

[0057] This application proposes a 3D vision-text localization method based on multiple perspectives and multiple texts. It simultaneously processes and understands both 3D vision and text modalities from multiple perspectives, and uses multi-perspective representative features to memorize scene-independent knowledge, enabling more accurate and robust 3D vision-text localization. Furthermore, this application method designs a multimodal fusion converter with inter-perspective interaction, as well as a perspective-guided attention module for the text modality. This facilitates the framework's fusion of multimodal knowledge and enhances the learning and extraction of perspective knowledge by the text modality.

[0058] The technical solution proposed in this application was quantitatively tested on two major data sets. The results showed that the model of this application achieved better results than other related technical methods, such as Figure 6 shown.

[0059] Compared with existing technologies, this application's advantages lie in: It fully utilizes information from both modalities, enabling the model to deeply explore the perspective information contained in 3D scenes and text, achieving more accurate and robust 3D visual-text localization. Furthermore, this technological advancement can further promote the development of fields and applications such as visual language navigation, autonomous driving, and embodied intelligence.

[0060] The technical solution proposed in this application is qualitatively tested on real data sets, such as Figure 7 As shown, through visualization of the results, this application found that the method of this application achieved the same level of results as other methods, and in some data, the method of this application showed better positioning effects.

[0061] Based on the above technical solution, an embodiment of the present application provides a three-dimensional vision-text positioning method based on multi-perspective and multi-text, which includes the following steps: first, extracting three-dimensional scene features under multiple perspectives from the input three-dimensional scene to obtain multi-perspective three-dimensional features; next, expanding the perspective and expression method of the text to obtain text with multi-perspective information, and extracting features from the text with multi-perspective information to obtain multi-perspective text features; then performing multi-modal fusion of the multi-perspective three-dimensional features and the multi-perspective text features to obtain multi-perspective fusion features; finally, based on a perspective-guided scoring mechanism, scoring multiple perspectives of the multi-perspective fusion features, using multi-perspective representative features to memorize scene-independent knowledge, and performing three-dimensional vision-text positioning.

[0062] This application proposes a 3D vision-text localization method based on multiple perspectives and multiple texts. It simultaneously processes and understands the two modalities of 3D vision and text from multiple perspectives, and uses multi-perspective representative features to memorize scene-independent knowledge, which can achieve more accurate and robust 3D vision-text localization. At the same time, the method of this application also designs a multimodal fusion converter with inter-perspective interaction and an attention module for perspective guidance of the text modality, which can promote the framework's fusion of multimodal knowledge and enhance the learning and extraction of perspective knowledge by the text modality.

[0063] Those skilled in the art will appreciate that the above-described embodiments are specific examples for implementing the present application, and that in actual applications, various changes in form and detail may be made thereto without departing from the spirit and scope of the present application. Any person skilled in the art may make changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be subject to the scope defined in the claims.

Claims

1. A 3D vision-text localization method based on multiple perspectives and multiple texts, characterized in that: The following steps are involved: Extracting 3D scene features from multiple perspectives from an input 3D scene to obtain multi-perspective 3D features; Expanding the perspective and expression of the text to obtain a text with multi-perspective information, and performing feature extraction on the text with multi-perspective information to obtain multi-perspective text features; Performing multimodal fusion on the multi-view three-dimensional features and the multi-view text features to obtain a multi-view fusion feature; Based on a perspective-guided scoring mechanism, multiple perspectives of the multi-perspective fusion feature are scored, and scene-independent knowledge is memorized using multi-perspective representative features to perform 3D vision-text positioning; The perspective-guided scoring mechanism scores the multi-perspective fusion features and uses the multi-perspective representative features to memorize scene-independent knowledge, including the following steps: Calculate the cosine similarity between 3D features and view representative features under multiple views; Scoring the multiple perspectives of the multi-perspective fusion feature based on the cosine similarity to obtain scores for the multiple perspectives; The scores are used as weights for the final prediction; After performing multimodal fusion on the multi-view 3D features and the multi-view text features, and before scoring the multiple views of the multi-view fusion features, the method further includes: A perspective-guided attention module extracts perspective information from perspective representative features; the perspective-guided attention module is used to enhance the learning and extraction of perspective knowledge in text modality.

2. The 3D vision-text localization method based on multiple perspectives and multiple texts according to claim 1, characterized in that: Extracting 3D scene features from multiple perspectives from an input 3D scene to obtain multi-perspective 3D features includes the following steps: The point cloud data of each object in the 3D scene is input into the point cloud encoder separately to extract the 3D features of each object; the center point of each object is rotated in 3D to obtain the position information of all objects at each viewing angle; The position information is directly added to the object features to obtain the three-dimensional scene features of the object corresponding to each perspective.

3. The 3D vision-text localization method based on multiple perspectives and multiple texts according to claim 1, characterized in that: The input three-dimensional scene includes point cloud data of each object and the three-dimensional coordinates of the center point of the object.

4. The 3D vision-text localization method based on multiple perspectives and multiple texts according to claim 1, characterized in that: A pre-trained large language model is used to expand the perspective and expression of the text; and a text feature extractor is used to extract features from the text with multi-perspective information to obtain multi-perspective text features.

5. The 3D vision-text localization method based on multiple perspectives and multiple texts according to claim 1, characterized in that: Before extracting 3D scene features from multiple perspectives from the input 3D scene, it also includes: A data set is obtained and divided into a training data set and a test data set; the data set is a data set of object positioning and language expression in a three-dimensional scene.

6. The 3D vision-text localization method based on multiple perspectives and multiple texts according to claim 1, characterized in that: After scoring the multiple perspectives of the multi-perspective fusion feature, memorizing scene-independent knowledge using the multi-perspective representative features, and performing 3D vision-text positioning, the method further includes: Verify the 3D vision-text localization results and apply the 3D vision-text localization method.

7. A 3D vision-text positioning system based on multiple perspectives and multiple texts, characterized by: It includes connected data processing modules and model training modules; among them, The data processing module is used to obtain three-dimensional scene data; The model training module includes a multi-view 3D feature acquisition unit, a multi-view text feature acquisition unit, a multimodal fusion unit, and a scoring unit based on view guidance, which are connected in sequence; The multi-view 3D feature acquisition unit is used to extract 3D scene features under multiple viewpoints from an input 3D scene to obtain multi-view 3D features; The multi-perspective text feature acquisition unit is used to expand the perspective and expression of the text to obtain a text with multi-perspective information, and perform feature extraction on the text with multi-perspective information to obtain multi-perspective text features; The multimodal fusion unit performs multimodal fusion on the multi-view three-dimensional features and the multi-view text features to obtain a multi-view fusion feature; The perspective-guided scoring unit is used for scoring the multiple perspectives of the multi-perspective fusion feature based on the perspective-guided scoring mechanism, and uses the multi-perspective representative features to memorize scene-independent knowledge and perform three-dimensional vision-text positioning; The perspective-guided scoring mechanism scores the multi-perspective fusion features and uses the multi-perspective representative features to memorize scene-independent knowledge, including the following steps: Calculate the cosine similarity between 3D features and view representative features under multiple views; Scoring the multiple perspectives of the multi-perspective fusion feature based on the cosine similarity to obtain scores for the multiple perspectives; The scores are used as weights for the final prediction; After performing multimodal fusion on the multi-view 3D features and the multi-view text features, and before scoring the multiple views of the multi-view fusion features, the method further includes: A perspective-guided attention module extracts perspective information from perspective representative features; the perspective-guided attention module is used to enhance the learning and extraction of perspective knowledge in text modality.

8. The multi-view and multi-text based 3D vision-text positioning system according to claim 7, characterized in that: Also included is a verification and application module connected to the model training module; The verification and application module is used to verify the 3D vision-text positioning result and apply the 3D vision-text positioning method; The data processing module includes a data set module and a data set partitioning module; The data set module is used to obtain a data set; the data set is a data set of object positioning and language expression in a three-dimensional scene; The data set division module is used to divide the data set into a training data set and a test data set.

Citation Information

Patent Citations

  • Text guidance image segmentation method based on cross-modal text retrieval attention mechanism

    CN113657400A

  • Scene character feature extraction method and device based on multi-modal information and application

    CN116469107A