Detailed three-dimensional directivity target segmentation method

Through the detailed three-dimensional directive target segmentation method, the limitations of sentence-level positioning in the existing 3D vision-language task are solved. By constructing the DetailRefer data set and the DetailBase baseline model, the sentence and phrase-level segmentation is realized, the model's context reasoning and semantic understanding ability is improved, and the development of 3D vision-language task is promoted.

CN120388032AActive Publication Date: 2025-07-29XIAMEN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510880991.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing 3D vision-language tasks are limited to sentence-level target positioning and cannot effectively evaluate the model's fine-grained language comprehension ability, limiting the modeling of internal relationships and semantic structures of sentences, resulting in limited performance when dealing with complex natural language instructions.

Method used

A detailed three-dimensional directive target segmentation method is proposed. By defining task forms and evaluation indicators, the DetailRefer data set is constructed and the DetailBase baseline model is designed to realize sentence-level and phrase-level segmentation, and multimodal information fusion is used to generate accurate mask segmentation.

Benefits of technology

It improves the performance of the model when processing complex natural language instructions, enhances the context reasoning ability of 3D vision-language tasks, provides finer-grained semantic understanding, and promotes the development of 3D vision-language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388032A_ABST
    Figure CN120388032A_ABST
Patent Text Reader

Abstract

The invention discloses a detailed three-dimensional directivity target segmentation method. The method comprises the following steps: S1, defining a task form and defining an evaluation index of a task; s2, modifying and enhancing the ScanRefer data set in combination with manual operation and a large model so as to generate a DetailRefer data set; s3, constructing a DetailBase baseline model, and segmenting a sentence-level language or a phrase-level language through the DetailBase baseline model; according to the method, by defining a task form, generating a DetailRefer data set and constructing a DetailBase baseline model, the ability of understanding and positioning text contexts in 3D vision and language tasks can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer data processing, and particularly relates to a detailed three-dimensional directional target segmentation method. Background Art

[0002] Vision-language interaction stands at the forefront of computer vision research, enabling machines to understand and reason about the visual world through natural language. In recent years, with the rapid development of 3D sensing technology and deep learning models, 3D vision-language tasks have become an important research focus in this field. These tasks play important roles in key applications such as robotics, autonomous navigation, mixed reality, and assistive technologies, as understanding the relationship between the 3D environment and natural language is essential in these applications. By combining the latest advancements, researchers can now develop more intelligent and intuitive human-computer interaction systems that can not only "see" the surrounding environment but also interpret and communicate this information in a human-understandable way. This brings new possibilities to many fields, from improving the quality of life of people with disabilities to enhancing the realism of virtual reality experiences.

[0003] Among these tasks, 3D vision grounding (3D-VG) tasks are particularly important because they require localizing targets in 3D scenes according to text instructions - a fundamental ability for embodied AI and autonomous systems. This field has developed systematically. Initially, it was 3D referential object captioning (3D-REC), which used rough 3D bounding boxes to localize objects and formulated the problem as coordinate regression (as shown in part (a) of Figure 1 ). Subsequently, to meet the need for more fine-grained localization, 3D referential object segmentation (3D-RES) was introduced, which requires point-level segmentation and transforms the task into an expression-to-point matching problem (as shown in part (b) of Figure 1 ). However, both of these methods are limited to handling one-to-one mappings between sentences and objects, restricting their practical applicability in cases where the instructions may involve multiple objects or no objects at all. To address this limitation, generalized 3D referential object segmentation (3D-GRES) (as shown in part (c) of Figure 1 ) extends the task description to accommodate the case of zero, one, or multiple targets for each text description. These developments have greatly advanced the 3D-VG field and made it a cornerstone of 3D scene understanding. In this way, researchers can create more flexible and powerful systems that can better understand and respond to complex natural language instructions, showing great potential in multiple application fields.

[0004] Despite these remarkable advancements, current 3D-VG tasks are still confined to sentence-level objectives, which hinders a comprehensive evaluation of the model's fine-grained language understanding capabilities. This sentence-centric approach fundamentally limits the modeling of intra-sentence relationships and semantic structures. As shown in Figure 1 part (b), traditional 3D-RES methods fail to provide a mechanism to determine whether the model correctly understands each element within a single instruction. This limitation creates a critical gap in interpretability, as effective referential expression understanding inherently relies on context reasoning abilities within the text. Constrained by the sentence-level supervision paradigm, current methods struggle to develop these necessary context reasoning skills. Therefore, a more meticulous, phrase-level understanding is needed to enable 3D vision-language models to develop strong context reasoning capabilities, which represents an important unexplored opportunity in this field. By introducing this finer level of understanding, researchers can better evaluate and enhance the model's performance in handling complex natural language instructions, thereby driving the development of 3D-VG technology. Summary of the Invention

[0005] To address the above problems, the present invention proposes a detailed three-dimensional pointing target segmentation method.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A detailed three-dimensional pointing target segmentation method, comprising the following steps: S1. Define the task form and the evaluation metrics of the task; In step S1, defining the task form is used to segment masks corresponding to each target phrase given in the sentence from the point cloud scene. The specific process is as follows: S11. Given a point cloud scene , where is the number of points, is the feature length; the feature length includes coordinates XYZ, color RGB, and normal vector; S12. Given a text description , where is the number of words in the text; S13. Given a set of indices , where is the index position of target phrases to be segmented in the text; the indices correspond to the positions of nouns that need to be segmented in the text, and the model outputs point cloud scene masks for all nouns to understand natural language descriptions and mark the corresponding objects or regions in three-dimensional space; S2. Modify and enhance the ScanRefer dataset by combining human and large models to generate the DetailRefer dataset; S3. Construct a DetailBase baseline model to segment language at the sentence level or phrase level through the DetailBase baseline model; In step S3, the specific processing process of the DetailBase baseline model is as follows: S31. Input the point cloud scene , text description and the index of the noun to be segmented , and input the point cloud into the 3D U-Net network to obtain point-level features; among them, the point cloud only uses the coordinates XYZ and color RGB as the initial features of each point; S32. Simplify the point-level features by superpoint pooling, and perform unsupervised over-segmentation on the point cloud scene to generate superpoints, where ; S33. Average the features of all points belonging to the same superpoint, and then through two independent linear transformations, convert the pooled features into visual features for multimodal information fusion and superpoint features for predicting masks , where represents the feature dimension; S34. For the given text description , add special words at the beginning and end and then input it into the MPNet network to obtain word features, and use the word features to generate an initial query through a multi-layer perceptron ; S35. Input the initial query into the decoder, use cross-attention in the decoder to integrate information from the visual modality, then use self-attention to focus on the information inside the sentence, and perform non-linear transformation through a feed-forward neural network; S36. Calculate the affinity between the query output by the last layer and the superpoint features , and binarize the affinity to obtain a superpoint mask corresponding to the query, and then broadcast the superpoint mask to obtain a point-level mask; for sentence-level segmentation, use the mask corresponding to the [CLS] token as the segmentation result; for phrase-level segmentation, use the mask generated by the query corresponding to the position provided in the index as the segmentation result.

[0007] Preferably, in step S1, the specific process of defining the evaluation index of the task is as follows: S14. Pay attention to the average intersection over union (IoU), Acc@0.25, and Acc@0.5 at the phrase level; at the phrase level, the average IoU is the average of the IoUs calculated for all phrases that need to be segmented, and Acc@0.25 and Acc@0.5 respectively represent the proportions of the IoUs greater than 0.25 and 0.5 among all segmented phrases; S15. Pay attention to the average IoU at the sentence level, that is, calculate the phrase IoU within each sentence and then average it, and then average all sentences at the dataset level; S16. Pay attention to the performance of the model on long texts and complex scene descriptions. Define texts containing more than 50 words as long texts, and evaluate the metrics of the model on long texts to reflect its understanding ability of long texts; define texts that need to segment four or more phrases as complex scene descriptions, and test the metrics of the model on complex scene descriptions to reflect its understanding ability of complex scene descriptions.

[0008] Preferably, the specific process of step S2 is as follows: S21. In the first stage, write a piece of code using the Python programming language to divide the descriptions in the ScanRefer dataset so that all descriptions of the same object are grouped together; then let the large language model (LLM) integrate the sentences into a more comprehensive new description, and after obtaining a threshold number of new descriptions, manually annotate all noun phrases in each text, associate all noun phrases with specific objects in the 3D scene, and during the annotation process, correct the inaccuracies in the new description to achieve the modification of the ScanRefer dataset; S22. In the second stage, convert the annotated text into a format that the LLM can understand, immediately add a bracket after each noun phrase, and place the object ID corresponding to the noun phrase inside the bracket; then input the sentence into the LLM and instruct the LLM to generate several different expressions while retaining the original semantic meaning and keeping the object ID closely following the corresponding noun phrase, so as to expand the dataset to five times its original size to achieve the enhancement of the ScanRefer dataset; S23. Traverse each object mentioned in the enhanced ScanRefer dataset, extract all texts related to the object from the dataset in the first stage, and input the text into the LLM for integration to generate a text with a more extensive description to obtain the DetailRefer dataset.

[0009] Preferably, the threshold number in step S21 is 10,000.

[0010] Preferably, the decoder described in step S35 adopts a cascaded multi-layer architecture, and the decoder includes structures of multiple cross-attention, self-attention, and feed-forward neural networks.

[0011] After adopting the above technical solutions, the present invention has the following beneficial effects: 1. The present invention proposes a new task, namely the detailed three-dimensional directional object segmentation method (3D-DRES), which is a novel fine-grained visual localization task aimed at enhancing the ability to understand and localize text context in 3D vision and language tasks. This task requires the model to segment the mask corresponding to each noun phrase in the sentence. Specifically, as Figure 1 shown in part (d) of the figure, the dataset provides the positions of all phrases to be segmented in the sentence (such as "trash can", "table", and "TV"), and the model needs to generate corresponding masks for each phrase respectively. The nature of the 3D-DRES task determines that the model trained under this task pays more attention to the fine-grained semantics within the sentence. This feature actually complements the 3D-RES task that emphasizes sentence-level semantics. In addition, 3D-RES and 3D-DRES can promote each other. This has also been proven in the experiments.

[0012] 2. The present invention constructs a new dataset called DetailRefer based on the ScanRefer dataset, combining fine manual annotation and the assistance of large language models. The DetailRefer dataset contains 55,432 descriptions, covering 11,054 different objects, with an average text length of 24.9 words, significantly exceeding the existing datasets (9.7 to 20.1 words). Different from the sentence-segmentation annotation format in which one sentence corresponds to one segmentation result commonly adopted in the current datasets, the DetailRefer dataset introduces a groundbreaking phrase-segmentation format, in which each noun phrase corresponds to an independent segmentation mask. This makes the dataset reach an unprecedented density, that is, there are 2.9 masks per piece of text. The dataset strategically includes 7.4% of "long" texts (more than 50 words) and a large number of "complex" samples that require segmenting four or more noun phrases, providing strong support for evaluating the model's fine-grained language understanding ability in a 3D environment.

[0013] 3. Most existing models in the present invention are designed for the "sentence-segmentation" sample format and cannot output multiple masks or specify masks for specific tokens. This limitation makes the current methods unable to be directly applied to the 3D-DRES task. Therefore, as the initiator of this task, the present invention proposes the DetailBase baseline model, a deliberately simplified but effective baseline model, to lay the foundation for future research. The design concept of the present invention gives priority to simplicity to ensure high scalability and adaptability, while demonstrating sufficient effectiveness to verify the potential of the task. The DetailBase baseline model has an elegant architecture that supports sentence-level and phrase-level segmentation. The results show that even on the traditional 3D-RES benchmark, training through this fine-grained task can bring significant improvements, indicating that phrase-level understanding can enhance the overall spatial reasoning ability. By introducing the DetailBase baseline model, not only a strong starting point is provided for the 3D-DRES task, but also the potential of fine-grained language understanding in improving the performance of 3D vision tasks is demonstrated. This paves the way for further exploration and development of more advanced 3D vision-language models and promotes the development of this field. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 Target comparison diagrams for different tasks; Figure 2 Flowchart of the present invention; Figure 3 Frame structure diagram of the DetailBase baseline model of the present invention; Figure 4 Visual comparison result diagram of the DetailBase baseline model of the present invention and the 3D-STMN model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0016] As Figures 1 to 4 shown, a detailed three-dimensional directed target segmentation method (abbreviated as 3D-DRES) includes the following steps: S1. Define the task form and the evaluation index of the task; In step S1, defining the task form is used to segment masks corresponding to each target phrase given in the sentence from the point cloud scene. The specific process is as follows: S11. Given a point cloud scene , where is the number of points, is the feature length; the feature length includes coordinates XYZ, color RGB, and normal vectors; S12. Given a text description , where is the number of words in the text; S13. Given a set of indices , where is the index position of the target phrases to be segmented in the text; the indices correspond to the positions of the nouns in the text that need to be segmented. The model outputs a point cloud scene mask for all nouns to understand the natural language description and mark the corresponding objects or regions in 3D space; In step S1, the specific process of defining the evaluation metrics for the task is as follows: S14. Pay attention to the mean Intersection over Union (mIoU), Acc@0.25, and Acc@0.5 at the phrase level; at the phrase level, the mIoU is the average of the Intersection over Union calculated for all phrases that need to be segmented, and Acc@0.25 and Acc@0.5 represent the proportions of the Intersection over Union greater than 0.25 and 0.5 among all segmented phrases, respectively; S15. Pay attention to the mean Intersection over Union at the sentence level, that is, after calculating and averaging the phrase Intersection over Union within each sentence, then averaging over all sentences at the dataset level; S16. Pay attention to the performance of the model on long texts and complex scene descriptions. Define texts with more than 50 words as long texts and evaluate the metrics of the model on long texts to reflect its understanding ability of long texts; define texts that need to segment four or more phrases as complex scene descriptions and test the metrics of the model on complex scene descriptions to reflect its understanding ability of complex scene descriptions; S2. Combine artificial and large models to modify and enhance the ScanRefer dataset to generate the DetailRefer dataset; The specific process of step S2 is as follows: S21. In the first stage, write a piece of code using the Python programming language to divide the descriptions in the ScanRefer dataset so that all descriptions of the same object are grouped together; then let the large language model (LLM) integrate the sentences into a more comprehensive new description. After obtaining the threshold number of new descriptions (the threshold number mentioned in step S21 is 10,000), manually annotate all noun phrases in each text, associate all noun phrases with specific objects in the 3D scene, and correct the inaccuracies in the new descriptions during the annotation process to achieve the modification of the ScanRefer dataset; The threshold number mentioned in step S21 is 10,000; S22. In the second stage, the annotated text is converted into a format that the LLM large language model can understand. A bracket is added immediately after each noun phrase, and the object ID corresponding to the noun phrase is placed in the bracket. The sentence is then input into the LLM large language model, and the LLM large language model is instructed to generate several different expressions while retaining the original semantic meaning and keeping the object ID closely following the corresponding noun phrase. This is used to expand the dataset to five times its original size, thereby enhancing the ScanRefer dataset. S23, traverse each object mentioned in the enhanced ScanRefer dataset, extract all texts related to the object from the dataset of the first stage, and input the text into the LLM large language model for integration to generate text with a broader description, and obtain the DetailRef dataset; S3. Build a DetailBase baseline model and use it to segment language at the sentence or phrase level. In step S3, the specific processing process of the DetailBase baseline model is as follows: S31. Input point cloud scene , text description and the index of the nouns that need to be split , by inputting the point cloud into the 3D U-Net network to obtain point-level features; where the point cloud only uses the coordinates XYZ and color RGB as the initial features of each point; S32, use super point pooling to simplify the point-level features, and Perform unsupervised over-segmentation to generate Super points, among which ; S33, average the features of all points belonging to the same superpoint, and then transform the pooled features into visual features for multimodal information fusion through two independent linear transformations and super-point features for predicting masks ,in, Represents feature dimension; S34. For a given text description , after adding special words at the beginning and end, input it into the MPNet network to obtain word features, and use the word features to generate the initial query through the multi-layer perceptron ; S35, the initial query The input is sent to the decoder, where cross-attention is used to integrate information from the visual modality, self-attention is used to focus on information within the sentence, and nonlinear transformation is performed through a feedforward neural network; The decoder described in step S35 adopts a cascaded multi-layer architecture, and the decoder contains structures of multiple cross-attention, self-attention, and feed-forward neural networks; S36. Calculate the affinity between the query output by the last layer and the superpoint features and binarize the affinity to obtain a superpoint mask corresponding to the query, and then broadcast the superpoint mask to obtain a point-level mask; for sentence-level segmentation, use the mask corresponding to the [CLS] token as the segmentation result; for phrase-level segmentation, use the mask generated by the query corresponding to the position provided in the index as the segmentation result.

[0017] Performance test: Since there is currently no model that can directly adapt to the task of a detailed three-dimensional directional object segmentation method (abbreviated as 3D-DRES) of the present invention, two existing models (i.e., the PNG model and the 3D-STMN model) were appropriately adjusted to make them suitable for the corresponding tasks. The PNG model originated from the panoramic narrative localization task in the 2D field, and this task is very similar to the task of the present invention. First, its input was changed by using point clouds instead of images, and then the SPFormer was used to extract instance candidates, and the point features within each instance were averaged as the instance features. Its result matching method was also modified, changing from selecting the candidate with the highest matching similarity to selecting the candidate with a matching similarity higher than a certain threshold. This adjustment is because in the task, a noun phrase may correspond to multiple instances. For the 3D-STMN model, its supervision method and result generation method were modified to align with the DetailBase dataset. In addition, in the original 3D-STMN model, 3D-UNet was frozen. To ensure fairness, it was selected to be included in the training process here.

[0018] Table 1 shows the quantitative comparison results of the PNG model, the 3D-STMN model, and the DetailBase baseline model on the DetailRefer dataset respectively. It can be seen that although the modified PNG model reached an average intersection over union of 40.4 on the test set, it still significantly lagged behind other models. The reason for this situation is that the current 3D instance segmentation network has an unsatisfactory effect; such two-stage models usually perform weaker than single-stage models. The modified 3D-STMN model demonstrated its ability in this task, but due to the low compatibility of its design module with the task, its performance was inferior to the DetailBase baseline model proposed by the present invention. The DetailBase baseline model reached an mIoU of 55.7 on the test set, and this result laid a good foundation for the further development of the 3D-DRES method. In addition, the model framework of the present invention is relatively simple and highly scalable, and is very suitable as an initial method for this task.

[0019] Table 1: Quantitative comparison results of the PNG model, 3D-STMN model, and DetailBase baseline model on the DetailRefer dataset respectively

[0020] Phrase-level segmentation emphasizes fine-grained semantic understanding, while sentence-level segmentation pays more attention to overall understanding. These two methods are not mutually exclusive but complementary. Here, joint training experiments were conducted to verify this view. By treating the [CLS] word (as the root node in the 3D-STMN model) as the "noun phrase" to be segmented, the formats of the two tasks were unified. Separate training was carried out on the ScanRefer and DetailRefer datasets respectively, and joint training was carried out on these two datasets. The mean intersection over union of the model on the validation set is reported in Table 2. Obviously, compared with separate training, joint training brings better results for both the 3D-STMN model and the DetailBase baseline model. It is worth noting that joint training significantly improves the performance of the 3D-RES task. The score of the DetailBase baseline model increases by 2.8 points, and the score of the 3D-STMN model increases by up to 3.2 points. In summary, the task of the present invention not only has its unique value but also can be synergistically complementary with traditional tasks.

[0021] Table 2: Comparison results of the DetailBase baseline model and the 3D-STMN model under different training strategies

[0022] In Figure 4 the visual comparison results between the DetailBase baseline model and the 3D-STMN model are shown. From Figure 4 it, the advantages of the 3D-DRES task can be more intuitively perceived. As shown in part (a) of Figure 4 , if it is a traditional 3D-RES task, the text points to the object "toilet", and both models successfully segment the target. In the traditional 3D-RES task, since the result is only evaluated based on a single entity in the text, it is difficult to evaluate the model's fine-grained understanding of the entire text. However, under the 3D-DRES task setting of the present invention, the model's overall understanding of the text can be observed in more detail. For example, the 3D-STMN model's understanding of the text in part (b) of Figure 4 is completely wrong, while its understanding of the text in part (c) of Figure 4 is partially correct and partially wrong.

[0023] As described above, it is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A detailed three-dimensional directional target segmentation method, characterized in that The following steps are involved: S1. Define the task form and the evaluation indicators of the task; In step S1, the task form is defined to segment the mask corresponding to each target phrase given in the sentence from the point cloud scene. The specific process is: S11. Given point cloud scene , where is the number of points, is the feature length; the feature length includes coordinates XYZ, color RGB, and normal vector; S12. Given a text description , where is the number of words in the text; S13. Given a set of indices , where are the index positions of target phrases to be segmented in the text; the indices correspond to the positions of nouns in the text that need to be segmented, and the model outputs a point cloud scene mask for all noun points to understand the natural language description and mark the corresponding objects or regions in 3D space; S2, modify and enhance the ScanRefer dataset by combining manual and large models to generate the DetailRefer dataset; S3. Build a DetailBase baseline model and use it to segment language at the sentence or phrase level. In step S3, the specific processing process of the DetailBase baseline model is as follows: S31. Input point cloud scene . Text description and the index of the nouns to be segmented , by inputting the point cloud into the 3D U-Net network to obtain point-level features; among them, the point cloud only uses the coordinates XYZ and color RGB as the initial features of each point; S32. Simplify the point-level features by using superpoint pooling, and perform unsupervised over-segmentation on the point cloud scene to generate superpoints, where ; S33. Average the features of all points belonging to the same superpoint, and then through two independent linear transformations, convert the pooled features into visual features for multimodal information fusion and superpoint features for predicting masks , where represents the feature dimension; S34. For the given text description , add special words at the beginning and end respectively and input it into the MPNet network to obtain word features, and use the word features to generate an initial query through a multi-layer perceptron ; S35. Input the initial query into the decoder. In the decoder, use cross-attention to integrate information from the visual modality, then use self-attention to focus on the information within the sentence, and perform a non-linear transformation through a feed-forward neural network; S36. Calculate the affinity between the query output by the last layer and the hyperpoint features and binarize the affinity to obtain a hyperpoint mask corresponding to the query. Then broadcast the hyperpoint mask to obtain a point-level mask. For sentence-level segmentation, use the mask corresponding to the [CLS] token as the segmentation result. For phrase-level segmentation, use the mask generated by the query corresponding to the position provided in the index as the segmentation result.

2. The detailed three-dimensional directional target segmentation method according to claim 1, characterized in that, In step S1, the specific process of defining the evaluation indicators of the task is as follows: S14. Focus on the average IoU, Acc@0.25, and Acc@0.5 at the phrase level. At the phrase level, the average IoU is the average of the IoUs calculated for all phrases to be segmented. Acc@0.25 and Acc@0.5 represent the proportion of all segmented phrases with an IoU greater than 0.25 and 0.5, respectively. S15: Focus on the sentence-level average IoU. That is, after calculating and averaging the phrase IoU within each sentence, we average it across all sentences at the dataset level. S16. Focus on the model's performance on long texts and complex scene descriptions. Define texts containing more than 50 words as long texts, and evaluate the model's indicators on long texts to reflect its ability to understand long texts. Define texts that need to be segmented into four or more phrases as complex scene descriptions, and test the model's indicators on complex scene descriptions to reflect its ability to understand complex scene descriptions.

3. The detailed three-dimensional directional target segmentation method according to claim 1, characterized in that, The specific process of step S2 is: S21. In the first stage, a piece of code is written in Python to partition the descriptions in the ScanRefer dataset so that all descriptions of the same object are grouped together. Then, the Large Language Model (LLM) is used to combine the sentences into a more comprehensive new description. After obtaining a threshold number of new descriptions, all noun phrases in each text are manually annotated, and all noun phrases are associated with specific objects in the 3D scene. During the annotation process, inaccuracies in the new descriptions are corrected to achieve the modification of the ScanRefer dataset. S22. In the second stage, the annotated text is converted into a format that the LLM large language model can understand. A bracket is added immediately after each noun phrase, and the object ID corresponding to the noun phrase is placed in the bracket. The sentence is then input into the LLM large language model, and the LLM large language model is instructed to generate several different expressions while retaining the original semantic meaning and keeping the object ID closely following the corresponding noun phrase. This is used to expand the dataset to five times its original size, thereby enhancing the ScanRefer dataset. S23. Traverse each object mentioned in the enhanced ScanRefer dataset, extract all texts related to the objects from the dataset in the first stage, and input the texts into the large language model (LLM) for integration to generate texts with more extensive descriptions, obtaining the DetailRefer dataset.

4. The detailed three-dimensional directional target segmentation method according to claim 3, characterized in that The threshold quantity described in step S21 is 10,000.

5. A detailed three-dimensional directional target segmentation method according to claim 1, characterized in that, The decoder described in step S35 adopts a cascaded multi-layer architecture, and the decoder contains multiple structures of cross-attention, self-attention, and feed-forward neural networks.

Citation Information

Patent Citations

  • Directive 3D instance segmentation method based on text information

    CN117634486A

  • Automated generation and presentation of visual data enhancements on camera view images captured in building

    CN117745983A

  • Three-dimensional directivity target segmentation method of image enhancement prompt decoding network

    CN119625011A

  • Three-dimensional directivity target segmentation method under weak supervision setting

    CN119649030A

  • Text segmentation with two-level transformer and auxiliary coherence modeling

    US11748571B1