3D visual object positioning method based on visual angle information and relation decoupling
By introducing viewing angle information and relationship decoupling modules into the 3D visual object positioning method, the crossover and self-attention mechanisms are used to solve the problems of viewing angle inconsistency and complex spatial description, and higher positioning accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510082319.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Existing 3D visual object positioning methods have problems with accuracy and robustness when dealing with viewing angle inconsistency and complex spatial descriptions, especially in text processing with viewing angle changes and high semantic complexity.
A 3D visual object positioning method based on decoupling of perspective information and relationships is proposed. By designing a simple relationship decoupling module and a viewing information transfer module, the cross attention mechanism and self-attention mechanism are used to enhance the transmission of perspective information and the understanding of spatial relationships, and improve the alignment of multimodal features.
Effectively eliminate inconsistencies in perspective, simplify the understanding of spatial relationships, improve the accuracy and robustness of 3D visual object positioning, and enhance the efficient learning ability of cross-modal tasks.
Smart Images

Figure CN120125655A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cross-modal 3D visual object localization, and in particular to a 3D visual object localization method based on perspective information and relationship decoupling. Background Art
[0002] Accurately identifying and localizing target objects in 3D space. With the development of computer vision and natural language processing technologies, more and more research has begun to focus on how to effectively combine spatial descriptions with 3D visual information. Although some progress has been made in this field in recent years, there are several technical challenges, mainly concentrated in perspective inconsistency and the semantic complexity of spatial descriptions. Most current 3D visual object localization methods rely on jointly modeling visual information and spatial descriptions. However, perspective inconsistency is an important challenge among them. Under different perspectives, the appearance and spatial relationships of objects will change, resulting in inconsistencies between spatial descriptions and visual information. For example, an object is described as "the bedside table is on the right side of the bed" in the front view, but in the back view, the same object will appear on the left side of the bed. This description inconsistency caused by perspective changes makes it difficult for traditional methods to maintain an accurate correspondence when combining visual and text information. Existing methods address perspective inconsistency by developing perspective-robust multimodal representations, which can effectively integrate visual information from different perspectives. However, complex spatial descriptions still pose a great challenge to the model. Spatial descriptions often contain high semantic complexity and long sentence structures, making it difficult for the model to extract key spatial relationships and object localization information. Existing language models have difficulty distinguishing which objects are anchor points and which are targets when processing such complex texts, thus affecting the effective fusion of text and visual features. Summary of the Invention
[0003] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and propose a 3D visual object localization method based on perspective information and relationship decoupling, which helps the model to more effectively eliminate perspective inconsistency when processing complex spatial descriptions and multi-perspective visual information, and at the same time simplifies the understanding of spatial relationships, thereby improving the accuracy and robustness of 3D visual object localization.
[0004] To achieve the above purpose, the technical solution provided by the present invention is: a 3D visual object localization method based on perspective information and relationship decoupling, comprising the following steps:
[0005] S1: Obtain multimodal data, and preprocess the data to obtain 3D scene point cloud data of a unified size and text with punctuation removed;
[0006] S2: A simple relationship decoupling module is designed. This module decouples the spatial relationships of the preprocessed text. Based on the spatial position relationship words in the text and leveraging the understanding ability of the large language model, it transforms the originally complex text into multiple sets of simplified spatial descriptions of points - relationships - targets;
[0007] S3: Feature extraction is performed on the preprocessed 3D scene point cloud data. Multiple point cloud objects are obtained from the 3D scene point cloud and feature extraction is carried out through a 3D object encoder to obtain 3D object features. At the same time, multiple sets of simplified spatial descriptions are transformed into multiple sets of spatial description features through a text encoder;
[0008] S4: A perspective information transfer module is designed. This module combines learnable multi - perspective markers with spatial description features using the cross - attention mechanism, and at the same time transfers the corresponding perspective markers to the 3D object features under the corresponding perspectives. According to the learnable characteristics of the perspective markers, this module can understand the relationships between spatial description features, thereby enhancing the information of different perspectives of the 3D object features and better aligning the cross - modal perspective information;
[0009] S5: A cross - modal decoder is designed. This module processes the spatial description features and 3D object features with perspective information based on the self - attention mechanism and the cross - attention mechanism. By using the 3D object features with enhanced perspective information as queries, the self - attention weights are assigned using the self - attention mechanism. Then, the spatial description features are used as keys and values, and the cross - attention mechanism is used to fuse these two modal features of spatial description features and 3D object features. The attention weights are dynamically adjusted according to the relationship between object features and description features, thereby optimizing the influence of perspective markers on each other between spatial description features and 3D object features, strengthening multi - modal information interaction, and finally generating fusion features for prediction;
[0010] S6: The obtained fusion features are classified and predicted through a classification head, thereby calculating the probability of each object in the scene and selecting the object with the maximum probability as the final localization result.
[0011] Furthermore, in step S2, a pre - trained label classifier is used to obtain the target labels mentioned in the text for the text. Then, known label sets are used to determine other anchor labels. Suitable prompts are constructed using the obtained text, target labels, and anchor labels and input into the large language model to generate a set of simple sentences as output.
[0012] Furthermore, the specific steps of step S3 are as follows:
[0013] S31: For the input 3D scene point cloud data, divide it into objects composed of multiple 3D points, and use a 3D object encoder to extract features to obtain 3D object features. The 3D object encoder performs point cloud sampling operations at different scales, by setting the radius Radius of the local neighborhood around each point i and the number of samples Samples i . The 3D object encoder will select a set number of neighborhood points. Different layers in the 3D object encoder will use different radii and numbers of neighborhood points for sampling, in order to capture the local features of the object at different scales. After sampling in each layer, the 3D object encoder will aggregate the features of the selected neighborhood points, using the max pooling Max Pooling operation to retain the most significant feature information in the local area. The aggregated features will be input into a multi-layer perceptron MLP for feature transformation. The output of the MLP will be processed by the ReLU activation function and the batch normalization Batch Normalization layer to enhance the expressive power of the features. Finally, after multiple samplings and MLP transformations, the 3D object encoder will output the high-dimensional object features F of each point or the entire point cloud obj , which can effectively capture the geometric shape and spatial distribution information of the object;
[0014] F obj = MLP(Pooling(P i ,Radius i ,Samples i ))
[0015] In the formula, P i represents the 3D scene point cloud data, and i represents that the data is in the i-th layer of the 3D object encoder;
[0016] S32: After obtaining the simplified spatial description, remove the redundant or semantically irrelevant parts, and retain the decoupled statement S that best matches the target and the anchor dec . Similarly, S dec will use a text encoder to extract the corresponding spatial description features F text :
[0017] F text = Encoder(S,S dec )
[0018] In the formula, S represents the preprocessed text; F text is the feature representation of the entire spatial description, containing the context information of each spatial description corresponding token.
[0019] Furthermore, in step S4, in order to introduce the perspective information, a trainable perspective marker is designed where Indicates that the view marker is a real number matrix, N view represents the number of views, d inner represents the hidden layer dimension of each view marker. For the 3D object feature F obj , this view marker is extended so that its matrix is extended to Concatenate the object features and view markers corresponding to the view along the marker number dimension, and finally obtain a combined feature with a shape of where N obj represents the number of objects in the scene point cloud, so as to ensure that the spatial description features under each corresponding view are aligned with the 3D object features under each view one by one. After the specific concatenation operation, the object feature Z with view information is obtained obj It is expressed as:
[0020] Z obj = concat(F obj , T vn )
[0021] In order to introduce view information into the spatial description features, a multi-modal feature fusion method based on the cross-attention mechanism, namely the Cross-Attention model, is designed. First, initialize a Cross-Attention model, which has a hidden layer dimension of d inner , k attention heads, and is expressed as:
[0022] Cross-Attention(d model = d inner , N heads = k)
[0023] In the formula, d model represents the dimension of the hidden layer, that is, the vector size of each layer inside the model; N heads represents the number of attention heads. Each head in the mechanism will calculate the attention separately, and then concatenate the results together to capture more features; in this mechanism, first, the spatial description features are obtained through a linear transformation to get the query vector Q, and these query vectors are used to match other information. In the Cross-Attention model, Q is obtained from the spatial description feature F text through the linear transformation matrix W Q as follows:
[0024] Q = F text W Q
[0025] Then, the view markers are respectively mapped to the key vector K and the value vector V, that is, calculated respectively through the linear transformation matrices W K and W V as follows:
[0026] K = T view W K , V = T view W V
[0027] In this mechanism, Q represents the representation of spatial description features in the high-dimensional space, K represents the content of perspective information, and V contains information associated with K; then, by calculating the similarity between Q and K, an attention score is obtained, and the attention matrix A is obtained through softmax normalization; after scaling the dot product result of Q and K, the attention weight between each object feature and the perspective information is calculated:
[0028]
[0029] Finally, the attention matrix A is used to weight the value vector V to obtain the spatial description feature Z that fuses the perspective labels text :
[0030] Z text = AV.
[0031] Furthermore, in step S5, the object feature Z obj is adjusted in dimension through a transpose operation to match the requirements of the cross-modal decoder:
[0032]
[0033] where Z o is the object feature after the transpose operation;
[0034] To further fuse the information of the two modalities of object features and spatial description features, first, through the self-attention mechanism, the object feature Z o that fuses the perspective labels is weighted as a query, that is, the self-attention mechanism simultaneously uses Z o as the key and value to complete the attention weight distribution between the query and finally obtains the attention-weighted object feature Z self , while in the cross-attention mechanism, the calculation process of the attention mechanism uses Z self as the query, Z text as the key and value to distribute the attention weight, and finally obtains the fusion feature Z cross by weighting the object feature with the attention matrix obtained from the spatial description feature, and the specific formula is as follows:
[0035]
[0036] Furthermore, in step S6, first, the fusion feature Z cross is subjected to perspective aggregation to integrate multiple perspectives to generate a unified aggregated feature Z agg, fully combining the information provided by the perspective markers to provide a comprehensive object feature, specifically by taking the fused feature Z cross divided by the number of perspectives N view to perform weighted average summation, and then for Z cross take the maximum value in the perspective dimension, and add the two to obtain the final Z agg , and the specific calculation process is as follows:
[0037]
[0038] In order to predict Z agg and output a classification result of the scene object, first, use a neural network classification head CLF with two fully connected layers to process the aggregated feature, generate the corresponding probability distribution LOGITS, and then select the category corresponding to the maximum probability from this probability distribution as the target object to be located:
[0039] LOGITS = CLF(Z agg )
[0040] Among them, the parameters of the neural network classification head CLF are optimized by calculating the difference between the predicted probability distribution LOGITS and the target values TARGETS, that is, using binary cross-entropy loss BCE as the loss function, and the final loss function L refer is expressed as:
[0041] L refer = BCE(LOGITS, TARGETS).
[0042] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0043] 1. By proposing learnable perspective markers, the present invention can effectively solve the inconsistency between different perspectives and spatial descriptions, enhance the feature alignment between different modalities, and thus improve the accuracy and effect in 3D visual object localization.
[0044] 2. By introducing a simple relationship decoupling module, which can decouple the positional relationship between the target and the anchor point in complex texts, enabling the large language model to extract more concise and effective spatial description features, thereby improving the understanding and processing ability of complex texts.
[0045] 3. The present invention can be widely applied to multiple fields involving the fusion of text and visual modalities, not limited to 3D visual object localization, but also extended to other cross-modal tasks in computer vision and natural language processing, with good generality and scalability.
[0046] In summary, by introducing innovative perspective marking and simple relationship decoupling modules, the present invention effectively improves the alignment between text and visual modalities, enhances the fusion and complementarity of cross-modal features. Using decoupling technology, the present invention simplifies the processing of complex texts, reduces the training complexity, and significantly improves the accuracy and robustness of 3D visual object localization by optimizing the consistency between perspectives and texts. In addition, while retaining the information association between modalities, the present invention promotes the efficient learning of cross-modal tasks, and has high adaptability and application potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the present invention.
[0048] Figure 2 It is a schematic diagram of the simple relationship decoupling module.
[0049] Figure 3 It is an architecture diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0051] As Figures 1 to 3 shown, the present invention discloses a 3D visual object localization method based on perspective information and relationship decoupling, using PointNet++ and Bert as the basic models for feature extraction, and the specific situation is as follows:
[0052] 1) Modal feature encoding:
[0053] For the extraction of 3D object features, the PointNet++ deep learning model is used as the 3D object encoder to hierarchically learn local and global features to extract geometric information. This method can effectively capture local details and long-range features to obtain 3D object features.
[0054] For the extraction of spatial description features, the Bert model is used as the text encoder, and the bidirectional self-attention mechanism is used to deeply understand the context relationship of text units, thereby improving the accuracy of spatial description feature extraction. Therefore, it is selected to obtain spatial description features.
[0055] 2) When dealing with complex spatial descriptions, especially in multimodal tasks, the degree of association and coupling between text and 3D object information is relatively high, and there are multiple dependencies. Traditional text processing methods usually regard each component in a sentence, such as the subject, object, modifier, etc., as a tightly coupled whole, which makes it difficult for the model to effectively capture the relative position relationship between each component, thus affecting the accuracy and robustness of the task. To overcome this problem, the present invention introduces a simple relationship decoupling module, which uses the powerful semantic understanding ability of the large language model to decouple the spatial position relationship words between the anchor points and targets in complex texts into more concise statements. Through decoupling, the method of the present invention can clearly understand the spatial position relationship existing between each target and anchor point. The specific simple relationship decoupling module is shown in Figure 2 as follows, and it includes the following steps:
[0056] 2.1) Parse the processed text description S to identify the targets in the text. The targets Tar and anchor points Anc in the text can be extracted through a pre-trained target classifier Classfier.
[0057] Tar, Anc = Classfier(S)
[0058] 2.2) After extracting the targets and anchor points, the next step is to construct a prompt Prompt. The Prompt is the input passed to the large language model, which can help the model understand the task objective and generate corresponding outputs. The design of the prompt aims to clearly express the relationship between the target and the anchor point and ensure that the large language model can process it effectively. According to the target, anchor point, and text, construct the following Prompt format:
[0059] Prompt = f(Tar, Anc, S)
[0060] where f is the function for constructing the prompt, Tar and Anc are the target and anchor point respectively, and S is the text description. For example, for the sentence "The cat is on the mat" with the target "cat" and the anchor point "mat", the constructed prompt is: According to the texts: {S}. What is the relationship between {Tar} and {Anc}, only output the result and here is the example: The desk is next to the chair.
[0061] 3) Feature extraction is performed on the preprocessed 3D scene point cloud data. Multiple point cloud objects are obtained from the 3D scene point cloud, and feature extraction is carried out through a 3D object encoder to obtain 3D object features. At the same time, multiple groups of simplified spatial descriptions are converted into multiple groups of spatial description features through a text encoder. The specific steps are as follows:
[0062] S31: For the input 3D scene point cloud data, it is divided into objects composed of multiple 3D points. Feature extraction is carried out using a 3D object encoder to obtain 3D object features. The 3D object encoder will perform point cloud sampling operations at different scales. By setting the radius Radius i of the local neighborhood around each point i and the sampling quantity Samples obj , the 3D object encoder will select a set number of neighborhood points. Different layers in the 3D object encoder will use different radii and numbers of neighborhood points for sampling in order to capture the local features of the object at different scales. After sampling at each layer, the 3D object encoder will aggregate the features of the selected neighborhood points, using the max pooling Max Pooling operation to retain the most significant feature information in the local area. The aggregated features will be input into a multi-layer perceptron MLP for feature transformation. The output of the MLP is processed through a ReLU activation function and a batch normalization Batch Normalization layer to enhance the expressive ability of the features. Finally, after multiple samplings and MLP transformations, the 3D object encoder will output the high-dimensional object feature F
[0063] F obj = MLP(Pooling(P i , Radius i , Samples i ))
[0064] In the formula, P i represents the 3D scene point cloud data, and i represents that the data is at the i-th layer of the 3D object encoder;
[0065] S32: After obtaining the result returned by the large language model (i.e., the simplified spatial description), through a sentence matching algorithm, the information most relevant to the text is screened out, redundant or semantically irrelevant parts are removed, and the decoupled statement S dec that best matches the target and the anchor point is retained. Similarly, S dec will also use a text encoder to extract the corresponding spatial description feature F text :
[0066] F text = Encoder(S, S dec )
[0067] In the formula, F text is the feature representation described in the entire space, which contains the context information of the corresponding markers for each space description.
[0068] 4) A perspective information transfer module is designed. This module combines the learnable multi-perspective markers with the spatial description features using the cross-attention mechanism, and at the same time transfers the corresponding perspective markers to the 3D object features under the corresponding perspectives. According to the learnable characteristics of the perspective markers, this module can understand the relationships between the spatial description features, thereby enhancing the information of different perspectives of the 3D object features and better aligning the cross-modal perspective information.
[0069] The description of the object's position is different in scenes with different perspectives. The model needs to be able to distinguish and integrate information from different perspectives. Perspective information can help the model determine from which direction a specific image comes, thus making the visual features more consistent. The present invention proposes learnable perspective markers. First, for each object m ∈ {1, 2,..., N view}}, there is a corresponding learnable perspective marker The perspective marker will be extended according to the perspective from which the object is observed:
[0070]
[0071] In the formula, T vn represents the extended perspective marker. The corresponding perspective markers are all aligned one-to-one with the object features under each perspective. After concatenation and transpose operations, the object feature Z with perspective information is obtained o which is expressed as:
[0072] Z obj = concat(F obj , T vn )
[0073]
[0074] To introduce the perspective information into the spatial description features, a multi-modal feature fusion method based on the cross-attention mechanism, namely the Cross-Attention model, is designed. First, a Cross-Attention model is initialized. This model has a hidden layer dimension of d inner , and k attention heads, which is expressed as:
[0075] Cross-Attention(d model = d inner , N heads = k)
[0076] In the formula, dmodel denotes the dimension of the hidden layer, that is, the vector size of each layer inside the model; N heads denotes the number of attention heads. Each head in the mechanism calculates attention separately and then concatenates the results to capture more features. In this mechanism, the spatial description features are first linearly transformed to obtain query vectors Q, which are used to match other information. In the Cross-Attention model, Q is obtained from the spatial description features F text through the linear transformation matrix W Q to get:
[0077] Q = F text W Q
[0078] Next, the perspective tokens are respectively mapped to key vectors K and value vectors V, that is, calculated respectively through the linear transformation matrices W K and W V to get:
[0079] K = T view W K , V = T view W V
[0080] In this mechanism, Q represents the representation of the spatial description features in the high-dimensional space, while K represents the content of the perspective information, and V contains the information associated with K. Then, by calculating the similarity between Q and K, the attention scores are obtained and normalized by softmax to get the attention matrix A. After scaling the dot product result of Q and K, the attention weights between each object feature and the perspective information are calculated:
[0081]
[0082] Finally, the attention matrix A is used to weight the value vector V to obtain the spatial description features Z that fuse the perspective tokens text :
[0083] Z text = AV
[0084] 5) A cross-modal decoder is designed. This module processes spatial description features with perspective information and 3D object features based on the self-attention mechanism and the cross-attention mechanism. By using the 3D object features strengthened with perspective information as queries, the self-attention mechanism assigns self-attention weights, and then takes the spatial description features as keys and values. The cross-attention mechanism is used to fuse these two modal features, namely the spatial description features and the 3D object features, and dynamically adjusts the attention weights according to the relationship between the object features and the description features, thereby optimizing the influence of the perspective markers on the spatial description features and the 3D object features on each other, strengthening the multi-modal information interaction, and finally generating the fused features for prediction.
[0085] To further fuse the information of these two modalities, namely the object features and the spatial description features, first, through the self-attention mechanism, the object features Z o with fused perspective markers are weighted as queries, that is, the self-attention mechanism simultaneously takes Z o as keys and values to complete the attention weight assignment between the queries, and finally obtains the attention-weighted object features Z self , and in the cross-attention mechanism, the calculation process of the attention mechanism assigns attention weights by taking Z self as queries, Z text as keys and values, and finally obtains the fused features Z cross by weighting the object features with the attention matrix obtained from the spatial description features. The specific formula is as follows:
[0086]
[0087] 6) The obtained fused features are classified and predicted through a classification head, thereby calculating the probability of each object in the scene, and selecting the object with the maximum probability as the final localization result, as follows:
[0088] First, perform perspective aggregation on the fused features Z cross to generate a unified aggregated feature Z agg by combining multiple perspective information, fully integrating the perspective information to provide a comprehensive object feature. Specifically, divide the fused features Z cross by the number of perspectives N view for weighted average summation, and then take the maximum value of Z cross in the perspective dimension. Therefore, Z agg The specific calculation process is as follows:
[0089]
[0090] To be able to perform operations on Z aggMake a prediction and output the classification prediction result of a scene object. First, use a neural network classification head CLF with two fully connected layers to process the aggregated features and generate the corresponding probability distribution LOGITS. Then, select the category corresponding to the maximum probability from this probability distribution as the target object to be located:
[0091] LOGITS = CLF(Z agg )
[0092] Among them, the parameters of the classification head are optimized by calculating the difference between the predicted probability distribution LOGITS and the target values TARGETS. That is, the binary cross-entropy loss BCE is used as the loss function, and the final loss function L refer is expressed as:
[0093] L refer = BCE(LOGITS, TARGETS)
[0094] The Nr3D and Sr3D datasets are designed for studying spatial language understanding in 3D environments and are mainly used for human-robot interaction and visual scene understanding tasks. The Nr3D dataset contains 45,503 human utterances, annotating 707 indoor scenes from the ScanNet dataset, which are subdivided into 76 fine-grained object types. Each scene contains at most 6 distractors, which requires the model to accurately identify the target among similar objects. The dataset is divided into Easy and Hard according to the difficulty of the task, and also includes two divisions: View Dependent and View Independent, with the latter distinguishing whether the description is affected by the speaker's perspective. Through this design, Nr3D aims to challenge the model's ability to handle multiple objects, spatial relationships, and view dependence in complex scenes.
[0095] In contrast, the Sr3D dataset provides a more simplified structure, containing 83,572 sentences based on the "target - space - relationship - anchored object" template. This structure makes the task simpler, reducing the complexity of the anchored parameters and descriptions, while still retaining the presence of distractors, enhancing the challenge of the task. Like Nr3D, Sr3D also has two difficulty divisions: Easy and Hard, and two view dependence divisions: View Dependent and View Independent. Both datasets aim to study how to effectively describe and locate objects and their relationships in 3D environments through natural language, promoting the language reasoning ability of agents to understand complex spatial layouts and interactions. These datasets provide valuable empirical data and test benchmarks for the fields of robot vision and artificial intelligence.
[0096] To verify the effectiveness of the method of the present invention, the accuracy ACC was used as the evaluation criterion. On these two commonly used datasets, Nr3D and Sr3D, a comparative analysis was respectively conducted with the methods of Overall, Easy, Hard, ViewDep, ViewInd, SAT, MVT, ViewRefer, MiKASA, and CoT3DRef. The experimental results are shown in Table 1.
[0097] Table 1 Analysis Table of Experimental Results
[0098]
[0099]
[0100] The experimental results can fully prove the effectiveness of the method of the present invention. The method of the present invention introduces learnable view markers and effectively integrates view information into 3D object features and spatial description features. By aggregating scene information from multiple views and aligning object features with description features, the method of the present invention can automatically learn and focus on the differences and correlations between different views and corresponding sentences, thereby having stronger adaptability and accuracy when processing multi-view data. In this way, the method of the present invention can dynamically adjust its language and visual representation space to better understand the semantic information of objects and scenes under different views. In addition, in order to improve the understanding of complex texts by the method of the present invention, the simple relationship decoupling module decouples the positional information relationship between the anchor point and the target in the text into simpler and easier-to-process components. The introduction of this module simplifies the computational complexity of the model when processing complex texts and enhances its ability to model fine-grained spatial relationships in the text. In this way, the method of the present invention can not only more accurately understand the target position in the text, but also effectively distinguish various spatial relationships and the dependency relationships in conditional sentences, thereby significantly improving the overall performance of the task and the response ability to complex texts.
[0101] Experimental conclusion: When conducting experiments using the Nr3D dataset, the method of the present invention performs superiorly among all existing methods. For example, under the same training conditions, the present invention has a 5.2% improvement in the overall accuracy compared to the optimal competing method CoT3DRef. The experimental results under the view-dependent conditions show that the present invention has a 6.7% improvement compared to CoT3DRef. These results indicate that the view markers effectively guide the method of the present invention to dynamically adjust the language and visual representation space, and by decoupling sentences, it shows that the present invention can identify whether there is view dependence in the sentence when understanding multiple spatial relationships. The present invention effectively solves the problem of how to apply in the multi-view statement object localization task, has good application prospects, and is worthy of promotion.
[0102] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A 3D visual object positioning method based on view information and relationship decoupling, characterized in that: The following steps are involved: S1: Acquire multimodal data and preprocess the data to obtain 3D scene point cloud data of uniform size and text with punctuation removed; S2: A simple relation decoupling module is designed to decouple the spatial relations of the preprocessed text. Based on the spatial position relation words in the text and with the help of the understanding ability of the large language model, the originally complex text is converted into multiple sets of simplified spatial descriptions of points, relations and targets. S3: extract features from the preprocessed 3D scene point cloud data, obtain multiple point cloud objects from the 3D scene point cloud and extract features through a 3D object encoder to obtain 3D object features, and convert multiple groups of simplified spatial descriptions into multiple groups of spatial description features through a text encoder; S4: We designed a view information transfer module, which combines learnable multi-view tags with spatial description features using a cross-attention mechanism and transfers the corresponding view tags to the 3D object features under the corresponding view. Based on the learnable characteristics of view tags, this module can understand the relationship between spatial description features, thereby enhancing the information of different viewpoints of 3D object features and better aligning cross-modal view information. S5: A cross-modal decoder is designed. This module processes spatial description features and 3D object features with view information based on self-attention mechanism and cross-attention mechanism. The 3D object features with enhanced view information are used as queries, and self-attention weights are assigned using self-attention mechanism. Then, the spatial description features are used as keys and values, and the cross-attention mechanism is used to fuse the two modal features of spatial description features and 3D object features. The attention weights are dynamically adjusted according to the relationship between object features and description features, thereby optimizing the influence of view labels on spatial description features and 3D object features, strengthening multimodal information interaction, and finally generating fused features for prediction. S6: The obtained fusion features are classified and predicted by the classification head to calculate the probability of each object in the scene, and the object with the highest probability is selected as the final positioning result.
2. The 3D visual object positioning method based on view information and relationship decoupling according to claim 1, characterized in that: In step S2, a pre-trained label classifier is used to obtain the target label mentioned in the text, and then other anchor labels are determined using a known label set. The obtained text, target label, and anchor label are used to construct a suitable prompt, which is input into the large language model to generate a set of simple sentences as output.
3. The 3D visual object positioning method based on view information and relationship decoupling according to claim 2, characterized in that: The specific steps of step S3 are as follows: S31: For the input 3D scene point cloud data, it is divided into objects composed of multiple 3D points, and the 3D object encoder is used to extract features to obtain 3D object features. The 3D object encoder performs point cloud sampling operations at different scales by setting the radius of the local neighborhood around each point. i and the number of samples i , the 3D object encoder will select a set number of neighborhood points. Different layers in the 3D object encoder will use different radii and numbers of neighborhood points for sampling in order to capture local features of objects at different scales. After sampling at each layer, the 3D object encoder will aggregate the selected neighborhood point features and use the maximum pooling operation to retain the most significant feature information in the local area. The aggregated features will be input into the multi-layer perceptron MLP for feature transformation. The output of the MLP is processed by the ReLU activation function and the batch normalization layer to enhance the feature expression ability. Finally, after multiple sampling and MLP transformation, the 3D object encoder will output high-dimensional object features F for each point or the entire point cloud. obj , can effectively capture the geometric shape and spatial distribution information of objects; F obj =MLP(Pooling(P i ,Radius i ,Samples i )) Where P i Represents 3D scene point cloud data, i means the data is in the i-th layer of the 3D object encoder; S32: After obtaining the simplified spatial description, remove the redundant or semantically irrelevant parts and retain the decoupled statement S that best matches the target and anchor point dec , the same S dec The text encoder will be used to extract the corresponding spatial description feature F text : F text =Encoder(S,S dec ) In the formula, S represents the preprocessed text; F text It is the feature representation of the entire spatial description, which contains the context information of each tag corresponding to the spatial description.
4. The 3D visual object positioning method based on view information and relationship decoupling according to claim 3, characterized in that: In step S4, in order to introduce the viewpoint information, a trainable viewpoint marker is designed in The view mark is a real number matrix, N view Represents the number of viewing angles, d inner Represents the hidden layer dimension of each view label, for 3D object features F obj , the perspective mark is expanded so that its matrix is expanded to The object features and viewpoint marks of the corresponding viewpoints are concatenated along the dimension of the number of marks, and finally a shape of The combination of features, where N obj Represents the number of objects in the scene point cloud, so as to ensure that the spatial description features under each corresponding perspective are aligned one by one with the 3D object features under each perspective. After the specific splicing operation, the object feature Z with perspective information is obtained. obj It is expressed as: Z obj =concat(F obj ,T vn ) In order to introduce perspective information into the spatial description features, a multimodal feature fusion method based on the cross-attention mechanism, namely the Cross-Attention model, is designed. First, a Cross-Attention model is initialized, which has a hidden layer dimension of d. inner , k attention heads, expressed as: Cross-Attention(d model =d inner ,N heads =k) Where, d model Represents the dimension of the hidden layer, that is, the vector size of each layer within the model; N heads Indicates the number of attention heads. Each head in the mechanism calculates attention separately, and then concatenates the results together to capture more features. In this mechanism, the spatial description features are first transformed through a linear transformation to obtain the query vector Q, which is used to match other information. In the Cross-Attention model, Q is composed of the spatial description features F text Through the linear transformation matrix W Q get: Q=F text W Q Next, the view tags are mapped to the key vector K and the value vector V, respectively, by the linear transformation matrix W K and W V The calculation results are: K=T view W K ,V=T view W V In this mechanism, Q represents the representation of spatial description features in high-dimensional space, K represents the content of perspective information, and V contains information associated with K. Then, by calculating the similarity between Q and K, the attention score is obtained, and the attention matrix A is obtained by softmax normalization. After the dot product result of Q and K is scaled, the attention weight between each object feature and perspective information is calculated: Finally, the attention matrix A is used to weight the value vector V to obtain the spatial description feature Z that integrates the view mark text : Z text =OFF。 5. The 3D visual object positioning method based on view information and relationship decoupling according to claim 4, characterized in that: In step S5, the object feature Z obj The transposition operation adjusts its dimensions to match the requirements of the cross-modal decoder: In the formula, Z o is the feature of the object after the transposition operation; In order to further integrate the information of the two modalities of object features and spatial description features, the self-attention mechanism is first used to integrate the object features Z marked by the fusion view. o As a query, the self-attention mechanism simultaneously regards Z o As the attention weight distribution between key, value completion and query, the object feature Z with weighted attention is finally obtained self , while in the cross attention mechanism, the calculation process of the attention mechanism is through Z self As a query, Z text As keys and values, attention weights are assigned, and finally the fusion feature Z is obtained by weighting the object features with the attention matrix obtained by the spatial description feature. cross , the specific formula is as follows:
6. The 3D visual object positioning method based on view information and relationship decoupling according to claim 5, characterized in that: In step S6, firstly, the fusion feature Z cross Perform perspective aggregation and integrate multiple perspectives to generate a unified aggregation feature Z agg , fully combines the information provided by the viewpoint mark to provide a comprehensive object feature. Specifically, the fusion feature Z cross Divide by the number of viewing angles N view Perform weighted average summation and then perform Z cross Find the maximum value in the viewing angle dimension, and add the two together to get the final Z agg , the specific calculation process is as follows: In order to Z agg To make a prediction and output the classification result of a scene object, first, a neural network classification head CLF containing two fully connected layers is used to process the aggregated features to generate the corresponding probability distribution LOGITS. Then, the category corresponding to the maximum probability is selected from the probability distribution as the target object to be located: LOGITS=CLF(Z agg ) Among them, the parameters of the neural network classification head CLF are optimized by calculating the difference between the predicted probability distribution LOGITS and the target value TARGETS, that is, the binary cross entropy loss BCE is used as the loss function, and the final loss function L refer It is expressed as: L refer =BCE(LOGITS,TARGETS)。
Citation Information
Patent Citations
Three-dimensional vision-text positioning method and system based on multiple visual angles and multiple texts
CN117197439A
Referring target detection and positioning method based on dynamic adaptive reasoning
WO2024037664A1