Open vocabulary three-dimensional object accessibility positioning method based on thinking chain reasoning and cross-modal fusion
Through the method based on thinking chain reasoning and cross-modal fusion, the positioning accuracy and robustness problems of the existing technology when dealing with categories and diversified scenarios are solved, and the precise positioning of three-dimensional objects and the generalization ability of the model is achieved.
Patent Information
- Application Number
- CN202510331144.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-27
AI Technical Summary
When handling the availability of open vocabulary three-dimensional objects, it is difficult to effectively process unknown categories and diversified scenarios outside the training data, resulting in poor positioning accuracy and robustness, and lack of adaptability and generalization capabilities.
Using a method based on thinking chain reasoning and cross-modal fusion, a multi-head affordable thinking chain reasoning strategy is constructed by fine-tuning the multi-modal large language model, and a cross-modal adaptive fusion module is used to fuse images, text and point cloud features to achieve accurate positioning of the affordability of three-dimensional objects.
It improves the generalization ability and adaptability of the model, can achieve higher positioning accuracy and robustness in unseen categories and diversified scenarios, form a structured dictionary of affordability knowledge, and enhances the model's inference efficiency.
Smart Images

Figure CN120219718A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and specifically relates to an open-vocabulary 3D object affordance localization method based on chain of thought reasoning and cross-modal fusion. Background Art
[0002] Open-vocabulary 3D object affordance localization aims to locate the "action possibilities" on objects, including visible and invisible scenes, and identify specific regions on the objects that support specific interactions. This connects the visual perception and physical operation of embodied agents and has rich application scenarios, such as robot operation, scene understanding, action anticipation, and imitation learning.
[0003] Existing methods all establish explicit mappings between semantic affordance categories and geometric structures, are limited to predefined affordance categories, and cannot locate object affordances in categories outside the training categories. Although some studies explore the localization of object affordances through additional instructions, including combining interactive images or languages to introduce external interaction priors and alleviate the generalization gap caused by affordance diversity, they are still vulnerable to the limited semantic space and cannot effectively handle unseen categories outside the training data. Affordances themselves have diversity, and existing methods are insufficient in dealing with this diversity, resulting in poor localization accuracy and robustness in complex or diverse scenarios. These methods lack the adaptive ability to new categories and cannot dynamically expand the semantic space or learn new affordance categories during training, limiting their wide applicability in practical applications. Therefore, there is an urgent need for new methods and technologies to improve the generalization ability and adaptability of the model to better cope with the challenges of unseen categories and diverse scenarios. Summary of the Invention
[0004] The present invention is proposed to solve the above-mentioned deficiencies of the existing technology, and provides an open-vocabulary 3D object affordance localization method based on chain of thought reasoning and cross-modal fusion, aiming to fully exploit the affordance knowledge inferred by the multi-modal large language model and perform cross-modal adaptive fusion of features to achieve accurate localization of 3D object affordances.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] An open-vocabulary 3D object affordance localization method based on chain of thought reasoning and cross-modal fusion of the present invention is characterized by including the following steps:
[0007] Step 1, obtain the preprocessed image-point cloud pair , where represents the point cloud of the object and is composed of the coordinates of the object and the affordance label . Represents an interactive image, represents the number of points in the point cloud , and respectively represent the height and width of the interactive image ;
[0008] Encode the interactive image and the point cloud respectively, and obtain the image feature and the point cloud feature , and reshape it into a new image feature ; represents the total number of spatial positions of the new image feature, and ; and respectively represent the height and width of the image feature; represents the number of sampled points, represents the number of channels;
[0009] Step 2: Construct and train an affordance model based on a fine-tuned multimodal large language model ;
[0010] Step 3: Construct a multi-head affordance thinking chain reasoning strategy based on the affordance model and use it to reason about the interactive image to obtain four affordance knowledge accordingly;
[0011] Step 4: The encoder encodes and fuses the four affordance knowledge obtained by reasoning to obtain the aligned object geometric knowledge feature and the aligned affordance intention knowledge feature ;
[0012] Step 5: Use the cross-modal adaptive fusion module to and as well as and for fusion, and obtain the features after text and point cloud fusion and the features after text and image fusion ;
[0013] Step 6: Use the decoder to and for decoding to obtain the three-dimensional affordance region ;
[0014] Step 7: Construct the total loss function of the three-dimensional object affordance localization network composed of an encoder, a cross-modal adaptive fusion module, and a decoder , and is used to train a three-dimensional object affordance localization network to obtain an optimal three-dimensional object affordance localization model, so as to realize the affordance localization of the input interactive image and point cloud.
[0015] Another feature of the open-vocabulary three-dimensional object affordance localization method based on chain-of-thought reasoning and cross-modal fusion according to the present invention is that the step 2 includes the following steps:
[0016] Step 2.1: Insert an adapter into each layer of the InternVL model to form a fine-tunable multi-modal large language model, where the parameters of the adapter are , where represents the left matrix of low-rank decomposition, represents the right matrix of low-rank decomposition, represents the rank, is the hidden layer dimension of each layer;
[0017] Step 2.2: Input the interactive image , and the text prompt into the multi-modal large language model for reasoning, and freeze the parameters of the InternVL model, and only train the adapter to obtain an affordance model, denoted as .
[0018] Furthermore, the step 3 includes the following steps:
[0019] Step 3.1: According to the reasoning of the invariant geometric attributes of the object, including: object interaction perception and geometric structure reasoning, design the first prompt "Point out which part of the object in the interactive image interacts with people"; design the second prompt "From the perspective of the geometric structure of the object, explain why this part of the object in the interactive image interacts with people";
[0020] Step 3.2: According to the reasoning of the potential interaction intention of the object, including: interaction detail description and interaction analogy reasoning, design the third prompt "Describe the interaction between the object and people in the interactive image"; design the fourth prompt "List other interactions that the object in the interactive image can have with people;
[0021] Step 3.3: Input the four designed prompts and the interactive image into the affordance model for reasoning, and correspondingly obtain four affordance knowledge, denoted as , , and .
[0022] Further, step 4 includes the following steps:
[0023] Step 4.1: Concatenate with to form object geometric knowledge, and concatenate with to form affordance intention knowledge;
[0024] Step 4.2: Use a text encoder to encode the object geometric knowledge and the affordance intention knowledge respectively, and obtain the object geometric knowledge feature and the affordance intention knowledge feature , where , represents the number of interactive objects and the number of interaction methods;
[0025] Step 4.3: Use equations (1) and (2) to obtain the aligned object geometric knowledge feature and the aligned affordance intention knowledge feature ;
[0026] (1)
[0027] (2)
[0028] In equations (1) and (2), is the cross-attention layer, and is the self-attention layer.
[0029] Further, step 5 includes the following steps:
[0030] Step 5.1: Represent and in the same feature space, so as to project into the first query , project into the first key and the first value , where , , are the 3 projection weights, , ; is the dimension of the projection;
[0031] Step 5.2: Use equation (3) to obtain the re-represented point cloud feature :
[0032] (3)
[0033] Step 5.3: For Redisplayed in the same feature space, so that is projected into a second query , and is projected into a second key and a second value , where , , are three projection weights, , ;
[0034] The geometric knowledge features of the object after redisplay are obtained by using Equation (4) :
[0035] (4)
[0036] Step 5.4, the fused point features are obtained by using Equation (5)
[0037] (5)
[0038] In Equation (5), represents two fully connected layers, represents the operation of flattening after pooling according to the dimension , represents concatenation, represents a convolutional layer with a kernel size of 1×1;
[0039] Step 5.5, upsample using Equation (6) to obtain the features after the fusion of text and point cloud :
[0040] (6)
[0041] In Equation (6), FP is a feature propagation layer;
[0042] Step 5.6, the features after the fusion of text and image are obtained by using Equation (7)
[0043] (7)
[0044] In Equation (7), represents a reshaping operation.
[0045] Furthermore, in Step 6, and are input into the decoder, and the predicted 3D object affordance is obtained by using Equation (8) and Equation (9)
[0046] (8)
[0047] (9)
[0048] In equations (8) and (9), represents the output head, represents the sigmoid function, represents the affordance feature.
[0049] Furthermore, in step 7, the total loss function is constructed using equation (10) :
[0050] (10)
[0051] In equation (10), represents the focal loss, represents the composition of the dice loss, and there is:
[0052] (11)
[0053] (12)
[0054] In equations (11) and (12), and are two hyperparameters of the focal loss, is a constant used to prevent division by zero.
[0055] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the open-vocabulary 3D object affordance localization method, and the processor is configured to execute the program stored in the memory.
[0056] A computer-readable storage medium according to the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the open-vocabulary 3D object affordance localization method.
[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0058] 1. The present invention proposes an affordance model based on fine-tuning a multi-modal large language model, which can utilize prior affordance knowledge to enhance the model's affordance understanding ability by adjusting LoRA while retaining the model weights, and avoids destroying the pre-trained knowledge of the multi-modal large language model, thereby improving the model's affordance understanding ability and generalization performance, and also achieving efficient parameter utilization and protection of pre-trained knowledge;
[0059] 2. The present invention proposes a multi-head affordance thinking chain reasoning strategy, which infers implicit invariant geometric properties and potential interaction intentions from a fine-tuned multi-modal large language model, and uses the attention mechanism to model the correlation between these interaction primitives to form an affordance knowledge dictionary, thereby improving the model's understanding ability of geometric properties and interaction intentions, and also forming a structured affordance knowledge dictionary, enhancing the model's generalization ability and reasoning efficiency;
[0060] 3. The present invention designs a cross-modal adaptive fusion module to better achieve cross-modal feature alignment and fusion of the geometric properties of an object and point cloud features, thereby making full use of the complementary information of image and point cloud data and enhancing the model's comprehensive understanding ability of the geometric properties and interaction intentions of the object. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a schematic diagram of the open-vocabulary 3D object affordance localization based on the thinking chain reasoning and cross-modal fusion of the present invention;
[0062] Figure 2 It is a schematic diagram of the experimental results of the point cloud affordance localization based on PIADv2 of the present invention;
[0063] Figure 3 It is a schematic diagram of the visualization results of the multi-angle analysis experiment based on PIADv2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] In this embodiment, an open-vocabulary 3D object affordance localization method based on thinking chain reasoning and cross-modal fusion comprehensively utilizes the hierarchical reasoning paradigm of the thinking chain and the cross-modal semantic integration ability of the adaptive fusion module, and designs a 3D open-vocabulary object affordance localization framework that can perform collaborative reasoning of geometric intentions based on interactive pictures. By fine-tuning the multi-modal large language model, it infers the inherent invariant geometric properties and potential interaction intentions of the object to form an affordance knowledge system, and fuses the knowledge with the point cloud features through a cross-modal adaptive fusion model to achieve accurate localization of the 3D object affordance. Specifically, the method includes the following steps:
[0065] Step 1. Obtain the input point cloud-image pair for encoding:
[0066] Step 1. Obtain the preprocessed image-point cloud pair , where, represents the point cloud of the object, and is composed of the coordinates of the object and the affordance label ; represents the interactive image, represents the point cloud of the number of points, is associated with respectively represent the interactive image height and width;
[0067] respectively encode the interactive image and the point cloud to obtain image features , point cloud features , and reshape into new image features ; represents the total number of spatial positions of the new image features, and ; and respectively represent the height and width of the image features; represents the number of sampled points, represents the number of channels.
[0068] Step 2. Construct and train an affordance model based on a fine-tuned multimodal large language model ;
[0069] Step 2.1. The multimodal large language model realizes semantic alignment and knowledge transfer of multi-dimensional information such as vision, language, and space through a cross-modal joint learning mechanism. Its core architecture is based on a cross-modal attention mechanism constructed by a deep neural network, which can map images, texts, and three-dimensional geometric features to a unified latent space to form a cross-modal semantic association network. At present, multimodal large language models focus on object recognition and description, but lack sufficient understanding of the actual uses of objects and their interactions with humans. Therefore, the present invention uses InternVL and an injectable learnable adapter to fine-tune the multimodal large language model, as shown in Figure 1 part (c); an adapter is inserted into each layer of the InternVL model to form a fine-tunable multimodal large language model, where the parameters of the adapter are , where represents the left matrix of the low-rank decomposition, represents the right matrix of the low-rank decomposition, represents the rank, is the hidden layer dimension of each layer;
[0070] Step 2.2. Input the interactive image , text prompt into the multimodal large language model for inference, and freeze the parameters of the InternVL model, and only train the adapter to obtain an affordance model, denoted as .
[0071] Step 3. Construct a multi-head affordance thinking chain inference strategy based on the affordance model and use it for the interactive image Perform reasoning to obtain four affordance knowledge accordingly;
[0072] Step 3.1, Reason according to the invariant geometric properties of the object, including: object interaction perception and geometric structure reasoning. The model needs to focus on understanding the interaction parts of specific objects in the image, refining its perception of the key parts of the object rather than the whole object, as shown in part (a) of Figure 1 ; Design the first prompt "Point out which part of the object in the interaction image interacts with people"; Due to similar geometric properties, similar regions of different objects can perform the same interaction. The model needs to reason about these features. This allows the model to go beyond the constraints of object categories and pay more attention to the relationship between structure and affordance; Design the second prompt "From the perspective of the geometric structure of the object, explain why this part of the object in the interaction image interacts with people."
[0073] Step 3.2, Reason according to the potential interaction intention of the object, including: interaction detail description and interaction analogy reasoning. The model needs to identify the entire interaction process between the object and people in the image, including the interaction parts between the object and people and the type of interaction; This allows the model to generate a fine-grained feature representation and capture the physical structure constraints of the interaction between people and objects, as shown in part (a) of Figure 1 ; Design the third prompt "Describe the interaction between the object and people in the interaction image"; In the human brain, after observing an object, it is usually associated with various potential interaction methods; Inspired by this, the present invention will use the world knowledge repository of the multimodal large model to explore other possible interaction intentions of the object, reduce the dependence on specific affordance instances, and enhance analogy reasoning; Design the fourth prompt "List other interactions that the object in the interaction image can have with people."
[0074] Step 3.3, Input the four designed prompts and the interaction image into the affordance model for reasoning, and obtain four affordance knowledge accordingly, denoted as 、 、 and in turn.
[0075] Step 4, The encoder encodes and fuses the four affordance knowledge obtained by reasoning to obtain the aligned object geometric knowledge feature and the aligned affordance intention knowledge feature :
[0076] Step 4.1, Combine with Stitch into object geometric knowledge, and with stitch into affordance intention knowledge;
[0077] Step 4.2: Use a text encoder to encode the object geometric knowledge and the affordance intention knowledge respectively, and obtain the object geometric knowledge feature and the affordance intention knowledge feature , where , represents the number of interactive objects and the number of interaction methods.
[0078] Step 4.3: Use Equation (1) and Equation (2) to obtain the aligned object geometric knowledge feature and the aligned affordance intention knowledge feature ;
[0079] (1)
[0080] (2)
[0081] In Equation (1) and Equation (2), is the cross-attention layer, is the self-attention layer.
[0082] Step 5: Use the cross-modal adaptive fusion module to fuse the object geometric knowledge feature and the point cloud feature , and directly fuse the affordance intention knowledge feature and the new image feature , and obtain the feature after text and point cloud fusion and the feature after text and image fusion respectively.
[0083] Step 5.1: To better promote the cross-modal fusion of the geometric attributes of the object interaction area and the point cloud feature, the present invention proposes that the cross-modal adaptive fusion module integrates the geometric attributes of the object into the deepest encoder layer of the point cloud encoder PointNet++, so as to refine the point feature map and achieve effective cross-modal feature alignment and fusion, as shown in Figure 1 part (b) of; and are re-represented in the same feature space, so as to project into the first query , project into the first key and the first value , where , , are 3 projection weights. , ; is the dimension of the projection.
[0084] Step 5.2: Obtain the re-represented point cloud features using Equation (3) :
[0085] (3)
[0086] Step 5.3: Re-represent under the same feature space according to the processes of Step 5.1 and Step 5.2, project into the second query , project into the second key and the second value , where , , are three projection weights, , .
[0087] Obtain the re-represented object geometric knowledge features using Equation (4) :
[0088] (4)
[0089] Step 5.4: Obtain the fused point features using Equation (5) :
[0090] (5)
[0091] In Equation (5), represents two fully connected layers, represents the operation of flattening after pooling and then along the dimension , represents concatenation, represents a convolutional layer with a kernel size of 1×1.
[0092] Step 5.5: Upsample using Equation (6) to obtain the features after text and point cloud fusion :
[0093] (6)
[0094] In Equation (6), FP is the feature propagation layer;
[0095] Step 5.6: Obtain the features after text and image fusion using Equation (7) :
[0096] (7)
[0097] In Equation (7), represents the reshaping operation.
[0098] Step 6: Use the decoder to and decode to obtain the three-dimensional affordance region ;
[0099] Input and into the input, and use Equations (8) and (9) to obtain the predicted three-dimensional object affordance :
[0100] (8)
[0101] (9)
[0102] In Equations (8) and (9), represents the output head, represents the sigmoid function, represents the affordance feature.
[0103] Step 7: Construct the total loss function of the three-dimensional object affordance localization network composed of an encoder, a cross-modal adaptive fusion module, and a decoder , and use it to train the three-dimensional object affordance localization network, so as to obtain the optimal three-dimensional object affordance localization model to realize the affordance localization of the input interactive image and point cloud; without being restricted by the affordance category label, the present invention focuses on the three-dimensional object affordance and the ground truth label and makes the model directly associate the three-dimensional object affordance with the interactive image through inference; therefore, the total loss consists of the focal loss and the dice loss, which supervise the point-level heat map.
[0104] Use Equation (10) to construct the total loss function :
[0105] (10)
[0106] In Equation (10), represents the focal loss, represents the dice loss, and there is:
[0107] (11)
[0108] (12)
[0109] In formulas (11) and (12), and are two hyperparameters of the focal loss, is a constant used to prevent division by zero.
[0110] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0111] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.
[0112] Embodiment:
[0113] To verify the effectiveness of the method of the present invention, three partitions of the dataset PIADv2 are selected for experiments in this embodiment. Four evaluation metrics, namely Area Under Curve (AUC), average Intersection OverUnion (aIOU), SIMilarity (SIM), and Mean Absolute Error (MAE), are used for the object affordance localization task. Among them, the smaller the MAE metric, the better, and the larger the other metrics, the better.
[0114] The proposed 3D object affordance localization network adopts the Pytorch architecture and is trained using two NVIDIA RTX 3090 GPUs. The optimizer uses Adam with momentum decay rates of and respectively. The model training process has a total of 65 epochs, and the initial learning rate is .
[0115] For a comprehensive comparison, 5 methods are selected in this embodiment for a performance comparison with the method of the present invention. The baseline method Baseline (Base.) directly connects the image features and point cloud features from the encoder. Two advanced image-point cloud cross-modal methods, FRCNN and XMF, are selected, as well as two state-of-the-art methods in the field of 3D object affordance localization, IAG and LASO.
[0116] Among them, the Baseline (Base.) method directly connects the features output by the image and point cloud encoders, and uses the output head for 3D object point cloud affordance localization without any intermediate steps to align the features; the FRCNN method addresses the challenges of object recognition and localization brought about by sparse point clouds in distant regions, and is a novel multi-modal two-stage method that effectively integrates the point cloud data and camera images within the region of interest, and adaptively combines the sparse lidar geometric information and camera images; the XMF method effectively fuses the features of two different modalities by combining self-attention and cross-attention mechanisms, and integrates the information of the two modalities into a local latent space; the IAG method utilizes the ability of humans to perceive object affordances in the physical world through demonstration images, and proposes a method for localizing 3D object affordances from 2D interactions in images, aligns the object region features from different sources, and decomposes the dynamic factors involved in affordance extraction into the interactions between the subject-object and object-scene, and realizes 3D object affordance localization through the context modeling of these interactions; the LASO method aims to localize 3D object affordances according to the questions proposed by experts, introduces a fusion module to identify the target affordance regions at different scales, and uses a set of affordance queries conditioned on language cues to generate dynamic kernels, and convolves these dynamic kernels with the point cloud features to achieve the localization of object affordances.
[0117] Comparing the above five methods with the method of the present invention, the experimental results are shown in Table 1:
[0118] Table 1 Comparison of the results of the method of the present invention and the 5 selected comparison methods on the PIADv2 dataset
[0119]
[0120] The results in Table 1 show that the 3D object affordance localization method proposed by the present invention is significantly superior to the existing methods in terms of key performance indicators, demonstrating the superiority and generalization of the method of the present invention.
[0121] Among them, PIADv2 has three partitions, including: Seen: The training set and the test set share the same objects and affordances; Unseen Object: The affordances are consistent between the training set and the test set, but some objects in the test set do not appear in the training set; Unseen Affordance: The affordances in the test set do not exist in the training set, and some object categories also do not exist; the experimental results show that the method of the present invention is better than the other 5 methods, thus demonstrating the feasibility of the method proposed by the present invention. The visualization results of the object point cloud are as Figure 2As shown. The present invention also conducts in-depth analysis experiments on a variety of objects, affordances, and examples. The visualization results of the experiments are respectively shown in parts (a), (b), and (c) of Figure 3 , and the results prove the robustness and generalization of the method proposed by the present invention.
Claims
1. An open vocabulary three-dimensional object affordance localization method based on thought chain reasoning and cross-modal fusion, characterized in that: The following steps are involved: Step 1: Get the preprocessed image-point cloud pair ,in, Represents the point cloud of the object, and consists of the coordinates of the object and affordance labels composition; Represents an interactive image, Representing point clouds The number of points, and Represents interactive images The height and width of For interactive images and point cloud Encode and obtain image features accordingly and point cloud features , and Reshape into new image features ; represents the total number of spatial locations of new image features, and ; and Respectively represent the height and width of the image feature; represents the number of points after sampling, Indicates the number of channels; Step 2: Build and train an affordance model based on a fine-tuned multimodal large language model ; Step 3: Build an Availability-Based Model The multi-head affordance thinking chain reasoning strategy is used to analyze interactive images. Perform reasoning and obtain four pieces of available knowledge accordingly; Step 4: The encoder encodes and fuses the four affordances acquired by reasoning to obtain the aligned object geometry knowledge features. and aligned affordance intention knowledge features ; Step 5: Use the cross-modal adaptive fusion module to and as well as and Fusion is performed to obtain the features of text and point cloud fusion. And the features after text and image fusion ; Step 6: Use the decoder to and Decode and get the three-dimensional availability area ; Step 7: Construct the total loss function of the 3D object affordance localization network consisting of the encoder, cross-modal adaptive fusion module and decoder , and is used to train the 3D object affordance localization network to obtain the optimal 3D object affordance localization model to achieve affordance localization of the input interactive image and point cloud.
2. The open vocabulary three-dimensional object affordance localization method based on thought chain reasoning and cross-modal fusion according to claim 1, characterized in that: The step 2 comprises the following steps: Step 2.1: Insert an adapter into each layer of the InternVL model to form a multimodal large language model that can be fine-tuned, where the parameters of the adapter are ,in, represents the left matrix of low-rank decomposition, represents the right matrix of the low-rank decomposition, represents the rank, is the hidden layer dimension of each layer; Step 2.2: Transform the interactive image , text prompt Input the multimodal large language model for inference, freeze the parameters of the InternVL model, and only train the adapter to obtain the affordance model, denoted as .
3. The open vocabulary three-dimensional object affordance localization method based on thought chain reasoning and cross-modal fusion according to claim 2, characterized in that: The step 3 comprises the following steps: Step 3.1: Design the first hint based on the invariant geometric properties of the object, including object interaction perception and geometric structure reasoning "Indicate which part of the object in the interactive image interacts with the person"; design the second prompt "From the perspective of the geometric structure of the object, explain why this part of the object in the interactive image interacts with the human"; Step 3.2: Design the third prompt based on the object’s potential interaction intention, including interaction detail description and interaction analogy reasoning "Describe the interaction between objects and people in interactive images"; Design the fourth prompt "List other interactions that objects in the interactive image can have with people;" Step 3.3: Design four prompts and interactive images Input to the Affordance Model Inference is performed in , and four pieces of affordance knowledge are obtained accordingly, which are recorded as , , and .
4. The open vocabulary three-dimensional object affordance localization method based on thought chain reasoning and cross-modal fusion according to claim 3, characterized in that: The step 4 comprises the following steps: Step 4.1: and Splicing into object geometry knowledge, and spliced into affordance intention knowledge; Step 4.2: Use the text encoder to encode the object geometry knowledge and affordance intention knowledge respectively, and obtain the object geometry knowledge features accordingly. and affordance intention knowledge features ,in, , Indicates the number of interactive objects and the number of interaction methods; Step 4.3: Use equations (1) and (2) to obtain the geometric knowledge features of the aligned objects. and aligned affordance intention knowledge features ; (1) (2) In formula (1) and formula (2), is the cross attention layer, is the self-attention layer.
5. The open vocabulary three-dimensional object affordance localization method based on thought chain reasoning and cross-modal fusion according to claim 4, characterized in that: The step 5 comprises the following steps: Step 5.1: and Re-express in the same feature space, so that Projection into the first query ,Will Projection as the first key and the first value ,in, , , are the 3 projection weights, , ; is the dimension of the projection; Step 5.2: Use formula (3) to get the re-expressed point cloud features : (3) Step 5.3: Re-express in the same feature space, so that Projected into the second query ,Will Projection as the second key and the second value ,in, , , are the 3 projection weights, , ; Using formula (4), we can get the geometric knowledge features of the re-expressed object: : (4) Step 5.4: Use formula (5) to get the fused point features : (5) In formula (5), represents two fully connected layers, Indicates that after pooling, the dimensions Perform the flattening operation. Indicates connection, represents a convolutional layer with a kernel size of 1×1; Step 5.5: Use formula (6) to Upsample to obtain the features after fusion of text and point cloud : (6) In formula (6), FP is the feature propagation layer; Step 5.6: Use formula (7) to get the features after text and image fusion : (7) In formula (7), Represents a reshape operation.
6. The open vocabulary three-dimensional object affordance localization method based on thought chain reasoning and cross-modal fusion according to claim 5, characterized in that: The step 6 is to and Input into the decoder, and use equations (8) and (9) to get the predicted 3D object availability : (8) (9) In formula (8) and formula (9), Indicates the output header, represents the sigmoid function, Represents affordance characteristics.
7. The open vocabulary three-dimensional object affordance localization method based on thought chain reasoning and cross-modal fusion according to claim 6, characterized in that: In step 7, the total loss function is constructed using formula (10): : (10) In formula (10), represents focal loss, Represents the dice loss composition, and has: (11) (12) In formula (11) and formula (12), and are the two hyperparameters of focal loss, is a constant used to prevent division by zero.
8. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the open vocabulary three-dimensional object availability positioning method as described in any one of claims 1-7, and the processor is configured to execute the program stored in the memory.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the open vocabulary three-dimensional object availability localization method described in any one of claims 1-7 are performed.
Citation Information
Cited By
Single-image three-dimensional character interaction generation method based on multi-modal deep learning
CN121788726A
Method for generating three-dimensional human interaction from a single image based on multi-modal deep learning
CN121788726B
Multi-agent-driven e-commerce dispute judgment method and system
CN121860647A
Reinforcement learning framework-based availability generalization reasoning method and system, computer equipment and medium
CN122021940A
Generative reasoning method, system, computer device and medium based on reinforcement learning framework
CN122021940B