A 3D result generation method, system, terminal and storage medium for hand-object interaction based on large language models

By preprocessing and text annotation of opponent-object interaction data, and using large language model training and optimization processing, high-quality 3D results of hand-object interaction are generated, solving the problem of insufficient optimization of hand-object interaction 3D models in the prior art, and achieving efficient and accurate interaction pose generation.

CN119861826BActive Publication Date: 2025-07-22GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510354110.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-22
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The prior art lacks an effective optimization mechanism after generating a 3D model of hand-object interaction, resulting in poor results of hand-object interaction and unable to meet user needs.

Method used

By obtaining hand-object interaction data for preprocessing and text annotation, a target 3D interaction data set is constructed, and a large language model is used for model training and fine-tuning, an interactive 3D generation model is generated, and the user description text is optimized to obtain optimized hand MANO parameters and generate high-quality hand-object interaction 3D results.

Benefits of technology

It realizes the rapid and accurate output of 3D results of hand-object interaction, improves the generation efficiency and quality of the model, ensures that the interactive posture conforms to physical laws and visual reality, and meets user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119861826B_ABST
    Figure CN119861826B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, terminal and storage medium for generating 3D results of hand-object interaction based on a large language model. The method includes: obtaining hand-object interaction data, and performing preprocessing and text annotation processing to obtain a target 3D interaction data set; determining a preset large language model, and training and fine-tuning the preset large language model according to the target 3D interaction data set to obtain an interaction 3D generation model; obtaining a user description text and inputting it into the interaction 3D generation model to output preliminary interaction 3D data; optimizing the hand parameters in the interaction 3D generation model according to the preliminary interaction 3D data to obtain optimized hand MANO parameters, and obtaining a target hand-object interaction 3D result according to the optimized hand MANO parameters. The present invention can quickly and accurately output the 3D result of hand-object interaction by constructing an interaction 3D generation model and optimizing the model through the user description text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method, system, terminal and computer-readable storage medium for generating 3D results of hand-object interaction based on a large language model. Background Art

[0002] A hand-object interaction 3D model refers to a model that can achieve natural and intuitive interaction between a user and a three-dimensional object in a virtual environment. Such a model can not only provide a rich visual experience, but also enhance the user's immersion and engagement through advanced interaction technologies.

[0003] However, in the prior art, after generating a hand-object interaction 3D model, there is a lack of an effective optimization mechanism to adjust the interaction posture between the user's hand and the object, resulting in poor 3D results of hand-object interaction generated in practical applications of the hand-object interaction 3D model and unable to meet the user's needs.

[0004] Therefore, the prior art still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method, system, terminal and computer-readable storage medium for generating 3D results of hand-object interaction based on a large language model, aiming to solve the problem that after generating a hand-object interaction 3D model, there is a lack of an effective optimization mechanism to adjust the interaction posture between the user's hand and the object, resulting in poor 3D results of hand-object interaction generated in practical applications of the hand-object interaction 3D model and unable to meet the user's needs.

[0006] To achieve the above object, the present invention provides a method for generating 3D results of hand-object interaction based on a large language model, and the method for generating 3D results of hand-object interaction based on a large language model includes the following steps:

[0007] Obtain hand-object interaction data, and perform preprocessing and text annotation processing on the hand-object interaction data to obtain a target 3D interaction data set;

[0008] Determine a preset large language model, and perform model training and model fine-tuning on the preset large language model according to the target 3D interaction data set to obtain an interactive 3D generation model;

[0009] Obtain a user description text, and input the user description text into the interactive 3D generation model to output preliminary interactive 3D data;

[0010] Optimize the hand parameters in the interactive 3D generation model according to the preliminary interactive 3D data to obtain optimized hand MANO parameters, and obtain a target hand-object interaction 3D result according to the optimized hand MANO parameters.

[0011] Optionally, in the method for generating 3D results of hand-object interaction based on a large language model, the steps of obtaining hand-object interaction data, preprocessing the hand-object interaction data, and performing text annotation processing to obtain a target 3D interaction dataset specifically include:

[0012] Obtain an open-source hand dataset, extract hand-object interaction data from the open-source hand dataset, and preprocess the hand-object interaction data to obtain a first 3D interaction dataset;

[0013] Perform text annotation processing on the hand-object interaction data to obtain a second 3D interaction dataset;

[0014] Merge and partition the first 3D interaction dataset and the second 3D interaction dataset to obtain a target 3D interaction dataset.

[0015] Optionally, in the method for generating 3D results of hand-object interaction based on a large language model, the hand-object interaction data includes an object 3D model and first hand MANO parameters;

[0016] The step of preprocessing the hand-object interaction data to obtain a first 3D interaction dataset specifically includes:

[0017] Perform model decimation, data filtering, and normalization on the object 3D model to obtain an integer object model;

[0018] Perform integerization on the first hand MANO parameters to obtain integer hand MANO parameters;

[0019] Merge the integer object model and the integer hand MANO parameters to obtain a first 3D interaction dataset.

[0020] Optionally, in the method for generating 3D results of hand-object interaction based on a large language model, the step of performing text annotation processing on the hand-object interaction data to obtain a second 3D interaction dataset specifically includes:

[0021] Perform multi-view rendering on the hand-object interaction data to obtain multiple 2D images of different views;

[0022] Perform text description processing on the multiple 2D images of different views to obtain multiple image description texts;

[0023] Extract image features and text features from the multiple image description texts, and calculate the similarity between the image features and the text features to obtain the target text description with the highest similarity in each image description text;

[0024] Construct a second 3D interaction dataset based on the target text description with the highest similarity in each of the image description texts.

[0025] Optionally, in the method for generating 3D results of hand-object interaction based on a large language model, the target 3D interaction dataset includes a training set and an evaluation set;

[0026] The determining of a preset large language model and the model training and fine-tuning of the preset large language model according to the target 3D interaction dataset to obtain an interactive 3D generation model specifically includes:

[0027] Determine a preset large language model and perform model training processing on the preset large language model according to the training set to obtain an initial interactive 3D generation model;

[0028] Use the LoRA algorithm to perform model fine-tuning processing on the initial interactive 3D generation model according to the evaluation set to obtain an interactive 3D generation model.

[0029] Optionally, in the method for generating 3D results of hand-object interaction based on a large language model, the obtaining of the user description text and the inputting of the user description text into the interactive 3D generation model to output preliminary interactive 3D data specifically includes:

[0030] Obtain the user description text and set custom hyperparameters, where the custom hyperparameters include text randomness and text length;

[0031] Input the user description text and the custom hyperparameters into the interactive 3D generation model, and perform word segmentation processing, encoding processing, and formatting processing on the user description text through the interactive 3D generation model to obtain serialized data;

[0032] Extract the context features of the serialized data through the self-attention mechanism and perform decoding processing on the context features to obtain preliminary interactive 3D data.

[0033] Optionally, in the method for generating 3D results of hand-object interaction based on a large language model, the optimizing of the hand parameters in the interactive 3D generation model according to the preliminary interactive 3D data to obtain optimized hand MANO parameters and obtaining the target hand-object interaction 3D result according to the optimized hand MANO parameters specifically includes:

[0034] Perform segmentation processing on the preliminary interactive 3D data to obtain an object OBJ model and second hand MANO parameters;

[0035] Calculate the corresponding object BPS distribution according to the object OBJ model, and calculate the surface distance between the hand vertices and the object according to the second hand MANO parameters;

[0036] Optimize the hand parameters of the interactive 3D generation model according to the object BPS distribution, the surface distance between the hand vertices and the object, and the second hand MANO parameters to obtain optimized hand MANO parameters;

[0037] Combine the hand MANO parameters and the object OBJ model to obtain the target hand-object interaction 3D result.

[0038] In addition, to achieve the above object, the present invention also provides a hand-object interaction 3D result generation system based on a large language model, wherein the hand-object interaction 3D result generation system based on a large language model includes:

[0039] A hand-object interaction data processing module for obtaining hand-object interaction data and performing preprocessing and text annotation processing on the hand-object interaction data to obtain a target 3D interaction data set;

[0040] An interactive 3D generation model construction module for determining a preset large language model and performing model training and model fine-tuning on the preset large language model according to the target 3D interaction data set to obtain an interactive 3D generation model;

[0041] A preliminary interactive 3D data output module for obtaining a user description text and inputting the user description text into the interactive 3D generation model to output preliminary interactive 3D data;

[0042] A target hand-object interaction 3D result generation module for optimizing the hand parameters in the interactive 3D generation model according to the preliminary interactive 3D data to obtain optimized hand MANO parameters, and obtaining a target hand-object interaction 3D result according to the optimized hand MANO parameters.

[0043] In the present invention, hand-object interaction data is acquired, and the hand-object interaction data is preprocessed and text-annotated to obtain a target 3D interaction dataset; a preset large language model is determined, and the preset large language model is trained and fine-tuned according to the target 3D interaction dataset to obtain an interactive 3D generation model; user description text is acquired, and the user description text is input into the interactive 3D generation model to output preliminary interactive 3D data; the hand parameters in the interactive 3D generation model are optimized according to the preliminary interactive 3D data to obtain optimized hand MANO parameters, and a target hand-object interaction 3D result is obtained according to the optimized hand MANO parameters. By collecting hand-object interaction data, preprocessing it, and text-annotating it, the present invention constructs a target 3D interaction dataset, trains and fine-tunes a large language model according to the target 3D interaction dataset to obtain an interactive 3D generation model, and then optimizes the interactive 3D generation model through user description text, enabling the rapid and accurate output of hand-object interaction 3D results. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a flowchart of a preferred embodiment of the method for generating a hand-object interaction 3D result based on a large language model of the present invention;

[0045] Figure 2 is a schematic diagram of the overall structural implementation process of a preferred embodiment of the method for generating a hand-object interaction 3D result based on a large language model of the present invention;

[0046] Figure 3 is a schematic diagram of the dataset generation of a preferred embodiment of the method for generating a hand-object interaction 3D result based on a large language model of the present invention;

[0047] Figure 4 is a schematic diagram of the dataset preprocessing of a preferred embodiment of the method for generating a hand-object interaction 3D result based on a large language model of the present invention;

[0048] Figure 5 is a schematic diagram of the text annotation process of a preferred embodiment of the method for generating a hand-object interaction 3D result based on a large language model of the present invention;

[0049] Figure 6 is a schematic diagram of the model fine-tuning of a preferred embodiment of the method for generating a hand-object interaction 3D result based on a large language model of the present invention;

[0050] Figure 7 is a schematic diagram of the processing flow of an interactive 3D generation model of a preferred embodiment of the method for generating a hand-object interaction 3D result based on a large language model of the present invention;

[0051] Figure 8It is a schematic diagram of the processing flow of the interactive 3D optimization model of the preferred embodiment of the method for generating 3D results of hand-object interaction based on the large language model of the present invention;

[0052] Figure 9 It is a schematic diagram of generating optimized MANO parameters of the hand of the preferred embodiment of the method for generating 3D results of hand-object interaction based on the large language model of the present invention;

[0053] Figure 10 It is a schematic diagram of generating the target 3D result of hand-object interaction of the preferred embodiment of the method for generating 3D results of hand-object interaction based on the large language model of the present invention;

[0054] Figure 11 It is a structural diagram of the preferred embodiment of the system for generating 3D results of hand-object interaction based on the large language model of the present invention;

[0055] Figure 12 It is a structural diagram of the preferred embodiment of the terminal of the present invention. Specific Embodiments

[0056] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0057] The existing 3D generation models have the following main problems in hand-object interaction 3D generation: 1. Poor annotation effect of the interactive 3D model: The existing annotation process directly describes the 3D model, and cannot comprehensively and detailedly describe the hand-object interaction details. There is a lack of an annotation process for interaction, and the system can only generally describe the overall 3D model, and cannot accurately describe the specific holding method of a specific object. 2. Insufficient pre-training of the large language model for hand-object interaction: The existing 3D generation models (such as MeshXL) are mainly pre-trained on a large number of 3D object models, lacking special training for hand-object interaction scenarios. Therefore, when these models generate hand-object interaction 3D models, it is difficult to accurately understand the interaction description input by the user (such as "pick up the water cup with the hand"), and they cannot generate 3D models that conform to physical laws and visual realism. 3. Limited ability of the interactive 3D optimization model: After generating the hand-object interaction 3D model, the existing technology lacks an effective optimization mechanism to adjust the interaction posture between the user's hand and the object. For example, key factors such as the distance between the hand vertices and the object surface and the influence of the object shape on the hand posture are not fully considered, resulting in obvious physical unreasonableness in the generated 3D models in practical applications (such as the hand penetrating the object, unnatural holding postures, etc.).

[0058] To solve the above problems, the present invention provides a hand-object interaction 3D generation method and system based on large language model fine-tuning technology. The specific objectives include: 1. Improving the 3D model annotation effect: Through an innovative multi-modal annotation process, calling BLIP (Bootstrapping Language-Image Pre-training, a visual pre-training framework) and CLIP (Contrastive Language-Image Pre-Training, a deep learning model that can process text and images simultaneously), generate detailed and accurate hand-object interaction 3D model annotations. Refined descriptions are made for specific grasping methods of specific objects (such as picking up, passing, using, etc.), solving the problems of incomplete and inaccurate annotation information in the prior art, and providing high-quality data support for model fine-tuning. 2. Enhancing the interaction understanding ability of the large language model: By fine-tuning the pre-trained large language model, enabling it to better understand the interaction descriptions input by users and initially generate 3D models that conform to physical laws and visual realism, solving the problem of insufficient pre-training of existing large language models in interaction scenarios. 3. Optimizing the interaction postures of the generation model: By introducing an optimization module, finely adjusting the generated hand-object interaction 3D models to ensure that the interaction postures between the hand and the object conform to physical laws (avoiding the hand penetrating the object and making the grasping posture natural, etc.), solving the problem of limited optimization ability of the generation model in the prior art, and improving the practicality and realism of the 3D models. 4. Constructing an end-to-end 3D generation system: By designing a complete process from text description to 3D model generation, including modules such as data collection, preprocessing, annotation, model fine-tuning, generation and optimization, construct an efficient and accurate hand-object interaction 3D generation system, solving the problems of scattered processes and low efficiency in the prior art, and significantly improving the generation efficiency and quality of 3D models.

[0059] The technologies adopted in the present invention include: 1. Natural Language Processing: Utilize a pre-trained large language model (such as Llama, Large Language Model Meta AI), and through fine-tuning technology, convert the user's text description into a 3D hand-object interaction model. 2. Computer Vision: Use multi-view rendering technology to render the 3D model into a 2D image, and run the BLIP model to generate an image description. 3. Multimodal Artificial Intelligence: By fusing text, images, and 3D data, combining the BLIP and CLIP models, generate multi-view 3D interaction data annotations, and utilize multimodal technology to improve the accuracy and diversity of 3D model annotations. 4. 3D Model Generation Technology: By fine-tuning the Llama model, directly generate a 3D object model (OBJ format) and hand pose parameters (MANO model) from the text description, and use an optimization module to adjust the interaction pose between the hand and the object to generate a high-quality hand-object interaction 3D model.

[0060] The method for generating a 3D result of hand-object interaction based on a large language model according to a preferred embodiment of the present invention is as Figure 1 shown, and the method for generating a 3D result of hand-object interaction based on a large language model includes the following steps:

[0061] Step S10: Obtain hand-object interaction data, and perform preprocessing and text annotation processing on the hand-object interaction data to obtain a target 3D interaction dataset, wherein the hand-object interaction data includes an object 3D model and first hand MANO parameters.

[0062] As Figure 2 shown, the present invention proposes a hand-object 3D interaction generation system based on the fine-tuning technology of a large language model LLM. Through two stages of generation + optimization, it realizes the generation from the user's text description to a 3D model. The system workflow is as follows: 1. Collect hand-object interaction data from multiple open-source hand datasets (i.e., the 3D data samples in Figure 2 ), preprocess the object and the hand, and construct an unannotated 3D interaction dataset (which is a part of the 3D model dataset in Figure 2 ). 2. Use BLIP and CLIP to perform text annotation on the hand-object interaction data to construct an annotated 3D interaction dataset (which is a part of the 3D model dataset in Figure 2Part of the 3D model dataset). 3. Use the labeled 3D model dataset to fine-tune and train the open-source large language model Llama as an interactive 3D generation model. 4. Generate a preliminary interactive 3D model (object and hand) using the interactive 3D generation model according to the user's text description. 5. Extract the distance between the hand vertices and the object surface in the preliminary interactive 3D model, the object BPS distribution (Basis Point Sets, a point cloud encoding technology aimed at efficiently converting three-dimensional point cloud information into a fixed-length feature vector. This method selects fixed base points in space and calculates the vectors from these base points to the nearest points in the point cloud as unique representative features), and the hand MANO parameters as the input of the optimization module to optimize the generated hand model as the final output.

[0063] The present invention fully combines the picture understanding and annotation capabilities of the (multimodal) large language model technology, designs a new 3D model annotation method, and innovatively uses the interactive 3D model as the fine-tuning target, enabling the Llama language model to directly generate a hand-object interactive 3D model for the first time.

[0064] Specifically, obtain an open-source hand dataset, extract the hand-object interaction data in the open-source hand dataset, and perform model decimation, data filtering, and normalization on the object 3D model to obtain an integer object model; perform integerization on the first hand MANO parameters to obtain integer hand MANO parameters; merge the integer object model and the integer hand MANO parameters to obtain a first 3D interaction dataset.

[0065] As Figure 3 shown, the dataset generation process is as follows: 1. Collect the hand-object interaction dataset: Widely collect the hand-object interaction data in the open-source hand dataset, including various holding methods of different objects by different subjects. 2. Call the data preprocessing process to generate 3D interaction data: Use an innovative data preprocessing process to generate a high-quality 3D interaction dataset for subsequent training and data annotation. 3. Call the text annotation process to generate a complete 3D interaction dataset: Use an innovative text annotation method to generate a detailed and complete 3D interaction text description to construct a complete 3D interaction dataset for model fine-tuning. 4. Call the dataset post-processing process: Merge the preprocessed 3D interaction data with the text annotation and organize the dataset in the Llama3 (a large language model) format, which can be used for Llama3 model fine-tuning.

[0066] Introduction to the data preprocessing process: The present invention collects a large number of high-quality and diverse 3D hand-object interaction sample data, and preprocesses and text-annotates these data, which plays a key role in improving the model's ability to generate 3D interaction models. The following is the specific process of data collection and dataset preprocessing:

[0067] 1. Data collection: The present invention extracts hand-object 3D interaction data from the Grab dataset (referring to the dataset of whole-body movements for interacting and grabbing 3D objects) and the Oakink dataset (referring to the dataset for image-to-3D tasks) (the dataset itself is composed of object 3D models and hand MANO parameters (hand MANO parameters can also be called hand posture parameters), and each sample contains one data item: hand-object contact area. In the present invention, the object model and hand parameters corresponding to the frame with the largest hand-object contact area in a video sequence in the dataset can be directly selected as training data). These datasets are composed of continuous video frames. In order to ensure the reliability of the collected data, the 3D interaction data of the frame with the largest contact area between the hand and the object in each video is extracted as available data. Considering the diversity of objects and hand movements, the dataset collects a total of up to 157 objects, such as water cups, mobile phones, hammers, mice, staplers, etc. that can interact with the hands, and at the same time, different interaction methods are performed on the objects, such as picking up, passing, using, etc., to avoid the model from overfitting to the characteristic interaction actions of specific objects.

[0068] 2. Dataset preprocessing: Dataset preprocessing includes two parts: object model processing and hand MANO parameter processing (the hand-object interaction data in the present invention includes two parts, one is the object 3D model, and the other is the hand MANO parameter (also called hand posture parameter). The two together are called hand-object interaction data, also called hand-object 3D interaction data). The overall process is as follows: Figure 4 As shown: The object model processing process is as follows: In the present invention, the object mesh is expressed in OBJ plain text format, and the preprocessing process is as follows: a. Model reduction (i.e. Figure 4 "Model simplification": Too many faces of an object will cause too many tokens to be occupied in the fine-tuning of the language model. The Quadric Edge Collapse Decimation method is used to reduce the face of the object model (the Quadric Edge Collapse Decimation method is used for model reduction. Quadric Edge Collapse Decimation is an algorithm for mesh simplification in 3D graphics. It improves rendering efficiency by reducing the number of vertices in the mesh while retaining the shape and details of the model as much as possible. The core idea of the algorithm is to simplify the mesh by gradually merging adjacent vertices. Each time a merge occurs, the algorithm calculates the cost of the edge, selects the edge with the smallest cost for contraction, and updates the error metric of the merged vertices to minimize the impact on the appearance of the model. It is widely used in 3D model optimization and real-time rendering). b. Data filtering (i.e. Figure 4("Removing low-quality objects"): For both decimation methods, it is possible to cause distortion of the object model. Filter the object model according to the original shape and representational semantics of the model, and screen out the data with model distortion and decimation failure. c. Data collation (i.e., Figure 4 ("Integerization to 0-64"): In order to adapt to the autoregressive characteristics of large language models, the model needs to predict the next token step by step according to the context of the sequence, and the data needs to be organized in a certain order to maintain logical consistency. Normalize and integerize the 3D coordinates of all object models to the range of 0-64, and at the same time sort the vertices of the OBJ file in ascending order of xyz coordinates to ensure the order and consistency of the coordinate data, which is convenient for subsequent tokenization of the data and use for fine-tuning training. Among them, the processing process of the hand model is as follows: In the present invention, MANO parameters are used to express the 3D hand model, and the MANO parameters of the hand are integerized to the range of 0-512, and merged with the object OBJ file (string concatenation) to construct interactive 3D data.

[0069] Introduction to 3D objects in OBJ format: In the present invention, the OBJ (Wavefront OBJ) plain text format is used to express the object mesh. It is a plain text file format for representing three-dimensional geometric shapes and is often used to store object meshes. It describes geometric information such as the vertices, faces, and normal vectors of the object, which is convenient for the storage and exchange of the model. Among them, the introduction to the geometric information is as follows: 1. Vertex (v): Represents the coordinates of a point in three-dimensional space through v, x, y, z. 2. Normal vector (vn): Represents the normal direction of the vertex through vn, x, y, z, which is used for lighting calculation. 3. Texture coordinate (vt): Represents the coordinates of the texture map through vt, u, v. 4. Face (f): Represents the polygon face through f v1 / vt1 / vn1, f v2 / vt2 / vn2..., where the vertex index is described in the order of vertex, texture, and normal.

[0070] The OBJ file is simple and highly readable. In order to retain the most important information of the 3D model and facilitate subsequent model fine-tuning, only the vertex (v) and face (f) data in the OBJ format are retained to construct the 3D interaction dataset.

[0071] Perform multi - perspective rendering processing on the hand - object interaction data to obtain multiple 2D images from different perspectives; perform text description processing on the multiple 2D images from different perspectives to obtain multiple image description texts; extract the image features and text features in the multiple image description texts, and calculate the similarity between the image features and the text features to obtain the target text description with the highest similarity in each image description text; construct a second 3D interaction dataset according to the target text description with the highest similarity in each image description text; perform merging processing and dataset partitioning processing on the first 3D interaction dataset and the second 3D interaction dataset to obtain a target 3D interaction dataset.

[0072] Introduction to the text annotation process: The present invention improves and innovates the Cap3D annotation method, combines models such as BLIP and clip, and adjusts the calling method and annotation merging method of the BLIP model, enabling it to accurately and detailedly describe the collected 3D interaction data and generate diverse and rich description texts, such as Figure 5 As shown, its steps include: 1. 3D object rendering: Observe and render the collected hand - object interaction 3D model from multiple perspectives (load the 3D model using Blender, an open - source 3D computer graphics software, set the perspectives of the camera at different positions, and then generate 8 images from different perspectives) into 2D images for subsequent text annotation. 2. Call the BLIP model to generate preliminary text descriptions: Use the QA mode (the QA (Question Answering) mode of the BLIP model is a multi - modal question - answering ability based on vision - language pre - training. It can combine image and text information to answer questions raised by users. In this mode, the user may input text questions into the model multiple times, and the model answers based on the text input and the image) to call the BLIP model multiple times to describe the pictures from different perspectives. First, ask about the object type in the image, use the object type as the prompt setting for the next question, and require the BLIP model to describe the object and hand - object interaction details in the picture. To ensure the correctness of the description generation, require the BLIP model to generate 5 results simultaneously for subsequent text description screening and merging. 3. Call CLIP for text screening and merging: For the 5 text descriptions generated by the BLIP model, use the CLIP model to extract the picture features and the 5 text description features from this perspective, and screen out the one text description that is closest to the picture features among the 5 text descriptions through similarity calculation as the description from this perspective.

[0073] For the text descriptions from 8 perspectives generated by a 3D interaction model, call the large - language model for merging, requiring the 8 - perspective text descriptions to be merged into one, emphasizing the hand - object interaction details and ignoring the background in the description, and fusing the text descriptions from multiple perspectives to make the description more detailed and complete.

[0074] Step S20: Determine a preset large language model, and perform model training and model fine-tuning on the preset large language model according to the target 3D interaction dataset to obtain an interactive 3D generation model. The target 3D interaction dataset includes a training set and an evaluation set.

[0075] The present invention generates a high-quality 3D hand-object interaction dataset through data collection, preprocessing, text annotation, and postprocessing, which conforms to the Llama 3.1 fine-tuning format. Key frames are extracted from the Grab and Oakink datasets, object models and hand parameters are processed, and detailed interaction descriptions are generated by combining BLIP, CLIP, and ChatGPT. Finally, the output is a standardized JSON file containing text descriptions and 3D data, which is directly used for model fine-tuning.

[0076] Specifically, determine a preset large language model, and perform model training on the preset large language model according to the training set to obtain an initial interactive 3D generation model; use the LoRA algorithm to perform model fine-tuning on the initial interactive 3D generation model according to the evaluation set to obtain an interactive 3D generation model.

[0077] Introduction to the postprocessing process of the dataset: 1. Dataset organization: For large model fine-tuning, the dataset needs to be organized in a certain format. For the Llama 3.1 model used in the present invention, the Llama 3 template is used to organize the data. Specifically, the descriptions obtained from text annotation are used as Instructions, the object OBJ format data and hand MANO parameters in the 3D dataset are combined as Outputs, the Input is set to an empty string, and multiple 3D data are organized into a json file according to the above format for large model fine-tuning. 2. Dataset division: The organized dataset is randomly shuffled, and the dataset as a whole is divided into an 80% training set and a 20% evaluation set. 3. Dataset storage: After data processing is completed, the training set and the evaluation set are stored separately for subsequent use.

[0078] As Figure 6 shown, the specific process of model fine-tuning is as follows: 1. Fine-tuning training: Take out the training set from the dataset and perform fine-tuning training for hand-object 3D interaction generation. 2. Evaluation and verification: Regularly select evaluation set data during the fine-tuning process for model evaluation to check the results of model fine-tuning.

[0079] Among them, the training process of the model is as follows: The Llama model is fine-tuned using SFT (Supervised Fine-Tuning), and LoRA is used to update and fine-tune some parameters of the Llama model. The fine-tuning strategy is autoregressive next-token prediction. The loss is calculated based on the probability distribution of the predicted next token and the actual next token each time, and the gradient descent algorithm is used to update the LoRA part of the parameters in the Llama model.

[0080] Introduction to the algorithm model: 1. Fine-tuning settings: According to the characteristics of the interactive 3D generation task, design a suitable fine-tuning scheme, select hyperparameters such as the number of fine-tuning layers, learning rate, Batch-Size, and number of training epochs, and determine the training objective and loss function. 2. Model fine-tuning: Load the pre-trained Llama model and the prepared hand-object interaction 3D data into the fine-tuning script using LoRA fine-tuning technology, and start the fine-tuning process. LoRA (Low-Rank Adaptation) technology is an efficient model fine-tuning method that fine-tunes the model behavior by introducing small, low-rank matrices in the key layers of the model without significantly modifying the entire model structure. The model learns the overall description in the hand-object interaction 3D data in a supervised manner, such as in what posture to hold what object, and performs Fine-tuning on the original parameters to adapt to the interactive 3D generation task. 3. Evaluation and verification: Regularly evaluate the generation effect of the model during the fine-tuning process, use methods such as cross-validation to evaluate the performance of the model in interactive 3D generation, and adjust the fine-tuning strategy according to the evaluation results to optimize the model. 4. Model saving: After the fine-tuning process is completed, save the fine-tuned model parameters for subsequent use and deployment.

[0081] The core function of model fine-tuning is to perform targeted fine-tuning on the Llama model to improve the generation ability for interactive 3D, serving as the input for subsequent interactive 3D optimization models. At the same time, the fine-tuned model adopts a unified data organization form, which is convenient for directly extracting the generated object meshes and hand parameters from the model output, facilitating subsequent optimization.

[0082] Step S30: Obtain the user description text and input the user description text into the interactive 3D generation model to output preliminary interactive 3D data.

[0083] As Figure 7 shown, the specific processing flow of the interactive 3D generation model is as follows: 1. Receive user input: The user calls the web Gradio platform developed by the present invention, enters a text description at the specified location, and sets custom hyperparameters (such as temperature, max_lengtht, etc.), and automatically obtains the parameters as the input of the interactive 3D generation model. 2. Interactive 3D model generation: Obtain the text description and hyperparameters input by the user, and the interactive 3D generation model generates preliminary interactive 3D data.

[0084] Among them, there are two custom hyperparameters, temperature and max_length, which control the randomness and length of the generated text respectively. Temperature is used to adjust the diversity of the generated text. The higher its value, the more random and diverse the generated text is, and it may generate a more flexible 3D interaction model according to the user's input. On the contrary, it is more fixed and generates a relatively conservative but accurate 3D interaction model. Max_length is used to control the maximum length of the generated text. If the interaction model to be generated is more complex, a longer output length needs to be adjusted to output the interaction model completely.

[0085] Specifically, obtain the user description text and set custom hyperparameters, where the custom hyperparameters include text randomness and text length. Input the user description text and the custom hyperparameters into the interactive 3D generation model, and perform word segmentation, encoding, and formatting processing on the user description text through the interactive 3D generation model to obtain serialized data. Extract the context features of the serialized data through the self-attention mechanism and perform decoding processing on the context features to obtain preliminary interactive 3D data.

[0086] Introduction to the algorithm model: Llama (Large Language Model Meta AI) is a high-performance large-scale language model for natural language processing tasks. Inside the Llama system, a series of preprocessing operations are first performed on the input text data, including word segmentation, encoding, and format standardization, etc., to convert the original text into serialized data in a unified format. Subsequently, these text data will be input into a multi-layer Transformer architecture, and the context features of the text are extracted through the self-attention mechanism to understand the semantic and structural relationships in the sentence.

[0087] Furthermore, the Llama model generates output content related to the task through a highly optimized decoding process. Whether it is text generation, translation, or question-and-answer tasks, Llama can provide high-quality results based on the context. Its core lies in the powerful pre-training mechanism and highly adjustable fine-tuning ability, enabling it to quickly adapt to different application scenarios.

[0088] Through a fine-tuned Llama-based 3D data generation model for hand-object interaction, the present invention can generate corresponding 3D objects (represented in OBJ format) and their hand interaction pose parameters (based on the MANO model) according to the text description input by the user. By combining semantic understanding and 3D data mapping capabilities, the model can accurately generate high-fidelity objects and hand interaction poses that match the description, providing reliable data support for application scenarios such as physical simulation, visual analysis, and virtual reality.

[0089] The interactive 3D generation model is the first generation in hand-object 3D interaction. Its main function is to generate a preliminary object OBJ model and the hand MANO parameters for interacting with it according to the user's text description, which is the key to optimizing the hand parameters based on the object later. Output the 3D OBJ format model of the object and the hand MANO parameters for subsequent hand parameter optimization.

[0090] Step S40: Optimize the hand parameters in the interactive 3D generation model according to the preliminary interactive 3D data to obtain optimized hand MANO parameters, and obtain the target hand-object interaction 3D result according to the optimized hand MANO parameters.

[0091] As Figure 8 shown, the specific processing flow of the interactive 3D optimization model is as follows: 1. Preliminary interactive 3D data segmentation: Segment the output of the interactive 3D generation model according to the standard form of the dataset organization, and extract the object OBJ model (i.e., the object obj model in Figure 8 and the hand MANO parameters (i.e., Figure 8(hand MANO parameters in it), and input them into the interactive 3D optimization model. 2. 3D model optimization: Calculate the object BPS distribution according to the generated object OBJ model, render the hand model according to the hand MANO parameters, and calculate the signed distance from the hand vertices to the object point-to-point (the 3D hand model rendered through the hand MANO parameters contains 778 vertices. For each hand vertex, select a vertex of the object model closest to it and calculate the signed distance between the two (the signed distance is positive if the hand vertex is outside the object model and negative if it is inside the object model)). Combine the generated hand MANO parameters as the input of the interactive 3D optimization model to optimize the hand parameters (the 3D optimization model is a neural network, and the input is "object BPS distribution, signed distance between hand and object vertices, hand MANO parameters", and the output is the optimized hand MANO parameters. For each training sample, the MANO parameters at the time of dataset collection are used as the true label. By calculating the error loss between the generated hand MANO parameters and the true MANO parameters, update the 3D optimization model parameters through the gradient descent method. Eventually, the model can optimize the hand parameters according to the input hand-object information and generate reasonable hand-object interaction results). 3. Interactive 3D result generation: Input the optimized hand MANO parameters of the interactive optimization model into the MANO model to render the hand model, combine the hand model with the object model to obtain the complete hand-object interaction 3D result, and use the Trimesh library (Trimesh is a pure Python library for loading and using triangular Mesh meshes) to load and display it at the specified position on the Web Gradio interface developed by the present invention (referring to the Web interface of the machine learning model created using the Gradio Python library).

[0092] Specifically, segment the preliminary interactive 3D data to obtain the object OBJ model and the second hand MANO parameters; calculate the corresponding object BPS distribution according to the object OBJ model, and calculate the surface distance between the hand vertices and the object according to the second hand MANO parameters; optimize the hand parameters of the interactive 3D generation model according to the object BPS distribution, the surface distance between the hand vertices and the object, and the second hand MANO parameters to obtain the optimized hand MANO parameters; combine the hand MANO parameters and the object OBJ model to obtain the target hand-object interaction 3D result.

[0093] The interactive 3D module is a brand-new deep learning neural network designed by the present invention, which is a neural network model for generating three-dimensional data of hand-object interaction. This model takes the hand-object distance feature, the hand pose rotation matrix and the translation vector as inputs, combines the object vertex data, and generates accurate hand parameters (including global direction, hand pose, translation vector and complete pose matrix) through iterative optimization.

[0094] As shown Figure 9 in the figure, after the fine-tuning training of the present invention, the interactive 3D module can learn more complex hand-object interaction patterns, optimize the preliminary hand-object 3D interaction model generated for a given text description, and generate a reliable hand-object interaction model that conforms to the text description. The specific steps are as follows: 1. MANO parameter processing: Use the MANO rendering model to render the generated MANO parameters into a 3D hand model, extract its vertex coordinate information, and use it for subsequent calculations and the directed distance between objects. 2. Object model processing: Encode the object model into a 4096-dimensional vector using the BPS encoding method. At the same time, calculate the directed distance between each vertex of the hand model and the object model, and both are used as inputs to the 3D optimization module. 3. Generate optimized hand parameters: The 3D optimization module receives three inputs: the BPS representation of the object, the hand-object vertex distance, and the hand MANO parameters. The neural network after fine-tuning can optimize the hand parameters to make the interaction with the object more realistic and reasonable.

[0095] Introduction to the MANO rendering model: The MANO rendering model is a parametric model for generating 3D hand meshes. Its inputs include 48-dimensional hand pose parameters and 3-dimensional spatial translation parameters. The pose parameters use axis-angle to represent joint rotations, and the translation parameters are used to control the spatial position of the hand. The model output is a 3D hand mesh containing 778 vertices, and a dynamic 3D hand model is generated through linear blend skinning.

[0096] Introduction to the BPS encoding method: The BPS encoding method (Basis Point Set) is a method for representing 3D objects as fixed-dimensional vectors. This method first defines a set of fixed base points in the 3D space of the object, and the number of these base points determines the final encoding dimension. For example, 4096 base points correspond to a 4096-dimensional vector. For each base point, BPS calculates the Euclidean distance from it to the nearest point on the object surface, and arranges these distances to form a feature vector. In this way, BPS can encode the shape information of 3D objects into fixed-length vectors while maintaining sensitivity to the spatial structure of the object. This encoding method has the property of being independent of the rotation and translation of the object, so it is suitable for various 3D shape analysis tasks.

[0097] It can be understood that the present invention optimizes the generated object OBJ model and hand MANO parameters. By using the BPS distribution of the object, the directed distance between hand-object vertices, and the initial hand MANO parameters as inputs, the hand parameters are iteratively adjusted to generate a more realistic and reasonable hand-object interaction result. The output of the model is the optimized hand MANO parameters, including global direction, hand pose, translation vector, and full pose matrix, which are used for the visualization of the final result.

[0098] AsFigure 10 As shown in the figure, the present invention uses the MANO rendering model to render the optimized hand MANO parameters into a 3D hand model, and uses the Trimesh library to combine the object model and the hand model to generate the final hand-object interaction 3D model.

[0099] Advantages of the present invention:

[0100] 1. 3D data annotation process: Combining large language models such as BLIP and CLIP, a new 3D interaction data annotation process is proposed. Through multi-view rendering and multi-model collaboration, high-quality and diverse 3D interaction data annotations are generated, significantly improving data usability.

[0101] 2. Fine-tuning of large language models: Using the hand-object interaction 3D data to fine-tune the Llama model, enabling the direct generation of 3D objects and hand postures from text descriptions, breaking through the limitations of traditional language models and expanding the application of language models in 3D generation tasks.

[0102] 3. Efficient hand parameter optimization module: Design a new hand optimization module to optimize the hand posture parameters through object BPS representation and hand-object distance, significantly improving the authenticity and rationality of the generation model and ensuring that the interaction postures conform to physical laws.

[0103] 4. Construction of high-quality 3D dataset: Through innovative preprocessing and annotation processes, a diverse 3D hand-object interaction dataset is constructed, covering a variety of objects and interaction methods, ensuring data reliability and diversity, and providing a solid foundation for model training.

[0104] 5. Flexible user interaction and generation: Develop a user interface based on the Web Gradio platform, supporting the rapid generation and optimization of 3D interaction models through text descriptions and parameter settings, meeting personalized needs and enhancing user experience and system usability.

[0105] The present invention proposes a hand-object interaction 3D generation system based on the fine-tuning technology of large language model LLM. By combining multi-modal large models such as BLIP, the generation and optimization of high-quality hand-object interaction 3D models from user text descriptions are realized. The core innovations are as follows:

[0106] 1. The present invention innovates in the 3D data annotation process. Combining multi-modal models such as BLIP, CLIP, and ChatGPT, a brand-new 3D data annotation method is proposed. By performing multi-view rendering on the 3D hand-object interaction model to generate 2D images, and using an improved BLIP call method to generate detailed and accurate 3D interaction text annotations, the annotation quality and diversity are significantly improved, providing high-quality data support for model training.

[0107] 2. The present invention has achieved a breakthrough in the fine-tuning technology of large language models. For the first time, hand-object interaction 3D data is used to fine-tune the Llama large language model, enabling it to directly generate 3D object models and hand MANO parameters from text descriptions. This innovation expands the application scope of large language models, enabling them to understand complex interaction semantics and generate high-fidelity 3D models, providing technical support for scenarios such as physical simulation and virtual reality.

[0108] 3. The present invention designs an efficient hand parameter optimization module. By calculating the directed distance between hand vertices and the object surface, and combining the object BPS distribution and hand MANO parameters, the hand posture is iteratively optimized. This process ensures that the generated hand-object interaction postures conform to physical laws, significantly improving the authenticity and rationality of the model.

[0109] 4. The present invention has made important contributions to the construction of high-quality 3D datasets. Through innovative preprocessing and annotation processes, a diverse dataset covering a large number of objects and various interaction methods is constructed. After preprocessing steps such as decimation, filtering, and sorting, the quality and diversity of the data are ensured, providing a solid foundation for model training and related research.

[0110] The proposal of the present invention marks the deep integration of large language models and 3D generation technology, opening up a new path for the generation and optimization of hand-object interaction 3D models. In the future, with the development of multimodal technologies, the technical framework of the present invention is expected to be extended to fields such as virtual reality and robot operation, promoting the innovation and progress of human-computer interaction technology.

[0111] Further, as Figure 11 shown, based on the above method for generating hand-object interaction 3D results based on a large language model, the present invention also correspondingly provides a system for generating hand-object interaction 3D results based on a large language model. Among them, the system for generating hand-object interaction 3D results based on a large language model includes:

[0112] A hand-object interaction data processing module 51, configured to obtain hand-object interaction data, and perform preprocessing and text annotation processing on the hand-object interaction data to obtain a target 3D interaction dataset;

[0113] An interactive 3D generation model construction module 52, configured to determine a preset large language model, and perform model training and model fine-tuning on the preset large language model according to the target 3D interaction dataset to obtain an interactive 3D generation model;

[0114] A preliminary interactive 3D data output module 53, configured to obtain a user description text, and input the user description text into the interactive 3D generation model to output preliminary interactive 3D data;

[0115] The target hand-object interaction 3D result generation module 54 is used to optimize the hand parameters in the interaction 3D generation model according to the preliminary interaction 3D data, obtain the optimized hand MANO parameters, and obtain the target hand-object interaction 3D result according to the optimized hand MANO parameters.

[0116] Further, as Figure 12 shown, based on the above hand-object interaction 3D result generation method and system based on a large language model, the present invention also correspondingly provides a terminal, which includes a processor 10, a memory 20, and a display 30. Figure 12 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0117] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as the hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as the program code for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a hand-object interaction 3D result generation program 40 based on a large language model is stored on the memory 20, and the hand-object interaction 3D result generation program 40 based on the large language model can be executed by the processor 10, so as to implement the hand-object interaction 3D result generation method in the present application.

[0118] The processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 20 or process data, such as executing the hand-object interaction 3D result generation method based on the large language model.

[0119] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and to display a visual user interface.

[0120] In one embodiment, when the processor 10 executes the large language model-based hand-object interaction 3D result generation program 40 in the memory 20, the steps of the large language model-based hand-object interaction 3D result generation method are implemented.

[0121] In summary, the present invention provides a large language model-based hand-object interaction 3D result generation method, system and terminal. The method includes: obtaining hand-object interaction data, and performing preprocessing and text annotation processing on the hand-object interaction data to obtain a target 3D interaction data set; determining a preset large language model, and performing model training and model fine-tuning on the preset large language model according to the target 3D interaction data set to obtain an interaction 3D generation model; obtaining a user description text, and inputting the user description text into the interaction 3D generation model to output preliminary interaction 3D data; optimizing the hand parameters in the interaction 3D generation model according to the preliminary interaction 3D data to obtain optimized hand MANO parameters, and obtaining a target hand-object interaction 3D result according to the optimized hand MANO parameters. By collecting hand-object interaction data, performing preprocessing and text annotation processing, the present invention constructs a target 3D interaction data set, and performs model training and model fine-tuning on the large language model according to the target 3D interaction data set to obtain an interaction 3D generation model. Then, by optimizing the model of the interaction 3D generation model with the user description text, the hand-object interaction 3D result can be output quickly and accurately.

[0122] It should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including the element.

[0123] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0124] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for generating 3D results of hand-object interaction based on large language models, characterized in that The method for generating 3D results of hand-object interaction based on a large language model includes: Obtaining hand-object interaction data, and performing preprocessing and text annotation processing on the hand-object interaction data to obtain a target 3D interaction dataset, including: Obtaining an open-source hand dataset, extracting hand-object interaction data from the open-source hand dataset, and performing preprocessing on the hand-object interaction data to obtain a first 3D interaction dataset; Performing text annotation processing on the hand-object interaction data to obtain a second 3D interaction dataset; Performing merging processing and dataset partitioning processing on the first 3D interaction dataset and the second 3D interaction dataset to obtain a target 3D interaction dataset; The hand-object interaction data includes an object 3D model and first hand MANO parameters; The performing preprocessing on the hand-object interaction data to obtain a first 3D interaction dataset specifically includes: Performing model decimation processing, data filtering processing, and normalization processing on the object 3D model to obtain an integer object model; Performing integerization processing on the first hand MANO parameters to obtain integer hand MANO parameters; Performing merging processing on the integer object model and the integer hand MANO parameters to obtain a first 3D interaction dataset; The performing text annotation processing on the hand-object interaction data to obtain a second 3D interaction dataset specifically includes: Performing multi-view rendering processing on the hand-object interaction data to obtain multiple 2D images of different views; Performing text description processing on the multiple 2D images of different views to obtain multiple image description texts; Extracting image features and text features from the multiple image description texts, and calculating the similarity between the image features and the text features to obtain a target text description with the highest similarity in each image description text; Constructing a second 3D interaction dataset according to the target text description with the highest similarity in each image description text; Determining a preset large language model, and performing model training and model fine-tuning on the preset large language model according to the target 3D interaction dataset to obtain an interactive 3D generation model; Obtaining a user description text, and inputting the user description text into the interactive 3D generation model to output preliminary interactive 3D data; Performing optimization processing on the hand parameters in the interactive 3D generation model according to the preliminary interactive 3D data to obtain optimized hand MANO parameters, and obtaining a target hand-object interaction 3D result according to the optimized hand MANO parameters.

2. The method for generating 3D results of hand-object interaction based on a large language model according to claim 1, wherein The target 3D interaction dataset includes a training set and an evaluation set; The determining a preset large language model, and performing model training and model fine-tuning on the preset large language model according to the target 3D interaction dataset to obtain an interactive 3D generation model specifically includes: Determining a preset large language model, and performing model training processing on the preset large language model according to the training set to obtain an initial interactive 3D generation model; Using the LoRA algorithm to perform model fine-tuning processing on the initial interactive 3D generation model according to the evaluation set to obtain an interactive 3D generation model.

3. The method for generating 3D results of hand-object interaction based on a large language model according to claim 1, wherein Obtaining the user description text and inputting the user description text into the interactive 3D generation model to output preliminary interactive 3D data, specifically including: Obtaining the user description text and setting custom hyperparameters, where the custom hyperparameters include text randomness and text length; Inputting the user description text and the custom hyperparameters into the interactive 3D generation model, and performing word segmentation processing, encoding processing, and formatting processing on the user description text through the interactive 3D generation model to obtain serialized data; Extracting the context features of the serialized data through the self-attention mechanism and performing decoding processing on the context features to obtain preliminary interactive 3D data.

4. The method for generating 3D results of hand-object interaction based on a large language model according to claim 1, wherein Optimizing the hand parameters in the interactive 3D generation model according to the preliminary interactive 3D data to obtain optimized hand MANO parameters, and obtaining the target hand-object interaction 3D result according to the optimized hand MANO parameters, specifically including: Performing segmentation processing on the preliminary interactive 3D data to obtain an object OBJ model and second hand MANO parameters; Calculating the corresponding object BPS distribution according to the object OBJ model, and calculating the surface distance between the hand vertices and the object according to the second hand MANO parameters; Optimizing the hand parameters of the interactive 3D generation model according to the object BPS distribution, the surface distance between the hand vertices and the object, and the second hand MANO parameters to obtain optimized hand MANO parameters; Combining the hand MANO parameters and the object OBJ model to obtain the target hand-object interaction 3D result.

5. A 3D result generation system for hand-object interaction based on large language models, characterized in that, The hand-object interaction 3D result generation system based on the large language model is applied to the hand-object interaction 3D result generation method according to any one of claims 1-4. The hand-object interaction 3D result generation system based on the large language model includes: A hand-object interaction data processing module, configured to obtain hand-object interaction data, and perform preprocessing and text annotation processing on the hand-object interaction data to obtain a target 3D interaction dataset; An interactive 3D generation model construction module, configured to determine a preset large language model, and perform model training and model fine-tuning on the preset large language model according to the target 3D interaction dataset to obtain an interactive 3D generation model; A preliminary interactive 3D data output module, configured to obtain the user description text and input the user description text into the interactive 3D generation model to output preliminary interactive 3D data; A target hand-object interaction 3D result generation module, configured to optimize the hand parameters in the interactive 3D generation model according to the preliminary interactive 3D data to obtain optimized hand MANO parameters, and obtain the target hand-object interaction 3D result according to the optimized hand MANO parameters.

6. A terminal, characterized in that, The terminal includes: a memory, a processor, and a large language model-based hand-object interaction 3D result generation program stored on the memory and executable on the processor. When the large language model-based hand-object interaction 3D result generation program is executed by the processor, it implements the steps of the large language model-based hand-object interaction 3D result generation method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a large language model-based hand-object interaction 3D result generation program. When the large language model-based hand-object interaction 3D result generation program is executed by a processor, it implements the steps of the large language model-based hand-object interaction 3D result generation method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Visual language model instruction fine tuning method and device

    CN117975475A

  • Real-time scene gesture generation method and device and digital human interaction system

    CN119648875A