Spatial position instruction fine tuning method based on multi-modal large language model

By fine-tuning the InternVL model using a multimodal large language model and a low-rank adaptation method, and combining this with optimization of the description using a large language model, the shortcomings of existing models in spatial location reasoning are addressed, resulting in more accurate and diverse spatial descriptions.

CN121960720APending Publication Date: 2026-05-01SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI JIAOTONG UNIV
Filing Date
2024-10-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing spatial location reasoning models lack comprehensive prior knowledge and text generation capabilities, resulting in inaccurate and detailed spatial descriptions that cannot be generalized across different scenarios.

Method used

We employ a multimodal large language model, transform the dataset through visual instruction format conversion, fine-tune the InternVL model using a low-rank adaptation method, and introduce a large language model to optimize the description, thereby constructing a visual spatial location reasoning dataset to enhance the model's generative capabilities.

Benefits of technology

It improves the accuracy and diversity of model generation in spatial location reasoning tasks, ensures the contextual appropriateness and linguistic richness of descriptions, and enhances computational efficiency and description quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121960720A_ABST
    Figure CN121960720A_ABST
Patent Text Reader

Abstract

The invention relates to a spatial position instruction fine tuning method based on a multi-modal large language model, and the method comprises the following steps: S1, converting a spatial position reasoning data set into a visual instruction format through employing a dialogue template, and obtaining a visual spatial position reasoning data set; s2, acquiring a large language model InternVL as a multi-modal large language model, performing pre-training on the general data set to obtain a pre-training model, reasoning the data set based on the visual spatial position, adjusting parameters of the pre-training model by adopting a low-rank adaptation method to obtain a trained large language model, and outputting a description corresponding to a spatial task by the large language model; and S3, introducing a text-based large language model, and optimizing the description corresponding to the space task based on the large language model. Compared with the prior art, the method has the advantages that the ability of the multi-modal large language model in understanding and generating context rich description is fully utilized, and the ability of the model in generating accurate and detailed description is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

A method for fine-tuning spatial location instructions based on a multimodal large language model Technical Field

[0001] This invention relates to the field of deep learning for spatial location reasoning, and in particular to a method for fine-tuning spatial location instructions based on a multimodal large language model. Background Technology

[0002] Spatial understanding is crucial for a wide range of applications and real-world scenarios, particularly in robotics and augmented reality. It enables a better understanding of visual scenes and object relationships, leading to more efficient planning and decision-making. The development of spatial understanding tasks has evolved to meet the growing demand for richer and more accurate representations of spatial relationships in images.

[0003] The initial focus was on Spatial Relationship Classification (VSRC), a task designed to identify spatial relationships between two objects in an image by selecting from a predefined set of relations. However, this approach only provides a shallow analysis of spatial semantics, limiting its applicability and expressive power. Recognizing the limitations of VSRC, researchers introduced a more advanced task called Spatial Reasoning. This task generates textual descriptions that convey the spatial semantics of an image, providing a deeper spatial analysis. Spatial Reasoning takes an image containing two specified objects as input and outputs a sentence describing their detailed spatial relationship.

[0004] Current research on spatial location reasoning tasks is primarily driven by pioneering work. These researchers pioneered the spatial location reasoning task by manually annotating spatial descriptions of images and utilizing existing spatial location classification datasets. This approach laid the foundation for generating richer and more accurate image spatial descriptions. Beyond establishing the spatial location reasoning task, these studies also treat it as a general image-to-text task, using Visual Language Pre-trained Models (VL-PTMs). These models are designed to take images as input and output text, making them suitable for generating spatial descriptions. By leveraging VL-PTMs, researchers are able to generate textual descriptions that effectively convey the spatial semantics within images.

[0005] Despite significant progress in spatial location reasoning tasks, notable limitations hinder the full potential of this approach. Two main drawbacks exist: a lack of sufficient prior knowledge and poor text generation capabilities of current models. First, existing spatial location reasoning models typically lack comprehensive prior knowledge, crucial for understanding and describing complex spatial relationships. These models are often trained on relatively small and more task-specific datasets. Consequently, they lack the broad world knowledge required to generate accurate and context-rich spatial descriptions. This limitation restricts their applicability to specific tasks and reduces their ability to generalize across diverse scenarios.

[0006] Secondly, current models' text generation capabilities have not yet reached a level that consistently produces high-quality, context-appropriate spatial descriptions. These models often lack the ability to follow complex instructions, which is crucial for generating detailed and accurate descriptions of spatial relationships. This deficiency results in outputs that are either too general or fail to capture the specific nuances of spatial context.

[0007] In summary, existing models are unable to generate accurate and detailed descriptions when used for spatial location reasoning tasks, resulting in poor reasoning performance. Summary of the Invention

[0008] The purpose of this invention is to provide a spatial location instruction fine-tuning method based on a multimodal large language model, which aims to fully utilize the ability of multimodal large language models to understand and generate rich contextual descriptions and enhance the model's ability to generate accurate and detailed descriptions.

[0009] The objective of this invention can be achieved through the following technical solutions:

[0010] A method for fine-tuning spatial location instructions based on a multimodal large language model, comprising the following steps:

[0011] S1. Use the dialogue template to convert the spatial location reasoning dataset into a visual instruction format to obtain the visual spatial location reasoning dataset.

[0012] S2. Obtain the large language model InternVL as a multimodal large language model, pre-train it on a general dataset to obtain a pre-trained model, and adjust the parameters of the pre-trained model using a low-rank adaptation method based on a visual spatial location reasoning dataset to obtain a trained large language model. The large language model outputs a description corresponding to the spatial task.

[0013] S3. Introduce a large text-based language model, and optimize the description corresponding to the spatial task based on the large text-based language model.

[0014] Furthermore, the specific steps of converting the spatial location reasoning dataset into a visual instruction format using a dialogue template are as follows:

[0015] For a data item (I, bbox) in the spatial location reasoning dataset O1 bbox O2 ,T1,T2,R,D), where I,bbox o1 bbox o2T1, T2, R, and D represent image, subject bounding box, object bounding box, subject category label, object category label, spatial relationship, and spatial description, respectively. D has two versions, D1 and D2, corresponding to a concise description and a detailed description, respectively. First, three dialogue templates are used to generate question X. q And answer X r Problem X q And answer X r This constitutes a visual spatial location reasoning dataset.

[0016] Furthermore, the question X corresponding to the three dialogue templates q And answer X r They are respectively:

[0017] Questions and answers corresponding to the classification templates for spatial location relationships, questions and answers corresponding to the description templates for single spatial relationships, and questions and answers corresponding to the descriptions of open spatial relationships.

[0018] Furthermore, the questions and answers corresponding to the classification template of the spatial location relationship are as follows:

[0019] Problem: Given an image, select the most appropriate preposition to complete the sentence: Subject category label T1 is blanked in the object category label T2, and the preposition to be filled in the blank is selected from multiple options;

[0020] Answer: Spatial relation R.

[0021] Furthermore, the specific questions and answers corresponding to the description template of a single spatial relationship are as follows:

[0022] Question: Based on the image, provide a concise text description or phrase describing a single spatial relationship between T1 and T2;

[0023] Answer: Briefly describe D1.

[0024] Furthermore, the specific questions and answers corresponding to the description of open spatial relationships are as follows:

[0025] Question: Based on the image, provide a detailed text description of the spatial relationship between T1 and T2;

[0026] Answer: Subject category label T1 Spatial relationship R Object category label T2\nConcise description D1\nDetailed description D2.

[0027] Furthermore, the low-rank adaptation method is used to adjust the parameters of the pre-trained model as follows:

[0028] The pre-trained model is connected to the visual language connector and the visual encoder. The visual spatial location reasoning dataset is input into the visual encoder and then fed into the pre-trained model for training via the visual language connector.

[0029] Furthermore, the specific steps for optimizing the description corresponding to the spatial task based on the large-scale language model are as follows:

[0030] Large-scale language models use prompt word templates to take the description of a single spatial relationship and the corresponding spatial task as the basic description, and optimize the basic description.

[0031] Furthermore, the specific details of the prompt word template are as follows:

[0032] Based on the concise spatial relationship description between the image and T1 and T2, i.e. the basic description, words or phrases in the basic description are replaced based on adjectives, verbs or synonyms, ensuring that no consecutive words in the basic description remain unchanged.

[0033] Furthermore, the large-scale language model is the Qwen2-7B model.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] (1) This invention constructs a visual spatial location reasoning dataset in a visual instruction format. The large language model, trained on a broad and diverse dataset including both visual and textual information, is able to better understand spatial relationships. This extensive training enables the model to utilize a wider range of prior knowledge, thereby generating more accurate and context-sensitive spatial descriptions, enhancing its ability to generate high-quality, coherent, and context-appropriate text. This capability is crucial for spatial location reasoning tasks because it ensures that the generated descriptions are not only accurate but also linguistically rich and information-rich.

[0036] (2) This invention introduces the InternVL model and fine-tunes it using LoRA to improve efficiency. This method significantly reduces the computational resources required for fine-tuning, enabling the InternVL model to be specifically tailored for spatial location reasoning tasks without incurring significant computational overhead. By utilizing LoRA, we can efficiently customize the InternVL model to meet the specific needs of spatial location reasoning tasks, thereby improving its performance while maintaining computational efficiency.

[0037] (3) To further improve the diversity and quality of the generated descriptions, this invention introduces a large-scale language model of plain text for description optimization. This method enhances the linguistic richness and diversity of the initial description. The optimization process ensures that the final description is not only accurate and context-appropriate, but also diverse and engaging. Attached Figure Description

[0038] Figure 1 is a schematic diagram of the structure of the present invention. Detailed Implementation

[0039] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0040] This invention discloses a spatial location instruction fine-tuning method based on a multimodal large language model (MLLM). The method comprises three modules: a spatial reasoning instruction tracking dataset construction module, a multimodal large language model fine-tuning module, and a description generation optimization module. The spatial reasoning instruction tracking dataset construction module constructs a spatial reasoning instruction tracking dataset for three tasks using given image-title pairs. The multimodal large language model fine-tuning module uses low-rank adaptation (LoRA) to fine-tune a multimodal large language model with 13 billion parameters and supporting high-resolution images for spatial reasoning. The description generation optimization module optimizes the generated sentences using the large language model, improving their diversity and accuracy. This invention addresses the problems of traditional spatial relationship classification methods neglecting world knowledge and lacking universal language capabilities, improving the ability to classify, describe, and provide open-ended descriptions of spatial relationships. Furthermore, the spatial location multimodal large language model of this invention exhibits excellent multimodal dialogue capabilities and can follow open-ended instructions to assist in querying object relationships in images. By using the method provided by this invention, spatial location reasoning tasks can be handled more effectively, improving the accuracy and diversity of generated descriptions, thereby achieving better performance in practical applications.

[0041] This invention proposes a spatial location instruction fine-tuning method based on a multimodal large language model. The flowchart of the method is shown in Figure 1. The method includes the following steps:

[0042] S1. Use the dialogue template to convert the spatial location reasoning dataset into a visual instruction format to obtain the visual spatial location reasoning dataset.

[0043] S2. Obtain the large language model InternVL as a multimodal large language model, pre-train it on a general dataset to obtain a pre-trained model, and adjust the parameters of the pre-trained model using a low-rank adaptation method based on a visual spatial location reasoning dataset to obtain a trained large language model. The large language model outputs a description corresponding to the spatial task.

[0044] S3. Introduce a large text-based language model, and optimize the description corresponding to the spatial task based on the large text-based language model.

[0045] Step 1: Visual instruction format conversion

[0046] 1.1) Convert the spatial location reasoning dataset into a visual instruction format, and create the first such dataset from existing spatial location reasoning data using a template-based approach;

[0047] 1.2) For the data items in the original spatial location reasoning dataset: (image, subject bounding box, object bounding box, subject category label, object category label, spatial relationship, spatial description), questions and answers are generated by sampling dialogue templates. These questions and answers respectively ask for and answer the spatial relationships required for a specific task.

[0048] 1.3) Create single-round instruction follow-up examples using (image, question, answer) to ensure that the data is presented in a visual dialogue manner, enabling the model to better understand the spatial relationships and nuances inherent in the visual data, thereby enhancing the model's ability to generate accurate and detailed descriptions.

[0049] Step 1 implements visual instruction format conversion, using dialogue templates to transform the spatial location reasoning dataset into visual instruction format. By framing spatial descriptions into visual instructions, this ensures that the data is presented in a way that allows the model to better understand the inherent spatial relationships and nuances within the visual data. This conversion fully leverages the capabilities of multimodal large language models (MLLMs) in understanding and generating context-rich descriptions, enhancing the model's ability to generate accurate and detailed descriptions.

[0050] Step 2: Low-rank adaptive fine-tuning

[0051] 2.1) We use InternVL as the basis for a multimodal dialogue model and fine-tune it using a low-rank adaptation method to improve efficiency and reduce the computational resources required for fine-tuning.

[0052] 2.2) By performing an initial two-stage training on our spatial location instruction following dataset, we use low-rank adaptive fine-tuning of the InternVL model to efficiently tailor the InternVL model to meet the specific needs of spatial location reasoning tasks while maintaining its generalization performance while keeping computational efficiency in place.

[0053] Step 2 involves fine-tuning the InternVL model using LoRA to improve efficiency. By leveraging LoRA, the computational resources required for fine-tuning are significantly reduced, allowing the InternVL model to be specifically tailored for spatial location reasoning tasks without incurring substantial computational overhead. This fine-tuning method efficiently customizes the InternVL model to meet the specific needs of spatial location reasoning tasks, thereby improving performance while maintaining computational efficiency.

[0054] Step 3 involves description optimization, which introduces a large text-based language model to refine the generated descriptions. By incorporating this large text-based language model, the richness and diversity of the language are increased; through the optimization process, the final descriptions are ensured to be not only accurate and context-appropriate, but also diverse and engaging. This step is crucial for applications that require high-quality, diverse descriptions.

[0055] The dialogue template design described in Step 1 aims to maximize the capabilities of MLLMs, enabling the model to better understand and generate spatial location instructions. The LoRA fine-tuning method described in Step 2 significantly reduces computational resource requirements, allowing the InternVL model to efficiently adapt to spatial location reasoning tasks. The description optimization process described in Step 3 further improves the quality and diversity of generated descriptions by introducing a large text-based language model.

[0056] Step 1 describes the visual instruction format conversion, which includes dataset creation: Due to the lack of multimodal spatial location reasoning datasets for training instruction-following assistants, we created the first such dataset from widely available spatial location reasoning data using a template process. For a single data item (I, bbox) in the original spatial location reasoning dataset... O1 bbox O2 The diagram (T1, T2, R, D) represents (image, subject bounding box, object bounding box, subject category label, object category label, spatial relationship, spatial description), respectively. Based on the spatial location inference dataset, D is further divided into two versions, D1 and D2, corresponding to a concise description and a detailed description, respectively. D1 may contain multiple sentences of description. We sample a dialogue template to generate a question X. q And an answer X r These respectively inquire into and answer the spatial relationships required for a specific task. Through (I,X) q ,X r We create a single-round instruction follow-up example: Human:IX q <stop>\n Assistant:X r <stop>\n.

[0057] The dialogue template described in step 1 is as follows: For a specific data item (I, Based on spatial location, the subtask type is inferred from T1, T2, R, D1, D2, and there are 3 dialogue templates:

[0058] 1. Task Type 1: Classification of Spatial Relationships

[0059] • Question: Given an image, choose the most appropriate preposition to complete the sentence: "The T1 is [BLANK] the T2." Choose from the following options: on, in, next to, above, behind, in front of, to the left of, to the right of.

[0060] ·Answer: R

[0061] 2. Task Type 2: Description of a Single Spatial Relationship

[0062] • Question: Based on the image, provide a concise text description or phrase describing a single spatial relationship between T1 and T2.

[0063] Answer: D1

[0064] 3. Task Type 3: Description of Open Spatial Relationships

[0065] • Question: Based on the image, provide a detailed text description of the spatial relationship between T1 and T2.

[0066] • Answer: T1RT2\nD1\n D2

[0067] Depending on the subtask of spatial location reasoning, the sampled questions may require describing spatial relationships using a word, a sentence, or multiple sentences. For Tasks 1 and 2, a phrase or sentence is sampled directly from the R or D of the data item for the answer. For Task 3, the answer contains a total of 3 sentences: the first sentence is generated by stacking T1, R, and T2; the second sentence is randomly sampled from D1; and the third sentence directly uses D2. For Tasks 1 and 3, one data example is generated for each image; for Task 2, one dialogue data item is generated for each description. The generated spatial location instruction follow-up dataset contains a total of 121,339 items, specifically: 20,490 for Task 1, 83,608 for Task 2, and 17,241 for Task 3. This transformation ensures that the data is presented in a way that allows the model to better understand the inherent spatial relationships and nuances in the visual data, thereby fully leveraging the capabilities of multimodal large language models (MLLMs) in understanding and generating context-rich descriptions, enhancing the model's ability to generate accurate and detailed descriptions.

[0068] The model selection and architecture used for fine-tuning the InternVL model described in step 2 are as follows: We utilize InternVL, a general-domain multimodal dialogue model, as the foundation for a large multimodal language model (MLLM), and gradually adapt it to the spatial location reasoning domain. The network architecture remains consistent, employing a linear projection layer to connect the visual encoder and the language model. This architectural design ensures that the model can effectively handle the fusion of visual and linguistic information. In the first stage of training, the model is pre-trained on a large-scale general dataset to acquire basic multimodal understanding capabilities. In the second stage of training, the model is fine-tuned on a specific spatial location reasoning dataset to enhance its performance on spatial location tasks.

[0069] Following the first stage of training, the LoRA fine-tuning described in step 2 uses Low-Rank Adaptation (LoRA) to fine-tune InternVL. The LoRA method adjusts model parameters by introducing a low-rank matrix, thereby improving the model's task-specific performance without significantly increasing computational resources. Through this fine-tuning method, we can develop a highly accurate, dataset-specific model, thereby improving the assistant's service quality.

[0070] The method described in step 2 offers significant advantages in terms of resources and efficiency. Our goal is to provide a cost-effective and practical solution that minimizes development overhead, rather than pursuing optimal performance by scaling up data or models. The fine-tuning process takes approximately 4 hours and uses 8 V100 GPUs for training. This efficient fine-tuning method ensures that the model still achieves excellent performance with limited resources.

[0071] The starting point for enhancing descriptive diversity described in step 3 lies in a problem with dataset integration: in Task 3, the limited availability of diverse training data for individual images poses a significant challenge. To address this, we integrate the simple descriptions, Task 2 descriptions, and Task 3 descriptions into our training dataset, as described above. However, this strategy results in highly similar generated descriptions, thus undermining the goal of achieving descriptive diversity.

[0072] The description optimization described in step 3 alleviates this problem. We leverage the powerful text generation capabilities of Large Language Models (LLMs) to introduce greater variability and richness into the generated descriptions. Through our measurements, we observed that the latter two sentences out of the three generated sentences exhibit higher BLEU4 scores.

[0073] Step 3, the optimization description, specifically employs the Qwen2-7B model. Using the second sentence as the base description, we construct prompt words. We instruct the model to generate a sentence to replace the base description while preserving the original spatial relationships. This method ensures that the final three generated descriptions are both diverse and accurate, thereby improving the overall quality and uniqueness of the output. The specific prompt word template is: "Based on the image and a concise spatial relationship description between T1 and T2: 'base description,' generate a sentence with a similar but simpler meaning. Maintain the main structure of the sentence, but replace words or phrases with simpler adjectives, verbs, or synonyms, ensuring that no consecutive words in the original description remain unchanged."

[0074] This invention demonstrates superior performance on spatial location reasoning datasets by converting spatial location reasoning datasets into a visual instruction format, fine-tuning the InternVL model using low-rank adaptation (LoRA), and introducing a large language model of pure text for description optimization.

[0075] 1) Advantages of Multimodal Large Language Models (MLLMs): MLLMs are trained on a wide and diverse range of datasets, including both visual and textual information, enabling them to better understand spatial relationships. This extensive training allows MLLMs to leverage broader prior knowledge, resulting in more accurate and context-sensitive spatial descriptions. The advanced architecture of MLLMs integrates visual and language processing, enhancing their ability to generate high-quality, coherent, and context-appropriate text. This capability is crucial for spatial location reasoning tasks, as it ensures that the generated descriptions are not only accurate but also linguistically rich and information-rich.

[0076] 2) Advantages of Efficient Fine-Tuning of the InternVL Model This invention introduces the InternVL model and fine-tunes it using LoRA to improve efficiency. This method significantly reduces the computational resources required for fine-tuning, enabling the InternVL model to be specifically tailored for spatial location reasoning tasks without incurring substantial computational overhead. By utilizing LoRA, we can efficiently customize the InternVL model to meet the specific needs of spatial location reasoning tasks, thereby improving its performance while maintaining computational efficiency.

[0077] 3) Advantages of Description Optimization: To further improve the diversity and quality of generated descriptions, this invention introduces a large-scale language model of plain text for description optimization. This method enhances the linguistic richness and diversity of the initial description. The optimization process ensures that the final description is not only accurate and context-appropriate, but also diverse and engaging. This step is crucial for applications requiring high-quality, diverse descriptions. The description optimization method delivers a 6-point advantage on BLEU-4 in Task 3, while the SPICE sacrifice is relatively negligible.

[0078] In summary, this invention significantly improves the descriptive diversity and accuracy in spatial location reasoning tasks by converting spatial location reasoning datasets into a visual instruction format, fine-tuning the InternVL model using LoRA, and introducing a large-scale language model based on pure text for description optimization. Through this method, the invention demonstrates excellent performance on multiple datasets and has significant application value.

[0079] Step 3: Description Optimization Based on Language Model

[0080] 3.1) Introduce a large text-based language model to optimize the generated descriptions, thereby improving the diversity and quality of the descriptions;

[0081] 3.2) By combining large-scale text-based language models, we can increase the richness and diversity of language, ensuring that the final description is not only accurate and context-appropriate, but also diverse and engaging.

[0082] 3.3) Specifically targeting open spatial location reasoning, the powerful text generation capabilities of the large language model are utilized to generate alternative sentences that retain the original spatial relationships through prompting models, ensuring that the final generated multi-sentence descriptions are improved in both diversity and accuracy.

[0083] The visual instruction format conversion step creates the first such dataset from existing spatial location reasoning data using a templated procedure, and generates questions and answers by sampling dialogue templates to ensure that the data is presented in a visual dialogue manner, enabling the model to better understand the spatial relationships and nuances inherent in the visual data.

[0084] The low-rank adaptation fine-tuning step described above fine-tunes the underlying multimodal dialogue model (i.e., InternVL) using a low-rank adaptation method, thereby reducing the computational resources required for fine-tuning and efficiently customizing the InternVL model to meet the specific needs of spatial location reasoning tasks.

[0085] The text model optimization step involves introducing a large text-based language model to optimize the generated descriptions, thereby improving their diversity and quality. This ensures that the final descriptions are not only accurate but also context-appropriate and varied. Specifically for open-ended spatial reasoning, the powerful text generation capabilities of the language model are utilized. A prompting model generates alternative sentences that preserve the original spatial relationships, ensuring that the three final generated descriptions are improved in both diversity and accuracy.

[0086] Referring to Figure 1, step 1 of this invention converts the spatial location reasoning dataset into a visual instruction format to fully utilize the capabilities of multimodal large language models (MLLMs) to generate context-rich descriptions. Converting the spatial location reasoning dataset into a visual instruction format using dialogue templates enables the model to better understand and generate descriptions of spatial relationships.

[0087] The spatial location reasoning dataset is converted into a visual instruction format, and dialogue templates are used to present the data so that the model can better understand the spatial relationships and subtleties in the visual data;

[0088] Each image and its corresponding description are converted into a visual instruction format using dialogue templates, ensuring that the data is presented in a way that enhances the model’s understanding.

[0089] Through this transformation, the model can generate more accurate and detailed descriptions, improving the diversity and contextual relevance of the descriptions.

[0090] Referring to Figure 1, step 2 of this invention uses the Low-Rank Adaptation (LoRA) method to fine-tune the InternVL model, thereby improving efficiency and reducing computational resource requirements. The LoRA method enables the InternVL model to be specifically tailored for spatial location reasoning tasks without significantly increasing computational overhead.

[0091] The LoRA method was used to fine-tune the InternVL model to improve its performance on spatial location reasoning tasks.

[0092] The LoRA method significantly reduces the computational resources required for fine-tuning, enabling efficient model fine-tuning even with limited computational resources.

[0093] The fine-tuned InternVL model is better suited to the needs of spatial location reasoning tasks, generating more accurate and context-sensitive descriptions.

[0094] Referring to Figure 1, step 3 of this invention introduces a large language model of plain text to optimize the generated description, thereby improving the diversity and quality of the description. This method enhances the linguistic richness and diversity of the initial description, ensuring that the final description is not only accurate and context-appropriate, but also diverse and engaging.

[0095] The generated descriptions are optimized using a large language model of plain text, enhancing the linguistic richness and diversity of the descriptions;

[0096] By optimizing the process, we ensure that the final description is not only accurate and context-appropriate, but also diverse and engaging.

[0097] This optimization method demonstrates superior performance on multiple spatial location inference datasets, further improving the quality and uniqueness of the descriptions.

[0098] In summary, this invention significantly improves the descriptive diversity and accuracy in spatial location reasoning tasks by converting spatial location reasoning datasets into a visual instruction format, fine-tuning the InternVL model using the LoRA method, and introducing a large-scale language model based on pure text for description optimization. Through this method, the invention demonstrates excellent performance on multiple datasets and has significant application value.

[0099] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.< / stop> < / stop>

Claims

1. A method for fine-tuning spatial location instructions based on a multimodal large language model, characterized in that, The method includes the following steps: S1. Use dialogue templates to convert the spatial location reasoning dataset into a visual instruction format to obtain a visual spatial location reasoning dataset; S2. Obtain the large language model InternVL as a multimodal large language model, pre-train it on a general dataset to obtain a pre-trained model, and adjust the parameters of the pre-trained model using a low-rank adaptation method based on the visual spatial location reasoning dataset to obtain a trained large language model, which outputs a description corresponding to the spatial task; S3. Introduce a text-based large language model, and optimize the description corresponding to the spatial task based on the large language model.

2. The method for fine-tuning spatial position instructions based on a multimodal large language model according to claim 1, characterized in that, The specific steps of converting the spatial location reasoning dataset into a visual instruction format using a dialogue template are as follows: For a data item in the spatial location reasoning dataset... Where I, T1, T2, R, and D represent image, subject bounding box, object bounding box, subject category label, object category label, spatial relationship, and spatial description, respectively. D has two versions, D1 and D2, corresponding to a concise description and a detailed description, respectively. First, three dialogue templates are used to generate question X. q And answer X r Problem X q And answer X r This constitutes a visual spatial location reasoning dataset.

3. The spatial position instruction fine-tuning method based on a multimodal large language model according to claim 2, characterized in that, The three dialogue templates correspond to question X. q And answer X r These are: questions and answers corresponding to the classification templates for spatial relationships, questions and answers corresponding to the description templates for single spatial relationships, and questions and answers corresponding to the descriptions of open spatial relationships.

4. The method for fine-tuning spatial position instructions based on a multimodal large language model according to claim 3, characterized in that, The specific questions and answers corresponding to the spatial location relationship classification template are as follows: Question: Given an image, select the most appropriate preposition to complete the sentence: Subject category label T1 is left blank in the object category label T2, and the preposition to be filled in the blank is selected from multiple options; Answer: Spatial relation R.

5. The method for fine-tuning spatial position instructions based on a multimodal large language model according to claim 4, characterized in that, The specific questions and answers corresponding to the template for describing a single spatial relationship are as follows: Question: Based on the image, provide a concise text description or phrase describing a single spatial relationship between T1 and T2; Answer: Briefly describe D1.

6. The method for fine-tuning spatial position instructions based on a multimodal large language model according to claim 5, characterized in that, The specific questions and answers for describing open spatial relationships are as follows: Question: Based on the image, provide a detailed text description of the spatial relationship between T1 and T2; Answer: Subject category label T1 Spatial relationship R Object category label T2\nConcise description D1\nDetailed description D2.

7. The method for fine-tuning spatial position instructions based on a multimodal large language model according to claim 1, characterized in that, The low-rank adaptation method is used to adjust the parameters of the pre-trained model as follows: the pre-trained model is connected to the visual language connector and the visual encoder, the visual spatial location inference dataset is input into the visual encoder, and then enters the pre-trained model for training through the visual language connector.

8. The spatial position instruction fine-tuning method based on a multimodal large language model according to claim 5, characterized in that, The specific steps for optimizing the description corresponding to the spatial task based on the large language model are as follows: the large language model uses the description of the spatial task corresponding to the description of a single spatial relationship as the basic description based on the prompt word template, and optimizes the basic description.

9. The method for fine-tuning spatial position instructions based on a multimodal large language model according to claim 8, characterized in that, The specific prompt word template is as follows: based on the image and the concise spatial relationship description between T1 and T2, i.e. the basic description, replace words or phrases in the basic description with adjectives, verbs or synonyms to ensure that no consecutive words in the basic description remain unchanged.

10. The method for fine-tuning spatial position instructions based on a multimodal large language model according to claim 9, characterized in that, The large-scale language model is the Qwen2-7B model.