Large language model space understanding ability enhancement method based on multi-modal fusion prompt
Through the cross-modal attention mechanism of target detection, deep semantic segmentation and CLIP model, a multimodal fusion prompt framework is constructed, which solves the shortcomings of multimodal models in understanding visual-language spatial relationships and improves the model's spatial understanding ability and application scope.
Patent Information
- Application Number
- CN202510735369.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-12
AI Technical Summary
Existing multimodal fusion methods find it difficult to effectively capture deep semantic associations across modalities, especially in the processing of visual-linguistic spatial relationships, due to insufficient understanding of the relative positions and motion trajectories of objects, resulting in poor performance of multimodal models in complex spatial understanding.
The target detection algorithm, deep semantic segmentation network and cross-modal attention mechanism of the CLIP model are used to perform multi-level association mapping between images and texts, and the image and text prompts are fused by weights to construct a multimodal fusion prompt framework to improve the spatial understanding ability of the model.
It significantly improves the multimodal large language model's ability to understand visual-spatial information, reduces prediction errors, and expands its application scope in multilingual support and domain adaptation.
Smart Images

Figure CN120633853A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intersection of artificial intelligence and computer vision, and in particular relates to a method for enhancing the spatial understanding capability of a large language model based on multimodal fusion prompts. Background Art
[0002] In recent years, the rapid development of artificial intelligence (AI) has driven its widespread application in fields such as natural language processing (NLP) and computer vision (CV). AI technologies, represented by deep learning, have evolved from traditional machine learning to pre-trained models and then to the era of large models. Large models with billions or even tens of billions of parameters have demonstrated remarkable capabilities, including advanced functions such as command following, contextual learning, and chain reasoning. However, these capabilities are primarily focused on text processing and still have significant limitations when it comes to visual data—they are unable to directly understand images, videos, and other content. To this end, researchers have begun exploring large-scale visual models, but these models are clearly lacking in reasoning capabilities.
[0003] To overcome the limitations of single modality, multimodal large language models (MLLMs) have emerged. These models are focused on areas such as intelligent driving, robotic navigation, and industrial quality inspection, which have a high demand for complex spatial understanding and aim to address the shortcomings of existing multimodal systems in understanding three-dimensional space.
[0004] Compared to traditional multimodal models, MLLMs, thanks to their large number of parameters and multimodal fine-tuning methods, demonstrate unprecedented performance in tasks such as image generation and video understanding. Current research focuses not only on improving performance but also on expanding application dimensions such as multilingual support and domain adaptation, while also exploring model interpretability and efficiency optimization.
[0005] However, in the context of the rapid development of multimodal models, improving spatial understanding capabilities still faces significant challenges. Multimodal data in real-world scenarios often contain complex spatial relationships, requiring models to have cross-modal reasoning capabilities. Specifically, current research faces the following key issues:
[0006] Existing multimodal fusion methods often rely on simple feature concatenation or shallow mapping, making it difficult to capture deep semantic connections across modalities. In particular, when processing visual-linguistic spatial relationships, models lack sufficient understanding of information such as the relative position and motion trajectory of objects. Designing more efficient cross-modal fusion mechanisms is crucial for improving spatial perception.
[0007] To address the above problems, it is urgent to propose a method to enhance the spatial understanding ability of large language models based on multimodal fusion cues. Summary of the Invention
[0008] To solve the above technical problems, the present invention proposes a method for enhancing the spatial understanding ability of a large language model based on multimodal fusion prompts to solve the problems existing in the above-mentioned prior art.
[0009] To achieve the above objectives, the present invention provides a method for enhancing the spatial understanding capability of a large language model based on multimodal fusion prompts, comprising the following steps:
[0010] Constructing a data set, the data set including a plurality of complete image-text pair data; wherein the data set comes from a physical entity;
[0011] Automatically analyzing and processing each image in the data set based on the target detection algorithm to obtain a first image;
[0012] Perform pixel-level analysis on each image in the dataset based on a deep semantic segmentation network to obtain a second image;
[0013] The cross-modal attention mechanism based on the CLIP model establishes an association mapping between each image and text in the dataset to obtain a third image;
[0014] fusing the first image, the second image, and the third image according to corresponding weights to obtain a final image;
[0015] Constructing a prompt template, combining the original text data with the prompt template and inputting it into a large language model to obtain the final text;
[0016] The final image and the final text are combined and input into a multimodal large language model to finally obtain optimized image-text pair data.
[0017] Optionally, the dataset includes two key subsets: counting task and relative depth task.
[0018] Optionally, each image in the data set is automatically analyzed and processed based on a target detection algorithm to obtain a first image, including:
[0019] The target object is identified by a bounding box annotation method, and a first image including the target object identification is obtained.
[0020] Optionally, performing pixel-level analysis on each image in the dataset based on a deep semantic segmentation network to obtain a second image includes:
[0021] Through semantic understanding, each image in the dataset is divided into several semantically meaningful regional units, and a corresponding color identifier is assigned to each regional unit to generate a second image; the second image includes a discriminative color segmentation mask while retaining the visual information of the original image.
[0022] Optionally, a cross-modal attention mechanism based on the CLIP model establishes an association mapping between each image and text in the dataset. The process of obtaining the third image includes:
[0023] The paired image and text are input into the pre-trained CLIP model, and the attention weight between the text and the image is calculated to generate a text-guided visual heat map, which is the third image.
[0024] The present invention also provides a large language model spatial understanding ability enhancement system based on multimodal fusion prompts, which is used to implement the method described above, including: a data acquisition module, an image processing module, an image fusion module, a text processing module and an image and text optimization module;
[0025] The data acquisition module is used to construct a data set, which includes a plurality of complete image-text pair data; wherein the data set comes from a physical entity;
[0026] The image processing module is used to analyze and process each image in the dataset based on the target detection algorithm, deep semantic segmentation network and cross-modal attention mechanism of the CLIP model to obtain the corresponding image;
[0027] The image fusion module is used to fuse corresponding images according to corresponding weights to obtain a final image;
[0028] The text processing module is used to construct a prompt template, combine the original text data with the prompt template and input it into the large language model to obtain the final text;
[0029] The image-text optimization module is used to combine the final image and the final text and input them into a multimodal large language model to ultimately obtain optimized image-text pair data.
[0030] Optionally, the image processing module includes:
[0031] The first processing unit is used to automatically analyze and process each image in the data set based on the target detection algorithm, identify the target object by a bounding box annotation method, and obtain a first image including the target object identification;
[0032] The second processing unit is used to perform pixel-level analysis on each image in the dataset based on a deep semantic segmentation network, divide each image into a number of semantically meaningful regional units through semantic understanding, and assign a corresponding color identifier to each regional unit to generate a second image;
[0033] The third processing unit is used to establish an association mapping between each image and text in the dataset based on the cross-modal attention mechanism of the CLIP model, input the paired images and texts into the pre-trained CLIP model, and generate a text-guided visual heat map by calculating the attention weight between the text and the image. The visual heat map is the third image.
[0034] The present invention also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method.
[0035] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method when the computer program is executed by a processor.
[0036] The present invention also provides a computer program product, comprising a computer program, which implements the steps of the method when executed by a processor.
[0037] Compared with the prior art, the present invention has the following advantages and technical effects:
[0038] The present invention proposes a multimodal prompt framework. This framework significantly improves the model's ability to understand visual spatial information through dual optimization of image prompts and text prompts. In terms of image prompts, the present invention constructs a fusion prompt mechanism that integrates three technologies: target detection, semantic segmentation, and attention visualization. Through the weight distribution strategy, the advantages of different visual prompt methods are complementary, enabling the model to more accurately capture key spatial information in the image. In terms of text prompts, the present invention proposes a fine-grained thinking chain guidance method. This method guides the model to focus on specific objects and their spatial relationships in the image through structured problem decomposition and progressive reasoning prompts, thereby improving the accuracy of understanding detailed information. The present invention improves the spatial understanding ability of large multimodal models through sufficient graphic and text dual-modal prompts, and can reduce the error level of model predictions.
[0039] The method proposed in this invention not only improves the spatial understanding ability of the multimodal large language model, but also expands its performance in application dimensions such as multilingual support and domain adaptation, enabling it to better adapt to the complex spatial understanding needs of different language environments and specific fields, further expanding the application scope and value of the multimodal large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0041] Figure 1Flowchart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0042] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0043] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0044] Example 1
[0045] like Figure 1 As shown, this embodiment provides a method for enhancing the spatial understanding ability of a large language model based on multimodal fusion prompts. The method proposed in this embodiment can be applied in the fields of intelligent driving, robot navigation, and industrial quality inspection. In the field of intelligent driving, the application of the method of this embodiment can solve the problem that the on-board camera has difficulty in accurately identifying the positional relationship between objects in the image, which increases the delay in the vehicle's lane change decision; in the field of robot navigation, especially when executing instructions such as "insert part A into the upper left corner hole of component B", the assembly error caused by the inaccurate geometric mapping between the visual scene and the language instruction can also be solved by the method for enhancing the spatial understanding ability of the large language model proposed in this embodiment; in the industrial quality inspection environment, especially in the chip pin spacing detection process, the application of the method of this embodiment can improve the missed detection caused by the ineffective combination of optical image analysis and spatial constraints in the process specification text. The above application scenarios reflect the lack of spatial understanding ability of the current multimodal system, and the method proposed in this embodiment can improve the spatial perception and processing capabilities of the system in these key areas.
[0046] The method of this embodiment specifically includes the following steps:
[0047] Constructing a dataset, the dataset including several complete image-text pairs; wherein the dataset is derived from a physical entity, which may be an object entity in fields such as intelligent driving, robot navigation, and industrial quality inspection;
[0048] Automatically analyzing and processing each image in the data set based on the target detection algorithm to obtain a first image;
[0049] Perform pixel-level analysis on each image in the dataset based on a deep semantic segmentation network to obtain a second image;
[0050] The cross-modal attention mechanism based on the CLIP model establishes an association mapping between each image and text in the dataset to obtain a third image;
[0051] fusing the first image, the second image, and the third image according to corresponding weights to obtain a final image;
[0052] Constructing a prompt template, combining the original text data with the prompt template and inputting it into a large language model to obtain the final text;
[0053] The final image and the final text are combined and input into a multimodal large language model to finally obtain optimized image-text pair data, wherein the optimized image-text pair data has enhanced spatial understanding capability.
[0054] As a specific implementation method, the following steps are specifically included:
[0055] Step 1: We select task-specific data from the publicly available BLINE dataset, including two key subsets: counting and relative depth. The counting task data is selected from question samples in the TallyQA dataset, while the relative depth task data is derived from the Depth in the Wild dataset. The constructed dataset contains complete image-text pairs, where each data sample consists of an image and its corresponding text description.
[0056] Furthermore, the BLINE dataset is a novel benchmark dataset, derived from physical objects, designed to comprehensively evaluate the visual perception capabilities of large, multimodal language models. It measures the performance of these models across a range of tasks, from low-level pattern matching to mid-level spatial reasoning to high-level visual understanding. It covers a variety of scenarios, from indoor home scenes to outdoor urban and natural environments.
[0057] Step 2: Use advanced target detection algorithms to automatically analyze and process each image in the dataset, accurately identify the main target objects contained in the image through bounding box annotation, and generate the first image;
[0058] Specifically, the YOLOv5 target detection algorithm is adopted. First, the original image dataset is preprocessed, which includes adjusting the image size to the input size required by the YOLOv5 model. Then, by loading the pre-trained YOLOv5 weights, the network parameters are initialized to ensure that the model can recognize target objects of multiple categories. The preprocessed image is then input into the YOLOv5 network, and its powerful feature extraction capability is used to automatically detect the main target objects in the image. YOLOv5 uses a single network to directly predict the position of the bounding box and the probability of its category, realizing end-to-end target detection. In order to obtain accurate bounding box annotations, in the non-maximum suppression (NMS) step, the best results are screened according to the confidence score of the predicted bounding box, thereby effectively removing redundant and overlapping bounding boxes. Finally, the image with precise bounding box annotations generated after processing by the YOLOv5 model is saved as the first image output;
[0059] Step 3: Use a deep semantic segmentation network to perform pixel-level analysis on the images in the dataset, dividing the images into several regions with clear semantic meaning through semantic understanding. Each semantic category is assigned a specific color identifier, and the generated second image has a highly discriminative color segmentation mask while preserving the visual information of the original image.
[0060] Specifically, this embodiment uses a deep semantic segmentation network (SAM) for pixel-level analysis. First, the input raw image is preprocessed, including resizing and format conversion to match the requirements of the SAM model, ensuring that the image can be effectively processed by the model. Next, the network is initialized by loading the pre-trained SAM model weights. This step is crucial for ensuring that the model accurately understands and distinguishes different semantic regions in the image. After initialization, the preprocessed image is fed into the SAM model, which leverages its powerful feature extraction and understanding capabilities to automatically identify and segment regions with clear semantic meaning. The SAM model not only accurately captures the detailed information in the image but also understands the actual meaning represented by these details, thereby achieving efficient semantic segmentation. To ensure that the generated second image is highly distinguishable and easy to analyze, a specific color identifier is assigned to each identified semantic category. This process involves determining the category to which each pixel belongs and assigning a corresponding color based on that category. This not only enhances the visual contrast between different semantic regions but also makes the resulting color segmentation mask more intuitive and easy to understand. At the same time, when generating the segmentation mask, we take measures to preserve the visual information of the original image, ensuring that all other visual features (such as texture and shape) are fully presented in addition to the added color identifier. This approach facilitates subsequent analysis because it combines the visual authenticity of the original image with the structured information brought by semantic segmentation.
[0061] Step 4: Leveraging the CLIP model's cross-modal attention mechanism, we establish an association mapping between key semantic elements in the text description and image regions. The paired image and text are fed into the pre-trained CLIP model, and the attention weights between the text tokens and the image patches are calculated to generate a text-guided visual heatmap, the third image. This method employs a progressive color coding strategy to visually annotate the attention regions, forming an explicit correspondence between key concepts in the text and specific image regions.
[0062] Specifically, before feeding the image and text into the CLIP model, the image is first segmented into a series of small image patches, each representing a local region of the image. Simultaneously, the text is converted into a series of tokens using a tokenizer, corresponding to different words or phrases in the text. The CLIP model then uses a transformer architecture to calculate attention weights between each pair of text tokens and image patches. This process involves numerous matrix operations and the application of activation functions to capture the most significant correlations between the text and the image. The heatmap visually demonstrates which image regions are receiving strong attention from specific words or concepts in the text description. To enhance this visualization, a progressive color coding strategy is employed to annotate the attention regions. Rather than simply using a single color to represent all attended regions, this approach uses varying color intensities based on the size of the attention weight, highlighting regions of higher attention with brighter colors. This approach not only clearly demonstrates the explicit correspondence between key concepts in the text and specific image regions, but also reflects the strength of this correspondence, providing a richer hierarchy of information.
[0063] Step 5: By designing a structured prompt template, the original text content is combined with the large language model instructions to construct an optimized text with a clear task orientation;
[0064] Specifically, we first need to design a structured prompt template, have an in-depth understanding of the target task, and conduct a detailed analysis of the type of text to be processed. Determine the key information types that need to be extracted or converted from the original text based on the needs, and design structured prompt templates. These templates not only contain fixed guiding words and sentences, but also reserve variable positions for inserting specific text content. The next step is to embed the specific text content into the pre-designed prompt template. Ensure that the text can be accurately parsed and understood by the large language model. The large model will identify the specific task requirements based on the content of the prompt and adjust its output strategy accordingly. In this way, the model can not only understand the basic meaning of the input text, but also accurately complete the specified task according to the guidance in the prompt;
[0065] Step 6: Fuse the three images processed in the above steps according to different weights to obtain the final image;
[0066] Step 7: Input the processed text and image into the multimodal large language model to generate optimized image-text pairs.
[0067] This embodiment also provides a large language model spatial understanding ability enhancement system based on multimodal fusion prompts, which is used to implement the method described above, including: a data acquisition module, an image processing module, an image fusion module, a text processing module, and an image and text optimization module;
[0068] The data acquisition module is used to construct a data set, which includes a plurality of complete image-text pair data; wherein the data set comes from a physical entity;
[0069] The image processing module is used to analyze and process each image in the dataset based on the target detection algorithm, deep semantic segmentation network and cross-modal attention mechanism of the CLIP model to obtain the corresponding image;
[0070] The image fusion module is used to fuse corresponding images according to corresponding weights to obtain a final image;
[0071] The text processing module is used to construct a prompt template, combine the original text data with the prompt template and input it into the large language model to obtain the final text;
[0072] The image-text optimization module is used to combine the final image and the final text and input them into a multimodal large language model to ultimately obtain optimized image-text pair data.
[0073] It is feasible that the image processing module includes:
[0074] The first processing unit is used to automatically analyze and process each image in the data set based on the target detection algorithm, identify the target object by a bounding box annotation method, and obtain a first image including the target object identification;
[0075] The second processing unit is used to perform pixel-level analysis on each image in the dataset based on a deep semantic segmentation network, divide each image into a number of semantically meaningful regional units through semantic understanding, and assign a corresponding color identifier to each regional unit to generate a second image;
[0076] The third processing unit is used to establish an association mapping between each image and text in the dataset based on the cross-modal attention mechanism of the CLIP model, input the paired images and texts into the pre-trained CLIP model, and generate a text-guided visual heat map by calculating the attention weight between the text and the image. The visual heat map is the third image.
[0077] Example 2
[0078] This embodiment further provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method.
[0079] Example 3
[0080] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method when executed by a processor.
[0081] Example 4
[0082] This embodiment also provides a computer program product, including a computer program, which implements the steps of the method when executed by a processor.
[0083] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for enhancing the spatial understanding ability of a large language model based on multimodal fusion prompts, characterized in that: The following steps are involved: Constructing a data set, the data set including a plurality of complete image-text pair data; wherein the data set comes from a physical entity; Automatically analyzing and processing each image in the data set based on the target detection algorithm to obtain a first image; Perform pixel-level analysis on each image in the dataset based on a deep semantic segmentation network to obtain a second image; The cross-modal attention mechanism based on the CLIP model establishes an association mapping between each image and text in the dataset to obtain a third image; fusing the first image, the second image, and the third image according to corresponding weights to obtain a final image; Constructing a prompt template, combining the original text data with the prompt template and inputting it into a large language model to obtain the final text; The final image and the final text are combined and input into a multimodal large language model to finally obtain optimized image-text pair data.
2. The method according to claim 1, characterized in that The dataset includes two key subsets: counting task and relative depth task.
3. The method according to claim 1, characterized in that Each image in the dataset is automatically analyzed and processed based on the target detection algorithm. The process of obtaining the first image includes: The target object is identified by a bounding box annotation method, and a first image including the target object identification is obtained.
4. The method according to claim 1, wherein The process of performing pixel-level analysis on each image in the dataset based on the deep semantic segmentation network to obtain the second image includes: Through semantic understanding, each image in the dataset is divided into several semantically meaningful regional units, and a corresponding color identifier is assigned to each regional unit to generate a second image; the second image includes a discriminative color segmentation mask while retaining the visual information of the original image.
5. The method according to claim 1, characterized in that The cross-modal attention mechanism based on the CLIP model establishes an association mapping between each image and text in the dataset. The process of obtaining the third image includes: The paired image and text are input into the pre-trained CLIP model, and the attention weight between the text and the image is calculated to generate a text-guided visual heat map, which is the third image.
6. A system for enhancing spatial understanding capabilities of a large language model based on multimodal fusion prompts, characterized in that: Used to implement the method according to any one of claims 1 to 5, comprising: a data acquisition module, an image processing module, an image fusion module, a text processing module and an image and text optimization module; The data acquisition module is used to construct a data set, which includes a plurality of complete image-text pair data; wherein the data set comes from a physical entity; The image processing module is used to analyze and process each image in the dataset based on the target detection algorithm, deep semantic segmentation network and cross-modal attention mechanism of the CLIP model to obtain the corresponding image; The image fusion module is used to fuse corresponding images according to corresponding weights to obtain a final image; The text processing module is used to construct a prompt template, combine the original text data with the prompt template and input it into the large language model to obtain the final text; The image-text optimization module is used to combine the final image and the final text and input them into a multimodal large language model to ultimately obtain optimized image-text pair data.
7. The system according to claim 6, characterized in that The image processing module includes: The first processing unit is used to automatically analyze and process each image in the data set based on the target detection algorithm, identify the target object by a bounding box annotation method, and obtain a first image including the target object identification; The second processing unit is used to perform pixel-level analysis on each image in the dataset based on a deep semantic segmentation network, divide each image into a number of semantically meaningful regional units through semantic understanding, and assign a corresponding color identifier to each regional unit to generate a second image; The third processing unit is used to establish an association mapping between each image and text in the dataset based on the cross-modal attention mechanism of the CLIP model, input the paired images and texts into the pre-trained CLIP model, and generate a text-guided visual heat map by calculating the attention weight between the text and the image. The visual heat map is the third image.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Cited By
Retrieval method and system based on semantic keyword classification and multi-language intelligent icons
CN120994855A
Visual inspection method and system based on industrial personal computer
CN121190425A