Image question and answer method based on flexible association control of multi-modal large model
Through flexible association control strategies, the correlation ability of multimodal large models is dynamically adjusted, and the problem of model illusion in image questions and answers is solved, accuracy and creativity are improved, and flexible control of model association capabilities is achieved.
Patent Information
- Application Number
- CN202510201200.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
AI Technical Summary
Existing multimodal large models are prone to hallucinations in image question-and-answer tasks, resulting in the generated responses that do not match the input image, affecting its reliability and credibility.
FlexAC strategy is adopted to dynamically adjust the correlation capability of the model by calculating the cosine distance of the non-association feature representation and correlation feature representation of each layer, filtering key layers, and calculating the correlation control vectors in these layers.
It significantly improves the accuracy and creativity of image Q&A, and realizes flexible control of the association capabilities of multimodal large models, reducing hallucinations and enhancing creativity.
Smart Images

Figure CN119992424A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image question answering, and more specifically, relates to an image question answering method based on flexible association control of a multimodal large model. Background Art
[0002] Multimodal large models (MLLMs) have demonstrated remarkable capabilities in integrating and understanding visual and linguistic information, and can be widely used in many fields, such as image captioning, visual question answering, and storytelling. Their powerful reasoning ability and ability to synthesize multimodal information demonstrate the potential of artificial general intelligence. However, these models are prone to hallucinations and generate false or misleading information that is inconsistent with the input image, which may undermine their reliability and credibility. To address this challenge, many researchers have focused on developing methods to reduce hallucinations in multimodal large models, aiming to improve their robustness and accuracy.
[0003] One of the main reasons for hallucination in large multimodal models is their tendency to form false associations, resulting in generated responses that do not match the given input image. Researchers have proposed various methods to suppress the association tendency to mitigate hallucinations. The two most common approaches are through fine-tuning strategies (such as the HA-DPO method) or decoding strategies (such as the VCD method): fine-tuning methods use preference datasets to align model outputs with user preferences, while decoding strategies reduce hallucinations by exploiting feature differences between states. However, over-suppressing hallucinations may weaken the model's "divergent thinking" ability - an ability that is crucial in tasks that require creativity, such as storytelling, art creation, and problem solving. Existing methods for hallucination control in MLLMs have mainly ignored the impact on the model's creative potential and do not provide the flexibility of two-way control. For example, DPO-based methods heavily rely on a large amount of high-quality manually annotated data and extensive retraining, which limits scalability. Similarly, contrastive decoding methods focus on hallucination reduction and only provide one-way control, lacking adaptability. These findings highlight the need for a flexible method to dynamically control the model's associative behavior to achieve factual accuracy or enhance creativity based on task requirements. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide an image question answering method based on flexible associative control of a multimodal large model. The flexible associative control (FlexAC) strategy is adopted to efficiently and dynamically adjust the associative ability of the multimodal large model without additional training, thereby solving the model hallucination problem in image question answering tasks and significantly enhancing the accuracy and creativity of image question answering.
[0005] In order to achieve the above-mentioned object of the invention, the image question answering method based on flexible association control of a multimodal large model of the present invention comprises the following steps:
[0006] S1: Input the image I and the corresponding text description T that need to be answered into the pre-trained multimodal large model, and the non-correlated features obtained at each layer are expressed as i = 1, 2, ..., L, where L represents the number of layers of the multimodal large model;
[0007] S2: Set customized prompts T′ according to actual needs, input image I and customized prompts T′ into the pre-trained multimodal large model, and express the correlation features obtained at each layer as
[0008] S3: Calculate the non-correlated feature representation of each layer and correlation feature representation The cosine distance Then according to the cosine distance Sort from large to small, select the layer corresponding to the first K cosine distances as the key layer, and record its serial number as i k , k=1,2,…,K;
[0009] S4: For each key layer i k , the associated control vector is calculated using the following formula
[0010]
[0011] S5: Input the image I and question Q into the pre-trained multimodal large model, each key layer i k The feature representation after the association control is generated in the following way
[0012]
[0013] in, represents the key layer i k The original generated feature representation, α represents the control coefficient;
[0014] Then the feature representation Output to the next layer, and finally generate the answer to question Q.
[0015] The present invention is based on an image question answering method with flexible associative control of a multimodal large model. First, a non-associated feature representation of an image and a corresponding text description is generated. Then, an associative feature representation of the image and a customized prompt is generated. The cosine distance between the non-associated feature representation and the associative feature representation of each layer is calculated. The key layer is obtained by screening according to the cosine distance. For each key layer, an associative control vector is calculated through its non-associated feature representation and associative feature representation. When performing image question answering, the corresponding associative control vector is applied in the key layer to perform associative control on the generated feature representation, thereby achieving dynamic control of the creativity and illusion level of the multimodal large model.
[0016] The present invention designs a flexible association control strategy, which is an effective method that does not require training. By analyzing the differences in the internal representation of the model to obtain the control vector, the multimodal large model can be flexibly adjusted according to the task requirements during the image question answering process, whether to follow factual accuracy or enhance creativity, thereby dynamically manipulating the associative ability of the multimodal large model and achieving flexible control between reducing hallucinations and enhancing creativity. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a visualization of the feature representation of the LLaVA multimodal large model simplified by PCA;
[0018] Figure 2 It is the distance map between the correlated features and the uncorrelated features in the middle layer of the LLaVA multimodal large model;
[0019] Figure 3 It is a schematic diagram of layer intervention of associated positioning in the present invention;
[0020] Figure 4 It is the distance map between the correlated features and the uncorrelated features in the middle layer after feature replacement in the LLaVA multimodal large model;
[0021] Figure 5 is a schematic diagram of the associated control vector in the present invention;
[0022] Figure 6 It is a flowchart of a specific implementation method of the image question answering method based on flexible association control of a multi-modal large model of the present invention;
[0023] Figure 7 This is an example diagram of the role of the association control vector in association generation in the present invention. DETAILED DESCRIPTION
[0024] The specific implementation of the present invention is described below in conjunction with the accompanying drawings so that those skilled in the art can better understand the present invention. It should be noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0025] In order to better illustrate the technical solution of the present invention, the technical principle of the present invention is first briefly described. In order to explore the reasons for the association generation of multimodal large models, the present invention generates two data distributions: one from the original output of the model (non-relevance), and the other from the output with induced associated content, which is generated through images and customized prompts. For example, the prompt of the model may be: "Describe the image and include some imaginary, non-existent objects as if they actually exist." According to this method, the present invention constructs a multiple-choice dataset to capture the feature distribution. The model receives an image and generates detailed answers based on the questions, which are compared with two predefined answers of non-relevance (factuality) and relevance (creativity). Features are then extracted from the hidden states corresponding to these inputs to obtain discriminative feature representations at different levels: non-relevance feature representation and correlation feature representation Where i represents the layer number, i = 1, 2, ..., L, and L represents the number of layers of the multimodal large model. Non-correlated feature representations and correlated feature representations can be extracted through different training objectives, where non-correlated feature representations are used to describe the general features of the image, while correlated feature representations adjust image features through specific customized prompts to enhance adaptability to specific tasks.
[0026] Figure 1 This is a visualization of the feature representation of the LLaVA multimodal large model simplified by PCA. Figure 1 As shown in the figure, these features show the internal response of the model under the relevant (red dots) and irrelevant (blue dots) prompts. The similar degree of separation of relevant and irrelevant features in each layer in the first and second rows shows that the multimodal large model is robust to changes in the order of multiple choices. However, the degree of separation of relevant and irrelevant features at different layers shows that the multimodal large model has a clearer expression tendency at deeper levels, reflecting that the associative differentiation deepens with the deepening of the model layers.
[0027] To identify the key layers that control the generation of correlation, the non-correlation features are calculated and association characteristics The cosine distance between and Euclidean distance The cosine distance is used to evaluate the degree of alignment of related and non-related features in direction, while the Euclidean distance is used to measure the difference in their spatial distribution. The specific calculation formula is as follows:
[0028]
[0029] Among them, N represents the number of samples, D represents the feature dimension, They represent the non-correlation features and correlation features of the nth sample in the i-th layer, respectively. Respectively represent characteristics The d-th eigenvalue of .
[0030] Figure 2 is the distance map between the correlated features and the uncorrelated features in the middle layer of the LLaVA multimodal large model. Figure 2 As shown in the figure, there are obvious differences between the cosine distance and Euclidean distance between different layers: the shallow distance is small, indicating that the correlation and non-correlation features are highly aligned, and the main basic feature extraction is performed. The distance in the middle layer increases significantly, indicating that with the introduction of correlation features, the model begins to distinguish between correlation and non-correlation features. The cosine distance in the deep layer decreases, but the Euclidean distance is still high, indicating that the correlation and non-correlation features are aligned in direction, but separated in space, in preparation for output.
[0031] In order to explore whether the deep feature separation is caused by the correlation features introduced in the middle layer, the present invention replaces the correlation features with non-correlation features layer by layer, and evaluates and observes their impact on the subsequent layers. Figure 3 Schematic diagram of layer intervention of associated positioning in the present invention. Figure 3 As shown, in the reasoning process of image question answering, the present invention uses non-correlation features at the target layer m Replace Associative Features At the same time, subsequent layers are allowed to process the modified features. It is expressed as:
[0032]
[0033] Among them, M i () represents the features extracted from the i-th layer parameters of the multimodal large model.
[0034] Then the final layer difference d is calculated by L and the average layer difference To evaluate the impact, the calculation formula is as follows:
[0035]
[0036] Figure 4 It is the distance map between the correlated features and non-correlated features in the middle layer after feature replacement in the LLaVA multimodal large model. Figure 4 As shown in the figure, “Last” and “Rest” represent the final layer difference d L and the average layer difference “Rest-ori” means the original average feature distance without replacement according to Figure 4It can be found that replacing the correlation features of different layers has an impact on subsequent layers: replacement of shallow layers has the least impact on subsequent layers, indicating that shallow layers have limited effect on correlation generation. Replacement of middle layers leads to a significant reduction in feature distance, indicating its key role in generating correlation features. Replacement of deep layers has a smaller impact, indicating that deep layers are mainly used to integrate information rather than introduce correlation elements. Based on these findings, it can be concluded that correlation distinctions are mainly introduced in the middle layers and then propagated to the deep layers. Therefore, in order to effectively control the correlation tendency of the model, the present invention focuses on adjusting the direction and spatial position of these middle layer features, and adjusts the correlation output of the multimodal large model by introducing correlation control vectors. Figure 5 Schematic diagram of the associated control vector in the present invention. Figure 5 As shown, the present invention is based on the associated feature F assoc and non-correlated features F non-assoc An associated control vector v is derived from the difference between assoc By applying the associated control vector v assoc Modify the target feature representation F test , making it possible to dynamically adjust the model’s relevance output, prioritizing creativity or factual accuracy based on demand.
[0037] Based on the above analysis, it can be seen that the multimodal large model perceives information at the shallow level, forms the concept of association at the middle level, and integrates information to generate responses at the deep level. Based on these insights, the present invention proposes an image question answering method based on flexible association control of the multimodal large model, which intervenes in the association ability of the multimodal large model at the middle key layer, thereby achieving flexible control of illusion and creativity. Figure 6 This is a flowchart of a specific implementation of the image question answering method based on flexible association control of a multi-modal large model of the present invention. Figure 6 As shown, the specific steps of the image question answering method based on flexible association control of a multi-modal large model of the present invention include:
[0038] S601: Generate non-associative feature representation:
[0039] The image I and the corresponding text description T that need to be answered are input into the pre-trained multimodal large model, and the non-correlated features obtained at each layer are expressed as i=1,2,…,L, where L represents the number of layers of the multimodal large model.
[0040] S602: Generate correlation feature representation:
[0041] Customized prompts T′ are set according to actual needs, and the image I and customized prompts T′ are input into the pre-trained multimodal large model. The correlation features obtained at each layer are expressed as In order to improve the quality of the correlation feature representation, the image I may be blurred based on a preset random probability.
[0042] In practical applications, in order to improve efficiency, multiple prompt templates can be preset according to actual needs, such as describing virtual objects, inferring causal relationships, or adding story descriptions, and then the prompt template most relevant to the current image question Q is selected as the customized prompt T′.
[0043] S603: Screening key layers:
[0044] According to the previous analysis, the present invention needs to perform correlation control in the middle layer with the most significant correlation, so it is necessary to first screen the key layer. The specific method is: calculate the non-correlation feature representation of each layer and correlation feature representation The cosine distance Then according to the cosine distance Sort from large to small, select the layer corresponding to the first K cosine distances as the key layer, and record its serial number as i k , k=1,2,…,K.
[0045] S604: Calculate the associated control vector:
[0046] For each key layer i k , the associated control vector is calculated using the following formula
[0047]
[0048] S605: Generate question answer:
[0049] The image I and question Q are input into the pre-trained multimodal large model, and each key layer i k The feature representation after the association control is generated in the following way
[0050]
[0051] in, represents the key layer i k The original generated feature representation, α represents the control coefficient.
[0052] The control coefficient α is mainly used to adjust the creativity and factual accuracy of the multimodal large model. In order to balance the relevance strength of the current layer of the model and the weight of user needs, the control coefficient α is calculated using a dynamic balance mechanism in this embodiment, and the formula is as follows:
[0053]
[0054] Among them, λ represents the user control weight, which is used to adjust the sensitivity of the control coefficient. Its value range is [0,1] and is set by the user. represents the key layer i k The cosine distance between the non-correlation feature representation and the correlation feature representation, It represents the average value of the cosine distance between the non-correlated feature representation and the correlated feature representation in all layers. If the cosine distance is greater than the mean, it means that the correlation of this layer is strong, and the model is more inclined to generate associative content, thereby improving creativity. If it is less than the mean, it means that the layer is more conservative and the model is more inclined to generate factual content. β represents the user preference factor, and its value range is [-1,1], where -1 represents complete factuality and 1 represents complete creativity. The specific value of β can be set according to specific needs.
[0055] Then the feature representation Output to the next layer, and finally generate the answer to question Q.
[0056] Figure 7 : is an example diagram of the role of the association control vector in association generation in the present invention. Figure 7 As shown, where -v assoc Indicates control coefficient α<0, +v assoc It indicates that the control coefficient α>0. It can be seen that adjusting the associated control vector can make the multimodal large model generate more factual or creative descriptions, which verifies the effectiveness of the control mechanism.
[0057] In order to better illustrate the technical effect of the present invention, a specific example is used to experimentally verify the present invention.
[0058] Example 1
[0059] The experimental conditions are set in this embodiment: system: Ubuntu 20.04, software: Python 3.8, parameter setting: 1,000 images from the MSCOCO dataset are used to extract the associative control vector. After screening, the 11th, 12th and 13th layers of the multimodal large model are subject to associative control as key layers. The present invention will be divided into two settings, FlexAC-P, which focuses on reducing hallucinations, and FlexAC-C, which enhances associative creativity, to reflect the control of model associations by the present invention. For FlexAC-P, the user preference factor β is set to -1; and for FlexAC-C, the user preference factor β is set to 1. The user control weight λ is set to 1 in both settings. In order to better evaluate the effectiveness of the present invention, this embodiment conducts comparative tests on the following evaluation indicators:
[0060] CHAIR: Caption Hallucination Assessment with Image Relevance (CHAIR) is a metric for evaluating the amount of hallucination generated in image description tasks. It compares the generated description with the real description and calculates the degree of inconsistency between the objects mentioned in the generated text and the real objects. CHAIR contains two metrics: CHAIR S and CHAIR I , and its calculation formula is as follows:
[0061]
[0062] VDAT: This embodiment proposes a new indicator Visual Divergent Association Task (VDAT) to evaluate the model's ability to generate diverse associations. VDAT evaluates the model's divergent thinking ability by presenting an image to the model and asking it to generate multiple nouns that are unrelated to the image, while ensuring that these nouns are also unrelated. Specifically, VDAT contains the following two key components: VDAT I-T :Calculate the semantic divergence between generated nouns and images, and evaluate the effectiveness of the model in generating novel associations. VDAT T-T : Calculate the semantic distance between generated nouns and measure the conceptual diversity of generated content.
[0063] Assume that given an image description C and a noun set T generated by the model, 1 ,t 2 ,…,t P}, the set of nouns extracted from the image description is denoted as C′={c 1 ,c 2 ,…,c Q}. Generate noun t p The semantic similarity between the image description C is defined as:
[0064]
[0065] Among them, ε() is the embedding function that maps words to vector space, p = 1, 2, …, P, q = 1, 2, …, Q.
[0066] The calculation formula of the overall indicator is:
[0067]
[0068] The VDAT index provides a structured approach to assessing the creativity of MLLMs, with lower scores indicating stronger remote association abilities.
[0069] This example compares the performance of different methods in reducing hallucinations and enhancing creativity in order to evaluate the control ability of the present invention. Table 1 is a comparison table of the performance of the present invention and the comparative method in reducing hallucinations and improving creativity in Example 1.
[0070]
[0071] Table 1
[0072] As shown in Table 1, although VCD and HA-DPO can effectively reduce hallucinations (measured by the CHAIR index), they perform poorly on creativity indicators (such as VDAT), inadvertently limiting the associative ability that is crucial in creative tasks. For example, VCD reduces the VDAT of the LLaVA1.5 model by 1. I-T and VDAT T-T The scores improved from 41.78 and 32.12 to 46.72 and 32.71, indicating that its main focus was on reducing illusions at the expense of creativity.
[0073] In contrast, our FlexAC method strikes a balance between reducing hallucinations and preserving creativity by effectively adjusting the associative power of the model using an associative control vector. S and CHAIR I The scores on VDAT and HA-DPO were 44.6 and 14.9 respectively, outperforming VCD and HA-DPO, which are specifically designed to reduce hallucinations. In the creativity-oriented task, FlexAC-C performed better than VDAT and HA-DPO, which were specifically designed to reduce hallucinations. I-T and VDAT T-T The scores on the 3D images were 39.54 and 31.72, respectively. This result was -2.24 and -0.40 higher than the original LLaVA model, respectively. These results show that the present invention can successfully adjust the associative function of the model according to task requirements to achieve the goal of reducing hallucinations or enhancing creativity.
[0074] Example 2
[0075] In order to further verify the effect of the present invention in reducing hallucinations in image question-answering tasks, this embodiment conducted experiments on the POPE benchmark. POPE: Polling-based Object Probing Evaluation (POPE) is an indicator for evaluating the object hallucination ability of multimodal large language models (MLLMs). It evaluates the model's cognitive ability of specific image objects in the form of "yes" or "no" questions and answers, avoiding problems caused by instruction sensitivity. POPE uses three sampling strategies - random sampling (Random), common object sampling (Popular) and adversarial sampling (Adversarial) to evaluate the model's tendency to generate common or co-occurring object hallucinations, and provide stable and reliable hallucination evaluation results. This embodiment constructs POPE on 500 randomly selected MSCOCO validation set images, each of which contains more than 3 real objects and 6 designed questions.
[0076] Table 2 is a comparison table of image classification performance of the present invention and the comparative method in Example 2.
[0077]
[0078] Table 2
[0079] As shown in Table 2, the original LLaVA1.5 has an F1 score of 89.71, 86.79, and 81.76 in the random, common, and adversarial settings, respectively. When optimizing for hallucination reduction, the F1 scores of the proposed FlexAC-P in these three settings reach 90.13, 87.32, and 82.30, respectively, which is better than the baseline model. When optimizing for creativity enhancement, the F1 scores of FlexAC-C drop to 89.30, 86.23, and 81.11, respectively. This result further verifies the ability of FlexAC to effectively control hallucinations by adjusting the associative ability of the model.
[0080] Through the above data analysis and comparison, it can be seen that the present invention is superior to other large-model image question-answering methods based on hallucination control. These results verify the effectiveness and superiority of the present invention.
[0081] Although the above describes the illustrative specific embodiments of the present invention to facilitate those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations using the concept of the present invention are protected.
Claims
1. An image question answering method based on flexible association control of a multimodal large model, characterized in that: The following steps are involved: S1: Input the image I and the corresponding text description T that need to be answered into the pre-trained multimodal large model, and the non-correlated features obtained at each layer are expressed as L represents the number of layers of the multimodal large model; S2: Set customized prompts T′ according to actual needs, input image I and customized prompts T′ into the pre-trained multimodal large model, and express the correlation features obtained at each layer as S3: Calculate the non-correlated feature representation of each layer and correlation feature representation The cosine distance Then according to the cosine distance Sort from large to small, select the layer corresponding to the first K cosine distances as the key layer, and record its serial number as i k , k=1,2,…,K; S4: For each key layer i k , the associated control vector is calculated using the following formula S5: Input the image I and question Q into the pre-trained multimodal large model, each key layer i k The feature representation after association control is generated in the following way in, represents the key layer i k The original generated feature representation, α represents the control coefficient; Then the feature representation Output to the next layer, and finally generate the answer to question Q.
2. The image question answering method according to claim 1, characterized in that: The method for generating the customized prompt T′ in step S2 is: presetting a plurality of prompt templates according to actual needs, and then selecting the prompt template most relevant to the question Q of the current image question and answer as the customized prompt T′.
3. The image question answering method according to claim 1, characterized in that: In the step S2, the image I is blurred based on a preset random probability.
4. The image question answering method according to claim 1, characterized in that: In step S5, the control coefficient α is calculated using the following formula: Among them, λ represents the user control weight, and its value range is [0,1]; represents the key layer i k The cosine distance between the non-correlation feature representation and the correlation feature representation, It represents the average value of the cosine distance between the non-correlated feature representation and the correlated feature representation in all layers; β represents the user preference factor, and its value range is [-1,1], where -1 represents complete factuality and 1 represents complete creativity.