Vision-text alignment system and method capable of relieving illusion

By introducing semantic supervision, alignment loss and semantic injection modules into the video-language model, the problem of visual and text features is solved, the accuracy and consistency of long video comprehension is improved, the hallucination phenomenon is reduced, and the computational overhead is reduced.

CN120472366APending Publication Date: 2025-08-12JINAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510550700.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing video-language models have problems with visual and text features in long video comprehension, resulting in hallucination phenomena and performance degradation, especially poor performance over long-term periods.

Method used

By introducing a semantic supervision module to generate text features, and aligning visual and text features in the intermediate feature space through the alignment loss module, using a frozen large language model for context reasoning, combining the semantic injection module to store spatiotemporal features and generate query embeddings through a long-term memory integration mechanism to ensure the consistency between vision and text features.

Benefits of technology

It effectively reduces the illusion phenomenon in the model output, improves the accuracy and semantic consistency of the model, enhances the generalization ability in long video comprehension tasks, reduces computational overhead and prevents overfitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472366A_ABST
    Figure CN120472366A_ABST
Patent Text Reader

Abstract

The invention provides a vision-text alignment system and method capable of relieving illusion. According to the vision-text alignment system and method, text features are generated through a semantic supervision module; taking the text features as targets of visual feature alignment, and aligning the visual features and the text features in the middle feature space by using a contrast learning strategy through an alignment loss module; performing context reasoning on the multi-modal features processed by the alignment loss module through a frozen large language model, and generating a reasoning text in an autoregression manner; storing the spatio-temporal characteristics by using a long-term memory integration mechanism through a semantic injection module, and generating query embedding through a query enhancement mechanism; according to the method, the illusion problem is relieved by explicitly aligning video and text representation in the middle feature space, so that the generalization ability of the model is improved; semantic supervision and an intermediate feature alignment mechanism are introduced, so that the illusion phenomenon in model output is reduced, and semantic consistency between visual and text modes in an intermediate feature space is effectively enhanced through alignment loss of similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and computer vision technology, and in particular to a vision-text alignment system and method capable of alleviating hallucinations. Background Art

[0002] Video understanding is widely used in many fields such as intelligent transportation systems, video surveillance, video retrieval, video subtitle generation, and cross-modal reasoning. Existing long video content understanding is mainly achieved through video-language models. However, existing video-language models still face the persistent challenge of visual hallucinations in video understanding tasks, that is, the responses generated by the model may be factually inaccurate, logically inconsistent, or completely unrelated to the input or real-world knowledge.

[0003] To date, research on mitigating model hallucinations has primarily focused on methods. For example, Video-LLaMA1 enables large language models to understand both visual and auditory content in videos, addressing these challenges by capturing temporal visual variations and integrating audiovisual signals. However, these studies have primarily focused on the hallucination problem over short timescales, typically limited to videos lasting 30 seconds or less.

[0004] Recently, progress has been made in long video understanding, for example by developing memory-efficient architectures that significantly reduce memory usage and accelerate the processing of large numbers of video frames. Specifically, the Memory-Augmented Large Multimodal Model (MA-LMM) extracts textual features from question-answer pairs using a pre-trained text encoder, while visual features are aggregated from video frames using a state-of-the-art (SOTA) vision model. These textual and visual features are then fed into a large language model for task-oriented text decoding. Thanks to its efficient architecture, MA-LMM achieves state-of-the-art performance in long video understanding. However, these methods still exhibit suboptimal performance by ignoring the misalignment between visual and textual features, a key issue in improving the capabilities of multimodal models. Specifically, the features encoded by the visual encoder often deviate from the semantic space of LLMs. This misalignment prevents LLMs from fully utilizing visual features, leading to hallucination artifacts and degraded performance. This problem is particularly severe in long video scenarios, where the expanded visual features exacerbate the misalignment. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a visual-text alignment system and method that can alleviate hallucinations. The present invention effectively solves the problem of misalignment of visual and text features that occurs in the existing technology when processing video and text data, thereby reducing the hallucination phenomenon in the model output and improving the accuracy and semantic consistency of the model.

[0006] The technical solution of the present invention is: a visual-text alignment system capable of alleviating hallucinations, comprising:

[0007] Semantic supervision module, used to generate text features and make visual features consistent with text features;

[0008] An alignment loss module is used to align visual and textual features in an intermediate feature space through a contrastive learning strategy to obtain multimodal features;

[0009] A frozen large language model, wherein the frozen large language model is used to perform contextual reasoning on the multimodal features processed by the alignment loss module and to autoregressively generate inference text;

[0010] The semantic injection module is used to store the spatiotemporal features extracted by the video encoder through a long-term memory integration mechanism, generate query embeddings through a query enhancement mechanism, and connect the visual features with the frozen large language model through the query embeddings.

[0011] Preferably, the semantic supervision module extracts text features through a pre-trained text encoder, and uses the extracted text features as targets for alignment with visual features generated by the query transformer, thereby ensuring consistency between the visual features and the text features.

[0012] Preferably, the expression for extracting text features by pre-trained text encoder is:

[0013] f text =TextEncoder(T);

[0014] Where, f text is the extracted text feature, T is the input text, and TextEncoder represents the text encoder.

[0015] As an example, the alignment loss module calculates the alignment loss L between the visual features and the text features by using cosine similarity. aligned ,Right now:

[0016]

[0017] Where, f visual represents the visual features, f text Represents text features.

[0018] Preferably, the total loss function of the alignment loss module is composed of task loss and alignment loss, that is:

[0019] L final =L task +L aligned ;

[0020] Where, L task For mission loss; L aligned is the alignment loss.

[0021] As a preference, the task loss L task The expression is:

[0022] L task =-∑ i logP(A i |V,P);

[0023] In the formula, P represents the prompt text, A i represents the i-th target text, and V represents the input video features.

[0024] Preferably, the frozen large language model uses visual features as prompts and combines them with tokenized text to generate inferred text.

[0025] Preferably, the semantic injection module stores spatiotemporal features through a long-term memory integration mechanism and generates query embeddings through a query enhancement mechanism to ensure the effective capture of short-term and long-term dependencies.

[0026] As an example, the semantic injection module integrates the spatiotemporal features {f1, f2, ..., f T} is stored in the visual memory and generated into query embeddings {q1,q2,...,q M}; The query transformer enhances the query through the self-attention mechanism, and the query transformer integrates temporal features through the cross-attention mechanism.

[0027] Preferably, the long-term memory integration mechanism compresses spatiotemporal features through cosine similarity to ensure the retention of key information, and the query enhancement mechanism enhances the interaction of query embeddings through self-attention and cross-attention.

[0028] The present invention also provides a visual-text alignment method capable of alleviating hallucinations, comprising the following steps:

[0029] S1), generate text features f through semantic supervision module text ; and use query transformer to extract visual features f from the video visual ; and the text feature f text As the visual feature f visual The goal of alignment is to ensure consistency between visual features and text features;

[0030] S2), aligning visual and textual features in the intermediate feature space using a contrastive learning strategy through an alignment loss module;

[0031] S3) Perform contextual reasoning on the multimodal features processed by the alignment loss module through a frozen large language model, and generate reasoning text in an autoregressive manner;

[0032] S4) The semantic injection module uses the long-term memory integration mechanism to store spatiotemporal features, and generates query embeddings through the query enhancement mechanism.

[0033] The beneficial effects of the present invention are:

[0034] 1. This paper alleviates the hallucination problem by explicitly aligning video and text representations in an intermediate feature space, thereby improving the generalization ability of the model;

[0035] 2. By introducing semantic supervision and an intermediate feature alignment mechanism, this paper effectively solves the problem of misalignment between visual and textual features that occurs in existing technologies when processing video and text data, thereby reducing the hallucination phenomenon in the model output. Through the similarity alignment loss, it can effectively enhance the semantic consistency between visual and textual modalities in the intermediate feature space.

[0036] 3. The present invention performs well in long video understanding tasks, can effectively handle complex visual scenes and ambiguous text prompts, and provide more accurate and semantically consistent video text representation;

[0037] 4. The present invention generates text output through a frozen language model without the need for additional fine-tuning, reducing computational overhead and preventing overfitting. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 Schematic diagram of the framework of the system of the present invention;

[0039] Figure 2 This is a visualization of the comparison results between the present invention and existing methods on the long video recognition task;

[0040] Figure 3 This is a visualization of the qualitative results of the long video recognition task on the LVU dataset of the present invention;

[0041] Figure 4 This is a performance comparison chart of different loss functions and regularization strategies on the LVU dataset in an embodiment of the present invention. DETAILED DESCRIPTION

[0042] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings:

[0043] Example 1

[0044] like Figure 1 As shown, this embodiment provides a visual-text alignment system capable of alleviating hallucinations, including:

[0045] Semantic supervision module, used to generate text features and make visual features consistent with text features;

[0046] An alignment loss module is used to align visual and textual features in an intermediate feature space through a contrastive learning strategy to obtain multimodal features;

[0047] A frozen large language model, wherein the frozen large language model is used to perform contextual reasoning on the multimodal features processed by the alignment loss module and to autoregressively generate inference text;

[0048] The semantic injection module is used to store the spatiotemporal features extracted by the video encoder through a long-term memory integration mechanism, generate query embeddings through a query enhancement mechanism, and connect the visual features with the frozen large language model through the query embeddings. In this embodiment, the spatiotemporal features are frame sequence features extracted by the video encoder.

[0049] As a preferred embodiment of this invention, the semantic supervision module extracts text features through a pre-trained text encoder and uses the extracted text features as the target for aligning the visual features generated by the query transformer, ensuring the consistency of visual features with text features, enhancing the alignment of visual and text representations, and ensuring the robustness of cross-modal interaction. In this embodiment, the text encoder is BERT; for a given text T = [question + answer], the expression for extracting text features through the pre-trained text encoder is:

[0050] f text =TextEncoder(T);

[0051] Where, f text is the extracted text feature, T is the input text, and TextEncoder represents the text encoder.

[0052] For a given video, a query transformer Q-Former is used to extract visual features f visual The query transformer Q-Former processes spatiotemporal embeddings from the visual memory bank to generate visual representations aligned with text features.

[0053] As preferred in this embodiment, the alignment loss module explicitly enhances the consistency of visual and text features in the shared semantic space by introducing CLIP loss, thereby reducing the hallucination phenomenon. The alignment loss module calculates the alignment loss L between visual features and text features by cosine similarity. aligned ,Right now:

[0054]

[0055] Where, f visual represents the visual features, f text Represents text features.

[0056] Preferably, the total loss function of the alignment loss module consists of task loss and alignment loss. This unified loss enforces fine-grained alignment across modalities, solves the hallucination problem, and enhances long-term video-language understanding, namely:

[0057] L final =L task +L aligned ;

[0058] Where, L task For mission loss; L aligned is the alignment loss.

[0059] As a preference, the task loss L task The expression is:

[0060] L task =-∑ i logP(A i |V,P);

[0061] In the formula, P represents the prompt text, A i represents the i-th target text, and V represents the input video features.

[0062] As preferred in this embodiment, the frozen large language model reduces computational overhead and prevents overfitting to a specific data set. The frozen large language model uses visual features as prompts and is combined with tokenized text.

[0063] As a preferred embodiment of the present invention, the semantic injection module stores spatiotemporal features through a long-term memory integration mechanism and generates query embeddings through a query enhancement mechanism to ensure the effective capture of short-term and long-term dependencies. The semantic injection module stores spatiotemporal features {f1, f2, ..., f T} is stored in the visual memory and generated into query embeddings {q1,q2,...,q M The query transformer enhances queries through a self-attention mechanism, and integrates temporal features through a cross-attention mechanism. The long-term memory integration mechanism compresses spatiotemporal features through cosine similarity to ensure the preservation of key information. The query enhancement mechanism enhances the interaction of query embeddings through self-attention and cross-attention.

[0064] Example 2

[0065] This embodiment provides a visual-text alignment method capable of alleviating hallucinations, comprising the following steps:

[0066] S1), generate text features f through semantic supervision module text ; and extract visual features f from the video using the query transformer visual ; and the text feature f text As the visual feature f visual The goal of alignment is to ensure consistency between visual features and text features;

[0067] S2), aligning visual and textual features in the intermediate feature space using a contrastive learning strategy through an alignment loss module;

[0068] The alignment loss module introduces CLIP loss to explicitly enhance the consistency of visual and text features in the shared semantic space and reduce the hallucination phenomenon. The alignment loss module calculates the alignment loss L between visual features and text features through cosine similarity. aligned ,Right now:

[0069]

[0070] Where, f visual represents the visual features, f text Represents text features.

[0071] Preferably, the total loss function of the alignment loss module consists of task loss and alignment loss. This unified loss enforces fine-grained alignment across modalities, solves the hallucination problem, and enhances long-term video-language understanding, namely:

[0072] L final =L task +L aligned ;

[0073] Where, L task For mission loss; L aligned is the alignment loss.

[0074] As a preference, the task loss L task The expression is:

[0075] L task =-∑ i logP(A i |V,P);

[0076] In the formula, P represents the prompt text, A i represents the i-th target text, and V represents the input video features.

[0077] S3) Perform contextual reasoning on the multimodal features processed by the alignment loss module through a frozen large language model, and generate reasoning text in an autoregressive manner;

[0078] S4) The semantic injection module uses the long-term memory integration mechanism to store spatiotemporal features, and generates query embeddings through the query enhancement mechanism.

[0079] The semantic injection module stores spatiotemporal features through the long-term memory integration mechanism and generates query embedding through the query enhancement mechanism to ensure the effective capture of short-term and long-term dependencies. T} is stored in the visual memory and generated into query embeddings {q1,q2,...,q M The query transformer enhances queries through a self-attention mechanism, and integrates temporal features through a cross-attention mechanism. The long-term memory integration mechanism compresses spatiotemporal features through cosine similarity to ensure the preservation of key information. The query enhancement mechanism enhances the interaction of query embeddings through self-attention and cross-attention.

[0080] Example 3

[0081] This example experiments on the LVU dataset, using the Vicuna-7B model as the backbone large language model. Training is performed for 20 epochs, with a learning rate of 1×10⁻4, the AdamW optimizer, a batch size of 64, and input frames resized to 224×224 resolution.

[0082] As shown in Table 1, Example 1 significantly outperforms the existing baseline model in the relationship, year, scene, and screenplay tasks in the LVU benchmark test, especially in the year and screenplay tasks, with improvements of 5.8% and 3.2% respectively.

[0083] Table 1 Comparison results between Example 1 and the latest method on the LVU dataset

[0084]

[0085] Table 2 shows the performance comparison study of different LLMs on the relational subtask based on the LVU dataset.

[0086]

[0087] like Figure 2 、 3 As shown, Example 1 demonstrates accuracy and robustness when handling complex semantic categories such as "romance," "airport," and "discussion." Specific examples include successfully identifying emotions and relationship dynamics ("romance" scene), accurately capturing the airport environment and related activities ("airport" scene), and identifying conversation scenes in different settings ("discussion" scene).

[0088] like Figure 4 As shown in Figure 2, comparative experiments using different loss functions and regularization strategies validate the effectiveness of the alignment mechanism and regularization strategy proposed in Example 1 in improving model performance. Combining CLIP loss with regularization achieves the highest Top-1 accuracy of 61.5%.

[0089] The above embodiments and descriptions are only for explaining the principles and best embodiments of the present invention. Without departing from the spirit and scope of the present invention, the present invention may be subject to various changes and improvements, which shall fall within the scope of the invention to be protected.

Claims

1. A visual-text alignment system capable of alleviating hallucinations, characterized in that: The system comprises: Semantic supervision module, used to generate text features and make visual features consistent with text features; An alignment loss module for aligning visual and textual features in an intermediate feature space via a contrastive learning strategy; A frozen large language model, wherein the frozen large language model is used to perform contextual reasoning on the multimodal features processed by the alignment loss module and to autoregressively generate inference text; The semantic injection module is used to store the spatiotemporal features extracted by the video encoder through a long-term memory integration mechanism, generate query embeddings through a query enhancement mechanism, and connect the visual features with the frozen large language model through the query embeddings.

2. The visual-text alignment system capable of alleviating hallucinations according to claim 1, characterized in that: The semantic supervision module extracts text features through a pre-trained text encoder and uses the extracted text features as targets for aligning visual features generated by a query transformer to ensure consistency between visual features and text features.

3. The visual-text alignment system capable of alleviating hallucinations according to claim 2, characterized in that: The expression for extracting text features through pre-trained text encoder is: f text =TextEncoder(T); Where, f text is the extracted text feature, T is the input text, and TextEncoder represents the text encoder.

4. The visual-text alignment system capable of alleviating hallucinations according to claim 1, characterized in that: The alignment loss module calculates the alignment loss L between visual features and text features through cosine similarity aligned ,Right now: Where, f visual represents the visual features, f text Represents text features.

5. The visual-text alignment system capable of alleviating hallucinations according to claim 4, characterized in that: The total loss function of the alignment loss module is composed of task loss and alignment loss, namely: L final =L task +L aligned ; Where, L task For mission loss; L aligned is the alignment loss.

6. The visual-text alignment system capable of alleviating hallucinations according to claim 5, characterized in that: The mission loss L task The expression is: L task =-∑ i logP(A i |V,P); In the formula, P represents the prompt text, A i represents the i-th target text, and V represents the input video features.

7. The visual-text alignment system capable of alleviating hallucinations according to claim 1, characterized in that: The semantic injection module stores spatiotemporal features through a long-term memory integration mechanism and generates query embeddings through a query enhancement mechanism.

8. The visual-text alignment system capable of alleviating hallucinations according to claim 7, characterized in that: The semantic injection module integrates the spatiotemporal features {f1,f2,...,f T } is stored in the visual memory and generated into query embeddings {q1,q2,...,q M }; The query transformer enhances the query through the self-attention mechanism, and the query transformer integrates temporal features through the cross-attention mechanism.

9. The visual-text alignment system capable of alleviating hallucinations according to claim 8, characterized in that: The long-term memory integration mechanism compresses spatiotemporal features through cosine similarity to ensure the retention of key information, and the query enhancement mechanism enhances the interaction of query embeddings through self-attention and cross-attention.

10. A visual-text alignment method capable of alleviating hallucinations, characterized in that: The method uses the system according to any one of claims 1 to 9 to achieve visual-text alignment, and the method comprises the following steps: S1), generate text features f through semantic supervision module text ; and extract visual features f from the video using the query transformer visual ; and the text feature f text As the visual feature f visual The goal of alignment is to ensure consistency between visual features and text features; S2) Aligning visual and text features in the intermediate feature space using a contrastive learning strategy through an alignment loss module; the alignment loss module calculates the alignment loss L between visual features and text features through cosine similarity aligned ,Right now: Where, f visual represents the visual features, f text Represents text features; S3) Perform contextual reasoning on the multimodal features processed by the alignment loss module through a frozen large language model, and generate reasoning text in an autoregressive manner; S4) Using the long-term memory integration mechanism to store spatiotemporal features through the semantic injection module, and generating query embeddings through the query enhancement mechanism; The semantic injection module stores spatiotemporal features through a long-term memory integration mechanism and generates query embedding through a query enhancement mechanism; the semantic injection module stores spatiotemporal features {f1, f2, ..., f T } is stored in the visual memory and generated into query embeddings {q1,q2,...,q M }; The query transformer enhances the query through the self-attention mechanism, and the query transformer integrates temporal features through the cross-attention mechanism.