Video question and answer method and system based on multi-level alignment

By adopting multi-level alignment method in video Q&A technology, multi-level visual features and text descriptions are generated, and modal alignment is established at object level, frame level and video level, the shortcomings of multi-modal information processing in the existing technology are solved, and high-performance video Q&A effect is achieved.

CN120104831APending Publication Date: 2025-06-06INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311647986.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing video Q&A technology has gradient vanishing problems when processing multimodal information, difficulty in capturing long-distance dependencies, high computing resources consumption, limited generalization capabilities, and neglecting the fine-grained interaction between local prominent information in the video and important text descriptions.

Method used

Using a video Q&A method based on multi-level alignment, by generating multi-level visual features and text descriptions, the alignment between visual modes and text modes is established at the object level, frame level and video level, and the language model is trained for video Q&A.

Benefits of technology

Effective alignment between multimodal information at the object level, frame level and video level is achieved, the performance of video Q&A is improved, and advanced performance is achieved better than the prior art with a smaller pre-trained data set and fewer trainable parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104831A_ABST
    Figure CN120104831A_ABST
Patent Text Reader

Abstract

The invention relates to a video question answering method and system based on multi-level alignment. The method comprises the following steps: generating multi-level visual features including global frame visual features and local frame visual features; generating multi-level text description according to the multi-level visual features, wherein the multi-level text description comprises global object description and local object description; and according to the generated multi-level visual features and the multi-level text description, establishing alignment between a visual mode and a text mode at an object level, a frame level and a video level, training a language model, and performing video question and answer by using the trained language model. According to the method, alignment between visual and text modes is established among object-level, frame-level and video-level multi-mode information, advanced performance can be obtained even if a small pre-training data set and few trainable parameters are adopted, and the method has wide practical value and application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer software and artificial intelligence technology, and specifically relates to a video question-answering method and system based on multi-level alignment. Background Art

[0002] Video question answering requires the intelligent agent to accurately understand the semantic information of the video and text, and to use questions as clues and guides to explore the key information in the video information, and finally get the answer through reasoning. There is a complex interactive relationship between video and text, and the model needs to be able to fully understand their interaction. At the same time, video question answering also needs to deal with cross-cutting issues in multiple fields such as natural language understanding, visual feature extraction, and reasoning. From the technical solution level, video question answering can be divided into attention-based methods, memory network-based methods, graph neural network-based methods, and Transformer-based methods.

[0003] 1. Methods based on attention mechanism: The basic idea of ​​the methods based on attention mechanism is to first model the video and question separately through the temporal modeling method (RNN, LSTM, etc.), and then capture the fusion information between the two modalities through the attention mechanism and output the answer. Attention is a human-inspired mechanism that locates the important parts of the input and selectively focuses on useful information. Self-attention has the ability to model long-range dependencies and can be used for intra-modal modeling, such as temporal information in videos and global dependency information in questions. Co-attention can focus on relevant and critical multimodal information, such as question-guided video representation and video-guided question representation. In video question answering, temporal attention and spatial attention are widely used to focus on specific parts of the video in spatial and temporal dimensions.

[0004] 2. Memory Network-based Methods: The basic idea of ​​Memory Network-based methods is to dynamically maintain memory modules. The question input will interact with the knowledge in the memory module in a weighted manner, and finally the memory content will be combined according to the weighted result to form the answer output. Memory networks can cache sequential inputs in memory modules and explicitly use earlier information. Among them, the problem of long video understanding involves not only the understanding of visual content, but also the long-distance dependency information they convey, so memory networks are particularly concerned in the problem of long video understanding.

[0005] 3. Graph neural network-based methods: Multimodal data is heterogeneous relational data. Graph neural networks are an effective means of modeling multimodal problems and an effective representation of relational data. Therefore, graph neural networks can be used to model video question-answering problems. The basic idea of ​​the graph neural network-based method is to first convert text and video into heterogeneous global graph neural network nodes, and then use the answer decoder to obtain the answer from the network.

[0006] 4. Transformer-based methods: Transformer has good temporal modeling capabilities and can focus on key information in text and video. Transformer-based methods have achieved success in both natural language processing and computer vision. From an architectural perspective, Transformer-based methods can be divided into encoder-only architecture and encoder-decoder architecture. In the field of video question answering, both text and video can be converted into one-dimensional sequence information that can be processed by Transformer. Therefore, the recent mainstream work is Transformer-based methods.

[0007] The above existing solutions have certain deficiencies in different aspects, as follows:

[0008] 1. Methods based on attention mechanism: Methods based on attention mechanism can associate multimodal information to a certain extent and have a memory effect. However, such methods rely on models such as RNN for time series modeling, have the problem of gradient vanishing, cannot capture long-distance dependencies, and are difficult to parallelize in engineering implementation.

[0009] 2. Memory network-based methods: Memory network-based methods can effectively model long-term problems and have good results in long video question answering. However, such methods can only provide combined answers based on the content of existing memory modules and are difficult to apply to open-ended tasks.

[0010] 3. Graph neural network-based methods: The difficulty of graph neural network-based methods lies in how to cleverly design a graph model for video representation. Such methods have good information communication capabilities between different modalities. However, such methods consume large computing resources, have limited generalization capabilities, and are weak in long-distance information transmission.

[0011] 4. Transformer-based methods: Transformer-based methods can effectively model long time series problems and can be calculated in parallel. Transformer-based methods usually globally align videos and texts to learn semantic relevance, ignoring the fine-grained interaction between local salient information in the video and important text descriptions. In addition, Transformer-based methods have high requirements on the quality of the dataset, while most of the subtitles in the dataset are very short and lack detailed text descriptions of important video content. Summary of the invention

[0012] In view of the above problems, the present invention provides a video question answering method and system based on multi-level alignment.

[0013] The technical solution adopted by the present invention is as follows:

[0014] A video question answering method based on multi-level alignment, comprising the following steps:

[0015] Generate multi-level visual features, including global frame visual features and local frame visual features;

[0016] Generate multi-level text descriptions based on multi-level visual features, including global object descriptions and local object descriptions;

[0017] Based on the generated multi-level visual features and multi-level text descriptions, alignment between the visual modality and the text modality is established at the object level, frame level, and video level, and a language model is trained. The trained language model is used to perform video question answering.

[0018] Furthermore, the generating of multi-level visual features includes:

[0019] For a given sample frame, an image encoder is used to extract global frame visual features;

[0020] A mask generator is used to generate a mask for the object in the frame, and an image encoder is used to obtain the object features to obtain the local frame visual features.

[0021] Furthermore, the generation process of the global frame visual features includes: representing the video as a series of frames obtained by uniform and sparse sampling, each frame is encoded separately using the image encoder CLIP to generate the global frame visual features; the generation process of the local frame visual features includes: first using the mask generator CutLER to generate a corresponding image mask for each frame, thereby obtaining a mask set of all frames; then applying cropping and masking operations to the images, and feeding them into the image encoder CLIP to obtain the local frame visual features:

[0022] Furthermore, generating a multi-level text description according to the multi-level visual features includes:

[0023] For global frame visual features, an image language model is used to generate global frame captions;

[0024] Extract noun phrases from global frame captions as frame-specific dynamic vocabulary, and combine the dynamic vocabulary with predefined static vocabulary to form the vocabulary input of the object filter;

[0025] Object filters are used to extract global object descriptions and local object descriptions from global frame visual features and local frame visual features.

[0026] Furthermore, the extracting of global object description and local object description from global frame visual features and local frame visual features using object filters includes: determining the object description of each visual feature by evaluating the cosine similarity between the global frame visual features or the local frame visual features and the text features.

[0027] Furthermore, establishing alignment between the visual modality and the textual modality at the object level, the frame level, and the video level includes:

[0028] In object-level alignment, a visual feature and its corresponding object description are regarded as alignment units, and the alignment units are iteratively combined to align global frame visual features with local frame visual features and global object descriptions with local object descriptions. At the same time, the object descriptions are sorted in descending order according to the cosine similarity between text features and visual features.

[0029] In frame-level alignment, all object-level information grouped by frames and global frame captions are considered as alignment units;

[0030] In video-level alignment, a scalable ordinal word cue method is adopted to integrate temporal information to construct video-level alignment.

[0031] Furthermore, the global frame visual features and local frame visual features, global object descriptions and local object descriptions, and global frame captions are combined and input into the language model with Adapter. The global frame visual features and local frame visual features are passed through a linear layer to obtain global feature prompts and local feature prompts, and a classifier head based on masked language modeling (MLM) is used to predict the answer from the vocabulary set constructed from the answers appearing in the training set.

[0032] A video question answering system based on multi-level alignment, comprising:

[0033] A multi-level visual feature generation module, used to generate multi-level visual features, including global frame visual features and local frame visual features;

[0034] A multi-level text description generation module is used to generate multi-level text descriptions based on multi-level visual features, including global object descriptions and local object descriptions;

[0035] The multi-level alignment and training module is used to establish alignment between visual modalities and textual modalities at the object level, frame level, and video level based on the generated multi-level visual features and multi-level text descriptions, train the language model, and use the trained language model for video question answering.

[0036] The beneficial effects of the present invention are as follows:

[0037] The present invention provides a video question-answering solution based on multi-level alignment, which establishes alignment between visual and textual modalities among multimodal information at the object level, frame level, and video level. Even with a smaller pre-training dataset and fewer trainable parameters, it can achieve advanced performance that is superior to the existing technology. It fills the gap in the current research in this direction to a certain extent and has a wide range of practical value and application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flowchart of the video question answering method based on multi-level alignment.

[0039] Figure 2 This is a schematic diagram of the specific implementation process of the video question answering method based on multi-level alignment. DETAILED DESCRIPTION

[0040] The present invention is further described in detail below through specific embodiments and drawings.

[0041] The goal of the present invention is to establish alignment between visual and textual modalities between multimodal information at the object level, frame level, and video level. The specific process of the present invention can be divided into three stages: multi-level visual feature generation, multi-level text description generation, and multi-level alignment and training. The process of the present invention is as follows: Figure 1 shown.

[0042] Stage one is multi-level visual feature generation. For a given sampling frame, first, the present invention uses an image encoder to extract global frame visual features. Then, the present invention uses a mask generator to generate a mask for the object in the frame, and uses an image encoder to obtain object features, which are local frame visual features. The global frame visual features capture the overall spatial information of the frame, while the local frame visual features accurately capture details such as the location and attributes of a specific object. In this process, since the output bounding box of the target detector is usually not specific enough and it is difficult to accurately identify objects in an open vocabulary environment, the present invention uses a mask generator instead of a target detector.

[0043] Stage two is multi-level text description generation. For global frame visual features, the present invention first uses the image language model BLIP to generate frame subtitles. Then, the part-of-speech tagging tool provided by spaCy is used to extract noun phrases as frame-specific dynamic vocabulary. These dynamic vocabulary are then combined with predefined static vocabulary to form the vocabulary input of the object filter. Then, the object filter is used to extract global and local object descriptions from the global and local frame visual features. In this process, due to the different information contained in the global and local frame visual features, the extracted global and local object descriptions may also be different. By using this direct and efficient method, the present invention achieves frame-level and object-level visual-language alignment.

[0044] Phase 3 is multi-level alignment and training. To perform the video question-answering task, the present invention combines global and local frame visual features, global and local object descriptions, and global frame captions into a language model with an adapter. In the information fusion process, in order to construct video-level alignment, the present invention uses an extensible ordinal word prompt method to integrate time information.

[0045] The following is a detailed description of the specific process of the above three stages (such as Figure 2 shown).

[0046] Stage 1: Multi-level visual feature generation

[0047] In order to capture multi-level visual information, it is crucial to understand the relationship between objects in the video frame (represented as global frame visual features) and the local semantic information of objects (represented as local frame visual features). Stage 1 describes the generation process of multi-level visual features.

[0048] 1. Global frame visual feature generation:

[0049] Video is represented as a series of frames obtained by uniform and sparse sampling Where T is the total number of sample frames. Each frame f i Both use image encoder φ CLIP Separately encode to generate global frame visual features v G :

[0050]

[0051]

[0052] Among them, D v is the dimension of visual features.

[0053] Because v G Take the original frame image as input, v G Contains the overall spatial information of the frame. Therefore, v GRepresents frame-level visual information.

[0054] 2. Local frame visual feature generation:

[0055] The goal of the present invention is to establish a multi-level visual-language alignment. Considering the effectiveness of contrastive learning-based image-language models in establishing a joint representation of images and text, the present invention adopts a contrastive learning model CLIP as an image encoder. However, since CLIP mainly focuses on capturing the global information of an image, CLIP is not very suitable for directly capturing local details. To solve this problem, the present invention uses an existing unsupervised mask generator CutLER to guide the image encoder CLIP to generate local frame visual features.

[0056] In order to obtain the local frame visual features, the present invention first uses the mask generator CutLER (with φ CutLER To represent) for each frame f i Generate the corresponding image mask m i , thus obtaining the mask set m of all frames:

[0057] m i =φ CutLER (f i )

[0058]

[0059] The present invention then applies cropping and masking operations to the images and feeds them to the image encoder φ CLIP To obtain the local frame visual feature v L :

[0060]

[0061]

[0062] in, represents the cropping and masking operation, and ⊙ is the Hadamard product operation. In this process, the present invention assumes that the number of masks of the i-th frame is N i .

[0063] Because v L The areas irrelevant to the object are removed, v L Only the target object itself is concerned. Therefore, the present invention considers that v L Represents object-level visual information.

[0064] Stage 2: Multi-level text description generation

[0065] In order to generate multi-level text information, it is necessary to simultaneously generate text describing the relationship between objects and text describing the detailed information of the objects. Phase 2 describes the generation process of multi-level text information.

[0066] 1. Global frame subtitle generation:

[0067] In order to establish text associations between objects in a frame and obtain an overall text description of the frame, the present invention uses an image language model BLIP to generate global frame subtitles. The global frame subtitle c contains the text information at the frame level.

[0068] 2. Global and local frame object description generation:

[0069] Using only global frame captions may not capture the detailed information of objects. Therefore, it is not enough to generate text descriptions only at the frame level. To solve this problem, the present invention further utilizes global and local frame visual features to generate global and local object descriptions.

[0070] First, the present invention constructs a predefined static vocabulary V S , static vocabulary V S Contains a set of candidate object names and attributes. Specifically, the present invention constructs a static vocabulary based on the category names in the OpenImage v7 dataset, which contains 21,000 noun phrases. For these noun phrases, the present invention calculates the cosine similarity between each pair, and if the cosine similarity between the two words exceeds 0.95, the word with a lower frequency is deleted. In addition, the present invention also manually removes meaningless phrases such as "video", "picture", "photo", etc. from the static vocabulary.

[0071] In order to solve the problem that the predefined static vocabulary cannot cover all objects and attributes, the present invention uses spaCy to extract noun phrases from the global frame subtitles c to construct a dynamic vocabulary V D The final vocabulary V is a static vocabulary V S With dynamic vocabulary V D The union of is defined as follows:

[0072] V=V S ∪V D

[0073] Then, the present invention uses the text encoder φ based on contrastive learning in the CLIP model CLIP-text Calculate the text feature r of the final vocabulary V:

[0074]

[0075] Where L is the size of the final vocabulary V.

[0076] Finally, the present invention evaluates the global frame visual feature v G Or local frame visual features v L The cosine similarity between the text feature r and the object description of each visual feature is determined. For the global frame visual feature v G , the present invention generates a total of M for each visual feature G A global object description, denoted as t G Similarly, for the local frame visual feature v L , the present invention generates a total of M for each visual feature L A local object description, denoted as t L .

[0077] Since the global frame visual feature v G and local frame visual features v L The content of the global object description is different. G and local object description t L There are also differences. The present invention believes that t G and t L Represents object-level text information.

[0078] Stage 3: Multi-level alignment and training

[0079] 1. Generation of multi-level alignment prompts:

[0080] To achieve visual-language alignment, the present invention first integrates the obtained object-level and frame-level visual features with their corresponding textual descriptions to align and fuse visual and textual information at the object level and frame level. Then, the present invention integrates scalable temporal tokens to establish video-level visual-language alignment.

[0081] First, if Figure 1 As shown, in order to achieve multi-level alignment, the present invention designs a multi-level alignment prompt as follows:

[0082] “Question: <question>?Answer:[MASK].Caption:.Global: <globalobjects>.Local:<Local Objects> .Subtitles:<additional description> .”

[0083] In object-level alignment, the present invention regards "a visual feature and its corresponding object description" as an alignment unit. The present invention iteratively combines object-level units to align global and local frame visual features (v G or v L ) and their respective object descriptions (t G or L ). At the same time, in order to utilize the neighborhood inductive bias of language, the present invention sorts the object descriptions in descending order according to the cosine similarity between text features and visual features. Specifically, for each global frame visual feature "[GLOBAL]" and local frame visual feature "[LOCAL]", the present invention constructs the following prompts:

[0084] "[GLOBAL]<Description 1> ,...”, where Description 1 represents the global object description;

[0085] "[LOCAL]<Description 1> ,...”, where Description 1 represents the local object description.

[0086] In frame-level alignment, the present invention regards all object-level information grouped by frame and the global frame caption "First,.Second,...." as alignment units, where Caption represents the text description in the global frame caption.

[0087] In video-level alignment, the present invention uses frame-level information and additional descriptions (additional video-level text descriptions that may exist in the dataset) as alignment units. At the video level, time information is crucial for video understanding. Therefore, the present invention introduces extensible ordinal numbers in the prompts, such as "First," "Second," etc.

[0088] 2. Language model with Adapter:

[0089] In order to complete the video question answering task, from the perspective of model structure, this paper inputs multi-level alignment cues into DeBERTa and uses the same Adapter layer as the FrozenBiLM model. Specifically, each visual feature is transformed by a linear layer and integrated into the language model as a separate token. Global and local visual features (v G and v L ) after the linear layer (P G and P L ) to obtain global and local feature hints (u G and u L ):

[0090]

[0091]

[0092] For the autoencoder language model such as DeBERTa, the present invention uses a classifier head based on MLM (masked language modeling), denoted as m θ , to predict the answer from a vocabulary set A constructed from the answers that appear in the training set.

[0093] 3. Cross-modal training:

[0094] During training, the present invention only updates the parameters of the vision-to-text projection module P and the Adapter module. To maintain consistency with the language model training method, the present invention uses a masked language modeling (MLM) objective, in which some tokens are randomly masked and the model must predict these masked tokens based on other tokens.

[0095] The key points of the present invention are:

[0096] 1. Multi-level alignment framework: This method consists of three stages: multi-level visual feature generation, multi-level text description generation, and multi-level alignment and training.

[0097] 2. The local frame visual feature generation process in stage one.

[0098] 3. The global frame subtitle generation process in stage 2.

[0099] 4. The global and local frame object description generation process in stage 2.

[0100] 5. The multi-level alignment prompt generation process in stage three.

[0101] Positive effects:

[0102] The hardware configuration of the experiment of the present invention is shown in Table 1.

[0103] Table 1

[0104] operating system Ubuntu 22.04LTS GPU 8 NVIDIA GeForce RTX 3090 GPUs Python 3.8 pytorch 1.12.1

[0105] The data sets used in the experiments of this invention are iVQA, MSRVTT-QA, MSVD-QA and TGIF-FrameQA. The specific information is as follows:

[0106] 1. iVQA: Focuses on objects, scenes, and people in teaching videos. Includes 10,000 video clips and 10,000 question-answer pairs, divided into 6,000 / 2,000 / 2,000 for training / validation / testing.

[0107] 2. MSVD-QA: includes 2,000 video clips and 51,000 question-answer pairs, divided into 32,000 / 6,000 / 13,000 for training / validation / testing. The question-answer pairs in MSVD-QA are automatically generated from the video descriptions.

[0108] 3.MSRVTT-QA: Includes 10,000 video clips and 243,000 question-answer pairs, divided into 158,000 / 12,000 / 73,000 for training / validation / testing. The question-answer pairs in MSRVTT-QA are automatically generated from the video descriptions.

[0109] 4. TGIF-FrameQA: An open-ended video question answering benchmark based on GIF videos. In the TGIF-FrameQA dataset, most videos are short in duration, usually no longer than 5 seconds. These videos are mainly divided into four categories: object, quantity, color, and position. It includes 46,000 GIF images and 53,000 question-answer pairs, divided into 39,000 / 13,000 for training / testing.

[0110] Experimental design: Use iVQA, MSRVTT-QA, MSVD-QA and TGIF-FrameQA datasets to conduct video question answering experiments. Compare the proposed method with various existing methods, and use the TOP-1 accuracy index to evaluate the performance of each method in zero-shot and fully supervised scenarios.

[0111] The performance test results in the zero-shot scenario are shown in Table 2. The performance test results in the full-supervision scenario are shown in Table 3.

[0112] Table 2. Statistics of zero-sample scenario experimental results

[0113]

[0114] Table 3. Statistics of experimental results of fully supervised scenarios

[0115]

[0116] Table 2 shows the comparison results of the present invention with the most advanced zero-shot methods. The present invention achieves state-of-the-art performance on the MSVD-QA, MSRVTT-QA, and TGIF-FrameQA datasets, while achieving competitive results on the iVQA dataset. Compared with all methods, the present invention uses the smallest pre-training dataset. In particular, the performance of the present invention exceeds that of the FrozenBiLM model, which also uses a language model with an Adapter. Compared with the fully supervised scenario, the size of the pre-training dataset is more critical in the zero-shot scenario. However, even though the present invention uses the smallest pre-training dataset compared to other methods, the present invention still achieves state-of-the-art performance. This can be attributed to the fact that the present invention effectively transfers the zero-shot capabilities of the mask generator, image encoder, and subtitle generator.

[0117] Table 3 shows the comparison results of the present invention with the most advanced fully supervised methods. The present invention achieves state-of-the-art performance on the iVQA, MSRVTT-QA, and MSVD-QA datasets, and achieves competitive results on the TGIF-FrameQA dataset. Similarly, compared with other comparison models, the present invention uses the smallest pre-training dataset and fewer trainable parameters. Similar to the conclusions in the zero-shot scenario, the present invention shows a relative advantage over FrozenBiLM in the fully supervised scenario.

[0118] Therefore, the present invention provides an effective video question answering method based on multi-level alignment, which has excellent performance indicators and has a wide range of practical value and application scenarios.

[0119] Another embodiment of the present invention provides a video question answering system based on multi-level alignment, comprising:

[0120] A multi-level visual feature generation module, used to generate multi-level visual features, including global frame visual features and local frame visual features;

[0121] A multi-level text description generation module is used to generate multi-level text descriptions based on multi-level visual features, including global object descriptions and local object descriptions;

[0122] The multi-level alignment and training module is used to establish alignment between visual modalities and textual modalities at the object level, frame level, and video level based on the generated multi-level visual features and multi-level text descriptions, train the language model, and use the trained language model for video question answering.

[0123] The specific implementation process of each module refers to the above description of the method of the present invention.

[0124] Based on the same inventive concept, another embodiment of the present invention provides a computer device (computer, server, smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention.

[0125] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, CD), which stores a computer program. When the computer program is executed by a computer, it implements the various steps of the method of the present invention.

[0126] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. It can be understood by those skilled in the art that various replacements, changes and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the contents disclosed in the embodiments of this specification, and the scope of protection of the present invention shall be subject to the scope defined in the claims.< / globalobjects> < / question>

Claims

1. A video question answering method based on multi-level alignment, It is characterized in that The following steps are involved: Generate multi-level visual features, including global frame visual features and local frame visual features; Generate multi-level text descriptions based on multi-level visual features, including global object descriptions and local object descriptions; Based on the generated multi-level visual features and multi-level text descriptions, alignment between the visual modality and the text modality is established at the object level, frame level, and video level, and a language model is trained. The trained language model is used to perform video question answering.

2. The method according to claim 1, It is characterized in that The generating of multi-level visual features comprises: For a given sample frame, an image encoder is used to extract global frame visual features; A mask generator is used to generate a mask for the object in the frame, and an image encoder is used to obtain the object features to obtain the local frame visual features.

3. The method according to claim 2, It is characterized in that The generation process of the global frame visual features includes: the video is represented as a series of frames obtained by uniform and sparse sampling, and each frame is individually encoded using an image encoder CLIP to generate global frame visual features; the generation process of the local frame visual features includes: first, using a mask generator CutLER to generate a corresponding image mask for each frame, thereby obtaining a mask set for all frames; then, cropping and masking operations are applied to the images, and they are fed into the image encoder CLIP to obtain local frame visual features.

4. The method according to claim 1, It is characterized in that The step of generating a multi-level text description according to the multi-level visual features includes: For global frame visual features, an image language model is used to generate global frame captions; Extract noun phrases from global frame captions as frame-specific dynamic vocabulary, and combine the dynamic vocabulary with predefined static vocabulary to form the vocabulary input of the object filter; Object filters are used to extract global object descriptions and local object descriptions from global frame visual features and local frame visual features.

5. The method according to claim 4, It is characterized in that The method of extracting global object description and local object description from global frame visual features and local frame visual features using object filters includes: determining the object description of each visual feature by evaluating the cosine similarity between the global frame visual features or the local frame visual features and the text features.

6. The method according to claim 1, It is characterized in that The alignment between the visual modality and the textual modality at the object level, the frame level and the video level is established, including: In object-level alignment, a visual feature and its corresponding object description are regarded as alignment units, and the alignment units are iteratively combined to align global frame visual features with local frame visual features and global object descriptions with local object descriptions. At the same time, the object descriptions are sorted in descending order according to the cosine similarity between text features and visual features. In frame-level alignment, all object-level information grouped by frames and global frame captions are considered as alignment units; In video-level alignment, a scalable ordinal word cue method is adopted to integrate temporal information to construct video-level alignment.

7. The method according to claim 6, It is characterized in that The global frame visual features and local frame visual features, global object descriptions and local object descriptions, and global frame captions are combined and input into the language model with Adapter. The global frame visual features and local frame visual features are passed through a linear layer to obtain global feature prompts and local feature prompts, and a classifier head based on masked language modeling (MLM) is used to predict the answer from the vocabulary set constructed from the answers that appear in the training set.

8. A video question answering system based on multi-level alignment, It is characterized in that include: A multi-level visual feature generation module, used to generate multi-level visual features, including global frame visual features and local frame visual features; A multi-level text description generation module is used to generate multi-level text descriptions based on multi-level visual features, including global object descriptions and local object descriptions; The multi-level alignment and training module is used to establish alignment between visual modalities and textual modalities at the object level, frame level, and video level based on the generated multi-level visual features and multi-level text descriptions, train the language model, and use the trained language model for video question answering.

9. A computer device, It is characterized in that The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Chest X-ray film multi-mode pre-training method and system based on graph perception learning

    CN120911533A