Super-long audio and video understanding method, system and equipment based on visual language model
Through multi-grained intention recognition and dynamic keyframe extraction technology based on visual language model, combined with multi-modal information fusion and space-time prompt mechanism, the real-time, multi-modal fusion and space-time modeling problems in ultra-long audio and video understanding are solved, and efficient and accurate long video understanding is achieved.
Patent Information
- Application Number
- CN202510444847.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-18
AI Technical Summary
The existing ultra-long audio and video understanding technology has problems such as insufficient real-time, insufficient multimodal fusion, limitations of space-time modeling, weak supervision and sparse labeling problems, small sample learning and long tail distribution, multi-scale event detection and lack of causal reasoning, resulting in high demand for computing resources, complex system architecture, limited timing modeling defects and generalization capabilities.
Using a method based on visual language model, through multi-grained intention recognition, dynamic keyframe extraction, multi-modal information fusion and spatiotemporal prompt mechanisms, combined with automatic speech recognition and hierarchical generation technology, lightweight heterogeneous architecture design is realized, supporting fine-grained timing alignment and semantic fusion of multi-modal information, and improving timing information dependence and generalization capabilities.
It reduces the computing resource requirements, simplifies the system architecture, improves the efficiency and accuracy of ultra-long audio and video understanding, meets the real-time processing needs of edge devices, and enhances the multimodal information fusion and timing understanding of long videos.
Smart Images

Figure CN120336483A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and relates to a method, system and device for understanding ultra-long audio-visual content, in particular to a method, system and device for understanding ultra-long audio-visual content based on a vision-language model. Background Art
[0002] Existing ultra-long audio-visual content understanding technologies mainly focus on aspects such as real-time performance, multi-modal fusion, and spatio-temporal modeling. However, they usually have the following problems:
[0003] 1. Insufficient real-time performance: Limitations in dynamic information and key frame sampling.
[0004] Existing long video understanding systems (such as Video Swin Transformer, TimeSformer, etc.) generally adopt an end-to-end model architecture to achieve long video modeling by uniformly compressing visual tokens in the video frame sequence. However, this design has significant drawbacks: (1) Fixed frame rate sampling leads to loss of dynamic information. Existing methods usually uniformly sample key frames at a fixed frame rate (such as 1 frame per second), making it difficult to adapt to the dynamic changes of video content. In fast-moving scenarios (such as shooting actions in a football game, vehicle lane changes in traffic monitoring), key actions may occur within the sampling interval, resulting in the model being unable to capture key temporal information and significantly reducing the recognition accuracy. (2) Information loss in token compression. The end-to-end model compresses high-dimensional visual tokens through global attention or pooling operations (for example, compressing hundreds of frames into dozens of tokens), but this process will lose fine-grained spatio-temporal details. For example, in medical endoscope video analysis, small changes in the lesion area may be ignored due to compression. (3) Conflict between computational efficiency and real-time performance. Existing solutions sacrifice real-time performance requirements to reduce computational costs (such as reducing the frame rate or resolution). For example, in an autonomous driving scenario, low frame rate processing may cause delays and cannot meet the safety requirements of millisecond-level responses. Processing ultra-long audio-visual content faces double limitations of video memory and computing power: a) Traditional 3D convolution requires caching more than 10^6 frame features to process 1 hour of high-definition video (1080p@30fps), exceeding the GPU video memory capacity. b) The time complexity of the self-attention mechanism grows quadratically with the video length (O(n 2 ^2)), and the delay in processing a 10-minute video reaches the minute level. Existing lightweight methods (such as local attention or low-rank decomposition) can reduce the computational amount, but they will break long-range dependencies, increasing the error rate of key argument extraction by 25% in scenarios such as courtroom debate video analysis.
[0005] 2. Deficiency in all modalities: Insufficient multi-modal fusion.
[0006] Most current video understanding models regard videos as silent visual sequences, ignoring the synergistic effects of multimodal information such as audio, text, and sensors, resulting in limited semantic understanding. The limitations of existing fusion schemes include: (1) The temporal alignment problem in early fusion. Visual and audio features (such as Mel spectrograms) are directly concatenated at the input end, but the temporal resolution differences between different modalities (such as millisecond-level sampling of audio vs. frame-level sampling of video) lead to fusion noise. For example, in music videos, the subtle temporal misalignment between instrument movements and sounds reduces classification accuracy. (2) The lack of interaction in late fusion. The visual and audio classification results are simply weighted at the end of the model (such as Two-Stream Networks), lacking fine-grained interaction between modalities. For example, in movie scene classification, the correlation between explosion sounds and fire scenes cannot be fully modeled. (3) The limitations of cross-modal pre-training. Existing multimodal models (such as CLIP, VideoCLIP) rely on large-scale aligned data, but lack domain adaptation capabilities in vertical domains (such as industrial inspection videos) and do not fully utilize sensor data (such as infrared, depth information).
[0007] 3. Limitations in spatio-temporal modeling: Long sequence dependencies and spatial fine-grainedness.
[0008] (1) Insufficient modeling of time dependencies, which specifically includes: a) The ineffectiveness of the attention mechanism for long videos. The self-attention computational complexity of Transformer-based models (such as ViT) grows quadratically with the number of frames, forcing the system to reduce the number of frames or process them in segments, resulting in the breakage of temporal correlations for long-range actions (such as continuous gymnastic movements). b) The problem of time forgetting. RNN / LSTM-based models are limited by memory capacity and are difficult to model dependencies exceeding hundreds of frames. For example, in a marathon race video, the trend of an athlete's physical fitness changes may be forgotten. c) Lack of temporal localization ability. Existing methods do not introduce a time cue mechanism and cannot perform targeted enhancement for specific time intervals (such as "the last 5 minutes of the game"), resulting in the omission of key events.
[0009] (2) Lack of spatial fine-grainedness in spatial modeling, which specifically includes: a) Imbalance in the trade-off between local and global. CNN-based models (such as C3D) rely on a fixed receptive field and are difficult to capture both local details (such as gestures) and global context (such as scene layout) simultaneously; Vision Transformer (such as Swin-T) reduces computational complexity through window attention, but cross-window interaction is limited, resulting in incomplete modeling of spatial relationships. b) Poor adaptability to dynamic scenes. In complex lighting, occlusion, or small object scenes (such as night pedestrian detection under a surveillance camera), the robustness of existing spatial feature extractors significantly decreases.
[0010] 4. Challenges in weak supervision and sparse annotation: Annotation efficiency and semantic noise.
[0011] Existing long - video supervised learning heavily relies on manual annotation. However, the annotation cost and data sparsity limit the generalization ability of the model. The specific defects include: (1) The contradiction between annotation cost and granularity. Mainstream datasets (such as ActivityNet, Charades) need to annotate the action categories and time boundaries frame - by - frame or segment - by - segment for videos. Annotating a 1 - hour video consumes an average of 6 person - hours of annotation resources. In industrial scenarios (such as surgical video analysis), the expert annotation cost is even higher and it is difficult to scale up, resulting in the model being unable to cover long - tail categories (such as rare complications). (2) Semantic dilution in weakly - supervised learning. Weakly - supervised methods based on video - text alignment (such as VideoCLIP) face semantic noise interference in long videos. For example, in an 8 - hour surveillance video, only 0.1% of the segments contain theft events. Directly applying contrastive learning will cause the model to overly focus on background features (such as stationary shelves) rather than dynamic actions (such as item hiding). (3) Data augmentation that violates physical laws. Traditional augmentation strategies (such as random cropping, time flipping) destroy the spatio - temporal coherence of actions. For example, in instrument operation videos, random cropping may generate illegal sequences like "disassembling first and then assembling", leading the model to learn incorrect causal logic.
[0012] 5. Few - shot learning and long - tail distribution: Data scarcity and generalization bottleneck.
[0013] Industrial - scenario video data shows an extremely long - tail distribution (e.g., normal events account for 99.9% in security). Existing methods have significant defects in data utilization and generalization ability: (1) Physical unreasonableness in few - shot learning. Traditional data augmentation (such as color jittering, spatial deformation) cannot generate action sequences that conform to kinematic laws. For example, in the fall - detection task, the "inverted human body" generated by rotation violates gravity constraints, resulting in a decrease in the model's discriminative ability for real fall scenarios (such as lateral imbalance). (2) Bias amplification in long - tail classification: The cross - entropy loss function is dominated by head classes in gradient updates, forcing the model to bias towards high - frequency classes. For example, in factory quality - inspection videos, the proportion of defective samples is less than 0.01%. Direct training will cause the model to predict all videos as "normal". (3) Semantic gap in cross - domain transfer. Pre - trained models (such as I3D pre - trained on Kinetics) have feature mismatches due to domain differences (lighting, perspective) in vertical domains (such as agricultural pest monitoring), and fine - tuning requires a large amount of annotated data.
[0014] 6. Multi - scale event detection and causal reasoning: Hierarchical fragmentation and logical absence.
[0015] Long video events have multi-level spatio-temporal characteristics. The existing single-scale modeling and causal reasoning are insufficient, resulting in the loss of key semantics: (1) The failure of multi-scale collaborative modeling. Mainstream frameworks (such as SlowFast) use dual paths to process spatio-temporal features respectively, but do not establish cross-granularity semantic associations. For example, in a conference video, the model may independently detect the action of "raising a hand" (micro) and the event of "agenda voting stage" (macro), but cannot infer the subordinate relationship between the two (such as raising a hand triggering a vote). (2) The lack of causal chain modeling. Existing methods (such as TCN, Transformer) only model the statistical correlation of temporally adjacent actions, ignoring causal logic constraints. For example, in a first aid training video, the model may detect "chest compressions" and "turning on the defibrillator" as independent events, rather than the necessary sequence of "turning on the machine first and then operating", resulting in incorrect assessment scores. (3) The interpretability limitation of black-box models. Explanation methods based on attention weights (such as Grad-CAM) in long videos lead to blurred heatmaps due to the explosion of token numbers, and cannot locate the root cause of the causal chain. For example, in the determination of traffic accident liability, the decision-making basis of "changing lanes and crossing the line → sudden braking → rear-end collision" cannot be traced.
[0016] Therefore, in view of the defects existing in the above-mentioned prior art, it is necessary to develop a new method, system and device for understanding ultra-long audio and video. Summary of the Invention
[0017] In order to overcome the defects of the prior art, the present invention proposes a method, system and device for understanding ultra-long audio and video based on a vision-language model, which can reduce the computational resource requirements, simplify the system architecture, enhance the dependence on temporal information, and improve the generalization ability, thereby effectively solving the technical problems of understanding ultra-long audio and video.
[0018] In order to achieve the above object, the present invention provides the following technical solutions:
[0019] A method for understanding ultra-long audio and video based on a vision-language model, characterized by comprising the following steps:
[0020] 1) Multi-granularity intention recognition: Using a fine-tuned large language model to perform multi-granularity intention recognition on the user's question to determine the inquiry mode of the user's question, and the inquiry mode includes a single-image inquiry mode, an audio content inquiry mode, and a video content inquiry mode;
[0021] 2) Content recognition: Based on the inquiry mode and the user's question, identify the pictures, audio, and video input by the user to obtain the recognized content;
[0022] 3) Multi-modal information fusion: Using a large language model to perform multi-modal information fusion on the recognized content based on a spatio-temporal prompt mechanism and a hierarchical generation mechanism;
[0023] 4) Answer generation: Input the fusion result of the user question and multimodal information into a vision-language model, and the vision-language model generates the corresponding answer to the user question.
[0024] Preferably, step 2) specifically includes:
[0025] 21) For the single-image query mode, use a vision-language model to recognize the user-input image based on the user question to obtain single-image recognition content;
[0026] 22) For the audio content query mode, use automatic speech recognition technology to recognize the user-input audio based on the user question to obtain audio recognition content;
[0027] 23) For the video content query mode, use dynamic key-frame extraction technology to extract key frames from the user-input video based on the user question to obtain key-frame pictures, and use a vision-language model to recognize the extracted key-frame pictures based on the user question to obtain video recognition content.
[0028] Preferably, the step of using dynamic key-frame extraction technology to extract key frames from the user-input video based on the user question in step 23) to obtain key-frame pictures specifically includes:
[0029] 231) Scene segmentation: Use the BaSSL model to segment the user-input video to obtain multiple segmented scenes, and obtain the segmented scene corresponding to the user question based on the user question;
[0030] 232) Key-frame extraction: Use FFmpeg to extract key frames from the segmented scene corresponding to the user question to obtain preliminarily extracted pictures;
[0031] 233) Semantic-aware redundancy removal: Calculate the similarity between the image feature vector of the preliminarily extracted picture and the text feature vector of the user question based on the CLIP algorithm, and discard the preliminarily extracted pictures with similarity lower than the threshold to obtain the finally extracted key-frame pictures.
[0032] Preferably, step 3) specifically includes:
[0033] 31) Set a time localization template, a spatial cue template, and a causal chain constraint template, and form a cue template based on the time localization template, the spatial cue template, and the causal chain constraint template. The time localization template is used to mark the key time interval, the spatial cue template is used to mark the key area, and the causal chain constraint template is used to annotate the event logical relationship;
[0034] 32) Set a context-aware template, which is used to dynamically adjust the prompt template according to the video duration and the user's question;
[0035] 33) Set a feedback correction template, which is used to optimize the prompt template based on the reinforcement learning method;
[0036] 34) Based on the optimized prompt template, use a large language model to perform multi-modal information fusion on the single-image recognition content, audio recognition content, and video recognition content.
[0037] Preferably, in step 34), an audio-video description fusion enhancement and complementary reasoning mechanism is adopted for multi-modal information fusion, that is, when the video recognition content is described vaguely, the audio recognition content is used to dominate the reasoning, and when the audio recognition content is missing, the video recognition content is activated for complementation.
[0038] Preferably, in step 34), a knowledge-enhanced consistency verification mechanism is used for multi-modal information fusion, that is, the knowledge bases of each segmented scene are accessed, and the knowledge of the knowledge bases is input into the large language model as context information during multi-modal information fusion, so that the large language model generates a multi-modal information fusion result based on the context information.
[0039] Preferably, the large language model is fine-tuned in the following way to obtain the fine-tuned large language model:
[0040] Set question-answer pairs including the user's question and its corresponding inquiry mode;
[0041] Fine-tune the large language model by using the question-answer pairs through the method of function call.
[0042] In addition, the present invention also provides a very long audio-visual video understanding system based on a vision-language model, which is characterized by including:
[0043] A multi-granularity intention recognition module, which is used to perform multi-granularity intention recognition on the user's question by using the fine-tuned large language model to determine the inquiry mode of the user's question, and the inquiry mode includes a single-image inquiry mode, an audio content inquiry mode, and a video content inquiry mode;
[0044] A content recognition module, which is used to recognize the pictures, audio, and video input by the user based on the inquiry mode and the user's question to obtain recognition content;
[0045] A multi-modal information fusion module, which is used to perform multi-modal information fusion on the recognition content by using the large language model based on the spatio-temporal prompt mechanism and the hierarchical generation mechanism;
[0046] An answer generation module, which is used to input the user question and the multi-modal information fusion result into a vision-language model, and the vision-language model generates the corresponding answer to the user question.
[0047] Moreover, the present invention also provides a very long audio-video understanding device based on a vision-language model, which is characterized by including:
[0048] One or more processors;
[0049] A memory for storing one or more programs;
[0050] When the one or more programs are executed by the one or more processors, the one or more processors implement the very long audio-video understanding method based on the vision-language model as described above.
[0051] Finally, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the program, when executed by a processor, implements the steps of the very long audio-video understanding method based on the vision-language model as described above.
[0052] Compared with the prior art, the very long audio-video understanding method, system and device based on the vision-language model of the present invention achieve a breakthrough improvement in the efficiency and accuracy of very long audio-video understanding through the collaborative design of multi-granularity intention driving, dynamic sampling and compression, and cross-modal fusion, and have one or more of the following beneficial technical effects:
[0053] 1. The present invention adopts a lightweight heterogeneous architecture design, constructs a visual-audio-text joint alignment mechanism, and adopts a hierarchical feature extraction strategy - the video branch extracts key frames, and the audio branch adopts ASR technology to synchronously generate cross-modal semantic anchors.
[0054] 2. The present invention adopts a dynamic multi-modal fusion mechanism, designs a multi-granularity intention recognition module, adaptively fuses audio, video and text descriptions through learnable call tools, designs a multi-modal information fusion module, realizes fine-grained temporal alignment and semantic fusion, maintains multi-modal consistency, and improves the recall rate of very long audio-video event retrieval.
[0055] 3. The present invention adopts a very long time series processing technology, innovatively introduces a spatio-temporal prompt template mechanism and a hierarchical prompt mechanism, and supports video understanding at the hour level.
[0056] 4. The present invention adopts a resource optimization strategy, extracts key frames based on methods such as CLIP similarity and scene segmentation, improves the compression rate, reduces the video memory, and meets the real-time processing requirements of edge devices. Description of the Drawings
[0057] Figure 1 is a flowchart of the very long audio-video understanding method based on the vision-language model of the present invention.
[0058] Figure 2 It is a schematic diagram of the composition of the ultra-long audio-visual understanding system based on the vision-language model of the present invention. Detailed implementation manners
[0059] Before detailing any embodiment of the present invention, it should be understood that in its application, the present invention is not limited to the construction and arrangement details of the components described in the following description or illustrated in the following drawings. The present invention is capable of having other embodiments and can be practiced or carried out in various ways. Additionally, it should be understood that the wording and terminology used herein are for the purpose of description and should not be considered restrictive. As used herein, the terms "including" or "having" and their variants are intended to cover the items listed hereinafter and their equivalents as well as additional items. Unless otherwise specified or limited, the terms "mounted", "connected", "supported" and "coupled" and their variants are used broadly and cover both direct and indirect mounting, connection, support and coupling. Further, "connected" and "coupled" are not limited to physical or mechanical connection or coupling.
[0060] Moreover, on the one hand, in the disclosure of the present invention, the orientation or positional relationship indicated by terms such as "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the above terms should not be construed as limiting the present invention; on the other hand, the term "one" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, while in other embodiments, the number of this element can be multiple. The term "one" should not be construed as limiting the quantity.
[0061] As multimedia content evolves towards ultra-long audio-visual forms of several hours, existing large vision-language models (LVLMs) face three technical challenges in video understanding tasks with ultra-long time series and multi-modal fusion. The current mainstream solutions mainly rely on two technical paths: fine-tuning and agent-based framework. However, they have the following essential defects:
[0062] (1) High computational resource dependence: Fine-tuning for ultra-long audio-visual content (≥2 hours) requires processing more than one million visual features and audio spectral units in the order of one hundred thousand, resulting in a super-linear growth in training costs;
[0063] (2) Excessive system architecture complexity: The agent framework needs to maintain independent subsystems such as visual parsing, speech recognition, and multimodal alignment simultaneously, resulting in significant communication overhead and coordination errors;
[0064] (3) Defects in temporal modeling: Existing solutions have core problems such as long-term dependence breakage, cross-modal attention drift, and audio-visual semantic mismatch in the processing of continuous audio-visual streams at the hour level;
[0065] (4) Limited generalization ability: Existing solutions are difficult to effectively handle cross-modal temporal correlations and fine-grained semantic understanding problems unique to long videos.
[0066] In view of these technical bottlenecks of the existing technology, the present invention proposes an ultra-long audio-visual understanding solution based on a vision-language model (VLM), aiming to reduce the computational resource requirements, simplify the system architecture, enhance the dependence on temporal information, and improve the generalization ability, so as to effectively solve the technical problems of ultra-long audio-visual understanding.
[0067] Figure 1 The flowchart of the ultra-long audio-visual understanding method based on the vision-language model of the present invention is shown. As Figure 1 shown, the ultra-long audio-visual understanding method based on the vision-language model of the present invention includes the following steps:
[0068] I. Multi-granularity intent recognition.
[0069] In the present invention, based on the semantic routing dynamic allocation mechanism, an innovative method of function call is adopted, combined with fine-tuning, to construct a multi-modal input adaptive processing framework. That is, the large language model (for example, ChatGLM) is fine-tuned by the method of function call, and the fine-tuned large language model is used to perform multi-granularity intent recognition on the user's question to determine the inquiry mode of the user's question.
[0070] At the same time, in the present invention, three inquiry modes are designed during multi-granularity intent recognition, namely, single-graph inquiry mode, audio content inquiry mode, and video content inquiry mode, so as to be able to dynamically and accurately allocate the processing path according to the characteristics of the user's question, greatly improving the efficiency and accuracy of intent recognition.
[0071] Among them, the large language model (for example, ChatGLM) is fine-tuned in the following manner to obtain the fine-tuned large language model:
[0072] 1. Set question-answer pairs including the user's question and its corresponding inquiry mode.
[0073] 2. Fine-tune the large language model by using the question-answer pairs through the method of function call to obtain the fine-tuned large language model.
[0074] Among them, the specific implementation of the function call is as follows: Using the question-and-answer pairs including the user's question and its corresponding inquiry pattern, that is, the QA data pairs to fine-tune the large language model. Examples of the said QA data are as follows:
[0075] {"prompt": "Is there a dog in the third picture?","response": "Single-picture inquiry"}
[0076] {"prompt": "Is there a dog in the video?","response": "Video content inquiry"}
[0077] {"prompt": "What is said at the 9th second of the audio?","response": "Audio content inquiry"}.
[0078] The user's question can be directly classified based on the fine-tuned large language model to determine whether the user's question is in the single-picture inquiry mode, the audio content inquiry mode or the video content inquiry mode.
[0079] II. Content recognition.
[0080] In the present invention, based on the inquiry mode and the user's question, the pictures, audio and video input by the user are recognized to obtain the recognition content.
[0081] Among them, if it is determined that the user's question is in the single-picture inquiry mode, then based on the user's question, a vision-language model is used to recognize the picture input by the user to obtain the single-picture recognition content. That is to say, for a single picture uploaded by the user, the relevant image recognition and semantic understanding functions are quickly started, and a vision-language model (VLM) is used to quickly understand the inquiry intention put forward by the user for the single picture, and perform fine-grained image parsing, supporting functions such as object recognition, scene understanding, and attribute reasoning. For example, when the user uploads a picture of a commodity, the user's desired commodity information, price, etc. can be accurately recognized.
[0082] How to specifically recognize a single picture is a basic function of the vision-language model (VLM) and does not need to be introduced in detail. Of course, in order to improve the recognition effect of the vision-language model (VLM) on a single picture, a single-sample prompt (One-shot) can be set for the vision-language model (VLM). For example:
[0083] [{"role":"user","content":"Please analyze the following user questions and extract the user's intent. Output this information in JSON format. Ensure that the JSON structure contains the keys: \"id\" and \"intent\".\nFor example,\nUser question: Is there a dog in the third picture?\nOutput: {'intent':'Single picture query'}\nUser question: What is the person wearing red clothes doing in the fifth picture?"}]
[0084] ```json
[0085] {
[0086] "intent":"Single picture query"
[0087] }。
[0088] If it is determined that the user question is in the audio content query mode, then based on the user question, automatic speech recognition technology is used to recognize the audio input by the user to obtain audio recognition content.
[0089] Specifically, an end-to-end speech recognition model (e.g., Whisper-Large model) can be used to perform speech recognition operations on the audio input by the user, such as transcribing the speech content in the audio and recognizing the audio theme, and can quickly give the audio recognition content.
[0090] If it is determined that the user question is in the video content query mode, then based on the user question, dynamic key frame extraction technology is used to extract key frames from the user input video to obtain key frame pictures, and based on the user question, a vision-language model is used to recognize the extracted key frame pictures to obtain video recognition content.
[0091] In the present invention, when obtaining video recognition content, semantic value evaluation and physical motion perception are fused to construct a content-sensitive dynamic frame extraction algorithm to achieve efficient representation of ultra-long videos; and through multi-level feature decoupling, redundant data is compressed while retaining key information, and dynamic frame rate adjustment and token compression strategies are implemented to balance computational efficiency and information integrity, and an adaptive frame extraction algorithm is developed to achieve efficient representation of ultra-long videos. Specifically, it includes:
[0092] 1. Scene segmentation.
[0093] The BaSSL (Boundary-Aware Self-Supervised Learning) model is used to perform scene segmentation on the user input video to obtain multiple segmented scenes, and based on the user question, the segmented scene corresponding to the user question is obtained.
[0094] The core of BaSSL lies in leveraging pseudo-boundary information to learn to capture context transitions between scenes during the pre-training phase, which is achieved through the following three novel boundary-aware pre-training tasks: (1) Shot-Scene Matching: to maximize similarity within a scene; (2) Context Group Matching: to further enhance similarity within a scene; (3) Pseudo-Boundary Prediction: to minimize similarity between scenes.
[0095] In the present invention, by returning the start and end times of the segmented scenes, scene selection can be performed according to the user's question, thereby obtaining the segmented scene corresponding to the user's question.
[0096] 2. Key frame extraction.
[0097] Use FFmpeg to extract key frames from the segmented scene corresponding to the user's question to obtain the initially extracted pictures.
[0098] FFmpeg will first parse the video stream corresponding to the segmented scene, read the encoded data in the video file, and understand the video format and encoding type. During the process of parsing the video stream, FFmpeg will check the type of each video frame, which is done by reading the frame header information in the video stream, and the frame header information contains the frame type flag. After identifying the frame type, FFmpeg will filter out I-frames (i.e., key frames), which do not depend on other frames and thus can be extracted independently.
[0099] 3. Semantic-aware redundancy removal.
[0100] Calculate the cosine similarity between the image feature vector of the initially extracted pictures and the text feature vector of the user's question based on the CLIP algorithm, and discard the initially extracted pictures with a cosine similarity lower than a threshold (e.g., lower than 85%), and only retain the initially extracted pictures with a high cosine similarity to obtain the final extracted key frame pictures.
[0101] Through scene segmentation, key frame extraction, and semantic-aware redundancy removal, the present invention can improve the retention rate of key information in ultra-long videos, increase the redundancy frame rejection rate, and support dynamic adjustment of the compression rate to adapt to different hardware computing powers, thereby effectively controlling the end-to-end latency and meeting industrial-level real-time requirements.
[0102] III. Multimodal information fusion.
[0103] In the present invention, when performing multimodal information fusion, a spatio-temporal prompting mechanism is used. That is, a time localization template, a spatial prompting template, and a causal chain constraint template are introduced.
[0104] Among them, the time positioning template is used to mark key time intervals. Through the time positioning template (such as "T-[start]-[end]"), key time intervals can be marked. For example, in a surveillance video, an abnormal time period from "21:30 to 21:35" can be located to guide the large language model to focus on the local features of this abnormal time period during multimodal information fusion.
[0105] The space prompt template is used to mark key areas, which guides the large language model to focus on the local features of this key area during multimodal information fusion based on coordinates (key area (x1, y1, x2, y2)).
[0106] The causal chain constraint template is used to annotate event logical relationships (such as "if A then B") and constrain the inference path of the large language model during multimodal information fusion. For example, in a traffic violation scenario, the large language model is forced to generate results in the order of "if a violation accident occurs → conduct a violation scenario analysis".
[0107] In this way, when designing the prompt template, a prompt template can be formed based on the time positioning template, space prompt template, and causal chain constraint template, so that the large language model can take into account time constraints, space constraints, and causal chain constraints when generating multimodal information fusion results, thereby improving the accuracy of the generated multimodal information fusion results.
[0108] At the same time, in the present invention, a hierarchical generation mechanism is used during multimodal information fusion. That is, a context-aware template and a feedback correction template are introduced.
[0109] Among them, the context-aware template is used to dynamically adjust the prompt template according to the video duration and the user's question. For example, if the user's question involves issues related to traffic accidents, the prompt template can be optimized to "Describe the content of the picture, ensure the content is coherent, and there are no line breaks. Please carefully observe the picture and describe the content in the picture in detail. Please carefully observe whether a traffic accident has occurred (such as vehicle collision / rear-end collision / scraping / breakdown, collision between motor vehicle and non-motor vehicle, collision between motor vehicle and pedestrian, collision between non-motor vehicle and pedestrian, collision between non-motor vehicles, etc.). If a traffic accident has occurred, please describe the accident in detail, including the type of accident, the time point of occurrence, the location of occurrence, the participants in the accident, the impact caused by the accident, and whether there are casualties, etc.". At the same time, according to the video duration, the granularity of frame extraction can be adjusted. The finer granularity is to extract one frame per second, and the prompt template is changed from "The 1st second describes... The 5th second describes...." to "The 1st second describes... The 2nd second describes....".
[0110] The feedback correction template is used to optimize the prompt template based on the reinforcement learning method. That is, if the multi-modal information fusion result generated by the large language model based on the prompt template does not meet the requirements, the feedback correction template can be used to fine-tune the prompt template for the corresponding segmentation scenario generated by the large language model according to the user preference data.
[0111] Therefore, the specific steps of multi-modal information fusion in the present invention include:
[0112] 1. Set up a time localization template, a spatial prompt template, and a causal chain constraint template, and form a prompt template based on the time localization template, the spatial prompt template, and the causal chain constraint template. The time localization template is used to mark the key time intervals, the spatial prompt template is used to mark the key regions, and the causal chain constraint template is used to annotate the event logical relationships.
[0113] 2. Set up a context awareness template, which is used to dynamically adjust the prompt template according to the video duration and the user's question.
[0114] 3. Set up a feedback correction template, which is used to optimize the prompt template based on the reinforcement learning method.
[0115] 4. Based on the optimized prompt template, use the large language model to perform multi-modal information fusion on the single-image recognition content, audio recognition content, and video recognition content. That is, input the optimized prompt template together with the single-image recognition content, audio recognition content, and video recognition content into the large language model, and the large language model generates the multi-modal information fusion result.
[0116] By introducing a spatio-temporal prompt mechanism and a hierarchical generation mechanism in multi-modal information fusion, the present invention can improve the F1-score (the harmonic mean of precision and recall) of the large language model in small-sample scenarios, improve the recall rate of long-video event retrieval, enhance ultra-long time series, and clearly process content at the hour level.
[0117] At the same time, in the present invention, an audio-video description fusion enhancement and complementary reasoning mechanism is adopted in multi-modal information fusion. That is, when the video recognition content is described vaguely, the audio recognition content is used to dominate the reasoning, that is, audio event detection → hypothesis generation → visual feature verification (such as "alarm sound" triggering infrared image analysis); when the audio recognition content is missing, the video recognition content is started to complete, that is, spatial attention locates the sound source area (such as pointing to the performer of the instrument) → generate a virtual spectrogram. Therefore, through cross-modal semantic alignment and description fusion, it is possible to achieve deep coordination and consistency enhancement of audio-visual information, break through the dynamic matching and complementary reasoning of audio-video descriptions, solve the problems of modal fragmentation and semantic conflicts in traditional methods, and improve the fine-grained alignment ability.
[0118] In addition, in the present invention, a knowledge-enhanced consistency verification mechanism is used for multimodal information fusion, that is, the knowledge bases of each segmentation scenario are accessed, and the knowledge of the knowledge bases is input into the large language model as context information during multimodal information fusion, so that the large language model generates a multimodal information fusion result based on the context information. For example, access the physical common sense knowledge base (such as the speed of sound is 340 m / s) to constrain the spatio-temporal logic of sound and picture: if "lightning" is detected but there is no thunder within 3 seconds, trigger the recalibration process. Thus, by introducing the knowledge-enhanced consistency verification mechanism, the problem of understanding long-tail concepts can be solved, the recognition rate of low-frequency entities can be improved, and the semantic completion efficiency can be improved.
[0119] IV. Answer generation.
[0120] The user question and the multimodal information fusion result are input into the vision-language model, and the vision-language model generates the corresponding answer to the user question.
[0121] Figure 2 Fig. shows the schematic composition diagram of the ultra-long audio and video understanding system based on the vision-language model of the present invention. As Figure 2 shown, the ultra-long audio and video understanding system based on the vision-language model of the present invention includes:
[0122] 1. Multigranularity intention recognition module.
[0123] The multigranularity intention recognition module is used to perform multigranularity intention recognition on the user question by using the fine-tuned large language model to determine the inquiry mode of the user question, and the inquiry mode includes a single-image inquiry mode, an audio content inquiry mode, and a video content inquiry mode.
[0124] 2. Content recognition module.
[0125] The content recognition module is used to recognize the pictures, audio, and video input by the user based on the inquiry mode and the user question to obtain the recognized content.
[0126] In the present invention, the content recognition module includes a single-image recognition sub-module, an audio recognition sub-module, and a video recognition sub-module.
[0127] For the single-image inquiry mode, the single-image recognition sub-module uses the vision-language model to recognize the pictures input by the user based on the user question to obtain the single-image recognition content.
[0128] For the audio content inquiry mode, the audio recognition sub-module uses automatic speech recognition technology to recognize the audio input by the user based on the user question to obtain the audio recognition content.
[0129] For the video content query mode, the video recognition sub-module extracts key frames from the user-input video using dynamic key frame extraction technology based on the user's question to obtain key frame pictures, and uses a vision-language model to recognize the extracted key frame pictures based on the user's question to obtain video recognition content.
[0130] 3. Multimodal information fusion module.
[0131] The multimodal information fusion module is used to perform multimodal information fusion on the recognition content using a large language model based on a spatio-temporal prompt mechanism and a hierarchical generation mechanism.
[0132] 4. Answer generation module.
[0133] The answer generation module is used to input the user's question and the multimodal information fusion result into a vision-language model, and the vision-language model generates the corresponding answer to the user's question.
[0134] In addition, the present invention also provides a very long audio-visual understanding device based on a vision-language model, which includes: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the very long audio-visual understanding method based on the vision-language model as described above.
[0135] Finally, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the very long audio-visual understanding method based on the vision-language model as described above are implemented.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the protection scope of the present invention. Those skilled in the art can modify or equivalently replace the technical solutions of the present invention according to the idea of the present invention, without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for understanding ultra-long audio-visual content based on a vision-language model, characterized in that, It includes the following steps: 1) Multi-granularity intent recognition: Use the fine-tuned large language model to perform multi-granularity intent recognition on the user's question to determine the inquiry pattern of the user's question. The inquiry pattern includes a single-image inquiry pattern, an audio content inquiry pattern, and a video content inquiry pattern; 2) Content recognition: Based on the inquiry pattern and the user's question, recognize the pictures, audio, and video input by the user to obtain the recognized content; 3) Multi-modal information fusion: Use the large language model to perform multi-modal information fusion on the recognized content based on the spatio-temporal prompt mechanism and the hierarchical generation mechanism; 4) Answer generation: Input the user's question and the multi-modal information fusion result into the vision-language model, and the vision-language model generates the corresponding answer to the user's question.
2. The method for understanding ultra-long audio-visual content based on a vision-language model according to claim 1, wherein The specific steps of step 2) include: 21) For the single-image inquiry pattern, use the vision-language model based on the user's question to recognize the picture input by the user to obtain the single-image recognized content; 22) For the audio content inquiry pattern, use automatic speech recognition technology based on the user's question to recognize the audio input by the user to obtain the audio recognized content; 23) For the video content inquiry pattern, use the dynamic key frame extraction technology based on the user's question to extract key frames from the video input by the user to obtain key frame pictures, and use the vision-language model based on the user's question to recognize the extracted key frame pictures to obtain the video recognized content.
3. The method for understanding ultra-long audio and video based on a vision-language model according to claim 2, wherein The specific steps of using the dynamic key frame extraction technology based on the user's question to extract key frames from the video input by the user to obtain key frame pictures in step 23) include: 231) Scene segmentation: Use the BaSSL model to perform scene segmentation on the video input by the user to obtain multiple segmented scenes, and obtain the segmented scene corresponding to the user's question based on the user's question; 232) Key frame extraction: Use FFmpeg to extract key frames from the segmented scene corresponding to the user's question to obtain the preliminarily extracted pictures; 233) Semantic-aware redundancy removal: Calculate the similarity between the image feature vector of the preliminarily extracted picture and the text feature vector of the user's question based on the CLIP algorithm, and discard the preliminarily extracted pictures with similarity lower than the threshold to obtain the finally extracted key frame pictures.
4. The method for understanding ultra-long audio-visual content based on a vision-language model according to claim 3, characterized in that The specific steps of step 3) include: 31) Set the time localization template, the spatial prompt template, and the causal chain constraint template, and form a prompt template based on the time localization template, the spatial prompt template, and the causal chain constraint template. The time localization template is used to mark the key time interval, the spatial prompt template is used to mark the key area, and the causal chain constraint template is used to annotate the event logical relationship; 32) Set the context-aware template, which is used to dynamically adjust the prompt template according to the video duration and the user's question; 33) Set the feedback correction template, which is used to optimize the prompt template based on the reinforcement learning method; 34) Use the large language model based on the optimized prompt template to perform multi-modal information fusion on the single-image recognized content, the audio recognized content, and the video recognized content.
5. The method for understanding ultra-long audio-visual content based on a vision-language model according to claim 4, wherein In step 34), when performing multi-modal information fusion, an audio-video description fusion enhancement and complementary reasoning mechanism is adopted, that is, when the description of the video recognition content is ambiguous, the audio recognition content is used to dominate the reasoning, and when the audio recognition content is missing, the video recognition content is activated for completion.
6. The method for understanding ultra-long audio-visual content based on a vision-language model according to claim 5, wherein In step 34), when performing multi-modal information fusion, a knowledge-enhanced consistency verification mechanism is used, that is, the knowledge bases of each segmented scene are accessed, and the knowledge of the knowledge bases is input into the large language model as context information during multi-modal information fusion, so that the large language model generates a multi-modal information fusion result based on the context information.
7. The method for understanding ultra-long audio and video based on a vision-language model according to any one of claims 1-6, characterized in that The large language model is fine-tuned in the following manner to obtain the fine-tuned large language model: Set up question-answer pairs including user questions and their corresponding query patterns; Fine-tune the large language model using the question-answer pairs by means of function calls.
8. An ultra-long audio-visual understanding system based on a vision-language model, characterized in that, It includes: A multi-granularity intent recognition module, which is used to perform multi-granularity intent recognition on user questions using the fine-tuned large language model to determine the query pattern of the user questions, and the query pattern includes a single-image query pattern, an audio content query pattern, and a video content query pattern; A content recognition module, which is used to recognize the pictures, audio, and video input by the user based on the query pattern and the user questions to obtain recognition content; A multi-modal information fusion module, which is used to perform multi-modal information fusion on the recognition content using the large language model based on a spatio-temporal prompt mechanism and a hierarchical generation mechanism; An answer generation module, which is used to input the user questions and the multi-modal information fusion result into a vision-language model, and the vision-language model generates the corresponding answers to the user questions.
9. An ultra-long audio-video understanding device based on a vision-language model, characterized in that, It includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method for understanding ultra-long audio-visual videos based on a vision-language model according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method for understanding ultra-long audio-visual videos based on a vision-language model according to any one of claims 1-7.
Citation Information
Cited By
Method and system for realizing real-time two-way visual interaction of digital human by calling camera
CN120523334A
Method and system for realizing real-time two-way visual interaction of digital human by calling camera
CN120523334B
Feedback force prediction method and device in virtual operation process
CN121075559A
A method and device for predicting feedback force during virtual surgery
CN121075559B
Long video visual question and answer method and device based on large model agent and storage medium
CN121094114A