A multi-modal inference method and device applied to streaming video

CN122842011APending Publication Date: 2026-09-29SHANGHAI MOUSHEN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610896915.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

该类方法能够利用视觉信息和文本信息完成视频问答,但当视频时间较长或视频以流式形式持续输入时,输入帧数量会不断增加,视觉Token数量也随之增长,从而带来显存占用升高和推理延迟增大的问题

Benefits of technology

本申请技术方案提供的应用于流式视频的多模态推理方法,获取目标文本信息以及目标流式视频后,从目标流式视频中确定候选帧集合,并根据候选帧集合中每个初始帧对应的多模态特征和相关性数据,确定至少一个关键帧,能够减少视频推理过程中所需要处理的视频帧数量,减少显存占用,以改善推理延迟,提高视频推理的实时性。同时,由于在关键帧确定的过程中引入多模态特征以及相关性数据,因此关键帧既能覆盖目标流式视频中的主要视觉内容和主要语义内容,又能更好服务于目标文本信息,从而在保证视频推理的实时性的同时,提高视频推理的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842011A_ABST
    Figure CN122842011A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a multi-modal inference method and device applied to streaming video. The method comprises: obtaining to-be-processed information; the to-be-processed information comprises target text information and a target streaming video corresponding to the target text information; determining a candidate frame set from the target streaming video based on a preset sampling condition; the candidate frame set comprises at least one initial frame; calculating multi-modal features and correlation data corresponding to the at least one initial frame; determining at least one key frame from the at least one initial frame based on the multi-modal features and the correlation data; performing block pre-padding processing based on the at least one key frame and the target text information to obtain a plurality of continuous cache blocks; and performing decoding processing on the plurality of cache blocks to obtain a video inference result corresponding to the to-be-processed information. Through the above technical features, the problem of excessive video memory occupation and excessive inference delay can be improved while ensuring the accuracy of video understanding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and multimodal video understanding technology, and in particular to a multimodal reasoning method and apparatus for streaming video. Background Technology

[0002] With the development of multimodal large-scale models, these models are now capable of simultaneously processing image, video, and natural language information, demonstrating strong capabilities in tasks such as visual question answering, video understanding, intelligent monitoring, autonomous driving, and real-time video assistants. Compared to static image and short video understanding, streaming video understanding is closer to real-world application scenarios. In scenarios such as video surveillance, intelligent transportation, robot visual perception, online meeting understanding, and real-time video question answering, video content is typically continuously arriving. The system needs to continuously receive new video frames over time and complete understanding and reasoning based on historical video content prior to the current moment.

[0003] In related technologies, video frames are typically first input into a visual encoder to obtain visual tokens. These visual tokens, along with the user's text question, are then input into a language model, which generates the answer. This type of method can complete video question answering using both visual and textual information. However, when the video is long or continuously input in a streaming format, the number of input frames increases, and the number of visual tokens also grows, leading to increased memory usage and inference latency. Summary of the Invention

[0004] In view of the above-mentioned deficiencies in the related technologies, the technical problem to be solved by the embodiments of this application is how to improve the real-time performance and accuracy of video reasoning.

[0005] According to a first aspect of this application, a multimodal inference method for streaming video is provided. The method includes: Acquire information to be processed; the information to be processed includes target text information and the target streaming video corresponding to the target text information. Based on preset sampling conditions, a candidate frame set is determined from the target streaming video; the candidate frame set includes at least one initial frame. Calculate the multimodal features and correlation data corresponding to at least one initial frame; Based on multimodal features and correlation data, at least one key frame is determined from at least one initial frame; Based on at least one keyframe and target text information, block pre-filling processing is performed to obtain multiple consecutive cache blocks; Multiple buffer blocks are decoded to obtain video inference results corresponding to the information to be processed.

[0006] In some embodiments, multimodal features include visual features and semantic features; Calculate the multimodal features and correlation data corresponding to at least one initial frame, including: Calculate the visual and semantic features corresponding to each initial frame; Semantic encoding is performed on the target text information to obtain a text semantic vector; Based on semantic features and text semantic vectors, the relevance data corresponding to each initial frame is determined.

[0007] In some embodiments, multimodal features include visual features and semantic features; After calculating the multimodal features and correlation data corresponding to at least one initial frame, the method further includes: Semantic aggregation is performed on the semantic features corresponding to each initial frame to obtain the overall semantic vector corresponding to the candidate frame set; Semantic similarity is determined based on semantic features and the overall semantic vector.

[0008] In some embodiments, determining at least one key frame from at least one initial frame based on multimodal features and correlation data includes: Based on the visual and semantic features corresponding to each initial frame, a fusion candidate set is constructed; the fusion candidate set includes at least one candidate frame. Calculate the comprehensive score data corresponding to at least one candidate frame; Based on the comprehensive scoring data, at least one candidate frame is filtered to obtain a set of key frames; the set of key frames includes at least one key frame.

[0009] In some embodiments, the keyframe set is initially an empty set; Calculate the comprehensive score data corresponding to at least one candidate frame, including: Calculate the redundancy parameters between candidate frames and keyframes; The semantic similarity, relevance data, and redundant parameters are weighted to obtain the comprehensive score data.

[0010] In some embodiments, based on comprehensive scoring data, at least one candidate frame is filtered to obtain a set of keyframes, including: Based on the comprehensive score data corresponding to at least one candidate frame, key frames are determined from at least one candidate frame; Update the fusion candidate set and the keyframe set; Based on the updated fusion candidate set and the updated keyframe set, update the redundant parameters and the comprehensive score data; Based on the updated redundancy parameters and the updated comprehensive score data, the next keyframe is determined until the number of keyframes in the keyframe set equals the preset number of frames.

[0011] In some embodiments, keyframes correspond to video timestamps; Before performing block pre-filling processing based on at least one keyframe and target text information to obtain multiple consecutive cache blocks, the method includes: Based on the video timestamps corresponding to at least one keyframe, sort at least one keyframe to obtain a set of sequential keyframes.

[0012] In some embodiments, block pre-filling processing is performed based on at least one keyframe and target text information to obtain multiple consecutive cache blocks, including: Visual encoding is performed on the keyframes included in the sequential keyframe set to obtain the first encoded information; The target text information is segmented into words to obtain the second encoded information; A unified input sequence is constructed based on the first and second encoded information; The unified input sequence is divided into blocks and pre-filled to obtain buffer blocks.

[0013] In some embodiments, the uniform input sequence is pre-padded in blocks to obtain a cache block, including: Based on a preset unit length, the unified input sequence is divided into blocks to obtain multiple data blocks; Multiple data blocks are pre-filled to obtain cache blocks.

[0014] In some embodiments, the method further includes: Generate runtime log data corresponding to the video inference results.

[0015] According to a second aspect of this application, a multimodal inference apparatus for streaming video is provided, comprising: The information acquisition module is used to acquire information to be processed; the information to be processed includes target text information and target streaming video corresponding to the target text information; The set determination module is used to determine a set of candidate frames from the target streaming video based on preset sampling conditions; the set of candidate frames includes at least one initial frame; The calculation module is used to calculate the multimodal features and correlation data corresponding to at least one initial frame; The keyframe determination module is used to determine at least one keyframe from at least one initial frame based on multimodal features and correlation data. The block pre-filling processing module is used to perform block pre-filling processing based on at least one keyframe and target text information to obtain multiple consecutive cache blocks. The decoding processing module is used to decode multiple buffer blocks to obtain video inference results corresponding to the information to be processed.

[0016] According to a third aspect of this application, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction and at least one program, the at least one instruction and at least one program being loaded and executed by the processor to implement the multimodal inference method for streaming video as described above.

[0017] According to a fourth aspect of this application, a computer storage medium is provided that stores at least one instruction and at least one program, wherein the at least one instruction and at least one program are loaded and executed by a processor to implement the multimodal inference method for streaming video as described above.

[0018] According to a fifth aspect of this application, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the multimodal reasoning method for streaming video as described above.

[0019] Compared with related technologies, the technical solution provided in this application has the following beneficial effects: The multimodal inference method for streaming video provided in this application acquires target text information and target streaming video, then determines a candidate frame set from the target streaming video, and determines at least one key frame based on the multimodal features and correlation data corresponding to each initial frame in the candidate frame set. This reduces the number of video frames that need to be processed during video inference, reduces memory usage, improves inference latency, and enhances the real-time performance of video inference. Furthermore, because multimodal features and correlation data are introduced during key frame determination, the key frames can cover both the main visual and semantic content of the target streaming video and better serve the target text information, thereby improving the accuracy of video inference while ensuring its real-time performance. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the multimodal reasoning method provided in the embodiments of this application. Figure 2 A flowchart illustrating the determination of the keyframe set provided in the embodiments of this application; Figure 3 Another flowchart illustrating the determination of the keyframe set provided in the embodiments of this application; Figure 4 This is a schematic diagram corresponding to the block pre-filling process provided in the embodiments of this application; Figure 5 Performance evaluation data tables corresponding to different basic models provided in the embodiments of this application; Figure 6This application provides an experimental data table corresponding to different sampling frame numbers in its embodiments. Figure 7 This application provides a trend graph showing the relationship between frame count, accuracy, and latency for embodiments of the application. Figure 8 The system architecture diagram corresponding to the multimodal inference system provided in the embodiments of this application; Figure 9 This is a schematic diagram of the structure of the multimodal inference device provided in the embodiments of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0022] For video understanding tasks involving long or streaming videos, streaming video understanding methods based on complete historical context can preserve as much historical video information as possible. However, because the input data size increases with video duration, this can easily lead to problems such as excessive memory usage, high inference latency, and difficulties in single-GPU deployment in long-duration streaming video scenarios. For online video understanding methods based on fixed time windows, only the most recent video content is typically retained to keep the input data size relatively stable, while historical information outside the window is discarded, which can easily result in inaccurate video understanding results.

[0023] Multimodal inference methods based on visual token compression compress the number of tokens that need to be processed after the video frame passes through the visual encoder. However, they mainly function within the visual representation of a single frame and are insufficient for handling the cross-frame redundancy between a large number of adjacent frames in streaming video. Furthermore, if visual token compression is too strong, it may damage key information such as target details, textual information, spatial relationships, and dynamic processes, affecting the video understanding effect.

[0024] In view of this, the technical solution of this application aims to solve the problems of poor reliability and low efficiency of reasoning results in long video reasoning tasks. The following description will be provided in conjunction with the accompanying drawings.

[0025] The multimodal inference method for streaming video disclosed in this application can be executed by any terminal equipped with a large multimodal model and capable of running the method; specifically, any terminal can be a server terminal, a mobile terminal, etc. Figure 1 As shown, the method may include: Step S101: Obtain information to be processed; the information to be processed includes target text information and target streaming video corresponding to the target text information.

[0026] The target text information is the input question that needs to be used for reasoning based on the target streaming video; in other words, the target text information represents the video reasoning task that needs to be performed based on the target streaming video.

[0027] Streaming video input refers to a form of input where video content arrives continuously over time, and at any given moment, understanding and reasoning can only be performed using video content that has been available up to that point. In a specific embodiment, the moment when reasoning about the streaming video is required, i.e., the moment when the target text information is input or received, can be considered the current moment. The target timestamp corresponding to the target text information is the current moment.

[0028] Any time before the current moment is considered a historical moment, and the streaming video of a historical moment is called historical streaming video. In historical streaming video, each video frame corresponds to a video timestamp. When it is necessary to perform video inference on historical streaming video in conjunction with target text information, the target streaming video can be extracted from the historical streaming video based on the target text information and its corresponding target timestamp.

[0029] The video reasoning process has a starting point and an ending point, both of which correspond to video timestamps. These timestamps are used to limit which part of the historical streaming video is being reasoned about. The video timestamp corresponding to the ending point is fixed to the target timestamp by default. The video timestamp corresponding to the starting point can be determined based on the target text information, but it must be less than or equal to the target timestamp.

[0030] Specifically, when the target text information does not explicitly mention the starting point for the video reasoning process, the first video frame of the historical streaming video can be considered the starting point. When the target text information explicitly states, or can be inferred, that the nth frame of the historical streaming video, or timestamp A1, is the starting point, then the nth frame or timestamp A1 is selected as the starting point. The target streaming video is then determined from the historical streaming video based on the starting and ending points of the reasoning.

[0031] By adaptively adjusting the inference starting point based on the target text information to determine the target streaming video, it is possible to avoid subsequent processing of irrelevant video frames, thereby reducing data volume and video memory usage, and saving computing resources.

[0032] As one implementation method, the number of video frames included in the historical streaming video can be greater than or equal to the number of video frames included in the target streaming video.

[0033] Step S102: Based on preset sampling conditions, determine a candidate frame set from the target streaming video; the candidate frame set includes at least one initial frame.

[0034] In one specific embodiment, the preset sampling conditions can correspond to sampling methods such as fixed frame rate, fixed time interval, or uniform sampling. Determining a candidate frame set from the target streaming video involves extracting frames from at least one video frame included in the target streaming video to obtain at least one initial frame. This at least one initial frame constitutes the candidate frame set.

[0035] As one implementation method, the preset sampling conditions can also correspond to sampling methods such as sliding window sampling or adaptive sampling based on changes in video content.

[0036] By extracting frames from the complete target streaming video, a candidate frame set is obtained as a candidate space of controllable size, providing a basis for subsequent sampling of keyframes.

[0037] Step S103: Calculate the multimodal features and correlation data corresponding to at least one initial frame.

[0038] In some embodiments, multimodal features may include at least visual features and semantic features. Multimodal features may also include image sharpness, motion intensity, etc. This embodiment uses the inclusion of visual and semantic features as an example for illustration.

[0039] Specifically, step S103 involves calculating the multimodal features and correlation data corresponding to at least one initial frame, including: Calculate the visual and semantic features corresponding to each initial frame.

[0040] Semantic encoding is performed on the target text information to obtain a text semantic vector.

[0041] Based on semantic features and text semantic vectors, the relevance data corresponding to each initial frame is determined.

[0042] As one implementation, visual features are calculated for each initial frame. These visual features measure the visual differences, or visual diversity, between the initial frames. RGB color histograms representing the underlying visual content of the initial frames can be used as visual features. Alternatively, image features output by the visual encoder, HSV color histograms, image gradient features, convolutional neural network features, etc., can also be used as visual features.

[0043] As one implementation method, for each initial frame, a visual semantic coding model such as CLIP (Contrastive Language-Image Pre-training) can be used to encode the initial frame to obtain the semantic features corresponding to the initial frame. The semantic features can be high-level semantic vectors used to represent the objects, scenes, actions, and event content in the initial frame, as well as to measure the semantic representativeness of the video frame and the relevance between the video frame and the target text information.

[0044] Relevance data characterizes the semantic features of each initial frame and their similarity to the text semantic vector. Higher relevance data, meaning greater similarity between the semantic features and the text semantic vector, indicates that the initial frame is more likely to contain the valid visual information needed to answer the current text question. It can be calculated using the following formula:

[0045] in, This represents the correlation data corresponding to the initial frame; Indicates the i-th initial frame; q represents the target text information; The semantic features representing the initial frame, This represents the text semantic vector corresponding to the target text information; This represents the cosine similarity.

[0046] As another implementation method, the semantic features of the initial frame can also be extracted using the visual encoder in other pre-trained image-text models or multimodal large models.

[0047] As another implementation method, relevance data can also be calculated using methods such as dot product similarity, cross-modal attention score, or image-text matching model score.

[0048] Furthermore, after calculating the multimodal features and correlation data corresponding to at least one initial frame, the method also includes: Semantic aggregation is performed on the semantic features corresponding to each initial frame to obtain the overall semantic vector corresponding to the candidate frame set; Semantic similarity is determined based on semantic features and the overall semantic vector.

[0049] Specifically, semantic vector aggregation can be performed on the semantic features of all initial frames to obtain an overall semantic vector corresponding to all initial frames, i.e., the candidate frame set. Then, the cosine similarity between the semantic features of each initial frame and the overall semantic vector is calculated, serving as the semantic similarity for each initial frame. This semantic similarity is used to subsequently select keyframes from the candidate frame set.

[0050] As another implementation method, local semantic uniqueness can be introduced to improve the similarity evaluation. For example, instead of directly applying semantic similarity to the selection of keyframes, it can be weighted with the average semantic distance between each initial frame and the remaining initial frames to obtain a fusion score, and then used in the selection of subsequent keyframes.

[0051] Step S104: Based on multimodal features and correlation data, determine at least one key frame from at least one initial frame.

[0052] In one specific embodiment, at least one keyframe is determined from at least one initial frame. Therefore, the determination of keyframes is based on at least two rounds of filtering of video frames. By identifying keyframes more valuable to the current video inference task from the target streaming video through two rounds of filtering, the amount of input data for subsequent model inference can be reduced, thereby reducing redundant computations in the video inference process from the source and reducing GPU memory usage. Simultaneously, determining keyframes by combining multimodal features and relevance data takes into account the visual differences between various video frames in the target streaming video, the semantic representativeness of the video content, the relevance between video frames and the current text question, and the degree of redundancy among selected frames. This ensures that the keyframe set covers both the main visual and semantic content of the video and better serves the target text information, thereby improving the reliability and accuracy of the subsequent video inference process.

[0053] Specifically, such as Figure 2 As shown, based on multimodal features and correlation data, at least one keyframe is determined from at least one initial frame, including: Step S201: Construct a fusion candidate set based on the visual and semantic features corresponding to each initial frame; the fusion candidate set includes at least one candidate frame.

[0054] In a specific embodiment, after obtaining the visual features and semantic features corresponding to each initial frame, the initial frames are sampled according to the visual features to obtain a first candidate set, and the initial frames are sampled according to the semantic features to obtain a second candidate set.

[0055] As one implementation method, the distance or similarity between different initial frames can be calculated, and initial frames with more dispersed visual content can be determined based on the inter-frame differences and visual features of different initial frames to form a first candidate set. Furthermore, initial frames that can cover the main semantic content can be selected based on the distribution and semantic features of each initial frame in the semantic space. For example, initial frames with more dispersed semantic features can be selected, or initial frames that are close to but do not overlap with the overall semantic center of the video can be selected to form a second candidate set.

[0056] Furthermore, the union of the first and second candidate sets is taken, and duplicate video frames are removed to obtain the fused candidate set. Determining the fused candidate set can narrow down the selection space for subsequent keyframes and ensure that the candidate frames included in the fused candidate set include both frames with significant visual differences and frames with important semantic content.

[0057] In this embodiment, determining the first candidate set reduces the occupancy of consecutive similar video frames in the final input and improves the keyframe set's coverage of different scenes and image changes. Determining the second candidate set ensures that the keyframe sets are not only visually distinct but also semantically cover the main content of the video. Therefore, fusing the candidate frames included in the candidate set not only covers the overall video content but also allows for adaptive selection based on target text information.

[0058] Step S202: Calculate the comprehensive score data corresponding to at least one candidate frame.

[0059] In one specific embodiment, the keyframe set S is initially empty. Alternatively, the keyframe set S can be initialized to an empty set before calculating the comprehensive score data corresponding to each candidate frame.

[0060] Step S202 may include: calculating the redundancy parameters between candidate frames and key frames; and weighting the semantic similarity, relevance data, and redundancy parameters to obtain comprehensive score data.

[0061] Redundancy parameters are used to represent the degree of redundancy in the candidate frame and keyframe sets. Redundancy can be defined as the maximum semantic similarity between a candidate frame and at least one keyframe. For candidate frames... The corresponding redundancy parameters can be determined using the following formula:

[0062] in, Indicates candidate frames Redundant parameters; This represents the j-th keyframe; Indicates candidate frames semantic features; Keyframe Semantic features.

[0063] When a candidate frame is too similar to a keyframe, its redundancy is high, and the overall score will be lowered, thus reducing the probability of duplicate video frames being selected. Additionally, when the keyframe set is empty, the redundancy parameter can be considered to be 0.

[0064] For candidate frames The overall score can be calculated using the following formula:

[0065] in, Indicates candidate frames Comprehensive scoring data; Indicates candidate frames Corresponding correlation data; Indicates candidate frames Corresponding semantic similarity; , , These are the weight parameters.

[0066] Step S203: Based on the comprehensive scoring data, at least one candidate frame is filtered to obtain a key frame set; the key frame set includes at least one key frame.

[0067] As one implementation method, a greedy selection strategy can be used to iteratively select keyframes from candidate frames by combining comprehensive scoring data.

[0068] As other implementation methods, keyframes can also be selected from candidate frames using methods such as Top-K sorting, cluster selection, maximum marginal relevance selection, or learnable sampler selection.

[0069] By calculating comprehensive scoring data, the selection of keyframes can be dynamically adjusted based on the target text information. Furthermore, by eliminating overly repetitive candidate frames, the information coverage and utilization rate of the keyframe set can be improved, thereby enhancing the reliability of the inference results output.

[0070] Specifically, step S203 may include: Based on the comprehensive score data corresponding to at least one candidate frame, key frames are determined from at least one candidate frame; Update the fusion candidate set and the keyframe set; Based on the updated fusion candidate set and the updated keyframe set, update the redundant parameters and the comprehensive score data; Based on the updated redundancy parameters and the updated comprehensive score data, the next keyframe is determined until the number of keyframes in the keyframe set equals the preset number of frames.

[0071] As one implementation method, a greedy selection strategy can be used to iteratively select keyframes. In each round, the comprehensive score data corresponding to the candidate frames that have not yet been selected is calculated, and the candidate frame corresponding to the maximum comprehensive score data is determined as a keyframe and added to the keyframe set. Redundancy items are updated until the number of keyframes in the keyframe set equals the preset number of frames.

[0072] Specifically, after initializing the keyframe set S to be empty, or in other words, when the keyframe set S is initially empty, the comprehensive score data for each candidate frame is calculated. At this point, the redundant parameter in the comprehensive score data calculation formula is 0. In the comprehensive score data corresponding to each candidate frame, the candidate frame corresponding to the maximum comprehensive score data is determined as the keyframe. This keyframe becomes the first element in the keyframe set.

[0073] Further, the keyframe is removed from the fusion candidate set, resulting in an updated fusion candidate set that includes the remaining candidate frames excluding the keyframe. Based on the updated keyframe set, redundancy parameters are calculated for each of the remaining candidate frames, i.e., the redundancy parameters are updated. These redundancy parameters are then used to calculate the comprehensive score data for each of the remaining candidate frames, and the candidate frame with the highest comprehensive score data is determined as the next keyframe. This keyframe becomes the second element in the keyframe set.

[0074] This process continues until the keyframe set includes a preset number of keyframes.

[0075] Step S105: Perform block pre-filling processing based on at least one keyframe and target text information to obtain multiple consecutive buffer blocks.

[0076] In one specific implementation, since each video frame corresponds to a video timestamp, and key frames are selected from the initial frame, that is, the video frames after frame extraction, key frames also correspond to video timestamps.

[0077] Since video understanding tasks depend on action changes and the order of events, after keyframe selection is completed, block pre-filling is not performed directly. Instead, the keyframes are reordered according to their video timestamps in the target streaming video. That is, before block pre-filling, the method includes: sorting at least one keyframe based on the video timestamps corresponding to at least one keyframe to obtain a set of sequential keyframes.

[0078] By organizing keyframes according to the video timestamp order after sampling, a set of sequential keyframes can be obtained, which can preserve the temporal structure of event development and action changes in the target streaming video and reduce the risk of temporal relationship disruption caused by sampling.

[0079] like Figure 3 As shown, in a specific embodiment, after obtaining the target text information, a candidate frame set F is determined from historical video clips. Visual diversity filtering is then performed on the candidate frame set F to obtain a first candidate set. And a second candidate set is obtained by semantic coverage filtering of the candidate frame set F. The first candidate set and the second candidate set Perform fusion and deduplication to obtain a fusion candidate set. The process involves calculating the problem relevance score (relevance data), semantic representativeness score (semantic similarity), and redundancy penalty (redundancy parameter) for each candidate frame in the fusion candidate set to obtain comprehensive score data. Then, a greedy iterative selection of keyframes is performed based on the comprehensive score data. When the number of keyframes equals the target number of frames K (the preset number of frames), a keyframe set S is obtained. This keyframe set S is then sorted according to the original video timestamps to obtain a sequential keyframe set S'.

[0080] Furthermore, a unified input sequence can be generated based on the sequential keyframe set, and the unified input sequence can be pre-filled in blocks to divide it into multiple consecutive buffer blocks. By dividing a long unified input sequence into multiple fragmented consecutive buffer blocks and processing them sequentially, the peak memory pressure caused by a one-time concentrated pre-filling of long sequences can be alleviated, and the stability of long video inference can be improved.

[0081] In one specific embodiment, block pre-filling processing is performed based on at least one keyframe and target text information to obtain multiple consecutive cache blocks, including: Visual encoding is performed on the keyframes included in the sequential keyframe set to obtain the first encoded information; The target text information is segmented into words to obtain the second encoded information; A unified input sequence is constructed based on the first and second encoded information; The unified input sequence is divided into blocks and pre-filled to obtain buffer blocks.

[0082] In one implementation, the keyframes in the sequential keyframe sequence are first input into a visual encoder for visual encoding processing to obtain the first encoded information corresponding to each keyframe. The first encoded information can be a visual token, which represents a discrete or continuous visual representation unit obtained after the keyframe has passed through the visual encoder.

[0083] The target text information is then input into a tokenizer to obtain a text token, and the visual token and text token are combined into a unified input sequence X. This sequence X may include multiple elements, represented in the following form:

[0084] in, The sequence length represents the unified input sequence. In some embodiments, the unified input sequence may include system prompts, keyframe visual information, user questions, and answer prompts.

[0085] As one implementation method, a chunk prefill mechanism can be used to prefill a uniform input sequence in blocks. The chunk prefill mechanism divides a long input sequence into multiple smaller segments and performs prefill calculations block by block sequentially. This mechanism can reduce the resource pressure caused by processing long sequences all at once. Figure 4 As shown, the unified input sequence is pre-padded in blocks to obtain a buffer block, which includes: Based on a preset unit length, the unified input sequence is divided into blocks to obtain multiple data blocks; Multiple data blocks are pre-filled to obtain cache blocks.

[0086] Specifically, given a unified input sequence containing N tokens, the N tokens can be divided into multiple data blocks, such as M data blocks, according to a preset unit length and the order of the N tokens. This can be done using... Indicates the first One input block.

[0087] As another implementation method, the data block division is not limited to the preset unit length, but can also be adaptively divided according to the input length, modal boundary, frame boundary or buffer capacity.

[0088] Each can be processed sequentially according to the original input order. Perform pre-filling to generate the corresponding cache blocks. After processing each... The multimodal large model updates the corresponding intermediate states and cached information before continuing to process the next one. .

[0089] Compared to processing the entire input sequence at once, the Chunk Prefill mechanism can reduce the peak memory pressure in a single forward computation, making the processing of long video inputs smoother.

[0090] It's important to note that Chunk Prefill itself does not reduce the total number of input tokens; its main function is to optimize the computational organization of long inputs. Therefore, by sampling keyframes to reduce the number of redundant video frames and visual tokens at the input source, and in conjunction with Chunk Prefill to alleviate the concentrated computational pressure on the remaining long input sequences, this collaborative optimization on both the input and computation sides can improve the efficiency and deployability of multimodal inference for streaming video while ensuring the reliability of video understanding.

[0091] Step S106: Decode multiple buffer blocks to obtain video inference results corresponding to the information to be processed.

[0092] As one implementation method, after completing the block pre-filling, the decoding stage begins. Multiple cache blocks can be decoded based on a multimodal large model to generate video inference results corresponding to the target text information. These video inference results can be natural language responses, object recognition results, action judgment results, event understanding results, or video summarization results, etc.

[0093] This application does not limit the specific type of multimodal large model used. The multimodal inference method applied to streaming video is adaptable to Qwen-VL, LLaVA series, MiniCPM-V or other multimodal models with video / image understanding capabilities, and has good engineering adaptability and practical deployment value.

[0094] Furthermore, after obtaining the video inference result, the method may also include: generating runtime log data corresponding to the video inference result.

[0095] Runtime log data can be used for subsequent performance analysis, result reproduction, parameter tuning, and other processes. Specifically, runtime log data may include target timestamps, index information corresponding to keyframes, video timestamps corresponding to keyframes, number of keyframes, number of first encoding information, number of data blocks, inference latency information, video memory usage information, and sampling parameters, etc.

[0096] Experiments have verified that this invention maintains good understanding performance in streaming video understanding tasks while reducing the number of input frames. Specifically, taking 10 keyframe input frames as an example... Figure 5As shown, the method disclosed in this application achieves an accuracy of 77.57% on the Real-Time Visual Understanding task under the StreamingBench benchmark and 69.76% on the Real-Time Visual Perception task under the OVO benchmark. Furthermore, compared to complete historical input or other input strategies, this application reduces the scale of redundant visual input data and achieves a better balance between memory usage and inference latency.

[0097] like Figure 6 As shown, in experiments with different sampling frame numbers, when the number of keyframes is set to 5, 10, 15, and 20 frames, the method disclosed in this application exhibits different accuracy and latency characteristics. Among them, 10 keyframes can achieve a better balance between accuracy and inference efficiency, thus demonstrating that the keyframe sampling mechanism can retain key visual information that is effective for the task with a limited number of frames.

[0098] Figure 7 The impact of the number of sampled frames on the accuracy and latency of StreamingBench is shown. These experimental results demonstrate that the method disclosed in this application does not simply reduce the number of input frames, but rather selects higher-value video frames through multi-metric fusion and redundancy control, and combines ChunkPrefill to improve the long input inference process, thereby improving inference efficiency while ensuring video understanding performance.

[0099] Correspondingly, this application also discloses a multimodal reasoning system. For example... Figure 8 As shown, the system includes an input module, a historical context management module, a candidate frame extraction module, a multi-index feature calculation module, a HybridSampler (multi-index fusion keyframe sampler) keyframe sampling module, a multimodal input organization module, a ChunkPrefill module, a multimodal large model inference module, and an answer output and log recording module, all connected in sequence.

[0100] The system comprises the following modules: an input module for receiving streaming video and text questions; a history context management module for identifying the target streaming video from historical streaming video; a candidate frame extraction module for identifying a set of candidate frames from the target streaming video based on preset sampling conditions; a multi-metric feature calculation module for calculating the multimodal features and correlation data corresponding to at least one initial frame; a Hybrid Sampler (multi-metric fusion keyframe sampler) keyframe sampling module for identifying keyframes based on multimodal features and correlation data, and generating visual and text tokens; a multimodal input organization module for combining visual and text tokens into a unified input sequence; a Chunk Prefill module for performing chunk prefilling on the unified input sequence to obtain cache blocks; a multimodal large model inference module for executing video inference tasks; and an answer output and log recording module for outputting video inference results and generating runtime log data.

[0101] Correspondingly, this application also discloses a multimodal inference device for streaming video. For example... Figure 9 As shown, the device includes: The information acquisition module 910 is used to acquire information to be processed; the information to be processed includes target text information and target streaming video corresponding to the target text information. The set determination module 920 is used to determine a candidate frame set from the target streaming video based on preset sampling conditions; the candidate frame set includes at least one initial frame; Calculation module 930 is used to calculate the multimodal features and correlation data corresponding to at least one initial frame; The keyframe determination module 940 is used to determine at least one keyframe from at least one initial frame based on multimodal features and correlation data. The block pre-filling processing module 950 is used to perform block pre-filling processing based on at least one keyframe and target text information to obtain multiple consecutive cache blocks. The decoding processing module 960 is used to decode multiple buffer blocks to obtain video inference results corresponding to the information to be processed.

[0102] In some embodiments, the computing module 930 includes: The feature calculation module is used to calculate the visual and semantic features corresponding to each initial frame. The semantic encoding module is used to perform semantic encoding on the target text information to obtain a text semantic vector; The correlation determination module is used to determine the correlation data corresponding to each initial frame based on semantic features and text semantic vectors.

[0103] In some embodiments, the device further includes: The aggregation processing module is used to perform semantic aggregation processing on the semantic features corresponding to each initial frame to obtain the overall semantic vector corresponding to the candidate frame set. The similarity determination module is used to determine semantic similarity based on semantic features and the overall semantic vector.

[0104] In some embodiments, the keyframe determination module 940 includes: The set construction module is used to construct a fusion candidate set based on the visual and semantic features corresponding to each initial frame; the fusion candidate set includes at least one candidate frame; The scoring calculation module is used to calculate the comprehensive score data corresponding to at least one candidate frame. The filtering module is used to filter at least one candidate frame based on the comprehensive scoring data to obtain a set of key frames; the set of key frames includes at least one key frame.

[0105] In some embodiments, the scoring calculation module includes: The parameter calculation module is used to calculate the redundant parameters between candidate frames and keyframes; The weighted processing module is used to weight semantic similarity, relevance data, and redundant parameters to obtain comprehensive score data.

[0106] In some embodiments, the filtering module includes: The first keyframe determination module is used to determine keyframes from at least one candidate frame based on the comprehensive score data corresponding to at least one candidate frame. The first update module is used to update the fusion candidate set and the keyframe set; The second update module is used to update redundant parameters and comprehensive score data based on the updated fusion candidate set and the updated keyframe set; The second keyframe determination module is used to determine the next keyframe based on the updated redundancy parameters and the updated comprehensive score data, until the number of keyframes in the keyframe set equals the preset number of frames.

[0107] In some embodiments, the device further includes: The sorting module is used to sort at least one keyframe based on the video timestamps corresponding to at least one keyframe, so as to obtain a set of sequential keyframes.

[0108] In some embodiments, the block pre-filling processing module 950 includes: The visual encoding processing module is used to perform visual encoding processing on the keyframes included in the sequential keyframe set to obtain the first encoded information. The word segmentation module is used to segment the target text information into words to obtain the second encoded information; A sequence construction module is used to construct a unified input sequence based on the first encoding information and the second encoding information; The cache block generation module is used to perform block pre-filling processing on the unified input sequence to obtain cache blocks.

[0109] In some embodiments, the cache block generation module further includes: The block segmentation module is used to segment a uniform input sequence into blocks based on a preset unit length, resulting in multiple data blocks. The pre-fill module is used to pre-fill multiple data blocks separately to obtain cache blocks.

[0110] In some embodiments, the device further includes: The log generation module is used to generate runtime log data corresponding to the video inference results.

[0111] The apparatus and method embodiments described above are based on the same inventive concept and are used to implement the above-described multimodal reasoning method applied to streaming video.

[0112] Correspondingly, this application also discloses an electronic device, which includes a processor and a memory. The memory stores at least one instruction and at least one program. The at least one instruction and at least one program are loaded and executed by the processor to implement the multimodal reasoning method for streaming video as described in any of the above embodiments.

[0113] Correspondingly, this application also discloses a computer storage medium storing at least one instruction and at least one program, wherein the at least one instruction and at least one program are loaded and executed by a processor to implement the multimodal reasoning method for streaming video as described in any of the above embodiments.

[0114] Correspondingly, this application also discloses a computer program product, including a computer program / instructions, which, when executed by a processor, implement the multimodal reasoning method for streaming video as described in any of the above embodiments.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A multimodal inference method applied to streaming video, characterized in that, include: Obtain information to be processed; The information to be processed includes target text information and target streaming video corresponding to the target text information; Based on preset sampling conditions, a candidate frame set is determined from the target streaming video; the candidate frame set includes at least one initial frame; Calculate the multimodal features and correlation data corresponding to at least one of the initial frames; Based on the multimodal features and the correlation data, at least one key frame is determined from at least one initial frame; Based on at least one of the keyframes and the target text information, block pre-filling processing is performed to obtain multiple consecutive cache blocks; The multiple cache blocks are decoded to obtain the video inference result corresponding to the information to be processed.

2. The multimodal inference method for streaming video according to claim 1, characterized in that, The multimodal features include visual features and semantic features; The calculation of at least one of the multimodal features and correlation data corresponding to the initial frame includes: Calculate the visual features and semantic features corresponding to each of the initial frames; The target text information is semantically encoded to obtain a text semantic vector; Based on the semantic features and the text semantic vector, the relevance data corresponding to each of the initial frames is determined.

3. The multimodal inference method for streaming video according to claim 1, characterized in that, The multimodal features include visual features and semantic features; After calculating the multimodal features and correlation data corresponding to at least one of the initial frames, the method further includes: The semantic features corresponding to each initial frame are subjected to semantic aggregation processing to obtain the overall semantic vector corresponding to the candidate frame set; Based on the semantic features and the overall semantic vector, semantic similarity is determined.

4. The multimodal inference method for streaming video according to claim 3, characterized in that, The step of determining at least one key frame from at least one initial frame based on the multimodal features and the correlation data includes: Based on the visual features and semantic features corresponding to each initial frame, a fusion candidate set is constructed; the fusion candidate set includes at least one candidate frame; Calculate the comprehensive score data corresponding to at least one of the candidate frames; Based on the comprehensive scoring data, at least one of the candidate frames is filtered to obtain a key frame set; the key frame set includes at least one key frame.

5. The multimodal inference method for streaming video according to claim 4, characterized in that, In the initial state, the set of keyframes is an empty set; The calculation of the comprehensive score data corresponding to at least one of the candidate frames includes: Calculate the redundancy parameters between the candidate frame and the key frame; The semantic similarity, the relevance data, and the redundant parameters are weighted to obtain the comprehensive score data.

6. The multimodal inference method for streaming video according to claim 5, characterized in that, The step of filtering at least one candidate frame based on the comprehensive scoring data to obtain a keyframe set includes: The keyframe is determined from at least one of the candidate frames based on the comprehensive score data corresponding to at least one of the candidate frames; Update the fusion candidate set and the keyframe set; Based on the updated fusion candidate set and the updated keyframe set, the redundant parameters and the comprehensive score data are updated; Based on the updated redundancy parameters and the updated comprehensive score data, the next key frame is determined until the number of key frames in the key frame set equals the preset number of frames.

7. The multimodal inference method for streaming video according to claim 1, characterized in that, Each keyframe corresponds to a video timestamp; Before performing block pre-filling processing based on at least one keyframe and the target text information to obtain multiple consecutive cache blocks, the method includes: Based on the video timestamps corresponding to at least one of the keyframes, the at least one keyframe is sorted to obtain a set of sequential keyframes.

8. The multimodal inference method for streaming video according to claim 7, characterized in that, The block pre-filling process based on at least one keyframe and the target text information yields multiple consecutive cache blocks, including: Visual encoding processing is performed on the keyframes included in the sequential keyframe set to obtain first encoded information; The target text information is segmented to obtain the second encoded information; A unified input sequence is constructed based on the first encoding information and the second encoding information; The unified input sequence is divided into blocks and pre-filled to obtain the cache block.

9. The multimodal inference method for streaming video according to claim 8, characterized in that, The step of performing block pre-filling processing on the unified input sequence to obtain the cache block includes: Based on a preset unit length, the unified input sequence is divided into blocks to obtain multiple data blocks; The cache blocks are obtained by pre-filling multiple data blocks respectively.

10. The multimodal inference method for streaming video according to claim 1, characterized in that, The method further includes: Generate runtime log data corresponding to the video inference results.

11. A multimodal inference device for streaming video, characterized in that, include: The information acquisition module is used to acquire information to be processed. The information to be processed includes target text information and target streaming video corresponding to the target text information; A set determination module is used to determine a set of candidate frames from the target streaming video based on preset sampling conditions; the set of candidate frames includes at least one initial frame; The calculation module is used to calculate at least one multimodal feature and correlation data corresponding to the initial frame; A keyframe determination module is used to determine at least one keyframe from at least one initial frame based on the multimodal features and the correlation data. The block pre-filling processing module is used to perform block pre-filling processing based on at least one of the keyframes and the target text information to obtain multiple consecutive cache blocks; The decoding processing module is used to decode multiple cache blocks to obtain video inference results corresponding to the information to be processed.

12. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one instruction and at least one program, the at least one instruction and the at least one program being loaded and executed by the processor to implement the multimodal inference method for streaming video as described in any one of claims 1-10.

13. A computer storage medium, characterized in that, The computer storage medium stores at least one instruction and at least one program, which are loaded and executed by a processor to implement the multimodal reasoning method for streaming video as described in any one of claims 1-10.

14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the multimodal reasoning method for streaming video as described in any one of claims 1-10.