A multi-decision collaborative control method and device for streaming video understanding, an electronic device, and a computer program product

By applying different compression techniques to streaming video frames and dynamically switching models, the issues of causality and real-time performance in streaming video understanding are resolved, achieving a high-performance balance between streaming video understanding and efficiency.

CN122132593APending Publication Date: 2026-06-02NINGBO ORIENTAL UNIVERSITY OF TECHNOLOGY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NINGBO ORIENTAL UNIVERSITY OF TECHNOLOGY
Filing Date
2026-05-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to meet the causal and real-time requirements in streaming video scenarios and cannot effectively address the performance issues of streaming video understanding, especially when optimizing a single stage, leading to a decrease in real-time perception capabilities or excessive computational overhead.

Method used

A multi-decision collaborative control method is adopted, which processes the data of the neighboring interval and the historical interval of the streaming video frame with different compression intensities, and combines the dynamic switching of the fast visual language model and the slow reasoning model to achieve a balance between real-time response and deep reasoning.

Benefits of technology

It significantly reduces the number of visual tokens, maintains the performance and efficiency of streaming video understanding under real-time constraints, dynamically improves the inference accuracy of complex problems, and meets the usage requirements of streaming video scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132593A_ABST
    Figure CN122132593A_ABST
Patent Text Reader

Abstract

This application discloses a multi-decision collaborative control method, apparatus, electronic device, and computer program product for streaming video understanding. The method includes performing visual token compression on neighboring interval data and historical interval data respectively to obtain the current memory state; calculating a readiness score based on the current query to be answered and the current memory state; in response to the readiness score not being lower than a score threshold, inputting the current query to be answered and the current memory state into a fast visual language model to obtain a first answer result; in response to the first answer result being marked as an upgraded answer, inputting the current query to be answered and the current memory state into a slow inference model to output a second answer result. In this application, controlled upgrades between the fast path and the slow inference path can maintain real-time constraints while also considering the inference accuracy of complex problems, dynamically improving the performance and efficiency balance of streaming video understanding, and better meeting the usage requirements of streaming video scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification belong to the field of streaming video data processing, and specifically relate to a multi-decision collaborative control method, apparatus, electronic device, and computer program product for streaming video understanding. Background Technology

[0002] In recent years, multimodal large models have demonstrated strong capabilities in tasks such as video question answering, event understanding, video description, and cross-modal reasoning. However, most methods are geared towards offline whole-segment video processing, which usually requires obtaining the complete video before unified encoding and reasoning, making it difficult to directly meet the causal and real-time requirements of streaming video scenarios.

[0003] In streaming video scenarios, video frames arrive continuously, and the system needs to continuously update the context and respond to the current text query within a limited latency budget. Existing solutions typically optimize only a single step, such as using a fixed memory compression strategy, answering directly with only a single model, or calling a heavy inference model at all times. These static strategies have a fixed processing method and cannot be adjusted according to the actual situation of the streaming information, resulting in poor performance in understanding streaming video in many cases, making it difficult to meet the usage requirements of streaming video scenarios. Summary of the Invention

[0004] Embodiments of this disclosure provide a multi-decision collaborative control method, apparatus, electronic device, and computer program product for streaming video understanding, designed to address one or more of the problems described above and other potential problems.

[0005] According to a first aspect of this disclosure, a multi-decision collaborative control method for streaming video understanding is provided. The method includes acquiring historical video frames of the streaming video within a preset historical duration and a current query to be answered; dividing the historical video frames into neighboring interval data and historical interval data based on the time distance from the current moment, and performing visual token compression with different compression intensities on the neighboring interval data and historical interval data respectively to obtain the current memory state; calculating a readiness score to characterize whether immediate response is supported based on the fusion features of the current query to be answered and the current memory state; in response to the readiness score not being lower than a score threshold, inputting the current query to be answered and the current memory state into a fast visual language model to obtain a first response result; in response to the first response result being marked as a direct response, outputting the first response result; in response to the first response result being marked as an upgraded response, inputting the current query to be answered and the current memory state into a slow inference model to obtain and output a second response result.

[0006] According to a second aspect of this disclosure, a multi-decision collaborative control device for streaming video understanding is provided. The device includes a data acquisition module configured to acquire historical video frames of the streaming video within a preset historical duration and a current query to be answered; a data compression module configured to divide the historical video frames into neighboring interval data and historical interval data based on the time distance from the current time, and to perform visual token compression with different compression intensities on the neighboring interval data and historical interval data respectively to obtain the current memory state; a first processing module configured to calculate a readiness score based on the fusion features of the current query to be answered and the current memory state to characterize whether immediate response is supported, and in response to the readiness score not being lower than a score threshold, input the current query to be answered and the current memory state into a fast visual language model to obtain a first response result; a second processing module configured to output the first response result in response to the first response result being marked as a direct response; and a third processing module configured to input the current query to be answered and the current memory state into a slow inference model in response to the first response result being marked as an escalation response, to obtain and output a second response result.

[0007] According to a third aspect of this disclosure, an electronic device is provided, including one or more processors and a memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform a method provided according to a first scheme.

[0008] According to a fourth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to the first aspect.

[0009] The solution provided in the embodiments of this specification can significantly reduce the number of visual tokens by compressing neighboring interval data and historical interval data with different visual tokens, while maintaining high fidelity of recent evidence and actively compressing old history. When answering queries that have passed the readiness determination, it can maintain real-time constraints while taking into account the reasoning accuracy of complex problems by performing controlled upgrades between fast and slow reasoning paths, thus dynamically improving the performance and efficiency balance of streaming video understanding and better meeting the usage needs of streaming video scenarios. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 A flowchart illustrating a multi-decision collaborative control method for streaming video understanding, according to some embodiments of this disclosure, is shown.

[0012] Figure 2 A schematic diagram illustrating the principle of a multi-decision collaborative control method for streaming video understanding, according to some embodiments of this disclosure, is shown.

[0013] Figure 3 A schematic diagram illustrating the principle of the target balance reinforcement learning training process of a fast visual language model according to some embodiments of the present disclosure is shown.

[0014] Figure 4 A schematic diagram of the structure of a multi-decision collaborative control device for streaming video understanding, according to some embodiments of the present disclosure, is shown.

[0015] Figure 5 A schematic block diagram of an electronic device according to some embodiments of the present disclosure is shown. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0017] The terms “comprising” and “having”, and any variations thereof, in this specification, claims, and the foregoing drawings are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. Depending on the context, the word “if” as it applies herein may be interpreted as “when”, “in response to determination”, or “in response to detection”.

[0018] As mentioned earlier, existing streaming video understanding solutions typically optimize only a single aspect. This has two drawbacks: firstly, applying uniform compression intensity to all historical frames can easily compress the most recent evidence strongly correlated with the current query, leading to a decline in real-time perception and tracking capabilities; secondly, consistently retaining high-fidelity historical information results in an excessive number of visual tokens, excessive memory consumption, and increased inference latency. Furthermore, relying solely on fast models can produce unreliable answers in complex queries or when evidence is insufficient, while constantly calling slow inference models incurs unnecessary computational overhead and may even underperform smaller, faster models in some real-time tasks. Therefore, a streaming video understanding approach that can jointly coordinate memory retention, response timing, and inference depth is urgently needed.

[0019] Figure 1 A flowchart illustrating a multi-decision collaborative control method 100 for streaming video understanding, based on some embodiments of this disclosure, is shown. Method 100 can be executed by a terminal, which may include, but is not limited to, mobile phones, tablets, desktop computers, servers, etc. Figure 1 As shown, in method 100, step 102 can obtain historical video frames of the streaming video within a preset historical duration and the current query to be answered.

[0020] In this embodiment, the current video frame corresponding to the current moment in the streaming video will be determined to obtain all historical video frames within a preset historical time period, as well as the current query to be answered at the current moment. The current query to be answered is a question, task instruction, or event query statement input into the system at the current moment, such as "What action is the person in the current frame performing?" or "Does the specified target appear in the video?", which can be represented in the form of text vectors.

[0021] In method 100, step 104 can divide historical video frames into neighboring interval data and historical interval data based on the time distance from the current time, and perform visual token compression with different compression intensities on the neighboring interval data and historical interval data respectively to obtain the current memory state.

[0022] In this embodiment, as Figure 2 As shown, for historical video frames within a preset historical duration, they will be further divided into adjacent interval data and historical interval data based on their time distance from the current moment. Adjacent interval data consists of historical video frames that are closer to the current moment, while historical interval data consists of historical video frames that are farther from the current moment. Specifically, the division can be based on a specified time distance W selected within the preset historical duration. That is, assuming the historical time span corresponding to the preset historical duration is H, then adjacent interval data consists of historical video frames from the current moment to W, and historical interval data consists of historical video frames from W to H. The values ​​of W and H can be adjusted and set according to actual needs.

[0023] Next, by applying visual token compression with varying intensities to the neighboring and historical interval data, the neighboring interval data can be used to retain the most recent evidence that has the greatest impact on the current query to be answered, while the historical interval data can be used to retain supplementary information related to cross-time periods. Visual token compression is the process of encoding the neighboring and historical interval data (e.g., by a visual encoder ViT) to obtain a visual feature sequence, followed by average pooling, max pooling, or attention pooling to transform the massive visual feature sequence into a shorter sequence. Since the neighboring interval data has a greater impact on the current query to be answered, the compression intensity of the neighboring interval data is generally greater than that of the historical interval data to maintain higher fidelity for the neighboring interval data and avoid over-compression of the neighboring interval data, which could affect the accuracy of the subsequent results. After visual token compression, the current memory state can be obtained, which can be represented by any of the following forms: a visual token sequence, a visual feature matrix, or a vector obtained by pooling the visual token sequence.

[0024] In method 100, step 106 can calculate a readiness score based on the fusion features of the current query to be answered and the current memory state to characterize whether an immediate answer is supported. In response to the readiness score not being lower than the score threshold, the current query to be answered and the current memory state are input into the fast visual language model to obtain the first answer result.

[0025] In this embodiment, before processing using the corresponding language model, a fusion feature of the current query to be answered and the current memory state is first generated to calculate a readiness score. The readiness score determines whether the current state can support an immediate response. The readiness score can be used to represent a preliminary judgment of the result hit rate. Only when the readiness score is not lower than a preset score threshold is it considered that the false response rate is not high, and only then will it be input into the corresponding language model for processing.

[0026] As an example, the formula for calculating fused features can be:

[0027]

[0028] in, As a feature of fusion, Represents the pooling function. Indicates the joint encoding function, This is the current query that needs an answer. This represents the current memory state.

[0029] The formula for calculating the readiness score is:

[0030]

[0031] in, For readiness score, For the sigmoid function, The weight parameters for the readiness header, The offset parameter for the readiness header. and It can be learned in advance from labeled training samples.

[0032] When the readiness score is not lower than the score threshold, the current time is considered to be ready to answer. The current query to be answered and the current memory state can be input into the fast visual language model to obtain the first answer result output by the fast visual language model. The first answer result can include the generated answer text and the tags set for the answer text. Among them, the fast visual language model can be Qwen2.5-VL-3B, for example.

[0033] In method 100, step 108 may output the first answer result in response to the first answer result being marked as a direct answer.

[0034] In this embodiment, if the first answer result is marked as a direct answer, it is considered that the answer at this time has sufficient accuracy and can be solved by the fast path without additional processing. The first answer result will be directly output to answer the current query.

[0035] In method 100, step 110 may respond to the first answer result being marked as an upgraded answer by inputting the current query to be answered and the current memory state into the slow inference model to obtain and output the second answer result.

[0036] In this embodiment, if the first answer is marked as an upgraded answer, it is considered that the current query to be answered is difficult or requires deeper reasoning to output a more accurate result. In this case, the first answer will not be used as the final output. Instead, the current query to be answered and the current memory state will be re-input into the slow inference model to obtain the second answer generated by the slow inference model, and the second answer will be used as the final answer for output. The slow inference model can be, for example, Qwen3-VL-48-Thinking or Qwen2.5-VL-32B. This allows for unified processing by the fast visual language model after the readiness score is met, and the slow inference model is only called when the output of the fast visual language model is marked as an upgraded answer. This avoids the computational waste caused by forcibly using a heavy model at all times, achieving controlled upgrades of the fast and slow inference paths. This maintains real-time constraints while considering the inference accuracy of complex problems, dynamically improving the performance and efficiency balance for streaming video understanding.

[0037] In one possible implementation, the method further includes:

[0038] In response to a readiness score falling below a threshold, the currently pending query is retained to postpone its evaluation.

[0039] In this embodiment, if the calculated readiness score is lower than a threshold, it is considered that the conditions for answering are not met, and the current query to be answered is retained. For example, a waiting flag can be set for the current query to be answered, waiting for a later time to be judged again. In this way, when it is determined that the conditions for answering are not met, an answer will not be forcibly generated, thereby reducing the false answer rate when there is insufficient evidence. The specific timing of the delayed judgment can be set according to actual needs, such as delaying it to the next or several times, or delaying it to a specified time.

[0040] In one possible implementation, based on the time distance from the current moment, historical video frames are divided into neighboring interval data and historical interval data, and visual token compression with different compression intensities is applied to the neighboring interval data and historical interval data respectively to obtain the current memory state, including:

[0041] Historical video frames whose time distance from the current time is not greater than the distance threshold are divided into neighboring interval data, and historical video frames whose time distance from the current time is greater than the distance threshold are divided into historical interval data.

[0042] Visual token compression with different compression intensities is applied to neighboring interval data and historical interval data respectively to obtain neighboring interval tokens and historical interval tokens. The compression threshold of neighboring interval data is greater than that of historical interval data. Visual token compression is a compression operation that is irrelevant to training.

[0043] By combining or merging neighboring zone tokens and historical zone tokens, the current memory state can be obtained.

[0044] In this embodiment, a distance threshold can be preset to classify historical video frames whose temporal distance is not greater than the threshold into neighboring interval data, and vice versa. During visual token compression, the compression threshold for neighboring interval data must be greater than that for historical interval data to avoid over-compression. Furthermore, visual token compression will employ a compression operation independent of training; that is, no new trainable parameters will be added to the compression module during deployment. Instead, preset rules, pre-selected compression operators, or existing encoding results will be directly invoked to complete the compression. Selectable compression operations include one or more of dynamic token deletion, average pooling, max pooling, cluster merging, or stride convolution. Active forgetting of historical video frames can be performed without additional training costs, thus the compression process does not rely on end-to-end retraining and can be more easily embedded into streaming video understanding systems. After visual token compression is completed, obtaining neighboring interval tokens and historical interval tokens, they can be merged by direct concatenation or weighted fusion to construct the current memory state.

[0045] Furthermore, by appropriately increasing the compression threshold of historical interval data, more long-term dependency information can be retained while continuing to reuse the waiting judgment and upgrade inference mechanisms, thereby improving the processing effect of medium- and long-term video tasks.

[0046] In one possible implementation, a readiness score, characterizing whether immediate response is supported, is calculated based on a fusion feature of the current query to be answered and the current memory state, including:

[0047] Based on the fusion features of the current query to be answered and the current memory state, a readiness score is calculated by the readiness decision head set on the fast visual language model to characterize whether an immediate response is supported.

[0048] In this embodiment, the readiness score can be calculated using a lightweight readiness determination head set on the fast visual language model.

[0049] In one possible implementation, the readiness decision head freezes the fast visual language model during training to update the parameters of the readiness decision head only based on training samples with waiting and answerable labels.

[0050] In this embodiment, after the fast visual language model completes its pre-training, it will temporarily freeze the fast visual language model and only use training samples with waiting labels (representing that the readiness score is lower than the score threshold) and answerable labels (representing that the readiness score is not lower than the score threshold) to supervise the training of the parameters of the readiness decision head. This achieves waiting decision with a low number of additional parameters while maintaining the existing routing capabilities of the fast visual language model.

[0051] In one possible implementation, the method further includes:

[0052] Construct a training set, which includes training query data and training response results corresponding to the training query data. The training response results are marked as direct responses when the average quality score of the presampled candidate responses of the training query data is not lower than the quality threshold, and as upgraded responses when the average quality score of the presampled candidate responses of the training query data is lower than the quality threshold.

[0053] Based on the training query data, the initial fast visual language model generates predicted answer results;

[0054] Using the training response results as a supervision signal, the initial fast visual language model is trained for at least one round to obtain a trained fast visual language model.

[0055] In this embodiment, the fast visual language model undergoes cold-start training before formal use. Several training query data sets are pre-set, including unanswered queries and memory states, and input into the initial fast visual language model for pre-sampling to obtain at least two candidate answers. Next, the candidate answers are scored. For open-ended questions, an external language model can be used to score the candidate answers, while for objective questions, the correctness of the answer can be judged based on the standard answer. The average quality score of the training query data is obtained by averaging the scores of all candidate answers for the same training query data. If the average quality score is not lower than a preset quality threshold, the result is considered highly reliable, and the corresponding training answer is marked as a direct answer; otherwise, it is marked as an upgraded answer. The fast visual language model is then trained under supervision using a training set consisting of the training query data and the training answer results. The training answer results may include at least one pre-sampled candidate answer.

[0056] In one possible implementation, the method further includes:

[0057] Determine the actual proportion of all predicted responses generated by the fast visual language model that are marked as upgrade responses. Adjust the upgrade path reward and non-upgrade path penalty based on the comparison between the actual proportion and the preset target proportion, so that the actual proportion remains within the preset target range.

[0058] In this embodiment, as Figure 3 As shown, after completing the cold-start supervised fine-tuning of the fast visual language model, it can be further optimized using objective-balanced reinforcement learning. Specifically, let the actual proportion of predicted responses marked as upgrade responses in the training sample group be... The preset target ratio is The upgrade path rewards can be adjusted by comparing the results of the two. For example, when Higher than When, reduce upgrade path rewards to curb excessive upgrades; when Below When, a penalty is imposed on non-promotion paths to prevent the model from almost never promoting; when While within the target range, a neutral reward is maintained. This mechanism allows the upgrade rate to be constrained within the deployable computing budget. Furthermore, a bandwidth tolerance mechanism can be introduced. , target ratio With tolerance bandwidth Combine and then compare with the actual ratio Compare them.

[0059] Figure 4 This document illustrates a structural schematic diagram of a multi-decision collaborative control device 400 for streaming video understanding, based on some embodiments of this disclosure. The various embodiments in this specification are described in a progressive manner; similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing its differences from other embodiments. In particular, the device embodiments are substantially similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments. Figure 4 As shown, the device 400 includes a data acquisition module 401, which is configured to acquire historical video frames of the streaming video within a preset historical duration and the current query to be answered;

[0060] The data compression module 402 is configured to divide historical video frames into neighboring interval data and historical interval data based on the time distance from the current time, and to perform visual token compression with different compression intensities on the neighboring interval data and historical interval data respectively to obtain the current memory state.

[0061] The first processing module 403 is configured to calculate a readiness score based on the fusion features of the current query to be answered and the current memory state to characterize whether an immediate answer is supported. In response to the readiness score not being lower than the score threshold, the current query to be answered and the current memory state are input into the fast visual language model to obtain the first answer result.

[0062] The second processing module 404 is configured to respond to the first answer result being marked as a direct answer and output the first answer result.

[0063] The third processing module 405 is configured to respond to the first answer result by marking it as an upgraded answer, inputting the current query to be answered and the current memory state into the slow reasoning model, and obtaining and outputting the second answer result.

[0064] In one possible implementation, the device further includes a fourth processing module configured to retain the currently pending query in response to a readiness score falling below a score threshold, thereby delaying the determination of the currently pending query.

[0065] In one possible implementation, the data compression module 402 is further configured to divide historical video frames whose time distance from the current moment is not greater than a distance threshold into neighboring interval data, and historical video frames whose time distance from the current moment is greater than a distance threshold into historical interval data; perform visual token compression with different compression intensities on the neighboring interval data and the historical interval data respectively to obtain neighboring region tokens and historical region tokens, wherein the compression threshold of the neighboring interval data is greater than the compression threshold of the historical interval data, and the visual token compression is a training-independent compression operation; and concatenate or fuse the neighboring region tokens and the historical region tokens to obtain the current memory state.

[0066] In one possible implementation, the first processing module 403 is further configured to calculate a readiness score, based on the fusion features of the current query to be answered and the current memory state, by a readiness determination head set on a fast visual language model, to characterize whether an immediate response is supported.

[0067] In one possible implementation, the readiness decision head freezes the fast visual language model during training to update the parameters of the readiness decision head only based on training samples with waiting and answerable labels.

[0068] In one possible implementation, the apparatus further includes a training module configured to construct a training set, which includes training query data and corresponding training response results. The training response results are marked as direct responses when the average quality score of the presampled candidate responses to the training query data is not lower than a quality threshold, and as upgraded responses when the average quality score of the presampled candidate responses to the training query data is lower than the quality threshold. Based on the training query data, a predicted response result is generated from an initial fast visual language model. Using the training response result as a supervision signal, the initial fast visual language model is trained for at least one round to obtain a trained fast visual language model.

[0069] In one possible implementation, the training module is further configured to determine the actual proportion of all predicted responses generated by the fast visual language model that are marked as upgrade responses, and adjust the upgrade path reward and non-upgrade path penalty based on the comparison between the actual proportion and the preset target proportion, so that the actual proportion remains within the preset target range.

[0070] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this specification is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).

[0071] Figure 5 A block diagram of an electronic device 500 that can implement various embodiments of the present disclosure is shown. For example... Figure 5 As shown, the electronic device 500 includes a processor 510, a disk drive 520, an input / output interface 530, a network interface 540, and a memory 550. The processor 510, disk drive 520, input / output interface 530, network interface 540, and memory 550 can communicate with each other via a communication bus 560.

[0072] The processor 510 can be implemented using a general-purpose CPU, microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits to execute relevant programs and implement the technical solution provided in this application.

[0073] The memory 550 can be implemented in the form of ROM (Read Only Memory), RAM (Read Access Memory), static memory, dynamic storage devices, etc. The memory 550 can store the operating system 551 used to control the operation of the electronic device 500, and the basic input / output system (BIOS) 552 used to control the low-level operations of the electronic device 500. Additionally, it can store a web browser 553, a data storage management system 554, etc. In summary, when implementing the technical solution provided in this application through software or firmware, the relevant program code is stored in the memory 550 and is called and executed by the processor 510.

[0074] Input / output interface 530 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, vibrators, indicator lights, etc.

[0075] Network interface 540 is used to connect a communication module (not shown in the figure) to enable communication and interaction between the device and other devices. The communication module can communicate via wired means (e.g., USB, Ethernet cable) or wireless means (e.g., mobile network, Wi-Fi, Bluetooth).

[0076] Bus 560 includes a pathway for transmitting information between various components of the device, such as processor 510, disk drive 520, input / output interface 530, network interface 540, and memory 550.

[0077] It should be noted that although the above-described device only shows the processor 510, disk drive 520, input / output interface 530, network interface 540, memory 550, bus 560, etc., in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the method of this application, and does not necessarily include all the components shown in the figures.

[0078] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0079] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. Furthermore, although operations are depicted in a specific order, this should be understood as requiring that such operations be performed in the specific order shown or in sequential order, or requiring that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.

[0080] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A multi-decision collaborative control method for streaming video understanding, characterized in that, The method includes: Retrieve historical video frames within a preset historical duration and the current unanswered query; Based on the time distance from the current moment, the historical video frames are divided into neighboring interval data and historical interval data, and visual token compression with different compression intensities is applied to the neighboring interval data and historical interval data respectively to obtain the current memory state. Based on the fusion features of the current query to be answered and the current memory state, a readiness score is calculated to characterize whether an immediate response is supported. In response to the readiness score not being lower than the score threshold, the current query to be answered and the current memory state are input into the fast visual language model to obtain the first response result. In response to the first answer result being marked as a direct answer, the first answer result is output; In response to the first answer result being marked as an upgraded answer, the current unanswered query and the current memory state are input into the slow reasoning model to obtain and output the second answer result.

2. The multi-decision collaborative control method for streaming video understanding according to claim 1, characterized in that, The method further includes: In response to the readiness score being lower than a score threshold, the currently pending query is retained to postpone the determination of the currently pending query.

3. The multi-decision collaborative control method for streaming video understanding according to claim 1, characterized in that, Based on the time distance from the current moment, the historical video frames are divided into neighboring interval data and historical interval data, and visual token compression with different compression intensities is applied to the neighboring interval data and historical interval data respectively to obtain the current memory state, including: Historical video frames whose time distance from the current time is no greater than the distance threshold are divided into neighboring interval data, and historical video frames whose time distance from the current time is greater than the distance threshold are divided into historical interval data. Visual token compression with different compression intensities is performed on the neighboring interval data and the historical interval data respectively to obtain neighboring interval tokens and historical interval tokens. The compression threshold of the neighboring interval data is greater than the compression threshold of the historical interval data. The visual token compression is a training-independent compression operation. The neighboring area token and the historical area token are concatenated or merged to obtain the current memory state.

4. The multi-decision collaborative control method for streaming video understanding according to claim 1, characterized in that, The readiness score, calculated based on the fusion features of the current query to be answered and the current memory state to characterize whether immediate response is supported, includes: Based on the fusion features of the current query to be answered and the current memory state, a readiness score is calculated by the readiness determination head set on the fast visual language model to characterize whether an immediate answer is supported.

5. A multi-decision collaborative control method for streaming video understanding according to claim 4, characterized in that, The readiness determination head freezes the fast visual language model during training so that the parameters of the readiness determination head are updated only based on training samples with waiting and answerable labels.

6. The multi-decision collaborative control method for streaming video understanding according to claim 1, characterized in that, The method further includes: Construct a training set, which includes training query data and training response results corresponding to the training query data. The training response results are marked as direct responses when the average quality score of the presampled candidate responses of the training query data is not lower than a quality threshold, and as upgraded responses when the average quality score of the presampled candidate responses of the training query data is lower than a quality threshold. Based on the training query data, a predicted answer is generated by the initial fast visual language model; Using the training response results as a supervision signal, the initial fast visual language model is trained for at least one round to obtain a trained fast visual language model.

7. A multi-decision collaborative control method for streaming video understanding according to claim 6, characterized in that, The method further includes: Determine the actual proportion of all predicted responses generated by the fast visual language model that are marked as upgrade responses. Adjust the upgrade path reward and non-upgrade path penalty based on the comparison between the actual proportion and the preset target proportion, so that the actual proportion remains within the preset target range.

8. A multi-decision collaborative control device for streaming video understanding, characterized in that, The device includes: The data acquisition module is configured to acquire historical video frames within a preset historical duration and the current query to be answered. The data compression module is configured to divide the historical video frames into neighboring interval data and historical interval data based on the time distance from the current time, and to perform visual token compression with different compression intensities on the neighboring interval data and historical interval data respectively to obtain the current memory state. The first processing module is configured to calculate a readiness score based on the fusion features of the current query to be answered and the current memory state to characterize whether an immediate response is supported. In response to the readiness score not being lower than a score threshold, the current query to be answered and the current memory state are input into a fast visual language model to obtain a first response result. The second processing module is configured to output the first answer result in response to the first answer result being marked as a direct answer. The third processing module is configured to respond to the first answer result being marked as an upgraded answer by inputting the current unanswered query and the current memory state into the slow reasoning model, and then obtain and output the second answer result.

9. An electronic device, characterized in that, include: One or more processors, and a memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of a multi-decision collaborative control method for streaming video understanding as described in any one of claims 1-7.

10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements a multi-decision collaborative control method for streaming video understanding according to any one of claims 1-7.