Evaluation Method, Device, Equipment and Medium for Streaming Video Understanding Model of Agent

By receiving the model evaluation benchmark, the evaluation video data is obtained and the input data is generated, the video stream input is simulated to the stream video understanding model, and the evaluation output data is generated, which solves the problem of insufficient evaluation of the stream video understanding model in the prior art, and realizes comprehensive and accurate evaluation and ability improvement of the stream video understanding model.

CN119741566BActive Publication Date: 2025-07-25BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510246878.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-25
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The prior art lacks systematic evaluation methods for convective video understanding models, and it is impossible to accurately evaluate its multi-step reasoning ability, perceive and predict future action ability, and interactivity ability.

Method used

It provides an evaluation method for the stream video understanding model of an agent. It obtains evaluation video data by receiving the model evaluation benchmark, segments and generates evaluation input data, simulates the video stream input to the stream video understanding model, generates evaluation output data, and obtains model evaluation data based on the model evaluation benchmark, and follows the data for training with preset instructions.

Benefits of technology

It realizes a comprehensive and accurate evaluation of the convective video understanding model, improves its active reasoning and interaction capabilities, and provides a more versatile evaluation standard.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741566B_ABST
    Figure CN119741566B_ABST
Patent Text Reader

Abstract

An evaluation method for a streaming video understanding model of an agent provided by an embodiment of the present invention includes, which can be applied to the field of artificial intelligence technology. The evaluation method for the streaming video understanding model of the agent includes: obtaining evaluation video data according to the received model evaluation benchmark; performing segmentation on the evaluation video data to generate evaluation input data; generating evaluation output data of the streaming video understanding model through the evaluation input data; and obtaining model evaluation data of the streaming video understanding model according to the evaluation output data and the model evaluation benchmark. An embodiment of the present invention also provides an evaluation device, a device, a storage medium, and a program product for the streaming video understanding model of the agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of image processing technology, and more specifically to an evaluation method, device, equipment, medium and product for a streaming video understanding model of an intelligent agent. Background Art

[0002] Artificial Intelligence (AI for short) is an important driving force for the new round of scientific and technological revolution and industrial transformation. It is a new key technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. As an important part of intelligent science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine (i.e., intelligent agent) that can react in a way similar to human intelligence.

[0003] In recent years, to improve the interaction experience between intelligent agents and humans, OpenAI has launched a chatbot model called ChatGPT (Chat Generative Pre-trained Transformer). It can generate answers based on the patterns and statistical laws seen in the pre-training stage, and can also interact with humans according to the context of the chat, achieving natural and fluent communication with humans, providing intelligent question answering and text generation capabilities for various scenarios. Its powerful natural language processing ability and multi-modal conversion ability make it applicable to multiple scenarios and fields. Subsequently, natural language processing (NPL) models developed based on ChatGPT began to focus on the interaction ability of video understanding models. However, there are still relatively few evaluation methods for the interaction ability of such video understanding models. For example, there is no conventional quantitative evaluation method for the evaluation of streaming video understanding models. Usually, only subjective judgments can be made based on the targeted tests of testers, and it is still not possible to accurately achieve the efficient evaluation of streaming video understanding models. Summary of the Invention

[0004] In view of at least one of the above problems, the embodiments of the present invention aim to provide an evaluation method, device, equipment, medium and product for a streaming video understanding model of an intelligent agent that can accurately and efficiently evaluate the data processing ability of the streaming video understanding model.

[0005] One aspect of an embodiment of the present invention provides a method for evaluating a streaming video understanding model of an agent, which includes: obtaining evaluation video data according to a received model evaluation benchmark; performing segmentation on the evaluation video data to generate evaluation input data; generating evaluation output data of the streaming video understanding model through the evaluation input data; and obtaining model evaluation data of the streaming video understanding model according to the evaluation output data and the model evaluation benchmark; wherein the model evaluation benchmark is evaluation standard data generated according to a streaming video understanding task and a preset question-and-answer data set required for the evaluation of the streaming video understanding model, and the preset question-and-answer data set includes a matching data set between questions and standard answers corresponding to different streaming video understanding tasks.

[0006] According to an embodiment of the present invention, in obtaining evaluation video data according to a received model evaluation benchmark, it includes: generating a model evaluation benchmark according to a preset evaluation understanding task; obtaining evaluation video data corresponding to the preset evaluation understanding task.

[0007] According to an embodiment of the present invention, in performing segmentation on the evaluation video data to generate evaluation input data, it includes: obtaining preset task questions in the preset evaluation understanding task; performing segmentation on the evaluation video data according to the time nodes corresponding to the preset task questions to generate evaluation input data.

[0008] According to an embodiment of the present invention, in generating evaluation output data of the streaming video understanding model through the evaluation input data, it includes: inputting the evaluation input data into the streaming video understanding model in a form simulating video stream input to generate evaluation output data.

[0009] According to an embodiment of the present invention, in obtaining model evaluation data of the streaming video understanding model according to the evaluation output data and the model evaluation benchmark, it includes: obtaining evaluation task answers corresponding to the preset task questions in the evaluation output data; generating model evaluation data according to the evaluation task answers and the evaluation benchmark answers in the model evaluation benchmark.

[0010] According to an embodiment of the present invention, the method for evaluating the streaming video understanding model of the agent further includes: training the streaming video understanding model according to the model evaluation data and preset instruction-following data, wherein the preset instruction-following data includes active inference data corresponding to interleaved text and image instructions, noise instructions, and interruption instructions.

[0011] Another aspect of an embodiment of the present invention provides an evaluation device for a streaming video understanding model of an agent, which includes a data acquisition module, a data segmentation module, an evaluation output module, and an evaluation acquisition module. The data acquisition module is configured to acquire evaluation video data according to a received model evaluation benchmark; the data segmentation module is configured to segment the evaluation video data to generate evaluation input data; the evaluation output module is configured to generate evaluation output data of the streaming video understanding model through the evaluation input data; and the evaluation acquisition module is configured to acquire model evaluation data of the streaming video understanding model according to the evaluation output data and the model evaluation benchmark.

[0012] Another aspect of an embodiment of the present invention provides an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned evaluation method of the streaming video understanding model of the agent.

[0013] Another aspect of an embodiment of the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned evaluation method of the streaming video understanding model of the agent.

[0014] Another aspect of an embodiment of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned evaluation method of the streaming video understanding model of the agent is implemented.

[0015] The evaluation method of the streaming video understanding model of the agent provided by the embodiment of the present invention can at least partially solve the problem of the lack of evaluation means for the streaming video understanding model in the related art, and thus can at least achieve one of the following technical effects:

[0016] Compared with the inability to systematically and accurately evaluate the streaming video understanding model in the traditional technology, the method of the embodiment of the present invention can establish an evaluation benchmark for the streaming video understanding model through two aspects of streaming video and active inference, and comprehensively evaluate 6 sub-tasks in total for the above two aspects of streaming video understanding and active inference according to the evaluation benchmark, so as to be able to achieve a comprehensive and accurate evaluation of the streaming video understanding model, and enable the streaming video understanding model to have a more general evaluation standard. Further, the interaction ability can be added to the streaming video understanding model at low cost in combination with the fine-tuning effect corresponding to the provided instruction-following dataset, so as to significantly improve the active inference ability of the streaming video understanding model.

[0017] It should be understood that the above general description and the following specific embodiments are only exemplary and explanatory, and cannot limit the scope claimed by the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0019] Figure 1 Schematically shows an application scenario diagram of an evaluation method, device, equipment, medium, and program product of a streaming video understanding model of an agent according to an embodiment of the present invention;

[0020] Figure 2 Schematically shows a flowchart of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention;

[0021] Figure 3A Schematically shows an application scenario diagram of an action prediction of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention;

[0022] Figure 3B Schematically shows an application scenario diagram of a dynamic state tracking task of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention;

[0023] Figure 3C Schematically shows an application scenario diagram of a multi-round dependent reasoning task of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention;

[0024] Figure 3D Schematically shows an application scenario diagram of an active reminder task of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention;

[0025] Figure 3E Schematically shows an application scenario diagram of a noise recognition task of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention;

[0026] Figure 3F Schematically shows an application scenario diagram of a speaker recognition task of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention;

[0027] Figure 4 Schematically shows another application scenario flowchart of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention;

[0028] Figure 5 Schematically shows a structural block diagram of an evaluation device of a streaming video understanding model of an agent according to an embodiment of the present invention; and

[0029] Figure 6 Schematically shows a block diagram of an electronic device suitable for implementing an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention.

[0030] The accompanying drawings mentioned above are a part of the specification of the embodiments of the present invention, which illustrate exemplary embodiments of the present invention. The accompanying drawings, together with the description in the specification, are used to explain the principles of the embodiments of the present invention. It should be understood that the above general description of the accompanying drawings and the following detailed description are only exemplary and explanatory, and they do not limit the scope of what the present invention intends to claim. Detailed Description of the Embodiments

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the spirit of what the present invention discloses will be clearly described below with reference to the accompanying drawings and in detail. After any person skilled in the art understands the embodiments of the content of the present invention, the techniques taught by the content of the present invention can be changed and modified, which does not deviate from the spirit and scope of the content of the present invention.

[0032] The schematic embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention. In addition, the same or similar reference numerals of elements / components used in the accompanying drawings and embodiments are used to represent the same or similar parts.

[0033] Regarding the use of "first", "second",... etc. in the present invention, it does not particularly refer to the meaning of order or sequence, nor is it used to limit the present invention. It is only used to distinguish elements or operations described with the same technical terms.

[0034] Regarding the directional terms used in the present invention, such as: up, down, left, right, front or back, etc., they are only references to the directions in the accompanying drawings. Therefore, the directional terms used are for explanation and not for limiting this creation.

[0035] Regarding the use of "comprising", "including", "having", "containing", etc. in the present invention, they are all open-ended terms, that is, they mean including but not limited to.

[0036] Regarding the use of "and / or" in the present invention, it includes any one or all combinations of the described things.

[0037] Regarding "a plurality" in the present invention, it includes "two" and "more than two"; regarding "multiple groups" in the present invention, it includes "two groups" and "more than two groups".

[0038] Regarding the terms "substantially", "about", etc. used in the present invention, they are used to modify any quantity or error that can vary slightly, but these slight variations or errors will not change their essence. Generally, the range of such slight variations or errors modified by such terms can be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments, or other values. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.

[0039] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0040] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning that those of ordinary skill in the art usually understand such expressions (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In cases where expressions similar to "at least one of A, B, or C, etc." are used, generally, it should be interpreted according to the meaning that those of ordinary skill in the art usually understand such expressions (for example, "a system having at least one of A, B, or C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). Those of ordinary skill in the art should also understand that substantially any disjunctive conjunction and / or phrase representing two or more alternative items, whether in the specification, claims, or drawings, should be understood as giving the possibility of including one of these items, either of these items, or both items. For example, the phrase "A or B" should be understood as including the possibility of "A" or "B", or "A and B".

[0041] Currently, large language models are developing rapidly, and more and more models are starting to focus on video understanding, especially for real-time video stream understanding, such as GPT-4o. The GPT-4o model is trained based on a large amount of data from the Internet, is better at processing text and audio, accepts any combination of text, audio, and images as input content, and generates any combination of text, audio, and images as output content, and supports 50 languages. It can respond to audio input in as fast as 232 milliseconds, almost reaching the human response level. Therefore, the GPT-4o model has "real-time" interaction, rich emotional expression, stronger visual functions, excellent multilingual performance, and a response speed almost indistinguishable from that of real people, setting a new benchmark in reasoning and audio translation. However, for streaming video understanding, if the GPT-4o model wants to reach an almost human understanding level, it still needs to further improve its multi-step reasoning ability, be able to perceive and predict future actions based on the current state, and at the same time have a certain interaction ability, with the ability to be interrupted and actively generate responses.

[0042] Therefore, for the understanding of streaming video, it is particularly crucial to evaluate the inference and interaction capabilities of the streaming video understanding model. In the prior art, there is no dedicated evaluation method for the streaming video understanding model. Therefore, it is impossible to provide a systematic, accurate, and intuitive feedback on the multi-step inference ability, the ability to perceive and predict future actions, and the interaction ability of the streaming video understanding model.

[0043] In view of at least one of the above technical problems existing in the prior art, embodiments of the present invention aim to provide an evaluation method, device, equipment, medium, and product for a streaming video understanding model of an intelligent agent that can accurately and efficiently evaluate the data processing ability of the streaming video understanding model.

[0044] One aspect of the embodiments of the present invention provides an evaluation method for a streaming video understanding model of an intelligent agent, which includes: obtaining evaluation video data according to the received model evaluation benchmark; performing segmentation on the evaluation video data to generate evaluation input data; generating evaluation output data of the streaming video understanding model through the evaluation input data; and obtaining model evaluation data of the streaming video understanding model according to the evaluation output data and the model evaluation benchmark.

[0045] Figure 1 Schematically shows an application scenario diagram of an evaluation method, device, equipment, medium, and program product for a streaming video understanding model of an intelligent agent according to an embodiment of the present invention.

[0046] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0047] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0048] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc.

[0049] The server 105 can be a server that provides various services, such as a background management server (merely for example) that supports the websites browsed by users using the terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0050] It should be noted that the evaluation method of the streaming video understanding model of the agent provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the evaluation device of the streaming video understanding model of the agent provided by the embodiments of the present invention can generally be set in the server 105. The evaluation method of the streaming video understanding model of the agent provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the evaluation device of the streaming video understanding model of the agent provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.

[0051] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in

[0052] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the Figures 2 to 4 scenario described below, the evaluation method of the streaming video understanding model of the agent of the disclosed embodiments will be described in detail through

[0053] Figure 2 FIG. schematically shows a flowchart of the evaluation method of the streaming video understanding model of the agent according to an embodiment of the present invention.

[0054] As Figure 2 shown, one aspect of the embodiments of the present invention provides an evaluation method of a streaming video understanding model of an agent, which includes operations S201 to S204.

[0055] In operation S201, evaluation video data is obtained according to the received model evaluation criteria;

[0056] In operation S202, the evaluation video data is segmented to generate evaluation input data; and

[0057] In operation S203, evaluation output data of the streaming video understanding model is generated through the evaluation input data.

[0058] In operation S204, model evaluation data of the streaming video understanding model is obtained according to the evaluation output data and the model evaluation benchmark.

[0059] The model evaluation benchmark is evaluation standard data generated according to the streaming video understanding tasks and the preset Q&A dataset required for the evaluation of the streaming video understanding model. The preset Q&A dataset includes a matching dataset between questions and standard answers corresponding to different streaming video understanding tasks.

[0060] The intelligent agent can be the execution subject of the evaluation method of the streaming video understanding model of the above intelligent agent in the embodiments of the present invention, or the execution party controlled by the evaluation method of the streaming video understanding model. Specifically, it can be a humanoid intelligent robot or other AI devices, and usually has an actuator to complete specific action tasks. For example, a humanoid robot can use a robotic manipulator to complete the action task of picking up an object. Among them, the intelligent agent in the embodiments of the present invention needs to have a streaming video understanding model and provide strong streaming video understanding interaction capabilities according to the streaming video understanding model.

[0061] The streaming video understanding model in the embodiments of the present invention can be understood as a multi-modal large model for video content understanding interaction for long-time video input streams. The streaming video understanding model can integrate various modal data such as text, images, videos, and audio, and perform comprehensive understanding and reasoning, with stronger understanding capabilities and a wider range of application scenarios.

[0062] The model evaluation benchmark is standard evaluation standard data generated according to the evaluation requirements of the streaming video understanding model, specifically, it can be model evaluation standard data generated according to the required streaming video understanding tasks and the preset Q&A dataset. The preset Q&A dataset includes a matching dataset between questions and standard answers corresponding to different streaming video understanding tasks.

[0063] The evaluation video data can be a streaming video set data with a question set data of the above preset Q&A dataset, and can have a large-scale video stream data, where these video stream data can include corresponding question set data.

[0064] The evaluation video data can be disassembled and extracted according to the time nodes in the video stream according to each task question in the question set data, and the question set data corresponding to specific streaming video understanding tasks is extracted. The evaluation input data is the segmented data of the evaluation video data, which belongs to the video stream data corresponding to certain time periods in the evaluation video data, and these video stream data can include at least one question set of at least one streaming video understanding task.

[0065] Input the evaluation input data into the streaming video understanding model, and output the evaluation output data through the inference and interaction capabilities of the streaming video understanding model. Among them, the evaluation output data is the inference data of the streaming video understanding model, including the set of inference answers corresponding to the set of questions in the evaluation input data.

[0066] According to the matching relationship between the set of standard answers in the preset Q&A dataset that can be provided by the model evaluation benchmark and matches the streaming video understanding task and the corresponding set of inference answers in the above evaluation output data, obtain the model evaluation data. Among them, the model evaluation data is the evaluation data of the streaming video understanding ability generated according to the matching relationship between the inference answers in the above set of inference answers and the corresponding standard answers in the set of standard answers, such as scoring data.

[0067] Therefore, compared with the traditional technology that cannot systematically and accurately evaluate the streaming video understanding model, the method of the embodiment of the present invention can establish an evaluation benchmark for the streaming video understanding model, and achieve a comprehensive evaluation of the above streaming video understanding task according to this evaluation benchmark, so as to be able to achieve a comprehensive and accurate evaluation of the streaming video understanding model, making the streaming video understanding model have a more general evaluation standard.

[0068] As Figures 2 - 4 shown, according to an embodiment of the present invention, in operation S201, obtaining the evaluation video data according to the received model evaluation benchmark includes:

[0069] Generate a model evaluation benchmark according to a preset evaluation understanding task;

[0070] Obtain the evaluation video data corresponding to the preset evaluation understanding task.

[0071] The preset evaluation understanding task can be an understanding task that needs to be evaluated preset by the user for the inference evaluation scenario of the streaming video understanding model. Specifically, the preset evaluation understanding task of the embodiment of the present invention can be divided into two aspects: streaming video understanding and active inference. At the same time, streaming video understanding can be divided into three subtasks: action prediction task, dynamic state grounding task, and multi-turn dependency reasoning task, etc. Active inference can also be divided into three subtasks: proactive alerting task, noise recognition task, and speaker identification task, etc.

[0072] Among them, each subtask of the preset evaluation understanding task can be used as the aforementioned streaming video understanding task, and can achieve understanding and reasoning for different question-and-answer interaction scenarios.

[0073] Figure 3A FIG. schematically shows an application scenario diagram of an action prediction of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention.

[0074] The action prediction task can be used to judge the actions that the protagonist will take for a certain purpose from the first-person perspective. For example Figure 3A As shown, according to the recognition of consecutive frame video images, after receiving the question "I am tasked with Peel, chop and cook the second potato· What is next step?", the streaming video understanding model can give the action speculation of "pick up potato" by combining the understanding of the streaming video.

[0075] Figure 3B FIG. schematically shows an application scenario diagram of a dynamic state tracking task of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention.

[0076] The dynamic state tracking task can be used to, as the streaming video progresses, the streaming video understanding model can combine the understanding of the streaming video to perceive the possible changes that will occur in the current state shown in the video, and give a reasonable answer. For example Figure 3B As shown, the answers to the same question are different at different times. The model needs to correctly answer these questions. Corresponding to "How many cars have shown up?", the streaming video understanding model needs to combine the understanding of the video image to give that the corresponding number of cars at the "00:05" moment is "1", the corresponding number of cars at the "00:10" moment is "2", and the corresponding number of cars at the "00:17" moment is "3".

[0077] Figure 3C FIG. schematically shows an application scenario diagram of a multi-round dependent reasoning task of an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention.

[0078] The multi-round dependent reasoning task is used for common front-back dependent relationships in the streaming video, setting different questions at different times, and the current question comes from the answer to the previous question. For example Figure 3C As shown, a continuous question-and-answer scenario:

[0079] Q: Who is organizing the back pack?

[0080] A:A woman

[0081] Q: Who is talk to?

[0082] A:A man

[0083] Q: What is doing?

[0084] A: Sitting on the bed.

[0085] Figure 3D Schematically shows an application scenario diagram of an active reminder task for the evaluation method of the streaming video understanding model of an agent according to an embodiment of the present invention.

[0086] The active reminder task is used to give an event description of the streaming video understanding model given at the beginning, hoping that the model can send a reminder to the user when the video progresses to this event. As Figure 3D shown, providing "Inform me when a cat is getting a shot", the streaming video understanding model will provide a reminder of "Informed" at the "00:25" moment according to the understanding of the streaming video image.

[0087] Figure 3E Schematically shows an application scenario diagram of a noise recognition task for the evaluation method of the streaming video understanding model of an agent according to an embodiment of the present invention.

[0088] The noise recognition task can identify user instructions that need to be replied to and noises that do not need to be replied to. As Figure 3E shown, for the question "What happens to the man in black clothes after he stands on the cliff?", the streaming video understanding model can determine that it belongs to a non-noise question and give a reply of "He is caught by a net laid by a helicopter", but for what the streaming video understanding model determines to be a noise question of "Oh well, it is what it is No point in overthinking it", the streaming video understanding model can remain silent ( <silence>) No response will be given.

[0089] Figure 3F FIG. schematically shows an application scenario diagram of a speaker recognition task in the evaluation method of the streaming video understanding model of an agent according to an embodiment of the present invention.

[0090] The speaker recognition task can be used in a multi-person conversation scenario. The streaming video understanding model needs to figure out which speaker issued the current instruction that appears in the streaming video. Usually, these people can all have self-introductions or introductions by others. For example Figure 3F As shown, according to the introductions by others of "This is Bobi” and "This is Wanita”, for the question "Who wants to serve drink?”, the streaming video understanding model can give the answer "Wanita (wants to serve drink)”.

[0091] Therefore, different preset evaluation understanding tasks can correspond to different reasoning and interaction capabilities of the streaming video understanding model. Considering that the evaluation video data can be a set of streaming video data with a question set data of the above-mentioned preset Q&A data set, it can have a large-scale video stream data, and these video stream data can contain the corresponding question set data. These question set data can be used to match the questions of the video images in the video stream according to the above-mentioned preset evaluation understanding tasks. For example Figure 3A As shown, the question corresponding to the action prediction task is "I am tasked with Peel, chop and cook the second potato· What is next step”.

[0092] Among them, the obtained evaluation video data can have 1121 videos, corresponding to 2290 Q&A data sets for 6 different streaming video understanding tasks. Therefore, the evaluation video data obtained according to the preset evaluation understanding tasks can have a wider task coverage, making the evaluation of the streaming video understanding model more targeted and persuasive.

[0093] For example Figures 2 - 4 As shown, according to an embodiment of the present invention, in operation S202, when splitting the evaluation video data to generate evaluation input data, it includes:

[0094] Obtain the preset task questions in the preset evaluation understanding tasks;

[0095] Perform splitting of the evaluation video data according to the time nodes corresponding to the preset task questions to generate evaluation input data.

[0096] The evaluation video data can be an offline video. According to different preset evaluation understanding tasks (such as the streaming video understanding task of action prediction), preset task questions matching the task can be determined. For example, Figure 3A As shown, the question corresponding to the action prediction task is "I am tasked with Peel, chop and cook the second potato· What is next step”.

[0097] Furthermore, the evaluation video data is segmented according to the time nodes matched by the preset task questions in the video stream, and only the video stream data within the time period directly related to the preset task questions can be obtained as the evaluation input data. It can be understood as intercepting the images of the video stream according to the time nodes corresponding to the preset task questions. Among them, the preset task questions are the preset questions matched with the streaming video understanding task in the above evaluation video data, and usually have known matching standard answers. For example, Figure 3B As shown, for the preset task question "How many cars have shown up?” of the dynamic state tracking task, cars appear in the images of the video stream at the "00:05” moment, "00:10” moment and "00:17” moment respectively. Then, at least the video between the "00:05” moment and the "00:17” moment can be used as the evaluation input data, or the streaming video images at the "00:05” moment, "00:10” moment and "00:17” moment can be used as the evaluation input data, thus completing the segmentation of the evaluation video data.

[0098] Therefore, even if most video understanding models do not support streaming video input, the above offline video can be segmented according to the time nodes corresponding to different tasks, so as to achieve segmented input into the streaming video understanding model to be evaluated, so as to complete the simulation input effect of the video stream. In this way, the evaluation of invalid video understanding can be greatly avoided, and the evaluation efficiency and accuracy can be improved.

[0099] For example, Figures 2 - 4 As shown, according to an embodiment of the present invention, in operation S203, when generating the evaluation output data of the streaming video understanding model through the evaluation input data, it includes:

[0100] The evaluation input data is input into the streaming video understanding model in the form of simulated video stream input to generate evaluation output data.

[0101] Since the streaming video understanding model needs to perform inference and understanding interactions on the continuous input video stream during the application process of the intelligent agent. Therefore, by obtaining the segmentation of the above-mentioned evaluation input data, the simulation of the video input can be realized, so that the evaluation input data can be input into the streaming video understanding model in the form of a video stream according to the order of time nodes. Among them, the video images at different time nodes in the video stream of the evaluation input data can correspond to different preset task questions.

[0102] The streaming video understanding model processes according to the video stream images of the input evaluation input data, and can generate inference answers to the preset task questions corresponding to the preset evaluation understanding tasks in the evaluation input data. These inference answers can form the processing result data set of the streaming video understanding model. Among them, the evaluation output data is the inference data of the streaming video understanding model, including the inference answer set corresponding to the question set in the evaluation input data.

[0103] Therefore, by simulating the input form of the video stream to input the evaluation input data into the streaming video understanding model, the processing process of the streaming video understanding model can be made more in line with the actual application scenario, the credibility of the output evaluation output data is higher, and the efficiency and accuracy of the model processing can be ensured at the same time.

[0104] As Figures 2 - 4 shown, according to an embodiment of the present invention, in operation S204 of obtaining the model evaluation data of the streaming video understanding model according to the evaluation output data and the model evaluation benchmark, it includes:

[0105] Obtain the evaluation task answers corresponding to the preset task questions in the evaluation output data;

[0106] Generate model evaluation data according to the evaluation task answers and the evaluation benchmark answers in the model evaluation benchmark.

[0107] For the preset task questions corresponding to different streaming video understanding tasks in the evaluation input data, after being processed by the streaming video understanding model, an inference answer set with the corresponding streaming video understanding task can be generated, that is, the above-mentioned preset task questions can all be answered by the model.

[0108] The evaluation task answers are the inference answers to different preset task questions corresponding to different streaming video understanding tasks, and the evaluation benchmark answers are the preset standard answers to different preset task questions corresponding to different streaming video understanding tasks. Among them, the evaluation benchmark answers can be obtained according to the question-standard answer matching data set of the preset Q&A data set of the model evaluation benchmark.

[0109] By comparing the evaluation task answers with the benchmark answers in the model evaluation benchmark, the accuracy of the evaluation task answers for the corresponding streaming video understanding task can be calculated. Further, the accuracy of each evaluation task answer for all 6 different streaming video understanding tasks can be calculated, and the average value of these accuracies can be used as a judgment basis for reflecting the strength of the streaming video understanding and reasoning ability of the streaming video understanding model. Among them, the model evaluation data can be the evaluation result of the above-mentioned streaming video understanding model for the strength of the streaming video understanding and reasoning ability.

[0110] For example, for the Figure 3A action prediction task shown, if the evaluation task answer given by the streaming video understanding model is "pick up potato", then the reasoning and understanding accuracy of the streaming video understanding model for this action prediction task is 100%. Correspondingly, if the evaluation task answer given by the streaming video understanding model is "pick up tomato", then the reasoning and understanding accuracy of the streaming video understanding model for this action prediction task can be 60%.

[0111] For the Figure 3B dynamic state tracking task shown, if the evaluation task answers given by the streaming video understanding model are: the number of cars corresponding to "00:05" is "1", the number of cars corresponding to "00:10" is "2", and the number of cars corresponding to "00:17" is "3", then the reasoning and understanding accuracy of the streaming video understanding model for this action prediction task is 100%. However, if the evaluation task answers given by the streaming video understanding model are: the number of cars corresponding to "00:05" is "1", the number of cars corresponding to "00:10" is "2", and the number of cars corresponding to "00:17" is "2", then the reasoning and understanding accuracy of the streaming video understanding model for this action prediction task can be 66%.

[0112] For the Figure 3C multi-round dependent reasoning task shown, if the evaluation task answers given by the streaming video understanding model do not involve the front-back dependency relationship of "the current question comes from the answer of the previous question", then even if the streaming video understanding model gives reasonable answers to the questions that appear at different times, the reasoning and understanding accuracy of the streaming video understanding model for this action prediction task may only be 20%.

[0113] For the Figure 3D For the active reminder task shown, if the evaluation task answer given by the streaming video understanding model does not provide an "Informed" reminder to the user at the "00:25" moment of "a cat is getting a shot", the inference and understanding accuracy rate of the streaming video understanding model for this action prediction task can be 0%. If the evaluation task answer given by the streaming video understanding model provides an "Informed" reminder at the "00:24" moment, which is one second before the "00:25" moment of "a cat is getting a shot", the inference and understanding accuracy rate of the streaming video understanding model for this action prediction task can be 90%.

[0114] For Figure 3E the noise recognition task shown, if the evaluation task answer given by the streaming video understanding model provides something other than <silence>For the reply, the inference and understanding accuracy rate of the streaming video understanding model for the action prediction task can be 0%, and the corresponding accuracy rate value can be matched according to the degree of the reply.

[0115] For Figure 3F the speaker recognition task shown, if the evaluation task answer given by the streaming video understanding model is not "Wanita" but "Bob", the inference and understanding accuracy rate of the streaming video understanding model for the action prediction task can be 50%. If the evaluation task answer given by the streaming video understanding model is neither "Wanita" nor "Bob", the inference and understanding accuracy rate of the streaming video understanding model for the action prediction task can be 0%.

[0116] Finally, take the average value of the inference and understanding accuracy rates of the above 6 different streaming video understanding tasks, which can be used as the overall inference and understanding accuracy rate of the streaming video understanding task, and constitute the final model evaluation data of the streaming video understanding task.

[0117] In summary, compared with the traditional technology where the streaming video understanding model cannot be systematically and accurately evaluated, the method of the embodiment of the present invention can establish an evaluation benchmark for the streaming video understanding model from two aspects of streaming video and active inference, which is the first evaluation standard in the field. Moreover, according to this evaluation benchmark, a comprehensive evaluation of a total of 6 sub-tasks in the above two aspects of streaming video understanding and active inference can be realized, so as to realize a comprehensive and accurate evaluation of the streaming video understanding model, enabling the streaming video understanding model to have a more general evaluation standard.

[0118] For Figures 2 - 4 shown, according to an embodiment of the present invention, the evaluation method of the streaming video understanding model of the intelligent agent further includes:

[0119] Train the streaming video understanding model according to the model evaluation data and the preset instruction following data, where the preset instruction following data includes the active inference data corresponding to the interleaved text and image instructions, noise instructions, and interruption instructions.

[0120] According to the above model evaluation data, the understanding and reasoning interaction ability of the streaming video understanding model can be intuitively judged. Therefore, when the interaction ability of the streaming video understanding model is poor, further training of the streaming video understanding model can be considered to overall improve the understanding and reasoning interaction ability of the model. On the other hand, considering that the current streaming video understanding model as a multi-modal large model develops rapidly, but seriously lacks the ability of active inference and interaction, such as the ability of interruption and noise recognition, resulting in a still high room for improvement in its human-computer interaction experience, further training of the streaming video understanding model with strong inference and understanding interaction ability for the above 6 tasks can also be considered to improve its human-computer interaction experience.

[0121] In view of the above situation, the preset instruction-following data in the embodiments of the present invention can be a dataset with a small amount of data and in a multi-modal instruction-following format for images, texts, and interleaved images and texts. By fine-tuning the preset instruction-following data, the interactive ability can be added to the streaming video understanding model of the multi-modal large model at a lower cost. The dataset of the preset instruction-following data mainly includes the following three parts: interleaved image and text instructions, noise instructions, and interruption instructions. Among them, the interleaved image and text instructions mean that training pictures and instructions will appear alternately in the training data to improve the active reasoning ability of the model. Secondly, the noise instructions can be used as negative samples to enhance the robustness of the model, such as "I understand. This video shows a flag fluttering". In addition, the interruption instructions add the interruption ability to the model, such as "Sorry to interrupt". Therefore, by training the above-mentioned streaming video understanding model that obtains the above-mentioned model evaluation data based on the dataset of the preset instruction-following data, the active reasoning and interactive abilities can be injected into the streaming video understanding model that has completed the evaluation, thereby significantly improving its level of streaming video understanding and reasoning.

[0122] Figure 4 Another application scenario flowchart of the evaluation method for the streaming video understanding model of the intelligent agent according to the embodiment of the present invention is schematically shown.

[0123] As Figure 4 shown, because the interleaved image and text instructions are used, the active reasoning of the streaming video understanding model that has completed the training of the preset instruction-following data can be regarded as the time localization of the future. This method can be applied to the KV-cache storage video streaming of the Omni Language Model, combined with the highlight spot max-heap algorithm. When the video segment of a certain event node continuously maintains a high attention score, the model can issue an active reminder at that place. For the interactive ability of the large model, because it is trained on the noise instructions and stop instructions, the fact interactive dialogue including proactive interruption can be realized through the way of parallel decoding.

[0124] Therefore, based on the above evaluation of the reasoning and understanding interactive abilities of the streaming video understanding model, the interactive ability can be added to the streaming video understanding model at a lower cost according to the fine-tuning effect corresponding to the provided preset instruction-following dataset, thereby significantly improving the active reasoning ability of the streaming video understanding model.

[0125] In order to endow the current multimodal model with active inference ability, the above method of the embodiments of the present invention provides an instruction-following dataset, and the active inference ability of the model can be realized by fine-tuning on this dataset. Further, by evaluating the streaming video understanding model and injecting an interactive function into the current multimodal large language model, the streaming video understanding model can realize the ability of intelligent "interruption and noise recognition".

[0126] Compared with the traditional technology that cannot systematically and accurately evaluate the streaming video understanding model, the method of the embodiments of the present invention can establish an evaluation benchmark for the streaming video understanding model from two aspects of streaming video and active inference, and comprehensively evaluate 6 sub-tasks in the above two aspects of streaming video understanding and active inference according to this evaluation benchmark, so as to realize a comprehensive and accurate evaluation of the streaming video understanding model, and enable the streaming video understanding model to have a more general evaluation standard. Further, the interactive ability can be added to the streaming video understanding model at low cost in combination with the fine-tuning effect corresponding to the provided instruction-following dataset, thereby significantly improving the active inference ability of the streaming video understanding model.

[0127] Based on the above evaluation method of the streaming video understanding model of the intelligent agent, the present invention also provides an evaluation device for the streaming video understanding model of the intelligent agent. The following will be combined with Figure 5 to describe this device in detail.

[0128] Figure 5 The structural block diagram of the evaluation device for the streaming video understanding model of the intelligent agent according to the embodiments of the present invention is schematically shown.

[0129] As Figure 5 shown, the evaluation device 500 for the streaming video understanding model of the intelligent agent in this embodiment includes a data acquisition module 510, a data segmentation module 520, an evaluation output module 530, and an evaluation acquisition module 540.

[0130] The data acquisition module 510 is used to obtain evaluation video data according to the received model evaluation benchmark. In one embodiment, the data acquisition module 510 can be used to perform the operation S201 described above, which will not be elaborated here.

[0131] The data segmentation module 520 is used to segment the evaluation video data to generate evaluation input data. In one embodiment, the data segmentation module 520 can be used to perform the operation S202 described above, which will not be elaborated here.

[0132] The evaluation output module 530 is used to generate evaluation output data of the streaming video understanding model through the evaluation input data. In one embodiment, the evaluation output module 530 can be used to perform the operation S203 described above, which will not be elaborated here.

[0133] The evaluation acquisition module 540 is configured to obtain model evaluation data of the streaming video understanding model according to the evaluation output data and the model evaluation benchmark. In one embodiment, the evaluation acquisition module 540 may be configured to perform the operation S204 described above, which will not be elaborated herein.

[0134] According to an embodiment of the present invention, any multiple of the data acquisition module 510, the data segmentation module 520, the evaluation output module 530, and the evaluation acquisition module 540 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the data acquisition module 510, the data segmentation module 520, the evaluation output module 530, and the evaluation acquisition module 540 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner that can integrate or package the circuit, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the data acquisition module 510, the data segmentation module 520, the evaluation output module 530, and the evaluation acquisition module 540 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.

[0135] Figure 6 A block diagram of an electronic device suitable for implementing an evaluation method of a streaming video understanding model of an agent according to an embodiment of the present invention is schematically shown.

[0136] The above-mentioned electronic device provided by the embodiment of the present invention includes one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned evaluation method of the streaming video understanding model of the agent.

[0137] As Figure 6 As shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. The processor 601 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 601 can also include on-board memory for caching purposes. The processor 601 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0138] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 executes various operations of the method flow according to an embodiment of the present invention by executing the programs in the ROM 602 and / or RAM 603. It should be noted that the program can also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 can also execute various operations of the method flow according to an embodiment of the present invention by executing the programs stored in the one or more memories.

[0139] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that the computer program read from it can be installed into the storage section 608 as needed.

[0140] The present invention also provides a computer-readable storage medium having executable instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned evaluation method of the intelligent agent's streaming video understanding model.

[0141] Among them, the computer-readable storage medium may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not be assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0142] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.

[0143] An embodiment of the present invention also includes a computer program product, which includes a computer program, and when the computer program is executed by a processor, the evaluation method of the streaming video understanding model of the above-mentioned agent is implemented.

[0144] Among them, the computer program contains program codes for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program codes are used to enable the computer system to implement the method provided by the embodiments of the present invention.

[0145] When the computer program is executed by the processor 601, the above-mentioned functions defined in the system / apparatus of the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0146] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or be installed from the removable medium 611. The program codes included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0147] In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above functions defined in the system of the embodiments of the present invention are executed. According to the embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc. described above can be implemented by computer program modules.

[0148] According to the embodiments of the present invention, the program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0150] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present invention can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0151] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present invention.< / silence> < / silence>

Claims

1. An evaluation method for a streaming video understanding model of an intelligent agent, characterized in that, Including: Obtain evaluation video data corresponding to a preset evaluation and understanding task according to the received model evaluation benchmark. Among them, the preset evaluation and understanding task is divided into two aspects: streaming video understanding and active reasoning. Streaming video understanding is divided into three subtasks: action prediction task, dynamic state tracking task, and multi-round dependency reasoning task. Active reasoning is divided into three subtasks: active reminder task, noise recognition task, and speaker recognition task; Perform segmentation on the evaluation video data to generate evaluation input data; Input the evaluation input data into the streaming video understanding model in the form of simulating video stream input to generate evaluation output data; and Obtain the model evaluation data of the streaming video understanding model according to the evaluation output data and the model evaluation benchmark; The model evaluation benchmark is evaluation standard data generated according to the streaming video understanding task and the preset Q&A dataset required for the evaluation of the streaming video understanding model. The preset Q&A dataset includes a matching dataset between questions and standard answers corresponding to different streaming video understanding tasks; Train the streaming video understanding model according to the model evaluation data and the preset instruction following data. The preset instruction following data includes active reasoning data corresponding to interleaved text and image instructions, noise instructions, and interruption instructions; Among them, the active reasoning of the streaming video understanding model that has completed the training of the preset instruction following data is applicable to the KV-cache of the Omni language model to store video streams, and combines the maximum heap algorithm to issue active reminders. In addition, a fact interactive dialogue including interruption at any time is realized through parallel decoding.

2. The method according to claim 1, wherein In the step of obtaining evaluation video data according to the received model evaluation benchmark, it includes: Generate the model evaluation benchmark according to the preset evaluation and understanding task.

3. The method according to claim 2, wherein In the step of performing segmentation on the evaluation video data to generate evaluation input data, it includes: Obtain the preset task questions in the preset evaluation and understanding task; Perform segmentation on the evaluation video data according to the time nodes corresponding to the preset task questions to generate evaluation input data.

4. The method according to claim 3, characterized in that, In the step of obtaining the model evaluation data of the streaming video understanding model according to the evaluation output data and the model evaluation benchmark, it includes: Obtain the evaluation task answers corresponding to the preset task questions in the evaluation output data; Generate the model evaluation data according to the evaluation task answers and the evaluation benchmark answers in the model evaluation benchmark.

5. An evaluation device for a streaming video understanding model of an intelligent agent, characterized in that, Including: A data acquisition module for obtaining evaluation video data corresponding to a preset evaluation and understanding task according to the received model evaluation benchmark. Among them, the preset evaluation and understanding task is divided into two aspects: streaming video understanding and active reasoning. Streaming video understanding is divided into three subtasks: action prediction task, dynamic state tracking task, and multi-round dependency reasoning task. Active reasoning is divided into three subtasks: active reminder task, noise recognition task, and speaker recognition task; A data segmentation module for performing segmentation on the evaluation video data to generate evaluation input data; An evaluation output module for inputting the evaluation input data into the streaming video understanding model in the form of simulating video stream input to generate evaluation output data; and An evaluation acquisition module, configured to obtain model evaluation data of the streaming video understanding model according to the evaluation output data and the model evaluation benchmark; The model evaluation benchmark is evaluation standard data generated according to the streaming video understanding task and the preset Q&A dataset required for the evaluation of the streaming video understanding model, and the preset Q&A dataset includes a matching dataset between questions and standard answers corresponding to different streaming video understanding tasks; The evaluation device is further configured to train the streaming video understanding model according to the model evaluation data and the preset instruction following data, where the preset instruction following data includes active inference data corresponding to interleaved text and image instructions, noise instructions, and interruption instructions; Among them, the active inference of the streaming video understanding model after completing the training of the preset instruction following data is applicable to the KV-cache of the Omni language model to store the video stream, and combined with the maximum heap algorithm, an active reminder is issued; in addition, a fact interactive dialogue including interruption at any time is realized through parallel decoding.

6. An electronic device, comprising: One or more processors; A memory for storing one or more programs, Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 4.

8. A computer program product, comprising a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Evaluation method and device for visual question and answer task, medium and computer program product

    CN118467709A