Knowledge enhanced sports video understanding method based on large model double-mode inference
By introducing a dual-mode reasoning system and a sports knowledge graph, the FineQuest framework addresses the dynamic and complex issues in sports video question answering, achieving efficient understanding and accurate responses to sports videos and improving the performance of multimodal large language models in sports video question answering tasks.
Patent Information
- Application Number
- CN202511263018.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing multimodal large language models face challenges in handling sports video question answering tasks due to their dynamic nature and complexity, differences in domain knowledge, and diversity of questions. They struggle to effectively understand the rapidly changing scenarios, rich domain knowledge, and complex user questions in sports videos.
A dual-mode reasoning method based on a large model is adopted, including reactive reasoning and deductive reasoning. By utilizing a dynamic motion segmenter, a key segment selector, and a fine-grained matcher based on a sports knowledge graph, the method adaptively segments the video, identifies key segments, and performs accurate matching based on the sports knowledge graph to achieve understanding of sports videos.
It significantly improves the accuracy and robustness of sports video question answering, especially performing well on complex questions, with an average accuracy improvement of 19.21% compared to existing methods, and demonstrates excellent adaptability and scalability on questions of varying difficulty.
Smart Images

Figure CN120745851B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of video understanding, in particular to a knowledge enhanced sports video understanding method based on large model double mode reasoning. BACKGROUND
[0002] Sports events attract global audiences due to their intense action, strategic gameplay, and outstanding athletic performances. With the continuous advancement of technology, there is an increasing demand for personalized and interactive viewing experiences. In this context, Video Question Answering (VideoQA) as an emerging technology, shows great potential in providing real-time insights and user-driven exploration in sports content. For example, audiences can ask natural language questions to obtain detailed explanations about game rules, athlete performance, or key events in real time, greatly enhancing the viewing experience.
[0003] Video Question Answering based on (Multimodal) Large Language Models (M)LLMs has shown great potential in the field of general video understanding in recent years, mainly due to its excellent performance in multi-modal information fusion, natural language understanding, and complex reasoning tasks. However, sports videos, as a special type of video with high dynamicity, domain specificity, and complex semantics, pose higher requirements on existing VideoQA methods. The rapid changes in action and scene, rich domain knowledge, and the diversity and complexity of user questions in sports videos make it challenging to directly apply large models to sports video question answering tasks. These challenges mainly manifest in the following three aspects:
[0004] 1) Dynamicity and complexity: The rapid changes in action and scene in sports videos increase the difficulty of video understanding;
[0005] 2) Domain knowledge difference: General large language models lack the ability to understand sports rules, terminology, and background knowledge;
[0006] 3) Question diversity: User questions often involve complex logical reasoning and context dependence, making it difficult for existing methods to provide accurate answers. SUMMARY
[0007] To address the problems existing in the prior art, the present application proposes a knowledge enhanced sports video understanding method based on large model double mode reasoning, which specifically includes the following steps:
[0008] Step 1: Obtain the sports video to be asked and the question text;
[0009] Step 2: input the sports video, question text and prompt words into the reactive reasoning agent, the reactive reasoning agent classifies the question according to the question text and prompt words, and determines whether the question belongs to a simple question or a complex question; if the question belongs to a simple question, go to step 3, if the question belongs to a complex question, go to step 4;
[0010] Step 3: the reactive reasoning agent answers the question according to the input sports video;
[0011] Step 4: input the sports video and question text into the constructed deliberate reasoning agent;
[0012] The deliberate reasoning agent is composed of a dynamic motion segmenter, a key segment selector and a fine-grained matcher based on a sports knowledge graph;
[0013] The dynamic motion segmenter adaptively reduces background interference and segments the sports video into multiple dynamic segments according to the motion intensity;
[0014] The key segment selector uses a multi-level contrast decoding strategy to identify key segments related to the question in the spatial, temporal and spatio-temporal dimensions;
[0015] The fine-grained matcher based on the sports knowledge graph realizes accurate matching of key segments and questions with the help of structured information extracted from the sports knowledge graph, thereby outputting corresponding knowledge.
[0016] Further, the prompt word content includes: 1) the relevance of the question and the video; 2) the question type; 3) the reasoning requirement; 4) whether external knowledge is needed.
[0017] Further, the dynamic motion segmenter includes a video content segmentation module, a motion intensity extraction module and a segment segmentation module;
[0018] The video content segmentation module segments each frame of image in the sports video and identifies the character subject and background in each frame of image;
[0019] The motion intensity extraction module obtains a plurality of consecutive frames of images in the sports video through a sliding window, and performs optical flow analysis on the images frame by frame to obtain the motion intensity of the character subject action in the sliding window; by moving the sliding window along the time axis direction of the sports video, the motion intensity of the character subject action in each sliding window is obtained, thereby obtaining the motion intensity change curve of the character subject action along the time axis direction of the sports video;
[0020] The segment segmentation module segments the sports video into multiple dynamic segments according to the motion intensity change curve of the character subject action and the set threshold of the motion intensity.
[0021] Further, the video content segmentation module is implemented by using a SAM2 model.
[0022] Further, the key segment selector comprises a multi-layer contrast decoding module, a correlation calculation module and a sorting module.
[0023] The multi-layer contrast decoding module respectively processes the original dynamic segment from the spatial, temporal and spatio-temporal dimensions.
[0024] The correlation calculation module respectively extracts features from an original dynamic segment and the corresponding three distorted dynamic segments, and extracts features from the input question text; calculates the correlation between the features of the four dynamic segments and the features of the question text, and performs weighted processing on the obtained correlation to obtain the correlation score of the original dynamic segment and the question text.
[0025] The sorting module sorts the correlation scores of each original dynamic segment and the question text to obtain the original dynamic segment most relevant to the question text as the key segment.
[0026] Further, the specific process of distortion processing from the spatial, temporal and spatio-temporal dimensions is as follows:
[0027] The distortion processing from the spatial dimension means adding Gaussian noise to each frame in the original dynamic segment.
[0028] The distortion processing from the temporal dimension means adjusting the duration of each frame while keeping the order of the frames unchanged.
[0029] The distortion processing from the spatio-temporal dimension means combining the spatial dimension distortion and the temporal dimension distortion.
[0030] Further, the weights corresponding to the distorted dynamic segments from the spatial, temporal and spatio-temporal dimensions are , and , wherein is the spatial weight, taking a value of 0.5, is the temporal weight, taking a value of 0.3, is the spatio-temporal weight, taking a value of 0.2.
[0031] Further, the specific process of obtaining the correlation score of the original dynamic segment and the question text is as follows: using the encoder in the multi-modal large model to extract features from each dynamic segment and the question text, and outputting the correlation probability distribution of the features of each dynamic segment and the features of the question text, using the correlation probability distribution of the features of the original dynamic segment and the features of the question text, subtracting the correlation probability distribution of each weighted distorted dynamic segment feature and the question text feature, and the corresponding weight is 、 and wherein is a spatial weight, is a temporal weight, is a spatio-temporal weight; and decoding the result obtained after subtraction to obtain a relevance score of the original dynamic segment and the question text.
[0032] Further, the fine-grained matcher based on the sports knowledge graph comprises a key segment description module, a multi-level matching module, and a knowledge extraction module.
[0033] The key segment description module describes the key segment by using a multi-modal large model to obtain a key segment description text.
[0034] The multi-level matching module calculates the following three cosine similarities:
[0035] (1) the cosine similarity between the key segment description text and the description text of each instance of the corresponding sports project in the sports knowledge graph;
[0036] (2) the cosine similarity between the key segment and the video of each instance of the corresponding sports project in the sports knowledge graph;
[0037] (3) the cosine similarity between the key segment description text and the scene information description text of the video of each instance of the corresponding sports project in the sports knowledge graph;
[0038] The three kinds of cosine similarities corresponding to each instance in the corresponding sports project are weighted and summed to obtain the similarity of the key segment and each instance of the corresponding sports project, and then the instance with the highest similarity is obtained.
[0039] The knowledge extraction module extracts the knowledge corresponding to the instance with the highest similarity and outputs it to the multi-modal large model, and answers the question through the multi-modal large model.
[0040] Further, the reaction formula reasoning intelligent agent and the multi-modal large model adopt a Video-LLaVA large model or a LLaVA-Next-Video large model.
[0041] Beneficial effects:
[0042] The FineQuest framework introduced by the present application provides a no-training solution for the complex task of sports video question answering (VideoQA) by innovatively introducing a dual-mode reasoning system (reactive reasoning and reflective reasoning). The design of FineQuest fully considers the dynamics, domain specificity of sports video, and the diversity and complexity of user questions, and significantly improves the performance of (multi-modal) large language models ((M)LLMs) in sports video understanding tasks by combining the reasoning mechanism inspired by cognitive science and the integration of domain knowledge. The core advantages of FineQuest are reflected in the following aspects:
[0043] (1) Flexibility and adaptability of dual-mode reasoning
[0044] Reactive reasoning can quickly handle simple problems, while reflective reasoning can effectively deal with complex problems through multi-step reasoning and context analysis. This dual-mode reasoning mechanism not only improves the performance of FineQuest on problems of different complexity, but also exhibits its flexibility in multiple scenarios.
[0045] (2) Combining sports knowledge graph to realize domain knowledge enhancement
[0046] By combining the sports knowledge graph with the reflective reasoning agent, FineQuest successfully bridges the gap between general (M)LLMs and sports domain knowledge. The structured domain knowledge and visual semantic information provided by the sports knowledge graph enable FineQuest to more accurately understand sports terminology, rules, and dynamic scenarios, thereby significantly improving the accuracy and context awareness of reasoning.
[0047] Experimental results show that FineQuest's performance on typical datasets has reached the current state-of-the-art level, with an average accuracy improvement of 19.21% compared to existing no-training methods. In addition, the modular design and multi-level reasoning mechanism enable FineQuest to exhibit excellent robustness and scalability on problems of different difficulty.
[0048] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter in the description of the application. BRIEF DESCRIPTION OF DRAWINGS
[0049] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings, in which:
[0050] Figure 1Motivation of the present invention and the proposed FineQuest framework overview; (a) Examples of Gym-QA and Diving-QA, (b) Training-based video question answering framework, (c) Existing training-free video question answering framework, (d) FineQuest inference process framework;
[0051] Figure 2 : Schematic diagram of deep thinking reasoning intelligent agent;
[0052] Figure 3 : Results of sports video question answering benchmark (Gym-QA and Diving-QA full set) test;
[0053] Figure 4 : Results of sports video question answering benchmark (Gym-QA and Diving-QA action set) test;
[0054] Figure 5 : Results of sports video question answering benchmark (SPORTU) test;
[0055] Figure 6 : Results of general video question answering benchmark test; DETAILED DESCRIPTION
[0056] The embodiments of the present application are described in detail below, which are exemplary and intended to explain the present application, and cannot be understood as a limitation of the present application.
[0057] In recent years, training-free VideoQA methods have gradually attracted attention. These methods usually use a combination of visual language models (VLMs) and large language models (LLMs) to generate descriptions through multi-modal pre-training models, and combine language models for reasoning. However, existing methods still have significant limitations when applied to sports videos, mainly in the following aspects:
[0058] (1) Complexity of dynamic scenes: Existing methods mostly rely on scene clustering-based strategies, but this strategy is difficult to cope with frequent camera cuts in sports broadcast. The actions in sports videos are usually continuous and complex, and can only be generated by single-frame or a small number of frame descriptions, which cannot capture the complete movement process and details.
[0059] (2) Lack of domain knowledge: general visual language models lack the ability to map visual clues to professional sports terminology. For example, terms such as "626B" need to be combined with specific movement rules and background knowledge to accurately explain, and existing models generally perform poorly on domain-specific tasks.
[0060] (3) Diversity and complexity of questions: sports video question answering has a variety of question forms, such as Figure 1(a) as shown, contains Q1 (simple questions) and Q2 (complex questions), i.e. from simple "What is this movement?" to complex "How many sets of action sub-set are completed?", which may appear in sports video question answering. Simple questions usually rely on basic pattern recognition, while complex questions require multi-step reasoning, including video segment identification, rule understanding, and performance analysis. Existing methods lack the ability to dynamically adjust reasoning strategies, making it difficult to handle both simple and complex questions. For example, Figure 1 (b) and Figure 1 (c) as shown, existing training-based video question answering frameworks and training-free video question answering frameworks can answer the correct answers when facing simple questions, but when facing complex questions, their performance is generally poor.
[0061] Based on the above analysis, the limitations of existing methods in handling sports video question answering further highlight the need to develop new methods, and the key to solving this problem lies in: (1) designing a framework that can dynamically adjust reasoning strategies according to the complexity of the question; (2) integrating domain-specific knowledge to make up for the shortcomings of general models; (3) designing a more effective scene understanding mechanism for the characteristics of sports videos. To this end, the present embodiment proposes a knowledge-enhanced sports video understanding method based on large model dual-mode reasoning, which designs a new training-free sports video question answering framework FineQuest, as shown in Figure 1 (d) inspired by the dual-process theory in cognitive science, it adopts two complementary reasoning modes: reactive reasoning and deliberative reasoning, specifically:
[0062] Reactive Reasoning: For simple, low-complexity questions, FineQuest adopts a single-step reasoning process, using pre-trained (multimodal) large language models to achieve fast response. This mode ensures efficiency and is suitable for most common questions.
[0063] Deliberative Reasoning: For complex, high-context-dependent questions, a multi-step reasoning process is adopted, involving the following steps: (1) identify relevant video segments through dynamic motion segmentation, (2) filter important clips through key clip selection based on the question, (3) use domain-specific knowledge for fine-grained matching; thereby systematically decomposing complex questions into sub-tasks and gradually reasoning to complete complex questions.
[0064] This dual-mode reasoning framework ensures the efficiency and robustness of FineQuest, enabling it to handle diverse questions with accuracy and adaptability.
[0065] Specifically, the knowledge enhancement type sports video understanding method based on large model double mode inference proposed in this embodiment includes the following steps:
[0066] Step 1: Obtain the sports video that needs to be asked and the question text;
[0067] Step 2: Input the sports video, question text and prompt words into the reactive inference agent. The reactive inference agent classifies the question according to the question text and prompt words, and determines whether the question belongs to a simple question or a complex question. If the question belongs to a simple question, go to step 3, if the question belongs to a complex question, go to step 4;
[0068] In this embodiment, the reactive inference agent uses the known Video-LLaVA large model or LLaVA-Next-Video large model; and the prompt word content used includes: 1) Question-video relevance; 2) Question Type; 3) Reasoning Requirements; 4) External Knowledge Dependency.
[0069] In this embodiment, the prompt words of the reactive inference agent provided are as follows:
[0070] You are a sports expert skilled in analyzing various sports videos. Your task is to evaluate sports videos and answer related questions based on them. Your goal is to determine the complexity of each question and decide on the appropriate action. Please follow the structured reasoning process below:
[0071] Step 1: Question difficulty analysis
[0072] Please evaluate the difficulty of the question from the following dimensions:
[0073] 1. Question-video relevance: Determine whether the video can directly answer the question. Consider the clarity and directness of the visual information in the video to the question.
[0074] 2. Question type: Identify whether the question is a static question (such as action matching) or a dynamic question (such as event inference). Static questions usually only require the identification of a specific element, while dynamic questions involve understanding a series of actions or events.
[0075] 3. Reasoning requirements: Evaluate whether the question requires single-step reasoning or multi-step reasoning. Single-step questions can get answers through direct reasoning, while multi-step questions require more complex logical chains.
[0076] 4. External knowledge dependency: Determine if the answer to the question requires domain knowledge beyond the visual information in the video.
[0077] Step 2: Decision
[0078] Based on the analysis in Step 1, choose one of the following actions:
[0079] Direct answer: If the question is relatively simple, provide the answer and explain your reasoning. Clearly explain how the video content supports your answer.
[0080] System switch: If the question is complex, trigger a system switch to delegate the task to a deliberative reasoning agent for more complex and in-depth reasoning.
[0081] Your response should include a "switch" instruction and explain why a switch is needed, highlighting the complexity factors identified in your analysis.
[0082] Guidelines:
[0083] Ensure your professional analysis and decision-making process is clear and accurate. Provide a full explanation of your actions and, where applicable, support them with video content as evidence. Maintain objectivity and neutrality in the reasoning process.
[0084] The above are the prompt words used by the reactive reasoning agent in this embodiment.
[0085] Step 3: The reactive reasoning agent answers the question based on the input sports video;
[0086] Step 4: Input the sports video and question text into the constructed deliberative reasoning agent;
[0087] Traditional training-free video question answering paradigms usually use clustering algorithms to identify scene transitions and divide the video into multiple segments. Key frames are selected based on the question and a visual language model is used as a description generation agent to generate descriptions for these key frames. These descriptions are input together into a large language model, which acts as a reasoning agent to predict the answer. Although this paradigm works well for general videos, it faces the following challenges when dealing with sports videos: (1) Sports videos are usually composed of a single and continuous scene, with complex backgrounds, making scene-based segmentation methods ineffective. (2) Frequent camera cuts in sports videos are often misjudged as scene changes by clustering algorithms. (3) Key frames cannot capture the complete semantics of complex sports actions.
[0088] To address these limitations, the present embodiment proposes a deliberate reasoning agent to gradually decompose the difficulty of the input question, as shown in Figure 2 The deliberate reasoning agent is composed of a dynamic motion segmenter, a key segment selector, and a fine-grained matcher based on a sports knowledge graph.
[0089] The dynamic motion segmenter adaptively reduces background interference and segments the sports video according to motion intensity, thereby enhancing the focus on key actions. The key segment selector uses a multi-level contrastive decoding strategy to identify key segments related to the question in spatial, temporal, and spatio-temporal dimensions, optimizing the segment selection process from multiple dimensions. The fine-grained matcher based on the sports knowledge graph uses structured information extracted from the sports knowledge graph to model knowledge from instance level to scene level, achieving accurate matching of key segments and questions, and thereby outputting corresponding knowledge.
[0090] The dynamic motion segmenter includes a video content segmentation module, a motion intensity extraction module, and a segment segmentation module.
[0091] Due to the lack of well-labeled sports video datasets, relying on existing timestamp annotations for pre-training may result in poor generalization ability. To alleviate the ambiguity problem caused by this, the sports video is input into the dynamic motion segmenter to adaptively segment sub-actions without relying on pre-defined timestamps. The video content segmentation module segments each frame of the sports video and identifies the human subjects and background in each frame, highlighting the athletes in the sports video. This method ensures the preservation of key background information, enabling accurate description generation later. In the present embodiment, the video content segmentation module is implemented using the well-known SAM2 (Segment Anything Model 2) model.
[0092] The dynamic motion segmenter segments the video based on changes in motion intensity, which naturally fluctuates during athletic performance. Specifically, athletes typically exhibit different motion intensities between action groups, accompanied by brief pauses. Therefore, to capture these dynamic changes, the motion intensity extraction module obtains a number of consecutive frames of the sports video through a sliding window and performs optical flow analysis on the frames to obtain the motion intensity of the human subject's actions within the sliding window. By moving the sliding window along the time axis of the sports video and obtaining the motion intensity of the human subject's actions within each sliding window, the motion intensity change curve of the human subject's actions along the time axis of the sports video is obtained.
[0093] The segment segmentation module divides the sports video into multiple dynamic segments according to the motion intensity change curve of the human subject's actions, with each dynamic segment corresponding to a sub-action, based on a set threshold of motion intensity.
[0094] For the segmented segments, the problem usually only involves a small part of the content in the video, although the existing method directly measures the similarity between the frames and the questions by using the visual language model, this method is vulnerable to the illusion phenomenon, which is caused by the bias in the training data and the over-reliance on the language priori. In order to solve this problem, the key segment selector is used here to improve the correlation score by comparing the output results of the original input and the distorted input.
[0095] The key segment selector comprises a multi-layer contrast decoding module, a correlation calculation module and a sorting module.
[0096] Unlike previous works that focus on images, the input of the present application is a video segment, so a distortion method that can maintain semantic integrity in spatial and temporal dimensions is needed. Specifically, the multi-layer contrast decoding module distorts the original dynamic segment from the spatial, temporal and spatio-temporal dimensions respectively; the specific process of distorting from the spatial, temporal and spatio-temporal dimensions is as follows:
[0097] Distorting from the spatial dimension means adding Gaussian noise to each frame in the original dynamic segment;
[0098] Distorting from the temporal dimension means adjusting the duration of each frame while keeping the order of the frames unchanged, for example, the original dynamic segment consists of 1, 2, 3, 4 frames, and after the temporal dimension distortion, it forms 1, 2, 2, 3, 4 frames;
[0099] Distorting from the spatio-temporal dimension means combining the spatial dimension distortion and the temporal dimension distortion;
[0100] Finally, for a certain original dynamic segment, three corresponding distorted dynamic segments are obtained. For each original dynamic segment C [i] , the distorted dynamic segments are denoted as C Spa [i] (spatial dimension), C Tem [i] (temporal dimension) and C ST [i] (spatio-temporal dimension), and the corresponding weights are , and . Experiments show that the spatial weight is set to 0.5, the temporal weight is 0.3, and the spatio-temporal weight When the weight is 0.2, the overall performance of FineQuest is optimal, which indicates that more attention is paid to spatial features in the multi-level decoding process, while the temporal and spatio-temporal elements are moderately balanced, which helps to optimize the decoding ability of the model; too high spatio-temporal weight will cause the accuracy of the model on difficult problems to decrease significantly, because too much emphasis on spatio-temporal interaction will introduce additional complexity and noise, thereby affecting the decoding effect.
[0101] The correlation calculation module performs feature extraction on a certain original dynamic segment and the corresponding three distorted dynamic segments, and performs feature extraction on the input question text; the correlation of the four dynamic segment features and the question text features is calculated respectively, and the obtained correlation is weighted to obtain the correlation score of the original dynamic segment and the question text;
[0102] The specific process is: using the encoder in the known multi-modal large model to extract features of each dynamic segment and question text, and output the correlation probability distribution of each dynamic segment feature and question text feature, using the correlation probability distribution of the original dynamic segment feature and the question text feature, subtracting the correlation probability distribution of each weighted distorted dynamic segment feature and the question text feature, and the corresponding weight is 、 and ; the result obtained by subtraction is decoded to obtain the correlation score of the original dynamic segment and the question text.
[0103] In this embodiment, the multi-modal large model uses Video-LlaVA or LLaVA-Next-Video.
[0104] The sorting module sorts the correlation scores of each original dynamic segment and question text to obtain the original dynamic segment most related to the question text as the key segment.
[0105] The fine-grained matcher based on the sports knowledge graph includes a key segment description module, a multi-level matching module, and a knowledge extraction module.
[0106] The key segment description module uses a known multi-modal large model to describe the key segment to obtain a key segment description text. The multi-modal large model uses Video-LlaVA or LLaVA-Next-Video.
[0107] The multi-level matching module calculates the following three cosine similarities:
[0108] (1) The cosine similarity between the key segment description text and the description text of each instance of the corresponding sports project in the sports knowledge graph;
[0109] (2) cosine similarity between the key segment and the video of each instance of the corresponding sports project in the sports knowledge graph;
[0110] (3) cosine similarity between the key segment description text and the scene information description text of the video of each instance of the corresponding sports project in the sports knowledge graph;
[0111] The sports knowledge graph includes a plurality of sports projects, each sports project has a plurality of instances, each instance has a description text, a video and scene information description text of the video;
[0112] The three kinds of cosine similarities corresponding to each instance in the corresponding sports project calculated are weighted and summed to obtain the similarity of the key segment and each instance of the corresponding sports project, and then the instance with the highest similarity is obtained.
[0113] The knowledge extraction module extracts the knowledge corresponding to the instance with the highest similarity and outputs it to the known multi-modal large model, and answers the question through the multi-modal large model; the knowledge includes the professional action name and its text description corresponding to the instance. The multi-modal large model uses Video-LLaVA or LLaVA-Next-Video.
[0114] Experimental verification:
[0115] In order to comprehensively verify the performance and effectiveness of FineQuest, the following aspects are systematically designed: selection of baseline model, setting of evaluation dataset and experimental index, and details of model deployment:
[0116] 1. Baseline model
[0117] The core of FineQuest design is to improve the performance of multi-modal video question answering through multi-model collaboration, combining reactive reasoning and reflective reasoning. Its architecture consists of multi-modal large model (MLLM), visual language model (VLM) and language model (LLM), which are used to realize fast response, video semantic generation and multi-step reasoning of complex problems. In the experiment, in order to ensure the fairness and clarity of the experiment, a single MLLM is selected to complete the three tasks at the same time, and FineQuest is applied to two 7B-scale MLLMs: Video-LLaVA and LLaVA-Next-Video. In addition, the results of the following models are provided as a reference: VideoChat2, Video-ChatGPT and Tarsier. The experimental results of these models provide a horizontal comparison for the performance of FineQuest, further verifying its advantages in the sports video question answering task.
[0118] 2. Evaluation dataset and experimental index
[0119] To comprehensively evaluate the performance of FineQuest, the experimental design covers both sports video question answering and general video question answering scenarios, aiming to verify its performance in specific domains and general tasks.
[0120] Sports video question answering benchmark. Sports video question answering is the core application scenario of FineQuest. For this purpose, the following datasets are used for evaluation: (1) Gym-QA: Gymnastics Question Answering Dataset, covering various gymnastics movements and related questions, focusing on evaluating the model's understanding of complex movement actions; (2) Diving-QA: Diving Question Answering Dataset, focusing on the analysis and understanding of diving movements, testing the model's ability to handle high-dynamic scenarios; (3) SPORTU: Video Question Answering Dataset containing basketball, football, ice hockey, tennis, baseball, badminton, and volleyball, covering a wide range of sports scenarios, and can comprehensively evaluate the model's domain adaptability.
[0121] General video question answering benchmark. To verify that the improvements of FineQuest in sports scenarios do not come at the expense of general question answering capabilities, this paper also evaluates on the following general video question answering datasets: (1) MSVD-QA: Video Question Answering Dataset based on the MSVD dataset, testing the model's understanding of short videos; (2) MSR-VTT-QA: A standard dataset widely used for general video question answering, covering a variety of video scenarios and question types; (3) TGIF-QA: A dataset focused on GIF dynamic video question answering, testing the model's reasoning ability for short dynamic videos. (4) ActivityNet-QA: Video Question Answering Dataset based on the ActivityNet dataset, covering long and complex video scenarios, evaluating the model's long video processing capabilities.
[0122] Figure 3 、 Figure 4 and Figure 5 respectively show the experimental results of FineQuest on the Gym-QA and Diving-QA full set, action set, and SPORTU benchmark dataset, further verifying the significant advantages of FineQuest in sports video question answering tasks. The bold values in the table represent the best performance, which indicates the corresponding percentage improvement (i.e., improvement amplitude) relative to the baseline model. ↑ indicates that the higher the value, the better.
[0123] The results show that:
[0124] (1) FineQuest significantly improves the performance of the base model in sports video question answering tasks. On the Gym-QA and Diving-QA datasets, FineQuest achieves significant performance improvements in all test dimensions, such as event-level, subset-level, and element-level. In particular, in the action analysis task, the overall accuracy of FineQuest (such as 70.4% and 63.6%) is much higher than the baseline model, fully demonstrating its strong ability in capturing complex action semantics. On the SPORTU dataset, FineQuest also performs well, especially on medium and difficult questions, with significant performance improvements. For example, based on Video-LLaVA, FineQuest's accuracy on medium difficulty questions increased from 55.9% to 69.6%, and on difficult questions from 27.5% to 59.0%, indicating its robustness in handling complex scenarios and questions.
[0125] (2) As the complexity of the question increases, FineQuest's performance improvement is more obvious. For example, in the Gym-QA and Diving-QA datasets, FineQuest's accuracy on difficult questions increased by 382.6% (LLaVA-Next-Video) and 418.7% (Video-LLaVA) compared to the baseline model. This significant improvement reflects the effectiveness of FineQuest's dual reasoning mode. By combining visual information with action semantic analysis, FineQuest can better understand the context of complex questions and provide accurate answers. In addition, this result highlights the importance of benchmark tests with diverse difficulty levels. FineQuest's superior performance shows that relying solely on simple question evaluation may not fully reflect the actual ability of the model, and introducing medium and difficult questions can more accurately reveal the potential of the model.
[0126] (3) FineQuest's outstanding performance not only lies in the improvement of accuracy, but also includes its consistency in different test dimensions. Whether it is action analysis or scene understanding, FineQuest has shown strong generalization ability and robustness. For example, in the SPORTU dataset, FineQuest's performance on simple, medium, and difficult questions is better than other models, especially on difficult questions, where its performance almost doubles. This multi-dimensional advantage further proves the effectiveness of FineQuest's design in dealing with diverse sports video question answering tasks.
[0127] The performance in the general video question answering benchmark is also important because it reflects the adaptability and robustness of the model in non-specific domain tasks. To verify that FineQuest improves the sports video question answering ability while its generalization is not weakened, the performance of FineQuest on four widely used open general video question answering benchmarks (MSVD-QA, MSR-VTT-QA, TGIF-QA and ActivityNet-QA) is evaluated in the zero-shot setting, and the results are shown in Figure 6
[0128] The experimental results show that FineQuest outperforms VideoTree on all benchmark tests and achieves significant performance improvement on two base models (LLaVA-Next-Video and Video-LLaVA). For example, on the MSVD-QA dataset, the performance of FineQuest is improved from 75.9% of VideoTree to 77.2% (based on LLaVA-Next-Video) and from 73.2% of VideoTree to 73.8% (based on Video-LLaVA). Similarly, on the MSR-VTT-QA, TGIF-QA and ActivityNet-QA datasets, FineQuest also shows stable improvement. These results fully demonstrate the effectiveness of the structured inference process of FineQuest. By introducing the reactive inference agent, FineQuest can better utilize the intrinsic capabilities of the base model, thereby achieving stronger general performance in different datasets and tasks. This improvement not only demonstrates the outstanding performance of FineQuest in specific domain tasks (such as sports video question answering), but also ensures its applicability and reliability in more general question answering tasks.
[0129] Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments without departing from the principles and spirit of the present application within the scope of the present application.
Claims
1. A knowledge-enhanced sports video understanding method based on large model dual-mode inference, characterized by: The method comprises the following steps: Step 1: obtaining a sports video and a question text that need to be asked; Step 2: inputting the sports video, the question text and a prompt word into a reactive reasoning intelligent agent, the reactive reasoning intelligent agent classifying the question according to the question text and the prompt word to determine whether the question is a simple question or a complex question, if the question is a simple question, entering step 3, if the question is a complex question, entering step 4; Step 3: the reactive reasoning intelligent agent answering the question according to the input sports video; Step 4: inputting the sports video and the question text into a constructed deliberative reasoning intelligent agent; The deliberative reasoning intelligent agent comprises a dynamic motion segmenter, a key segment selector and a fine-grained matcher based on a sports knowledge graph; The dynamic motion segmenter adaptively reduces background interference and segments the sports video into multiple dynamic segments according to motion intensity; The key segment selector identifies a key segment related to the question in the spatial, temporal and spatio-temporal dimensions by using a multi-level contrast decoding strategy; The fine-grained matcher based on the sports knowledge graph realizes accurate matching of the key segment and the question by means of structured information extracted from the sports knowledge graph, thereby outputting corresponding knowledge.
2. The knowledge enhanced sports video understanding method based on large model dual-mode inference according to claim 1, characterized in that: The prompt word content adopted includes: 1) relevance of the question and the video; 2) question type; 3) reasoning requirement; 4) whether external knowledge is needed.
3. The knowledge enhanced sports video understanding method based on large model dual-mode inference according to claim 1, characterized in that: The dynamic motion segmenter comprises a video content segmentation module, a motion intensity extraction module and a segment segmentation module; The video content segmentation module segments each frame of image in the sports video and identifies a person subject and a background in each frame of image; The motion intensity extraction module obtains a plurality of continuous frames of image in the sports video through a sliding window, performs optical flow analysis on the frames of image one by one to obtain motion intensity of the person subject action in each sliding window, moves the sliding window along the time axis direction of the sports video, obtains motion intensity of the person subject action in each sliding window, thereby obtaining a motion intensity change curve of the person subject action along the time axis direction of the sports video; The segment segmentation module segments the sports video into multiple dynamic segments according to a set threshold of motion intensity according to the motion intensity change curve of the person subject action.
4. The knowledge enhanced sports video understanding method based on large model dual-mode inference according to claim 3, characterized in that: The video content segmentation module is realized by using a SAM2 model.
5. The knowledge enhanced sports video understanding method based on large model dual-mode inference according to claim 1, characterized in that: The key segment selector comprises a multi-level contrast decoding module, a correlation calculation module and a sorting module; The multi-level contrast decoding module respectively performs distortion processing on an original dynamic segment from the spatial, temporal and spatio-temporal dimensions; The correlation calculation module respectively extracts features of a certain original dynamic segment and the three distortion-processed dynamic segments corresponding to the original dynamic segment, extracts features of the input question text, respectively calculates correlations of the four dynamic segment features and the question text features, and performs weighted processing on the obtained correlations to obtain a correlation score of the original dynamic segment and the question text; The sorting module sorts the correlation scores of the original dynamic segments and the question text to obtain an original dynamic segment most relevant to the question text as a key segment.
6. The knowledge enhanced sports video understanding method based on large model dual-mode inference according to claim 5, characterized in that: The specific process of distortion processing from the three dimensions of space, time and space-time is as follows: The distortion processing from the spatial dimension refers to adding Gaussian noise to each frame in the original dynamic segment; The distortion processing from the time dimension refers to adjusting the duration of each frame while keeping the order of the frames unchanged; The distortion processing from the space-time dimension refers to combining the distortion from the spatial dimension and the distortion from the time dimension.
7. The knowledge enhanced sports video understanding method based on large model dual-mode inference according to claim 5, characterized in that: The weight corresponding to the dynamic segment after the distortion processing from the three dimensions of space, time, and space-time is , and , wherein is a space weight, taking a value of 0.5, is a time weight, taking a value of 0.3, is a space-time weight, taking a value of 0.
2.
8. The knowledge-enhanced sports video understanding method based on large model dual-mode inference according to claim 5, characterized in that: The specific process of obtaining the relevance score of the original dynamic segment and the question text is: using the encoder in the multi-modal large model to extract features of each dynamic segment and the question text, and outputting the relevance probability distribution of the features of each dynamic segment and the features of the question text, using the relevance probability distribution of the features of the original dynamic segment and the features of the question text, subtracting the relevance probability distribution of the weighted distorted dynamic segment features and the features of the question text, and the corresponding weight is , and wherein is a spatial weight, is a temporal weight, is a spatio-temporal weight; and decoding the result obtained after subtraction to obtain the relevance score of the original dynamic segment and the question text.
9. The knowledge enhanced sports video understanding method based on large model dual-mode inference according to claim 1, characterized in that: The fine-grained matcher based on the sports knowledge graph comprises a key segment description module, a multi-level matching module and a knowledge extraction module; The key segment description module uses a multi-modal large model to describe the key segment to obtain a key segment description text; The multi-level matching module calculates the following three cosine similarities: (1) The cosine similarity between the key segment description text and the description text of each instance of the corresponding sports project in the sports knowledge graph; (2) The cosine similarity between the key segment and the video of each instance of the corresponding sports project in the sports knowledge graph; (3) The cosine similarity between the key segment description text and the scene information description text of the video of each instance of the corresponding sports project in the sports knowledge graph; The three kinds of cosine similarities corresponding to each instance of the corresponding sports project are weighted and summed to obtain the similarity of the key segment and each instance of the corresponding sports project, and then the instance with the highest similarity is obtained; The knowledge extraction module extracts the knowledge corresponding to the instance with the highest similarity and outputs it to the multi-modal large model, and answers the question through the multi-modal large model.
10. The knowledge-enhanced sports video understanding method based on large model dual-mode inference according to claim 1, characterized in that: The reaction formula reasoning intelligent agent adopts a Video-LLaVA large model or a LLaVA-Next-Video large model.