Video understanding question and answer method and device and storage medium

By building an image vector library, a knowledge graph, and a text vector library, and combining it with the multimodal data joint reasoning of a large language model, we have solved the problem of low accuracy in video question answering in existing technologies, and achieved a deep understanding of video content and accurate answers.

CN120821871APending Publication Date: 2025-10-21BEIJING JIZHI DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510805662.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing video retrieval and question-answering technologies are unable to effectively analyze the spatiotemporal structure and semantic relationships of videos when dealing with movies with complex plots and multiple themes, resulting in low question-answering accuracy.

Method used

By acquiring video data, extracting key frame images, building an image vector library, knowledge graph, and text vector library, and using a large language model to perform joint reasoning on multimodal data, answers are generated.

Benefits of technology

It improves the accuracy of question and answer, can more comprehensively understand user questions, capture important time nodes and plot turning points in the video, enhances the pertinence and discrimination of image features, and integrates the spatiotemporal structure of video semantic knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821871A_ABST
    Figure CN120821871A_ABST
Patent Text Reader

Abstract

The invention discloses a video understanding question answering method and device and a storage medium, and belongs to the technical field of question answering. The method comprises the steps of obtaining video data and a problem of a user for the video data; extracting a plurality of key frame images according to feature differences among video frames in the video data; extracting image features of the key frame images to construct an image vector library, and extracting content information of the key frame images to construct a knowledge graph and a text vector library; and generating answers to the questions based on the image vector library, the knowledge graph and the text vector library through a large language model. By constructing the image vector library, the knowledge graph and the text vector library, not only is the video content quantitatively described from the visual perspective, but also the semantic relationship and the logic structure of the video content are mined through the knowledge graph, the video semantic knowledge is more comprehensively and deeply mined and integrated, and the multi-modal data are combined for joint reasoning, so that the video semantic knowledge is more comprehensively and deeply mined. User questions can be more comprehensively understood, and the question and answer accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of question-answering technology, and in particular relates to a video understanding question-answering method, device and storage medium. Background Art

[0002] With the rapid development of the internet and multimedia technologies, video content has become a crucial vehicle for information dissemination and entertainment consumption. As the depth and breadth of user consumption of video content continues to increase, understanding video content and enabling intelligent question-answering based on this understanding have become highly valuable research areas. For example, in education, accurate understanding of video content can help students quickly grasp key knowledge points.

[0003] Among related technologies, video retrieval and question-answering technologies mostly use fixed feature extraction and fusion methods. When processing movies with complex plots and multiple themes, they may not be able to effectively analyze the spatiotemporal structure and semantic relationships of the video, resulting in low accuracy in question answering and making it difficult to meet users' needs for in-depth understanding of video content and question answering. Summary of the Invention

[0004] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a video understanding question answering method, device and storage medium to improve the accuracy of question answering.

[0005] In a first aspect, the present application provides a video understanding question answering method, comprising:

[0006] Obtaining video data and user questions regarding the video data;

[0007] extracting a plurality of key frame images according to feature differences between video frames in the video data;

[0008] Extracting image features of the key frame images to construct an image vector library, and extracting content information of the key frame images to construct a knowledge graph and a text vector library;

[0009] An answer to the question is generated by a large language model based on the image vector library, the knowledge graph, and the text vector library.

[0010] According to the video understanding question-answering method of the present application, video data and user questions regarding the video data are obtained; multiple key frame images are extracted based on the feature differences between video frames in the video data; image features of the key frame images are extracted to construct an image vector library, and content information of the key frame images is extracted to construct a knowledge graph and a text vector library; and answers to the questions are generated based on the image vector library, the knowledge graph and the text vector library through a large language model. The embodiment of the present application extracts key frame images by analyzing the feature differences between video frames, which can capture important time nodes and plot turning points in the video, reducing the problem of missing key information due to fixed feature extraction in traditional methods, constructing an image vector library, a knowledge graph and a text vector library, not only quantitatively describing the video content from a visual perspective, but also mining the semantic relationship and logical structure of the video content through the knowledge graph, more comprehensively and deeply mining and integrating video semantic knowledge, and utilizing the powerful semantic understanding and generation capabilities of the large language model, combined with the above-mentioned multimodal data joint reasoning, to more comprehensively understand user questions and improve the accuracy of question-answering.

[0011] According to one embodiment of the present application, extracting image features of the key frame image to construct an image vector library includes:

[0012] Performing object detection on the key frame image to obtain an object detection result; the object detection result includes object coordinates and object category;

[0013] Inputting the object detection result and the key frame image into an image encoder to extract image features;

[0014] Image features of different key frame images are hierarchically stored according to the scene type, object category and time stamp of the key frame images to construct an image vector library.

[0015] In this embodiment, by performing object detection on key frame images, objects in the image and their coordinates and categories can be identified, providing a semantic anchor for image feature extraction, enhancing the pertinence and discrimination of features. This feature extraction method that combines semantic information can more comprehensively reflect the visual content and semantic meaning of the image, and improve the quality of image features; image features are stored in layers according to the scene type, object category and timestamp of the key frame image, which not only facilitates rapid retrieval and matching, but also better reflects the spatiotemporal structure and semantic hierarchy of the video content, providing richer data support for the understanding of complex video content and question-answering tasks.

[0016] According to one embodiment of the present application, extracting content information of the key frame image to construct a knowledge graph includes:

[0017] Extracting content information of the key frame image through an image understanding model to obtain image description text;

[0018] Inputting the image description text into the large language model to extract knowledge triples of the image description text;

[0019] Using the object detection result of the key frame image as information supplement, and constructing a knowledge subgraph of the key frame image according to the knowledge triples;

[0020] The knowledge graph is obtained by fusing the knowledge subgraphs of different key frame images through the large language model.

[0021] In this embodiment, the key frame image content information is extracted through the image understanding model to generate image description text, realizing the initial transformation from visual information to natural language description. With the help of the large language model, knowledge triples are extracted from the image description text, the core semantic relationships in the text are fully mined, the image information is further abstracted into structured knowledge, and the object detection results are incorporated as supplementary information, which enriches the details and accuracy of the knowledge subgraph, so that the knowledge subgraph can more comprehensively reflect the key frame image content. The knowledge subgraphs of each key frame image are fused through the large language model, and the scattered knowledge is integrated into a complete knowledge graph to form a systematic global knowledge graph. The semantic associations between different key frames in the video content are established, which can more comprehensively and deeply mine and integrate video semantic knowledge than traditional methods.

[0022] According to one embodiment of the present application, the method further includes:

[0023] Performing time series analysis on the knowledge subgraphs corresponding to the consecutive key frame images to obtain analysis results; the analysis results include event chain recognition results, event trend analysis results, and abnormal event detection results;

[0024] extracting event features and event relationships from the analysis results;

[0025] According to the event features and event relationships, event nodes are constructed in the knowledge subgraph and association edges between events and entities in the knowledge subgraph are established.

[0026] In this embodiment, by performing time series analysis on the knowledge subgraph corresponding to continuous key frame images, it is possible to identify event chains, analyze event development trends and detect abnormal events from the time dimension, so that the originally static knowledge subgraph has time context and dynamic evolution characteristics, extract event characteristics and event relationships in the analysis results, and further refine the core elements and association logic of events in the video content. Based on these characteristics and relationships, event nodes are constructed in the knowledge subgraph and association edges with entities are established, and event information in the time dimension is deeply integrated into the knowledge graph system to achieve spatiotemporal fusion of local knowledge graphs. It can not only present video content more comprehensively and three-dimensionally, but also provide users with richer knowledge support for timeline-based video content questions and answers.

[0027] According to one embodiment of the present application, fusing knowledge subgraphs of different keyframe images using the large language model to obtain a knowledge graph includes:

[0028] Add timestamps to entities and relationships in different knowledge subgraphs, and merge entities and relationships with event associations through a large language model to obtain a knowledge graph; the merging includes:

[0029] For repeated parts of different knowledge subgraphs, retain the latest or most credible information;

[0030] For contradictory information in different knowledge subgraphs, the time series and the logical relationship between elements in the knowledge subgraphs are combined to remove erroneous content.

[0031] In this embodiment, by adding timestamps to entities and relationships in different knowledge subgraphs, the knowledge graph can better reflect the changes and development of video content in the time series. By merging entities and relationships with event associations through a large language model, it is possible to integrate scattered knowledge and establish closer semantic connections. For repeated parts, the latest or most credible information is retained, which reduces redundant information interference and improves the timeliness of information in the knowledge graph. For contradictory information, the erroneous content is removed by combining the time series and the logical relationship between elements. Through deep verification and screening of knowledge, the accuracy of information in the knowledge graph is improved.

[0032] According to one embodiment of the present application, generating an answer to the question based on the image vector library, the knowledge graph, and the text vector library using a large language model includes:

[0033] Perform semantic analysis on the problem and decompose it into multiple sub-problems;

[0034] According to the plurality of sub-questions, first search results corresponding to the sub-questions are retrieved from the knowledge graph, second search results corresponding to the sub-questions are retrieved from the text vector library, and third search results corresponding to the sub-questions are retrieved from the image vector library;

[0035] Analyzing the first and second search results using the large language model, and if the analysis shows that the first and second search results cannot serve as answers to corresponding sub-questions, reasoning on the first, second, and third search results using the large visual language model to obtain answers to corresponding sub-questions;

[0036] The answers to the sub-questions are integrated to obtain the answer to the question.

[0037] In this embodiment, by semantically parsing the question and breaking it down into sub-questions, complex questions can be refined, making subsequent retrieval and reasoning more targeted. The answers to each sub-question are retrieved based on the knowledge graph, text vector library, and image vector library respectively, and relevant data is obtained from multiple dimensions of semantic knowledge, text information, and image features. When the retrieval results of the knowledge graph and text vector library analyzed by the large language model cannot meet the requirements of answering the sub-questions, the visual large language model is introduced to perform reasoning in combination with the retrieval results of the image vector library. With the help of the multimodal fusion capability of the visual large language model, the shortcomings of single modality information are compensated, and accurate answers to users' questions are achieved.

[0038] According to one embodiment of the present application, the method further includes:

[0039] When it is analyzed that the first search result and the second search result can serve as answers to the corresponding sub-questions, the first search result and the second search result are integrated into answers to the corresponding sub-questions through the large language model.

[0040] In this embodiment, when the large language model determines that the retrieval results of the knowledge graph and text vector library are sufficient to answer the sub-question, the answers to the sub-questions are directly integrated to generate the answers. This reduces the reasoning process of the large visual language model, thereby reducing the consumption of computing resources and processing time, and improving the efficiency of question answering.

[0041] In a second aspect, the present application provides a video understanding question-answering device, comprising:

[0042] An acquisition module, configured to acquire video data and user questions regarding the video data;

[0043] An extraction module, configured to extract a plurality of key frame images based on feature differences between video frames in the video data;

[0044] A construction module, configured to extract image features of the key frame images to construct an image vector library, and to extract content information of the key frame images to construct a knowledge graph and a text vector library;

[0045] A generation module is used to generate an answer to the question based on the image vector library, the knowledge graph and the text vector library through a large language model.

[0046] According to the video understanding question-answering device of the present application, video data and user questions regarding the video data are obtained; multiple key frame images are extracted based on the feature differences between video frames in the video data; image features of the key frame images are extracted to construct an image vector library, and content information of the key frame images is extracted to construct a knowledge graph and a text vector library; and answers to the questions are generated based on the image vector library, the knowledge graph and the text vector library through a large language model. The embodiment of the present application extracts key frame images by analyzing the feature differences between video frames, which can capture important time nodes and plot turning points in the video, reducing the problem of missing key information due to fixed feature extraction in traditional methods, constructing an image vector library, a knowledge graph and a text vector library, not only quantitatively describing the video content from a visual perspective, but also mining the semantic relationship and logical structure of the video content through the knowledge graph, more comprehensively and deeply mining and integrating video semantic knowledge, and utilizing the powerful semantic understanding and generation capabilities of the large language model, combined with the above-mentioned multimodal data joint reasoning, to more comprehensively understand user questions and improve the accuracy of question-answering.

[0047] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the video understanding question-answering method as described in the first aspect above is implemented.

[0048] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video understanding question-answering method as described in the first aspect above.

[0049] In a fifth aspect, the present application provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the video understanding question-answering method described in the first aspect above.

[0050] In a sixth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the video understanding question-answering method as described in the first aspect above.

[0051] The above one or more technical solutions in the embodiments of the present application have at least one of the following technical effects:

[0052] According to the video understanding question-answering method of the present application, video data and user questions regarding the video data are obtained; multiple key frame images are extracted based on the feature differences between video frames in the video data; image features of the key frame images are extracted to construct an image vector library, and content information of the key frame images is extracted to construct a knowledge graph and a text vector library; and answers to the questions are generated based on the image vector library, the knowledge graph and the text vector library through a large language model. The embodiment of the present application extracts key frame images by analyzing the feature differences between video frames, which can capture important time nodes and plot turning points in the video, reducing the problem of missing key information due to fixed feature extraction in traditional methods, constructing an image vector library, a knowledge graph and a text vector library, not only quantitatively describing the video content from a visual perspective, but also mining the semantic relationship and logical structure of the video content through the knowledge graph, more comprehensively and deeply mining and integrating video semantic knowledge, and utilizing the powerful semantic understanding and generation capabilities of the large language model, combined with the above-mentioned multimodal data joint reasoning, to more comprehensively understand user questions and improve the accuracy of question-answering.

[0053] Furthermore, in some embodiments, by performing object detection on keyframe images, objects in the image and their coordinates and categories can be identified, providing a semantic anchor for image feature extraction, enhancing the pertinence and discrimination of features. This feature extraction method that combines semantic information can more comprehensively reflect the visual content and semantic meaning of the image, and improve the quality of image features; hierarchical storage of image features according to the scene type, object category and timestamp of the keyframe image not only facilitates rapid retrieval and matching, but also better reflects the spatiotemporal structure and semantic hierarchy of the video content, providing richer data support for the understanding of complex video content and question-answering tasks.

[0054] Furthermore, in some embodiments, the key frame image content information is extracted through the image understanding model to generate image description text, thereby achieving a preliminary transformation from visual information to natural language description. With the help of a large language model, knowledge triples are extracted from the image description text, and the core semantic relationships in the text are fully mined. The image information is further abstracted into structured knowledge, and the object detection results are incorporated as supplementary information, which enriches the details and accuracy of the knowledge subgraph, so that the knowledge subgraph can more comprehensively reflect the key frame image content. The knowledge subgraphs of each key frame image are fused through the large language model, and the scattered knowledge is integrated into a complete knowledge graph to form a systematic global knowledge graph. The semantic associations between different key frames in the video content are established, which can more comprehensively and deeply mine and integrate video semantic knowledge than traditional methods.

[0055] Furthermore, in some embodiments, by performing time series analysis on the knowledge subgraph corresponding to continuous key frame images, it is possible to identify event chains, analyze event development trends and detect abnormal events from the time dimension, so that the originally static knowledge subgraph has time context and dynamic evolution characteristics, extract event characteristics and event relationships in the analysis results, and further refine the core elements and association logic of events in the video content. Based on these characteristics and relationships, event nodes are constructed in the knowledge subgraph and association edges with entities are established, and event information in the time dimension is deeply integrated into the knowledge graph system to achieve spatiotemporal fusion of local knowledge graphs. It can not only present video content more comprehensively and three-dimensionally, but also provide users with richer knowledge support for timeline-based video content questions and answers.

[0056] Furthermore, in some embodiments, by adding timestamps to entities and relationships in different knowledge subgraphs, the knowledge graph can better reflect the changes and development of video content in a time series. By merging entities and relationships with event associations through a large language model, scattered knowledge can be integrated to establish closer semantic connections. For repeated parts, the latest or most credible information is retained, which reduces redundant information interference and improves the timeliness of information in the knowledge graph. For contradictory information, erroneous content is removed by combining time series and logical relationships between elements. Through deep verification and screening of knowledge, the accuracy of information in the knowledge graph is improved.

[0057] Furthermore, in some embodiments, by semantically parsing the question and breaking it down into sub-questions, complex questions can be refined, making subsequent retrieval and reasoning more targeted. The answers to each sub-question are retrieved based on the knowledge graph, text vector library, and image vector library respectively, and relevant data are obtained from multiple dimensions of semantic knowledge, text information, and image features. When the retrieval results of the knowledge graph and text vector library analyzed by the large language model cannot meet the requirements of answering the sub-questions, the visual large language model is introduced to perform reasoning in combination with the retrieval results of the image vector library. With the help of the multimodal fusion capability of the visual large language model, the shortcomings of single modality information are compensated, and accurate answers to users' questions are achieved.

[0058] Furthermore, in some embodiments, when the large language model determines that the retrieval results of the knowledge graph and text vector library are sufficient to answer the sub-question, the answers to the sub-questions are directly integrated to generate the answers. This reduces the reasoning process of the large visual language model, thereby reducing the consumption of computing resources and processing time, and improving the efficiency of question answering.

[0059] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The above and / or additional aspects and advantages of the present application will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:

[0061] Figure 1 Schematic diagram of the process of video comprehension and question answering provided in the embodiment of the present application;

[0062] Figure 2 This is a schematic diagram of the construction process of the image vector library, knowledge graph, and text vector library provided in the embodiment of the present application;

[0063] Figure 3 Schematic diagram of the structure of the video understanding question-answering device provided in an embodiment of the present application;

[0064] Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0065] The following will be combined with the accompanying drawings in the embodiments of this application to clearly describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0066] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0067] The following, in conjunction with the accompanying drawings, describes in detail the video understanding question-answering method, device, and storage medium provided in the embodiments of the present application through specific embodiments and their application scenarios.

[0068] Among them, the video understanding question answering method can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.

[0069] The terminal includes, but is not limited to, a portable communication device such as a mobile phone or tablet computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad).

[0070] In the following embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and a joystick.

[0071] The video understanding question-answering method provided in the embodiments of the present application may be executed by an electronic device or a functional module or functional entity in the electronic device that can implement the video understanding question-answering method. The electronic devices mentioned in the embodiments of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices. The video understanding question-answering method provided in the embodiments of the present application is described below using an electronic device as an example of the execution subject.

[0072] like Figure 1 As shown, the video comprehension question answering method includes: step 110, step 120, step 130 and step 140.

[0073] Step 110: Obtain video data and user questions regarding the video data.

[0074] In the embodiments of the present application, video data is multimedia content consisting of continuous video frames, including multimodal information such as images, audio, and subtitles, and can be presented in various forms such as movies, educational courses, news reports, documentaries, etc. The format of the video data can be any video file format, such as MP4, AVI, MOV, etc.

[0075] Based on the video content they watch or their interest in, users can ask questions that require answers from the video. These questions can involve plot comprehension, knowledge point extraction, character relationship analysis, and more. User questions can be text in natural language, such as "Who is the protagonist in the video?" or "Where is the plot turning point in the video?"

[0076] In some embodiments, video data uploaded by a user through an upload button on a user interface may be obtained, video data may be obtained by specifying a path to a video file in a local storage device, or video data may be obtained through other means, which are not limited in the embodiments of the present application. Questions may be obtained by user input in a text box on the user interface, or by voice-to-text conversion of questions entered by the user, or through other means, which are not limited in the embodiments of the present application.

[0077] Step 120: extract multiple key frame images based on feature differences between video frames in the video data.

[0078] Keyframe images are representative and important frames extracted from a video sequence. They capture crucial information such as important time points, plot turning points, scene changes, and action details. For example, in a movie, keyframes can mark plot climaxes, important dialogue moments, or scene transitions. In sports videos, keyframes can capture athletes' spectacular moves, such as goals and spikes. By selecting keyframe images, we can reduce redundant data while preserving the core semantic information of the video.

[0079] The process of extracting keyframe images requires comprehensive consideration of the dynamic nature of the video content, scene changes, and action details. To improve the accuracy of keyframe image extraction, this application uses an improved adaptive K-Means++ clustering algorithm that automatically adjusts the keyframe extraction strategy based on the complexity of the video content. Specifically, it dynamically determines the number and location of keyframes by analyzing the differences in visual features between video frames.

[0080] Specifically, visual features can be extracted from video frames in video data. These visual features can include color histograms, texture features, edge features, and optical flow features, which can reflect the visual differences between video frames. For example, color histograms can capture the changes in color distribution in a scene, while optical flow features can reflect the direction and speed of movement of objects or people. By quantifying these features, each frame of the image is converted into a feature vector that can be understood and compared by computers.

[0081] The traditional K-Means++ algorithm relies heavily on initial values ​​when determining cluster centers. The improved K-Means++ algorithm optimizes its center selection strategy, resulting in more stable and accurate clustering results. During algorithm execution, the number and placement of keyframes are dynamically determined based on the differences in visual features between video frames. For scenes with rich action and rapid changes, such as intense sports sequences, which contain a wealth of action details and require more keyframes to capture, the number of extracted keyframes can be increased. For example, in a football match, the athletes' rapid running, passing, and shooting movements all require more keyframes to fully capture them. In relatively static scenes, such as static distant shots in landscape videos, where there is less change and more redundant frames, the number of extracted keyframes can be reduced, improving processing efficiency without losing the core content.

[0082] In some embodiments, after extracting keyframe images, each keyframe image can be timestamped to record its specific location within the video data. The timestamps can be used to determine the temporal order and context of the keyframe images within the video. For example, the timestamps can be used to determine whether a keyframe image is at the beginning, middle, or end of the video, thereby better understanding the overall structure and plot development of the video.

[0083] Step 130: extract image features of the key frame images to construct an image vector library, and extract content information of the key frame images to construct a knowledge graph and a text vector library.

[0084] Image features are quantitative expressions of image visual information, which can capture key information such as color, texture, shape, and object position in the image. In some embodiments, the image features of key frame images can be extracted by using traditional algorithms such as SIFT (Scale-invariant Feature Transform), SURF (Speeded-Up Robust Features), or convolutional neural network (CNN) methods based on deep learning. By storing the feature vectors corresponding to each key frame image, an image vector library is formed, and the Faiss library can be used to build an efficient vector index structure to support fast nearest neighbor search and improve retrieval efficiency.

[0085] In addition to image features, keyframe images also contain rich semantic content information. To further explore this information, we can extract the content information of keyframe images to construct a knowledge graph. A knowledge graph is a structured semantic knowledge base that represents knowledge in the form of a graph through a combination of entities, relationships, and attributes. In video content, entities can be people, objects, scenes, etc., and relationships can be interactions between entities, location relationships, or temporal sequences.

[0086] In some embodiments, target detection, image segmentation and other technologies can be used to parse the content information of key frame images, identify entity objects in the image, such as people, objects, scenes, etc., and classify and label these entities. Then analyze the spatial relationship between entities (such as "on the left" and "above"), action relationship (such as "attack" and "hug"), causal relationship (such as "because it rains, so open an umbrella"), etc. These entities and their relationships are stored in the form of triples (entity 1, relationship, entity 2), and the knowledge graph is gradually constructed. For example, in the key frame images of a movie, the two entities "protagonist" and "car" are identified, as well as the relationship "driving" between them, forming a triple such as (protagonist, driving, car). With the analysis of multiple key frame images, the knowledge graph is continuously enriched and improved.

[0087] In some embodiments, an image understanding model can be used to convert the content information in the keyframe image into a natural language description to generate an image description. For example, for an image containing a person, a scene, and an action, the image understanding model can generate an image description such as "a boy playing soccer in a park." In some embodiments, the image understanding model used can be a BLIP (Bootstrapping Language-Image Pre-training) model.

[0088] After obtaining the image description text corresponding to each keyframe image, natural language processing technology can be used to extract entities and relationships from the image description text and construct a knowledge graph. For example, from the description text "A person running on the beach", the entities "person" and "beach" and the relationship "running" can be extracted. These entities and relationships can be organized into nodes and edges in the knowledge graph, forming a structured knowledge network.

[0089] In some embodiments, the text vector library is constructed by vectorizing the text content. Figure 2As shown, the image description text corresponding to each key frame image can be converted into a vector representation through a text embedding model, and the text feature vectors extracted from all key frame images can be integrated to build a text vector library. For example, for the image description text "A person is running on the beach", the image description text can be converted into a vector and stored in a text vector library. In addition, the Faiss library can be used to build an efficient vector index structure to support fast nearest neighbor search and improve retrieval efficiency. Among them, the text embedding model is the core tool for realizing text vectorization. It is based on deep learning architectures, such as variants of the BERT (Bidirectional Encoder Representations from Transformers) and GPT (Generative Pre-trained Transformer) series models, etc., which can capture the semantic features of text and map them to points in the vector space.

[0090] Step 140: Generate an answer to the question through a large language model based on the image vector library, the knowledge graph, and the text vector library.

[0091] In an embodiment of the present application, a large language model (LLM) such as the GPT series and Wenxin Yiyan can be used to perform natural language processing on the question to extract key information and semantic intent in the question. For example, if a user asks "What brand of laptop is the protagonist using in the cafe in the video?", the large language model will identify key entities such as "protagonist", "cafe", "laptop", and the intention of asking about the brand. Then, based on these key information, the large language model performs targeted searches in the image vector library, knowledge graph, and text vector library. In the image vector library, by calculating the similarity between the question keywords and the image feature vector, the key frame image containing the protagonist, cafe, and laptop is located; in the knowledge graph, the entity relationship related to "protagonist" and "laptop" is searched to obtain possible brand clues; in the text vector library, text information matching the question keywords is retrieved to see if there is a description of the laptop brand.

[0092] The retrieved multimodal information undergoes deep reasoning by the large language model to generate accurate answers. Specifically, the large language model will conduct comprehensive analysis and logical integration of the information obtained from the image vector library, knowledge graph, and text vector library. For example, the appearance features of the laptop computer in the image vector library are compared with the product features of common brands in the knowledge graph, and cross-validation and reasoning are performed with reference to possible brand mentions in the text vector library. During the reasoning process, the large language model uses its powerful language understanding and logical reasoning capabilities to handle contradictions or missing parts between information. If the brand logo cannot be clearly identified in the image, but the knowledge graph shows that the protagonist often uses a certain brand of computer, and there are relevant hints in the text vector library, the large language model will generate a reasonable answer based on probability and logical judgment.

[0093] The large language model converts the inference results into natural and fluent natural language expressions, generates answers to user questions, and returns them to the user.

[0094] According to the video understanding question-answering method of the present application, video data and user questions regarding the video data are obtained; multiple key frame images are extracted based on the feature differences between video frames in the video data; image features of the key frame images are extracted to construct an image vector library, and content information of the key frame images is extracted to construct a knowledge graph and a text vector library; answers to questions are generated based on the image vector library, the knowledge graph, and the text vector library through a large language model. The embodiment of the present application extracts key frame images by analyzing the feature differences between video frames, which can capture important time nodes and plot turning points in the video, reducing the problem of missing key information due to fixed feature extraction in traditional methods, constructing an image vector library, a knowledge graph, and a text vector library, not only quantitatively describing the video content from a visual perspective, but also mining the semantic relationship and logical structure of the video content through the knowledge graph, more comprehensively and deeply mining and integrating video semantic knowledge, and utilizing the powerful semantic understanding and generation capabilities of the large language model, combined with the above-mentioned multimodal data joint reasoning, to more comprehensively understand user questions and improve the accuracy of question-answering.

[0095] In some embodiments, extracting image features of key frame images to construct an image vector library includes:

[0096] Perform object detection on the key frame image to obtain object detection results; the object detection results include object coordinates and object categories;

[0097] Input the object detection results and key frame images into the image encoder to extract image features;

[0098] The image features of different key-frame images are stored hierarchically according to the scene type, object category and timestamp of the key-frame images to construct an image vector library.

[0099] In this embodiment, deep learning algorithms, such as the YOLO (You Only Look Once) series and Faster R-CNN models, can be used to perform object detection tasks. These object detection models, trained on large amounts of image data, can accurately identify a variety of common object categories, ranging from people, vehicles, and animals to everyday items and architectural structures. For example, using keyframe images from a movie, an object detection model can quickly identify objects such as the main character, props, and iconic buildings in the scene.

[0100] In some embodiments, the appropriate object detection model can be selected by sharing and understanding the image scenes in the video data. For example, in complex scenes, GroundingDINO can be used to focus on the semantic localization of fine-grained objects; in fast-moving scenes, YOLO-X can be used to leverage its lightweight advantages to achieve real-time detection.

[0101] During the detection process, the object detection model outputs object detection results, which can include object coordinates and object categories. Object coordinates, measured in pixels, represent the object's specific location within the keyframe image, typically represented by the coordinates of the top-left and bottom-right corners of the bounding box. The object category assigns a label to the identified object, such as "person," "car," or "cup." Object detection can break down complex images into distinct objects with clear identification and location.

[0102] In this embodiment, Figure 2 As shown in Figure 1, the object detection results and the corresponding keyframe images can be input into an image encoder for image feature extraction. The image encoder is a deeply trained neural network module, such as the Contrastive Language-Image Pretraining (CLIP) model.

[0103] When extracting image features, the image encoder uses the object coordinates and category information from the object detection results to perform targeted feature extraction on each object in the keyframe image. For example, for detected human objects, the image encoder focuses on extracting features such as the person's posture and facial expression; for buildings, it extracts shape and texture features. In this way, the image encoder generates a comprehensive image feature vector that not only contains the overall visual information of the image but also incorporates the semantic information provided by the object detection results.

[0104] In order to achieve efficient management and fast retrieval of these features, it is necessary to store image features in layers according to the scene type, object category and timestamp of the key frame image, so as to build a complete image vector library.

[0105] Specifically, the first layer of hierarchical storage categorizes image features by scene type. Scene type summarizes the overall context and subject matter of an image, such as scenes with people, nature, indoors, or outdoors. By grouping image features of the same scene type into the same category, it's possible to quickly filter out image information that meets specific scene requirements. For example, when a user requests content related to a specific scene in a video, they can directly search the storage layer for that scene type, significantly improving query efficiency.

[0106] The second layer further subdivides the scene type by object category. Within the same scene type, there may be multiple different object categories. By layering the object categories, we can more accurately locate image features containing specific objects. For example, within an "outdoor scene," the image can be further subdivided into sub-storage layers containing different object categories, such as people, vehicles, and buildings, making it easier for users to search for specific objects.

[0107] The third layer stores the image features of each object category using timestamps as clues, deeply integrating the time dimension into the data structure and giving the image vector library dynamic logical coherence.

[0108] In this embodiment, by performing object detection on key frame images, objects in the image and their coordinates and categories can be identified, providing a semantic anchor for image feature extraction, enhancing the pertinence and discrimination of features. This feature extraction method that combines semantic information can more comprehensively reflect the visual content and semantic meaning of the image, and improve the quality of image features; image features are stored in layers according to the scene type, object category and timestamp of the key frame image, which not only facilitates rapid retrieval and matching, but also better reflects the spatiotemporal structure and semantic hierarchy of the video content, providing richer data support for the understanding of complex video content and question-answering tasks.

[0109] In some embodiments, extracting content information of key frame images to construct a knowledge graph includes:

[0110] Extract the content information of key frame images through the image understanding model to obtain image description text;

[0111] Input the image description text into the large language model to extract the knowledge triples of the image description text;

[0112] The object detection results of the key frame images are used as information supplement, and the knowledge subgraph of the key frame images is constructed based on the knowledge triples;

[0113] The knowledge graph is obtained by fusing the knowledge subgraphs of different key frame images through a large language model.

[0114] In this example, BLIP-2 can be used as an image understanding model to perform in-depth analysis of keyframe images and obtain rich and accurate image description text. In addition to mining basic visual element descriptions, BLIP-2 also explores the semantic information and emotional tendencies behind the image. For example, in an image of people celebrating a festival, it not only identifies visual elements such as people and decorations, but also analyzes the emotional tendency of joy, providing a more comprehensive information foundation for subsequent processing.

[0115] In this embodiment, the image description text can be analyzed by a large language model. The large language model can perform basic natural language processing operations such as word segmentation, part-of-speech tagging, and syntactic analysis on the image description text, disassemble the grammatical structure of the text, and extract various entities from the text, such as people, objects, places, etc., through named entity recognition technology. In addition, the large language model can also analyze the semantic relationship between entities, determine appropriate predicates to connect entities, and form knowledge triples, namely entity-relationship-entity, such as protagonist (entity)-located (relationship)-street corner (entity). In this way, the semantic information in the image description text is converted into structured knowledge triples.

[0116] In this embodiment, after extracting knowledge triples, the object detection results need to be integrated to construct a knowledge subgraph for the keyframe image. The object information obtained from object detection is combined to further supplement the elements and relationships of the knowledge subgraph. For example, spatial relationships (such as adjacency and inclusion) and action relationships (such as pushing and picking) between objects can be added.

[0117] By supplementing the object detection results with information and organizing them in a graph structure, a knowledge subgraph of the keyframe image can be constructed. In the knowledge subgraph, each entity serves as a node, and the relationships between entities serve as edges, forming a visual semantic network. A keyframe image may contain multiple entities and multiple relationships. These elements, connected by nodes and edges, fully present the semantic information in the image. For example, with the protagonist as the core node, edges related to other people, objects, and scenes extend outward, constructing a local knowledge network reflecting the content of the keyframe image, laying the foundation for subsequent graph fusion.

[0118] The knowledge subgraph of a single keyframe image only reflects the local information of the frame. To fully present the semantic relationship of the entire video, it is necessary to fuse the knowledge subgraphs of different keyframe images and construct a unified knowledge graph.

[0119] In this embodiment, a large language model can be used to analyze the entities and relationships within each knowledge subgraph, identifying duplicate entities and similar relationships. For example, if the knowledge subgraphs for different keyframe images all contain the entity "protagonist," the model will merge these identical entities and integrate their relationship information across the subgraphs. For similar relationships, such as "protagonist walking on the street" and "protagonist running on the street," the model will analyze their semantic differences and make appropriate summaries and distinctions during fusion.

[0120] In this embodiment, the key frame image content information is extracted through the image understanding model to generate image description text, realizing the initial transformation from visual information to natural language description. With the help of the large language model, knowledge triples are extracted from the image description text, the core semantic relationships in the text are fully mined, the image information is further abstracted into structured knowledge, and the object detection results are incorporated as supplementary information, which enriches the details and accuracy of the knowledge subgraph, so that the knowledge subgraph can more comprehensively reflect the key frame image content. The knowledge subgraphs of each key frame image are fused through the large language model, and the scattered knowledge is integrated into a complete knowledge graph to form a systematic global knowledge graph. The semantic associations between different key frames in the video content are established, which can more comprehensively and deeply mine and integrate video semantic knowledge than traditional methods.

[0121] In some embodiments, the method further comprises:

[0122] Perform time series analysis on the knowledge subgraphs corresponding to the continuous key frame images to obtain analysis results; the analysis results include event chain recognition results, event trend analysis results, and abnormal event detection results;

[0123] Extract event features and event relationships from analysis results;

[0124] According to event features and event relationships, event nodes are constructed in the knowledge subgraph and association edges between events and entities in the knowledge subgraph are established.

[0125] In this embodiment, in order to present the video content more comprehensively and provide users with richer knowledge support, it is necessary not only to construct a static knowledge subgraph, but also to perform time series analysis on the knowledge subgraphs corresponding to continuous key frame images so that the knowledge subgraphs have time context and dynamic evolution characteristics.

[0126] Specifically, the knowledge subgraphs of consecutive keyframes can be arranged in strict chronological order to form a complete "knowledge timeline," with each knowledge subgraph representing the semantic information at a specific moment in the video. Time series analysis can be performed using LSTM-CRF (Long Short-Term Memory with Conditional Random Field) to obtain multi-dimensional analysis results. These results can include event chain identification, event trend analysis, and abnormal event detection.

[0127] Among them, the result of event chain recognition is that the model mines a sequence of events with causal relationships. For example, in a video of a conference scene, through the analysis of the knowledge subgraph, the event chain of "people enter → meeting starts → keynote speech" is identified, clearly presenting the logical context of the development of the event. The results of event trend analysis include predictive analysis of continuous events. For example, in weather videos, the knowledge subgraph related to weather conditions in continuous keyframe images is analyzed to predict weather change trends; in data visualization videos, the knowledge subgraph related to data growth is analyzed to predict the subsequent direction of the data. The result of abnormal event detection is to compare normal event patterns with the current knowledge subgraph information to identify event changes that do not conform to conventional logic. For example, in a surveillance video, if a frame of the knowledge subgraph shows that an object suddenly moves in the picture, and there is no information about the existence of the relevant object before, it can be determined as an abnormal event.

[0128] In this embodiment, event features and event relationships can be extracted from the analysis results. Event features may include

[0129] Event type, duration, participating entities, scope of influence, etc. Specifically, the event type specifies the nature of the event, such as "meeting" or "sports." The duration can be determined by analyzing the timestamp span corresponding to the knowledge subgraph, reflecting the length of the event. Participating entities can be extracted from the nodes of the knowledge subgraph to identify people, objects, etc. related to the event. The scope of influence can be determined by evaluating the nodes and edges involved in the event to determine the breadth and depth of the event's impact. For example, for a "keynote speech" event, the features that can be extracted are: event type - conference activity, duration - from the speech start timestamp to the end timestamp, participating entities - speakers and audience, and scope of influence - people at the meeting and related discussion content.

[0130] Event relationships can include temporal relationships (sequential, concurrent), causal relationships, and dependency relationships between events. Temporal relationships clarify the order in which events occur or whether they occur concurrently. For example, "entry of personnel" precedes "start of the meeting," while "group discussion" and "material distribution" during the meeting may be concurrent events. Causal relationships explore the inherent driving factors between events, such as "equipment failure" leading to "production suspension." Dependency relationships indicate that the occurrence or development of an event depends on other events. For example, a "keynote speech" depends on "start of the meeting"; the speech can only proceed after the meeting begins.

[0131] In this embodiment, event features and event relationships can be integrated into the knowledge subgraph to supplement event-related nodes and edges. For example, based on event features and event relationships, event nodes such as "Meeting Starts" and "Keynote Speech" can be added, and edges can be established between events and entities, such as: Person A - Participation - Meeting Starts. This integration of event information not only enriches the content of the knowledge subgraph but also gives it the dynamic evolution characteristics of the time dimension.

[0132] In this embodiment, by performing time series analysis on the knowledge subgraph corresponding to continuous key frame images, it is possible to identify event chains, analyze event development trends and detect abnormal events from the time dimension, so that the originally static knowledge subgraph has time context and dynamic evolution characteristics, extract event characteristics and event relationships in the analysis results, and further refine the core elements and association logic of events in the video content. Based on these characteristics and relationships, event nodes are constructed in the knowledge subgraph and association edges with entities are established, and event information in the time dimension is deeply integrated into the knowledge graph system to achieve spatiotemporal fusion of local knowledge graphs. It can not only present video content more comprehensively and three-dimensionally, but also provide users with richer knowledge support for timeline-based video content questions and answers.

[0133] In some embodiments, a knowledge graph is obtained by fusing knowledge subgraphs of different keyframe images through a large language model, including:

[0134] Add timestamps to entities and relationships in different knowledge subgraphs, and merge entities and relationships with event associations through a large language model to obtain a knowledge graph. The merging includes:

[0135] For repeated parts of different knowledge subgraphs, retain the latest or most credible information;

[0136] For contradictory information in different knowledge subgraphs, the time series and the logical relationship between elements in the knowledge subgraphs are combined to remove erroneous content.

[0137] In this embodiment, the fusion process may include the following three stages:

[0138] Timestamp alignment: Video content is dynamic, and each keyframe image's knowledge subgraph reflects semantic information at a specific moment. To integrate these scattered knowledge subgraphs into a coherent whole, timestamps can be added to the entities and relationships in each knowledge subgraph, accurately indicating the moment the event occurred, marking the time period in which it was effective or existed. For example, in a movie video, a keyframe image might depict the protagonist entering a room at a certain moment, while another keyframe image depicts the protagonist conversing with someone at a later moment. By adding timestamps to these events, their order on the timeline can be determined.

[0139] Logical Relationship Sorting Stage: Video content is not just sequential in time; it also contains complex logical structures, such as cause and effect, and sequential order. The goal of this stage is to construct a logical chain and clarify the types of logical relationships between elements. For example, in a video, one event may occur because of another (cause and effect), or one event may occur only after another (sequential order). By sorting out these logical relationships, we can better understand the inherent logic of the video content.

[0140] Fusion and optimization stage: The goal of this stage is to merge entities and relationships with time association, and to handle duplicate parts and contradictory information differently during the merging process to obtain a complete, accurate and time-series-based fusion knowledge graph. For entities and relationships with time association, they can be merged into a unified knowledge graph based on timestamps and logical relationships. For example, if multiple keyframe images describe the state or behavior of the same entity at different time points, this information can be integrated into a time series to form a complete entity trajectory. For duplicate parts, the most recent or most credible information can be retained first. For example, if two keyframe images describe the same event, but the keyframe image with a later timestamp provides more detailed or more accurate information, the newer information can be used first to improve the timeliness and accuracy of the knowledge graph.

[0141] For conflicting information, we can leverage the reasoning capabilities of large language models (such as GPT-4) and combine time series and logical rules to make judgments. For example, if one keyframe image depicts the subject indoors at a certain moment, while another keyframe depicts the subject outdoors at the same moment, we can use the logical relationship and timestamp to determine which information is more credible. If the logical relationship indicates that the subject cannot be in two places at the same time, we can combine time series and contextual information to remove the erroneous content and retain the correct information.

[0142] In this embodiment, by adding timestamps to entities and relationships in different knowledge subgraphs, the knowledge graph can better reflect the changes and development of video content in the time series. By merging entities and relationships with event associations through a large language model, it is possible to integrate scattered knowledge and establish closer semantic connections. For repeated parts, the latest or most credible information is retained, which reduces redundant information interference and improves the timeliness of information in the knowledge graph. For contradictory information, the erroneous content is removed by combining the time series and the logical relationship between elements. Through deep verification and screening of knowledge, the accuracy of information in the knowledge graph is improved.

[0143] In some embodiments, generating answers to questions using a large language model based on an image vector library, a knowledge graph, and a text vector library includes:

[0144] Perform semantic analysis on the problem and break it down into multiple sub-problems;

[0145] According to the multiple sub-questions, first search results corresponding to the sub-questions are retrieved from the knowledge graph, second search results corresponding to the sub-questions are retrieved from the text vector library, and third search results corresponding to the sub-questions are retrieved from the image vector library;

[0146] Analyze the first and second search results using the large language model. If the analysis shows that the first and second search results cannot serve as answers to the corresponding sub-questions, use the large visual language model to infer the first, second, and third search results to obtain answers to the corresponding sub-questions.

[0147] The answer to each sub-question is integrated to get the answer to the question.

[0148] In some embodiments, when it is analyzed that the first search result and the second search result can serve as answers to corresponding sub-questions, the first search result and the second search result are integrated into the answers to the corresponding sub-questions through a large language model.

[0149] In this embodiment, after the user asks a question, it may be difficult to accurately obtain comprehensive and appropriate information by directly searching. Because user questions may be complex and diverse, containing multiple semantic focuses or sub-topics, it is necessary to first process the questions and break them down into more fine-grained sub-questions. Specifically, a large language model, such as GPT-4, can be used in combination with the characteristics of the video content to break down complex questions into multiple logically clear and interrelated sub-questions. For example, for the question "What is the motivation for the behavior of the characters in a specific scene in the video", it is broken down into sub-questions such as "Which characters are in the video", "What is the specific scene", and "What behaviors do the characters have", so that each sub-question focuses on a certain aspect of the video content, which is convenient for subsequent targeted processing.

[0150] For the multiple sub-problems obtained by decomposition, for each sub-problem, you can search from the knowledge graph, text vector library and image vector library respectively to obtain the corresponding search results.

[0151] Specifically, entity recognition and relationship matching can be performed in the knowledge graph based on the keywords and entities in the sub-question, relevant nodes and relationships can be found, and these nodes and relationships related to the sub-question can be combined to form the first search result of the sub-question.

[0152] Before searching from the text vector library and image vector library, the sub-question can be input into the CLIP text encoder to convert the natural language question into a vector representation and obtain the corresponding text features.

[0153] Utilize the text features corresponding to the sub-questions to search in the text vector library. Specifically, a vector similarity calculation algorithm, such as the cosine similarity algorithm, can be used to filter text feature vectors with high relevance to the sub-questions from the text vector library based on the similarity. When the similarity between a text feature vector and the text feature corresponding to the sub-question exceeds a preset threshold, it indicates that the corresponding text feature vector is related to the sub-question. These text feature vectors related to the sub-question can be filtered out, and the image description text corresponding to the text feature vector can be extracted to form a second search result. Furthermore, GPT-4 can be used as a verifier to filter the retrieved information based on semantic understanding to ensure that the retrieval results are highly relevant to the sub-questions and to obtain accurate related images.

[0154] Utilize the text features corresponding to the sub-questions to search in the image vector library. Specifically, a vector similarity calculation algorithm, such as the cosine similarity algorithm, can be used to filter image feature vectors that are highly relevant to the sub-questions from the image vector library based on the similarity. When the similarity between an image feature vector and the text features corresponding to the sub-question exceeds a preset threshold, it indicates that the corresponding image feature vector is relevant to the sub-question. These image feature vectors related to the sub-question can be filtered out, and the key frame images corresponding to the image feature vectors can be extracted to form a third search result. Furthermore, GPT-4 can be used as a verifier to filter the retrieved information based on semantic understanding to ensure that the retrieval results are highly relevant to the sub-questions and to obtain accurate relevant images.

[0155] In this way, the first search result, the second search result and the third search result corresponding to each sub-question can be obtained.

[0156] After searching the knowledge image and text vector libraries, the large language model can perform a comprehensive analysis of the first and second search results to determine whether they can serve as answers to the sub-questions. For example, for the sub-question "Based on the video explanation, what important historical renovations has the Hall of Supreme Harmony undergone?", if the first and second search results only provide a small amount of vague information and cannot clearly and accurately answer the question, the large language model will determine that the first and second search results cannot serve as answers.

[0157] If the first and second search results can serve as answers to corresponding sub-questions, the large language model can be used to combine them into the answer to the sub-question. Specifically, due to the characteristics of different data sources, the first and second search results may contain a large amount of duplicate information. The large language model can identify this duplicate content through semantic analysis and similarity calculation, compare the semantic core of different expressions, and merge information with similar meanings and duplicate descriptions, retaining the most representative and accurate expressions. In this way, redundant information is removed, and the answer to the sub-question is obtained.

[0158] In the case where the first and second search results cannot serve as answers to the corresponding sub-questions, the first, second, and third search results can be input into a visual general language model, such as VisualGLM (Visual General Language Model). Through the fusion and joint reasoning of multimodal information, the complementary information between images and texts can be fully mined to generate more comprehensive and accurate answers to sub-questions. For example, for the sub-question "What emotions does the protagonist's expression convey?", the first and second search results may not contain content that directly describes the emotions, but the key frame image of the third search result shows the protagonist's facial expression. VisualGLM combines the expression features in the image (such as frowns and eyes) and the contextual information in the text (such as conversation content and event background) to perform comprehensive analysis and reasoning to generate a more accurate answer.

[0159] After obtaining the answers to each sub-question, the answers to each sub-question can be integrated through a large language model, such as GPT-4. Based on the logical relationship and semantic coherence of the questions, a generative summary algorithm can be used to organize the scattered answers into a complete and fluent answer.

[0160] For questions involving the time dimension, relevant time information can be extracted from the knowledge graph and the answers can be presented in the form of a visual timeline. The timeline can intuitively show the sequence and duration of events, helping users better understand the temporal logic of the video content. For example, when a user asks "What is the process of the meeting in the video?", the answer will not only describe each link in text, but also generate a timeline, marking the time points and duration of key events such as "check-in time", "opening speech", and "keynote speech", so that the information can be presented visually.

[0161] After fusion and time-series processing, the answer is fed back to the user in natural language. This output not only includes the direct answer to the question but also provides relevant context and explanations as needed, enhancing the completeness and comprehensibility of the answer. For example, when answering the question "Why did the protagonist make that decision?", the answer not only provides the direct reason but also briefly reviews the relevant plot developments, helping the user better understand the whole process of the event.

[0162] In this embodiment, by semantically parsing the question and breaking it down into sub-questions, complex questions can be refined, making subsequent retrieval and reasoning more targeted. The answers to each sub-question are retrieved based on the knowledge graph, text vector library, and image vector library respectively, and relevant data is obtained from multiple dimensions of semantic knowledge, text information, and image features. When the retrieval results of the knowledge graph and text vector library analyzed by the large language model cannot meet the requirements of answering the sub-questions, the visual large language model is introduced to perform reasoning in combination with the retrieval results of the image vector library. With the help of the multimodal fusion capability of the visual large language model, the shortcomings of single modality information are compensated, and accurate answers to users' questions are achieved.

[0163] In this embodiment, when the large language model determines that the retrieval results of the knowledge graph and text vector library are sufficient to answer the sub-question, the answers to the sub-questions are directly integrated to generate the answers. This reduces the reasoning process of the large visual language model, thereby reducing the consumption of computing resources and processing time, and improving the efficiency of question answering.

[0164] The video understanding question-answering method provided in the embodiment of the present application can be executed by a video understanding question-answering device. In the embodiment of the present application, the video understanding question-answering device executing the video understanding question-answering method is used as an example to illustrate the video understanding question-answering device provided in the embodiment of the present application.

[0165] The embodiment of the present application also provides a video understanding question and answer device.

[0166] like Figure 3 As shown, the video understanding question answering device includes:

[0167] An acquisition module 310 is used to acquire video data and user questions regarding the video data;

[0168] An extraction module 320 is configured to extract a plurality of key frame images based on feature differences between video frames in the video data;

[0169] A construction module 330 is configured to extract image features of key frame images to construct an image vector library, and to extract content information of key frame images to construct a knowledge graph and a text vector library;

[0170] The generation module 340 is used to generate answers to questions based on the image vector library, knowledge graph and text vector library through a large language model.

[0171] According to the video understanding question-answering device of the present application, video data and user questions regarding the video data are obtained; multiple key frame images are extracted based on the feature differences between video frames in the video data; image features of the key frame images are extracted to construct an image vector library, and content information of the key frame images is extracted to construct a knowledge graph and a text vector library; answers to questions are generated based on the image vector library, the knowledge graph, and the text vector library through a large language model. The embodiment of the present application extracts key frame images by analyzing the feature differences between video frames, which can capture important time nodes and plot turning points in the video, reducing the problem of missing key information due to fixed feature extraction in traditional methods, constructing an image vector library, a knowledge graph, and a text vector library, not only quantitatively describing the video content from a visual perspective, but also mining the semantic relationship and logical structure of the video content through the knowledge graph, more comprehensively and deeply mining and integrating video semantic knowledge, and utilizing the powerful semantic understanding and generation capabilities of the large language model, combined with the above-mentioned multimodal data joint reasoning, to more comprehensively understand user questions and improve the accuracy of question-answering.

[0172] In some embodiments, the building block 330 is further configured to:

[0173] Perform object detection on the key frame image to obtain object detection results; the object detection results include object coordinates and object categories;

[0174] Input the object detection results and key frame images into the image encoder to extract image features;

[0175] The image features of different key-frame images are stored hierarchically according to the scene type, object category and timestamp of the key-frame images to construct an image vector library.

[0176] In some embodiments, the building block 330 is further configured to:

[0177] Extract the content information of key frame images through the image understanding model to obtain image description text;

[0178] Input the image description text into the large language model to extract the knowledge triples of the image description text;

[0179] The object detection results of the key frame images are used as information supplement, and the knowledge subgraph of the key frame images is constructed based on the knowledge triples;

[0180] The knowledge graph is obtained by fusing the knowledge subgraphs of different key frame images through a large language model.

[0181] In some embodiments, the building block 330 is further configured to:

[0182] Perform time series analysis on the knowledge subgraphs corresponding to the continuous key frame images to obtain analysis results; the analysis results include event chain recognition results, event trend analysis results, and abnormal event detection results;

[0183] Extract event features and event relationships from analysis results;

[0184] According to event features and event relationships, event nodes are constructed in the knowledge subgraph and association edges between events and entities in the knowledge subgraph are established.

[0185] In some embodiments, the building block 330 is further configured to:

[0186] Add timestamps to entities and relationships in different knowledge subgraphs, and merge entities and relationships with event associations through a large language model to obtain a knowledge graph. The merging includes:

[0187] For repeated parts of different knowledge subgraphs, retain the latest or most credible information;

[0188] For contradictory information in different knowledge subgraphs, the time series and the logical relationship between elements in the knowledge subgraphs are combined to remove erroneous content.

[0189] Perform semantic analysis on the problem and break it down into multiple sub-problems;

[0190] In some embodiments, the generating module 340 is further configured to:

[0191] According to the multiple sub-questions, first search results corresponding to the sub-questions are retrieved from the knowledge graph, second search results corresponding to the sub-questions are retrieved from the text vector library, and third search results corresponding to the sub-questions are retrieved from the image vector library;

[0192] Analyze the first and second search results using the large language model. If the analysis shows that the first and second search results cannot serve as answers to the corresponding sub-questions, use the large visual language model to infer the first, second, and third search results to obtain answers to the corresponding sub-questions.

[0193] The answer to each sub-question is integrated to get the answer to the question.

[0194] In some embodiments, the generating module 340 is further configured to:

[0195] When it is analyzed that the first search result and the second search result can be used as answers to the corresponding sub-questions, the first search result and the second search result are integrated into the answers to the corresponding sub-questions through the large language model.

[0196] The video understanding question-answering device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or a device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.

[0197] The video understanding question-answering device in the embodiments of the present application may be a device having an operating system. The operating system may be a Microsoft (Windows) operating system, an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0198] In some embodiments, as Figure 4 As shown, an embodiment of the present application also provides an electronic device 400, including a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401. When the program is executed by the processor 401, each process of the above-mentioned video understanding question-answering method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0199] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0200] An embodiment of the present application also provides a non-transitory computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the various processes of the above-mentioned video understanding question-answering method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0201] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0202] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above-mentioned video understanding question-answering method when executed by a processor.

[0203] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0204] An embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned video understanding question-answering method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0205] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0206] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0207] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0208] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

[0209] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0210] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.

Claims

1. A video understanding question answering method, characterized in that: include: Obtaining video data and user questions regarding the video data; extracting a plurality of key frame images according to feature differences between video frames in the video data; Extracting image features of the key frame images to construct an image vector library, and extracting content information of the key frame images to construct a knowledge graph and a text vector library; An answer to the question is generated by a large language model based on the image vector library, the knowledge graph, and the text vector library.

2. The method according to claim 1, characterized in that The step of extracting image features of the key frame image to construct an image vector library includes: Performing object detection on the key frame image to obtain an object detection result; the object detection result includes object coordinates and object category; Inputting the object detection result and the key frame image into an image encoder to extract image features; Image features of different key frame images are hierarchically stored according to the scene type, object category and time stamp of the key frame images to construct an image vector library.

3. The method according to claim 1, characterized in that Extracting content information of the key frame image to construct a knowledge graph includes: Extracting content information of the key frame image through an image understanding model to obtain image description text; Inputting the image description text into the large language model to extract knowledge triples of the image description text; Using the object detection result of the key frame image as information supplement, and constructing a knowledge subgraph of the key frame image according to the knowledge triples; The knowledge graph is obtained by fusing the knowledge subgraphs of different key frame images through the large language model.

4. The method according to claim 3, characterized in that The method further comprises: Performing time series analysis on the knowledge subgraphs corresponding to the consecutive key frame images to obtain analysis results; the analysis results include event chain recognition results, event trend analysis results, and abnormal event detection results; extracting event features and event relationships from the analysis results; According to the event features and event relationships, event nodes are constructed in the knowledge subgraph and association edges between events and entities in the knowledge subgraph are established.

5. The method according to claim 3, characterized in that The method of fusing knowledge subgraphs of different key frame images using the large language model to obtain a knowledge graph includes: Add timestamps to entities and relationships in different knowledge subgraphs, and merge entities and relationships with event associations through a large language model to obtain a knowledge graph; the merging includes: For repeated parts of different knowledge subgraphs, retain the latest or most credible information; For contradictory information in different knowledge subgraphs, the time series and the logical relationship between elements in the knowledge subgraphs are combined to remove erroneous content.

6. The method according to claim 1, characterized in that Generating the answer to the question based on the image vector library, the knowledge graph, and the text vector library using a large language model includes: Perform semantic analysis on the problem and decompose it into multiple sub-problems; According to the plurality of sub-questions, first search results corresponding to the sub-questions are retrieved from the knowledge graph, second search results corresponding to the sub-questions are retrieved from the text vector library, and third search results corresponding to the sub-questions are retrieved from the image vector library; Analyzing the first and second search results using the large language model, and if the analysis shows that the first and second search results cannot serve as answers to corresponding sub-questions, reasoning on the first, second, and third search results using the large visual language model to obtain answers to corresponding sub-questions; The answers to the sub-questions are integrated to obtain the answer to the question.

7. The method according to claim 6, characterized in that The method further comprises: When it is analyzed that the first search result and the second search result can serve as answers to the corresponding sub-questions, the first search result and the second search result are integrated into answers to the corresponding sub-questions through the large language model.

8. A video comprehension question-answering device, characterized in that: include: An acquisition module, configured to acquire video data and user questions regarding the video data; An extraction module, configured to extract a plurality of key frame images based on feature differences between video frames in the video data; A construction module, configured to extract image features of the key frame images to construct an image vector library, and to extract content information of the key frame images to construct a knowledge graph and a text vector library; A generation module is used to generate an answer to the question based on the image vector library, the knowledge graph and the text vector library through a large language model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the video understanding question-answering method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the video understanding question answering method according to any one of claims 1 to 7 is implemented.