A structured video understanding method, apparatus, computer device, and readable storage medium based on a multimodal large model.

By extracting focal entities from videos and generating spatiotemporal scene graphs using a multimodal large model, the accuracy problem of traditional video understanding methods in complex scenes is solved, and user needs are met efficiently.

CN119478769BActive Publication Date: 2026-01-06DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411501242.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2026-01-06
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

Traditional video understanding methods struggle to handle complex dynamic scenes and lack targeted analysis of user needs, leading to inaccurate understanding results.

Method used

A structured video understanding method based on a multimodal large model is adopted. By acquiring the video and interactive text input by the user, focusing keywords are extracted, focusing entities are determined by combining the pre-trained video multimodal large model, and a focusing spatiotemporal scene graph is generated, and finally, dialogue feedback is provided.

Benefits of technology

It achieves efficient and accurate video content understanding, meets user needs, and provides highly targeted video answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478769B_ABST
    Figure CN119478769B_ABST
Patent Text Reader

Abstract

The application discloses a structured video understanding method and device based on a multimodal large model, computer equipment and a readable storage medium, which comprises the following steps: first, obtaining a to-be-processed video of a user and interactive text of a description requirement; then, extracting keywords from the interactive text and determining a focused entity by combining a first video multimodal large model; then, inputting the video into a second video multimodal large model to obtain a focused spatio-temporal scene graph of the focused entity; finally, performing dialogue feedback on the interactive text according to the scene graph, realizing efficient and accurate video content understanding and meeting the user requirement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing, and more specifically, to a structured video understanding method, apparatus, computer device, and readable storage medium based on a multimodal large model. Background Technology

[0002] In today's digital age, video, as a crucial information carrier, contains a vast amount of content. With the development of intelligent technologies, accurately understanding video content has become key to meeting the needs of numerous application scenarios, such as intelligent surveillance requiring the analysis of abnormal behavior in videos and autonomous driving needing to interpret road conditions. However, traditional video understanding methods face many challenges, such as difficulty handling complex dynamic scenes and a tendency to produce inaccurate understanding results. Existing video understanding technologies often lack targeted analysis of user needs and are not efficient or accurate enough in processing video and text interactions. Summary of the Invention

[0003] The purpose of this invention is to provide a structured video understanding method, apparatus, computer device, and readable storage medium based on a multimodal large model.

[0004] In a first aspect, embodiments of the present invention provide a structured video understanding method based on a multimodal large model, including:

[0005] Obtain the video to be processed and interactive text input by the user, wherein the interactive text is a description of the user's needs for the video to be processed;

[0006] Keyword extraction is performed on the interactive text to obtain focusing keywords, and the focusing entity corresponding to the focusing keywords in the video to be processed is determined by combining the pre-trained first video multimodal large model.

[0007] The video to be processed is input into a pre-trained second video multimodal large model to obtain a spatiotemporal scene map of the focused person and the focused object, which are included in the focused entity.

[0008] The interactive text is used to provide dialogue feedback based on the focused spatiotemporal scene diagram.

[0009] In one possible implementation, the step of extracting key keywords from the interactive text to obtain focused keywords includes:

[0010] The pre-trained large language model or NLTK library is used to extract keywords from the interactive text to obtain focused keywords.

[0011] In one possible implementation, determining the focus entity corresponding to the focus keyword in the video to be processed by combining a pre-trained first video multimodal large model includes:

[0012] The focus entity corresponding to the focus keyword in the video to be processed is determined by combining the first video multimodal large model obtained by training the fine-tuned VILA model.

[0013] In one possible implementation, the step of inputting the video to be processed into a pre-trained second video multimodal large model to obtain a spatiotemporal scene map of the focused entities, including focused people and focused objects, includes:

[0014] The second video multimodal large model, which is pre-trained from the video to be processed, is used by frame extraction to obtain the spatiotemporal scene map of the focused entity, including the focused person and the focused object.

[0015] Based on the intersection-union ratio of the focused entities, each focused person and each focused item is determined, resulting in multiple focused spatiotemporal scene graphs that include the interaction relationships between each focused person and each focused item. Each focused spatiotemporal scene graph includes a unique identifier ID for each focused entity, used to track each focused entity on the video timeline.

[0016] In one possible implementation, the focused spatiotemporal scene graph is represented in JSON format text form. The focused spatiotemporal scene graph includes frame information, node information, and edge information. The node information represents the focused entity and its position, and the edge information represents the interaction relationship or spatial relationship between the focused entities.

[0017] In one possible implementation, the step of providing dialogue feedback to the interactive text based on the focused spatiotemporal scene graph includes:

[0018] The focused spatiotemporal scene graph is analyzed to extract key video details and timing information;

[0019] The extracted information is input into a pre-trained dialogue generation model to generate a response text to the interactive text.

[0020] In one possible implementation, the method further includes:

[0021] The first video multimodal large model and the second video multimodal large model are jointly trained. The first video multimodal large model and the second video multimodal large model share some training data, and the model parameters are continuously optimized through cross-validation.

[0022] Secondly, embodiments of the present invention provide a structured video understanding device based on a multimodal large model, characterized in that it includes:

[0023] The acquisition module is used to acquire the video to be processed and the interactive text input by the user, wherein the interactive text is a description of the user's needs for the video to be processed;

[0024] The understanding module is used to extract keywords from the interactive text to obtain focus keywords, and combine them with a pre-trained first video multimodal big data model to determine the focus entities corresponding to the focus keywords in the video to be processed; input the video to be processed into a pre-trained second video multimodal big data model to obtain a focus spatiotemporal scene map of the focus entities, including focus people and focus objects; and provide dialogue feedback to the interactive text based on the focus spatiotemporal scene map.

[0025] Thirdly, embodiments of the present invention provide a computer device, the computer device including a processor and a non-volatile memory storing computer instructions, wherein when the computer instructions are executed by the processor, the computer device performs the method described in the first aspect.

[0026] Fourthly, embodiments of the present invention provide a readable storage medium, the readable storage medium including a computer program, wherein the computer program, when running, controls the computer device where the readable storage medium is located to execute the method described in the first aspect.

[0027] Compared with existing technologies, the beneficial effects provided by this invention include: using a structured video understanding method, apparatus, computer device, and readable storage medium based on a multimodal large model disclosed in this invention, the method acquires the user's video to be processed and interactive text describing the user's needs; then, it extracts keywords from the interactive text and combines them with a first video multimodal large model to determine the focus entity; then, it inputs the video into a second video multimodal large model to obtain a focus spatiotemporal scene map of the focus entity; and finally, it provides dialogue feedback to the interactive text based on this scene map, thereby achieving efficient and accurate video content understanding and meeting user needs. Attached Figure Description

[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a flowchart illustrating the steps of a structured video understanding method based on a multimodal large model provided in an embodiment of the present invention.

[0030] Figure 2 This is a schematic block diagram of the structured video understanding device based on a multimodal large model provided in an embodiment of the present invention;

[0031] Figure 3 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0033] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0034] In order to solve the technical problems mentioned in the background art Figure 1 This is a flowchart illustrating the structured video understanding method based on a multimodal large model provided in this disclosure. The following is a detailed description of this structured video understanding method based on a multimodal large model.

[0035] Step S201: Obtain the video to be processed and interactive text input by the user, wherein the interactive text is a description of the user's needs for the video to be processed;

[0036] Step S201: Extract keywords from the interactive text to obtain focusing keywords, and combine them with a pre-trained first video multimodal big model to determine the focusing entity corresponding to the focusing keywords in the video to be processed;

[0037] Step S201: Input the video to be processed into a pre-trained second video multimodal large model to obtain a spatiotemporal scene map of the focused entity, including the focused person and the focused object;

[0038] Step S201: Provide dialogue feedback to the interactive text based on the focused spatiotemporal scene diagram.

[0039] In this embodiment of the invention, for example, it is assumed that the server is a core component of an intelligent video analysis service platform. A user sends a request to the server through a client (such as a mobile app or webpage). The user wants to know specific information from a video about traffic conditions on a city street, so they upload a 30-second video file (the video to be processed) and simultaneously enter the interactive text in the input box: "I want to know the interaction between the traffic police officer in the blue uniform and the car that ran the red light in the video."

[0040] After receiving the video file and interactive text from the user, the server first performs a format check and preliminary integrity verification on the video file to ensure that the video can be read and processed normally. For the interactive text, the server performs simple encoding conversion (if necessary) to convert it into an internally processable format for use in subsequent steps.

[0041] After receiving the interactive text, the server begins keyword extraction. First, it calls a pre-trained large language model (e.g., GPT-3.5) and the NLTK library to extract keywords.

[0042] For large language models, the server constructs a specific prompt, such as "Please extract key nouns related to people, things, and behaviors from the following text: I want to know the interaction between the traffic police officer in the blue uniform and the car that ran a red light in the video." Based on this prompt, the large language model might return keywords such as "traffic police officer," "blue uniform," "car," "running a red light," and "interaction."

[0043] Meanwhile, the NLTK library also begins working. The server preprocesses the interactive text, such as performing word segmentation, and then uses the NLTK library's part-of-speech tagging feature to identify nouns. The NLTK library might identify nouns such as "traffic police" and "car".

[0044] The server will merge and deduplicate the keywords extracted from the large language model or NLTK library, and finally obtain the focused keywords: "traffic police", "blue uniform", "car", "running a red light", and "interaction".

[0045] The server combines these focus keywords with a pre-trained first video multimodal large model (assuming it is a finely tuned VILA model) to determine the focus entity.

[0046] The server inputs the focus keywords and relevant information about the video to be processed (such as basic video metadata, including video duration, resolution, etc.) into the first video multimodal large model. The model first analyzes the video, using its visual feature recognition capabilities learned during pre-training, and begins searching for entities related to the focus keywords in each frame of the video.

[0047] For the keyword "traffic police," the model will search for individuals wearing blue uniforms in the video (further refining the location by combining the keyword "blue uniform"), and determine their position and related attributes in the video using object detection algorithms. For the keyword "car," the model will identify car objects in the video and, based on the behavioral keyword "running a red light," attempt to analyze whether the car's trajectory matches the characteristics of running a red light (e.g., by comprehensively judging based on information such as traffic light status and the time it takes for the vehicle to cross the stop line).

[0048] After analysis by the model, the identified focus entities were: the traffic police officer in blue uniform in the video (the specific location and appearance characteristics were identified by the model) and the car that was judged to be likely to run a red light (the car's location, color, model, and other information were also identified).

[0049] The server inputs the video to be processed into a pre-trained second video multimodal large model by sampling frames (2 frames per second). This model can also be a model fine-tuned for a specific dataset, such as the VILA model mentioned earlier or other suitable video multimodal large models.

[0050] Once the video frames are input into the model, the model begins to analyze each frame. For the traffic police and cars in the focus area, the model uses its object detection capabilities to accurately locate the positions of the traffic police and cars in each frame. For example, in the first frame, the traffic police are located in the lower left corner of the screen (the position is represented by coordinate information, such as [10, 20, 30, 40], which represents the upper left and lower right corner coordinates of the traffic police in a rectangular area in the screen), and the car is located in the middle-right position of the screen [100, 120, 150, 180].

[0051] As video frames are continuously input, the model begins to assign unique identifiers (IDs) to the person in focus (traffic police officer) and the object in focus (car). Let's assume the traffic police officer is assigned ID 1 and the car is assigned ID 2.

[0052] When constructing a focused spatiotemporal scene graph, the model needs to consider the interaction and spatial relationships between entities. For spatial relationships, the model analyzes the relative positions of the traffic police officer and the car in each frame; for example, in a certain frame, the traffic police officer is standing to the left of the car (this spatial relationship is recorded). For interaction relationships, since the user mentioned "interaction situations," the model pays special attention to whether there are specific interactive behaviors between the traffic police officer and the car. For example, starting from frame 10 of the video, the traffic police officer makes a gesture indicating that the car should stop. After the model detects this action, it records this interaction relationship in the focused spatiotemporal scene graph, and the edge information is recorded as [1, "indicating to stop", 2].

[0053] A focused spatiotemporal scene diagram represented in JSON format might look like this:

[0054]

[0055] This focused spatiotemporal scene map is continuously updated as video frames are processed, completely recording the spatiotemporal information and interaction relationships of the focused entity in the video.

[0056] After receiving the generated focused spatiotemporal scene graph, the server first performs a parsing operation. The server reads the node and edge information frame by frame according to the JSON format structure.

[0057] From the node information, the server obtains the location, category, and ID information of the traffic police and the car in each frame. For example, by reading the information in the aforementioned focused spatiotemporal scene graph, the server knows the initial positions of the traffic police and the car in frame 1, and their position changes in frame 10. From the edge information, the server obtains the interaction relationship between the traffic police and the car, such as the traffic police signaling the car to stop in frame 10.

[0058] The server inputs the parsed key video details (positional changes of traffic police and cars, interaction relationships, etc.) and temporal information (what happened in which frame) into a pre-trained dialogue generation model.

[0059] The dialogue generation model generates response text based on this input information and the user's interactive text content. For example, the response text might be: "In the video, a traffic policeman in a blue uniform is standing in the lower left corner of the frame, and the car is located slightly to the right of the center of the frame. At frame 10, the traffic policeman stands to the left of the car and signals the car to stop."

[0060] This response text accurately answers the user's question about the interaction between traffic police and cars, and is generated based on a precise understanding of the video content (represented by a focused spatiotemporal scene diagram). In this way, the server provides the user with an effective answer regarding the video content, thus meeting the user's needs.

[0061] On the server side, the first and second video multimodal large models are jointly trained. The server prepares a large dataset containing a large number of videos and their corresponding interactive texts.

[0062] During training, the first and second video multimodal large models share some training data. For example, a portion of the data is used to train the first model's ability to identify focused entities, while another portion is used to train the second model's ability to generate focused spatiotemporal scene maps. However, some intermediate data is shared, such as basic annotation information about entities in the video.

[0063] The model parameters are continuously optimized through cross-validation. For example, the server divides the dataset into multiple subsets. Each time, a part of the subsets is used as the validation set, and the remaining subsets are used as the training set. When using the first subset as the validation set, the second model is trained on the other subsets and then its performance is verified using the first subset, and vice versa. In this way, the parameters of the model are continuously adjusted to improve the accuracy and robustness of the model, enabling the entire structured video understanding method based on the multimodal large model to better handle various videos and interactive text inputs from users.

[0064] In the embodiment of the present invention, for the extraction of focus keywords from the interactive text, the following examples can be used for implementation.

[0065] Call a pre-trained large language model or the NLTK library to extract focus keywords from the interactive text.

[0066] In the embodiment of the present invention, exemplarily, the server, as the core of the intelligent video understanding service, continuously listens for requests from users. When the user inputs a video to be processed and interactive text (such as "I want to know what the child holding a red balloon in the video is doing"), the server first receives and stores this information.

[0067] For the interactive text, the server needs to perform preprocessing for subsequent keyword extraction operations. The server converts the interactive text from the format input by the user (which may be UTF-8 encoded, etc.) into a standard format for internal processing. At the same time, simple cleaning operations are performed on the text, such as removing extra spaces and punctuation marks (but retaining punctuation marks that have an impact on semantics, such as the word "de" in "the child holding a red balloon" is important for semantic understanding and cannot be removed).

[0068] The server calls a pre-trained large language model (taking GPT-3.5 as an example) for keyword extraction. The server constructs a request specifically for keyword extraction and includes the preprocessed interactive text in the request content. For example, construct a request like this: "Please perform keyword extraction on the following text, focusing on nouns related to people, objects, actions, etc.: I want to know what the child holding a red balloon in the video is doing".

[0069] After receiving the server's request, the large language model analyzes it according to its pre-trained algorithm and a large amount of corpus knowledge. It will identify important elements in the text. For this example, the large language model may identify words such as "child", "red balloon", and "what to do" as keywords. This is because the large language model has learned various language structures and semantic information during the pre-training process and can understand the importance of these words in describing a scene.

[0070] After receiving the keyword results returned by the large language model, the server will perform preliminary processing. For example, the results returned by the large language model may contain some auxiliary explanations or other irrelevant formatting information. The server needs to remove this information and keep only the pure keyword content to obtain a preliminary keyword set, such as {"child", "red balloon", "what to do"}.

[0071] Simultaneously, the server utilizes the NLTK library for keyword extraction. First, the server further processes the preprocessed interactive text to conform to the input requirements of the NLTK library. This may include converting the text into a specific part-of-speech tagging format (e.g., splitting the text into a list of words). Then, the server initializes the NLTK library, loading relevant corpora and tools, such as part-of-speech taggers.

[0072] The NLTK library performs part-of-speech tagging on each word in the interactive text. For example, for the sentence "I want to know what the child holding the red balloon in the video is doing," the NLTK library might tag "child" as a noun, "red" as an adjective, and "balloon" as a noun, etc. Then, the server uses the tagging results to filter nouns as candidate keywords. In this example, "child" and "balloon" would be identified as candidate keywords.

[0073] The server further filters and organizes the keyword candidates identified by the NLTK library. Since "red" is an adjective describing "balloon," it will not be ultimately selected as a keyword if only nouns are considered. The final set of keywords identified by the NLTK library is obtained, such as {"child", "balloon"}.

[0074] The server merges the keyword set {"child", "red balloon", "what to do"} obtained from the large language model with the keyword set {"child", "balloon"} obtained from the NLTK library. The merged keyword set becomes {"child", "red balloon", "what to do", "balloon"}.

[0075] Because "red balloon" and "balloon" have semantic overlap, the server performs deduplication to obtain concise and accurate focused keywords. The server uses certain semantic analysis rules (e.g., "red balloon" is a more specific description of "balloon," and if "red balloon" already exists, the broader keyword "balloon" can be removed) to deduplicate the merged keyword set, ultimately resulting in the focused keyword set {"child", "red balloon", "what to do"}.

[0076] In this way, the server calls a pre-trained large language model or NLTK library to extract keywords from the interactive text, obtaining focusing keywords. This provides an important foundation for subsequent identification of focusing entities, generation of a focusing spatiotemporal scene graph, and final dialogue feedback. This step fully utilizes the semantic understanding capabilities of the large language model and the natural language processing tools of the NLTK library to ensure the comprehensiveness and accuracy of keyword extraction.

[0077] Suppose the interactive text is "I want to know if the man in the hat and the woman in the white dress in the video are talking, and what the little dog next to them is doing."

[0078] Large Language Model Analysis: After receiving this interactive text, the large language model will identify keywords such as "man wearing a hat," "woman in a white dress," "puppy," "conversation," and "what to do." This is because the large language model can understand the different entities described in the text and the possible relationships (conversational relationships) and actions (the puppy's behavior) between them.

[0079] NLTK library processing: After performing part-of-speech tagging on this interactive text, the NLTK library identified "man," "woman," and "puppy" as candidate noun keywords. After further filtering, the descriptive phrases "wearing a hat" and "wearing a white dress" were removed (because the more specific phrases "man wearing a hat" and "woman wearing a white dress" had already been identified by the large language model when determining the focus keywords), resulting in {"man," "woman," "puppy"}.

[0080] Merging and Deduplication: The keywords obtained from the large language model or NLTK library are merged into {"man wearing a hat", "woman in a white dress", "puppy", "talk", "what to do", "man", "woman"}. After deduplication, based on semantic analysis, the final focused keywords are {"man wearing a hat", "woman in a white dress", "puppy", "talk", "what to do"}.

[0081] For example, the interactive text could be, "I want to see if the young man running in the park in the video jumped over that bench."

[0082] Large language model analysis: The large language model can identify keywords such as "young people," "park," "running," "bench," and "jump." These keywords cover important elements such as people, scenes, and actions.

[0083] NLTK library processing: The NLTK library will identify "young people" and "bench" as candidate noun keywords.

[0084] Merging and deduplication: The merged keyword set is {"young people", "park", "running", "bench", "skip", "young people", "bench"}, and the deduplicated set is {"young people", "park", "running", "bench", "skip"}.

[0085] These examples of different types of interactive text scenarios demonstrate that when the server calls the large language model or NLTK library for keyword extraction, it can adapt to various complex interactive text contents and accurately extract the key keywords, thus laying a solid foundation for the entire structured video understanding process based on a multimodal large model.

[0086] In this embodiment of the invention, the step of determining the focus entity corresponding to the focus keyword in the video to be processed by combining a pre-trained first video multimodal large model can be implemented through the following example.

[0087] The focus entity corresponding to the focus keyword in the video to be processed is determined by combining the first video multimodal large model obtained by training the fine-tuned VILA model.

[0088] In this embodiment of the invention, for example, after the server completes keyword extraction, it obtains a set of focused keywords such as {"child", "red balloon", "what to do"} (continuing with the previous example). The server organizes these keywords and converts them into an input format that the first video multimodal large model (a finely tuned VILA model) can understand. This may involve encoding the keywords or constructing a specific data structure. For example, the keywords can be organized into a dictionary form containing the keywords and their related attributes (such as part-of-speech, importance weight, etc., here assumed to be a simple example with all weights set to 1): [{"keyword":"child", "part-of-speech":"noun", "weight":1}, {"keyword":"red balloon", "part-of-speech":"noun", "weight":1}, {"keyword":"what to do", "part-of-speech":"verb", "weight":1}].

[0089] Simultaneously, the video to be processed is preprocessed before being input into the VILA model. The server will perform format conversion on the video (if necessary) to ensure that the video's encoding format, resolution, etc., meet the requirements of the VILA model. For example, the video may be converted to a specific encoding format that the model can handle (such as H.264 encoding), and the video resolution may be adjusted to the model's predefined input size (such as 224x224 pixels). The server will also extract some basic metadata from the video, such as video duration (e.g., 30 seconds) and frame rate (e.g., 25fps), and include this metadata along with the video data as part of the input.

[0090] The server inputs the organized keyword-focused data structure and the preprocessed video to be processed (including metadata) into the fine-tuned VILA model. The VILA model is a multimodal model capable of processing video and text data. After being fine-tuned based on a specific task (such as the video understanding task in this scenario) after pre-training, it has the ability to analyze specific types of video content.

[0091] The VILA model first extracts features from the video to be processed. It utilizes the convolutional neural network (CNN) structure learned during pre-training to extract visual features from each frame of the video. For example, for the first frame of the video, the model identifies various visual elements in the image, such as the outline of a person, the shape of an object, and its color. Simultaneously, the model extracts textual features from the input keywords, converting them into vector representations corresponding to the video's visual features for matching and association analysis in subsequent steps.

[0092] Based on the extracted visual features of the video and the textual features of the keywords, the VILA model begins searching for entities in the video that correspond to the keywords. For the keyword "child," the model searches for objects in each frame of the video that match the characteristics of a child. It makes this judgment based on previously extracted visual features, such as the person's height, body shape (relatively small, matching the characteristics of a child), and facial features. Similarly, for the keyword "red balloon," the model searches for objects in the video frames that are red in color and have a shape that matches the characteristics of a balloon. Throughout this process, the model comprehensively considers multiple cues in the video, such as the color distribution and texture of objects, to accurately identify entities.

[0093] After a comprehensive analysis of the video, the VILA model identified entities corresponding to the focal keywords. For example, in frame 5 of the video, a short person with facial features matching a child was found holding a red balloon. The model identified this child as the focal entity corresponding to the keywords "child" and "red balloon." Furthermore, since this was a query about the child's behavior ("what they are doing"), the model continuously tracked the child's movements within the video.

[0094] Once the focus entity is identified, the VILA model labels it. The labels include the entity's category (e.g., "child"), its location in the video (e.g., represented by coordinates in each frame; in frame 5, the child's position might be represented as [100, 200, 150, 250], indicating the top-left and bottom-right coordinates of a rectangular area within the frame), and its relationship to other entities (currently only the child's relationship with the red balloon is considered). This labeled information will be used in subsequent steps, such as constructing a focus spatiotemporal scene graph.

[0095] Suppose the interactive text is "I want to know if the man in the hat and the woman in the white dress in the video are talking, and what the little dog next to them is doing." After keyword extraction, the focused keywords are {"man in hat", "woman in white dress", "little dog", "talk", "doing"}.

[0096] Model Input and Analysis: The server inputs these keywords and preprocessed videos into the VILA model. The VILA model extracts features from the video, searching for individuals matching the characteristics of "man wearing a hat," possibly based on the shape and color of the hat, as well as the man's physical features. For "woman in a white dress," it locates the person based on the color of the dress and the woman's physical features. For "puppy," it searches based on the animal's physical characteristics.

[0097] Entity Identification and Labeling: In different frames of the video, the model identified three focal entities: a man wearing a hat, a woman in a white dress, and a dog. For the man wearing a hat, the model labeled his position in the video (e.g., in frame 10, at [200, 300, 250, 350]) and the characteristics of his hat (e.g., black color, baseball cap style). For the woman, the model labeled the details of her dress (e.g., white, long) and its position. For the dog, the model labeled its breed characteristics (e.g., small breed, white fur) and its position. The model also paid particular attention to whether there was any interaction between the man and woman, and the dog's behavior ("what it did").

[0098] For example, the interactive text is "I want to see if the young man running in the park in the video jumped over that bench", with the focused keywords being {"young man", "park", "running", "bench", "jump"}.

[0099] Model Input and Analysis: After the server inputs keywords and videos into the VILA model, the model searches for people with the physical characteristics of young people who are running in the video (by analyzing the person's movements and posture to determine if they are running), and at the same time searches for park scenes (possibly by identifying typical elements in the park such as trees and grass) and benches (based on the shape of the benches and their position in the scene).

[0100] Entity Identification and Labeling: The model identifies the young man and the bench as the focused entities. For the young man, information such as his position in different frames (e.g., in frame 15, he is located at [300, 400, 350, 450]) and his running speed (estimated by analyzing changes in his position across consecutive frames) are labeled. For the bench, features such as its position, length, and color are labeled. Furthermore, the model focuses on analyzing whether the young man jumped over the bench (corresponding to the keyword "jump").

[0101] Through these detailed examples of different scenarios, we can see how the server, combined with a finely tuned VILA model, accurately identifies the focus entity in the video to be processed based on the focus keywords, providing key entity information for the subsequent structured video understanding process based on a multimodal large model.

[0102] In this embodiment of the invention, the step of inputting the video to be processed into a pre-trained second video multimodal large model to obtain a spatiotemporal scene map of the focused entity, including the focused person and the focused object, can be implemented through the following example.

[0103] The second video multimodal large model, which is pre-trained from the video to be processed, is used by frame extraction to obtain the spatiotemporal scene map of the focused entity, including the focused person and the focused object.

[0104] Based on the intersection-union ratio of the focused entities, each focused person and each focused item is determined, resulting in multiple focused spatiotemporal scene graphs that include the interaction relationships between each focused person and each focused item. Each focused spatiotemporal scene graph includes a unique identifier ID for each focused entity, used to track each focused entity on the video timeline.

[0105] In this embodiment of the invention, for example, after receiving the video to be processed, the server performs frame extraction on the video according to a prescribed frame extraction method (2 frames per second). For example, for a 30-second video, a total of 60 frames will be extracted. During the frame extraction process, the server accurately records the timestamp of each frame, which will be used to represent the time dimension information when constructing the focused spatiotemporal scene graph later.

[0106] The server sequentially inputs the extracted video frames into a pre-trained second video multimodal large model. This model, also fine-tuned with a specific dataset, is capable of processing video frames and information related to the focused entity. For each frame, the model performs a series of processing steps, including visual feature extraction and analysis related to the focused entity.

[0107] Taking the previously mentioned video containing a child and a red balloon as an example, for a video focusing on the actual child and red balloon, the model first performs initial localization on them in each frame. Using object detection algorithms, the model identifies the approximate area of ​​the child in the frame (e.g., represented by a rectangle), and also determines the position of the red balloon. Then, the model extracts the visual features of the child and the red balloon, such as the child's physical features (facial expressions, hairstyle, etc.) and the color and shape features of the red balloon.

[0108] Inter-frame comparison: As video frames are continuously input, the model begins to calculate the Intersection over Union (IoU) of the focused entities between adjacent frames. For the focused figure, the child, assuming the child is located in the coordinate region [100, 200, 150, 250] in frame 1 and [105, 203, 153, 253] in frame 2, calculates the intersection and union areas of these two rectangular regions (representing the child's position in the two frames) to obtain the IoU. If the IoU exceeds a preset threshold (e.g., 0.5), and the visual features such as appearance of the child in the two frames are similar, the model determines that the entities in the two frames belong to the same child.

[0109] Multi-entity comparison: For the red balloon as the focal object, the intersection-union (IU) ratio is also calculated. In some frames, there may be multiple objects with similar balloon shapes. The model accurately distinguishes the red balloon as the focal object through IU calculation and visual feature comparison. For example, in frame 3, there is a yellow balloon with a similar color to the red balloon, but through IU calculation and comprehensive judgment of features such as color and shape, the model can accurately identify the location and range of the red balloon.

[0110] Focusing on a child node: For a focused child identified as the same child, the model assigns it a unique identifier ID, assuming the child's ID is 1. In the focused spatiotemporal scene graph, the node information for each frame records the child's category ("child"), ID (1), and location information. For example, the node information for frame 1 is {"category":"child","id":1,"location information":[100,200,150,250]}.

[0111] Focusing on the item node: For the red balloon, assume its ID is 2. The node information in frame 1 is {"Category":"Red Balloon","id":2,"Location Information":[120,220,130,230]}. Thus, each frame of the focused spatiotemporal scene graph contains node information for the focused character and the focused item.

[0112] Relationship Analysis: When analyzing video frames, the model focuses on whether there are interaction relationships between the focused person and the focused object. For example, if a child is always holding a red balloon, this holding relationship is an interaction relationship. In the focused spatiotemporal scene graph, this relationship is represented by edge information. In the edge information of each frame (assuming it starts from frame 1), [1,"holding",2] is recorded, indicating that the child with ID 1 is holding the red balloon with ID 2.

[0113] Handling Complex Relationships: In more complex scenarios, such as relationships between multiple focused characters and objects. Suppose the interactive text mentions an interaction between a man (focused character, ID 3) and a child (ID 1) in a video. The model will detect their interactions while analyzing video frames. If, in a frame (e.g., frame 10), the model finds that the man hands a toy to the child (a new focused object, ID 4), then the edge information in frame 10 will add interaction relationship records such as [3,"handed over", 1] and [1,"received", 4].

[0114] Temporal integration: As video frames are processed, the server integrates the node and edge information corresponding to each frame to form a complete spatiotemporal scene graph. A spatiotemporal scene graph represented in JSON format might look like this:

[0115]

[0116]

[0117] Examples of constructing focused spatiotemporal scene graphs in different scenarios:

[0118] 1. Scenario 1: The relationship between vehicles and pedestrians in a traffic scenario:

[0119] Video frame extraction and input: Assume the video to be processed is a video of a traffic intersection, and the interactive text is "I want to know the relationship between the pedestrian in the yellow shirt and the car that ran the red light in the video". The server extracts video frames and inputs them into a second video multimodal large model.

[0120] Entity identification: The focus entities are the pedestrian wearing yellow clothes (ID 1) and the car running a red light (ID 2). The model accurately identifies the positions of the pedestrian and the car in each frame by calculating the intersection-union ratio. For example, in frame 1, the pedestrian is located at [50, 100, 80, 130], and the car is located at [200, 300, 250, 350].

[0121] Scene graph construction: When constructing the focused spatiotemporal scene graph, node information records the category, ID, and location of pedestrians and cars. Edge information reflects the relationship between them. If a car is found to be approaching a pedestrian in frame 5 and the pedestrian makes a dodge action, the edge information will record [2,"approach",1] and [1,"dodge",2].

[0122] 2. Scene Two: Character Interaction in an Indoor Scene:

[0123] Video frame extraction and input: For a video of an indoor party, the interactive text is "I want to know if there is any interaction between the man wearing glasses and the woman holding the cake in the video." The server extracts frames and inputs them into the model.

[0124] Entity identification: The focus is on the man wearing glasses (ID 1) and the woman holding a cake (ID 2). The model determines their positions in each frame, such as the man being located at [100, 150, 130, 180] and the woman being located at [180, 230, 210, 260] in frame 1.

[0125] Scene graph construction: The node information of the spatiotemporal scene graph includes the character's category, ID, and location. Edge information records interaction relationships. If a man is found smiling at a woman in frame 10, the edge information will be increased by [1,"smiling gesture",2].

[0126] Through these detailed examples of different scenarios, we can see how the server uses frame extraction to input the video to be processed into the second video multimodal large model, determines the focused people and objects based on the intersection-union ratio of the focused entities, and constructs a focused spatiotemporal scene graph containing interactive relationships, thereby providing rich structured information for subsequent dialogue feedback.

[0127] In this embodiment of the invention, the focused spatiotemporal scene graph is represented in JSON format text form. The focused spatiotemporal scene graph includes frame information, node information and edge information. The node information represents the focused entity and its position, and the edge information represents the interaction relationship or spatial relationship between the focused entities.

[0128] In this embodiment of the invention, for example, when constructing a focused spatiotemporal scene graph, the server first determines to organize the data in JSON format. JSON (JavaScript Object Notation) is a lightweight data-interchange format, well-suited for representing this type of structured information. The entire focused spatiotemporal scene graph is wrapped in a large JSON object.

[0129] Frame Number: Each frame in the video has a corresponding key-value pair representing its information. For example, "Frame 1", "Frame 2", etc., serve as keys, and their corresponding values ​​are objects containing node and edge information. This frame number is determined according to the video playback order, starting from the first frame and incrementing sequentially.

[0130] The significance of frame order: The order of frames is crucial in representing the temporal information of a video. It reflects the state changes of the focused entity during video playback. For example, in a surveillance video, the first frame might show a person (the focused entity) standing at a doorway (represented by position information in node information), while as the frames progress, this person might begin to move and interact with other entities (such as doors, surrounding objects, etc.) in subsequent frames (represented by edge information).

[0131] Character Entity: Taking an office scene video as an example, suppose the interactive text is "I want to know the movement trajectory of the employee holding the file in the office and their relationship with surrounding objects." In this scene, there is a focal character: the employee holding the file. In the focal spatiotemporal scene graph, for this employee (assuming their unique identifier ID is 1), the node information in a certain frame (e.g., frame 5) might be as follows:

[0132] {"Category":"Employee","id":1,"Location Information":[100,200,150,250]}. Here, "Category" clearly identifies the entity as "Employee," "id" uniquely identifies the entity across the entire video timeline, and "Location Information" uses coordinates (assuming these are relative to the video frame, such as the coordinates of the top-left and bottom-right corners of a rectangular area within the frame) to represent the employee's position within the frame.

[0133] Item Entities: In this office scene, there are also some focused items, such as the desk (let's assume ID 2) and the filing cabinet (let's assume ID 3). For the desk, the node information in frame 5 might be {"Category":"Desk","id":2,"Location Information":[300,100,400,200]}; for the filing cabinet, the node information might be {"Category":"Filing Cabinet","id":3,"Location Information":[500,150,600,300]}.

[0134] Person Movement: As the video plays, the position of the person in focus (employee) changes. For example, in frame 10, the employee walks to their desk, and their node information changes to {"Category":"Employee","id":1,"Location Information":[250,150,300,200]}. This change in location information accurately reflects the employee's movement trajectory in the video.

[0135] The state of the items remains unchanged: For items such as desks and filing cabinets, their positions are relatively fixed in the video in most cases (unless there are special circumstances, such as being moved). Therefore, in different frames, their node information, apart from the frame number, may remain unchanged in terms of the entity's category, ID, and location information (assuming no position movement or other changes occur). For example, in frame 10, the node information of the desk is still {"Category":"Desk","id":2,"Location Information":[300,100,400,200]}.

[0136] Interaction between people and objects: Continuing with the office scene example above, in frame 10, when an employee walks to their desk, there might be an interaction where they place a file on the desk. In the focused spatiotemporal scene graph, this interaction is represented by edge information. The edge information takes the form of triplets, and in frame 10, the edge information will contain [1,"Place the file",2], where 1 is the employee's ID, 2 is the desk's ID, and "Place the file" describes the interaction between the employee and the desk.

[0137] Interactions between characters: Suppose there is another employee in the office (ID 4), and at some point in the video (e.g., frame 15), the two employees have a conversation. Then, the side information in frame 15 will be added with a record like [1,"talking with...",4], indicating the conversation relationship between employee ID 1 and employee ID 4.

[0138] Relative Positional Relationships: In addition to interaction relationships, edge information can also represent spatial relationships between focused entities. For example, in frame 5, there is a spatial relationship between the employee (ID 1) and the filing cabinet (ID 3), with the employee to the left of the filing cabinet. This spatial relationship can be represented in the edge information as [1,"to the left of...", 3]. Recording this spatial relationship helps to more comprehensively describe the layout relationships between entities in the scene.

[0139] Complex Spatial Relationships: In some scenarios, there may be complex spatial relationships between multiple entities. For example, in a conference room scene, there is a conference table (ID 5), chairs (multiple, one of which is assumed to be ID 6), and a projector (ID 7). In frame 1, the chairs are arranged around the conference table, and the projector is at one end of the conference table. Edge information may contain triples representing spatial relationships such as [6,"around",5] and [7,"at one end",5].

[0140] Below is a JSON representation example of a relatively complete spatiotemporal scene diagram built based on the above office scenario:

[0141]

[0142]

[0143]

[0144] This focused spatiotemporal scene graph fully records the state (node ​​information) of the focused entities (employees, desks, filing cabinets, etc.) in the office scene in different frames, as well as the relationships between them (edge ​​information). It is represented in a structured, easy-to-understand and process JSON format, providing a rich information foundation for the server to provide dialogue feedback based on the user's interactive text.

[0145] In this embodiment of the invention, the step of providing dialogue feedback to the interactive text based on the focused spatiotemporal scene graph can be implemented through the following example.

[0146] The focused spatiotemporal scene graph is analyzed to extract key video details and timing information;

[0147] The extracted information is input into a pre-trained dialogue generation model to generate a response text to the interactive text.

[0148] In this embodiment of the invention, for example, we assume that we continue to use the previously mentioned example of a focused spatiotemporal scene diagram of an office scene video.

[0149] After receiving this focused spatiotemporal scene diagram (represented in JSON format), the server begins parsing it.

[0150] Frame information parsing: The server first reads the information of each frame sequentially. Starting with "Frame 1," the server knows that this is the initial state of the video. In this frame, the node information shows an employee (id=1, location information [100,200,150,250]), a desk (id=2, location information [300,100,400,200]), and a filing cabinet (id=3, location information [500,150,600,300]). The edge information indicates that the employee is to the left of the filing cabinet. Through this parsing, the server obtains the crucial information of the initial positions of each focused entity at the beginning of the video and the spatial relationships between them.

[0151] Analysis of Entity Movement and Relationship Changes: As subsequent frames are analyzed, such as "frame 10," the server observes that the employee's position changes to [250, 150, 300, 200], and the edge information increases to [1, "Place file", 2]. This indicates that the employee moved near the desk and placed a file in frame 10. The server is able to clearly extract this change in entity position and the newly generated interaction relationship, which are important video details.

[0152] Timing information integration: By sequentially parsing the information in each frame, the server constructs the timing information of the entire video. For example, it knows that the employee is first on the left side of the filing cabinet, then moves to the desk and places the file, and finally (assuming there are more actions in subsequent frames) other interactions may occur. This timing information is crucial for accurately responding to the user's interactive text.

[0153] Let's look at another example of a traffic scene video. The interactive text is "I want to know the interaction between the car that ran the red light and the surrounding vehicles and pedestrians."

[0154] Analyzing the initial states of vehicles and pedestrians: In the focused spatiotemporal scene graph, the node information of "frame 1" shows the position information of the car (assuming id=1), other vehicles (such as id=2, id=3, etc.) and pedestrians (assuming id=4, id=5, etc.). The edge information may show some initial spatial relationships, such as the front-back, left-right position relationship of the car and other vehicles, and the distance relationship between the car and nearby pedestrians.

[0155] Analyzing the interactions in a traffic incident: In subsequent frames, a car might run a red light. The server analyzes that around frame 10, the car crosses the stop line (determined by changes in the car's position). Simultaneously, side information might indicate the car swerving to avoid other vehicles (e.g., [1,"leading to...emergency avoidance", 2]) or dangerously approaching pedestrians (e.g., [1,"approaching", 4]). The server accurately extracts this crucial video detail—the interactions between the car running the red light and surrounding vehicles and pedestrians—from these frames and constructs complete temporal information based on the frame order.

[0156] Input dialogue generation model: The server extracts information from the spatiotemporal scene map of the office scene video, including employee position changes, interactions with desks and filing cabinets, and the temporal information of the entire video, and inputs it into a pre-trained dialogue generation model. This dialogue generation model is trained on a large amount of text-video data pairs and can generate reasonable response text based on the relevant video information.

[0157] Generate response text: If the interaction text is "What did the employee with the file do near the desk?", the dialogue generation model generates response text based on the input information: "In the video, the employee with the file is initially on the left side of the filing cabinet, then moves to the desk and places the file on the desk at frame 10." This response accurately answers the user's question about the employee's behavior near the desk and is generated based on information extracted after parsing the focused spatiotemporal scene graph.

[0158] Input dialogue generation model: For traffic scene videos, the server inputs the extracted interaction information and time sequence information between the car running a red light and surrounding vehicles and pedestrians into the dialogue generation model.

[0159] Generate response text: If the interaction text is "Did the car that ran the red light affect other vehicles and pedestrians?", the dialogue generation model may generate the response: "Yes, when the car ran the red light and crossed the stop line around frame 10, it caused vehicle number 2 to swerve and approached pedestrian number 4, affecting the surrounding vehicles and pedestrians."

[0160] Data Source: The server acquires data from a large dataset of videos and corresponding interactive text for model training. These videos cover various scenarios, such as indoor scenes, outdoor scenes, and traffic scenes, while the interactive text consists of various questions posed in response to these videos, such as questions about human behavior and object relationships.

[0161] Data preprocessing: For video data, the server performs preprocessing operations such as format standardization and resolution adjustment to make it suitable for the model's input requirements. For interactive text, preprocessing such as keyword extraction and semantic analysis is performed to convert it into a format that the model can understand.

[0162] Determine shared data: From the prepared training data, select a portion of the data as shared data between the first and second video multimodal large models. For example, some basic video feature annotation data (such as annotations of common objects in the video, simple human action annotations, etc.) are helpful for both models to understand the video content, so this data is set as shared data.

[0163] Use of shared data: When the first video multimodal large model is trained for focusing entity identification, the video feature annotations in the shared data can help the model better identify key elements in the video, thereby accurately identifying the focusing entity corresponding to the focusing keyword. Similarly, when the second video multimodal large model is constructing a focusing spatiotemporal scene graph, this basic annotation information in the shared data helps the model more accurately locate the focusing person and object, as well as analyze the relationships between them.

[0164] Data subset partitioning: The server divides the entire training data (including shared data and individual data sets) into multiple subsets. For example, the data is divided into three subsets: A, B, and C.

[0165] Cross-validation process:

[0166] When the first video multimodal large model is trained using A and shared data as the training set and B as the validation set, the second video multimodal large model is trained using B and shared data as the training set and A as the validation set. In this way, the two models are trained under different training-validation combinations.

[0167] During training, the model adjusts its parameters based on its performance on the validation set (evaluation metrics such as accuracy and recall). For example, if the first video multimodal large model has a low accuracy in identifying the focused entity on the validation set B, the model will adjust its internal neural network weights and other parameters to improve its ability to identify the focused entity.

[0168] As the number of training rounds increases, this cross-validation process is repeated, using different subset combinations for training and validation each time. For example, in the next round, the first video multimodal large model can use B and the shared data as the training set and C as the validation set, while the second video multimodal large model can use C and the shared data as the training set and B as the validation set.

[0169] Continuous optimization: Through repeated cross-validation, the parameters of the two models are continuously optimized. This joint training approach enables the first and second video multimodal large models to work together better, improving the overall performance of the structured video understanding method based on multimodal large models. For example, after multiple cross-validation training sessions, the first video multimodal large model can more accurately identify the focus entity, providing more precise input to the second video multimodal large model. This results in a more accurate spatiotemporal scene map generated by the second video multimodal large model, ultimately improving the accuracy of dialogue feedback.

[0170] In this embodiment of the invention, the following implementation methods are also provided.

[0171] The first video multimodal large model and the second video multimodal large model are jointly trained. The first video multimodal large model and the second video multimodal large model share some training data, and the model parameters are continuously optimized through cross-validation.

[0172] In this embodiment of the invention, for example, in order to jointly train the first video multimodal large model and the second video multimodal large model, the server first needs to collect a large amount of video and corresponding interactive text data. These data come from a wide range of sources, such as obtaining videos from publicly available video datasets (e.g., Kinetics-400, UCF-101, etc.), manually annotated question-and-answer datasets, or collecting interactive text from various question-and-answer platforms via web crawlers.

[0173] Video Data Preprocessing: The server performs a series of preprocessing operations on the collected video data. First, format conversion is performed, unifying videos of different formats (such as MP4, AVI, etc.) into a format that the model can process, for example, converting all videos to H.264 encoding. Next, the video resolution is adjusted, as videos from different sources may have different resolutions. The server uniformly adjusts them to a resolution suitable for the model input, such as 224x224 pixels, or other sizes determined according to model requirements. Simultaneously, the server extracts metadata from the videos, such as frame rate, duration, and video encoding parameters, and associates this information with the video data for use as auxiliary information during training.

[0174] Interactive text data preprocessing: For interactive text data, the server first performs text cleaning, removing special characters, redundant spaces, and other elements that affect text processing. Then, lexical analysis is performed, such as word segmentation, breaking down each sentence into words or phrases, which aids in subsequent keyword extraction and semantic understanding. Next, the server performs part-of-speech tagging on the interactive text, determining the part of speech (e.g., noun, verb, adjective), which is crucial for understanding the semantic structure of the text. Furthermore, the server categorizes the interactive text according to predefined rules, such as classifying text based on question type (e.g., questions about entity relationships, entity behaviors, etc.), to provide more targeted supervision information for the model during training.

[0175] The server determines which data from the processed video and interactive text data will be shared between the first and second video multimodal large model. The selection of shared data is based on an analysis of the functional requirements of both models. For example, labeled data of basic visual elements in the video (such as common object categories and basic human actions) is very suitable as shared data. This is because the first video multimodal large model needs to identify these basic elements when determining the focus entity, and the second video multimodal large model also needs these basic elements to accurately identify the focus person and object and the relationships between them when constructing the focus spatiotemporal scene graph.

[0176] Taking a dataset containing indoor scene videos as an example, the video annotations in the shared data might include annotations for furniture (such as tables, chairs, sofas, etc.) and people (such as basic actions like standing, sitting, and walking) appearing in the videos. This annotation information is stored in a structured format, such as storing the object category, location coordinates, and person actions for each video frame in JSON format. For the shared data in the interactive text portion, it might consist of general question templates related to basic visual elements, such as "Where is the table in the video?" or "What action is the person taking?" This shared data is compiled into a dedicated dataset for use by both models during training.

[0177] The server divides the entire training dataset (including shared data and individual data) into multiple subsets. For example, assuming there are a total of 1000 video-interactive text data pairs, the server divides them into four subsets: A, B, C, and D, each containing 250 data pairs.

[0178] First Video: Training and Validation of Multimodal Large Models

[0179] Training Set Selection: The server selects a subset A and shared data as the training set for the first video multimodal large model. On this training set, the model learns to identify the focus entity based on keywords in the video content and interactive text. For example, for an indoor scene video, when the interactive text mentions "find the person holding the cup," the model learns to identify the person holding the cup in the video as the focus entity by learning from the video-text pairs in the training set.

[0180] Validation Set Selection and Evaluation: The server uses subset B as the validation set for the first video multimodal large model. During validation, the model processes video-interactive text pairs in the validation set based on knowledge learned from the training set to identify focus entities. The server then calculates evaluation metrics such as precision and recall by comparing the model-identified focus entities with manually labeled correctly identified focus entities. For example, if the model correctly identifies 80% of the manually labeled focus entities, the precision is 80%. If the model identifies focus entities covering 90% of the entities that should be identified, the recall is 90%.

[0181] Second video: Training and validation of a multimodal large model:

[0182] Training Set Selection: The server selects subset B and shared data as the training set for the second video multimodal large model. On this training set, the model learns to construct a focused spatiotemporal scene graph based on video content and identified focal entities. For example, for the focal entity (the person holding a cup) in the aforementioned indoor scene video, the model learns to accurately track the person's position in different frames of the video and identify their relationships with surrounding objects (such as tables, chairs, etc.), constructing a focused spatiotemporal scene graph that includes node information (the category, location, and ID of the person and object) and edge information (the interaction or spatial relationship between the person and the object).

[0183] Validation Set Selection and Evaluation: The server uses subset A as the validation set for the second video multimodal large model. During validation, the model constructs a spatiotemporal scene graph of the video-interaction text pairs in the validation set based on the knowledge learned on the training set. The server then compares the model-generated spatiotemporal scene graph with a manually annotated correctly labeled spatiotemporal scene graph. The comparison includes the accuracy of node information (e.g., whether the entity category and location are correct) and the accuracy of edge information (e.g., whether the relationship between people and objects is correct). The performance metrics of the model, such as the structural similarity metric (SSIM), are determined by calculating the matching degree of these contents.

[0184] First Video Multimodal Large Model Parameter Tuning: If the accuracy of the First Video Multimodal Large Model on the validation set B is low, for example, below 70%, the server will analyze the types of errors the model makes when processing the data in the validation set. If it is found that inaccurate identification of certain object categories (such as small ornaments) is causing the entity identification error, the server will adjust the parameters of the neural network layers in the model related to object category identification. For example, the weights of some convolutional kernels in a convolutional neural network (CNN) may be adjusted so that the model can better identify these easily confused object categories, thereby improving the accuracy of entity identification.

[0185] Second-stage video multimodal large-scale model parameter tuning: For the second-stage video multimodal large-scale model, if the structural similarity index on the validation set A is low, for example, below 0.6, the server will check for problems in the model's construction of the focused spatiotemporal scene graph. If it is found that the judgment of the interaction relationship between people and objects is inaccurate (such as judging "close" relationship as "far away"), the server will adjust the parameters of the neural network module in the model used to analyze the interaction relationship. For example, the weight parameters in the Long Short-Term Memory (LSTM) network or the attention mechanism module will be adjusted to improve the model's ability to understand and judge the interaction relationship.

[0186] Dataset selection for the next round of cross-validation: In the second round of cross-validation, the server changes the selection of the training and validation sets. The first video multimodal large model uses subset C and the shared data as the training set, and subset D as the validation set; the second video multimodal large model uses subset D and the shared data as the training set, and subset C as the validation set. This allows the model to be trained and validated on different data subsets, ensuring that the model can generalize to different types of video-interactive text data.

[0187] Continuous optimization process: As the number of cross-validation rounds increases, both the first and second video multimodal large-scale models continuously adjust their parameters. For example, after multiple rounds of cross-validation, the accuracy of the first video multimodal large-scale model in identifying the focused entity may improve to over 90%, while the structural similarity index of the second video multimodal large-scale model in constructing the focused spatiotemporal scene graph may improve to over 0.8. This continuous optimization process enables the two models to work better together after joint training, improving the overall performance of the structured video understanding method based on multimodal large-scale models.

[0188] Surveillance scenario: In the joint training scenario of surveillance video, the interactive text may be about the activity trajectory of a specific person (such as a suspect wearing a specific color of clothing) in the monitored area and its relationship with the surrounding environment (such as entrances and exits, vehicles, etc.).

[0189] Training Process: During joint training, the shared data includes annotations of common objects in surveillance scenarios (such as cameras, fences, etc.) and basic human actions (such as walking, running, etc.). Through cross-validation, the first video multimodal model learns to accurately identify the suspect as a focused entity, while the second video multimodal model learns to construct a focused spatiotemporal scene graph that includes the relationship between the suspect and the surrounding environment. For example, on the validation set, the first video multimodal model can accurately locate the suspect in complex surveillance scenes, and the second video multimodal model can accurately construct the spatial and interactive relationships (such as approaching the entrance / exit, entering the vehicle, etc.) between the suspect and entrances / exits and vehicles at different times.

[0190] Optimization Results: After multiple rounds of joint training, the performance of the two models was significantly improved when processing video-interactive text pairs related to surveillance scenes. In practical applications, they can answer user questions about the content of surveillance videos more quickly and accurately, such as "Did the suspect pass by a certain vehicle?" or "Where did the suspect appear near which entrance / exit?"

[0191] Educational Scenario: In educational video scenarios (such as online course videos), interactive text may be about the interaction between teachers and students in the classroom, as well as the use of specific teaching tools (such as blackboards, projectors, etc.).

[0192] Training Process: The shared data includes annotations of common characters (teachers, students), teaching tools, and basic interactive actions (such as asking questions, answering questions, and writing) in educational scenarios. During cross-validation, the first video multimodal model accurately identifies the focal entities such as teachers and students, while the second video multimodal model constructs a focal spatiotemporal scene graph containing the relationships between teachers, students, and teaching tools. For example, on the validation set, the model can accurately identify the interaction between the teacher and students when the teacher uses a projector to display content at a certain moment (such as asking students questions about the projected content).

[0193] Optimization Results: As joint training progresses, the two models are able to provide more detailed and accurate answers when processing video-interactive text pairs related to educational scenarios. For example, for the interactive text "How do students react when the teacher is writing on the blackboard?", the model can accurately answer based on the optimized performance, describing the students' concentration level, whether they asked questions, and other relevant information.

[0194] Through this joint training approach, the first and second video multimodal large models share some training data and continuously optimize model parameters through cross-validation, thereby improving the performance and accuracy of the entire structured video understanding method based on multimodal large models in different scenarios.

[0195] To more clearly describe the solutions provided in the embodiments of the present invention, a more complete implementation method is provided below.

[0196] Prerequisite definition:

[0197] Scene graph: A structured data representation used to describe objects in an image or scene and the relationships between them. Scene graphs are typically represented as "graphs," where "graph" refers to a topological structure. A graph consists of nodes and edges. Nodes in a scene graph represent objects or entities in the image, such as people, cars, and trees. Each node contains category and location information (usually represented by a bounding box). Edges in a scene graph represent the relationships between two nodes. These relationships can be spatial (e.g., up, down, left, right) or interactive (e.g., "holding").

[0198] Spatiotemporal scene graphs: The scene graphs mentioned above are for static images. To extend scene graphs to video analysis, spatiotemporal scene graphs were developed. Compared to scene graphs, spatiotemporal scene graphs add the time dimension. They need to track the position and relationship changes of all nodes in the video. Therefore, while recording the category of each node, a unique ID needs to be recorded for each node, which is used throughout the entire timeline (similar to the ID in video tracking tasks). Different objects have different IDs, but the ID of the same object in the video remains unchanged.

[0199] Text-based spatiotemporal scene diagrams: These diagrams are represented using plain text. Although in text format, they can be formatted in a structured way, similar to JSON. For example:

[0200] {

[0201] Frame 1: {Node: [{Category: Person, id: 1, Location: [11, 22, 33, 44]}, {Category: Tree, id: 2, Location: [15, 25, 35, 45]}]}], Edge: [[1, "leaning against", 2]]

[0202] }

[0203] Here, "edge" is represented in the form of triples. In the example above, the triple means that the entity with id 1 is attached to the entity with id 2.

[0204] Focused Spatiotemporal Scene Graph: Traditional scene graphs usually try to include all entities in the image, while the focused spatiotemporal scene graph defined here refers to a spatiotemporal scene graph that focuses only on all entities that the user is concerned with and entities that have interactive relationships with these entities, while ignoring other entities in the video (which can be regarded as the background).

[0205] Note: Content not described in detail is prior art known to those skilled in the art. Specific implementation examples:

[0207] Note: This specific embodiment refers to the steps of using each module for inference after all modules have been trained.

[0208] This invention discloses a structured video understanding method based on a multimodal large model. A concrete example of its overall process is as follows: Figure 1 As shown. Taking a video and a user's question as input, the code sequentially passes through three modules: a focused information generation module, a focused spatiotemporal scene graph generation module, and a dialogue generation module, ultimately generating a highly accurate response text. The specific implementation steps are as follows:

[0209] Step 1: Obtain a piece of text, which comes from the user's input and is usually a user instruction indicating the completion of a task or the request for help. Obtain a piece of video, which also comes from the user's input. The user's input text is related to the input video, and the user expects to obtain some information or results related to the input video through this method.

[0210] Step 2: Input the video and text obtained in Step 1 into the focus information generation module. The main function of this module is to output the people or things worth paying attention to in the video based on the text and video input by the user, and output them in the form of text. Different attributes are used to represent different people or things, such as "the person wearing green clothes" and "the person wearing red clothes" to distinguish two different people.

[0211] Step 3: Input the video obtained in Step 1 and the focus information obtained in Step 2 into the focus spatiotemporal scene graph generation module. The main function of this module is to generate a focus spatiotemporal scene graph in text form based on the video content and the information that needs to be focused.

[0212] Step 4: Input the user input text and video obtained in Step 1, as well as the focused spatiotemporal scene graph obtained in Step 3, into the dialogue generation module. The focused spatiotemporal scene graph provides relatively rich video details and timing information. This module will use this information to better understand the video content and finally generate a response text for the user.

[0213] Detailed description of the solution:

[0214] This invention comprises three modules, which work together in use. The following will describe in detail the relevant content of each module.

[0215] The focus information generation module is responsible for generating noteworthy people and objects from the video based on user-input text and video, outputting them as text to provide conditional information for subsequent modules. First, the user-input text enters the noun extractor to extract nouns. Noun extraction is a crucial step in this module because nouns typically represent important entities in the text, such as people, objects, and locations, which are often the focus of attention in the video content. There are many ways to implement a noun extractor, and users can choose the most suitable technology based on their needs. One simple and effective implementation is through large language models. Large language models (such as the GPT series) can accomplish this task by setting appropriate prompts, accurately extracting the nouns of interest from the text. However, a potential problem with this method is its slow processing speed, especially when processing long texts or requiring fast responses, which may affect overall efficiency. To improve processing speed, another common implementation is to use the NLTK library for noun extraction. NLTK (Natural Language Toolkit) is one of the most popular natural language processing libraries in Python, providing a rich set of language processing tools. The NLTK library allows users to perform part-of-speech tagging and extraction on text, including noun identification and extraction. Compared to large language models, NLTK has a significant advantage in processing speed, especially when dealing with large-scale texts, where it can extract important nouns more efficiently. However, NLTK's drawback is that its accuracy may not be as high as that of large language models, particularly when dealing with complex sentence structures and texts with ambiguous meanings, where its noun extraction accuracy may decrease.

[0216] It's important to note that the user-input text may not contain any meaningful nouns, or these nouns may not fully encompass all the people or objects that should be focused on; there may be other people or objects in the video that should also be focused on. To make this method robust, a fine-tuned video multimodal large model is used here, capable of outputting the actual people or objects that should be focused on based on the input video and potentially related nouns. Video multimodal large models are a class of models that have made significant progress in computer vision and natural language processing in recent years. They can process both video and text data simultaneously, understanding and generating relevant information that matches the video content. Any publicly available video multimodal large model can be selected for fine-tuning, such as VILA, Qwen2-VL, etc. These models have been pre-trained on large-scale video and text data, possessing strong cross-modal understanding capabilities, and therefore perform excellently in various application scenarios. Fine-tuning the video multimodal large model mainly involves constructing the corresponding dataset. One way to construct the dataset is to simulate users generating various different questions using the video multimodal large model, and then annotators label the actual people that should be focused on based on the questions and video content. The focus should be on the people and things that are strongly relevant to the question or interact with the main characters in the video, while ignoring background information.

[0217] The module focuses on generating spatiotemporal scene maps. Its function is to output a text-based spatiotemporal scene map containing all people and objects from the focused information in the video, based on the focused information and the video. The video input is processed by frame-by-frame extraction, at 2 frames per second. One implementation method is to use a large video multimodal model to generate the text-based focused spatiotemporal scene map. This large video multimodal model can be the same model mentioned in the focused information generation module, but fine-tuned for different data; alternatively, it can be the same large model fine-tuned using a dataset that combines both models. There are two reasons why the large model only generates spatiotemporal scene maps containing all people and objects from the focused information: firstly, the model has a limited input / output context length, and without this limitation, it may output excessively long text exceeding the context length limit, especially in outdoor scenes; secondly, generating too many spatiotemporal scene maps of people and objects unrelated to the main video content and the user's concerns may interfere with the generation of information about the people and objects that are truly relevant. Furthermore, while learning to generate focused spatiotemporal scene maps, the model also learns object detection, video tracking, and understanding of human interaction and spatial relationships, as these are elements explicitly included in the spatiotemporal scene map. The dataset used in this module primarily involves constructing a spatiotemporal scene map of a video. The video dataset built in the previous stage already contains focus information. Using this focus information, object detection algorithms automatically detect the position and category of people in each focused frame. Then, the relationships between people are manually labeled. Since object detection results cannot guarantee that the same person has a unique ID in different frames, meaning that the position and behavior of the same person in different frames cannot be directly tracked in video processing, an additional tracking mechanism is needed to address this issue. In video object tracking tasks, a common approach is to determine whether they belong to the same person by comparing the overlap of object detection boxes in adjacent frames. Specifically, if two detection boxes in adjacent frames have a high degree of overlap, and the targets they represent belong to the same category (e.g., both detected people), then these two detection boxes can be considered to belong to the same person in the video. In this case, a unique ID can be assigned to this person, and the person can be continuously tracked in subsequent frames. The core of this tracking method lies in calculating the overlap of detection boxes and judging the consistency of categories. The degree of overlap can usually be measured by calculating the Intersection over Union (IoU) of two bounding boxes. IoU is the ratio of the area of ​​intersection of two bounding boxes to the area of ​​their union. When the IoU exceeds a certain preset threshold (such as 0.5 or 0.7), the two boxes are generally considered to be highly overlapping. Furthermore, category consistency is also a key factor. If two overlapping bounding boxes belong to different categories (e.g., one for people and the other for vehicles), then they obviously cannot be considered the same target.This implementation method is based on a basic and reasonable assumption that the short interval between two frames in the video prevents the character from making large displacement movements.

[0218] The dialogue generation module's main function is to answer user questions based on the input video and text-based spatiotemporal scene graphs. These spatiotemporal scene graphs are a structured information representation that effectively identifies noteworthy people, objects, and their relationships within the video, providing this information to the dialogue generation model in text form. This approach allows the dialogue generation module to better understand the video content and reduces the "illusion" phenomenon that large models might produce when generating answers that don't accurately reflect the video content. One implementation method also uses a large video multimodal model, which can be trained simultaneously with the spatiotemporal scene graph generation module. Since their tasks are similar—one outputs a spatiotemporal scene graph, and the other inputs one—the dataset can be reused from previous datasets, requiring only additional annotation of the answers to each question.

[0219] The three modules described above together constitute a structured video understanding method based on a multimodal large model. This method is a robust video content understanding approach with high response accuracy. Video modalities are common in the real world and can handle complex dynamic scenes to meet intelligent needs. This method can accurately understand and parse key video content and provide high-quality response analysis, applicable to various intelligent application scenarios such as intelligent monitoring, autonomous driving, and intelligent assistants, providing users with a precise and efficient service experience. The focused spatiotemporal scene graph used is a structured information representation form with high organization and readability. This method can both generate and process and understand this structured information, through which the model can clearly demonstrate its understanding of video content, accurately capturing key people, objects, and their relationships. Structured information helps improve the model's reasoning ability, enabling it to efficiently process complex scenes and derive accurate and reasonable inferences, making the model more robust, intelligent, and applicable to a wide range of scenarios when dealing with diverse video content. The focusing strategy of this method can highlight the advantages of useful information and ignore useless background information, thereby increasing the focus on key content. This helps the large model reduce the illusion phenomenon during text generation, avoid the generation of information that does not match the actual content, and significantly reduce the amount of output text. This not only improves the accuracy and relevance of the generated text, but also reduces the consumption of computing resources, indirectly reducing usage costs and enabling a more economical operating mode, suitable for scenarios that efficiently process large-scale data.

[0220] Please refer to the following: Figure 2 , Figure 2 A structured video understanding device 110 based on a multimodal large model, provided in an embodiment of the present invention, includes:

[0221] The acquisition module 1101 is used to acquire the video to be processed and interactive text input by the user, wherein the interactive text is a description of the user's needs for the video to be processed;

[0222] The understanding module 1102 is used to extract keywords from the interactive text to obtain focus keywords, and combine them with a pre-trained first video multimodal large model to determine the focus entities corresponding to the focus keywords in the video to be processed; input the video to be processed into a pre-trained second video multimodal large model to obtain a focus spatiotemporal scene map of the focus entities including focus people and focus objects; and provide dialogue feedback to the interactive text based on the focus spatiotemporal scene map.

[0223] It should be noted that the implementation principle of the aforementioned structured video understanding device 110 based on a multimodal large model can refer to the implementation principle of the aforementioned structured video understanding method based on a multimodal large model, and will not be repeated here. It should be understood that the division of the various modules in the above device is merely a logical functional division; in actual implementation, they can be fully or partially integrated into a single physical entity, or physically separated. Furthermore, these modules can all be implemented in software through processing element calls; they can all be implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, the structured video understanding device 110 based on a multimodal large model can be a separately established processing element, or it can be integrated into a chip within the aforementioned device. Alternatively, it can be stored as program code in the memory of the aforementioned device, and called and executed by a processing element of the aforementioned device. The implementation of other modules is similar. Furthermore, these modules can be fully or partially integrated together, or implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step or module of the above method can be completed by the integrated logic circuit in the hardware of the processor element or by instructions in the form of software.

[0224] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together to implement a system-on-a-chip (SOC).

[0225] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned structured video understanding device 110 based on a multimodal large model. Figure 3 As shown, Figure 3 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a structured video understanding device 110 based on a multimodal large model, a memory 111, a processor 112, and a communication unit 113.

[0226] To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The multimodal large-model-based structured video understanding device 110 includes at least one software function module that can be stored in the memory 111 or embedded in the operating system (OS) of the computer device 100 in the form of software or firmware. The processor 112 is used to execute the multimodal large-model-based structured video understanding device 110 stored in the memory 111, such as the software function modules and computer programs included in the multimodal large-model-based structured video understanding device 110.

[0227] This invention provides a readable storage medium, which includes a computer program. When the computer program runs, it controls the computer device where the readable storage medium is located to execute the aforementioned structured video understanding device 110 based on a multimodal large model.

[0228] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.

Claims

1. A multi-modal large model based structured video understanding method, characterized in that, The method comprises the following steps: obtaining user inputted video to be processed and interactive text, the interactive text being user demand description text for the video to be processed; extracting focus keywords from the interactive text, and determining focus entities corresponding to the focus keywords in the video to be processed by combining a pre-trained first video multi-modal large model; inputting the video to be processed into a pre-trained second video multi-modal large model to obtain focus spatio-temporal scene graphs of focus characters and focus objects included in the focus entities; performing dialogue feedback on the interactive text according to the focus spatio-temporal scene graphs; the step of inputting the video to be processed into the pre-trained second video multi-modal large model to obtain the focus spatio-temporal scene graphs of the focus characters and the focus objects included in the focus entities comprises the following steps: adopting a frame extraction manner to input the video to be processed into the pre-trained second video multi-modal large model to obtain the focus spatio-temporal scene graphs of the focus characters and the focus objects included in the focus entities; determining each focus character and each focus object based on an intersection-over-union ratio of the focus entities to obtain a plurality of focus spatio-temporal scene graphs including interaction relationships between each focus character and each focus object, wherein each focus spatio-temporal scene graph includes a unique identifier ID of each focus entity for tracking each focus entity on a time line of the video.

2. The method of claim 1, wherein, the step of extracting focus keywords from the interactive text comprises the following steps: calling a pre-trained large language model and an NLTK library to extract focus keywords from the interactive text.

3. The method of claim 1, wherein, the step of determining focus entities corresponding to the focus keywords in the video to be processed by combining the pre-trained first video multi-modal large model comprises the following steps: determining focus entities corresponding to the focus keywords in the video to be processed by combining a first video multi-modal large model trained by a fine-tuned VILA model.

4. The method of claim 1, wherein, The focus spatio-temporal scene graph is in the form of a text in JSON format, the focus spatio-temporal scene graph includes frame information, node information and edge information, the node information represents the focus entities and their positions, and the edge information represents interaction relationships or spatial relationships between the focus entities.

5. The method of claim 1, wherein, the step of performing dialogue feedback on the interactive text according to the focus spatio-temporal scene graphs comprises the following steps: analyzing the focus spatio-temporal scene graphs to extract key video detail information and time sequence information; inputting the extracted information into a pre-trained dialogue generation model to generate reply text for the interactive text.

6. The method of claim 1, wherein, The method further comprises the following steps: jointly training the first video multi-modal large model and the second video multi-modal large model, wherein the first video multi-modal large model and the second video multi-modal large model share part of training data, and model parameters are continuously optimized through cross-validation.

7. An apparatus for structured video understanding based on a multi-modal large model, comprising: The method comprises the following steps: an obtaining module, configured to obtain user inputted video to be processed and interactive text, the interactive text being user demand description text for the video to be processed; An understanding module is configured to perform keyword extraction on the interactive text to obtain a focus keyword, and determine a focus entity corresponding to the focus keyword in the to-be-processed video by combining a pre-trained first video multi-modal large model; input the to-be-processed video into a pre-trained second video multi-modal large model to obtain a focus spatio-temporal scene graph of focus characters and focus objects included in the focus entity; and perform dialogue feedback on the interactive text according to the focus spatio-temporal scene graph. The understanding module is specifically configured to: The to-be-processed video is input into the pre-trained second video multi-modal large model in a frame extraction manner to obtain a focus spatio-temporal scene graph of focus characters and focus objects included in the focus entity; each focus character and each focus object are determined based on an intersection-over-union ratio of the focus entity to obtain a plurality of focus spatio-temporal scene graphs including an interaction relationship between each focus character and each focus object, and the focus spatio-temporal scene graph includes a unique identifier ID of each focus entity, which is used to track each focus entity on a time line of a video.

8. A computer device, comprising: The computer device includes a processor and a non-volatile memory storing computer instructions, and when the computer instructions are executed by the processor, the computer device performs the method in any one of claims 1-6.

9. A readable storage medium, characterized by, The readable storage medium includes a computer program, and when the computer program runs, controls a computer device where the readable storage medium is located to perform the method in any one of claims 1-6.

Citation Information

Patent Citations

  • Video analysis method and device based on time boundary perception and large language model

    CN117636217A

  • Generative large model-based natural language interactive security video retrieval system and device

    CN118551077A