Long video event prediction method, system, device and storage medium

By segmenting and building event chains for long videos, combined with common sense knowledge model, the problem of difficult-to-understand event logic chains in long videos is solved, and more reliable future event prediction is achieved.

CN119904787BActive Publication Date: 2025-07-04UNIV OF SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510391427.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The prior art is difficult to effectively understand and predict the macro event logic chain in long videos, resulting in poor event prediction results in long video scenarios.

Method used

By segmenting long videos into fragments, extracting dialogue text and character images, building event chains using visual description generation and common sense knowledge expert models, and making future event predictions based on the evolutionary pattern of events.

Benefits of technology

It realizes the macro-event level of long videos and reliable future event predictions, solving the complex problems of space-time information explosion and inter-event connections in long videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904787B_ABST
    Figure CN119904787B_ABST
Patent Text Reader

Abstract

The present invention discloses a long video event prediction method, system, device and storage medium, which are corresponding solutions. In the solutions, potential evolution patterns are mined and utilized from the highest-level macroscopic situations to guide more reliable future event predictions. Specifically: semantic abstraction and induction are performed layer by layer on a large amount of spatio-temporal information in the original long video, and then it delves into the macroscopic situation development patterns to guide future event predictions. The hierarchical framework effectively captures and refines the key semantics related to event understanding from a large amount of information, while the situation evolution pattern reveals the macroscopic trend of the future development of events. These designs effectively solve the difficulties of the explosion of spatio-temporal information in long videos and the intricate connections between events, and generate more reliable event predictions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of long - video event prediction, and in particular, to a method, system, device and storage medium for long - video event prediction. Background Art

[0002] Event prediction aims to predict the future situation of an event based on historical circumstances, which helps to identify and avoid potential risks, and thus provides strong support for event decision - making and emergency response. In the past, event prediction work often centered around the textual description of historical events. With the enrichment of information collection and interaction means, more and more events are presented in the form of videos such as live streaming. However, the characteristics of video expression, such as multi - modality, rich semantics, high noise and abstract expression, make it difficult to apply text - based analysis methods to the video scenario. Therefore, how to effectively meet the needs of event understanding and situation awareness in the video scenario has become an urgent need.

[0003] Around this need, early work often tried to understand and predict video situations by focusing on some shallow details, such as object or action recognition. However, these techniques often lack the ability of semantic abstraction and are difficult to understand the video from a more macroscopic perspective and summarize the logical chain therein. More seriously, this defect leads to the fact that these methods can often only process short videos, and lack the processing ability for long videos, which are more common in reality and have more complete and rich semantic information, and are difficult to form a significant event logical chain. This undoubtedly severely limits their application value.

[0004] With the development of large language models and vision - language technologies, the semantic understanding ability of long videos has been greatly improved. These methods usually directly understand the video in an end - to - end manner or convert the video into text and then make predictions. Although they can effectively locate details and summarize events, these methods still cannot further rise to the level of the situation to grasp the macroscopic trend of the event, and thus are prone to getting lost in the complex logical relationships between a large number of events in the long video, which greatly weakens their effectiveness.

[0005] In view of this, the present invention is specifically proposed. Summary of the Invention

[0006] The object of the present invention is to provide a method, system, device and storage medium for long - video event prediction, which can effectively understand long videos layer by layer to the macroscopic situation level and accurately predict future events based on the development law of the situation.

[0007] The object of the present invention is achieved by the following technical solutions:

[0008] A method for long - video event prediction includes:

[0009] The input original video is segmented into a number of consecutive video clips, and the dialogue text and video character images in each video clip are extracted;

[0010] Each video clip, the dialogue text and video character images in the video clip are encoded respectively, and a visual description text is generated by fusion;

[0011] The dialogue text and the corresponding visual description text in each video clip are summarized into an event, and based on a common sense knowledge expert model, a knowledge - promoted retrieval strategy is adopted to concatenate different events into a coherent and ordered event chain according to logical associations;

[0012] The evolution pattern of the situation is captured from the event chain, and the next situation stage is predicted. Then, the event chain is combined with the predicted next situation stage to predict future events.

[0013] A long - video event prediction system for implementing the foregoing method, the system includes:

[0014] A multi - modal data pre - processing unit for segmenting the input original video into a number of consecutive video clips, and extracting the dialogue text and video character images in each video clip;

[0015] A key visual description generation unit for encoding each video clip, the dialogue text and video character images in the video clip respectively, and generating a visual description text by fusion;

[0016] A knowledge - promoted event chain construction unit for summarizing the dialogue text and the corresponding visual description text in each video clip into an event, and based on a common sense knowledge expert model, adopting a knowledge - promoted retrieval strategy to concatenate different events into a coherent and ordered event chain according to logical associations;

[0017] A future event prediction unit for capturing the evolution pattern of the situation from the event chain, predicting the next situation stage, and then combining the event chain with the predicted next situation stage to predict future events.

[0018] A processing device includes: one or more processors; a memory for storing one or more programs;

[0019] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing method.

[0020] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.

[0021] As can be seen from the technical solution provided by the present invention above, semantic abstraction and induction are performed layer by layer on a large amount of spatio-temporal information in the original long video, and then it delves into the macroscopic event development pattern to guide future event prediction. The hierarchical framework effectively captures and refines the key semantics related to event understanding from a large amount of information, while the event evolution pattern reveals the macroscopic trend of the future development of events. These designs effectively solve the difficulties of the explosion of spatio-temporal information and the intricate connections between events in long videos, and generate more reliable event predictions. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a flowchart of a long video event prediction method provided by an embodiment of the present invention;

[0024] Figure 2 It is a schematic diagram of the overall architecture of a long video event prediction method provided by an embodiment of the present invention;

[0025] Figure 3 It is a schematic diagram of a long video event prediction system provided by an embodiment of the present invention;

[0026] Figure 4 It is a schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0028] First, the terms that may be used in this article are described as follows:

[0029] Descriptions using terms such as "comprising", "including", "containing", "having" or other similar semantics shall be construed as non-exclusive inclusion. For example, including a certain technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) shall be construed as not only including the explicitly listed certain technical feature element, but also including other technical feature elements well-known in the art that are not explicitly listed.

[0030] The term "consisting of" means excluding any technical feature element that is not explicitly listed. If this term is used in a claim, it will make the claim a closed type, making it not contain technical feature elements other than the explicitly listed ones, except for conventional impurities related thereto. If this term only appears in a certain clause of a claim, then it only limits the elements explicitly listed in that clause, and the elements recorded in other clauses are not excluded from the overall claim.

[0031] The following provides a detailed description of a long video event prediction method, system, device and storage medium provided by the present invention. Contents not described in detail in the embodiments of the present invention belong to the prior art well-known to those skilled in the art. Conditions not specified in the embodiments of the present invention are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. Instruments not specified as the manufacturer in the embodiments of the present invention are all conventional products that can be obtained through commercial purchase.

[0032] Embodiment 1

[0033] The embodiment of the present invention provides a long video event prediction method, as Figure 1 shown, which mainly includes the following steps:

[0034] Step 1: Divide the input original video into a number of consecutive video segments, and extract the dialogue text and video character images in each video segment.

[0035] Preferably, the input original video can be segmented based on scene changes, and then the segmentation results are connected in chronological order according to a set segment length threshold to obtain a number of consecutive video segments.

[0036] Preferably, extracting the dialogue text and video character images in each video segment includes: for each video segment, combining speaker recognition technology and audio transcription technology to identify video characters and generate corresponding dialogue texts; classifying the video character images based on face similarity, each category corresponding to a video character, sorting each category according to the appearance frequency of the video characters, retaining the top L categories, and for each retained category, randomly selecting any video character image belonging to the category.

[0037] Step 2: Encode each video clip, the dialogue text in the video clip, and the video character images respectively, and fuse them to generate visual description text.

[0038] Preferably, in this step, each video clip is encoded by a video encoder, and the video clip features are extracted; each video character image is encoded by an image encoder to extract image features; a visual description text is generated by a first large language model in combination with the video clip features, the image features, and the dialogue text in the video clip.

[0039] Step 3: Summarize the dialogue text and the corresponding visual description text in each video clip into an event, and based on a common sense knowledge expert model, adopt a knowledge-promoted retrieval strategy to concatenate different events into a coherent and orderly event chain according to logical associations.

[0040] Preferably, the common sense knowledge expert model can be a language model fine-tuned by a large-scale common sense knowledge graph.

[0041] Preferably, the step of concatenating different events into a coherent and orderly event chain based on the common sense knowledge expert model and adopting a knowledge-promoted retrieval strategy includes:

[0042] (1) For the current event, the events at previous times are called historical events; usually, based on the long video timestamp, all the times before the current event in the long video are historical events;

[0043] (2) Input each historical event into the common sense knowledge expert model to obtain the impacts of each historical event, connect each historical event with its corresponding impact and then encode, and save the encoding result in the historical event database.

[0044] In the embodiments of the present invention, the impacts of each historical event may directly cause the occurrence of the current event, and revealing this potential impact helps to improve the accuracy of subsequent discrimination of the causal relationship between the current event and the historical events.

[0045] (3) Retrieve the top k historical events most relevant to the text similarity of the current event from the historical event database, input the current event and the retrieved top k historical events into a second large language model, and classify the current event into the corresponding event chain through the second large language model.

[0046] Step 4: Capture the pattern of the evolution of the situation from the event chain, predict the next situation stage, and then combine the event chain with the predicted next situation stage to predict future events.

[0047] Preferably, a third largest language model and a fourth largest language model are set, and the fourth largest language model is fine-tuned with the data constructed by the third largest language model, and the fine-tuned fourth largest language model is used as a future event prediction unit to execute step 4; wherein, the process of fine-tuning the fourth largest language model with the data constructed by the third largest language model includes:

[0048] Collect a multi-modal data set for video understanding and narrative analysis, which condenses the video into a series of segments, each segment serving as an event with a textual description of the event.

[0049] Obtain the summary of the video and each event in the data set as the summary of the entire video narrative, and summarize the summary of the entire video narrative into the pattern of the evolution of the situation of the entire video through the third largest language model.

[0050] Input the summary of the entire video narrative, the pattern of the evolution of the situation, and the set target future event into the fourth largest language model, and set prompt words to guide the fourth largest language model to focus on the pattern of the evolution of the situation. According to the difference between the output of the fourth largest language model and the set target future event, construct a loss function to fine-tune the fourth largest language model (i.e., fine-tune the parameters of the fourth largest language model).

[0051] The above solution provided by the embodiments of the present invention, compared with the traditional method of predicting future events by identifying low-semantic information in videos such as actions and objects, explores the way of imitating the hierarchical understanding of long video events by humans, semantically abstracting and generalizing a large amount of spatio-temporal information in the original long video layer by layer, and further delving into the macroscopic situation development pattern to guide future event prediction. The hierarchical framework effectively captures and refines the key semantics related to event understanding from a large amount of information, while the pattern of the evolution of the situation reveals the macroscopic trend of the future development of the event. These designs effectively solve the difficulties of the explosion of spatio-temporal information and the intricate connections between events in long videos, and generate more reliable event predictions.

[0052] In order to more clearly present the technical solution provided by the present invention and the technical effects produced, the method provided by the embodiments of the present invention will be described in detail below with specific embodiments.

[0053] I. Overall overview of the solution.

[0054] An embodiment of the present invention provides a long - video event prediction method. It adopts a hierarchical structure corresponding to the way humans understand videos to understand long videos, and guides more reliable future event predictions by mining and utilizing potential evolution patterns from the highest - level macro - situations. First, the present invention uses a specially designed vision - large model to convert key visual information into a text representation with low noise and high semantic density, thereby efficiently extracting the detailed information necessary for understanding events. Then, the present invention uses a large - language model to summarize the low - level details into events, and adopts a knowledge - promoted retrieval strategy to concatenate all the isolated events discovered in the long video into a coherent and orderly event chain according to their internal logical relationships. Finally, the present invention designs a chain of thought to imitate the way humans predict events, training the language model to gradually refine the event chain into macro - situations, mine the potential evolution patterns of macro - situations, and combine the situation evolution patterns with specific scenarios to make future event predictions.

[0055] II. Detailed introduction of the solution.

[0056] As Figure 2 shown, four processing parts of the present invention are presented: (1) The present invention uses some professional tools to segment the original long video into appropriately - sized segments, and extracts basic information such as text dialogues and character portraits for use by subsequent modules. (2) The present invention specially designs and trains a vision - language model. Compared with ordinary video subtitle annotation models, it can generate key visual descriptions with less noise, thereby providing more accurate and concise detailed information for higher - level modules. (3) The present invention summarizes the details into events and organizes the interrelated events into a structured chain to enhance the coherence of event - level semantics. To achieve this, the present invention introduces a common - sense knowledge expert model to provide causal guidance and adopts a retrieval strategy to screen relevant events. In this way, the present invention constructs a coherent event chain to facilitate the analysis at the final situation level. (4) The present invention designs a pattern - centered chain of thought to explicitly capture potential situational evolution patterns and reason based on them. Corresponding data is collected to fine - tune the large - language model so that it masters the situational evolution patterns. Under the guidance of the situational evolution patterns, more reliable predictions of future events in the long video can be made. The various models involved in the above process will be introduced later. The flame symbol marked in the box where the model is located indicates that the corresponding model parameters need to be optimized, and the snowflake indicates that the model parameters are frozen.

[0057] The following will introduce the above four parts separately.

[0058] 1. Multimodal pre - processing.

[0059] (1.1) Video segmentation.

[0060] Considering the length of the video, the present invention first uses relevant tools to segment the long video into many short video clips based on scene changes. Considering that too short clips may lead to missing context, the present invention further connects consecutive clips in chronological order into segments, ensuring that the duration of each segment is not less than a set length (e.g., 4 minutes). It should be noted that this length is only the duration adopted in the current implementation case, and the specific length can be adjusted accordingly according to the data characteristics and requirements. Correspondingly, the number N of the finally segmented segments can also be determined according to the actual situation.

[0061] (1.2) Transcription of video character dialogues.

[0062] The dialogues between video characters are crucial for understanding video events. To facilitate the effective use of character dialogues to assist event understanding, the present invention uses professional audio transcription tools to transcribe the original video to obtain character dialogues in text form. At the same time, considering that there are often multiple characters in a long video and character identity is information indispensable for accurately understanding events, the present invention combines speaker recognition technology during transcription to clarify the identity of the speaker of each dialogue, thus complementing character identity information at the text end.

[0063] (1.3) Collection of video character portraits.

[0064] To further complement the lack of character identity information on the visual side, the present invention collects portraits of the main video characters based on the original long video. First, the present invention uses face recognition technology to collect face images in the video at a set frequency; subsequently, all the collected face images are classified based on face similarity, and each category corresponds to a character; finally, the present invention only retains the top L characters with the highest occurrence frequencies, and randomly samples one face image corresponding to a character as the character portrait for subsequent steps.

[0065] In the embodiment of the present invention, only the top L video character portraits are retained, but all the previously obtained dialogue texts are retained. The reason is that the top L video character portraits contain the identity information of the main characters. Except for the top L video character portraits, the remaining character portraits are all secondary character portraits. Injecting them into the network will weaken the identity information of the main characters. However, considering that the deletion of the dialogues of secondary character portraits is likely to cause local semantic loss and affect the understanding of segment events, all the previously obtained dialogue texts are retained.

[0066] Exemplarily, the frequency can be set to 1 frame per 5 seconds. Of course, this frequency is only the parameter adopted in this embodiment, and the specific frequency can be set according to the actual situation or experience.

[0067] Exemplarily, L can be set to 10, that is, the top 10 characters with the highest frequencies are retained. Of course, the value of L provided here is only for illustration and does not constitute a limitation. Its specific value can also be set according to the actual situation or experience.

[0068] 2. Generation of key visual descriptions.

[0069] After obtaining a clip of appropriate length, the present invention plans to generate key visual descriptions therefrom to extract the details required for event understanding. However, common video captions usually contain a lot of noise unrelated to the event, such as meaningless objects; and lack key information, such as character identities. These deficiencies may all lead to misunderstandings of the event. In order to retain key information while minimizing noise, the present invention introduces audio transcription technology to generate an audio description (AudioDescription) of the text, which is a data type that provides short audio descriptions of key visual facts during pauses in character dialogue, aiming to help visually impaired viewers appreciate the video. This specific scenario ensures that the audio description is more accurate and concise than ordinary video captions.

[0070] The present invention specifically designs a vision-language model to generate text-based audio descriptions. Since character portraits provide important references for character identities and dialogues contain important context information, the present invention introduces an additional image encoder to represent character portraits visually and combines nearby dialogues textually. Subsequently, a dedicated audio description generation dataset is introduced to train the model. Finally, high-quality visual descriptions can be generated to provide accurate and concise details for subsequent event hierarchies.

[0071] Figure 2 In the [description of the key visual description generation dashed box], the video encoder, image encoder, Q-Former (which is a new type of neural network architecture that focuses on improving information retrieval and representation learning through query mechanisms), mapping layer, and the first large language model together constitute the vision-language model. The Q-Former and mapping layer behind the video encoder and image encoder are responsible for mapping the features extracted by the corresponding encoders into a unified vector space. It should be noted that Figure 2 Only a feasible model structure example is provided. In actual applications, users can adjust it according to the actual situation or requirements.

[0072] Exemplarily, both the video encoder and the image encoder can use the EVA-CLIP model, which is a model improved based on the contrastive language-image pre-training (CLIP) technology and can extract general visual representations containing high-level semantic information from visual signals, providing a perceptual basis for a wide range of visual understanding and vision-language multimodal tasks.

[0073] 3. Knowledge - promoted event - chain construction.

[0074] In the foregoing process of the present invention, the generated text conversations and key visual descriptions constitute the detailed semantics required to understand the semantic level of events. However, simply summarizing the details into events is not sufficient to construct a coherent event - level semantics. Logically, a long video usually contains multiple event sequences, and the events in different sequences appear alternately in chronological order. If the logical connections between events are not considered, it may lead to the confusion of irrelevant events, thus affecting the clarity of the narrative.

[0075] Therefore, in addition to simply summarizing events from details, the present invention also connects interrelated events into chains to improve coherence. Although directly using a large - language model is a simple method, when dealing with a large number of events simultaneously, the complex logical relationships make it difficult to obtain satisfactory results. To solve this problem, the present invention introduces a language model fine - tuned by a large - scale commonsense knowledge graph (ATOMIC) as a commonsense knowledge expert model to provide causal - logic guidance; and adopts a retrieval strategy to narrow the scope of the event window.

[0076] First, the present invention uses a second large - language model to summarize the text conversations and key visual descriptions into current events; at the same time, each historical event is input into the commonsense knowledge expert model to obtain its possible impacts as causal relationships, which is effective in highlighting the similarities between related events; then, each historical event and its corresponding causal relationship are connected and encoded, and the encoding results are saved in the historical event database. Subsequently, the top k historical events with the highest text similarity to the current event are retrieved from the historical event database, thus significantly narrowing the scope of events. Finally, the current event and the top k retrieved historical events are input into the second large - language model to classify it into a specific event chain.

[0077] In the embodiments of the present invention, multiple event chains can be obtained (for example, it can be denoted as M, and its specific number is determined according to the actual situation). Initially, the event chains are empty. As the video timestamp advances, the events that occur are iteratively added to the event chains. For the current event to be judged, there are two possibilities based on its causal relationship with the events in the existing event chains: if it is determined that the current event has a clear causal relationship with the events in a certain event chain, the current event is added to the corresponding event chain; if the model believes that the event has no clear causal relationship with any historical event, then the event forms a new event chain alone.

[0078] 4. Pattern - based chain - of - thought prediction.

[0079] The event chain provides a coherent event-level narrative for the semantic abstraction of the situation level. At the situation level, long videos show potential situation evolution patterns (scenario evolution patterns), which can provide key guidance for future event prediction. To utilize these patterns, the present invention introduces a third large language model to capture potential situation evolution patterns from existing events and perform reasoning based on these situation evolution patterns. Therefore, the present invention designs a pattern-based chain of thought and specifically trains a fourth large language model to enhance the ability to capture scenario evolution patterns and perform reasoning. Specifically, this chain of thought divides future event prediction into three steps:

[0080] (4.1) Summarize the situation evolution patterns from the existing event chain;

[0081] (4.2) Based on the situation evolution patterns, predict the next situation stage from the historical situations;

[0082] (4.3) Combine the predicted next situation stage with the specific event scenario (i.e., the event chain obtained above) to predict specific future events.

[0083] Since it is different from the chain of thought commonly used in mathematical or symbolic reasoning, the situation evolution pattern is essentially empirical. To achieve empirical learning of situation evolution, the present invention introduces the Condensed Movies Dataset (CMD) as a guide. This dataset is a multi-modal dataset for video understanding and narrative analysis. The dataset condenses long videos into a series of segments. Each segment is accompanied by a textual event description and highlights the key points of the long video narrative. When applying the present invention, appropriate publicly available datasets can be used according to data characteristics and scenario requirements, or similar annotations can be made to the collected long videos to meet the needs of situation evolution pattern mining.

[0084] After that, for the long videos in the CMD dataset and each target future event therein, the corresponding summaries are retrieved from Internet resources as the benchmark for the entire long video narrative. Next, the present invention uses a powerful large language model (the third large language model) to summarize the summaries into the situation evolution patterns of the entire long video to ensure its accuracy. Subsequently, the present invention inputs the summaries, the situation evolution patterns, and the target future events into the fourth large language model. Since the fourth large language model actually already knows the standard answers, it only needs to trace back the specific reasoning process to reach the answers. To ensure that the reasoning process meets the expectations of the present invention, the present invention uses specially designed prompt words to guide the fourth large language model to focus on the situation evolution patterns. In this way, the present invention effectively extracts the knowledge of the situation evolution patterns and its reasoning ability from an omniscient and powerful large language model.

[0085] In the embodiments of the present invention, various models involved can be implemented using existing models. In practical applications, users can select the required models according to the actual situation or experience. Exemplarily: the key visual description generation can be developed based on the VideoLLaMA (video multi-modal large model), a visual language large model; the common sense knowledge expert model in the knowledge-promoted event chain construction can be obtained by fine-tuning the FLAN-T5 (text-to-text conversion model fine-tuned based on instructions), a language model. The first large language model can be implemented using the Llama2-7b model (an open-source model by Meta, 7b indicating 7 billion model parameters), the second large language model can be implemented by the GPT3.5-Turbo-0613 (an enhanced dialogue model released by OpenAI), a closed-source large language model, and the fourth large language model can be implemented through Llama3.1-8b-instruc (a Chinese optimized version of the open-source model Llama3 by Meta).

[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0087] Embodiment 2

[0088] The present invention also provides a long video event prediction system, which is mainly used to implement the methods provided in the foregoing embodiments, such as Figure 3 shown. The system mainly includes:

[0089] A multi-modal data preprocessing unit, configured to segment the input original video into a plurality of consecutive video segments, and extract the dialogue text and video character images in each video segment;

[0090] A key visual description generation unit, configured to encode each video segment, the dialogue text and video character images in the video segment respectively, and fuse them to generate visual description text;

[0091] A knowledge-promoted event chain construction unit, configured to summarize the dialogue text and the corresponding visual description text in each video segment into an event, and based on the common sense knowledge expert model, adopt a knowledge-promoted retrieval strategy to concatenate different events into a coherent and orderly event chain according to logical associations;

[0092] A future event prediction unit is used to capture the evolution pattern of the situation from the event chain, predict the next situation stage, and then combine the event chain with the predicted next situation stage to predict future events.

[0093] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.

[0094] Embodiment III

[0095] The present invention also provides a processing device, such as Figure 4 shown, which mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.

[0096] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.

[0097] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:

[0098] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.;

[0099] The output device can be a display terminal;

[0100] The memory can be a random access memory (RAM), or a non-volatile memory, such as a disk memory.

[0101] Embodiment IV

[0102] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiments when the computer program is executed by a processor.

[0103] In the embodiments of the present invention, the readable storage medium as a computer-readable storage medium can be set in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc.

[0104] As described above, it is only the preferred specific embodiment of the present invention. However, the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background art of the present invention, and should not be regarded as an admission or imply in any form that this information constitutes the prior art already known to those skilled in the art.

Claims

1. A long video event prediction method, characterized in that Including: Segment the input original video into a number of consecutive video clips, and extract the dialogue text and video character images in each video clip; Encode each video clip, the dialogue text and video character images in the video clip respectively, and fuse them to generate visual description text; Summarize the dialogue text and the corresponding visual description text in each video clip into an event, and based on the common sense knowledge expert model, adopt a knowledge-promoted retrieval strategy to concatenate different events into a coherent and orderly event chain according to logical associations; Capture the evolution pattern of the situation from the event chain, predict the next situation stage, and then combine the event chain with the predicted next situation stage to predict future events; It also includes: setting a third large language model and a fourth large language model, fine-tuning the fourth large language model with the data constructed by the third large language model, using the fine-tuned fourth large language model as a future event prediction unit to capture the evolution pattern of the situation from the event chain, predict the next situation stage, and then combine the event chain with the predicted next situation stage to predict future events; Among them, the process of fine-tuning the fourth large language model with the data constructed by the third large language model includes: Collect a multi-modal data set for video understanding and narrative analysis. This data set condenses the video into a series of segments, each segment serving as an event and having a text description of the event; Obtain the summary of the video and each event in the data set as the summary of the entire video narrative, and summarize the summary of the entire video narrative into the evolution pattern of the situation of the entire video through the third large language model; Input the summary of the entire video narrative, the evolution pattern of the situation, and the set target future event into the fourth large language model, and set prompt words to guide the third large language model to focus on the evolution pattern of the situation. According to the difference between the output of the fourth large language model and the set target future event, construct a loss function to fine-tune the fourth large language model.

2. The long video event prediction method according to claim 1, wherein The segmenting the input original video into a number of consecutive video clips includes: Based on scene transformation, segment the input original video, and then connect the segmentation results in chronological order according to the set segment length threshold to obtain a number of consecutive video clips.

3. A long video event prediction method according to claim 1, characterized in that, The extracting the dialogue text and video character images in each video clip includes: For each video clip, combine the speaker recognition technology and the audio transcription technology to identify the video characters and the corresponding dialogue text; Classify the video character images based on face similarity, each category corresponding to a video character. Sort each category according to the appearance frequency of the video character, retain the top L categories, and for each retained category, randomly select any video character image belonging to the category.

4. A long video event prediction method according to claim 1, characterized in that The encoding each video clip, the dialogue text and video character images in the video clip respectively, and fusing them to generate visual description text includes: Encode each video clip through a video encoder and extract video clip features; Encode each video character image through an image encoder and extract image features; Generate a visual description text through the first large language model, in combination with video clip features, image features, and dialogue text in the video clip.

5. A long video event prediction method according to claim 1, characterized in that The common sense knowledge expert model is a language model fine-tuned through a common sense knowledge graph.

6. A long video event prediction method according to claim 1 or 5, characterized in that Based on the common sense knowledge expert model, adopting a knowledge-promoted retrieval strategy, connecting different events in series according to logical associations into a coherent and orderly event chain includes: For the current event, the events at previous moments are called historical events; Input each historical event into the common sense knowledge expert model to obtain the impacts of each historical event, connect each historical event with its corresponding impact and then encode, and save the encoding result in the historical event database; Retrieve the top k historical events with the most relevant text similarity to the current event from the historical event database, input the current event and the retrieved top k historical events into the second large language model, and classify the current event into the corresponding event chain through the second large language model.

7. A long video event prediction system, characterized in that, For implementing the method described in any one of claims 1 to 6, the system includes: A multi-modal data preprocessing unit, configured to segment the input original video into a number of continuous video clips, and extract the dialogue text and video character images in each video clip; A key visual description generation unit, configured to encode each video clip, the dialogue text in the video clip, and the video character images respectively, and fuse them to generate a visual description text; A knowledge-promoted event chain construction unit, configured to summarize the dialogue text and the corresponding visual description text in each video clip into an event, and based on the common sense knowledge expert model, adopt a knowledge-promoted retrieval strategy to connect different events in series according to logical associations into a coherent and orderly event chain; A future event prediction unit, configured to capture the evolution pattern of the situation from the event chain, predict the next situation stage, and then combine the event chain with the predicted next situation stage to predict future events.

8. A processing device, characterized in that, Including: One or more processors; A memory, configured to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of claims 1 to 6.

9. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video event description and attribution generation method, system and device and storage medium

    CN117557946A

  • Method and device for automatically sensing emergencies

    CN118428375A

  • Video future event prediction method and device, storage medium and program product

    CN118823635A