Weak supervision time sequence action positioning method based on multi-modal large language model
By constructing key semantic matching and complete semantic reconstruction modules, combined with a dual-priority interactive distillation strategy, the problems of incomplete localization and over-localization in weakly supervised temporal action localization are solved, improving localization accuracy and efficiency, and is applicable to video analysis and computer vision fields.
Patent Information
- Application Number
- CN202510981116.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-31
AI Technical Summary
Existing weakly supervised temporal action localization methods suffer from incomplete localization and over-localization issues, and multimodal large language models have high computational and memory requirements, which cannot meet the needs of real-time processing.
We adopt a weakly supervised temporal action localization method based on a multimodal large language model. By constructing a key semantic matching module and a complete semantic reconstruction module, and combining a dual-prior interactive distillation strategy, we optimize the action localization process, reduce incomplete localization and over-localization, and lower computational overhead.
It significantly improves the accuracy and completeness of weakly supervised temporal action localization, reduces computing resource requirements, and supports real-time processing and long video analysis.
Smart Images

Figure CN120877183A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of video analysis, computer vision, and deep learning, and in particular to a weakly supervised temporal action localization method based on a multimodal large language model. Background Technology
[0002] Temporal Action Localization (TAL) is an important research area in computer vision, aiming to accurately locate instances of interest from undressed videos. Unlike traditional video classification tasks, TAL not only predicts the category of an action but also determines the precise time segment of the action within the video. TAL tasks are generally divided into two categories: fully supervised TAL and weakly supervised TAL.
[0003] Fully supervised temporal action localization methods rely on detailed annotation information for each video segment, including action categories and their temporal boundaries. While this approach achieves high-precision action localization, the need for precise annotation of every frame in every video significantly increases the cost and labor intensity of the annotation process. Furthermore, as video length increases, the annotation workload and computational resource requirements of fully supervised methods also increase dramatically.
[0004] Weakly supervised temporal action localization (WTAL) methods attempt to complete action localization tasks using video-level labels without requiring precise temporal labels. Video-level labels typically indicate the presence of a certain type of action in a video, but do not provide specific start and end times. By avoiding frame-level annotation for each video segment, WTAL methods significantly reduce annotation costs and can handle longer video datasets.
[0005] Existing WTAL methods typically generate temporal class activation maps (T-CAMs) by training a classifier, enabling prediction of action start and end times based on video-level labels. However, current weakly supervised methods still face two main problems: A) Incomplete localization: Some low-discrimination action instances may be ignored, leading to incomplete localization results. B) Overlocalization: Background segments may be misclassified as foreground actions, resulting in overlocalization.
[0006] Multimodal large language models (MLLMs) have demonstrated powerful video understanding capabilities in recent years, but their high computational and memory requirements make them unsuitable for real-time processing. Therefore, combining existing traditional methods with the advantages of MLLMs to improve the performance of weakly supervised temporal action localization has become a pressing issue.
[0007] It should be noted that the information disclosed in the background section above is only for understanding the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0008] The main objective of this invention is to overcome the deficiencies in the aforementioned background technology and provide a weakly supervised temporal action localization method based on a multimodal large language model.
[0009] To achieve the above objectives, the present invention adopts the following technical solution:
[0010] In a first aspect of the present invention, a weakly supervised temporal action localization method based on a multimodal large language model includes the following steps:
[0011] S1. Key semantic matching: Construct a key semantic matching module, and use the key semantic prior information generated by the multimodal large language model to match video segments and locate the action time interval;
[0012] S2. Complete Semantic Reconstruction: Construct a complete semantic reconstruction module, which uses complete semantic prior information generated by a multimodal large language model to reconstruct action descriptions, thereby enhancing the comprehensive understanding of the temporal intervals of action instances;
[0013] S3. Dual-prior interactive distillation: Construct a dual-prior interactive distillation strategy module to optimize the collaboration between the key semantic matching module and the complete semantic reconstruction module through mutual distillation, thereby reducing incomplete localization and over-localization.
[0014] S4. Training and Inference: Jointly train the key semantic matching module and the complete semantic reconstruction module, and use the trained module to directly perform temporal action localization during inference, avoiding the use of the multimodal large language model.
[0015] Further, in step S1, the key semantic matching specifically includes:
[0016] Construct key semantic prompt templates and input them into a multimodal large language model to generate key semantic descriptions of action categories;
[0017] Extract the spatiotemporal features of the video and embed them into a representation;
[0018] Generate text embedding representations of key semantic queries;
[0019] Calculate the similarity matrix between video segments and key semantics, and apply attention weights to suppress the response of background segments;
[0020] The matching loss is calculated through a multi-instance learning mechanism to optimize the action localization results.
[0021] Furthermore, in step S2, the complete semantic reconstruction specifically includes:
[0022] Construct a complete semantic prompt template and input it into a multimodal large language model to generate a complete semantic description that details the action;
[0023] Random masking is applied to key action words in the complete semantic description;
[0024] Extract video features and embed them into a representation, then apply an attention mechanism to assign weights to time segments;
[0025] Use the Transformer model to predict masked action words to reconstruct the complete semantic description;
[0026] The semantic reconstruction process is optimized by reconstructing the loss function.
[0027] Further, in step S3, the dual-priority interactive distillation specifically includes:
[0028] In the key semantic matching stage, optimization is performed by combining matching loss, mean square error of attention weights, and localization loss.
[0029] In the complete semantic reconstruction stage, optimization is performed by combining reconstruction loss and mean squared error of attention weights;
[0030] Through a mutual distillation mechanism, the output of the key semantic matching module is used as the optimization target of the complete semantic reconstruction module, and the output of the complete semantic reconstruction module is used as the optimization target of the key semantic matching module.
[0031] Furthermore, the localization loss includes a combination of a focus loss function, a DIOU loss function, and a multi-instance learning loss function to focus on key action areas and improve localization accuracy.
[0032] Further, in step S4, the training includes:
[0033] An optimization algorithm is used to jointly train the key semantic matching module and the complete semantic reconstruction module;
[0034] Minimize localization loss and reconstruction loss;
[0035] Optimize the timing action positioning accuracy through a multi-view learning mechanism.
[0036] Further, in step S4, the reasoning includes:
[0037] The input video is processed independently using the trained key semantic matching module and the complete semantic reconstruction module.
[0038] Generate action time-series interval prediction results without calling a multimodal large language model throughout the process.
[0039] Furthermore, the method also includes video input and feature extraction steps:
[0040] Extract multimodal features from the original video, including RGB features and optical flow features;
[0041] Spatiotemporal features are fused to construct a video embedding representation, providing input features for the key semantic matching module and the complete semantic reconstruction module.
[0042] In a second aspect of the invention, a weakly supervised temporal action localization system based on a multimodal large language model includes:
[0043] The key semantic matching module is used to match key semantic prior information generated by a multimodal large language model with video segments to locate the action time interval;
[0044] The complete semantic reconstruction module is used to reconstruct action descriptions using complete semantic prior information generated by a multimodal large language model, so as to enhance the comprehensive understanding of the instance order interval of the action.
[0045] The dual-priority interactive distillation strategy module is used to optimize the collaboration between the key semantic matching module and the complete semantic reconstruction module through mutual distillation, thereby reducing incomplete localization and over-localization.
[0046] The training and inference module is used to jointly train the key semantic matching module and the complete semantic reconstruction module, and to directly perform temporal action localization using the trained module during inference, avoiding the need to call the multimodal large language model.
[0047] In a third aspect of the invention, a computer program product includes a computer program that, when executed by a processor, implements the described weakly supervised temporal action localization method based on a multimodal large language model.
[0048] The present invention has the following beneficial effects:
[0049] This invention proposes a framework, MLLM4WTAL, to enhance traditional weakly supervised temporal action localization (WTAL) methods. By innovatively integrating the semantic priors of a multimodal large language model (MLLM) with a dual-module collaborative mechanism, it significantly improves the performance and practicality of weakly supervised temporal action localization. Its core advantage lies in accurately addressing the localization deficiencies of traditional methods: on the one hand, the Key Semantic Matching (KSM) module extracts core features of action categories and dynamically aligns them with video segments, effectively suppressing over-localization caused by background interference; on the other hand, the Complete Semantic Reconstruction (CSR) module reconstructs the temporal context using detailed action descriptions, compensating for the omission of low-discriminative actions and thus improving the completeness of localization. Both are collaboratively optimized through a dual-prior interactive distillation strategy (DPID)—during the training phase, attention weights and localization information are mutually distilled, allowing the modules to focus on the same action regions, further enhancing localization accuracy.
[0050] Furthermore, this invention achieves breakthroughs in efficiency and versatility. During the training phase, semantic priors generated by MLLM enhance the model's understanding capabilities, while the inference phase is completely independent of MLLM, requiring only lightweight KSM and CSR modules for localization, significantly reducing computational overhead and supporting real-time processing. This design not only overcomes the high resource consumption bottleneck of MLLM but also allows for seamless integration into existing weakly supervised frameworks.
[0051] Experimental results show that the method of this invention exhibits optimal performance on multiple public datasets and can be seamlessly integrated into existing WTAL methods, improving the accuracy and completeness of video temporal action localization and providing a highly robust solution for long video analysis scenarios.
[0052] Other beneficial effects of the embodiments of the present invention will be further described below. Attached Figure Description
[0053] Figure 1 This is a schematic diagram of the structure of an embodiment of the present invention.
[0054] Figure 2 This is the overall framework of MLLM4WTAL in this embodiment of the invention.
[0055] Figure 3 This is a comparison chart of the prediction results and the actual labels in an embodiment of the present invention. Detailed Implementation
[0056] The embodiments of the present invention will be described in detail below. It should be emphasized that the following description is merely exemplary and not intended to limit the scope and application of the present invention.
[0057] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of the present invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0058] This invention aims to address the incomplete and over-localization problems in existing weakly supervised temporal action localization (WTAL) methods, and provides a weakly supervised temporal action localization method (MLLM4WTAL) based on a multimodal large language model (MLLM). By combining key semantic and complete semantic prior information generated by the multimodal large language model (MLLM), the accuracy and robustness of weakly supervised temporal action localization (WTAL) are improved.
[0059] See Figure 1 This invention provides a weakly supervised temporal action localization method based on a multimodal large language model, comprising the following steps:
[0060] Step S1, Key Semantic Matching: Construct a key semantic matching module, and use the key semantic prior information generated by the multimodal large language model to match with video segments to locate the action time interval.
[0061] In some embodiments, step S1, the key semantic matching specifically includes: constructing a key semantic prompt template and inputting it into a multimodal large language model to generate a key semantic description describing the action category; extracting the spatiotemporal features of the video and embedding them; generating a text embedding representation of the key semantic query; calculating the similarity matrix between the video segment and the key semantic, and applying attention weights to suppress the response of background segments; and calculating the matching loss through a multi-instance learning mechanism to optimize the action localization result.
[0062] Step S2, Complete Semantic Reconstruction: Construct a complete semantic reconstruction module, and reconstruct the action description using complete semantic prior information generated by the multimodal large language model, so as to enhance the comprehensive understanding of the temporal interval of the action instance.
[0063] In some embodiments, step S2 specifically includes: constructing a complete semantic prompt template and inputting it into a multimodal large language model to generate a complete semantic description that describes the action in detail; performing random masking on key action words in the complete semantic description; extracting video features and embedding them, and applying an attention mechanism to assign time segment weights; using a Transformer model to predict masked action words to reconstruct the complete semantic description; and optimizing the semantic reconstruction process through a reconstruction loss function.
[0064] Step S3, Dual Prior Interactive Distillation: Construct a dual prior interactive distillation strategy module to optimize the collaboration between the key semantic matching module and the complete semantic reconstruction module through mutual distillation, thereby reducing incomplete localization and over-localization.
[0065] In some embodiments, step S3, the dual-prior interactive distillation specifically includes: in the key semantic matching stage, optimizing by combining matching loss, attention weight mean square error, and localization loss; in the complete semantic reconstruction stage, optimizing by combining reconstruction loss and attention weight mean square error; through a mutual distillation mechanism, using the output of the key semantic matching module as the optimization target of the complete semantic reconstruction module, and simultaneously using the output of the complete semantic reconstruction module as the optimization target of the key semantic matching module. In a further preferred embodiment, the localization loss includes a combination of a focus loss function, a DIOU loss function, and a multi-instance learning loss function to focus on key action regions and improve localization accuracy.
[0066] Step S4, Training and Inference: Jointly train the key semantic matching module and the complete semantic reconstruction module, and use the trained module to directly perform temporal action localization during inference, avoiding the use of the multimodal large language model.
[0067] In some embodiments, step S4, the training includes: jointly training the key semantic matching module and the complete semantic reconstruction module using an optimization algorithm; minimizing the localization loss and reconstruction loss; and optimizing the temporal action localization accuracy through a multi-view learning mechanism.
[0068] In some embodiments, in step S4, the reasoning includes: independently processing the input video using a trained key semantic matching module and a complete semantic reconstruction module; generating action temporal interval prediction results without calling a multimodal large language model throughout the process.
[0069] In some embodiments, the weakly supervised temporal action localization method based on a multimodal large language model further includes a video input and feature extraction step: extracting multimodal features from the original video, including RGB features and optical flow features; fusing spatiotemporal features to construct a video embedding representation, providing input features for the key semantic matching module and the complete semantic reconstruction module.
[0070] This invention also provides a weakly supervised temporal action localization system based on a multimodal large language model, comprising:
[0071] The key semantic matching module is used to match key semantic prior information generated by a multimodal large language model with video segments to locate the action time interval;
[0072] The complete semantic reconstruction module is used to reconstruct action descriptions using complete semantic prior information generated by a multimodal large language model, so as to enhance the comprehensive understanding of the instance order interval of the action.
[0073] The dual-priority interactive distillation strategy module is used to optimize the collaboration between the key semantic matching module and the complete semantic reconstruction module through mutual distillation, thereby reducing incomplete localization and over-localization.
[0074] The training and inference module is used to jointly train the key semantic matching module and the complete semantic reconstruction module, and to directly perform temporal action localization using the trained module during inference, avoiding the need to call the multimodal large language model.
[0075] The technical solution of this invention can effectively solve the problems of incomplete localization and over-localization in existing weakly supervised methods. Experimental results (which will be further detailed below) show that this invention exhibits state-of-the-art performance on multiple datasets, effectively improving the accuracy and completeness of weakly supervised temporal action localization, and providing a highly robust solution for long video analysis scenarios.
[0076] The following further describes specific embodiments, algorithm examples, and experimental verifications of the present invention.
[0077] A weakly supervised temporal action localization method and system based on a multimodal large language model (MLLM) (MLLM4WTAL) mainly includes the following core modules:
[0078] S1: Key Semantic Matching (KSM) Module: Utilizing key semantic prior information generated by a multimodal large language model (MLLM), this module matches video segments to accurately locate the time intervals of actions within the video. By comparing with various segments in the video, the KSM module can identify the most representative action segments and locate them as target action regions. The goal of this module is to determine the start and end times of each action through semantic description and video feature matching.
[0079] S2: Complete Semantic Reconstruction (CSR) Module: This module reconstructs action descriptions in videos using complete semantic prior information generated by a multimodal large language model (MLLM). By reconstructing the action words in the mask, this module restores the complete semantics of the video, ensuring the model has a more comprehensive understanding of the temporal intervals of action instances. The CSR module fills in the localization gaps caused by incomplete semantics, thereby reducing missed or over-labeling and enhancing the completeness of localization.
[0080] S3: Dual-Prior Interactive Distillation (DPID) Strategy: To optimize the collaboration between the Key Semantic Matching (KSM) and Complete Semantic Reconstruction (CSR) modules, this invention proposes a Dual-Prior Interactive Distillation (DPID) strategy. In the DPID strategy, the KSM and CSR modules optimize each other through mutual distillation. The output generated by the KSM serves as the optimization target of the CSR module, and the output generated by the CSR, in turn, serves as the optimization target of the KSM module. Through this mutual optimization, the KSM and CSR modules can cooperate more effectively, addressing the incomplete localization and overlocalization problems in existing methods.
[0081] S4: Model Training and Inference Process: During model training, the KSM and CSR modules are jointly trained by minimizing appropriate loss functions (such as localization loss and reconstruction loss). During training, the model optimizes the localization accuracy of temporal actions through multi-view learning. In the inference phase, the trained KSM and CSR modules perform action localization and generate accurate temporal action ranges.
[0082] The MLLM4WTAL method proposed in this invention improves the weakly supervised temporal action localization (WTAL) method by combining key semantic and complete semantic prior information generated by multimodal large language model (MLLM) and introducing a dual prior interactive distillation strategy.
[0083] The specific implementation steps and technical details of the method are described in detail below.
[0084] 1. System Architecture
[0085] like Figure 2 As shown, the architecture of the weakly supervised temporal action localization system based on a multimodal large language model (MLLM) includes the following key modules:
[0086] Video input and feature extraction module;
[0087] Key Semantic Matching (KSM) module;
[0088] Complete Semantic Reconstruction (CSR) module;
[0089] Dual Prior Interactive Distillation (DPID) strategy module;
[0090] Training and inference modules.
[0091] Through the collaboration of these modules, the entire system improves the performance of weakly supervised temporal action localization tasks. The implementation steps of each module are described in detail below.
[0092] 2. Video Input and Feature Extraction Module
[0093] In the initial stage of the system, the input is an untrimmed video. After preprocessing, the video is processed by the video input and feature extraction module. The task of this module is to extract useful features from the original video for use in subsequent steps.
[0094] Video input: The raw video file is input into the system via a decoder.
[0095] Feature extraction: Features are extracted from each frame and key time periods of the video using methods such as Convolutional Neural Networks (CNNs) or FlowNet. Common features include RGB image features, optical flow features, and depth features. These features enable the system to capture the spatial and temporal information of actions in the video.
[0096] Multimodal feature fusion: At this stage, information from other modalities (such as audio, subtitles, etc.) can be fused and further feature extraction can be performed through corresponding multimodal neural networks. Ultimately, the resulting feature vector includes temporal and spatial information, providing a foundation for subsequent temporal action localization.
[0097] 3. Key Semantic Matching (KSM) Module
[0098] The main goal of the Key Semantic Matching (KSM) module is to match key semantic prior information generated by a multimodal large language model (MLLM) with video segments to locate action instances in the video.
[0099] Key semantic generation: Key semantic cues are generated using a multimodal large language model (MLLM), primarily including semantic information such as action categories and key time periods. These semantic cues are generated based on the video content and describe the types of actions that may occur in the video. Cue template P key The format is as follows:
[0100] "This video contains the actions of {Cls}. Please describe the actions of {Cls} in one sentence."
[0101] Please use the prompt template P key Input the original video V into MLLM to obtain the key semantic description D of the video. key :
[0102] D key =Φ MLLM (P key ,V)
[0103] Video Embedding Module: Similar to other WTAL models, the video embedding module Φ emb It consists of two one-dimensional convolutions followed by ReLU and Dropout layers. It fuses RGB features F. RGB and optical flow characteristics F flowTo obtain the input F of the video embedding module:
[0104] F = Φ fuse (F RGB ,F flow )
[0105] Then through the video embedding module Φ emb Obtain the corresponding video embedding features F e :
[0106] F e =Φ emb (F)
[0107] In addition, attention mechanisms are used. attention Generate attention weights A for each video segment KSM :
[0108] A KSM =sigmoid(Φ attention (F e ))
[0109] Text embedding module: Text embedding module Φ trans The aim is to use key semantic descriptions to describe D key Generate key semantic query F query The [START] flag T will be randomly initialized. start Learnable textual context T context Key text features T embedded via GloVe key By concatenating these parts, we obtain the input T of the text embedding module. query :
[0110] T key =Φ GloVe (D key )
[0111] T query =[T start ,T context ,T key ]
[0112] F query =Φ trans (T query )
[0113] Semantic Matching Module: The semantic matching module aims to activate video segment features relevant to key semantic queries. Specifically, it performs an inner product operation on video embedding features and text embedding features to generate a segment-level video-text similarity matrix M. Furthermore, attention weights A are also employed. KSM To suppress background fragment responses, the final fragment-level matching result with background suppression is obtained.
[0114]
[0115] Top-k multi-instance learning is used to calculate the matching loss L. KSM Specifically, the average of the top-k similarities for each specific text query category along the response time dimension is calculated as the video-to-text similarity. For the j-th action category, the video-to-text similarity S... j and respectively composed of M and generate.
[0116]
[0117] Then, apply softmax to S j and Processing is performed to generate a video-level similarity score p. j and
[0118] p j =softmax(S j )
[0119]
[0120] Finally, regarding the true label y j and With p j and Calculate the cross-entropy loss and optimize the KSM module.
[0121]
[0122] 4. Complete Semantic Reconstruction (CSR) Module
[0123] The Complete Semantic Reconstruction (CSR) module utilizes complete semantic prior information generated by MLLM to help the system locate the complete action time interval by reconstructing missing key action words in the video. The proposed Complete Semantic Reconstruction (CSR) includes not only video embedding and text embedding modules, but also a transformer reconstructor for multimodal interaction and complete text description reconstruction.
[0124] Complete Semantic Generation: The CSR module first generates complete semantic prior information using MLLM (Multimodal Large Language Model). In this module, a complete semantic cue is first constructed using MLLM, prompting MLLM to generate complete information about the action instance. When constructing the semantic cue, it is necessary to consider the known action categories contained in the video and pre-introduce these action categories into the cue template to guide the generation of correct descriptions. Cue Template P complete The format is as follows:
[0125] "This video contains the {Cls} action. Please describe the {Cls} action in detail."
[0126] Please use the prompt template P complete Input the original video V into MLLM to obtain a complete semantic description of the video D. complete :
[0127] D complete =Φ MLLM (P complete ,V)
[0128] Video embedding module: In this module, given the original video features F, the features are embedded through a fully connected layer Φ. FC Obtain the corresponding video feature embedding F complete :
[0129] F complete =Φ FC (F)
[0130] To explore the complete video time interval related to text semantics, this embodiment employs an attention mechanism Φ attention Through this mechanism, the CSR (Complete Semantic Reconstruction) module can effectively assign weights to each temporal segment in the video, thereby closely associating video segments with textual semantics. The attention weight A of the CSR module... CSR It can be calculated using the following formula:
[0131] A CSR =sigmoid(Φ attention (F e ))
[0132] Text embedding module: This embodiment first masks the key action words in the complete video description. Specifically, by using GloVe embedding and a fully connected layer, the complete semantic embedding F of the masked sentence is obtained. c The purpose of this step is to obtain a complete semantic representation of the text data, providing a foundation for subsequent multimodal processing.
[0133] Complete semantic reconstruction: The Transformer model is used to predict missing action words, thereby recovering the complete semantic description. The specific steps are as follows:
[0134] First, randomly select video description D. complete One-third of the words are masked as alternative descriptions. Then, they are processed by the Transformer encoder Φ. trans-enc Extracting foreground video features F from video data fg This is used in the subsequent reconstruction process.
[0135] Ffg =Φ trans-enc (F complete A CSR )
[0136] Then, through the Transformer decoder Φ trans-dec Generate a multimodal representation H, and use this representation to reconstruct the mask description.
[0137] H = Φ trans-dec (F c ,F fg A CSR )
[0138] By analyzing the i-th word ω i The probability distribution is calculated to obtain the probability value of the reconstruction result.
[0139] P(ω i |F complete ,F c[0:i-1] ) = softmax(Φ FC (H))
[0140] Finally, by calculating the loss function L CSR This loss function is used to optimize the model. It guides the model to gradually adjust its parameters during the reconstruction process to achieve the best semantic reconstruction results.
[0141]
[0142] 5. Dual Prior Interactive Distillation (DPID) Strategy Module
[0143] The Dual Prior Interactive Distillation (DPID) strategy is one of the key innovations of this invention. This module addresses the challenges of traditional semantic matching and semantic reconstruction by designing a two-stage optimization process. Furthermore, it strengthens the joint collaboration between the Key Semantic Matching (KSM) and Complete Semantic Reconstruction (CSR) modules through an interactive distillation strategy, making them more focused on the same action regions in the video, thereby reducing localization errors and improving the accuracy of semantic reconstruction.
[0144] Loss function: The optimization process consists of two stages: key semantic matching and complete semantic reconstruction. By introducing a prior-guided localization head, the model can better focus on key regions in the video.
[0145] The positioning head is optimized using the following loss function:
[0146] L loc =L focal +L DIOU +L MIL
[0147] Loss function for the key semantic matching stage:
[0148] L match =L KSM +λ1MSE(A KSM ,ψ(A CSR ))+μ1L loc
[0149] Loss function for the complete semantic reconstruction stage:
[0150] L recon =L CSR +λ2MSE(A CSR ,ψ(A KSM ))
[0151] 6. Training and Reasoning Process
[0152] During the training phase, the system uses standard optimization algorithms (such as the Adam optimizer) to jointly train the KSM and CSR modules. During training, the system optimizes the accuracy of temporal action localization through multi-view learning, with the goal of minimizing localization and reconstruction losses.
[0153] During the inference phase, the trained KSM and CSR modules perform action localization on the input video. Unlike the training phase, the computationally expensive Multimodal Large Language Model (MLLM) is no longer used during inference; instead, the optimized KSM and CSR modules are directly used for temporal action localization.
[0154] Experimental results
[0155] Experimental results show that the method of the present invention exhibits optimal performance on multiple public datasets and can be seamlessly integrated into existing WTAL methods, improving the accuracy and completeness of video temporal action localization.
[0156] Table 1 shows the performance evaluation of MLLM4WTAL in the embodiments of the present invention. As can be seen from Table 1, after applying the framework proposed in this invention to DELU and Zhou et al., state-of-the-art performance was achieved, surpassing the previous state-of-the-art method PivoTAL, and even surpassing the fully supervised methods SSN and P-GCN.
[0157] Table 2 presents the performance evaluation of each module in the embodiments of the present invention. As can be seen from Table 2, after adding the key semantic matching and complete semantic reconstruction modules proposed in this invention to the baseline method in sequence, significant performance improvements were achieved, ultimately resulting in an average mAP of 39.2%.
[0158] Figure 3 This is a comparison chart of the prediction results and the actual labels in an embodiment of the present invention. Figure 3As can be seen, by incorporating the weakly supervised temporal action localization framework assisted by the multimodal large model of this invention, the localization incompleteness and overcompleteness problems of the previous baseline methods are effectively alleviated, and more accurate temporal action localization results are achieved.
[0159] Table 1
[0160]
[0161] Table 2
[0162]
[0163] In summary, this invention proposes a weakly supervised temporal action localization method and system (MLLM4WTAL) based on a multimodal large language model (MLLM). Compared with existing technologies, the key innovations and significant technical advantages of this invention include:
[0164] This invention utilizes key semantic and complete semantic prior information generated by a multimodal large language model (MLLM) to enhance weakly supervised temporal action localization (WTAL). By introducing a multimodal large language model into the weakly supervised temporal action localization method, it overcomes the problem of existing methods relying on single-modal information and improves the model's localization accuracy in complex video scenes.
[0165] This invention proposes a dual-module structure combining Key Semantic Matching (KSM) and Complete Semantic Reconstruction (CSR). The KSM module identifies action segments in a video by matching video clips with key semantic prior information, while the CSR module compensates for localization defects caused by incomplete semantics by reconstructing a complete semantic description. The combination of the two effectively improves the accuracy and completeness of the localization results.
[0166] This invention proposes a dual prior interactive distillation strategy (DPID), which optimizes the performance of the KSM and CSR modules through mutual distillation, thereby reducing incomplete localization and overlocalization. The DPID strategy enables these two modules to mutually enhance and optimize each other during training, improving the overall system performance.
[0167] This invention improves inference efficiency and reduces computational resource requirements by using optimized KSM and CSR modules during the inference process, thus avoiding the direct use of computationally expensive MLLM.
[0168] This invention also provides a storage medium for storing a computer program, which, when executed, performs at least the methods described above.
[0169] This invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein the processor executes the computer program by performing at least the method described above.
[0170] This invention also provides a processor that executes a computer program, at least performing the methods described above.
[0171] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc or CD-ROM; magnetic surface memory can be disk storage or magnetic tape storage. The storage media described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable types of memory.
[0172] In the several embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0173] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0174] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0175] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0176] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
[0177] The methods disclosed in the several method embodiments provided by this invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0178] The features disclosed in the several product embodiments provided by this invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0179] The features disclosed in the several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0180] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various equivalent substitutions or obvious modifications can be made without departing from the concept of the present invention, and all such modifications, achieving the same performance or application, should be considered within the scope of protection of the present invention.
Claims
1. A weakly supervised temporal action localization method based on a multimodal large language model, characterized in that, Includes the following steps: S1. Key semantic matching: Construct a key semantic matching module, and use the key semantic prior information generated by the multimodal large language model to match video segments and locate the action time interval; S2. Complete Semantic Reconstruction: Construct a complete semantic reconstruction module, which uses complete semantic prior information generated by a multimodal large language model to reconstruct action descriptions, thereby enhancing the comprehensive understanding of the temporal intervals of action instances; S3. Dual-prior interactive distillation: Construct a dual-prior interactive distillation strategy module to optimize the collaboration between the key semantic matching module and the complete semantic reconstruction module through mutual distillation, thereby reducing incomplete localization and over-localization. S4. Training and Inference: Jointly train the key semantic matching module and the complete semantic reconstruction module, and use the trained module to directly perform temporal action localization during inference, avoiding the use of the multimodal large language model.
2. The weakly supervised temporal action localization method as described in claim 1, characterized in that, In step S1, the key semantic matching specifically includes: Construct key semantic prompt templates and input them into a multimodal large language model to generate key semantic descriptions of action categories; Extract the spatiotemporal features of the video and embed them into a representation; Generate text embedding representations of key semantic queries; Calculate the similarity matrix between video segments and key semantics, and apply attention weights to suppress the response of background segments; The matching loss is calculated through a multi-instance learning mechanism to optimize the action localization results.
3. The weakly supervised temporal action localization method as described in claim 1, characterized in that, In step S2, the complete semantic reconstruction specifically includes: Construct a complete semantic prompt template and input it into a multimodal large language model to generate a complete semantic description that details the action; Random masking is applied to key action words in the complete semantic description; Extract video features and embed them into a representation, then apply an attention mechanism to assign weights to time segments; Use the Transformer model to predict masked action words to reconstruct the complete semantic description; The semantic reconstruction process is optimized by reconstructing the loss function.
4. The weakly supervised temporal action localization method as described in claim 1, characterized in that, In step S3, the dual-priority interactive distillation specifically includes: In the key semantic matching stage, optimization is performed by combining matching loss, mean square error of attention weights, and localization loss. In the complete semantic reconstruction stage, optimization is performed by combining reconstruction loss and mean squared error of attention weights; Through a mutual distillation mechanism, the output of the key semantic matching module is used as the optimization target of the complete semantic reconstruction module, and the output of the complete semantic reconstruction module is used as the optimization target of the key semantic matching module.
5. The weakly supervised temporal action localization method as described in claim 4, characterized in that, The localization loss includes a combination of the focus loss function, the DIOU loss function, and the multi-instance learning loss function to focus on key action areas and improve localization accuracy.
6. The weakly supervised temporal action localization method as described in any one of claims 1 to 5, characterized in that, In step S4, the training includes: An optimization algorithm is used to jointly train the key semantic matching module and the complete semantic reconstruction module; Minimize localization loss and reconstruction loss; Optimize the timing action positioning accuracy through a multi-view learning mechanism.
7. The weakly supervised temporal action localization method as described in any one of claims 1 to 6, characterized in that, In step S4, the reasoning includes: The input video is processed independently using the trained key semantic matching module and the complete semantic reconstruction module. Generate action time-series interval prediction results without calling a multimodal large language model throughout the process.
8. The weakly supervised temporal action localization method as described in any one of claims 1 to 7, characterized in that, It also includes video input and feature extraction steps: Extract multimodal features from the original video, including RGB features and optical flow features; Spatiotemporal features are fused to construct a video embedding representation, providing input features for the key semantic matching module and the complete semantic reconstruction module.
9. A weakly supervised temporal action localization system based on a multimodal large language model, characterized in that, include: The key semantic matching module is used to match key semantic prior information generated by a multimodal large language model with video segments to locate the action time interval; The complete semantic reconstruction module is used to reconstruct action descriptions using complete semantic prior information generated by a multimodal large language model, so as to enhance the comprehensive understanding of the instance order interval of the action. The dual-priority interactive distillation strategy module is used to optimize the collaboration between the key semantic matching module and the complete semantic reconstruction module through mutual distillation, thereby reducing incomplete localization and over-localization. The training and inference module is used to jointly train the key semantic matching module and the complete semantic reconstruction module, and to directly perform temporal action localization using the trained module during inference, avoiding the need to call the multimodal large language model.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the weakly supervised temporal action localization method based on a multimodal large language model as described in any one of claims 1 to 8.