Multi-stage zero-sample video action positioning method based on multi-modal large model framework
Through the multi-stage zero-sample video action localization method of the multimodal large model framework, the video action category and temporal position are automatically identified, which solves the problem of existing methods' dependence on labeled information, achieves high-accuracy and stable action localization, and enhances the model's adaptability in complex scenarios.
Patent Information
- Application Number
- CN202510706659.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing video action localization methods rely on a large amount of labeled information and are difficult to adapt to complex and changeable real-world scenes. The action boundary positioning is inaccurate and the model generalization ability is insufficient, making it difficult to accurately locate actions in different scenarios.
A multi-stage zero-shot video action localization method based on a multimodal large model framework is adopted. The multimodal large language model is used to obtain candidate video action categories and key action stages. Combined with frame-level confidence scoring and similarity calculation, the action category and temporal position are automatically identified. Through the semantic alignment of images and texts and the similarity calculation mechanism, the dependence on manually labeled data is eliminated.
It significantly improves the automation level of video motion analysis, improves the accuracy and stability of motion positioning, can accurately identify motion boundaries and categories in complex multi-action scenarios, and enhances the generalization ability of the model.
Smart Images

Figure CN120748033A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video understanding technology, and in particular to a multi-stage zero-sample video action localization method based on a multimodal large model framework. Background Art
[0002] With the rapid development of Internet technology, video data has exploded. As a key technology in the field of computer vision, video action localization plays an important role in many practical scenarios.
[0003] In the smart grid sector, video motion positioning can detect abnormal behavior in surveillance videos in real time, such as climbing towers and switching on power switches, and issue timely alerts, providing strong support for grid operation safety. By locating worker movements in operation videos, it can be determined whether workers are following standard procedures, thereby improving production efficiency and reducing the occurrence of safety accidents.
[0004] Traditional methods for video action localization are mainly divided into fully supervised and weakly supervised methods. However, these methods rely on a large amount of annotated information, which is not only labor-intensive, resource-intensive, and time-consuming, but also covers a limited number of actions, making them difficult to adapt to complex and changing real-world scenarios. Therefore, developing a zero-shot video action localization method that does not require extensive annotated information is of great practical significance.
[0005] The research on video action localization under zero-shot settings not only deepens the research on traditional video action localization, but also has stronger scalability and can cope with the ever-increasing scale of video data. Currently, this research mainly faces the following two challenges:
[0006] (1) Accurate determination of action boundaries. Accurately determining the start and end time points of an action in a video, that is, locating the action boundary, is another key challenge in zero-shot video action localization. Since there is no guidance from labeled information, the model needs to independently determine the start and end of the action by understanding and analyzing the video content. However, the action transitions in actual videos may be relatively smooth, without obvious boundaries, which makes it difficult to accurately delineate the action boundaries. In addition, some complex action sequences may contain multiple sub-actions. How to correctly identify these sub-actions and their boundaries is a problem that needs to be solved.
[0007] (2) Generalization ability of the model. The zero-shot video action localization model needs to have good generalization ability and be able to accurately locate actions in different scenarios and different data sets. However, in reality, video data comes from a wide range of sources and scenes are rich and diverse. Video data in different scenes may have large differences in feature distribution, action patterns, etc. If the model only learns the features of a specific dataset and fails to capture the general features of the action, it will be difficult to achieve good performance in new scenes and data. Therefore, how to improve the generalization ability of the model so that it can adapt to various complex practical application scenarios is one of the important challenges facing this field. Summary of the Invention
[0008] To solve the above technical problems, the present invention provides a multi-stage zero-sample video action localization method based on a multimodal large model framework.
[0009] In a first aspect, the present invention provides a multi-stage zero-sample video action localization method based on a multimodal large model framework, the technical solution of which is as follows:
[0010] Using a multimodal large language model, obtain at least one candidate video action category corresponding to the video to be tested, and determine multiple key action stages corresponding to each candidate video action category, arranged in chronological order;
[0011] For any candidate video action category, the multimodal large language model is used to obtain the confidence of each key action stage corresponding to each candidate video action category for each video frame of the video to be tested, and the key action stage corresponding to the highest confidence of each video frame is determined as the target action stage of each video frame. According to the video frames with the highest confidence greater than the confidence threshold, at least one candidate time segment is constructed, and the target action stage with the largest number of occurrences in each candidate time segment is determined as the final action stage of each candidate time segment. According to the occurrence time sequence of the final action stage of each candidate time segment and the interval length between adjacent candidate time segments, all candidate time segments are merged to obtain the video action segment positioning result corresponding to any candidate video action category, until the video action segment positioning result corresponding to each candidate video action category is obtained.
[0012] The beneficial effects of the multi-stage zero-sample video action localization method based on a multimodal large model framework of the present invention are as follows:
[0013] The method of the present invention effectively realizes the discrimination of action categories and the annotation of temporal positions in videos by introducing a multimodal large model, utilizing image-text semantic alignment and similarity calculation mechanisms, and combining frame-level confidence scoring. It completely gets rid of the dependence on manually labeled data, significantly improves the automation level of video action analysis, and enhances the accuracy of action positioning and stability in complex multi-action scenarios.
[0014] Based on the above solution, the multi-stage zero-sample video action localization method based on a multimodal large model framework of the present invention can also be improved as follows.
[0015] In an optional manner, the step of merging all candidate time segments according to the occurrence time sequence of the final action phase of each candidate time segment and the interval duration between adjacent candidate time segments to obtain a video action segment positioning result corresponding to any candidate video action category includes:
[0016] For two adjacent candidate time segments, if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is before the occurrence time sequence of the final action stage of the next adjacent candidate time segment, the two adjacent candidate time segments are merged; if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is the same as the occurrence time sequence of the final action stage of the next adjacent candidate time segment, and the time interval between the two adjacent candidate time segments is less than the interval duration threshold, the two adjacent candidate time segments are merged;
[0017] Repeat the steps for two adjacent candidate time segments until all candidate time segments are merged to obtain a video action segment positioning result corresponding to any candidate video action category.
[0018] In an optional manner, the step of obtaining at least one candidate video action category corresponding to the video to be tested using a multimodal large language model includes:
[0019] Based on a plurality of preset action categories and in combination with action category prompt words, the at least one candidate video action category corresponding to the video to be tested is obtained using the multimodal large language model.
[0020] In an optional manner, the step of determining a plurality of key action stages corresponding to each candidate video action category and arranged in chronological order of occurrence includes:
[0021] Based on the video to be tested and the at least one candidate video action category, and in combination with text prompt words, using the multimodal large language model, generate text description information corresponding to each candidate video action category;
[0022] The multimodal large language model is used to divide the text description information corresponding to each candidate video action category into a plurality of key action stages arranged in chronological order.
[0023] In a second aspect, the present invention provides a multi-stage zero-sample video action localization system based on a multimodal large model framework, the technical solution of which is as follows:
[0024] Includes: processing module and operation module;
[0025] The processing module is used to: use a multimodal large language model to obtain at least one candidate video action category corresponding to the video to be tested, and determine a plurality of key action stages corresponding to each candidate video action category and arranged in chronological order;
[0026] The operation module is used to: for any candidate video action category, use the multimodal large language model to obtain the confidence of each key action stage corresponding to each candidate video action category for each video frame of the video to be tested, determine the key action stage corresponding to the highest confidence of each video frame as the target action stage of each video frame, construct at least one candidate time segment based on the video frame with the highest confidence greater than the confidence threshold, determine the target action stage with the largest number of occurrences in each candidate time segment as the final action stage of each candidate time segment, and merge all candidate time segments according to the occurrence time sequence of the final action stage of each candidate time segment and the interval length between adjacent candidate time segments to obtain the video action segment positioning result corresponding to any candidate video action category, until the video action segment positioning result corresponding to each candidate video action category is obtained.
[0027] The beneficial effects of the multi-stage zero-sample video action localization system based on the multimodal large model framework of the present invention are as follows:
[0028] The system of the present invention effectively realizes the discrimination of action categories and the annotation of temporal positions in videos by introducing a multimodal large model, utilizing image-text semantic alignment and similarity calculation mechanisms, and combining frame-level confidence scoring. It completely gets rid of the dependence on manually labeled data, significantly improves the automation level of video action analysis, and enhances the accuracy of action positioning and stability in complex multi-action scenarios.
[0029] Based on the above solution, the multi-stage zero-sample video action localization system based on a multimodal large model framework of the present invention can also be improved as follows.
[0030] In an optional manner, the operation module is specifically configured to:
[0031] For two adjacent candidate time segments, if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is before the occurrence time sequence of the final action stage of the next adjacent candidate time segment, the two adjacent candidate time segments are merged; if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is the same as the occurrence time sequence of the final action stage of the next adjacent candidate time segment, and the time interval between the two adjacent candidate time segments is less than the interval duration threshold, the two adjacent candidate time segments are merged;
[0032] Repeat the steps for two adjacent candidate time segments until all candidate time segments are merged to obtain a video action segment positioning result corresponding to any candidate video action category.
[0033] In an optional manner, the processing module is specifically configured to:
[0034] Based on a plurality of preset action categories and in combination with action category prompt words, the at least one candidate video action category corresponding to the video to be tested is obtained using the multimodal large language model.
[0035] In an optional manner, the processing module is specifically configured to:
[0036] Based on the video to be tested and the at least one candidate video action category, and in combination with text prompt words, using the multimodal large language model, generate text description information corresponding to each candidate video action category;
[0037] The multimodal large language model is used to divide the text description information corresponding to each candidate video action category into a plurality of key action stages arranged in chronological order.
[0038] In a third aspect, the technical solution of an electronic device of the present invention is as follows:
[0039] The invention comprises a memory, a processor and a program stored in the memory and running on the processor. When the processor executes the program, the steps of the multi-stage zero-sample video action localization method based on a multimodal large model framework of the present invention are implemented.
[0040] In a fourth aspect, the present invention provides a computer-readable storage medium having the following technical solution:
[0041] Instructions are stored in the computer-readable storage medium. When the computer-readable storage medium reads the instructions, the computer-readable storage medium executes the steps of the multi-stage zero-sample video action localization method based on the multimodal large model framework of the present invention.
[0042] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings are only used to illustrate the embodiments and are not to be considered as limiting the present invention. In addition, the same reference symbols are used to represent the same components throughout the drawings. In the drawings:
[0044] Figure 1 1 is a flow chart of an embodiment of a multi-stage zero-sample video action localization method based on a multimodal large model framework of the present invention;
[0045] Figure 2 Schematic diagram of the overall principle of video action positioning;
[0046] Figure 3 A schematic diagram of the comparison results on the Activitynet v1.2 dataset;
[0047] Figure 4 Schematic diagram of the comparison results on the Thumos14 dataset;
[0048] Figure 5 Schematic diagram of the structure of an embodiment of a multi-stage zero-sample video action localization system based on a multimodal large model framework of the present invention;
[0049] Figure 6 The figure is a schematic structural diagram of an embodiment of an electronic device of the present invention. DETAILED DESCRIPTION
[0050] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein.
[0051] Figure 1A flow chart of an embodiment of a multi-stage zero-sample video action localization method based on a multimodal large model framework provided by the present invention is shown. The multi-stage zero-sample video action localization method based on a multimodal large model framework can be executed by an electronic device such as a terminal device or a server. Among them, the terminal device can be any fixed or mobile terminal such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be a single server or a server cluster composed of multiple servers. Any electronic device can implement a multi-stage zero-sample video action localization method based on a multimodal large model framework by calling computer-readable instructions stored in a memory through a processor. As Figure 1 As shown, the following steps are included:
[0052] S1. Using a multimodal large language model, obtain at least one candidate video action category corresponding to the video to be tested, and determine multiple key action stages corresponding to each candidate video action category, which are arranged in chronological order.
[0053] The video to be tested is the video for which video action positioning is required in this embodiment. Candidate video action categories refer to action categories obtained by encoding the content of the video to be tested using a multimodal large language model, including but not limited to: standing, running, walking, and waving.
[0054] In S1, the step of using a multimodal large language model to obtain at least one candidate video action category corresponding to the video to be tested includes:
[0055] Based on multiple preset action categories and combined with action category prompt words, a multimodal large language model is used to obtain at least one candidate video action category corresponding to the video to be tested.
[0056] Where, given an action category list C = {c1, c2, ..., c N}, N represents the total number of preset action categories. The action category list contains N preset action categories. The default action category prompt is "Based on <action category list C>, select the corresponding action from the video to be tested". The candidate video action category set is: P MLLM (c i |V) represents the given video to be tested V and the i-th preset action category c i , the probability value output by the multimodal large language model. τ c is the threshold parameter used to determine c i Can you enter?
[0057] In S1, the step of determining a plurality of key action stages corresponding to each candidate video action category and arranged in chronological order includes:
[0058] Based on the video to be tested and at least one candidate video action category, combined with text prompt words, a multimodal large language model is used to generate text description information corresponding to each candidate video action category.
[0059] The default text prompt is "Describe the candidate video action category in the video to be tested" Specifically, for the video to be tested V and any candidate video action category in the candidate video action category set Generate a text description of each candidate video action category of the video to be tested through a multimodal large language model desc is a text prompt word that guides MLLM to generate natural language.
[0060] Using a multimodal large language model, the text description information corresponding to each candidate video action category is divided into multiple key action stages arranged in chronological order.
[0061] Specifically, the candidate video action categories Text description of Divided into multiple action semantic stages, the process is expressed as: Represents the i-th preset action category c i Multiple action semantic stages, s x Indicates the xth action semantic stage, key x ∈{0,1} indicates whether the xth action semantic stage can be judged as an action The identifier of an action semantic phase; assign serial numbers to multiple action semantic phases in the order of occurrence time and generate key action phases based on the identifiers. The key action phase set is expressed as: s k Indicates the kth key action stage, key_num k Indicates s k The assigned sequence number (a number generated based on the chronological order of occurrence, starting from 1 and increasing).
[0062] S2. For any candidate video action category, use the multimodal large language model to obtain the confidence of each key action stage corresponding to each candidate video action category for each video frame of the video to be tested, and determine the key action stage corresponding to the highest confidence of each video frame as the target action stage of each video frame. According to the video frames with the highest confidence greater than the confidence threshold, construct at least one candidate time segment, determine the target action stage with the largest number of occurrences in each candidate time segment as the final action stage of each candidate time segment, and merge all candidate time segments according to the occurrence time sequence of the final action stage of each candidate time segment and the interval duration between adjacent candidate time segments to obtain the video action segment positioning result corresponding to any candidate video action category, until the video action segment positioning result corresponding to each candidate video action category is obtained.
[0063] In S2, specifically:
[0064] S21. For a video to be tested V, if the total duration of the video to be tested is T and the total number of frames of the video to be tested is J, then the frame sequence of the video to be tested is: in, t j represents the jth video frame F j The corresponding real video time. For any candidate video action category Using a multimodal large language model, obtain the video frame F j respectively Corresponding to each key action stage s k The confidence level of is expressed as: Represents the video frame F j Corresponding to the confidence (probability value) at the kth key action stage; finally, each key action stage s is obtained k Frame-by-frame confidence distribution of
[0065] S22. For the frame-by-frame confidence distribution corresponding to all key action stages, the key action stage corresponding to the highest confidence of each video frame is determined as the target action stage of each video frame, which is specifically expressed as: represents the highest confidence of the j-th video frame, Indicates the number of the key action stage corresponding to the highest confidence of the j-th video frame.
[0066] S23, for each video frame with the highest confidence, select the video frame with the highest confidence greater than the confidence threshold, and construct at least one candidate time segment S raw , each candidate time segment S raw Expressed as: F s represents the starting frame of the candidate time segment, F e represents the end frame of the candidate time segment, and θ represents the confidence threshold, which is used to determine whether the video frame is a high-confidence frame.
[0067] S24. For each candidate time segment, the number that appears most frequently among the key action stage numbers corresponding to all video frames in each candidate time segment is used as the final action stage number of the corresponding candidate time segment. This process is expressed as: in represents the candidate time segment [F s ,F e ] corresponds to the final action stage number, I(.) is the indicator function, when The function takes the value 1 when , and zero otherwise. Indicates that in the candidate time segment [F s ,F e ], the frequency of occurrence of the key action stage number m. Finally, the candidate time segment set S raw Expressed as: L represents the total number of candidate time segments.
[0068] S25. For two adjacent candidate time segments, if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is before the occurrence time sequence of the final action stage of the next adjacent candidate time segment, the two adjacent candidate time segments are merged; if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is the same as the occurrence time sequence of the final action stage of the next adjacent candidate time segment, and the time interval between the two adjacent candidate time segments is less than the interval duration threshold, the two adjacent candidate time segments are merged.
[0069] Among them, for two adjacent candidate time segments and If the key action stage numbers corresponding to these two adjacent candidate time segments satisfy The candidate time segment Considered as candidate time segments About Action Categories Therefore, the two adjacent candidate time segments are merged to obtain a new candidate time segment. The process is expressed as:
[0070] Among them, for two adjacent candidate time segments and If the two adjacent candidate time segments belong to the same action stage and the time interval between them is small, then and It is considered that these two adjacent candidate time segments should be merged. The process is expressed as: Seg y Represents the y-th video action segment after merging.
[0071] It should be noted that tIoU is used to calculate the proportion of two adjacent candidate time segments in the time period, which is expressed as θ IoU The interval duration threshold controls whether two adjacent candidate time segments are merged.
[0072] S26 , repeatedly executing S25 until all candidate time segments are merged, thereby obtaining a video action segment positioning result corresponding to any candidate video action category.
[0073] The video action segment location results include action interval predictions and corresponding video action categories. The video under test contains at least one action interval prediction value, each corresponding to a video action category. For example, the video action category corresponding to the 1s-3s segment of the video under test is "running," while the video action category corresponding to the 5s-10s segment is "kicking a ball."
[0074] S27 . Repeat S21 - S26 for each candidate video action category to obtain a video action segment positioning result corresponding to each candidate video action category.
[0075] It should be noted that Figure 2 The overall principle diagram of this embodiment is shown. The performance of the video action positioning method in this embodiment is compared with the positioning accuracy of the international leading similar models. Figure 3 Shows the comparison results on the Activitynet v1.2 dataset. Figure 4 The comparison results on the Thumos14 dataset are shown. Through the comparison, it can be seen that the positioning accuracy of the video action positioning method in this embodiment is significantly superior.
[0076] The technical solution of this embodiment effectively realizes the discrimination of action categories and the annotation of temporal positions in videos by introducing a multimodal large model, utilizing image-text semantic alignment and similarity calculation mechanisms, and combining frame-level confidence scoring. It completely gets rid of the dependence on manually labeled data, significantly improves the automation level of video action analysis, and improves the accuracy of action positioning and stability in complex multi-action scenarios.
[0077] Figure 5 FIG. 1 shows a schematic diagram of a multi-stage zero-sample video action localization system 200 based on a multimodal large model framework provided by the present invention. Figure 5As shown, the system 200 includes: a processing module 210 and an operation module 220;
[0078] The processing module 210 is configured to: utilize a multimodal large language model to obtain at least one candidate video action category corresponding to the video to be tested, and determine a plurality of key action stages corresponding to each candidate video action category, which are arranged in chronological order;
[0079] The operation module 220 is used to: for any candidate video action category, use the multimodal large language model to obtain the confidence of each key action stage corresponding to each candidate video action category for each video frame of the video to be tested, determine the key action stage corresponding to the highest confidence of each video frame as the target action stage of each video frame, construct at least one candidate time segment based on the video frame with the highest confidence greater than the confidence threshold, determine the target action stage with the largest number of occurrences in each candidate time segment as the final action stage of each candidate time segment, and merge all candidate time segments according to the occurrence time sequence of the final action stage of each candidate time segment and the interval length between adjacent candidate time segments to obtain the video action segment positioning result corresponding to any candidate video action category, until the video action segment positioning result corresponding to each candidate video action category is obtained.
[0080] In an optional manner, the operation module 220 is specifically configured to:
[0081] For two adjacent candidate time segments, if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is before the occurrence time sequence of the final action stage of the next adjacent candidate time segment, the two adjacent candidate time segments are merged; if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is the same as the occurrence time sequence of the final action stage of the next adjacent candidate time segment, and the time interval between the two adjacent candidate time segments is less than the interval duration threshold, the two adjacent candidate time segments are merged;
[0082] Repeat the steps for two adjacent candidate time segments until all candidate time segments are merged to obtain a video action segment positioning result corresponding to any candidate video action category.
[0083] In an optional manner, the processing module 210 is specifically configured to:
[0084] Based on a plurality of preset action categories and in combination with action category prompt words, the at least one candidate video action category corresponding to the video to be tested is obtained using the multimodal large language model.
[0085] In an optional manner, the processing module 210 is specifically configured to:
[0086] Based on the video to be tested and the at least one candidate video action category, and in combination with text prompt words, using the multimodal large language model, generate text description information corresponding to each candidate video action category;
[0087] The multimodal large language model is used to divide the text description information corresponding to each candidate video action category into a plurality of key action stages arranged in chronological order.
[0088] It should be noted that the beneficial effects of the multi-stage zero-sample video action localization system 200 based on the multimodal large model framework provided in the above embodiment are the same as the beneficial effects of the multi-stage zero-sample video action localization method based on the multimodal large model framework, and will not be repeated here. In addition, when the system provided in the above embodiment realizes its functions, it only uses the division of the above-mentioned functional modules as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to actual conditions to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0089] Among them, the multi-stage zero-sample video action localization system 200 based on the multimodal large model framework of the present invention can be a computer program (including program code) running in a computer device. For example, the multi-stage zero-sample video action localization system based on the multimodal large model framework of the present invention is an application software that can be used to execute the corresponding steps in the multi-stage zero-sample video action localization method based on the multimodal large model framework of the present invention.
[0090] In some embodiments, the multi-stage zero-sample video motion localization system based on the multimodal large model framework of the present invention can be implemented in a combination of software and hardware. As an example, the multi-stage zero-sample video motion localization system based on the multimodal large model framework of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the multi-stage zero-sample video motion localization method based on the multimodal large model framework of the present invention. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0091] The modules described in the embodiments of the present invention may be implemented in software or hardware, and the name of a module does not necessarily limit the module itself.
[0092] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, any of the above-mentioned multi-stage zero-sample video action localization methods based on a multimodal large model framework is implemented. That is, an electronic device according to an embodiment of the present invention may include but is not limited to: a processor and a memory; a memory for storing a computer program; a processor for executing the multi-stage zero-sample video action localization method based on a multimodal large model framework shown in any embodiment of the present invention by calling the computer program.
[0093] In an alternative embodiment, an electronic device is provided, such as Figure 6 As shown, Figure 6 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.
[0094] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0095] Bus 4002 may include a path for transmitting information between the above components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 In the figure, only one thick line is used to represent the bus 4002, but this does not mean that there is only one bus or one type of bus.
[0096] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0097] The memory 4003 is used to store application code (computer program) for executing the solution of the present invention, and is controlled by the processor 4001. The processor 4001 is used to execute the application code stored in the memory 4003 to implement the content shown in the above method embodiment.
[0098] Among them, the electronic device can also be a terminal device, and the terminal device can be any terminal device that can install applications and access web pages through applications, including at least one of a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, smart TV, and smart car-mounted device.
[0099] It should be noted that Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.
[0100] A computer-readable storage medium according to an embodiment of the present invention stores a computer program, which, when executed by a processor, implements any of the above-mentioned multi-stage zero-sample video action localization methods based on a multimodal large model framework.
[0101] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, or the like.
[0102] In an exemplary embodiment, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the multi-stage zero-shot video action localization method based on the multimodal large model framework.
[0103] Computer program code for performing the operations of the present invention may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0104] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0105] The computer-readable storage medium provided in the embodiments of the present invention may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or component.
[0106] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0107] The above description is merely a preferred embodiment of the present invention and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present invention.
[0108] It should be noted that the terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects and to define a specific order or precedence. Where appropriate, the order used for similar objects may be interchanged, such that the embodiments of the present application described herein can be implemented in an order other than the order shown or described.
[0109] Those skilled in the art will appreciate that the present invention may be implemented as a system, method, or computer program product. Therefore, the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the present invention may be implemented in the form of a computer program product embodied in one or more computer-readable media containing computer-readable program code.
[0110] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A multi-stage zero-shot video action localization method based on a multimodal large model framework, characterized by: include: Using a multimodal large language model, obtain at least one candidate video action category corresponding to the video to be tested, and determine multiple key action stages corresponding to each candidate video action category, arranged in chronological order; For any candidate video action category, the multimodal large language model is used to obtain the confidence of each key action stage corresponding to each candidate video action category for each video frame of the video to be tested, and the key action stage corresponding to the highest confidence of each video frame is determined as the target action stage of each video frame. According to the video frames with the highest confidence greater than the confidence threshold, at least one candidate time segment is constructed, and the target action stage with the largest number of occurrences in each candidate time segment is determined as the final action stage of each candidate time segment. According to the occurrence time sequence of the final action stage of each candidate time segment and the interval length between adjacent candidate time segments, all candidate time segments are merged to obtain the video action segment positioning result corresponding to any candidate video action category, until the video action segment positioning result corresponding to each candidate video action category is obtained.
2. The multi-stage zero-shot video action localization method based on a multimodal large model framework according to claim 1 is characterized in that: The step of merging all candidate time segments according to the occurrence time sequence of the final action phase of each candidate time segment and the interval duration between adjacent candidate time segments to obtain a video action segment positioning result corresponding to any candidate video action category includes: For two adjacent candidate time segments, if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is before the occurrence time sequence of the final action stage of the next adjacent candidate time segment, the two adjacent candidate time segments are merged; if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is the same as the occurrence time sequence of the final action stage of the next adjacent candidate time segment, and the time interval between the two adjacent candidate time segments is less than the interval duration threshold, the two adjacent candidate time segments are merged; Repeat the steps for two adjacent candidate time segments until all candidate time segments are merged to obtain a video action segment positioning result corresponding to any candidate video action category.
3. The multi-stage zero-shot video action localization method based on a multimodal large model framework according to claim 1 is characterized in that: The step of obtaining at least one candidate video action category corresponding to the video to be tested using the multimodal large language model includes: Based on a plurality of preset action categories and in combination with action category prompt words, the at least one candidate video action category corresponding to the video to be tested is obtained using the multimodal large language model.
4. The multi-stage zero-shot video action localization method based on a multimodal large model framework according to claim 1 is characterized in that: The steps of determining multiple key action stages corresponding to each candidate video action category and arranged in chronological order include: Based on the video to be tested and the at least one candidate video action category, and in combination with text prompt words, using the multimodal large language model, generate text description information corresponding to each candidate video action category; The multimodal large language model is used to divide the text description information corresponding to each candidate video action category into a plurality of key action stages arranged in chronological order.
5. A multi-stage zero-shot video action localization system based on a multimodal large model framework, characterized by: include: Processing module and operation module; The processing module is used to: use a multimodal large language model to obtain at least one candidate video action category corresponding to the video to be tested, and determine a plurality of key action stages corresponding to each candidate video action category and arranged in chronological order; The operation module is used to: for any candidate video action category, use the multimodal large language model to obtain the confidence of each key action stage corresponding to each candidate video action category for each video frame of the video to be tested, determine the key action stage corresponding to the highest confidence of each video frame as the target action stage of each video frame, construct at least one candidate time segment based on the video frame with the highest confidence greater than the confidence threshold, determine the target action stage with the largest number of occurrences in each candidate time segment as the final action stage of each candidate time segment, and merge all candidate time segments according to the occurrence time sequence of the final action stage of each candidate time segment and the interval length between adjacent candidate time segments to obtain the video action segment positioning result corresponding to any candidate video action category, until the video action segment positioning result corresponding to each candidate video action category is obtained.
6. The multi-stage zero-shot video action localization system based on a multimodal large model framework according to claim 5 is characterized in that: The operation module is specifically used for: For two adjacent candidate time segments, if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is before the occurrence time sequence of the final action stage of the next adjacent candidate time segment, the two adjacent candidate time segments are merged; if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is the same as the occurrence time sequence of the final action stage of the next adjacent candidate time segment, and the time interval between the two adjacent candidate time segments is less than the interval duration threshold, the two adjacent candidate time segments are merged; Repeat the steps for two adjacent candidate time segments until all candidate time segments are merged to obtain a video action segment positioning result corresponding to any candidate video action category.
7. The multi-stage zero-shot video action localization system based on a multimodal large model framework according to claim 5, characterized in that: The processing module is specifically used for: Based on a plurality of preset action categories and in combination with action category prompt words, the at least one candidate video action category corresponding to the video to be tested is obtained using the multimodal large language model.
8. The multi-stage zero-shot video action localization system based on a multimodal large model framework according to claim 5, characterized in that: The processing module is specifically used for: Based on the video to be tested and the at least one candidate video action category, and in combination with text prompt words, using the multimodal large language model, generate text description information corresponding to each candidate video action category; The multimodal large language model is used to divide the text description information corresponding to each candidate video action category into a plurality of key action stages arranged in chronological order.
9. An electronic device, characterized in that: The electronic device includes a processor, which is coupled to a memory, and the memory stores at least one computer program, which is loaded and executed by the processor so that the electronic device implements the multi-stage zero-sample video action localization method based on a multimodal large model framework as described in any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor so that the computer-readable storage medium implements the multi-stage zero-sample video action localization method based on a multimodal large model framework as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Action recognition method and system based on multi-modal sequence fusion
CN115937975A
Cross-modal power video positioning method and system, electronic equipment and storage medium
CN119888563A
Unsupervised power video action positioning method, system and equipment and storage medium
CN119888566A
Multimodal heterogeneous feature fusion-based compact video event description method
WO2023050295A1
Cited By
Video stream processing method and system based on character detection, program product and medium
CN121309755A