Multi-stage zero-shot video action localization method based on multi-modal large model framework

By employing a multi-stage zero-shot video action localization method based on a multimodal large model framework, and utilizing a multimodal large language model to identify video action categories and key stages, this method solves the problems of action boundary determination and model generalization, achieving high accuracy and stable action localization.

CN120748033BActive Publication Date: 2026-02-13LEQING POWER IND CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510706659.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2026-02-13
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing video motion localization methods rely on a large amount of labeled information, making it difficult to adapt to complex and ever-changing real-world scenarios. They also face difficulties in defining motion boundaries, have insufficient model generalization ability, and struggle to accurately locate motions in different scenarios.

Method used

A multi-stage zero-shot video action localization method based on a multimodal large model framework is adopted. The method uses a multimodal large language model to obtain candidate video action categories and key action stages. Combined with image-text semantic alignment and similarity calculation, action segments are merged through frame-level confidence scoring to achieve action category discrimination and temporal position labeling.

Benefits of technology

Completely eliminates the reliance on manually labeled data, significantly improves the automation level of video motion analysis, enhances the accuracy and stability of motion localization, and adapts to complex scenarios with multiple actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748033B_ABST
    Figure CN120748033B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of video understanding, and specifically discloses a multi-stage zero-sample video action positioning method based on a multi-modal large model framework, which comprises the following steps: a multi-modal large language model is used to obtain candidate video action categories of a to-be-detected video and a plurality of corresponding key action stages; for any candidate video action category, the confidence of a video frame of the to-be-detected video in each key action stage is obtained; according to the video frame with the highest confidence greater than a threshold value, a candidate time segment is constructed and merged to obtain a positioning result corresponding to the candidate video action category, until the positioning result corresponding to each candidate video action category is obtained. By introducing the multi-modal large model, using a picture-text semantic alignment and a similarity calculation mechanism, and combining frame-level confidence scoring, the application realizes the discrimination of action categories in a video and the labeling of time sequence positions, breaks away from the dependence on artificial labeling data, and improves the action positioning accuracy and stability in a multi-action complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video understanding, and particularly relates to a multi-stage zero-shot video action localization method based on a multi-modal large model framework. BACKGROUND

[0002] With the rapid development of Internet technology, video data is growing explosively. Video action localization, as a key technology in the field of computer vision, plays an important role in many practical scenarios.

[0003] In the field of smart grid, video action localization can monitor abnormal behaviors in surveillance videos in real time, such as climbing towers, pulling circuit breakers, etc., and timely issue alarms to provide strong support for the safety of power grid operations. Through action localization of worker operation videos, it can be determined whether the workers are operating according to the standard process, thereby improving production efficiency and reducing the occurrence of safety accidents.

[0004] Traditional video action localization methods are mainly divided into fully supervised and weakly supervised methods. However, these methods all rely on a large amount of labeled information, and the labeling work not only consumes a lot of manpower, material resources and time, but also the number of actions covered by the labeled information is limited, which is difficult to adapt to complex and variable real scenarios. Therefore, it is of great practical significance to develop a zero-shot video action localization method without a large amount of labeled information.

[0005] Through the research of video action localization under zero-shot setting, not only the research of traditional video action localization is deepened, but also it has stronger scalability and can cope with the growing scale of video data. At present, this research mainly faces the following two challenges:

[0006] (1) Accurate determination of action boundaries. Accurately determining the starting and ending time points of actions in videos, i.e. the localization of action boundaries, is another key challenge of zero-shot video action localization. Without the guidance of labeled information, the model needs to understand and analyze the content of the video to autonomously determine the start and end of the action. However, the transition of actions in actual videos may be relatively smooth without obvious boundaries, which makes it difficult to accurately divide the action boundaries. In addition, some complex action sequences may contain multiple sub-actions, and how to correctly identify these sub-actions and their boundaries is a problem to be solved.

[0007] (2) Generalization ability of the model. Zero-shot video action localization models need to have good generalization ability and accurately localize actions in different scenes and different datasets. However, video data in reality comes from a wide range of sources and has a rich variety of scenes. Video data in different scenes may have large differences in feature distribution and action patterns. If the model only learns the features of a specific dataset and cannot capture the general features of actions, it will be difficult to achieve good performance on new scenes and data. Therefore, how to improve the generalization ability of the model to adapt to various complex practical application scenarios is one of the important challenges in this field. SUMMARY

[0008] To solve the above technical problems, the present application provides a multi-stage zero-shot video action localization method based on a multi-modal large model framework.

[0009] In the first aspect, the present application provides a multi-stage zero-shot video action localization method based on a multi-modal large model framework. The technical scheme of the method is as follows:

[0010] Using a multi-modal large language model, at least one candidate video action category corresponding to the to-be-tested video is obtained, and a plurality of key action stages corresponding to each candidate video action category in chronological order are determined;

[0011] For any candidate video action category, using the multi-modal large language model, the confidence of each video frame of the to-be-tested video in each key action stage corresponding to the any candidate video action category is obtained, the key action stage corresponding to the highest confidence of each video frame is determined as the target action stage of each video frame, at least one candidate time segment is constructed according to the video frame with a highest confidence greater than a confidence threshold, the final action stage of each candidate time segment is determined as the target action stage with the highest occurrence frequency in each candidate time segment, and all candidate time segments are merged according to the chronological order of the final action stage of each candidate time segment and the interval duration between adjacent candidate time segments, to obtain the video action segment localization result corresponding to the any candidate video action category, until the video action segment localization result corresponding to each candidate video action category is obtained.

[0012] The multi-stage zero-shot video action localization method based on a multi-modal large model framework of the present application has the following beneficial effects:

[0013] The method of the application effectively realizes the discrimination of action categories and the labeling of time sequence positions in the video by introducing a multimodal large model, using a text and image semantic alignment and similarity calculation mechanism, and combining frame-level confidence scores, completely gets rid of the dependence on artificial labeling data, significantly improves the automation level of video action analysis, and improves the action positioning accuracy and stability in a multi-action complex scene.

[0014] On the basis of the above scheme, the multi-stage zero-shot video action positioning method based on the multimodal large model framework of the application can be further improved as follows.

[0015] In an optional manner, the step of merging all candidate time segments according to the time sequence of occurrence of the final action stage of each candidate time segment and the interval duration between adjacent candidate time segments to obtain the video action segment positioning result corresponding to any candidate video action category comprises:

[0016] For two adjacent candidate time segments, if the time sequence of occurrence of the final action stage of the former adjacent candidate time segment is located before the time sequence of occurrence of the final action stage of the latter adjacent candidate time segment, the two adjacent candidate time segments are merged; if the time sequence of occurrence of the final action stage of the former adjacent candidate time segment is the same as the time sequence of occurrence of the final action stage of the latter adjacent candidate time segment, and the time interval between the two adjacent candidate time segments is less than the interval duration threshold, the two adjacent candidate time segments are merged.

[0017] The step of performing the operation on two adjacent candidate time segments is repeated until the merging of all candidate time segments is completed, and the video action segment positioning result corresponding to any candidate video action category is obtained.

[0018] In an optional manner, the step of obtaining at least one candidate video action category corresponding to the to-be-tested video by using a multimodal large language model comprises:

[0019] Based on a plurality of preset action categories and in combination with an action category prompt word, the multimodal large language model is used to obtain the at least one candidate video action category corresponding to the to-be-tested video.

[0020] In an optional manner, the step of determining a plurality of key action stages corresponding to each candidate video action category and arranged in a time sequence order comprises:

[0021] Based on the to-be-tested video and the at least one candidate video action category, and in combination with a text prompt word, the multimodal large language model is used to generate text description information corresponding to each candidate video action category.

[0022] The text description information corresponding to each candidate video action category is divided into a plurality of key action stages arranged in chronological order by using the multi-modal large language model.

[0023] In a second aspect, the present application provides a multi-stage zero-shot video action localization system based on a multi-modal large model framework, and the technical scheme of the system is as follows:

[0024] The system comprises a processing module and a running module.

[0025] The processing module is configured to obtain at least one candidate video action category corresponding to a to-be-tested video by using a multi-modal large language model, and determine a plurality of key action stages corresponding to each candidate video action category in chronological order.

[0026] The running module is configured to, for any candidate video action category, obtain, by using the multi-modal large language model, a confidence degree of each video frame of the to-be-tested video in each key action stage corresponding to the any candidate video action category, determine a target action stage corresponding to the highest confidence degree of each video frame as the target action stage of each video frame, construct at least one candidate time segment according to the video frame with a confidence degree higher than a confidence threshold, determine a final action stage of each candidate time segment as the action stage that appears most frequently in the candidate time segment, and merge all candidate time segments according to the chronological order of the final action stage of each candidate time segment and the interval duration between adjacent candidate time segments to obtain a video action segment localization result corresponding to the any candidate video action category, until a video action segment localization result corresponding to each candidate video action category is obtained.

[0027] The multi-stage zero-shot video action localization system based on the multi-modal large model framework has the following beneficial effects.

[0028] The system introduces a multi-modal large model, uses a text-image semantic alignment and similarity calculation mechanism, and combines frame-level confidence scoring to effectively realize the discrimination of action categories and the labeling of time sequence positions in the video, completely eliminates the dependence on artificial labeled data, significantly improves the automation level of video action analysis, and improves the action localization accuracy and stability in a multi-action complex scene.

[0029] On the basis of the above scheme, the multi-stage zero-shot video action localization system based on the multi-modal large model framework can be further improved as follows.

[0030] In an optional manner, the running module is specifically configured to:

[0031] For two adjacent candidate time segments, if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is located before the occurrence time sequence of the final action stage of the subsequent adjacent candidate time segment, the two adjacent candidate time segments are merged; if the occurrence time sequence of the final action stage of the previous adjacent candidate time segment is the same as the occurrence time sequence of the final action stage of the subsequent adjacent candidate time segment, and the time interval between the two adjacent candidate time segments is less than the interval duration threshold, the two adjacent candidate time segments are merged.

[0032] The step of performing for two adjacent candidate time segments is repeatedly executed until the merging of all candidate time segments is completed, and the video action segment positioning result corresponding to any candidate video action category is obtained.

[0033] In an optional manner, the processing module is specifically configured to:

[0034] Based on a plurality of preset action categories and in combination with an action category prompt word, the multi-modal large language model is used to obtain the at least one candidate video action category corresponding to the to-be-tested video.

[0035] In an optional manner, the processing module is specifically configured to:

[0036] Based on the to-be-tested video and the at least one candidate video action category, and in combination with a text prompt word, the multi-modal large language model is used to generate text description information corresponding to each candidate video action category;

[0037] The multi-modal large language model is used to divide the text description information corresponding to each candidate video action category into a plurality of key action stages arranged in an occurrence time sequence.

[0038] In a third aspect, a technical solution of an electronic device of the present application is as follows:

[0039] The processor executes the program to implement the steps of the multi-stage zero-shot video action positioning method based on the multi-modal large model framework of the present application.

[0040] In a fourth aspect, a technical solution of a computer readable storage medium provided by the present application is as follows:

[0041] The computer readable storage medium stores instructions, and when the computer readable storage medium reads the instructions, the computer readable storage medium executes the steps of the multi-stage zero-shot video action positioning method based on the multi-modal large model framework of the present application.

[0042] The above description is only a summary of the technical solutions of the present application. In order to enable a clearer understanding of the technical means of the present application, the content of the specification can be implemented, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0043] The accompanying drawings are included to provide a further understanding of the present application, and are incorporated herein and constitute a part of the detailed description. It should be noted that the accompanying drawings illustrate only exemplary embodiments of the present application and therefore should not be considered to limit the present application. In the drawings:

[0044] Figure 1 A flowchart of an embodiment of a multi-stage zero-shot video action localization method based on a multi-modal large model framework of the present application;

[0045] Figure 2 A schematic diagram of the overall principle of video action localization;

[0046] Figure 3 A schematic diagram of comparative results on the Activitynet v1.2 dataset;

[0047] Figure 4 A schematic diagram of comparative results on the Thumos14 dataset;

[0048] Figure 5 A schematic diagram of the structure of an embodiment of a multi-stage zero-shot video action localization system based on a multi-modal large model framework of the present application;

[0049] Figure 6 A schematic diagram of the structure of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION

[0050] Exemplary embodiments of the present application will be described in greater detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein.

[0051] Figure 1This diagram illustrates a flowchart of an embodiment of a multi-stage zero-shot video motion localization method based on a multimodal large model framework provided by the present invention. This method can be executed by electronic devices such as terminal devices or servers. The terminal device can be any fixed or mobile terminal, such as a user equipment (UE), mobile device, user terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, or wearable device. The server can be a single server or a server cluster consisting of multiple servers. Any electronic device can implement the multi-stage zero-shot video motion localization method based on the multimodal large model framework by having its processor call computer-readable instructions stored in its memory. Figure 1 As shown, it includes the following steps:

[0052] S1. Using a multimodal large language model, obtain at least one candidate video action category corresponding to the video to be tested, and determine multiple key action stages corresponding to each candidate video action category, arranged in chronological order of occurrence.

[0053] In this embodiment, the video to be tested is the video for which video action localization needs to be performed. The candidate video action category refers to the action category obtained by encoding the content of the video to be tested through a multimodal large language model, including but not limited to: standing, running, walking, and waving.

[0054] In S1, the step of obtaining at least one candidate video action category corresponding to the video to be tested using a multimodal large language model includes:

[0055] Based on multiple preset action categories and combined with action category prompts, a multimodal large language model is used to obtain at least one candidate video action category corresponding to the video under test.

[0056] Wherein, a list of action categories C = {c1, c2, ..., c N}, where N represents the total number of preset action categories, and the action category list contains N preset action categories. The action category prompt defaults to "Based on <action category list C>, select the corresponding action from the video to be tested". The candidate video action category set is: P MLLM (c i |V) represents a given video V to be tested and the i-th preset action category c. i τ is the probability value output by the multimodal large language model. c The threshold parameter is used to determine c. i Can I enter?

[0057] In S1, the step of determining a plurality of key action stages corresponding to each candidate video action category in chronological order includes:

[0058] Based on the to-be-tested video and at least one candidate video action category, and in combination with the text prompt word, a multi-modal large language model is used to generate text description information corresponding to each candidate video action category.

[0059] Wherein, the text prompt word is "describe the < candidate video action category in the to-be-tested video > by default. Specifically, for the to-be-tested video V and any candidate video action category in the candidate video action category set The text description of each candidate video action category of the to-be-tested video is generated by a multi-modal large language model desc is a text prompt word to guide the MLLM to generate natural language.

[0060] Using a multi-modal large language model, the text description information corresponding to each candidate video action category is divided into a plurality of key action stages in chronological order.

[0061] Specifically, the text description of the candidate video action category is divided into a plurality of action semantic stages, and the process is represented as: represents a plurality of action semantic stages of the i-th preset action category c i , s x represents the x-th action semantic stage, key x ∈{0,1} represents whether the x-th action semantic stage can be determined as an action semantic stage of the action ; a key action stage is generated by assigning a serial number to the plurality of action semantic stages in chronological order and in combination with the identifier, and the key action stage set is represented as: s k represents the k-th key action stage, key_num k represents the assigned serial number of s k (based on the chronological order, the number is generated from 1 and increases).

[0062] S2. For any candidate video action category, using the multimodal large language model, obtain the confidence level of each video frame of the test video for each key action stage corresponding to any candidate video action category. Determine the key action stage corresponding to the highest confidence level of each video frame as the target action stage of each video frame. Based on video frames with the highest confidence level greater than the confidence threshold, construct at least one candidate time segment. Determine the target action stage that appears most frequently in each candidate time segment as the final action stage of each candidate time segment. Merge all candidate time segments according to the occurrence time order of the final action stages of each candidate time segment and the interval between adjacent candidate time segments to obtain the video action segment localization result corresponding to any candidate video action category, until the video action segment localization result corresponding to each candidate video action category is obtained.

[0063] In S2, specifically:

[0064] S21. For the video V to be tested, if the total duration of the video is T and the total number of frames is J, then the frame sequence of the video is: in, t j F represents the j-th video frame. j The corresponding actual video time. For any candidate video action category. Using a multimodal large language model, obtain video frame F j In respectively Each corresponding key action phase s k The confidence level is expressed as: Indicates video frame F j The confidence level (probability value) corresponds to the k-th key action stage; finally, the confidence level (probability value) for each key action stage is obtained. k Frame-by-frame confidence distribution

[0065] S22. For the frame-by-frame confidence distribution corresponding to all key action stages, the key action stage corresponding to the highest confidence level in each video frame is determined as the target action stage for each video frame, specifically as follows: This represents the highest confidence level of the j-th video frame. This represents the key action stage number corresponding to the highest confidence level of the j-th video frame.

[0066] S23. For the highest confidence score of each video frame, select video frames with a highest confidence score greater than the confidence threshold to construct at least one candidate time segment S. raw Each candidate time segment S raw Represented as: F s F represents the start frame of the candidate time segment. e This represents the end frame of the candidate time segment, and θ represents the confidence threshold, used to determine whether a video frame is a high-confidence frame.

[0067] S24. For each candidate time segment, the number that appears most frequently among the key action stage numbers corresponding to all video frames in each candidate time segment is taken as the final action stage number of the corresponding candidate time segment. This process is expressed as follows: in Represents candidate time segments [F] s ,F e The corresponding final action stage number, I(.) is the indicator function, when The function takes a value of 1 if the condition is met, and zero otherwise. Indicates the candidate time segment [F] s ,F e Within [the context], the frequency of occurrence of the key action phase number m. Finally, the candidate time segment set S. raw Represented as: L represents the total number of candidate time segments.

[0068] S25. For two adjacent candidate time segments, if the occurrence time sequence of the final action stage of the preceding adjacent candidate time segment is earlier than the occurrence time sequence of the final action stage of the following adjacent candidate time segment, then the two adjacent candidate time segments are merged; if the occurrence time sequence of the final action stage of the preceding adjacent candidate time segment is the same as the occurrence time sequence of the final action stage of the following adjacent candidate time segment, and the time interval between the two adjacent candidate time segments is less than the interval duration threshold, then the two adjacent candidate time segments are merged.

[0069] Among them, for two adjacent candidate time segments and If the key action stage numbers corresponding to these two adjacent candidate time segments satisfy Then the candidate time segments Considered as candidate time segments Regarding action categories Therefore, in the preceding stage, these two adjacent candidate time segments are merged to obtain a new candidate time segment. This process is represented as follows:

[0070] Among them, for two adjacent candidate time segments and If these two adjacent candidate time segments belong to the same action phase and the time interval between them is small, then the following conditions are met: and it is considered that the two adjacent candidate time segments should be merged, and the process is represented as: Seg y represents the yth video action segment after merging.

[0071] It should be noted that tIoU is used to calculate the proportion of two adjacent candidate time segments in the time period, represented as θ IoU is the interval duration threshold, which controls whether two adjacent candidate time segments are merged.

[0072] S26, repeatedly performing S25 until the merging of all candidate time segments is completed, obtaining the video action segment positioning result corresponding to any candidate video action category.

[0073] Among them, the video action segment positioning result includes: the action interval prediction value and the corresponding video action category. The video to be tested contains at least one action interval prediction value, and each action interval prediction value corresponds to a video action category. For example, the video to be tested corresponds to the video action category of "running" from 1s to 3s, and the video action category of "kicking the ball" from 5s to 10s.

[0074] S27, for each candidate video action category, repeatedly performing S21-S26, obtaining the video action segment positioning result corresponding to each candidate video action category.

[0075] It should be noted that, Figure 2 The overall principle schematic diagram of the embodiment is shown. The performance of the video action positioning method in the embodiment is compared with the positioning accuracy of the international leading same kind model, Figure 3 The comparison results on the Activitynet v1.2 dataset are shown, Figure 4 The comparison results on the Thumos14 dataset are shown. Through comparison, it can be known that the positioning accuracy of the video action positioning method in the embodiment has significant superiority.

[0076] The technical scheme of the embodiment introduces a multi-modal large model, uses a picture-text semantic alignment and similarity calculation mechanism, and combines frame-level confidence score, effectively realizes the discrimination of action categories and the labeling of time sequence positions in the video, completely gets rid of the dependence on artificial labeled data, significantly improves the automation level of video action analysis, and improves the action positioning accuracy and stability in a multi-action complex scene.

[0077] Figure 5 The structure schematic diagram of an embodiment of a multi-stage zero-shot video action positioning system 200 based on a multi-modal large model framework provided by the application is shown. As Figure 5As shown, the system 200 comprises a processing module 210 and a running module 220.

[0078] The processing module 210 is configured to: acquire at least one candidate video action category corresponding to a to-be-tested video by using a multi-modal large language model, and determine a plurality of key action stages arranged in chronological order corresponding to each candidate video action category.

[0079] The running module 220 is configured to: for any candidate video action category, acquire a confidence of each video frame of the to-be-tested video in each key action stage corresponding to the any candidate video action category by using the multi-modal large language model, determine a target action stage corresponding to a highest confidence of each video frame as the target action stage of each video frame, construct at least one candidate time segment according to video frames with a highest confidence greater than a confidence threshold, determine a final action stage appearing most frequently in each candidate time segment as the final action stage of each candidate time segment, and merge all candidate time segments according to a chronological order of the final action stages of each candidate time segment and an interval duration between adjacent candidate time segments to obtain a video action segment positioning result corresponding to the any candidate video action category, until a video action segment positioning result corresponding to each candidate video action category is obtained.

[0080] In an optional manner, the running module 220 is specifically configured to:

[0081] For two adjacent candidate time segments, if a chronological order of the final action stage of a previous adjacent candidate time segment is located before a chronological order of the final action stage of a subsequent adjacent candidate time segment, the two adjacent candidate time segments are merged; if the chronological order of the final action stage of the previous adjacent candidate time segment is the same as the chronological order of the final action stage of the subsequent adjacent candidate time segment, and a time interval between the two adjacent candidate time segments is less than an interval duration threshold, the two adjacent candidate time segments are merged.

[0082] The step of performing for two adjacent candidate time segments is repeatedly executed until the merging of all candidate time segments is completed, and the video action segment positioning result corresponding to the any candidate video action category is obtained.

[0083] In an optional manner, the processing module 210 is specifically configured to:

[0084] Based on a plurality of preset action categories and in combination with an action category prompt word, the multi-modal large language model is used to acquire the at least one candidate video action category corresponding to the to-be-tested video.

[0085] In an optional mode, the processing module 210 is specifically used for:

[0086] Based on the to-be-tested video and the at least one candidate video action category, and in combination with the text prompt word, the multi-modal large language model is used to generate text description information corresponding to each candidate video action category;

[0087] The multi-modal large language model is used to divide the text description information corresponding to each candidate video action category into a plurality of key action stages arranged in chronological order.

[0088] It should be noted that the beneficial effects of the multi-stage zero-shot video action positioning system 200 based on the multi-modal large model framework provided in the above embodiments are the same as those of the multi-stage zero-shot video action positioning method based on the multi-modal large model framework, and will not be repeated here. In addition, when the system provided in the above embodiments implements its functions, only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to the needs, that is, the system is divided into different functional modules according to the actual situation to complete all or part of the above described functions. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is shown in the method embodiments, which will not be repeated here.

[0089] Among them, the multi-stage zero-shot video action positioning system 200 based on the multi-modal large model framework of the present application can be a computer program (including program code) running in a computer device, for example, the multi-stage zero-shot video action positioning system based on the multi-modal large model framework of the present application is an application software, which can be used to execute the corresponding steps in the multi-stage zero-shot video action positioning method based on the multi-modal large model framework of the present application.

[0090] In some embodiments, the multi-stage zero-shot video motion localization system based on a multimodal large model framework of the present invention can be implemented in a combination of hardware and software. As an example, the multi-stage zero-shot video motion localization system based on a multimodal large model framework of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the multi-stage zero-shot video motion localization method based on a multimodal large model framework of the present invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0091] The modules described in the embodiments of this invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.

[0092] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned multi-stage zero-sample video motion localization methods based on a multimodal large model framework. That is, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store the computer program; the processor is used to execute the multi-stage zero-sample video motion localization method based on a multimodal large model framework shown in any embodiment of the present invention by calling the computer program.

[0093] In one alternative embodiment, an electronic device is provided, such as Figure 6 As shown, Figure 6 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.

[0094] The processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in connection with the present disclosure. The processor 4001 can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0095] The bus 4002 can include a path for transmitting information between the above-mentioned components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 can be divided into an address bus, a data bus, a control bus, and the like. For convenience of representation, Figure 6 The bus 4002 is represented by only one thick line, but it does not mean that there is only one bus or only one type of bus.

[0096] The memory 4003 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited thereto.

[0097] The memory 4003 is configured to store application code (computer program) for implementing the solutions of the present application, and the processor 4001 is configured to control the execution. The processor 4001 is configured to execute the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0098] The electronic device can also be a terminal device, and the terminal device can be any terminal device that can install an application and access a webpage through the application, including at least one of a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart television, and a smart vehicle device.

[0099] It should be noted that, Figure 6 The electronic device shown is only an example and should not limit the functions and use range of the embodiments of the present application.

[0100] The computer readable storage medium of the embodiment of the present application, the computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize any one of the above-mentioned multi-stage zero-shot video action positioning methods based on the multi-modal large model framework.

[0101] Optionally, the computer readable storage medium can be a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a read-only compact disc (Compact Disc Read-Only Memory, CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0102] In the exemplary embodiments, a computer program product or computer program is also provided, which includes computer instructions stored in a computer readable storage medium. The processor of the electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the electronic device execute the above-mentioned multi-stage zero-shot video action positioning method based on the multi-modal large model framework.

[0103] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0104] It should be understood that the flowchart and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of various embodiments of the present application. In this regard, each block in the flowchart and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.

[0105] The computer readable storage medium of embodiments of the present application can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to, the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, the computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0106] The computer readable storage medium described above bears one or more programs, when the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0107] The above description is merely exemplary of the application and the application principles of the technology used. Those skilled in the art should understand that the disclosed range of the application is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the disclosed concept. For example, the above features are replaced with the technical features disclosed in the application (but not limited to) having similar functions to form technical solutions.

[0108] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application are used to distinguish similar objects, and represent a specific order or sequence. The order of use of similar objects can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described.

[0109] Those skilled in the art know that the application can be implemented as a system, a method or a computer program product, so the application can be specifically implemented as follows: it can be a complete hardware, a complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, which is generally referred to as "circuit", "module" or "system" in this paper. In addition, in some embodiments, the application can also be implemented as a computer program product in one or more computer readable media, which contains computer readable program code.

[0110] Although the embodiments of the application have been shown and described above, it should be understood that the above embodiments are exemplary and cannot be understood as limiting the application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the application.

Claims

1. A multi-stage zero-shot video action localization method based on a multi-modal large model framework, characterized in that, The method comprises the following steps: acquiring at least one candidate video action category corresponding to the to-be-tested video by using a multimodal large language model, and determining a plurality of key action stages corresponding to each candidate video action category in the order of occurrence time; for any candidate video action category, acquiring the confidence of each video frame of the to-be-tested video in each key action stage corresponding to the any candidate video action category by using the multimodal large language model, determining the key action stage corresponding to the highest confidence of each video frame as the target action stage of each video frame, constructing at least one candidate time segment according to the video frame whose highest confidence is greater than a confidence threshold, determining the final action stage of each candidate time segment as the target action stage appearing most frequently in each candidate time segment, and merging all candidate time segments according to the occurrence time sequence of the final action stage of each candidate time segment and the interval duration between adjacent candidate time segments to obtain the video action segment positioning result corresponding to the any candidate video action category, until the video action segment positioning result corresponding to each candidate video action category is obtained.

2. The multi-stage zero-shot video action localization method based on the multi-modal large model framework according to claim 1, characterized in that, The step of merging all candidate time segments according to the occurrence time sequence of the final action stage of each candidate time segment and the interval duration between adjacent candidate time segments to obtain the video action segment positioning result corresponding to the any candidate video action category comprises the following steps: for two adjacent candidate time segments, if the occurrence time sequence of the final action stage of the former adjacent candidate time segment is located before the occurrence time sequence of the final action stage of the latter adjacent candidate time segment, the two adjacent candidate time segments are merged; if the occurrence time sequence of the final action stage of the former adjacent candidate time segment is the same as the occurrence time sequence of the final action stage of the latter adjacent candidate time segment, and the time interval between the two adjacent candidate time segments is less than an interval duration threshold, the two adjacent candidate time segments are merged; the steps of the two adjacent candidate time segments are repeatedly executed until the merging of all candidate time segments is completed, and the video action segment positioning result corresponding to the any candidate video action category is obtained.

3. The multi-stage zero-shot video action localization method based on the multi-modal large model framework according to claim 1, characterized in that, The step of acquiring at least one candidate video action category corresponding to the to-be-tested video by using a multimodal large language model comprises the following steps: based on a plurality of preset action categories and in combination with an action category prompt word, the multimodal large language model is used to acquire the at least one candidate video action category corresponding to the to-be-tested video.

4. The multi-stage zero-shot video action localization method based on the multi-modal large model framework of claim 1, wherein, The step of determining a plurality of key action stages corresponding to each candidate video action category in the order of occurrence time comprises the following steps: based on the to-be-tested video and the at least one candidate video action category, and in combination with a text prompt word, the multimodal large language model is used to generate text description information corresponding to each candidate video action category; the multimodal large language model is used to divide the text description information corresponding to each candidate video action category into a plurality of key action stages arranged in the order of occurrence time.

5. A multi-stage zero-shot video action localization system based on a multi-modal large model framework, characterized in that, The method comprises the following steps: a processing module and a running module; The processing module is configured to: acquire at least one candidate video action category corresponding to the to-be-tested video by using a multi-modal large language model, and determine a plurality of key action stages corresponding to each candidate video action category in an order of occurrence time; The running module is configured to: for any candidate video action category, acquire a confidence degree of each video frame of the to-be-tested video in each key action stage corresponding to the any candidate video action category by using the multi-modal large language model, determine a target action stage corresponding to a highest confidence degree of each video frame as the target action stage of each video frame, construct at least one candidate time segment according to the video frame whose confidence degree is greater than a confidence threshold, determine a final action stage of each candidate time segment as the action stage that appears most frequently in each candidate time segment, and merge all candidate time segments according to an order of occurrence time of the final action stage of each candidate time segment and an interval duration between adjacent candidate time segments to obtain a video action segment positioning result corresponding to the any candidate video action category, until a video action segment positioning result corresponding to each candidate video action category is obtained.

6. The multi-stage zero-shot video action localization system based on a multi-modal large model framework of claim 5, wherein, The running module is specifically configured to: For two adjacent candidate time segments, if an order of occurrence time of the final action stage of a previous adjacent candidate time segment is located before an order of occurrence time of the final action stage of a subsequent adjacent candidate time segment, the two adjacent candidate time segments are merged; if the order of occurrence time of the final action stage of the previous adjacent candidate time segment is the same as the order of occurrence time of the final action stage of the subsequent adjacent candidate time segment, and a time interval between the two adjacent candidate time segments is less than an interval duration threshold, the two adjacent candidate time segments are merged; The step of performing the operation on the two adjacent candidate time segments is repeatedly executed until the merging of all candidate time segments is completed, and the video action segment positioning result corresponding to the any candidate video action category is obtained.

7. The multi-stage zero-shot video action localization system based on a multi-modal large model framework of claim 5, wherein, The processing module is specifically configured to: Based on a plurality of preset action categories and in combination with an action category prompt word, the multi-modal large language model is used to acquire the at least one candidate video action category corresponding to the to-be-tested video.

8. The multi-stage zero-shot video action localization system based on a multi-modal large model framework of claim 5, wherein, The processing module is specifically configured to: Based on the to-be-tested video and the at least one candidate video action category, and in combination with a text prompt word, the multi-modal large language model is used to generate text description information corresponding to each candidate video action category; The multi-modal large language model is used to divide the text description information corresponding to each candidate video action category into a plurality of key action stages arranged in an order of occurrence time.

9. An electronic device, comprising: The electronic device includes a processor coupled with a memory, and the memory stores at least one computer program, which is loaded and executed by the processor, so that the electronic device implements the multi-stage zero-shot video action positioning method based on the multi-modal large model framework as claimed in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor, so that the computer readable storage medium implements the multi-stage zero-shot video action localization method based on the multi-modal large model framework as claimed in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Action recognition method and system based on multi-modal sequence fusion

    CN115937975A

  • Cross-modal power video positioning method and system, electronic equipment and storage medium

    CN119888563A