A first-person video positioning method and system

By extracting fine-grained object semantic information from the first-view video and performing lens-text comparison learning, the information capture and understanding problems in the first-view video positioning task are solved, and positioning accuracy and video comprehension capabilities of the model are improved.

CN120032301BActive Publication Date: 2025-07-04HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN) +4
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510510087.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-04
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture fine-grained semantic information in first-view videos, and handle the positioning difficulty caused by low video quality and long video length, resulting in poor performance of the model in first-view video positioning task.

Method used

By extracting fine-grained object semantic information and performing lens-text comparison learning, combining multimodal fusion module and contrast learning methods, video representation and understanding capabilities are enhanced.

Benefits of technology

The positioning accuracy of fine-grained query and the model's understanding of first-view videos is significantly improved, and the challenges brought by long-sequence video features and lens movements are better handled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032301B_ABST
    Figure CN120032301B_ABST
Patent Text Reader

Abstract

The present invention provides a first-person video localization method and system, which acquire a first-person video and a query text; use a pre-trained object detector to extract object annotations from the first-person video, and filter out object categories related to the query by matching nouns in the query text; use a pre-trained feature encoder to encode video, object, and text information, extract video features, object features, and text features, perform text feature context modeling, and execute feature interaction between text and objects; use a multi-modal fusion module that includes a linear time series model using a selective state space and cross-attention to understand and fuse video feature sequences, and obtain a multi-modal feature representation; use the multi-modal feature representation to perform first-person video segment localization. The present invention overcomes the defects in the prior art of lacking fine-grained semantic information and being difficult to understand first-person videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video localization, and particularly relates to a first-person video localization method and system. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] The first-person video grounding task is to locate a specific video segment corresponding to a given natural language query text and an unedited first-person video. Locating the video segment requires fully understanding the meaning of the natural language query text and at the same time fully understanding visual cues such as the behavior of the people in the first-person video, so as to accurately retrieve the corresponding video segment. This technology can transform personal daily life videos into "external memories", helping users quickly locate the required segments and solve common problems such as forgetting the location of items. Therefore, it has broad application value in various first-person scenarios, such as using it as a smart assistant to retrieve the user's first-person videos in real time, or as a memory module to help a household robot recall the user's activities.

[0004] Therefore, the first-person video grounding task has attracted extensive attention in the academic and industrial communities. Many researchers have made a series of improvements to the original deep learning methods and achieved many progresses. The existing methods mainly develop in two directions: 1) Video-text pre-training. Due to the difference in feature distributions between first-person videos and third-person videos, some studies have fine-tuned the video backbone encoder using first-person videos, thereby improving the robustness of the model in first-person video understanding tasks. 2) Data augmentation. Narration as Query constructs a large video-narration pair dataset using existing narration data, pre-trains the downstream model on this video-narration pair dataset first, and then fine-tunes it on the first-person video grounding dataset, achieving significant improvements.

[0005] However, previous methods usually regard the first-person video grounding task as a general long video localization problem, but this also makes it difficult for existing methods to cope with the following new challenges posed by first-person videos:

[0006] 1) Limited video subject information. First-person videos are usually captured by wearable devices, and the videos only contain a small part of the body (such as hands, feet) or local scenes, lacking complete visual information. This limitation makes it difficult for the model to extract sufficient semantic information from the video to support natural language queries. For example, when the user queries "where is the key placed", previous methods may not be able to accurately identify the location of the key in the video because the key may only appear at the edge or in the background of the picture. In addition, the expression of the subject's behavior in the first-person video is also relatively implicit. For example, the hand movement may only occupy a small part of the picture, further increasing the difficulty for the model to understand the video content. Therefore, how to capture sufficient information in the first-person scenario to support fine-grained queries has become a key problem that current video localization methods urgently need to solve.

[0007] 2) Low video quality. Due to the movement of the first-person video shooting device and the head movement of the cameraman, problems such as perspective jitter, blurring, and large-scale movement often occur in the video, resulting in a significant decline in video quality. This low-quality video data brings additional difficulties to model learning. For example, jitter and blurring may cause key objects or actions to not be clearly presented, and large-scale movement may make it difficult for the model to track continuous actions or scene changes in the video. Therefore, when facing first-person videos, the performance of existing models is often unsatisfactory.

[0008] 3) Long video length. First-person videos usually record the user's daily activities for a long time, and the video length is much longer than that of traditional third-person videos. This long video characteristic makes the localization task more complex because the model needs to search for relevant segments within a longer time range.

[0009] In summary, previous methods have not effectively solved the difficulties in video understanding, low video quality, and high localization difficulty in the first-person video localization task. Therefore, more targeted solutions are needed. Summary of the Invention

[0010] To solve the above problems, the present invention proposes a first-person video localization method and system. The present invention enhances video representation by mining fine-grained item semantic information from the video and inputting it into the model, and improves the model's understanding ability of the video by segmenting the shots and performing shot-text contrast learning, overcoming the defects of lacking fine-grained semantic information and being difficult to understand first-person videos in the prior art.

[0011] According to some embodiments, the present invention adopts the following technical solutions:

[0012] A first-person video localization method, comprising the following steps:

[0013] Obtain a first-person video and a query text;

[0014] Extract object annotations from the first-person video using a pre-trained object detector, and filter out the object categories related to the query by matching with the nouns in the query text;

[0015] Use a pre-trained feature encoder to encode video, object, and text information, extract video features, object features, and text features, perform context modeling on the text features, and execute feature interaction between the text and the objects;

[0016] Use a multi-modal fusion module that includes a linear time series model using a selective state space and cross-attention to understand and fuse the video feature sequence, and obtain a multi-modal feature representation;

[0017] Use the multi-modal feature representation to perform first-person video segment localization.

[0018] As an alternative implementation, the process of extracting object annotations from the first-person video using a pre-trained object detector and filtering out the object categories related to the query by matching with the nouns in the query text includes: extracting object annotations frame by frame from the first-person video using a pre-trained object detector, and generating structured information including object categories, confidence levels, and their appearance timestamps;

[0019] By performing natural language processing on the query text, extract the noun phrases as key query words, and use semantic similarity calculation to match the extracted object categories with the query words;

[0020] Filter out the object categories with confidence levels and semantic similarity to the query higher than the set threshold as the object annotations for this frame.

[0021] As an alternative implementation, the process of using a pre-trained feature encoder to encode video, object, and text information, and extract video features, object features, and text features includes: using a pre-trained video backbone model to extract segment-level video features, and then project them into the feature space to obtain video representations;

[0022] Use a pre-trained text backbone model to extract text features, and then use a text encoder composed of Transformer to perform context interaction on the text features;

[0023] Use a pre-trained text backbone model to extract object features, and use an object encoder to refine the object features related to the query.

[0024] As a further step, the process of refining item features related to a query using an item encoder includes encoding the text of the detected item category, using an item encoder composed of multiple layers of Transformers to refine the item features related to the query, using the item features as queries, and the query text features as keys and values, to obtain item features related to the query text.

[0025] As an alternative implementation, the process of using a multi-modal fusion module that includes a linear time series model using a selective state space and cross-attention for video feature sequence understanding and feature fusion includes: using a linear time series model with a bidirectional selective state space to enhance video features to capture long-range dependencies in video data; applying a cross-attention mechanism and a feed-forward layer to aggregate the enhanced video features and query text information; using a parallel cross-attention mechanism to aggregate video features and item features;

[0026] Combining the two features that aggregate different information through a gating mechanism.

[0027] As an alternative implementation, the process of using the multi-modal feature representation for first-person video clip localization includes: using a multi-scale network composed of multiple layers of Transformers to generate a feature pyramid, where each layer of Transformer includes a 1D depth convolution before the self-attention and FFN modules (feed-forward layer, Feed Forward Network, FFN) to achieve sequence downsampling and obtain representations of multi-scale candidate clips;

[0028] Decoding the multi-scale feature pyramid into the final prediction of video localization through a task head, where the classification head predicts the confidence score of each candidate clip, and the regression head predicts the offset of the candidate clip boundary relative to the anchor point.

[0029] As an alternative implementation, it further includes the following steps: enhancing the first-person video feature representation using contrastive learning in the shot side branch, specifically including: using a pre-trained narrative model to segment shots;

[0030] Extracting the video features and query-level text features of each segmented shot;

[0031] Aggregating the shots and the query-level text features, projecting the text and video features into a joint semantic space, and using them for contrastive learning.

[0032] A first-person video localization system includes:

[0033] An acquisition module configured to acquire a first-person video and a query text;

[0034] A preprocessing module, configured to extract item annotations from a first-person view video using a pre-trained object detector, and filter out item categories related to the query by matching nouns in the query text;

[0035] A feature extraction module, configured to encode video, item, and text information using a pre-trained feature encoder, extract video features, item features, and text features, perform text feature context modeling, and perform feature interaction between text and items;

[0036] A multimodal fusion module, configured to perform video feature sequence understanding and feature fusion using a multimodal fusion module that includes a linear time series model using a selective state space and cross-attention to obtain a multimodal feature representation;

[0037] A localization module, configured to use the multimodal feature representation to perform first-person view video segment localization.

[0038] A computer-readable storage medium for storing computer instructions that, when executed by a processor, complete the steps in the above method.

[0039] An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor, which, when run by the processor, complete the steps in the above method.

[0040] Compared with the prior art, the beneficial effects of the present invention are:

[0041] The present invention discloses a video localization method and system for first-person view videos, including a first-person view video localization method based on fine-grained item semantic enhanced video representation and a feature representation enhancement method based on contrast learning for the side branch of the shot, considering multiple key factors, including difficult video understanding, low video quality, and large localization difficulty, etc.

[0042] The first-person view video localization method based on fine-grained item semantic enhanced video representation of the present invention can localize fine-grained query texts; in order to ensure that the model can understand and localize fine-grained query texts, by integrating fine-grained item information into the video localization task, the model can obtain information such as item categories in video frames, significantly improving the localization accuracy of queries related to background items.

[0043] The feature representation enhancement method based on contrast learning for the side branch of the shot of the present invention can enhance the model's understanding of first-person view videos; in order to ensure that the model can better understand the frequent shot movements in the first-person view, query-level text features and shot-level video features are extracted from text features and video features, and the alignment ability between text representation and video representation is enhanced through contrast learning.

[0044] In the present invention, by using a linear time series model with a bidirectional selective state space (i.e., the Mamba network) in the multi-modal fusion module to perform self-interaction processing on videos, historical information can be better memorized, so that more context information can be fused when processing long-sequence video features.

[0045] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following preferred embodiments are specifically given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings

[0046] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0047] Figure 1 It is a flowchart of a video localization method for the characteristics of the first-person view video in the first embodiment of the present invention;

[0048] Figure 2 It is a flowchart of the process of extracting item annotations in the first embodiment of the present invention;

[0049] Figure 3 It is a schematic diagram of the multi-modal feature fusion process in the first embodiment of the present invention. Detailed Embodiments

[0050] The present invention will be further described below in conjunction with the drawings and embodiments.

[0051] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0052] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0053] Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0054] Embodiment 1:

[0055] The first embodiment of the present invention provides a video localization method for the characteristics of the first-person view video, as Figure 1 shown, which specifically includes the following steps:

[0056] Step 1: Enhance the video representation using fine-grained object semantics and perform first-person video localization.

[0057] Step 1.1: As Figure 2 shown, extract fine-grained object information from the video.

[0058] In a specific implementation, use an object detector to extract object annotations from the video. Then, use the spaCy library to extract nouns from the natural language query as the objects to be queried. Then, also use the spaCy library to calculate the text similarity between the queried object and the object category name. If it is higher than a specific threshold, it is considered that the object category contained in the object to be queried by the natural language query contains this object category. Finally, filter out the object categories with high confidence in the video frames and contained in the natural language query as the object category information and input it into the model.

[0059] Step 1.2: Encode the video, object information, and text into features. As Figure 1 shown, extract features from the input video, query text, and object information (i.e., object category information) through a pre-trained base model and a designed encoder, perform text context interaction, and refine the object features related to the query.

[0060] Step 1.2.1: Use a pre-trained video base model to extract video features.

[0061] In a specific implementation, assume the video consists of frames, denoted as . Divide it into non-overlapping segments, denoted as , where represents the size of the sliding window, and represents the number of segments in the video. The pre-trained video base model is used to extract segment-level video features, and then project them into the feature space to obtain the video representation as .

[0062] Step 1.2.2: Use a pre-trained text base model to extract text features, and then use a text encoder composed of Transformer to perform context interaction on the text features.

[0063] In a specific implementation, assume the query text contains words. Use the pre-trained text base model to extract query text features (which can be simply referred to as text features) to obtain word-level features . To capture the relationships between words in the query, this embodiment uses a text encoder for encoding and outputs the query representation .

[0064] Step 1.2.3: Extract item features using a pre-trained text base model and use an item encoder to refine the item features related to the query.

[0065] In a specific implementation, encode the text of the detected item category to obtain , and then use an item encoder composed of multiple layers of Transformer to refine the item features related to the query, where the item features serve as the query, and the query text features serve as the key and value, and finally obtain the item features related to the query text .

[0066] Step 1.3: Use a multi-modal fusion module, as Figure 3 shown, to obtain a multi-modal feature representation by fusing text, item, and video features through a multi-layer structure.

[0067] Step 1.3.1: Use a linear time series model (Mamba) layer with bidirectional selective state space to enhance the video features to capture long-range dependencies in the video data.

[0068] In a specific implementation, denote the video feature output of the th layer as , and . The formal process of obtaining the context-enhanced video feature is as follows:

[0069] ;

[0070] Step 1.3.2: Apply a cross-attention module (Cross Attention, CA) and a feed-forward network (FeedForward Network, FFN) to aggregate the context-enhanced video feature and the query text information to obtain the text-enhanced video feature , and the formal process is as follows:

[0071] ;

[0072] ;

[0073] Step 1.3.3: Similarly, use a parallel cross-attention module to aggregate the context-enhanced video feature and the item feature to obtain the item-enhanced video feature , and the formal process is as follows:

[0074] ;

[0075] ;

[0076] Step 1.3.4, combine the two differently information-enhanced video features through a gating mechanism to obtain the output of the th layer, which is also the input of the th layer , and the formalization process is as follows:

[0077] ;

[0078] ;

[0079] where represents the Sigmoid function, represents vector concatenation, and MLP represents the Multilayer Perceptron.

[0080] Step 1.4, use a multi-scale network composed of multiple layers of Transformer to generate a feature pyramid, and each layer includes a 1D depth convolution before the self-attention and FFN modules to achieve sequence downsampling and obtain multi-scale candidate segment representations.

[0081] Step 1.5, decode the multi-scale feature pyramid into the final prediction of video localization through the task head. The classification head predicts the confidence score of each candidate segment, and the regression head predicts the offset of the candidate segment boundary relative to the anchor point.

[0082] Step 1.6, use the video segment localization loss to supervise the model learning, and the formalization is as follows:

[0083] ;

[0084] where, is the focal loss, which is used to supervise the learning of the classification head, is the Distance-IoU loss, which is used to supervise the learning of the regression head.

[0085] Some embodiments further include Step 2, enhancing the first-person video feature representation by using a contrastive learning side branch of the shot.

[0086] Step 2.1, use the pre-trained narrative model to segment the shots.

[0087] In a specific embodiment, first, a pre-trained narrative model LAVILA is used to generate a narrative for the video, and in the generated video captions, expressions such as "looks around" or "turns around" are usually used to describe the head movement of the wearer. These cues reflect the changes in the wearer's attention during the video shooting process and segment the video into multiple semantically different shots.

[0088] Step 2.2, extract the video features at the shot level and the text features at the query level.

[0089] Step 2.2.1, to effectively perform contrastive learning, a bidirectional Mamba block integrated with an MLP that shares parameters with the multimodal fusion module is used to extract the context video features representing each shot content from the video, which are irrelevant to the query and item information. 。

[0090] Step 2.2.2, use a feature aggregator of the Transformer architecture to aggregate the shot and query features.

[0091] In a specific embodiment, a learnable embedding is used as the query, while the shot context features and the query text features are used as keys and values, and then, one-dimensional convolution is applied to obtain the final shot representation as well as the text representation , is the number of shots, is the number of queries.

[0092] Step 2.3, use an MLP to project the shot representation and the text representation into the joint semantic space and use them for contrastive learning.

[0093] In a specific embodiment, a contrastive learning loss is used to supervise the learning of the shot branch, which is formalized as follows:

[0094] ;

[0095] where, represents the cosine similarity, is the set of positive query-shot pairs where the video segment corresponding to the query intersects with the corresponding shot, is the temperature coefficient.

[0096] Embodiment 2:

[0097] Embodiment 2 of the present invention provides a video localization system for the characteristics of first-person perspective videos, including:

[0098] An acquisition module, configured to acquire a first - perspective video and a query text;

[0099] A pre - processing module, configured to extract item annotations from the first - perspective video using a pre - trained item detector, and filter out item categories related to the query by matching nouns in the query text;

[0100] A feature extraction module, configured to encode video, item, and text information using a pre - trained feature encoder, extract video features, item features, and text features, perform text - feature context modeling, and execute feature interaction between text and items;

[0101] A multi - modal fusion module, configured to perform video - feature sequence understanding and feature fusion using a multi - modal fusion module that includes a linear time - series model using a selective state space and cross - attention to obtain a multi - modal feature representation;

[0102] A localization module, configured to use the multi - modal feature representation to perform first - perspective video segment localization.

[0103] Each step involved in the second embodiment above corresponds to that in the first method embodiment. For specific implementation manners, reference may be made to the relevant description part of the first embodiment. The term "computer - readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0104] Those skilled in the art should understand that the above - mentioned modules or steps of the present invention can be implemented by a general - purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device for execution by the computing device, or they can be separately fabricated into individual integrated - circuit modules, or multiple modules or steps among them can be fabricated into a single integrated - circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0105] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made by those skilled in the art without creative efforts within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A first-person video positioning method, characterized in that, It includes the following steps: Obtain a first-person view video and a query text; Use a pre-trained object detector to extract object annotations from the first-person view video, and filter out the object categories related to the query by matching with the nouns in the query text; Utilize a pre-trained feature encoder to encode video, object, and text information, extract video features, object features, and text features, perform context modeling on the text features, and execute feature interaction between the text and the object; Utilize a multi-modal fusion module that includes a linear time series model using a selective state space and cross-attention to perform video feature sequence understanding and feature fusion to obtain a multi-modal feature representation; Use the multi-modal feature representation to perform first-person view video segment localization; Among them, the process of using a multi-modal fusion module that includes a linear time series model using a selective state space and cross-attention to perform video feature sequence understanding and feature fusion includes: using a linear time series model with a bidirectional selective state space to enhance video features to capture long-range dependencies in the video data; applying a cross-attention mechanism and a feed-forward layer to aggregate the enhanced video features and query text information; using a parallel cross-attention mechanism to aggregate video features and object features; Combine the features that aggregate different information through a gating mechanism.

2. The first - perspective video positioning method according to claim 1, wherein The process of using a pre-trained object detector to extract object annotations from the first-person view video and filtering out the object categories related to the query by matching with the nouns in the query text includes: using a pre-trained object detector to extract object annotations frame by frame from the first-person view video to generate structured information including object categories, confidence levels, and their appearance timestamps; Extract the noun phrases in the query text as key query words through natural language processing of the query text, and use semantic similarity calculation to match the extracted object categories with the query words; Filter out the object categories with confidence levels and semantic similarities to the query higher than the set threshold as the object annotations for this frame.

3. A first - perspective video positioning method according to claim 1, characterized in that, The process of using a pre-trained feature encoder to encode video, object, and text information and extract video features, object features, and text features includes: using a pre-trained video backbone model to extract segment-level video features, and then projecting them into the feature space to obtain video representations; Use a pre-trained text backbone model to extract text features, and then use a text encoder composed of Transformers to perform context interaction on the text features; Use a pre-trained text backbone model to extract object features, and use an object encoder to refine the object features related to the query.

4. The first - perspective video positioning method according to claim 3, wherein, The process of using an object encoder to refine the object features related to the query includes encoding the text of the detected object categories, using an object encoder composed of multiple layers of Transformers to refine the object features related to the query, using the object features as queries, and the query text features as keys and values to obtain object features related to the query text.

5. The first - perspective video positioning method according to claim 1, wherein, The process of performing first-person video clip localization using the described multi-modal feature representation includes: generating a feature pyramid using a multi-scale network composed of multiple layers of Transformer. Each layer of Transformer includes a 1D depth convolution before the self-attention and FFN modules to achieve sequence downsampling and obtain the representation of multi-scale candidate clips. Decoding the multi-scale feature pyramid into the final prediction of video localization through a task head. The classification head predicts the confidence score of each candidate clip, and the regression head predicts the offset of the candidate clip boundary relative to the anchor point.

6. A first - perspective video positioning method according to any one of claims 1 - 5, characterized in that, It also includes the following steps: Enhancing the first-person video feature representation using the contrastive learning shot side branch, specifically including: segmenting shots using a pre-trained narrative model; Extracting the video features and query-level text features of each segmented shot; Aggregating the shot and the query-level text features, projecting the text and video features into a joint semantic space, and using them for contrastive learning.

7. A first-person video positioning system, characterized in that, It includes: An acquisition module configured to acquire a first-person video and a query text; A preprocessing module configured to extract object annotations from the first-person video using a pre-trained object detector and filter out the object categories related to the query by matching with the nouns in the query text; A feature extraction module configured to encode video, object, and text information using a pre-trained feature encoder, extract video features, object features, and text features, perform text feature context modeling, and execute feature interaction between text and objects; A multi-modal fusion module configured to perform video feature sequence understanding and feature fusion using a multi-modal fusion module that includes a linear time series model with a selective state space and cross-attention to obtain a multi-modal feature representation. Among them, the process of performing video feature sequence understanding and feature fusion using a multi-modal fusion module that includes a linear time series model with a selective state space and cross-attention includes: enhancing video features using a bidirectional selective state space linear time series model to capture long-range dependencies in video data; applying a cross-attention mechanism and a feed-forward layer to aggregate the enhanced video features and query text information; using a parallel cross-attention mechanism to aggregate video features and object features; Combining the features that aggregate different information through a gating mechanism; A localization module configured to perform first-person video clip localization using the described multi-modal feature representation.

8. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by a processor, the steps in the method described in any one of claims 1-6 are completed.

9. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the method described in any one of claims 1-6 are completed.

Citation Information

Patent Citations

  • Time sequence behavior detection method and device, equipment, medium and program product

    CN117218572A

  • First view video description system based on retrieval enhancement

    CN119226567A