First view angle video positioning method and system
By using fine-grained object semantic information and multimodal feature fusion technology in the first-view video positioning, the problems of difficult video comprehension, low video quality and high positioning are solved, and a more accurate and efficient video positioning effect is achieved.
Patent Information
- Application Number
- CN202510510087.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The prior art is difficult to capture sufficient information in a first-view video to support fine-grained query, and faces problems with low video quality and long video length, which makes positioning difficult.
By mining fine-grained object semantic information from the video, and extracting video, item and text features using pre-trained item detectors and feature encoders, performing multimodal feature fusion, using a linear time series model of selective state space and a multimodal fusion module of cross attention to video feature sequence understanding and feature fusion, and finally using a multi-scale network composed of multi-layer Transformer for video positioning.
It significantly improves the positioning accuracy of background items related queries, enhances the model's understanding of first-view videos, can better process long-sequence video features, and achieve more accurate positioning in low-quality videos.
Smart Images

Figure CN120032301A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video positioning, and in particular relates to a first-person perspective video positioning method and system. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] The task of egocentric video grounding is to locate a specific video clip corresponding to a given natural language query text and an unedited first-person video. To locate the video clip, it is necessary to fully understand the meaning of the natural language query text, and at the same time fully understand the visual clues such as the behavior of the characters in the first-person video, so as to accurately retrieve the corresponding video clip. This technology can convert videos of personal daily life into "external memory", helping users to quickly locate the required clips and solve common problems such as forgetting the location of items. Therefore, it has a wide range of application value in various first-person scenarios, such as as an intelligent assistant to retrieve users' first-person videos in real time, or as a memory module to help home robots trace back user activities.
[0004] Therefore, the first-person video localization task has attracted widespread attention from academia and industry. Many researchers have made a series of improvements to the original deep learning methods and made a lot of progress. Existing methods are mainly developed in two directions: 1) Video text pre-training. Due to the difference in feature distribution between first-person and third-person videos, some studies have used first-person videos to fine-tune the video base encoder, thereby improving the robustness of the model in the first-person video understanding task. 2) Data enhancement. The Narration as Query strategy uses existing narrative data to build a large video narrative pair dataset, pre-train the downstream model on the video narrative pair dataset, and then fine-tune it on the first-person video localization dataset, achieving significant improvements.
[0005] However, previous methods usually regard the first-person video localization task as a general long video localization problem, which also makes it difficult for existing methods to cope with the following new challenges raised by first-person videos: 1) Limited information about the subject of the video. First-person videos are usually shot by wearable devices, and the video only contains a small part of the limbs (such as hands and feet) or local scenes, lacking complete visual information. This limitation makes it difficult for the model to extract enough semantic information from the video to support natural language queries. For example, when a user queries "where are the keys", previous methods may not be able to accurately identify the location of the keys in the video because the keys may only appear at the edge or in the background of the screen. In addition, the expression of the subject's behavior in the first-person video is also relatively obscure. For example, hand movements may only occupy a small part of the screen, further increasing the difficulty of the model to understand the video content. Therefore, how to capture enough information to support fine-grained queries in the first-person scenario has become a key problem that needs to be solved urgently in current video positioning methods.
[0006] 2) Low video quality. Due to the movement of the first-person video shooting device and the cameraman's head movement, the video often has problems such as perspective jitter, blur, and large-scale movement, resulting in a significant decrease in video quality. This low-quality video data brings additional difficulties to model learning. For example, jitter and blur may cause key objects or actions to not be presented clearly, while large-scale movement may make it difficult for the model to track continuous actions or scene changes in the video. Therefore, when faced with first-person videos, the performance of existing models is often unsatisfactory.
[0007] 3) Long video length. First-person videos usually record users’ daily activities for a long time, and the video length is much longer than traditional third-person videos. This long video feature makes the positioning task more complicated because the model needs to search for query-related segments over a longer time range.
[0008] In summary, previous methods have failed to effectively solve the difficulties in first-person video positioning tasks, such as difficult video understanding, low video quality, and high positioning difficulty. Therefore, more targeted solutions are needed. Summary of the invention
[0009] In order to solve the above problems, the present invention proposes a first-person video positioning method and system. The present invention enhances the video representation by mining fine-grained object semantic information from the video and inputting it into the model, and improves the model's understanding of the video by splitting the shots and performing shot-text comparative learning, thereby overcoming the defects of the prior art of lacking fine-grained semantic information and difficulty in understanding the first-person video.
[0010] According to some embodiments, the present invention adopts the following technical solutions: A first-person video positioning method comprises the following steps: Get the first-person video and query text; Extract object annotations from the first-person video using a pre-trained object detector, and filter out object categories related to the query by matching with nouns in the query text; Encode video, object, and text information using a pre-trained feature encoder, extract video features, object features, and text features, perform context modeling on the text features, and perform feature interaction between the text and objects; Use a multimodal fusion module that includes a linear time series model using a selective state space and cross-attention to understand and fuse the video feature sequence, and obtain a multimodal feature representation; Use the multimodal feature representation to perform first-person video segment localization.
[0011] As an alternative implementation, the process of extracting object annotations from the first-person video using a pre-trained object detector and filtering out object categories related to the query by matching with nouns in the query text includes: extracting object annotations frame by frame from the first-person video using a pre-trained object detector, and generating structured information including object categories, confidence levels, and their appearance timestamps; Perform natural language processing on the query text, extract noun phrases as key query words, and match the extracted object categories with the query words using semantic similarity calculation; Filter out object categories with a confidence level and a semantic similarity to the query higher than a set threshold as the object annotation for this frame.
[0012] As an alternative implementation, the process of encoding video, object, and text information using a pre-trained feature encoder to extract video features, object features, and text features includes: using a pre-trained video backbone model to extract segment-level video features, and then projecting them into the feature space to obtain a video representation; Use a pre-trained text backbone model to extract text features, and then use a text encoder composed of Transformer to perform context interaction on the text features; Use a pre-trained text backbone model to extract object features, and use an object encoder to refine the object features related to the query.
[0013] As a further step, the process of using an object encoder to refine the object features related to the query includes encoding the text of the detected object categories, using an object encoder composed of multiple layers of Transformer to refine the object features related to the query, using the object features as the query, and the query text features as the key and value to obtain object features related to the query text.
[0014] As an optional implementation, the process of performing video feature sequence understanding and feature fusion using a multimodal fusion module including a linear time series model using a selective state space and a cross-attention includes: enhancing video features using a linear time series model of a bidirectional selective state space to capture long-range dependencies in video data; applying a cross-attention mechanism and a feed-forward layer to aggregate the enhanced video features and query text information; aggregating video features and item features using a parallel cross-attention mechanism; Two features that aggregate different information are combined through a gating mechanism.
[0015] As an optional implementation, the process of locating the first-person video clip using the multimodal feature representation includes: generating a feature pyramid using a multi-scale network composed of multiple layers of Transformers, each layer of Transformer including a 1D deep convolution before the self-attention and FFN modules (Feed Forward Network, FFN) to achieve sequence downsampling and obtain a representation of multi-scale candidate clips; The task head decodes the multi-scale feature pyramid into the final prediction for video localization, the classification head predicts the confidence score of each candidate segment, and the regression head predicts the offset of the candidate segment boundary relative to the anchor point.
[0016] As an optional implementation, the following steps are also included: using contrastive learning to enhance the first-person video feature representation, specifically including: using a pre-trained narrative model to segment the shots; Extract the video features of each segmented shot and the query-level text features; By aggregating the shot- and query-level text features, the text and video features are projected into a joint semantic space for contrastive learning.
[0017] A first-person video positioning system, comprising: An acquisition module, configured to acquire a first-person perspective video and a query text; A preprocessing module, configured to extract object annotations from the first-person video using a pre-trained object detector, and filter out the object categories related to the query by matching them with the nouns in the query text; A feature extraction module is configured to encode video, object and text information using a pre-trained feature encoder, extract video features, object features and text features, perform text feature context modeling, and perform feature interaction between text and objects; A multimodal fusion module is configured to perform video feature sequence understanding and feature fusion using a multimodal fusion module including a linear time series model using a selective state space and cross attention to obtain a multimodal feature representation; The positioning module is configured to use the multimodal feature representation to locate the first-person perspective video clip.
[0018] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps in the above method are completed.
[0019] An electronic device comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps in the above method are completed.
[0020] Compared with the prior art, the present invention has the following beneficial effects: The present invention discloses a video positioning method and system targeting the characteristics of first-person perspective videos, including a first-person perspective video positioning method based on fine-grained object semantics-enhanced video representation and a feature representation enhancement method for shot branches based on contrastive learning, taking into account multiple key factors, including difficulty in video understanding, low video quality, and great positioning difficulty.
[0021] The first-person video localization method based on fine-grained object semantically enhanced video representation of the present invention can locate fine-grained query text; in order to ensure that the model can understand and locate fine-grained query text, by integrating fine-grained object information into the video localization task, the model can obtain information such as the object category in the video frame, thereby significantly improving the localization accuracy of background object-related queries.
[0022] The feature representation enhancement method of shot side branches based on contrastive learning of the present invention can enhance the model's understanding of first-person perspective videos; in order to ensure that the model can better understand the frequent lens movements of the first-person perspective, query-level text features and shot-level video features are extracted from text features and video features, and the alignment capability between text representation and video representation is enhanced through contrastive learning.
[0023] The present invention uses a linear time series model of a bidirectional selective state space (i.e., Mamba network) in a multimodal fusion module to perform self-interaction processing of videos, which can better memorize historical information and thus integrate more contextual information when processing long-sequence video features.
[0024] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0026] Figure 1 This is the flowchart of the video localization method for the characteristics of the first-person view video in the first embodiment of the present invention; Figure 2 This is the flowchart of the process of extracting item annotations in the first embodiment of the present invention; Figure 3 This is the schematic diagram of the multi-modal feature fusion process in the first embodiment of the present invention. Detailed implementation manners
[0027] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0028] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0029] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0030] In the case of no conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0031] Embodiment 1: The first embodiment of the present invention provides a video localization method for the characteristics of the first-person view video, as Figure 1 shown, specifically including the following steps: Step 1, use fine-grained item semantics to enhance the video representation and perform first-person view video localization.
[0032] Step 1.1, as Figure 2 shown, extract fine-grained item information from the video.
[0033] In a specific implementation manner, use an item detector to extract item annotations from the video. Then, use the spaCy library to extract nouns from the natural language query as the items to be queried. Then, also use the spaCy library to calculate the text similarity between the queried items and the item category names. If it is higher than a specific threshold, it is considered that the item category contained in the items queried by the natural language query contains this item category. Finally, filter out the item categories with high confidence in the video frames and contained in the natural language query as the item category information and input it into the model.
[0034] Step 1.2, encode the video, item information and text into features. Figure 1 As shown, the pre-trained base model and the designed encoder extract features from the input video, query text, and item information (i.e., item category information), perform contextual interaction with the text, and refine item features related to the query.
[0035] Step 1.2.1, extract video features using the pre-trained video base model.
[0036] In a specific implementation, suppose the video is The frame composition is expressed as , which is divided into non-overlapping segments, represented as ,in represents the size of the sliding window, Represents the number of segments in the video. The pre-trained video base model is used to extract segment-level video features, which are then projected into the feature space to obtain the video representation as .
[0037] In step 1.2.2, the pre-trained text base model is used to extract text features, and then the text encoder composed of Transformer is used to perform contextual interaction on the text features.
[0038] In a specific implementation, suppose the query text contains Use the pre-trained text base model to extract query text features (which can be referred to as text features for short) and obtain word-level features In order to capture the relationship between words in the query, this embodiment uses a text encoder to encode and output the query representation .
[0039] In step 1.2.3, the pre-trained text base model is used to extract item features, and the item encoder is used to refine the item features relevant to the query.
[0040] In a specific implementation, the text of the detected item category is encoded to obtain , and then use the item encoder composed of multiple layers of Transformer to refine the item features related to the query, where the item features As a query, query text features As keys and values, we finally get item features related to the query text. .
[0041] Step 1.3, using the multimodal fusion module, such as Figure 3 As shown in the figure, multimodal feature representation is obtained by fusing text, object, and video features through a multi-layer structure.
[0042] In step 1.3.1, video features are enhanced using a bidirectional selective state-space linear time series model (Mamba) layer to capture long-range dependencies in video data.
[0043] In a specific embodiment, The video feature output of the layer is ,and . Get the context-enhanced video features The formalization of the process is as follows: ; Step 1.3.2: Apply Cross Attention (CA) and FeedForward Network (FFN) to aggregate context-enhanced video features. and query text information , get the video features after text enhancement , the formalization process is as follows: ; ; Step 1.3.3, similarly use the parallel cross-attention module to aggregate the context-enhanced video features and item features , get the video features after object enhancement , the formalization process is as follows: ; ; Step 1.3.4, two video features with different information enhancement are processed through the gating mechanism Combine and get The output of the layer is also the Layer Input , the formalization process is as follows: ; ; in represents the Sigmoid function, stands for Vector Connection and MLP stands for Multilayer Perceptron.
[0044] In step 1.4, a multi-scale network consisting of multiple layers of Transformer is used to generate a feature pyramid. Each layer contains a 1D depth convolution before the self-attention and FFN modules to achieve sequence downsampling and obtain multi-scale candidate fragment representations.
[0045] In step 1.5, the multi-scale feature pyramid is decoded into the final prediction of video localization by the task head. The classification head predicts the confidence score of each candidate segment, while the regression head predicts the offset of the candidate segment boundary relative to the anchor point.
[0046] Step 1.6, supervised model learning using video clip localization loss, is formalized as follows: ; in, is the focal loss, used to supervise the learning of the classification head, is the distance intersection over union (Distance-IoU) loss, which is used to supervise the learning of the regression head.
[0047] Some embodiments further include step 2, enhancing the first-person perspective video feature representation by using shot branches of contrastive learning.
[0048] Step 2.1, use the pre-trained narrative model to segment the shots.
[0049] In a specific implementation, the pre-trained narrative model LAVILA is first used to generate a narrative for a video, and the generated video captions usually use expressions such as "looks around" or "turns around" to describe the wearer's head movements. These cues reflect the wearer's changes in attention during the video shooting process, segmenting the video into multiple semantically different shots.
[0050] Step 2.2, extract shot-level video features and query-level text features.
[0051] Step 2.2.1: To effectively perform contrastive learning, contextual video features representing the content of each shot that are independent of the query and item information are extracted from the video by using a bidirectional Mamba block integrated with MLP that shares parameters with the multimodal fusion module. .
[0052] In step 2.2.2, the feature aggregator of the Transformer architecture is used to aggregate the shot and query features.
[0053] In one specific implementation, the learnable embedding is used as a query, and the shot context features and query text features Used as keys and values, we then apply a 1D convolution to obtain the final shot representation And text representation , is the number of lenses, is the number of queries.
[0054] Step 2.3, use MLP to represent the lens and text representation Projected into a joint semantic space and used for contrastive learning.
[0055] In a specific implementation, contrastive learning loss is used to supervise the learning of the shot branch, which is formalized as follows: ; in, represents the cosine similarity, is the set of positive query-shot pairs where the video clip corresponding to the query intersects with the corresponding shot, is the temperature coefficient.
[0056] Embodiment 2: Embodiment 2 of the present invention provides a video positioning system targeting the characteristics of first-person perspective video, including: An acquisition module, configured to acquire a first-person perspective video and a query text; A preprocessing module, configured to extract object annotations from the first-person video using a pre-trained object detector, and filter out the object categories related to the query by matching them with the nouns in the query text; A feature extraction module is configured to encode video, object and text information using a pre-trained feature encoder, extract video features, object features and text features, perform text feature context modeling, and perform feature interaction between text and objects; A multimodal fusion module is configured to perform video feature sequence understanding and feature fusion using a multimodal fusion module including a linear time series model using a selective state space and cross attention to obtain a multimodal feature representation; The positioning module is configured to use the multimodal feature representation to locate the first-person perspective video clip.
[0057] The steps involved in the above embodiment 2 correspond to the method embodiment 1. For the specific implementation, please refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0058] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0059] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made by those skilled in the art without creative efforts within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A first-person video positioning method, characterized in that: The following steps are involved: Get the first-person video and query text; Use pre-trained object detectors to extract object annotations from first-person videos and filter out query-related object categories by matching them with nouns in the query text. Use pre-trained feature encoders to encode video, object, and text information, extract video features, object features, and text features, perform text feature context modeling, and perform feature interaction between text and objects; A multimodal fusion module including a linear time series model using a selective state space and cross-attention is used to perform video feature sequence understanding and feature fusion to obtain multimodal feature representation. The multimodal feature representation is used to locate the first-person perspective video clip.
2. A first-view video positioning method as claimed in claim 1, characterized in that: The process of extracting object annotations from the first-person video using a pre-trained object detector and filtering out the object categories related to the query by matching with the nouns in the query text includes: extracting the object annotations frame by frame from the first-person video using the pre-trained object detector to generate structured information including the object category, confidence and its appearance timestamp; By performing natural language processing on the query text, noun phrases are extracted as key query words, and semantic similarity calculation is used to match the extracted item categories with the query words; Item categories with confidence and semantic similarity with the query higher than the set threshold are selected as the item annotations of the frame.
3. A first-view video positioning method as claimed in claim 1, characterized in that: The process of encoding video, object and text information by using a pre-trained feature encoder and extracting video features, object features and text features includes: using a pre-trained video base model to extract segment-level video features, and then projecting the features into a feature space to obtain a video representation; Use the pre-trained text base model to extract text features, and then use the text encoder composed of Transformer to perform contextual interaction on the text features; Utilize the pre-trained text base model to extract item features and use the item encoder to refine the item features relevant to the query.
4. A first-view video positioning method as claimed in claim 3, characterized in that: The process of using an item encoder to refine item features related to a query includes encoding the text of a detected item category, and using an item encoder composed of a multi-layer Transformer to refine item features related to the query, using item features as queries and query text features as keys and values to obtain item features related to the query text.
5. A first-view video positioning method as claimed in claim 1, characterized in that: The process of video feature sequence understanding and feature fusion using a multimodal fusion module including a linear time series model using a selective state space and a cross-attention module includes: enhancing video features using a linear time series model using a bidirectional selective state space to capture long-range dependencies in video data; applying a cross-attention mechanism and a feed-forward layer to aggregate the enhanced video features and query text information; aggregating video features and item features using a parallel cross-attention mechanism; Two features that aggregate different information are combined through a gating mechanism.
6. A first-view video positioning method as claimed in claim 1, characterized in that: The process of locating first-person video clips using the multimodal feature representation includes: generating a feature pyramid using a multi-scale network consisting of multiple layers of Transformers, each layer of Transformers including a 1D deep convolution before the self-attention and FFN modules to achieve sequence downsampling and obtain representations of multi-scale candidate clips; The task head decodes the multi-scale feature pyramid into the final prediction for video localization, the classification head predicts the confidence score of each candidate segment, and the regression head predicts the offset of the candidate segment boundary relative to the anchor point.
7. A first-view video positioning method according to any one of claims 1 to 6, characterized in that: The following steps are also included: Using contrastive learning to enhance the feature representation of first-person video, specifically including: using a pre-trained narrative model to segment the shots; Extract the video features of each segmented shot and the query-level text features; By aggregating the shot- and query-level text features, the text and video features are projected into a joint semantic space for contrastive learning.
8. A first-person video positioning system, characterized in that: include: An acquisition module, configured to acquire a first-person perspective video and a query text; A preprocessing module, configured to extract object annotations from the first-person video using a pre-trained object detector, and filter out the object categories related to the query by matching them with the nouns in the query text; A feature extraction module is configured to encode video, object and text information using a pre-trained feature encoder, extract video features, object features and text features, perform text feature context modeling, and perform feature interaction between text and objects; A multimodal fusion module is configured to perform video feature sequence understanding and feature fusion using a multimodal fusion module including a linear time series model using a selective state space and cross attention to obtain a multimodal feature representation; The positioning module is configured to use the multimodal feature representation to locate the first-person perspective video clip.
9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the steps of the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The method comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the steps in the method according to any one of claims 1 to 7 are completed.
Citation Information
Patent Citations
Method for predicting next interactive object based on transform first view angle
CN114764899A
Time sequence behavior detection method and device, equipment, medium and program product
CN117218572A
First view video description system based on retrieval enhancement
CN119226567A