First-person video question answering method and system based on cross-view semantic alignment
Through cross-perspective semantic alignment and multimodal large language model training strategies, the problems of model differences and data set limitations in first-person video question-and-answer tasks are solved, and the high performance performance and universality of the model in multi-task scenarios are achieved.
Patent Information
- Application Number
- CN202510147029.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The first-person video Q&A task faces limitations in model differences and data set size and quality, resulting in inconsistent performance of the model on multiple tasks and low generalization ability.
Through the cross-view semantic alignment method, the viewpoint alignment data and instruction fine-tuning data are constructed, the cross-viewpoint alignment mapping learning algorithm is designed, and a three-stage multimodal large language model training strategy is adopted to improve the model's understanding of video semantics and its performance ability in multi-task scenarios.
It significantly improves the performance of the model on various video Q&A tasks, enhances the universality of the model and the expansion of application scenarios, and improves the understanding and generalization capabilities of the model.
Smart Images

Figure CN119622024B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a first-person video question-answering method and system based on cross-view semantic alignment, and belongs to the technical field of video question-answering. Background Art
[0002] With the rapid development of portable camera equipment, virtual reality technology and sensor technology, first-person video (FPV) has gradually become an important form of visual expression. In particular, the popularity of equipment such as sports cameras, smart glasses, head-mounted displays (HMDs), etc. has greatly promoted the widespread application of first-person video in multiple fields, including games, education, medical care, sports, security, etc. These innovative devices enable users to record and replay their activities from a personal perspective, thus presenting a more immersive, intuitive and realistic experience than traditional third-person perspective video.
[0003] Video Question Answering (VideoQA) is a core task in the field of video understanding and the basis of many other advanced video understanding tasks. By extracting semantic information from videos and answering natural language questions, Video Question Answering not only verifies the model's ability to fully understand video content, but also provides technical support for other complex tasks (such as behavior prediction, causal reasoning, event generation, etc.). Compared with traditional third-person video question answering tasks, first-person video question answering (FPVQA) faces more technical challenges. This is mainly because first-person videos have more complex dynamic features, including environmental changes, user behaviors, and diverse motion patterns. The changes in these dynamic features not only increase the difficulty of video analysis, but also put forward higher requirements for video processing and model design. Especially when the camera is closely related to the user's line of sight, the first-person video must consider multiple factors such as the dynamic changes in the perspective, the real-time interaction between the user and the environment, and the timing of the user's behavior and actions during processing.
[0004] In order to promote the technical development of FPVQA, existing technologies have constructed multiple high-quality datasets, including EPIC-KITCHENS, Ego4D, Charades-Ego, and Ego-Exo4D. These datasets cover a variety of scenarios from daily life to specific tasks, and provide rich human behavior data. They not only provide a basis for training and verifying video question answering models, but also promote the comparison and improvement of model performance through standardized test sets. In particular, datasets such as Ego4D have added annotation information for video question answering, such as target localization, action recognition, causal reasoning, etc., which makes multi-task learning and cross-task migration possible.
[0005] Although the above research has made significant progress, the first-person video question answering task still faces the following core challenges:
[0006] (1) Model differences caused by task diversity: First-person video question answering involves a wide range of understanding tasks. Different tasks often use different model architectures. Due to the differences between tasks, the correlation between models is poor and the model transferability is low. This makes the performance of the same model on different tasks vary greatly, affecting the universality of the model and the expansion of its application scenarios.
[0007] (2) Limitations in dataset size and quality: Although there are some datasets for different video question answering tasks, such as Ego4D and Charades-Ego, the size of these datasets is usually relatively limited due to the high cost of collecting first-person videos, which restricts the training effect and generalization ability of the model on various tasks. In order to overcome the problem of insufficient data, some studies have tried to use unmatched third-person perspective videos as auxiliary data to enhance the first-person video question answering effect. However, due to the large difference in perspective between the two videos, training data that has not been precisely paired often affects the learning effect of the model, resulting in low robustness and accuracy in actual video question answering applications. Summary of the invention
[0008] In view of the shortcomings of the prior art, the present invention provides a first-person video question-answering method and system based on cross-view semantic alignment, which constructs view alignment data and instruction fine-tuning data through an automated data processing flow, designs a cross-view alignment mapping learning algorithm and a three-stage multimodal large language model (LLM) training strategy, which not only improves the model's ability to understand video semantics, but also significantly enhances the model's performance on various video question-answering tasks through task unified design.
[0009] Terminology explanation:
[0010] Ego4D dataset: The Ego4D dataset is one of the most comprehensive large-scale first-person video datasets in the world. It contains 3,670 hours of high-quality videos, which record in detail the scenes of people performing various activities in their daily lives. These video data were shot by 931 photographers wearing head-mounted cameras or other wearable devices, covering diverse backgrounds from 74 different locations and 9 countries around the world, including home environments, city streets, outdoor natural scenes, workplaces, and entertainment venues, showing a rich and diverse culture, lifestyle, and social environment.
[0011] EgoClip dataset: The Ego4D dataset only provides text annotations at the video moment level, such as "At 3.7 seconds, the photographer put the tomato down." Since this annotation corresponds to one frame of image, it cannot meet the data requirements of video training. Therefore, the researchers expanded the annotation to the video clip level and obtained the EgoClip dataset including 3.8M video clip-text annotation pairs.
[0012] Ego-Exo4D dataset: Currently, it is the largest dataset with both first-person and third-person paired videos. It was shot by 839 participants from 15 universities at home and abroad, with a total duration of more than 1,400 hours. It includes 8 categories such as dance, football, basketball, rock climbing, music, cooking, and bicycle repair, and 131 complex scene actions. Similar to the Ego4D dataset, this dataset only provides text annotations at the video moment level.
[0013] The technical solution of the present invention is as follows:
[0014] The first-person video question answering method based on cross-view semantic alignment has the following steps:
[0015] (1) Collect and screen massive amounts of high-quality third-person video-text pairing data from the Internet for subsequent preliminary training;
[0016] (2) Using the filtered data, we train a third-person video encoder to accurately align video and text features;
[0017] (3) Using the EgoClip dataset, we trained a first-person video encoder to adapt to the semantic characteristics of videos from the first-person perspective.
[0018] (4) Screen and expand the cross-view synchronized data Ego-Exo4D to construct the pre-training dataset Ego-ExoClip;
[0019] (5) Use the third-person video data in Ego-ExoClip to further train the third-person encoder to improve its adaptability to cross-view synchronization data, so that it can better handle complex video scenes and ensure that accurate knowledge is delivered to the first-person video understanding branch in the subsequent process. The training process is the same as steps (2) and (3);
[0020] (6) Construct a cross-view mapping function to learn the transformation relationship between third-person data and first-person data;
[0021] (7) Create an instruction fine-tuning dataset and use it to further train the first-person video encoder to improve the performance of the model in multi-task scenarios;
[0022] (8) Through the cross-view feature fusion strategy, the first-person video features and the estimated third-person features are deeply fused in the multimodal common space to generate a unified multimodal video semantic feature representation;
[0023] (9) Combine the question input by the user with the fused multimodal video features to generate an answer response.
[0024] According to the preferred embodiment of the present invention, in step (1), specifically:
[0025] (11) Establish specific criteria for pairing videos and texts, including video clarity, scene diversity, and accuracy and completeness of text descriptions, and then conduct screening;
[0026] In a specific implementation, high-quality public datasets, such as Panda, WebVid, and VIDAL datasets, can be selected as preliminary data sources.
[0027] (12) Organize the filtered data into a unified structured format, including video files and corresponding text annotations, to lay the foundation for subsequent model training.
[0028] According to the preferred embodiment of the present invention, in step (2), specifically:
[0029] Design the basic model framework and training scheme, use the cascaded video encoder TimeSformer and the large language model Mistral-7B as the model architecture, freeze the parameters of the large language model, optimize the video encoder, use the text generation loss function, and train the model through the screened third-person video-text pairing data to enable it to have accurate semantic expression capabilities.
[0030] Input a third-person video and the question "Please describe the content of the video", decode the predicted text through the large language model, calculate the loss between it and the real text, and update the parameters of the video encoder through the gradient descent method for optimization.
[0031] In step (3), the same model architecture as step (2) is used to train based on the EgoClip dataset, so that the first-person video encoder can obtain basic first-person video understanding capabilities. The input is replaced with the first-person video-text data from the EgoClip dataset. The first-person video and the question "Please describe the content of the video" are input. The predicted text is decoded by the large language model, and the loss between it and the real text is calculated. The parameters of the video encoder (the encoder structure is exactly the same as the third-person video encoder, but it is optimized independently) are updated and optimized by the gradient descent method.
[0032] According to the preferred embodiment of the present invention, in step (4), specifically:
[0033] (41) Screen samples from the Ego-Exo4D dataset, remove videos that cannot be loaded normally, have a resolution lower than 224×224, and have missing or erroneous text annotations to ensure data accuracy and reliability;
[0034] (42) The filtered data must retain all annotated text descriptions to provide a basis for subsequent time segment expansion. The same video is often annotated by multiple annotators at the same time. In order to increase diversity, the text annotations of all the annotators after filtering are retained;
[0035] (43) Extend the narrative annotations corresponding to the timestamps to the video clip level to form the pre-training dataset Ego-ExoClip;
[0036] (44) Combine manual and automatic verification to ensure the quality of annotation and the balance of data distribution.
[0037] Automatic verification is performed for the scene types corresponding to the data set. Different scenes, such as cooking and bicycle repair, have different corresponding numbers of samples. To avoid the long-tail distribution problem, the minimum number of samples in all scenes is used as the benchmark, and samples in each scene are sampled based on this benchmark to ensure that the number of samples remains consistent.
[0038] According to the present invention, further preferably, step (43) is specifically:
[0039] Organize narrative annotations into a sequence of sentences , corresponding to the timestamp , where the sentence S i Describes event i, which occurs at time t i , n represents the total number, for the occurrence at time t i Narrative S i , define the corresponding time segment C i Start time and end time :
[0040] (1)
[0041] Among them, β i represents the average time distance between adjacent sentences in the i-th narrative, that is, , α is the total number of β in the entire Ego-ExoClip dataset i To avoid the description spanning multiple video clips, the preceding and following timestamps are used as boundaries to ensure that each narration corresponds to a single video clip.
[0042] According to the preferred embodiment of the present invention, in step (6), specifically:
[0043] (61) Establish a mapping relationship F between the first-person feature X and the third-person feature Y, thereby learning the semantic knowledge of the third-person data;
[0044] The UNet network is used as the transformation mapping F: X→Y. The encoder part of the U-Net network extracts the first-person features, and the decoder generates a representation consistent with the third-person perspective features through a symmetric path, thereby achieving perspective transformation;
[0045] (62) The inverse transformation mapping G: Y→X is introduced for coupling to strengthen the constraint of F: X→Y. Through the mutual mapping of F and G, the mutual restoration ability of the first-person features and the third-person features is guaranteed during the transformation of the feature space, thereby improving the robustness and accuracy of the mapping.
[0046] According to the present invention, further preferably, based on the data set Ego-ExoClip, the learning of the mapping function is optimized, specifically:
[0047] A. Optimize the learning of the mapping function through supervised loss to improve the mapping accuracy. Apply the cycle consistency loss (CCL), which is defined as:
[0048] (2)
[0049] Among them, x∈X, y∈Y, ‖‖1 represents the L1 norm, and the cycle consistency loss CCL includes forward cycle consistency, i.e., x→F(x)→G(F(x))≈x, and backward cycle consistency, i.e., y→G(y)→F(G(y))≈y, IE x , IE y They represent the corresponding mathematical expectations respectively;
[0050] B. Introduce feature distribution alignment constraints to associate first-person video features with third-person video features to reduce feature differences between perspectives;
[0051] The third-person sample y and the predicted third-person sample The Kullback-Leibler (KL) divergence is introduced to align the true and estimated third-person perspective feature distributions:
[0052] (3)
[0053] Among them, P (y i ) represents the probability distribution of the true third-person sample, Represents the predicted probability distribution, log represents the logarithmic function;
[0054] C. Aggregate the semantic information of the two perspectives and input them into the large language model together with the text, and use the vision-grounded text generation (VTG) loss to optimize the first-person video understanding branch. Specifically;
[0055] A multi-layer perceptron (MLP) is used to aggregate semantic information from two perspectives:
[0056] (4)
[0057] Among them, o is a feature representation that integrates the semantics of the two perspectives. In addition, since the third-person visual encoder has fully learned the corresponding third-person video data, it is frozen to ensure that this branch is not affected. In other words, only the video encoder, transform map F, inverse transform map G, and MLP fusion layer in the first-person video understanding branch are optimized.
[0058] According to the preferred embodiment of the present invention, in step (7), specifically:
[0059] (71) Determine the scope of data selection and use ChatGPT to generate multiple task instructions for each type of data, including task descriptions and execution requirements in different expressions. By increasing the types of task instructions, the diversity of data and the task adaptability of the model are improved;
[0060] (72) Standardize video and text samples into a unified structured format to ensure consistency of question-answering data.
[0061] A first-person video question answering system based on cross-view semantic alignment, comprising:
[0062] The video feature extraction unit is used to collect and filter massive amounts of high-quality third-person video-text pairing data from the Internet, use the filtered data to train the third-person video encoder, use the EgoClip dataset to train the first-person video encoder, adapt to the video semantic characteristics of the first-person perspective, filter and expand the cross-perspective synchronization data Ego-Exo4D, construct the pre-training dataset Ego-ExoClip, and use the third-person video data in Ego-ExoClip to further train the third-person encoder;
[0063] A perspective mapping unit is used to construct a cross-perspective mapping function, learn the transformation relationship between third-person data and first-person data, and create an instruction fine-tuning dataset. Through the instruction fine-tuning dataset, the first-person video encoder is further trained.
[0064] A semantic fusion unit, which is used to deeply fuse the first-person video features with the estimated third-person features in a multimodal common space through a cross-view feature fusion strategy to generate a unified multimodal video semantic feature representation;
[0065] The reply generation unit is used to combine the questions input by the user with the fused multimodal video features, and use the multimodal large language model to generate accurate and targeted answer replies.
[0066] The beneficial effects of the present invention are:
[0067] 1. The present invention uses an automated data annotation process, combined with manual and automatic screening and verification, to construct a high-quality and diverse perspective alignment dataset, which not only reduces the cost of manual annotation, but also significantly improves the accuracy and consistency of the annotated data, providing better data support for model training.
[0068] 2. The cross-view mapping function designed by the present invention can accurately learn the feature transformation relationship between the first-person and third-person data, ensure feature stability by freezing the third-person encoder, and combine the transformation and inverse transformation optimization process to significantly improve the perspective alignment effect and cross-view feature mapping capability.
[0069] 3. The three-stage multimodal model training process proposed in the present invention targets third-person encoder pre-training, first-person encoder adaptation, and instruction fine-tuning optimization, respectively, so that the model gradually has powerful video question-answering capabilities and significantly enhances its adaptability and performance in multi-task scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a flow chart of the first-person video question-answering training method of the present invention;
[0071] Figure 2 is the cross-view mapping transformation function of the present invention;
[0072] Figure 3 Fine-tune the automated annotation process of datasets for the instructions of the present invention. DETAILED DESCRIPTION
[0073] The present invention will be further described below by way of embodiments in conjunction with the accompanying drawings, but is not limited thereto.
[0074] Embodiment 1:
[0075] This embodiment provides a first-person video question answering method based on cross-view semantic alignment, and the steps are as follows:
[0076] (1) Collect and filter a large amount of high-quality third-person video-text pairing data from the Internet for subsequent preliminary training. Specifically:
[0077] (11) Establish specific criteria for pairing videos and texts, including video clarity, scene diversity, and accuracy and completeness of text descriptions, and then conduct screening;
[0078] In a specific implementation, high-quality public datasets, such as Panda, WebVid, and VIDAL datasets, can be selected as preliminary data sources.
[0079] (12) Organize the filtered data into a unified structured format, including video files and corresponding text annotations, to lay the foundation for subsequent model training.
[0080] (2) Using the filtered data, we train a third-person video encoder to accurately align video and text features. Specifically:
[0081] Design the basic model framework and training scheme, use the cascaded video encoder TimeSformer and the large language model Mistral-7B as the model architecture, freeze the parameters of the large language model, optimize the video encoder, use the text generation loss function, and train the model through the screened third-person video-text pairing data to enable it to have accurate semantic expression capabilities.
[0082] like Figure 1 As shown in stage 1 in the middle left figure, a third-person video and the question "Please describe the content of the video" are input, the predicted text is decoded by the large language model, and the loss between it and the real text is calculated. The parameters of the video encoder are updated and optimized through the gradient descent method.
[0083] (3) Use the EgoClip dataset to train a first-person video encoder to adapt to the semantic characteristics of videos from the first-person perspective; follow the same model architecture as step (2) and train based on the EgoClip dataset to enable the first-person video encoder to acquire basic first-person video understanding capabilities. Replace the input with the first-person video-text data from the EgoClip dataset, input the first-person video and the question "Please describe the content of the video", decode the predicted text through the large language model, and calculate the loss between it and the real text. Update the parameters of the video encoder (the encoder structure is exactly the same as the third-person video encoder, but is optimized independently) through the gradient descent method for optimization.
[0084] (4) Screen and expand the cross-view synchronization data Ego-Exo4D to construct the pre-training dataset Ego-ExoClip. Specifically:
[0085] (41) Screen samples from the Ego-Exo4D dataset, remove videos that cannot be loaded normally, have a resolution lower than 224×224, and have missing or erroneous text annotations to ensure data accuracy and reliability;
[0086] (42) The filtered data must retain all annotated text descriptions to provide a basis for subsequent time segment expansion. The same video is often annotated by multiple annotators at the same time. In order to increase diversity, the text annotations of all the annotators after filtering are retained;
[0087] (43) Extend the narrative annotations corresponding to the timestamps to the video clip level to form the pre-training dataset Ego-ExoClip;
[0088] Organize narrative annotations into a sequence of sentences , corresponding to the timestamp , where sentence S i Describes event i, which occurs at time t i , n represents the total number, for the occurrence at time t i Narrative S i , define the corresponding time segment C i Start time and end time :
[0089] (1)
[0090] Among them, β i represents the average time distance between adjacent sentences in the i-th narrative, that is, , α is the total number of β in the entire Ego-ExoClip dataset i To avoid the description spanning multiple video clips, the preceding and following timestamps are used as boundaries to ensure that each narration corresponds to a single video clip.
[0091] (44) Combine manual and automatic verification to ensure the quality of annotation and the balance of data distribution.
[0092] Automatic verification is performed for the scene types corresponding to the data set. Different scenes, such as cooking and bicycle repair, have different corresponding numbers of samples. To avoid the long-tail distribution problem, the minimum number of samples in all scenes is used as the benchmark, and samples in each scene are sampled based on this benchmark to ensure that the number of samples remains consistent.
[0093] (5) Use the third-person video data in Ego-ExoClip to further train the third-person encoder to improve its adaptability to cross-view synchronization data, so that it can better handle complex video scenes and ensure that accurate knowledge is delivered to the first-person video understanding branch in the subsequent process. The training process is the same as steps (2) and (3);
[0094] (6) Construct a cross-view mapping function to learn the transformation relationship between third-person data and first-person data;
[0095] (61) Establish a mapping relationship F between the first-person feature X and the third-person feature Y, thereby learning the semantic knowledge of the third-person data;
[0096] Using UNet network, such as Figure 2 As shown, as the transformation mapping F: X→Y, the encoder part of the U-Net network extracts the first-person features, and the decoder generates a representation consistent with the third-person perspective features through a symmetric path, thereby realizing perspective transformation;
[0097] (62) The inverse transformation mapping G: Y→X is introduced for coupling to strengthen the constraint of F: X→Y. Through the mutual mapping of F and G, the mutual restoration ability of the first-person features and the third-person features is guaranteed during the transformation of the feature space, thereby improving the robustness and accuracy of the mapping.
[0098] Based on the dataset Ego-ExoClip, we optimize the learning of mapping functions. Specifically:
[0099] A. Optimize the learning of the mapping function through supervised loss to improve the mapping accuracy. Apply the cycle consistency loss (CCL), which is defined as:
[0100] (2)
[0101] Among them, x∈X, y∈Y, ‖‖1 represents the L1 norm, and the cycle consistency loss CCL includes forward cycle consistency, i.e., x→F(x)→G(F(x))≈x, and backward cycle consistency, i.e., y→G(y)→F(G(y))≈y, IE x , IE y They represent the corresponding mathematical expectations respectively;
[0102] B. Introduce feature distribution alignment constraints to associate first-person video features with third-person video features to reduce feature differences between perspectives;
[0103] The third-person sample y and the predicted third-person sample The Kullback-Leibler (KL) divergence is introduced to align the true and estimated third-person perspective feature distributions:
[0104] (3)
[0105] Among them, P (y i ) represents the probability distribution of the true third-person sample, Represents the predicted probability distribution, log represents the logarithmic function;
[0106] C. Aggregate the semantic information of the two perspectives and input them into the large language model together with the text, and use the vision-grounded text generation (VTG) loss to optimize the first-person video understanding branch. Specifically;
[0107] A multi-layer perceptron (MLP) is used to aggregate semantic information from two perspectives:
[0108] (4)
[0109] Among them, o is a feature representation that integrates the semantics of the two perspectives. In addition, since the third-person visual encoder has fully learned the corresponding third-person video data, it is frozen to ensure that this branch is not affected. In other words, only the video encoder, transform map F, inverse transform map G, and MLP fusion layer in the first-person video understanding branch are optimized.
[0110] (7) Create an instruction fine-tuning dataset and use it to further train the first-person video encoder to improve the performance of the model in multi-task scenarios;
[0111] (71) Determine the scope of data selection and use ChatGPT to generate multiple task instructions for each type of data, including task descriptions and execution requirements in different expressions. By increasing the types of task instructions, the diversity of data and the task adaptability of the model are improved;
[0112] In a specific implementation manner, Figure 3 As shown in the figure, each type of data consists of the following three parts: (1) Dataset description: a detailed description of the source, content and characteristics of the dataset, such as video type, scene range and annotation information; (2) Task description: a description of the main task objectives for each type of data, such as "identify the main actions in the video and generate a narrative" or "analyze the video content and answer related questions"; (3) Instruction example: an example of the instruction to be generated, providing more accurate template information such as the format.
[0113] (72) Standardize video and text samples into a unified structured format to ensure consistency of question-answering data.
[0114] In a specific implementation, each sample contains the following four parts: (1) Video path: clearly specify the storage location or unique identifier of the video file to facilitate the model to access the video content; (2) Task instructions: clearly describe the specific task that the model needs to complete, such as "describe the behavior of the characters in the video" or "answer questions based on the video content"; (3) Question: specify the execution details of the task instruction, such as "What did the main character in the video do?"; (4) Answer: provide a standard answer generated based on the video content, which is used to supervise the correctness of the model output during training.
[0115] (8) Through the cross-view feature fusion strategy, the first-person video features and the estimated third-person features are deeply fused in the multimodal common space to generate a unified multimodal video semantic feature representation;
[0116] (9) Combine the question input by the user with the fused multimodal video features to generate an answer response.
[0117] Embodiment 2:
[0118] This embodiment provides a first-person video question answering system based on cross-view semantic alignment, including:
[0119] The video feature extraction unit is used to collect and filter massive amounts of high-quality third-person video-text pairing data from the Internet, use the filtered data to train the third-person video encoder, use the EgoClip dataset to train the first-person video encoder, adapt to the video semantic characteristics of the first-person perspective, filter and expand the cross-perspective synchronization data Ego-Exo4D, construct the pre-training dataset Ego-ExoClip, and use the third-person video data in Ego-ExoClip to further train the third-person encoder;
[0120] A perspective mapping unit is used to construct a cross-perspective mapping function, learn the transformation relationship between third-person data and first-person data, and create an instruction fine-tuning dataset. Through the instruction fine-tuning dataset, the first-person video encoder is further trained.
[0121] A semantic fusion unit, which is used to deeply fuse the first-person video features with the estimated third-person features in a multimodal common space through a cross-view feature fusion strategy to generate a unified multimodal video semantic feature representation;
[0122] The reply generation unit is used to combine the questions input by the user with the fused multimodal video features, and use the multimodal large language model to generate accurate and targeted answer replies.
[0123] It is understandable that the above-mentioned units can be separately or completely combined into one or several other units to constitute, or some of the units can be further divided into multiple functionally smaller units to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In practical applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the video behavior segment candidate set generation system can also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0124] The system described in this embodiment and the method for generating a candidate set of video behavior segments of the embodiment of the present application can be constructed by running a computer program (including program code) capable of executing the steps involved in the corresponding method described in the embodiment on a general computing device such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read only memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.
[0125] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made by those skilled in the art within the spirit and principle of the present invention without creative labor shall be included in the protection scope of the present invention.
Claims
1. A first-person video question answering method based on cross-view semantic alignment, characterized in that: Here are the steps: (1) Collect and filter massive amounts of high-quality third-person video-text pairing data from the Internet; (2) Using the filtered data, train the third-person video encoder; (3) Using the EgoClip dataset, we trained a first-person video encoder to adapt to the semantic characteristics of videos from the first-person perspective. (4) Screen and expand the cross-view synchronized data Ego-Exo4D to construct the pre-training dataset Ego-ExoClip; (5) Use the third-person video data in Ego-ExoClip to further train the third-person encoder; (6) Construct a cross-view mapping function to learn the transformation relationship between third-person data and first-person data; (7) Create an instruction fine-tuning dataset and further train the first-person video encoder using the instruction fine-tuning dataset; (8) Through the cross-view feature fusion strategy, the first-person video features and the estimated third-person features are deeply fused in the multimodal common space to generate a unified multimodal video semantic feature representation; (9) Combine the question input by the user with the fused multimodal video features to generate an answer response.
2. The first-person video question answering method based on cross-view semantic alignment as claimed in claim 1, characterized in that: In step (1), specifically: (11) Establish specific criteria for pairing videos and texts, including video clarity, scene diversity, and accuracy and completeness of text descriptions, and then conduct screening; (12) Organize the filtered data into a unified structured format, including video files and corresponding text annotations.
3. The first-person video question answering method based on cross-view semantic alignment as claimed in claim 2, characterized in that: In step (2), specifically: Design the basic model framework and training scheme, use the cascaded video encoder TimeSformer and the large language model Mistral-7B as the model architecture, use the text generation loss function, and train the model through the screened third-person video-text pairing data to enable it to have accurate semantic expression capabilities.
4. The first-person video question answering method based on cross-view semantic alignment as claimed in claim 3, characterized in that: In step (4), specifically: (41) Samples were screened from the Ego-Exo4D dataset, and videos that could not be loaded normally, had a resolution lower than 224 × 224, and text annotations with missing or incorrect annotations were removed; (42) The screened data must retain all annotated text descriptions; (43) Extend the narrative annotations corresponding to the timestamps to the video clip level to form the pre-training dataset Ego-ExoClip; (44) Combine manual and automatic verification to ensure the quality of annotation and the balance of data distribution.
5. The first-person video question answering method based on cross-view semantic alignment as claimed in claim 4, characterized in that: Step (43) is specifically as follows: Organize narrative annotations into a sequence of sentences , corresponding to the timestamp , where the sentence S i Describes event i, which occurs at time t i , n represents the total number, for the occurrence at time t i Narrative S i , define the corresponding time segment C i Start time and end time : (1) Among them, β i represents the average time distance between adjacent sentences in the i-th narrative, that is, , α is the total number of β in the entire Ego-ExoClip dataset i To avoid the description spanning multiple video clips, the preceding and following timestamps are used as boundaries to ensure that each narration corresponds to a single video clip.
6. The first-person video question answering method based on cross-view semantic alignment as claimed in claim 5, characterized in that: In step (6), specifically: (61) Establish a mapping relationship F between the first-person feature X and the third-person feature Y, thereby learning the semantic knowledge of the third-person data; The UNet network is used as the transformation mapping F: X→Y. The encoder part of the U-Net network extracts the first-person features, and the decoder generates a representation consistent with the third-person perspective features through a symmetric path, thereby achieving perspective transformation; (62) The inverse transformation mapping G: Y→X is introduced for coupling to strengthen the constraint of F: X→Y. Through the mutual mapping of F and G, the mutual restoration ability of the first-person features and the third-person features is guaranteed during the transformation of the feature space.
7. The first-person video question answering method based on cross-view semantic alignment as claimed in claim 6, characterized in that: Based on the dataset Ego-ExoClip, we optimize the learning of mapping functions. Specifically: A. Optimize the learning of the mapping function through supervised loss to improve the mapping accuracy. Apply cycle consistency loss, defined as: (2) Among them, x∈X, y∈Y, ‖‖1 represents the L1 norm, and the cycle consistency loss CCL includes forward cycle consistency, i.e., x→F(x)→G(F(x))≈x, and backward cycle consistency, i.e., y→G(y)→F(G(y))≈y, IE x , IE y They represent the corresponding mathematical expectations respectively; B. Introduce feature distribution alignment constraints to associate first-person video features with third-person video features to reduce feature differences between perspectives; The third-person sample y and the predicted third-person sample The Kullback-Leibler (KL) divergence is introduced to align the true and estimated third-person perspective feature distributions: (3) Among them, P (y i ) represents the probability distribution of the true third-person sample, Represents the predicted probability distribution, log represents the logarithmic function; C. Aggregate the semantic information of the two perspectives and input them into the large language model together with the text, and use the text generation loss to optimize the first-person video understanding branch. Specifically; A multi-layer perceptron is used to aggregate semantic information from two perspectives: (4) Among them, o is a feature representation that integrates the semantics of two perspectives.
8. The first-person video question answering method based on cross-view semantic alignment as claimed in claim 7, characterized in that: In step (7), specifically: (71) Determine the scope of data selection and use ChatGPT to generate multiple task instructions for each type of data, including task descriptions and execution requirements in different expressions; (72) Standardize video and text samples into a unified structured format to ensure consistency of question-answering data.
9. A first-person video question answering system based on cross-view semantic alignment, characterized in that: include: The video feature extraction unit is used to collect and filter massive amounts of high-quality third-person video-text pairing data from the Internet, use the filtered data to train the third-person video encoder, use the EgoClip dataset to train the first-person video encoder, adapt to the video semantic characteristics of the first-person perspective, filter and expand the cross-perspective synchronization data Ego-Exo4D, construct the pre-training dataset Ego-ExoClip, and use the third-person video data in Ego-ExoClip to further train the third-person encoder; A perspective mapping unit is used to construct a cross-perspective mapping function, learn the transformation relationship between third-person data and first-person data, and create an instruction fine-tuning dataset. Through the instruction fine-tuning dataset, the first-person video encoder is further trained. A semantic fusion unit, which is used to deeply fuse the first-person video features with the estimated third-person features in a multimodal common space through a cross-view feature fusion strategy to generate a unified multimodal video semantic feature representation; The reply generation unit is used to combine the question input by the user with the fused multimodal video features to generate an answer reply.
Citation Information
Patent Citations
Pedestrian trajectory prediction method and device based on first person view angle video
CN114581488A
First-person view angle fixation point prediction method based on multi-modal deep learning
CN118821047A