Method for obtaining video-based model based on unmasked alignment
By using the non-mask alignment method in the video base model and using the CLIP model for linear projection alignment, the problem of reducing knowledge generalization of the video base model in the video field and difficulty in learning time-related behaviors is solved, and the rapid convergence and good scalability of the video base model are achieved.
Patent Information
- Application Number
- CN202310315252.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-03-27
AI Technical Summary
When converting from image-based models to video fields, existing video-based models face problems such as reduced knowledge generalization, difficulty in learning time-related behaviors, and poor scalability.
A video base model acquisition method based on non-mask alignment is proposed. By masking the image frame of the original video, a non-mask video image block is generated, and a linear projection alignment is used as a teacher model and a student model is performed to optimize the mean square error to train the video base model.
It realizes the fast convergence and good generalization and scalability of the video-based model, can effectively handle cross-modal video tasks, and reduces training costs.
Smart Images

Figure CN116310995B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision, and particularly to a method for obtaining a video-based model based on unmasked alignment. Background Art
[0002] Due to the high computational cost and data scarcity, the exploration of video-based models is very limited. Previous video-based models rely on image-based models, but transferring to the video domain faces formidable challenges. Building a video-based model on a well-pre-trained image-based model can significantly reduce the training cost, but it faces an important challenge of transferring knowledge from the image domain to the video domain. First, due to limited video data and significant domain differences, secondary video pre-training may weaken the generalization inherited from the image-based model. In addition, a strong spatial initialization provides a simple strategy for perceiving scenes in videos (e.g., "grassland" in "riding a horse"), which limits the video-based model's ability to learn to recognize and localize time-related behaviors (such as "open" and "close"). Finally, this mode is difficult to scale because it requires a well-pre-trained image feature model.
[0003] Recently, VideoMAE has successfully learned effective spatio-temporal features from scratch and can effectively handle complex time behavior recognition and detection tasks. Although VideoMAE can train a powerful ViT from limited data, this is achieved through a long pre-training for high data efficiency. For example, it has 2400 iterations on 160k videos. Moreover, its low-level reconstruction task converges with difficulty and conflicts with the high-level cross-modal alignment task. It is not applicable to video language tasks because the low-level pixel reconstruction task conflicts with the high-level cross-modal alignment task. Furthermore, the additional decoder processes masked and unmasked visual tokens, which leads to excessive memory overhead due to the global self-attention mechanism, making the expansion of this mode also challenging. Summary of the Invention
[0004] To solve the above problems existing in the prior art, the object of the present invention is to propose a method for obtaining a video-based model based on unmasked alignment. The obtained video-based model is further trained in combination with other modal models, and the obtained final model can achieve cross-modal video task processing.
[0005] To achieve the above object, the technical solution of the present invention is as follows.
[0006] In a first aspect, the present invention proposes a method for obtaining a video-based model based on unmasked alignment, and the method includes the following steps:
[0007] Take the set of video image patches of the original video used as training data as the first dataset, and use a masking strategy to occlude the image frames of the original video to obtain a set of unmasked video image patches (patches) as the second dataset;
[0008] Use the visual encoder of the CLIP (Contrastive Language-Image Pre-Training) model as the teacher model, and use an untrained visual encoder as the student model;
[0009] During training, input the first dataset into the teacher model, input the second dataset into the student model, select the corresponding unmasked video image patch outputs of the two models for linear projection alignment, calculate the mean square error between the two after normalization, and continuously optimize to reduce the mean square error;
[0010] Use the trained student model as the video base model, combine it with other modality models for further training, and the obtained final model can achieve cross-modal video task processing.
[0011] In one implementation of the above technical solution, the masking strategy is to sequentially perform masking on the image frames to be masked using semantic guidance, and the masking ratio is 80%.
[0012] In one implementation of the above technical solution, the sequential use of semantic guidance to perform masking on the image frames to be masked is specifically as follows:
[0013] In the last self-attention layer of the teacher model, obtain the class token z of each frame cls ∈R 1×C and the spatial tokens Z ∈ R L×C , where L = H × W is the number of tokens, C is the dimension of the token, and H and W are the height and width of the frame;
[0014] Calculate the attention score A ∈ R 1×L , which is used to represent the semantic importance of each token:
[0015]
[0016]
[0017] In the formula: N is the number of attention heads, Q n (·) and K n (·) are the linear projections of the nth attention head, and softmax is the normalized exponential function;
[0018] Select non-masked tokens in a frame of image based on A, and mask those not selected.
[0019] In an implementation of the above technical solution, the image frame to be masked is obtained from the original video image through sparse sampling.
[0020] In an implementation of the above technical solution, linear projection alignment is specifically implemented as follows:
[0021] In the teacher model, use the pre-trained mapping layer to establish the semantic connection between vision and text editing;
[0022] In the student model, use the mapping layer to align the channel dimensions;
[0023] Align each output token in the student model with the relevant output token in the teacher model.
[0024] In an implementation of the above technical solution, the method further includes configuring a text encoder, a cross-modal decoder for the trained vision encoder to form a cross-modal task model;
[0025] The output ends of the text encoder and the trained vision encoder are connected to the cross-modal decoder;
[0026] The functions that the cross-modal decoder can achieve according to its input include: video action recognition and action detection, video retrieval, video question answering.
[0027] In an implementation of the above technical solution, when training the cross-modal task model, the training data includes video-text data and image-text data. The training objectives include, in addition to aligning the non-masked output tokens of the teacher model and the student model, three training objectives: video-text contrast, video-text matching, and video-guided masked language modeling;
[0028] The video-text contrast is to align the non-masked video and text embeddings;
[0029] The video-text matching is to fuse and classify the non-masked video tokens and text tokens;
[0030] The video-guided masked language modeling is to use the cross-modal decoder to predict the masked text from the text and non-masked video tokens.
[0031] In one implementation of the above technical solution, after pre-training in different stages, it further includes fine-tuning a variety of downstream task models; the fine-tuning is to selectively adjust the parameters of the visual encoder, text encoder, and cross-modal decoder using downstream task data according to the expected functions after completing the pre-training. Specifically, for pure video tasks, fine-tune the initially trained visual encoder; for video retrieval tasks, fine-tune the visual encoder and text encoder; for video question answering tasks, fine-tune the visual encoder, text encoder, and cross-modal decoder.
[0032] In a second aspect, the present invention proposes a cloud server on which a computer program capable of executing any of the above methods is deployed. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0034] Figure 1 、 one Schematic diagrams of the teacher model and student model structures in a specific implementation manner;
[0035] Figure 2 、 one Schematic diagram of the progressive video base model pre-training framework in a specific implementation manner;
[0036] Figure 3 、 one Schematic diagram of the application in a specific implementation manner. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] Existing video base models are mainly developed based on well-pre-trained image base models. Building a video base model on a well-pre-trained image base model can greatly reduce the training cost, but it faces important challenges in transferring knowledge from the image domain to the video domain. First, due to limited video data and significant domain differences, secondary video pre-training may weaken the generalization inherited from the image base model. In addition, the powerful spatial initialization provides a simple strategy for perceiving the scenes in the video (e.g., "grassland" in "riding a horse"), which limits the video base model's ability to learn to recognize and localize time-related behaviors (such as "open" and "close"). Finally, this mode is difficult to scale because it requires a well-pre-trained image feature model.
[0038] To solve this problem, this case proposes an efficient video-based model training method that uses masked video pre-training to make the model time-sensitive. The video-based model obtained based on this not only converges quickly but also has good generalization and scalability. By masking most of the low-semantic video tokens, only the video tokens retained in the student model are aligned with the corresponding video tokens in the teacher model, where the image-based model acts as the unmasked teacher. Through the semantic guidance provided by the teacher, faster convergence and multi-modal friendliness are achieved. Further, through a progressive pre-training framework, models for handling various video-related tasks can be obtained, including scene-related, time-related, and complex video language understanding, etc. The model obtained using the method of this case is denoted as ViT-L / 16. It was pre-trained on 32 A100 GPUs for 6 days using only publicly available data and achieved state-of-the-art performance levels on various video tasks. The video tokens in this case are one or more image patches that make up a frame of an image.
[0039] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0040] The terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The method steps involved in the following process may not be implemented in sequence. On the contrary, the method steps may be implemented in reverse order or simultaneously. In addition, one or more other steps may be added to the method steps. Further improvements may also be made based on the method of this case by removing one or more steps to optimize the method implementation.
[0041] (I) Core
[0042] The following steps are used to obtain an unmasked video-based model. The method includes the following steps:
[0043] Use the set of video image patches of the original video as the training data as the first data set, and use a masking strategy to mask the image frames of the original video to obtain a set of unmasked video image patches as the second data set;
[0044] Use the CLIP (Contrastive Language-Image Pre-Training) model as the teacher model and the vision encoder as the student model;
[0045] During training, the first dataset is input into the teacher model, and the second dataset is input into the student model. The corresponding unmasked video image patches output by the two models are selected for linear projection alignment, the mean square error between the two after normalization is calculated, and the mean square error is continuously optimized and reduced. The determination to stop training can be to meet the set threshold or to meet the set number of iterations.
[0046] The trained student model is used as the video base model, which is combined with other modality models for further training, and the obtained final model can achieve cross-modal video task processing.
[0047] As Figure 1 shown, different from the existing method of secondary development based on the image base model, the visual encoder ViT (CLIP-ViT) of the image base model CLIP is used as the unmasked teacher model, and ViT is used as the student model, and a simple ViT model is trained from scratch using the teacher model. The ViT model, full name Vision Transformer, is an image classification model based on the self-attention mechanism (Self-Attention). The CLIP-ViT model can make full use of the rich semantic information learned under the guidance of natural language, which is beneficial to subsequent multi-modal learning. For the student model, specifically, a ViT model based on average pooling is adopted.
[0048] During training, by occluding most of the low-semantic video units, only the unoccluded units are linearly projected and aligned with the corresponding output units of the teacher. This method not only inherits the high data utilization rate of VideoMAE, but also makes the learned video encoder friendly to multi-modal. Moreover, compared with VideoMAE, only the unmasked units are used for training, that is: in this case, the student model only inputs the image patches not masked by the mask, and there is no additional decoder in the student model during the training process, which greatly saves the GPU video memory. In addition, the guidance of the teacher representation rich in semantics makes the student model converge faster. The obtained ViT model can not only handle actions related to the scene and actions related to time well, but also make the model compatible with subsequent cross-modal learning.
[0049] In the specific implementation of the above technical solution, for the masking strategy: First, considering that an overly aggressive random occlusion scheme may only retain background units, and the meaningless information contained in these units may hinder the knowledge distillation of the teacher, semantic guidance is sequentially adopted for the image frames to be masked; Second, considering the redundancy of the video and adopting a relatively high masking ratio, the masking ratio is set to 80%. Among them, the sequential adoption of semantic guidance for the image frames to be masked is to use the similarity matrix generated in the last self-attention layer of the teacher model, and use the attention scores of each spatial unit by the classification unit as weights to generate the retention probabilities of different units through polynomial distribution, so as to retain the salient objects in each frame with a higher probability. The specific calculation is as follows:
[0050] (1) In the last self-attention layer of the teacher model, obtain the class token z of each frame cls ∈R 1×C and the spatial tokens Z ∈ R L×C , where L = H × W is the number of tokens, C is the dimension of the tokens, and H and W are the height and width of the frame.
[0051] (2) Calculate the attention score A ∈ R 1×L , which is used to represent the semantic importance of each token, so as to occlude the video units with low semantics and improve the data usage efficiency:
[0052]
[0053]
[0054] In the formula: N is the number of attention heads, Q n (·) and K n (·) are the linear projections of the nth attention head, and softmax is the normalized exponential function. By using the spatio-temporal attention mechanism, sufficient interaction between non-masked units can be promoted. And when using the spatio-temporal attention mechanism, no temporal downsampling operation is adopted to ensure that the student model and the teacher model can be frame-aligned, that is: each image patch input to the student model must be in the set of image patches input to the teacher model.
[0055] (3) Based on A, select the non-masked block tokens in a frame of image, and mask the unselected ones.
[0056] For the teacher model, all image patches are input, while for the student model, only the unmasked image patches are input. As a further improvement, the image frames to be masked are obtained from the original video images through sparse sampling, providing more complex action contexts with larger frame intervals, which prompts the model to learn longer-term object spatio-temporal relationships. To more fully distill the semantic information of the teacher model, linear projection alignment is adopted, and the specific implementation is as follows:
[0057] In the teacher model, a pre-trained mapping layer is used to establish the semantic connection between vision and text editing;
[0058] In the student model, a mapping layer is used to align the channel dimensions;
[0059] Each output token in the student model is aligned with the relevant output token in the teacher model, and with the semantic guidance provided by the teacher model, faster convergence of the student model is achieved.
[0060] Through the above implementation, it can be seen that in this case, only video data is used for masked video modeling, and a video-based model capable of completing pure video tasks can be obtained.
[0061] (2) Obtaining a cross-modal task model through progressive training
[0062] Next, the above-trained video-based model can be used in combination with other modal models to form a cross-modal task model. Through continuous training and learning, the cross-modal task model can handle more complex video-related cross-modal tasks. Taking the combination of the video-based model and the text model as an example, common vision-language data is used for multi-modal learning, so that the newly formed cross-modal task model can handle complex video language tasks such as video retrieval and video question answering, etc.
[0063] Figure 2 Schematically shows a progressive video-based model pre-training framework for obtaining a cross-modal task model, which is divided into two stages.
[0064] In the first stage, the trained student model obtained in the above implementation is a visual encoder.
[0065] In the second stage, the pre-trained visual encoder from the first stage is used, and an open-source language model is introduced as the text encoder and cross-modal decoder to form a cross-modal task model. The output ends of the text encoder and the trained visual encoder are connected to the cross-modal decoder. The functions that the cross-modal decoder can achieve according to its input include: video action recognition and action detection, video retrieval, and video question answering. Multimodal learning is carried out using public visual-text data, enabling the model to handle complex video language tasks such as video retrieval and video question answering, etc. Considering the scarcity of video-text data, image-text data is introduced for joint training. The training objectives include, in addition to aligning the non-masked output tokens of the teacher model and the student model, three training objectives: video-text contrast, video-text matching, and video-guided masked language modeling. The video-text contrast is to align the non-masked video and text embeddings; the video-text matching is achieved by fusing the non-masked video tokens and text tokens; the video-guided masked language modeling uses the cross-modal decoder to predict the masked text from the text and non-masked video tokens.
[0066] The above training in the second stage is a type of pre-training. In order to enable the obtained cross-modal task model to efficiently complete different video tasks in a targeted manner, the present case also fine-tunes the obtained cross-modal task model.
[0067] In an implementation manner of the above technical solution, after pre-training, it further includes fine-tuning the cross-modal task model; the fine-tuning is to selectively adjust the parameters of the visual encoder, text encoder, and cross-modal decoder using downstream task data according to the expected functions after completing the pre-training. Specifically, for pure video tasks, the visual encoder trained in the first stage is fine-tuned. For video retrieval tasks, the visual encoder and text encoder are fine-tuned; for video question answering tasks, the visual encoder, text encoder, and cross-modal decoder are fine-tuned.
[0068] Through the description of the above implementation manners, those skilled in the art can clearly understand that the present disclosure can be implemented by means of software plus necessary general-purpose hardware, and of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, in more cases for the present disclosure, software program implementation is a better implementation manner.
[0069] In another embodiment, a computer program capable of being executed, such as any of the above methods, is deployed on a cloud server. Users on different terminals can upload videos to the cloud server through an application and select the desired functions, such as behavior classification, behavior detection, and behavior question answering. The cloud server will call the corresponding deployed algorithm and return the corresponding results after the inference is completed, such as Figure 3 as shown
[0070] In one embodiment, the method proposed in this case was verified on popular pure video tasks and video-language task datasets, and state-of-the-art performance was achieved in all of them. Some of the results are listed below, including scene-related behavior recognition (Kinetics) (see Table 1), temporal-related behavior recognition (Something-Something) (see Table 2), temporal detection (AVAv2.2) (see Table 3), video retrieval (MSTVTT, SSV2-label) (see Table 4), and video question answering (MSRVTT-QA, MSVD-QA) (see Table 5).
[0071] Table 1
[0072]
[0073] Table 2
[0074]
[0075] Table 3
[0076] method number of frames computational amount AVA v2.2 MviTv2-L 40 2.8 33.5 VideoMAE-L 16 0.6 39.3 Ours-L 8 0.6 39.8
[0077] Table 4
[0078] method MSRVTT SSV2-label CLIP4Clip 44.5 42.8 Singularity 42.7 47.4 VINDLU 46.5 53.1 Ours-L 58.8 70.4
[0079] Table 5
[0080] method MSRVTT-QA MSVD-QA All-in-one 44.3 47.9 Violet 43.9 47.9 OmniVL 44.1 51.0 Ours-L 47.9 55.2
[0081] In summary, the method for training and obtaining a video base model proposed in this case greatly reduces the training cost compared with the video base model developed based on an image base model. On the basis of the obtained video base model, by further combining the training of other modality models, a cross-modal task model capable of simultaneously processing scene-related and temporal-related behaviors can be obtained, and the obtained cross-modal task model is excellent in the behavior detection task. Compared with VideoMAE based on masked modeling, the method for obtaining a video base model proposed in this case has a better convergence speed and better results, is friendly to multi-modal learning, and the model performs excellently in various video-language evaluation criteria.
[0082] Although the embodiments of the present invention have been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present invention, and all of these fall within the scope of protection of the present invention.
Claims
1. Method for obtaining video-based model based on non-masked alignment Characterized in that The method includes the following steps Taking the set of video image blocks of the original video as training data as the first data set, and masking the image frames of the original video using a masking strategy to obtain a set of non-masked video image blocks as the second data set Taking the visual encoder of the CLIP model as the teacher model, and taking the untrained visual encoder as the student model During training, inputting the first data set into the teacher model, inputting the second data set into the student model, selecting the corresponding non-masked video image block outputs of the two models for linear projection alignment, calculating the mean square error after normalization between the two, and continuously optimizing to reduce the mean square error Taking the trained student model as the video-based model, and combining it with other modality models for further training, and the obtained final model can achieve cross-modal video task processing The masking strategy is to mask the image frames to be masked successively using semantic guidance, and the masking ratio is 80% The successive masking of the image frames to be masked using semantic guidance is specifically In the self-attention of the last layer of the teacher model, obtain the class tokens for each frame and the spatial tokens , where is the number of tokens, and are the height and width of the frame; Calculate the attention score using the following formula , which is used to represent the semantic importance of each token: Wherein: is the number of attention heads; is the nth function, and are its parameters; and are the linear projections of the nth attention head, is the normalized exponential function; T is matrix transpose Based on Select the non-mask block markers in a frame of image and mask those not selected.
2. The method according to claim 1 Characterized in that The image frames to be masked are obtained from the original video images through sparse sampling 3. The method according to claim 1 Characterized in that The linear projection alignment is specifically implemented as Using a pre-trained mapping layer in the teacher model to establish a semantic connection between vision and text editing Using a mapping layer in the student model to align the channel dimensions Aligning each output token in the student model with the corresponding output token in the teacher model 4. The method according to claim 1 Characterized in that The method further includes configuring a text encoder and a cross-modal decoder for the trained visual encoder to form a cross-modal task model The output ends of the text encoder and the trained visual encoder are connected to the cross-modal decoder The functions that the cross-modal decoder can achieve according to its input include: video action recognition and action detection, video retrieval, video question answering 5. The method according to claim 4 Characterized in that During the training of the cross-modal task model, the training data includes video-text data and image-text data, and the training objectives include, in addition to aligning the non-masked output tokens of the teacher model and the student model, three training objectives: video-text contrast, video-text matching, and video-guided masked language modeling The video-text contrast is to align the non-masked video and text embeddings The video-text matching is to fuse and classify the non-masked video tokens and text tokens The video-guided masked language modeling is to use the cross-modal decoder to predict the masked text from the text and non-masked video tokens 6. The method according to claim 5 Characterized in that After the training of the cross-modal task model, it further includes fine-tuning the cross-modal task model The fine-tuning refers to, after the completion of pre-training, selectively adjusting the parameters of the visual encoder, text encoder, and cross-modal decoder using downstream task data according to the expected functions to be achieved 7. The method according to claim 6, wherein, the fine-tuning includes: for a pure video task, fine-tuning the initially trained visual encoder; for a video retrieval task, fine-tuning the visual encoder and the text encoder; for a video question answering task, fine-tuning the visual encoder, the text encoder, and the cross-modal decoder.
8. A cloud server, characterized in that: a computer program capable of executing any one of the methods according to claims 1 to 7 is deployed on the cloud server.
Citation Information
Patent Citations
Visual language model obtaining method and device, visual language task processing method and device, equipment and storage medium
CN113792113A
Self-supervised document representation learning
US20220382975A1