Self-supervised video representation learning method and system based on time sequence correspondence

By adopting the timing sampling strategy of dual auxiliary frames and the auxiliary branch of the image block matching module in video representation learning, the reconstruction uncertainty and insufficient information compression caused by the timing sampling strategy in masked video modeling are solved, and a high-level semantic representation with time perception is achieved.

CN120198832APending Publication Date: 2025-06-24INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510202615.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing masked video modeling methods have reconstruction uncertainty problems in timing sampling strategies and are insufficient in information compression, which leads to the model being easily confused in complex dynamic scenarios and is difficult to learn advanced semantic representations.

Method used

A self-supervised video representation learning method based on timing correspondence is proposed. The timing sampling strategy of dual auxiliary frames and auxiliary branches containing image block matching module are adopted. The supervision signal is generated through the self-distillation mechanism, and the big model is guided to perform characterization and reconstruction in hidden space to generate a high-level semantic representation with time perception.

Benefits of technology

The timing sampling strategy of dual auxiliary frames reduces the uncertainty of representation reconstruction, and the model performs more stably in complex dynamic scenarios; through the auxiliary branches of the image block matching module, the dependence on redundant low-level visual information is reduced, and the model's learning ability for advanced semantic representation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198832A_ABST
    Figure CN120198832A_ABST
Patent Text Reader

Abstract

The invention discloses a self-supervised video representation learning method based on time sequence correspondence, and the method comprises the steps: carrying out the random sampling of a frame as a current frame for each video based on a given video data training set, carrying out the random masking of an image block of the current frame, and respectively sampling a frame from the past moment and the future moment of the current frame as auxiliary frames; inputting the auxiliary frame into an auxiliary branch, inputting the current frame of the mask into a student branch, retrieving an image block most similar to a mask image block in the current frame from the auxiliary frame, performing representation reconstruction, and establishing a time sequence corresponding relation between frames; and inputting the maskless current frame into a teacher branch, generating a supervision signal through a self-distillation mechanism, guiding the large model to carry out representation reconstruction on the masked current frame in a hidden space, and generating a high-level semantic representation with time perception. According to the method, the uncertainty of representation reconstruction is reduced, and advanced semantic representation with time perception can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video representation learning, and particularly to a self-supervised video representation learning method and system based on temporal correspondence. Background Art

[0002] Self-supervised representation learning can efficiently utilize rich visual information without manual annotation and has made remarkable progress in recent years. Among them, Masked Image Modeling (MIM) is considered to be one of the most promising self-supervised paradigms. This method realizes the learning of representations by using an encoder-decoder architecture to reconstruct the masked image patches. The success of Masked Image Modeling provides inspiration for its extension to the video field, thus giving rise to the Masked Video Modeling (MVM) method. The main difference between video data and image data is the introduction of an additional time dimension. Masked Image Modeling uses spatially adjacent image patches to establish spatial correspondence, thereby restoring the masked image patches, while Masked Video Modeling extends this idea to the time dimension, by using temporally adjacent frames to establish temporal correspondence, thereby reconstructing the image patches in the masked frames.

[0003] This extension from Masked Image Modeling to Masked Video Modeling brings two key problems: First, how to design a reasonable frame sampling strategy to effectively establish temporal correspondence; Second, how to efficiently reconstruct the masked video representation to capture the inter-frame correlation and learn temporal information.

[0004] Regarding the first problem, early studies mainly adopted a dense sampling strategy, that is, adding a three-dimensional mask to the video, and then directly inputting all frames into a spatio-temporal encoder for processing. These methods implicitly encode the sampled frames by forcing the model to restore the masked image patches. However, due to the lack of a clear mechanism to guide this process, these methods rely on large models and large datasets to capture the temporal structure in the video, resulting in a high computational cost. Recent studies have turned to a random sampling strategy, that is, only randomly masking the image patches of one frame, using another frame as a conditional frame to provide prior information, and reconstructing the image patches of the masked frame through a conditional decoder. Although this method reduces the computational cost, since a single conditional frame may produce multiple potential prediction results, increasing the uncertainty of representation reconstruction, the model is prone to confusion in complex dynamic scenarios.

[0005] Regarding the second problem, existing masked video modeling methods mainly focus on restoring the masked content in the pixel space. Due to the high redundancy of video content, restoring all masked pixels usually forces the model to over - focus on low - level visual information, thus limiting its ability to learn high - level semantic representations. This method lacking in irrelevant information compression leads to poor performance of the model in downstream video tasks because the model tends to over - fit to surface visual features.

[0006] In summary, the following drawbacks exist in the prior art:

[0007] The reconstruction uncertainty problem caused by the temporal sampling strategy and the insufficient information compression problem caused by the representation reconstruction method. First, the random temporal sampling strategy adopted by existing methods introduces uncertainty into the representation reconstruction process, resulting in the model being prone to confusion in complex dynamic scenes; second, existing methods mainly restore the masked image patches in the pixel space, with insufficient information compression ability, causing the model to over - focus on low - level visual information and limiting the learning of high - level semantic representations.

[0008] Therefore, there is an urgent need to propose a self - supervised video representation learning model based on temporal correspondence to learn high - level semantic representations with time awareness, so as to ensure effectiveness in various video downstream tasks. Summary of the Invention

[0009] To solve the above problems such as the reconstruction uncertainty problem caused by the temporal sampling strategy in the prior art and the insufficient information compression problem caused by the representation reconstruction method, a self - supervised video representation learning method based on temporal correspondence is proposed.

[0010] In a first aspect, an embodiment of the present application provides a self - supervised video representation learning method based on temporal correspondence, the method comprising:

[0011] Temporal sampling step: Based on a given video data training set, for each video, randomly sample one frame as the current frame, and after randomly masking the image patches of the current frame, sample one frame from the past and future moments of the current frame respectively as the auxiliary frames;

[0012] Temporal correspondence construction step: Input the auxiliary frames into the auxiliary branch, input the masked current frame into the student branch, retrieve the image patches in the auxiliary frames that are most similar to the masked image patches in the current frame, perform representation reconstruction, and establish the temporal correspondence between frames;

[0013] Representation reconstruction step: Input the unmasked current frame into the teacher branch, generate a supervision signal through the self - distillation mechanism, and guide the large model to perform representation reconstruction on the masked current frame in the latent space to generate high - level semantic representations with time awareness.

[0014] In the specific embodiments of the present invention, the above-mentioned self-supervised video representation learning method based on temporal correspondence further includes:

[0015] Steps for training and updating the model: Optimize the student branch by the gradient descent method to minimize the representation reconstruction error, and update the teacher branch by the exponential moving average method to ensure the stability and consistency of the teacher branch.

[0016] In the specific embodiments of the present invention, the above-mentioned temporal sampling step further includes:

[0017] Step of sampling the current frame: In the given video, select the current frame based on the time step, and then randomly mask some image patches of the current frame to obtain the masked current frame, generating a set of masked image patches;

[0018] Step of sampling auxiliary frames: In the given video, relative to the time step, randomly sample the past frame time step and the future frame time step according to a uniform distribution within a predefined offset interval;

[0019] Encoding step: The masked current frame is encoded by the encoder of the student branch to obtain the representation of the student branch; the past frame and the future frame are respectively encoded by the encoders of the auxiliary branch to obtain their respective past frame representations and future frame representations.

[0020] In the specific embodiments of the present invention, the above-mentioned step of constructing temporal correspondence further includes:

[0021] Time matching step based on the cross-attention mechanism: Retrieve the image patches in the auxiliary frames that are most similar to the masked image patches in the current frame, use the masked representation as the query condition, the auxiliary representation as the key and value, and use the cross-attention method to establish temporal correspondence between the image patches of different frames;

[0022] Spatial matching step based on the self-attention mechanism: Adopt the self-attention mechanism to process for spatial matching and refinement, and learn the spatial dependence relationship between the image patches of the current frame;

[0023] Dimension transformation step based on the feed-forward neural network: For the past frame representation and the future frame representation, through the feed-forward neural network processing method, adjust the representation in the latent space to obtain the reconstructed representation of the masked current frame containing the information of the past frame and the future frame.

[0024] In the specific embodiments of the present invention, the above-mentioned representation reconstruction step further includes:

[0025] Representation reconstruction constraint step: Introduce projection heads after the student and teacher models respectively, and after applying centering and sharpening techniques to the outputs of the projection heads, reconstruct the masked image patches in the current frame based on the representation reconstruction loss, calculate the representation reconstruction loss, and obtain the image patch-level loss value;

[0026] Temporal Compression Constraint Step: An additional constraint term is introduced on the values of the projection heads of the student and teacher models after centering and sharpening.

[0027] Frame-Level Alignment Constraint Step: For the sharpened and normalized embedding values of the student and teacher models, the DINO loss is used to align the global semantics, and at the same time, the KoLeo loss is used to promote the diversity of the representations within the sample batch, obtaining the frame-level loss value.

[0028] In the specific embodiments of the present invention, the training and updating steps of the above model further include:

[0029] Define Total Loss Function Step: Combine the patch-level loss and the frame-level loss, calculate the total loss function, and achieve collaborative learning between local and global semantic representations.

[0030] Student Branch Update Step: Update the model parameters of the student branch using the gradient descent method based on the optimizer.

[0031] Teacher Branch Update Step: Update the model parameters of the teacher branch using the exponential moving average method.

[0032] In a second aspect, an embodiment of the present application provides a self-supervised video representation learning system based on temporal correspondence, which adopts the self-supervised video representation learning method as described above. The system includes:

[0033] Temporal Sampling Module: Based on a given video data training set, for each video, randomly sample one frame as the current frame, randomly mask the image patches of the current frame, and then sample one frame from the past and future moments of the current frame as the auxiliary frames.

[0034] Temporal Correspondence Relationship Construction Module: Input the auxiliary frames into the auxiliary branch, input the masked current frame into the student branch, retrieve the image patches in the auxiliary frames that are most similar to the masked image patches in the current frame, perform representation reconstruction, and establish the temporal correspondence relationship between frames.

[0035] Representation Reconstruction Module: Input the unmasked current frame into the teacher branch, generate a supervision signal through the self-distillation mechanism, guide the large model to perform representation reconstruction on the masked current frame in the latent space, and generate high-level semantic representations with time awareness.

[0036] In the specific embodiments of the present invention, the self-supervised video representation learning system based on temporal correspondence further includes:

[0037] Model Training and Updating Module: Optimize the student branch through the gradient descent method to minimize the representation reconstruction error, and update the teacher branch through the exponential moving average method to ensure the stability and consistency of the teacher branch.

[0038] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the self-supervised video representation learning method based on temporal correspondence are implemented.

[0039] In a fourth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the self-supervised video representation learning method based on temporal correspondence as described above are implemented.

[0040] Compared with the related prior art, the following outstanding beneficial effects are achieved:

[0041] 1) The method of the present invention provides a temporal sampling strategy of dual auxiliary frames to reduce the uncertainty of representation reconstruction. For a given masked frame, an auxiliary frame is sampled from the past and future moments respectively based on a set offset. Subsequently, the encoder reconstructs the image patch representation of the masked frame under the guidance of these two auxiliary frames. Through this bidirectional compression design, obvious temporal clues are provided for the reconstruction of the masked frame image patches, reducing the uncertainty of representation reconstruction, while modeling the temporal correspondence between frames, thereby providing more accurate temporal information for video representation learning;

[0042] 2) The method of the present invention provides an auxiliary branch adapted to the existing self-distillation architecture to learn robust high-level semantic representations. Based on the current advanced self-distillation architecture, an auxiliary branch including an image patch matching module is designed and introduced. The auxiliary frame is used to provide a temporal reference for representation reconstruction, and the representation of the masked image patches is reconstructed in the latent space, thereby reducing the dependence on redundant low-level visual information, enabling the model to focus on the learning of high-level semantic representations, and enhancing the robustness and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0044] Figure 1 Schematic diagram of the self-supervised video representation learning method of the present invention Figure 1 ;

[0045] Figure 2 Schematic diagram of the self-supervised video representation learning method of the present invention Figure 2 ;

[0046] Figure 3 Schematic diagram of the system of the self-supervised video representation learning method of the present invention;

[0047] Figure 4 Schematic diagram of the computer hardware of the present invention. Detailed implementation manners

[0048] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or a similar expression means any combination of these items, including any combination of single item or plural items. For example, at least one of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c may be single or multiple.

[0049] It should also be understood that the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship. The specific meaning can be understood by referring to the context before and after.

[0050] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above - mentioned processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0051] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0052] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0053] In addition, in each embodiment of the present invention, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0054] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0055] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given and will be described in detail in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0056] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0057] The method of the present invention aims to propose a self-supervised video representation learning model based on temporal correspondence (referred to as T-CoRe). Specifically, aiming at the problem of reconstruction uncertainty, the inventor proposes a temporal sampling strategy of dual auxiliary frames, which reduces the uncertainty of representation reconstruction by means of bidirectional compression; aiming at the problem of insufficient information compression, the inventor designs and introduces an auxiliary branch containing an image patch matching module, integrates it into the existing self-distillation architecture, and generates high-level semantic representations with temporal awareness by performing representation reconstruction in the latent space.

[0058] Embodiment 1

[0059] As Figure 1 shown, the embodiment of the present application provides a self-supervised video representation learning method based on temporal correspondence. The method includes:

[0060] Temporal sampling step 101: Based on a given video data training set, for each video, randomly sample one frame as the current frame, randomly mask the image patches of the current frame, and then sample one frame from the past and future moments of the current frame as the auxiliary frames;

[0061] In a specific embodiment of the present invention, double - auxiliary - frame temporal sampling is adopted. Given a video data training set, for each video, first randomly sample a frame as the current frame, and randomly mask the image patches of this frame. Subsequently, sample a frame from each of the past and future moments as auxiliary frames.

[0062] Temporal correspondence construction step 102: Input the auxiliary frame into the auxiliary branch, input the masked current frame into the student branch, retrieve the image patch in the auxiliary frame that is most similar to the masked image patch in the current frame, perform feature reconstruction, and establish the temporal correspondence between frames;

[0063] In a specific embodiment of the present invention, image - patch matching is used to construct the temporal correspondence. Input the auxiliary frame into the auxiliary branch, input the masked current frame into the student branch. The masked region in the current frame queries the most similar image patch in the auxiliary frame through the image - patch matching module for feature reconstruction, thereby effectively introducing temporal information and establishing the temporal correspondence between frames.

[0064] Feature reconstruction step 103: Input the un - masked current frame into the teacher branch, generate a supervision signal through the self - distillation mechanism, guide the large model to perform feature reconstruction on the masked current frame in the latent space, and generate high - level semantic features with time awareness.

[0065] In a specific embodiment of the present invention, feature reconstruction based on the self - distillation architecture is adopted. Input the un - masked current frame into the teacher branch, generate a supervision signal through the self - distillation mechanism, and guide the model to perform feature reconstruction on the masked current frame, thereby enhancing the quality and robustness of the features.

[0066] Embodiment 2

[0067] As Figure 2 shown, an embodiment of the present application provides a self - supervised video feature learning method based on temporal correspondence. The method includes:

[0068] Temporal sampling step 201: Based on a given video data training set, for each video, randomly sample a frame as the current frame, and after randomly masking the image patches of the current frame, sample a frame from each of the past and future moments of the current frame as auxiliary frames;

[0069] Temporal correspondence construction step 202: Input the auxiliary frame into the auxiliary branch, input the masked current frame into the student branch, retrieve the image patch in the auxiliary frame that is most similar to the masked image patch in the current frame, perform feature reconstruction, and establish the temporal correspondence between frames;

[0070] Characterization reconstruction step 203: Input the current frame without a mask into the teacher branch, generate a supervision signal through the self-distillation mechanism, guide the large model to perform characterization reconstruction on the masked current frame in the latent space, and generate a high-level semantic characterization with temporal awareness.

[0071] Model training and update step 204: Optimize the student branch through the gradient descent method to minimize the characterization reconstruction error, and update the teacher branch through the exponential moving average method to ensure the stability and consistency of the teacher branch.

[0072] In a specific embodiment of the present invention, the model is trained and updated by optimizing the student branch through the gradient descent method to minimize the characterization reconstruction error, and updating the teacher branch through the exponential moving average method to ensure the stability and consistency of the teacher branch.

[0073] In a specific embodiment of the present invention, the above-mentioned temporal sampling step 201 further includes:

[0074] Current frame sampling step: In a given video, select the current frame based on the time step, and then randomly mask some image patches of the current frame to obtain the masked current frame, generating a set of masked image patches;

[0075] In a specific embodiment of the present invention, when sampling the current frame, in a given video V, based on the time step t c select the current frame v c = V(t c ), and then randomly mask some image patches of the current frame to obtain the masked current frame The set of masked image patches is denoted as

[0076] Auxiliary frame sampling step: In a given video, relative to the time step, randomly sample the past frame time step and the future frame time step according to a uniform distribution within a predefined offset interval;

[0077] In a specific embodiment of the present invention, when sampling the auxiliary frame, in a given video V, relative to the time step t c , randomly sample the past frame time step and the future frame time step according to a uniform distribution within a predefined offset interval [α, β]:

[0078] t p ~U(t c - α·T, t c - β·T)

[0079] t f ~U(t c + α·T, t c + β·T)

[0080] where U(·, ·) is a uniform distribution and T is the number of frames in the video. Subsequently, the past frame v p = V(t p ) and the future frame v f = V(t f ).

[0081] Encoding step: The masked current frame is encoded by the encoder of the student branch to obtain the representation of the student branch; the past frame and the future frame are respectively encoded by the encoders of the auxiliary branch to obtain their respective past frame representations and future frame representations.

[0082] In a specific embodiment of the present invention, the masked current frame is encoded by the encoder f of the student branch to obtain the representation The past frame v p and the future frame v f are respectively encoded by the encoder f of the auxiliary branch to obtain their respective representations z p = f(v p ) and z f = f(v f ). During the encoding process, the Vision Transformer model is used as the encoder.

[0083] In a specific embodiment of the present invention, the above-mentioned temporal correspondence construction step 202 further includes:

[0084] Time matching step based on the cross-attention mechanism: Retrieving the image block most similar to the masked image block in the current frame from the auxiliary frames, using the masked representation as the query condition, the auxiliary representation as the key and value, and using the cross-attention method to establish temporal correspondence between the image blocks of different frames;

[0085] In a specific embodiment of the present invention, for the time matching based on the cross-attention mechanism, the image block most similar to the masked image block in the current frame is retrieved from the auxiliary frames. Taking the use of the past frame representation z p as an example, using the masked representation as the query, the auxiliary representation z p as the key and value, and using the cross-attention module to establish temporal correspondence between the image blocks of different frames:

[0086]

[0087] where is the query matrix, K p = z p W k is the key matrix, V p = z p W v is the value matrix, W q , Wk , is the weight matrix, d is the representation dimension, τ is a temperature coefficient greater than 0, and the softmax function is as follows:

[0088]

[0089] Intuitively, represents the similarity matrix. In this way, the model retrieves the most similar image patches from the auxiliary frame in the latent space to fill in the masked image patches in the current frame. Different from the traditional method of inferring the missing image patches, this method retrieves from the candidate image patches provided by the auxiliary frame, thus more accurately restoring the masked image patches.

[0090] Spatial matching step based on the self-attention mechanism: The self-attention mechanism is used to process for spatial matching and refinement, and learn the spatial dependence relationship between the image patches of the current frame;

[0091] In the specific embodiment of the present invention, for spatial matching based on the self-attention mechanism, in order to further enhance the spatial co-occurrence relationship between the image patches, the self-attention module is used to process z c ' for spatial matching and refinement, so as to learn the spatial dependence relationship between the image patches of the current frame:

[0092]

[0093] where Q c ' = z c 'W q ' is the query matrix, K c ' = z c 'W k ' is the key matrix, V c ' = z c 'W v ' is the value matrix, W q ', W k ', is the weight matrix.

[0094] Dimension transformation step based on the feed-forward neural network: For the past frame representation and the future frame representation, through the feed-forward neural network processing method, the representation is adjusted in the latent space to obtain the masked current frame reconstruction representation containing the information of the past frame and the future frame.

[0095] In the specific embodiment of the present invention, for the dimension transformation based on the feed-forward neural network, in order to enhance the expression ability of the representation, z c ” is processed by the feed-forward neural network to adjust the representation in the latent space, thereby improving the adaptability of the model to complex tasks:

[0096]

[0097] That is, a masked current frame reconstruction representation containing past frame information. Similarly, for the future frame representation z f , it is processed by the same method as from the time matching step based on the cross-attention mechanism to the dimension transformation step based on the feed-forward neural network, so as to obtain a masked current frame reconstruction representation containing future frame information

[0098] In a specific embodiment of the present invention, the above representation reconstruction step 203 further includes:

[0099] Representation reconstruction constraint step: Introduce projection heads after the student and teacher models respectively, and after applying centering and sharpening techniques to the outputs of the projection heads, reconstruct the masked image patches in the current frame based on the representation reconstruction loss, calculate the representation reconstruction loss, and obtain the image patch-level loss value;

[0100] In a specific embodiment of the present invention, the representation reconstruction constraint: To prevent the model from collapsing, projection heads h and h are introduced after the student and teacher models t , and centering and sharpening techniques are applied to the outputs of the projection heads:

[0101]

[0102] where τ t is a temperature coefficient greater than 0, and μ t is an empirical estimator of the average value of the representation. Subsequently, the masked image patches in the current frame are reconstructed based on the representation reconstruction loss, and the representation reconstruction loss is designed as follows:

[0103]

[0104] Temporal compression constraint step: An additional constraint term is introduced respectively on the values of the student and teacher model projection heads after centering and sharpening;

[0105] In a specific embodiment of the present invention, the temporal compression constraint, in and an additional constraint term is introduced to further reduce the uncertainty during training:

[0106]

[0107] Intuitively, without the reconstruction representation may overly rely on past frames and future frames. Therefore, even if and are both small, but and may exhibit obvious inconsistent phenomena. This deviation will introduce additional uncertainty, resulting in the model being confused in complex scenarios.

[0108] Frame-level alignment constraint step: For the sharpened and normalized embedding values of the student model and the teacher model, use the DINO loss to align the global semantics, and at the same time use the KoLeo loss to promote the diversity of representations within the sample batch to obtain the frame-level loss value.

[0109] In a specific embodiment of the present invention, for frame-level alignment constraint, let p and z represent the sharpened and L2-normalized [CLS] embeddings of the student model, representing the centered and sharpened [CLS] embeddings of the teacher model. Use the DINO loss to align the global semantics:

[0110]

[0111] At the same time, use the KoLeo loss to promote the diversity of representations within the sample batch:

[0112]

[0113] where B represents the batch size, representing the distance between z i within the batch and its nearest neighbor.

[0114] In a specific embodiment of the present invention, the training and updating step 204 of the above model further includes:

[0115] Step of defining the total loss function: Combine the patch-level loss and the frame-level loss to calculate the total loss function to achieve collaborative learning between local and global semantic representations;

[0116] In a specific embodiment of the present invention, define the total loss function, combine the above patch-level loss and frame-level loss, and design the total loss function to achieve collaborative learning between local and global semantic representations:

[0117]

[0118] where λ i is a hyperparameter that controls the weight of each sub-loss function.

[0119] Student branch update step: Update the model parameters of the student branch using the gradient descent method based on the optimizer; In a specific embodiment of the present invention, the student branch update uses the gradient descent method based on the AdamW optimizer to update the model parameters of the student branch

[0120] Teacher branch update step: Update the model parameters of the teacher branch using the exponential moving average method. In a specific embodiment of the present invention, the teacher branch update uses the exponential moving average method to update the model parameters of the teacher branch:

[0121] θ t ← m·θt +(1 - m)·θ

[0122] where θ and θ t are the parameters of the student model f and the teacher model f t respectively.

[0123] After the training phase ends, the trained teacher model will be used to interface with a variety of video downstream tasks, thereby performing a performance evaluation of the model's expressive ability, verifying the effectiveness of the model in fine-grained video tasks such as video object segmentation and human pose detection, and providing a more accurate and robust representation for the actual scenarios of video understanding.

[0124] Experimental data of specific embodiments of the present invention

[0125] Table 1 shows the performance comparison with existing methods on three video downstream tasks: video object segmentation (DAVIS-2017 dataset), semantic part segmentation (VIP dataset), and human pose propagation (JHMDB dataset). Missing values indicate the lack of relevant reference data. The best and second-best results are highlighted in bold and underlined respectively. The method proposed in the present invention is superior to existing methods in most cases.

[0126]

[0127]

[0128] Table 1

[0129] Table 2 shows the performance comparison with existing methods using a larger model on three video downstream tasks: video object segmentation (DAVIS-2017 dataset), semantic part segmentation (VIP dataset), and human pose propagation (JHMDB dataset). The best and second-best results are highlighted in bold and underlined respectively. The method proposed in the present invention is superior to existing methods.

[0130]

[0131] Table 2

[0132] In summary, the method of the present invention proposes a self-supervised video representation learning model based on temporal correspondence, aiming to solve the problems of reconstruction uncertainty and insufficient information compression existing in the existing masked video modeling methods. The present invention first proposes a temporal sampling strategy of dual auxiliary frames, which reduces the reconstruction uncertainty of the representation in a bidirectional compression manner and provides more accurate temporal information for video representation learning. Subsequently, by designing an auxiliary branch including an image patch matching module, the reconstruction of the masked representation is completed based on the temporal correspondence relationship in the latent space, generating high-level semantic representations with time awareness, thereby showing good performance in specific video downstream tasks such as video object segmentation and human pose detection, and further enhancing the application value of the model in video understanding and analysis.

[0133] Embodiment III

[0134] As Figure 3 shown, an embodiment of the present application provides a self-supervised video representation learning system based on temporal correspondence, adopting the self-supervised video representation learning method as described above. The system includes:

[0135] Temporal sampling module 301: Based on a given video data training set, for each video, randomly sample one frame as the current frame, and after randomly masking the image patches of the current frame, sample one frame from the past and future moments of the current frame as the auxiliary frames respectively;

[0136] Temporal correspondence relationship construction module 302: Input the auxiliary frames into the auxiliary branch, input the masked current frame into the student branch, retrieve the image patches in the auxiliary frames that are most similar to the masked image patches in the current frame, perform representation reconstruction, and establish the temporal correspondence relationship between frames;

[0137] Representation reconstruction module 303: Input the unmasked current frame into the teacher branch, generate a supervision signal through the self-distillation mechanism, guide the large model to perform representation reconstruction on the masked current frame in the latent space, and generate high-level semantic representations with time awareness.

[0138] Model training and updating module 304: Optimize the student branch by the gradient descent method, minimize the representation reconstruction error, and update the teacher branch by the exponential moving average method to ensure the stability and consistency of the teacher branch.

[0139] Embodiment IV

[0140] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the self-supervised video representation learning method based on temporal correspondence are implemented.

[0141] Embodiment V

[0142] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the self-supervised video representation learning method based on temporal correspondence as described are implemented.

[0143] In addition, combined with Figure 1 the self-supervised video representation learning method according to the embodiment of the present application described can be implemented by an electronic device, such as a computer device. Figure 4 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application.

[0144] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 4 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.

[0145] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits implementing the embodiments of the present application.

[0146] The memory 82 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.

[0147] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the self-supervised video representation learning methods based on temporal correspondence in the above embodiments.

[0148] The technical features of the above-mentioned embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as these combinations of technical features do not conflict, they should all be considered as the scope recorded in this specification.

[0149] The above-mentioned embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A self-supervised video representation learning method based on temporal correspondence, characterized in that: The method comprises: Time series sampling step: based on a given video data training set, for each video, randomly sample a frame as the current frame, and after randomly masking the image block of the current frame, sample a frame from the past moment and the future moment of the current frame respectively as an auxiliary frame; The step of constructing a temporal correspondence relationship is as follows: inputting the auxiliary frame into the auxiliary branch, inputting the masked current frame into the student branch, retrieving the image block most similar to the masked image block in the current frame from the auxiliary frame, performing representation reconstruction, and establishing a temporal correspondence relationship between frames; Representation reconstruction step: input the unmasked current frame into the teacher branch, generate a supervision signal through the self-distillation mechanism, guide the large model to reconstruct the representation of the masked current frame in the latent space, and generate a high-level semantic representation with time perception.

2. The self-supervised video representation learning method based on temporal correspondence according to claim 1, characterized in that: The method further comprises: The training and updating steps of the model are as follows: the student branch is optimized by the gradient descent method to minimize the representation reconstruction error, and the teacher branch is updated by the exponential moving average method to ensure the stability and consistency of the teacher branch.

3. The self-supervised video representation learning method based on temporal correspondence according to claim 1 or 2, characterized in that: The timing sampling step further includes: Sampling the current frame step: In a given video, the current frame is selected based on the time step, and then some image blocks of the current frame are randomly masked to obtain a masked current frame, generating a set of masked image blocks; Sampling auxiliary frame step: In a given video, relative to the time step, randomly sample past frame time steps and future frame time steps according to a uniform distribution according to a predefined offset interval; Coding step: the masked current frame is encoded by the encoder of the student branch to obtain the representation of the student branch; the past frame and the future frame are respectively encoded by the encoder of the auxiliary branch to obtain their own past frame representation and future frame representation.

4. The self-supervised video representation learning method based on temporal correspondence according to claim 1 or 2, characterized in that: The step of constructing the time sequence correspondence relationship further includes: Temporal matching step based on cross-attention mechanism: retrieve the image patch most similar to the mask image patch in the current frame from the auxiliary frame, use the mask representation as the query condition, the auxiliary representation as the key and value, and use the cross-attention method to establish temporal correspondence between image patches in different frames; Spatial matching step based on self-attention mechanism: self-attention mechanism is used to perform spatial matching and refinement, and learn the spatial dependency between image blocks in the current frame; Dimension transformation steps based on feedforward neural network: For the past frame representation and the future frame representation, the representation is adjusted in the latent space through the feedforward neural network processing method to obtain the masked current frame reconstructed representation containing the past frame and the future frame information.

5. The self-supervised video representation learning method based on temporal correspondence according to claim 1 or 2, characterized in that: The characterization reconstruction step further includes: Representation reconstruction constraint step: introduce a projection head after the student and teacher models respectively, and after applying centering and sharpening techniques to the output of the projection head, reconstruct the mask image block in the current frame based on the representation reconstruction loss, calculate the representation reconstruction loss, and obtain the image block level loss value; Timing compression constraint step: an additional constraint term is introduced on the centering and sharpening values ​​of the projection heads of the student and teacher models respectively; Frame-level alignment constraint step: For the sharpened and normalized embedding values ​​of the student model and the teacher model, DINO loss is used to align the global semantics, and KoLeo loss is used to promote the diversity of representation within the sample batch to obtain the frame-level loss value.

6. The self-supervised video representation learning method based on temporal correspondence according to claim 1 or 2, characterized in that: The training and updating steps of the model also include: Defining a total loss function step: combining the image block level loss and the frame level loss, calculating the total loss function, and realizing collaborative learning between local semantics and global semantic representation; Student branch update step: Update the model parameters of the student branch using the gradient descent method based on the optimizer; Teacher branch update step: Use exponential moving average method to update the model parameters of the teacher branch.

7. A self-supervised video representation learning system based on temporal correspondence, using the self-supervised video representation learning method based on temporal correspondence as claimed in any one of claims 1 to 6, characterized in that: The system comprises: Time series sampling module: Based on a given video data training set, for each video, randomly sample a frame as the current frame, and after randomly masking the image block of the current frame, sample a frame from the past moment and the future moment of the current frame respectively as an auxiliary frame; A temporal correspondence building module: inputting the auxiliary frame into the auxiliary branch, inputting the masked current frame into the student branch, retrieving the image block most similar to the masked image block in the current frame from the auxiliary frame, performing representation reconstruction, and establishing a temporal correspondence between frames; Representation reconstruction module: The unmasked current frame is input into the teacher branch, and a supervision signal is generated through the self-distillation mechanism to guide the large model to reconstruct the representation of the masked current frame in the latent space, thereby generating a high-level semantic representation with time perception.

8. The self-supervised video representation learning system based on temporal correspondence according to claim 7, characterized in that: The system further comprises: The training and updating module of the model: The student branch is optimized by the gradient descent method to minimize the representation reconstruction error, and the teacher branch is updated by the exponential moving average method to ensure the stability and consistency of the teacher branch.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the self-supervised video representation learning method based on temporal correspondence described in any one of claims 1 to 6 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the self-supervised video representation learning method based on temporal correspondence are implemented as described in any one of claims 1 to 6.