Video-based self-supervised training method and related device
Through the self-supervised training method, the feature extraction and image reconstruction of video samples are used, combined with multi-faceted loss information constraint model training, the problem of insufficient generalization ability in video self-supervised pre-training is solved, and the output of new labels is supported under a small amount of data and the performance of the model is improved.
Patent Information
- Application Number
- CN202410034660.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-11
AI Technical Summary
When processing video data, the existing video self-supervised pre-training method is difficult to effectively utilize the video timing information and semantic information, resulting in insufficient generalization ability of the model in video understanding tasks and requires a large amount of labeling data to support the output of newly added labels.
By obtaining at least two frames of image samples in the video sample, the training model is self-supervised, and the feature extraction module and the image reconstruction module are used to perform feature extraction and image reconstruction. Combining the first loss information between the target image features and the comparison image features, as well as the second loss information between the target input image and the target output image, the model parameters are updated until the training is completed, and the pre-trained model is generated.
The generalization capability and performance of the model are improved, so that the output of new tags can be supported when using a small amount of data, and the number of iterations of downstream tasks and the amount of training data is reduced, and the rapid changes in market demand are adapted to the rapidly changing market demand.
Smart Images

Figure CN120298940A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, this application relates to a self-supervised training method based on video and related devices. Background Art
[0002] With the continuous growth of video data, the richness of video content has also increased day by day. However, overly complex video content poses challenges to the model. Currently, a huge amount of video data is generated every day. For example, in news applications, downstream tasks need to identify and recommend tags for all videos, which is a task that requires processing a large amount of unlabeled data. In order to enable the model to perform optimally on new data and continuously support new tags, it is necessary to use as little labeled data as possible.
[0003] Therefore, there is an urgent need for a model with strong generalization ability so that a small amount of data can be used to support the output of new tags. Summary of the Invention
[0004] Embodiments of this application provide a self-supervised training method based on video and related devices, which can achieve the technical effect of using a small amount of data to support the output of new tags.
[0005] On the one hand, embodiments of this application provide a self-supervised training method based on video, including:
[0006] Obtain at least two first image samples in the video sample;
[0007] Perform self-supervised training on the model to be trained based on at least two first image samples:
[0008] Extract features of the target input image through the feature extraction module in the model to be trained to obtain target image features; and reconstruct the target image features through the image reconstruction module in the model to be trained to obtain the target output image; and extract features of the target output image through the feature extraction module to obtain comparison image features; wherein, the target input image is obtained based on at least two first image samples;
[0009] Determine the first loss information between the target image features and the comparison image features, and determine the second loss information between the target input image and the target output image;
[0010] If it is determined based on the first loss information and the second loss information that the model to be trained has not finished training, update the model parameters of the model to be trained and continue to perform self-supervised training on the model to be trained;
[0011] If it is determined that the training of the model to be trained is completed based on the first loss information and the second loss information, the trained model to be trained is used as a pre-trained model to process the target video through the pre-trained model to obtain a target video processing result.
[0012] On the other hand, an embodiment of the present application further provides a self-supervised training device based on video, including:
[0013] An acquisition module for acquiring at least two first image samples in a video sample;
[0014] A training module for performing self-supervised training on the model to be trained based on at least two first image samples:
[0015] Feature extraction is performed on the target input image through a feature extraction module in the model to be trained to obtain target image features; and image reconstruction is performed on the target image features through an image reconstruction module in the model to be trained to obtain a target output image; and feature extraction is performed on the target output image through the feature extraction module to obtain comparison image features; wherein, the target input image is obtained based on at least two first image samples;
[0016] Determine the first loss information between the target image features and the comparison image features, and determine the second loss information between the target input image and the target output image;
[0017] If it is determined that the training of the model to be trained is not completed based on the first loss information and the second loss information, continue to perform self-supervised training on the model to be trained;
[0018] If it is determined that the training of the model to be trained is completed based on the first loss information and the second loss information, the trained model to be trained is used as a pre-trained model to process the target video through the pre-trained model to obtain a target video processing result.
[0019] Optionally, when the training module obtains the target input image based on at least two first image samples, it can be used for at least one of the following:
[0020] Use at least two first image samples as the target input image;
[0021] Divide at least two first image samples into blocks to obtain a plurality of first image sample blocks, and use the plurality of first image sample blocks as the target input image.
[0022] Optionally, when the training module performs image reconstruction on the target image features through the image reconstruction module in the model to be trained to obtain a target output image, it can be used for:
[0023] The target image features are reconstructed at the pixel level through the image reconstruction module in the model to be trained, and a target output image matching the target input image is obtained;
[0024] Among them, if the target input image includes at least two first image samples, the target output image includes at least two second image samples; if the target input image includes multiple first image sample blocks, the target input image includes multiple second image sample blocks.
[0025] Optionally, the image to be extracted includes multiple image sample blocks, and the feature extraction module includes an image occlusion network and a backbone network. The feature extraction module extracts image features from the image to be extracted through the following method by the training module:
[0026] The multiple image sample blocks are occluded through the image occlusion network to obtain the occluded multiple image sample blocks, and the occluded multiple image sample blocks include unoccluded image pixels;
[0027] The image features of the occluded multiple image sample blocks are extracted through the backbone network to obtain the image features output by the backbone network;
[0028] Among them, if the image to be extracted includes the target input image, the image features output by the image backbone network include the target image features; if the image to be extracted includes the target output image, the image features output by the image backbone network include the comparison image features.
[0029] Optionally, the multiple image sample blocks include at least two image sample blocks corresponding to each frame of the image sample. When the training module occludes the multiple image sample blocks through the image occlusion network, it can be used for:
[0030] Occlude all the image sample blocks corresponding to at least one frame of the image sample and retain all the image sample blocks corresponding to at least one frame of the image sample;
[0031] Occlude some of the image sample blocks corresponding to each frame of the image sample;
[0032] Randomly select some of the image sample blocks in the multiple image sample blocks for occlusion.
[0033] Optionally, when the training module extracts the image features of the occluded multiple image sample blocks through the backbone network to obtain the target image features, it can be used for:
[0034] Extract the image features of at least some of the unoccluded image sample blocks in the multiple image sample blocks through the backbone network to obtain the target image features.
[0035] Optionally, the backbone network includes at least one level of feature extraction sub-network, and each level of feature extraction sub-network includes a non-convolutional downsampling unit and a feature extraction unit based on an attention mechanism. When the training module performs image feature extraction on multiple occluded image sample blocks through the backbone network, it can be used for:
[0036] For each level of feature extraction sub-network, perform downsampling on the feature extraction sub-network through the non-convolutional downsampling unit of the feature extraction sub-network to obtain a downsampling result, and perform feature extraction on the downsampling result through the feature extraction unit of the feature extraction sub-network to obtain a feature extraction result output by the feature extraction sub-network.
[0037] Optionally, the backbone network includes at least two cascaded levels of feature extraction sub-networks. When the training module performs image feature extraction on multiple occluded image sample blocks through the backbone network, it can also be used for:
[0038] The input of the first-level feature extraction sub-network includes the output of the image occlusion network, the input of the nth-level feature extraction sub-network includes the output of the (n - 1)th-level feature extraction sub-network, the output of the nth-level feature extraction sub-network serves as the input of the (n + 1)th-level feature extraction sub-network, and the output of the last-level feature extraction sub-network serves as the image features output by the image backbone network;
[0039] where 1 < n < m, and m is the total number of at least two cascaded levels of feature extraction sub-networks.
[0040] Optionally, when the training module performs image reconstruction on target image features through the image reconstruction module in the model to be trained to obtain a target output image, it can be used for:
[0041] Perform at least one of upsampling, linear mapping, or non-linear mapping on the target image features through the image reconstruction module in the model to be trained to reconstruct the target image features into a target output image.
[0042] Optionally, the target image features include the first image features corresponding to each first image sample block in multiple first image sample blocks, and the comparison image features include the second image features corresponding to the second image sample blocks corresponding to each first image sample block. When the training module determines the first loss information between the target image features and the comparison image features, it can be used for:
[0043] For each first image sample block, determine the first loss sub-information between the first image feature corresponding to the first image sample block and the second image feature corresponding to the second image sample corresponding to the first image sample block;
[0044] Fuse based on the first loss sub-information corresponding to each first image sample block to determine the first loss information between the target image features and the comparison image features.
[0045] Optionally, when determining the second loss information between the target input image and the target output image, the training module can be used to:
[0046] For each first image pixel in the target input image, determine the second image pixel corresponding to the first image pixel in the target output image, and determine the second loss sub-information between the first image pixel and the corresponding second image pixel;
[0047] Based on the fusion of the second loss sub-information corresponding to each first image pixel in the target input image, determine the second loss information between the target input image and the target output image.
[0048] On the other hand, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method according to any embodiment of the present application.
[0049] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any embodiment of the present application are implemented.
[0050] On the other hand, an embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method according to any embodiment of the present application are implemented.
[0051] The technical solution of this embodiment is to obtain at least two first image samples in a video sample; perform self-supervised training on a model to be trained based on the at least two first image samples. When performing self-supervised training on the model to be trained based on the at least two first image samples, feature extraction is performed on a target input image through a feature extraction module in the model to be trained to obtain target image features; image reconstruction is performed on the target image features through an image reconstruction module in the model to be trained to obtain a target output image; and feature extraction is performed on the target output image through the feature extraction module to obtain comparison image features. Wherein, the target input image is obtained based on the at least two first image samples; determine first loss information between the target image features and the comparison image features, and determine second loss information between the target input image and the target output image; if it is determined based on the first loss information and the second loss information that the model to be trained has not finished training, update the model parameters of the model to be trained and continue to perform self-supervised training on the model to be trained; if it is determined based on the first loss information and the second loss information that the model to be trained has finished training, use the trained model to be trained as a pre-trained model. Then, the pre-trained model can be used to generate a new target video, that is, a new target video with the same label or a similar label as the target video can be generated using the pre-trained model, and thus a small amount of data can support the output of new labels. In addition, since the first loss information between the target image features and the comparison image features and the second loss information between the target input image and the target output image are combined to constrain the model, and the comparison image features are obtained based on the target output image, the training of the model through multi-faceted mutual constraints can improve the generalization ability and performance of the pre-trained model. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present application.
[0053] Figure 1 FIG. is a schematic flowchart of a self-supervised pre-training method based on contrastive learning in a related technology provided by an embodiment of the present application;
[0054] Figure 2 FIG. is a schematic flowchart of a self-supervised pre-training method based on generative learning in a related technology provided by an embodiment of the present application;
[0055] Figure 3 FIG. is a schematic diagram of an implementation environment of a self-supervised training method based on video provided by an embodiment of the present application;
[0056] Figure 4 FIG. is a schematic flowchart of a self-supervised training method based on video provided by an embodiment of the present application;
[0057] Figure 5 Schematic diagram of a training framework for a video-based self-supervised training method provided by an embodiment of the present application;
[0058] Figure 6 Schematic diagram of occluding the same positions of image samples in different frames provided by an embodiment of the present application;
[0059] Figure 7 Schematic diagram of occluding different positions of image samples in different frames provided by an embodiment of the present application;
[0060] Figure 8 Schematic diagram of a framework of a feature extraction module provided by an embodiment of the present application;
[0061] Figure 9 Schematic diagram of another framework of a model to be trained provided by an embodiment of the present application;
[0062] Figure 10 Schematic diagram of the structure of a video-based self-supervised training device provided by an embodiment of the present application;
[0063] Figure 11 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0064] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0065] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the art of the present technology. It should be understood that when an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include a wireless connection or a wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" indicates being implemented as "A", or being implemented as "B", or being implemented as "A and B". "Plural" means not less than two.
[0066] To make the objectives, technical solutions, and advantages of this application more clear, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0067] First, several terms related to this application are introduced and explained:
[0068] Feature-Reconstruction-loss: The local information is reconstructed through an encoder-decoder network (encoding-decoding), and then the global information is generated through the decoder. Specifically, the encoder encodes the global information into a low-dimensional vector, and then the decoder reconstructs the local information into global information. The reconstructed global information is encoded again through the encoder, and trained using the contrastive loss in the feature space to minimize the similarity between the features of the reconstructed global image and the local image. This method can not only improve the model's understanding and modeling ability of global information, but also promote the extraction and utilization of local information by the model. Compared with traditional self-supervised methods, this method can better utilize the relationship between local information and global information, thereby improving the generalization ability and performance of the model.
[0069] Image-Reconstruction-loss: The local information is reconstructed through an encoder-decoder network (encoding-decoding), and then supervised by pixel-level contrastive loss to minimize the pixel reconstruction loss.
[0070] In the related art, although self-supervised pre-training methods have achieved considerable success in the field of machine learning, there are relatively few self-supervised pre-training methods based on videos. Usually, image-based self-supervised pre-training methods, such as MOCO, BYOL, and MOBY, etc., or pre-trained models of 2D pictures are used to initialize the 3D video backbone. The advantages of these methods are that they have been widely studied and verified, and have reliable performance. In addition, self-supervised pre-training methods based on text and images are also very mature, and can be roughly divided into two categories: one is the method based on contrastive learning, and the other is the method based on generative learning. These methods can effectively improve the model's understanding and learning ability of data, thereby achieving better performance in various tasks. However, there are still some challenges and difficulties in self-supervised pre-training methods based on videos, such as how to handle complex dynamics and changes in videos, etc.
[0071] Please refer to Figure 1 , Figure 1A flowchart of a self-supervised pre-training method based on contrastive learning in a related art provided by an embodiment of the present application.
[0072] As Figure 1 shown, the self-supervised pre-training method based on contrastive learning: This method defines positive and negative sample pairs and uses the distance relationship in the representation space to learn the representations of images or videos. In the case of no classification labels, this method constructs positive sample pairs by performing different data augmentations on the same sample and uses different samples as negative sample pairs. By maximizing the distance between positive sample pairs and minimizing the distance between negative sample pairs, the effect of attracting similar classes and repelling different classes is achieved. Methods such as MOCO, BYOL, and MOBY are typical representatives of the self-supervised pre-training method based on contrastive learning. These methods have achieved very good results in the processing tasks of natural images and videos and have been widely applied in various fields. In addition, the advantage of this method is that it can effectively improve the model's ability to understand and learn data, thus providing a better foundation for subsequent supervised learning tasks.
[0073] Please refer to Figure 2 , Figure 2 A flowchart of a self-supervised pre-training method based on generative learning in a related art provided by an embodiment of the present application.
[0074] The self-supervised pre-training method based on generative learning aims to improve the model's performance by reconstructing the input, making the input and output as similar as possible. These methods include Variational AutoEncoder (VAE) and Generative Adversarial Networks (GAN), etc. In the field of Natural Language Processing (NLP), the more common mode is masked autoencoding, which is also widely used in the visual field. The generative learning method is implemented by defining an encoder and a decoder. A certain proportion of partial frames or blocks in the picture or video are randomly masked, and then the original picture is reconstructed through the encoder-decoder, allowing the model to predict the masked content. This method uses the data itself as supervision and does not require complex manual annotation, so it is an efficient self-supervised learning method. Some specific generative self-supervised pre-training methods include BEIT, MAE, and MaskFeat.
[0075] Taking MAE as an example, this method adopts an effective self-supervised pre-training strategy. The model is trained by randomly masking some image patches in the picture and reconstructing the masked pixels. The network structure consists of two parts: an encoder and a decoder. The encoder only encodes the visible (unmasked) image patches, while the decoder takes all image patches as input. At the same time, the decoder can adopt a relatively lightweight structure (for example, only a few layers or even 1 layer are required), while the encoder is usually a multi-layer stacked Transformer. The advantage of this method is that it does not require any manually labeled data because it can utilize the information of the image itself for self-supervised learning.
[0076] However, for the self-supervised scheme based on contrastive learning in the related technology, contrastive learning, as a self-supervised pre-training method, pays more attention to semantic-level features, which is very effective for image-level self-supervised pre-training. However, when the task extends to the pre-training of 3D video frames, this method will show its limitations. First, compared with images, the semantic information in videos is more abundant. The idea of contrastive learning often cannot accurately learn and process the details of videos, making it difficult to achieve fine-grained video content learning. Second, if the model parameters of the 2D pre-trained model are used to initialize the model of 3D video frames, part of the temporal dimension information is often ignored, and it is impossible to better understand the video. Therefore, a method more suitable for video pre-training is needed, which can make full use of the temporal information of the video to better understand and analyze the video, which is crucial for the development of the video processing field.
[0077] For the self-supervised scheme of generative learning in the related technology, generative learning, as a self-supervised pre-training method, pays more attention to image-level features. It learns the features of the samples by mapping the input samples to a high-dimensional space and then restoring the original image through a decoder. However, this method has two main limitations. First, there is a lot of noise and invalid pixels in video data. The pixel-level information restored by the decoder may fit with this noise, resulting in being not conducive to learning the pixel-level features of the video. Second, it is difficult for this method to learn the semantic representation of the video, so it is relatively difficult to converge.
[0078] Specifically, with the explosion of online videos and the continuous progress of neural network technology, videos have become the most important information dissemination medium. A video contains a large amount of information, and more and more researchers are devoting themselves to video analysis and research. How to extract and utilize effective information from massive video data is an important topic in the field of deep learning research and a difficult problem that both academia and industry are concerned about. Providing a pre-trained model suitable for downstream tasks can not only improve the performance of the model, but also shorten the iteration cycle of downstream tasks, thus better adapting to the rapidly changing market demands.
[0079] Currently, a huge amount of video data floods into major video applications every day. For example, in news applications, downstream tasks require identifying and recommending tags for all videos, which is a task that needs to process a large amount of unlabeled data. Labeling a noise-free and substantial dataset consumes a huge amount of manpower and financial resources. To enable the model to perform optimally on new data and continuously support new tags, as little labeled data as possible needs to be used. How to provide a model with strong generalization ability to reduce the iteration times and training data volume of downstream models is a problem that both academia and industry are currently concerned about. By using a pre-trained model, the downstream task model can be trained quickly, the recognition performance can be improved, and it also helps to adapt to new tags faster, enabling the output of new tags to be easily supported with a small amount of data.
[0080] The pre-training method proposed in this application can be applied not only to downstream multi-label recognition, but also to video understanding tasks such as video classification and video retrieval.
[0081] In view of the above at least one technical problem or area for improvement in the related art, the present application proposes a video-based self-supervised training method and related device. The solution includes obtaining at least two first image samples in a video sample; performing self-supervised training on a model to be trained based on the at least two first image samples: extracting features of a target input image through a feature extraction module in the model to be trained to obtain target image features; reconstructing an image of the target image features through an image reconstruction module in the model to be trained to obtain a target output image; and extracting features of the target output image through the feature extraction module to obtain comparison image features, where the target input image is obtained based on the at least two first image samples; determining first loss information between the target image features and the comparison image features, and determining second loss information between the target input image and the target output image; if it is determined based on the first loss information and the second loss information that the model to be trained has not finished training, updating the model parameters of the model to be trained and continuing to perform self-supervised training on the model to be trained; if it is determined based on the first loss information and the second loss information that the model to be trained has finished training, using the trained model to be trained as a pre-trained model to process a target video to obtain a target video processing result, which can achieve the output of new labels with a small amount of data. In addition, compared with the self-supervised training solutions in the related art, the solution of the embodiments of the present application can also enable the pre-trained model to have better generalization ability and performance.
[0082] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0083] The embodiments of the present application may relate to the field of Artificial Intelligence (AI).
[0084] Artificial Intelligence is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, Artificial Intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial Intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0085] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0086] The technical solution of this embodiment particularly relates to computer vision technology, such as processing video samples.
[0087] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing graphic processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0088] It should be noted that in the optional embodiments of this application, for relevant data such as object information (such as video sample data of the application installed on the terminal used by the object), when the embodiments in this application are applied to specific products or technologies, object permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments in this application involve data related to the object, it needs to be obtained under the authorization and consent of the object, the authorization and consent of the relevant department, and compliance with the relevant laws, regulations, and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject needs to be obtained, and the embodiments also need to be implemented under the authorization and consent of the object.
[0089] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the implementation environment of a video-based self-supervised training method provided by an embodiment of the present application. As Figure 3 shown, among them, the terminal 310 communicates with the server 320 through a network. The data storage system can store the data that the server 320 needs to process. The data storage system can be integrated on the server 320, or placed in the cloud or on other servers 320.
[0090] The video-based self-supervised training method is executed independently by the terminal 310 or the server 320, or jointly executed by the terminal 310 and the server 320. In some embodiments, the video-based self-supervised training method is executed by the server 320. The server 320 obtains at least two first image samples in the video sample; performs self-supervised training on the model to be trained based on the at least two first image samples. When performing self-supervised training, the feature extraction module in the model to be trained is used to extract features from the target input image to obtain target image features; and the image reconstruction module in the model to be trained is used to reconstruct the target image features to obtain a target output image; and the feature extraction module is used to extract features from the target output image to obtain comparison image features; wherein, the target input image is obtained based on the at least two first image samples; determine the first loss information between the target image features and the comparison image features, and determine the second loss information between the target input image and the target output image; if it is determined based on the first loss information and the second loss information that the model to be trained has not finished training, update the model parameters of the model to be trained, and continue to perform self-supervised training on the model to be trained; if it is determined based on the first loss information and the second loss information that the model to be trained has finished training, use the trained model to be trained as a pre-trained model to process the target video to obtain a target video processing result.
[0091] Among them, the terminal 310 can be but is not limited to at least one of various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices or portable wearable devices, etc. The Internet of Things device can be at least one of a smart speaker, a smart TV, a smart air conditioner or a smart vehicle-mounted device, etc. The portable wearable device can be at least one of a smart watch, a smart bracelet or a head-mounted device, etc.
[0092] The server 320 can be an independent physical server 320, or a server cluster or distributed system composed of multiple physical servers 320, or a cloud server 320 that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 310 can be at least one of a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, or a smart watch, etc., but is not limited thereto. The terminal 310 and the server 320 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0093] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a self-supervised training method based on video provided by an embodiment of this application. As Figure 4 shown, the method can be executed by an electronic device, and the electronic device can include at least one of a server or a terminal. As Figure 4 shown, the method can include:
[0094] S410. Obtain at least two first image samples in the video sample.
[0095] Among them, the video sample can refer to a sample used for self-supervised training. In this embodiment, the video sample can be sourced from various applications, such as news applications, etc. At least two first image samples can refer to some frames or all frame image samples in the video sample, that is, all frame images in the video sample can be used as at least two first image samples; or the images in the video sample can be frame-extracted to obtain at least two first image samples in the video sample. Among them, the frame extraction process can be to extract one frame image every n frame images, where n is a natural number. In addition, the frame extraction process can also be to select some consecutive frame images in the video sample as at least two first image samples.
[0096] S420. Perform self-supervised training on the model to be trained based on at least two first image samples.
[0097] Among them, the model to be trained can be a pre-built model framework. In this embodiment, the model framework of the model to be trained can be built according to needs, and no specific limitations are made here.
[0098] Among them, S420. Perform self-supervised training on the model to be trained based on at least two first image samples, which can include:
[0099] S421. Extract features from the target input image through the feature extraction module in the model to be trained to obtain target image features; reconstruct the target image features through the image reconstruction module in the model to be trained to obtain a target output image; and extract features from the target output image through the feature extraction module to obtain comparison image features.
[0100] Among them, the target input image is obtained based on at least two first image samples. In this embodiment, the target input image may include an image sequence obtained based on at least two first image samples, and an image sequence may include one or more images or image blocks.
[0101] In this embodiment, the target output image may match the target output image. The target output image may match the target output image, which may mean that at least one of the number of images or image blocks of the target output image is the same as that of the target output image or the size of the images or image blocks of the target output image is the same as that of the target output image.
[0102] S422. Determine the first loss information between the target image features and the comparison image features, and determine the second loss information between the target input image and the target output image.
[0103] S423. If it is determined based on the first loss information and the second loss information that the model to be trained has not finished training, update the model parameters of the model to be trained and continue the self-supervised training of the model to be trained.
[0104] S424. If it is determined based on the first loss information and the second loss information that the model to be trained has finished training, use the trained model to be trained as a pre-trained model.
[0105] Optionally, the first loss information may include a first loss value, and the second loss information may include a second loss value. The method for determining that the model to be trained has finished training based on the first loss information and the second loss information may be that the first loss value is less than a preset first loss threshold, and the second loss value is less than a second loss threshold of the element sum, then it is determined that the model to be trained has finished training; if at least one of the first loss value is less than the preset first loss threshold or the second loss value is less than the second loss threshold of the element sum is not satisfied, continue to train the model to be trained.
[0106] In this embodiment, after using the trained model to be trained as a pre-trained model, the target video can be processed by the pre-trained model to obtain the target video processing result. Optionally, processing the target video may involve extracting features from the target video or at least two target images in the target video to obtain target video features, and then reconstructing the target video features to obtain a target reconstructed video. The target video processing result may include the target reconstructed video. The target reconstructed video can be regarded as a new target video.
[0107] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the training framework of a video-based self-supervised training method provided by an embodiment of the present application.
[0108] As Figure 5 shown, the model to be trained may include a feature extraction module and an image reconstruction module. As Figure 5 shown, during self-supervised training, the feature extraction module in the model to be trained extracts features from the target input image to obtain target image features; the image reconstruction module in the model to be trained reconstructs the target image features to obtain a target output image; and the feature extraction module extracts features from the target output image to obtain comparison image features; then the first loss information between the target image features and the comparison image features is determined, and the second loss information between the target input image and the target output image is determined; furthermore, based on the first loss information and the second loss information, it is determined whether the model to be trained has completed training. If the training is completed, the trained model to be trained is used as a pre-trained model.
[0109] The technical solution of this embodiment is to obtain at least two first image samples in the video sample; perform self-supervised training on the model to be trained based on at least two first image samples. When performing self-supervised training on the model to be trained based on at least two first image samples, feature extraction is performed on the target input image through the feature extraction module in the model to be trained to obtain target image features; the target image features are reconstructed into a target output image through the image reconstruction module in the model to be trained; and feature extraction is performed on the target output image through the feature extraction module to obtain comparison image features, where the target input image is obtained based on at least two first image samples; determine the first loss information between the target image features and the comparison image features, and determine the second loss information between the target input image and the target output image; if it is determined based on the first loss information and the second loss information that the model to be trained has not finished training, update the model parameters of the model to be trained and continue to perform self-supervised training on the model to be trained; if it is determined based on the first loss information and the second loss information that the model to be trained has finished training, use the trained model to be trained as a pre-trained model. Then, the pre-trained model can be used to generate a new target video, that is, a new target video with the same label or a similar label as the target video can be generated using the pre-trained model, and thus a small amount of data can support the output of new labels. In addition, since the first loss information between the target image features and the comparison image features and the second loss information between the target input image and the target output image are combined to constrain the model, and the comparison image features are obtained based on the target output image, the training of the model through mutual constraints in multiple aspects can improve the generalization ability and performance of the pre-trained model.
[0110] In a possible implementation manner, the method for obtaining the target input image based on at least two first image samples includes at least one of the following:
[0111] Use at least two first image samples as the target input image;
[0112] Divide at least two first image samples into blocks to obtain a plurality of first image sample blocks, and use the plurality of first image sample blocks as the target input image.
[0113] In this embodiment, a plurality of first image sample blocks can be input into the model to be trained as an image block sequence for training. Specifically, in the image block sequence, the plurality of first image sample blocks are arranged in the playing order of the frame image samples to which they belong in the video sample. For example, the earlier the playing order of the frame image sample in the video sample, the earlier the corresponding first image sample block of the frame image sample is arranged in the image block sequence.
[0114] The technical solution of this embodiment can select at least two first image samples as the target input image as needed, or block at least two first image samples to obtain multiple first image sample blocks, and use the multiple first image sample blocks as the target input image, which can improve the flexibility of model training.
[0115] In a possible implementation, the target image features are reconstructed into a target output image through the image reconstruction module in the model to be trained, including:
[0116] The target image features are reconstructed at the pixel level through the image reconstruction module in the model to be trained to obtain a target output image that matches the target input image;
[0117] Among them, if the target input image includes at least two first image samples, the target output image includes at least two second image samples; if the target input image includes multiple first image sample blocks, the target input image includes multiple second image sample blocks.
[0118] In this embodiment, the target image features are reconstructed at the pixel level through the image reconstruction module in the model to be trained to obtain a target output image that matches the target input image, which can improve the accuracy of model training and thus improve the performance of the obtained pre-trained model.
[0119] In a possible implementation, the image to be extracted includes multiple image sample blocks, and the feature extraction module includes an image occlusion network and a backbone network. The feature extraction module extracts image features from the image to be extracted in the following way:
[0120] The multiple image sample blocks are occluded by the image occlusion network to obtain multiple occluded image sample blocks, and the multiple occluded image sample blocks include unoccluded image pixels;
[0121] The multiple occluded image sample blocks are subjected to image feature extraction through the backbone network to obtain the image features output by the backbone network;
[0122] Among them, if the image to be extracted includes the target input image, the image features output by the image backbone network include the target image features; if the image to be extracted includes the target output image, the image features output by the image backbone network include the comparison image features.
[0123] In this embodiment, the image occlusion network and the backbone network can perform the same processing flow on the target input image and the target output image, and thus obtain the target image features corresponding to the target input image and the comparison image features corresponding to the target output image.
[0124] Specifically, when occluding multiple image sample blocks, it can be done in units of image sample blocks or in units of image pixels, and this is not limited here.
[0125] The technical solution of this embodiment extracts image features from the occluded multiple image sample blocks through a backbone network to obtain the image features output by the backbone network. Among them, if the image to be extracted includes a target input image, the image features output by the image backbone network include target image features, which can enable the trained pre-trained model to reconstruct a complete new target video for an incomplete target video.
[0126] In a possible implementation, the multiple image sample blocks include at least two image sample blocks corresponding to each frame of image samples. Occluding the multiple image sample blocks through an image occlusion network includes at least one of the following:
[0127] Occlude all the image sample blocks corresponding to at least one frame of image samples and retain all the image sample blocks corresponding to at least one frame of image samples;
[0128] Occlude some of the image sample blocks corresponding to each frame of image samples;
[0129] Randomly select some of the multiple image sample blocks for occlusion.
[0130] In this embodiment, if some of the multiple image sample blocks are randomly selected for occlusion, the generalization ability of the pre-trained model can be improved, enabling the pre-trained model to adapt to the reconstruction of incomplete videos in different situations. If all the image sample blocks corresponding to at least one frame of image samples are occluded and all the image sample blocks corresponding to at least one frame of image samples are retained, the model to be trained can be made to have the feature extraction of at least one frame of complete image samples, thereby improving the video reconstruction performance of the pre-trained model. If some of the image sample blocks corresponding to each frame of image samples are occluded, the model to be trained can learn the context information of at least two frames of image samples of the video sample, which can improve the video reconstruction performance of the pre-trained model.
[0131] It should be noted that when occluding some of the image sample blocks corresponding to each frame of image samples, for the occlusion of different frames of image samples, the same positions of different frames of image samples can be occluded. In addition, for the occlusion of different frames of image samples, different positions of different frames of image samples can be occluded, and the occlusion positions of any two frames of image samples are different.
[0132] Please refer to Figure 6 and Figure 7 . Figure 6Schematic diagram for occluding the same positions of different-frame image samples provided by an embodiment of this application. Figure 7 Schematic diagram for occluding different positions of different-frame image samples provided by an embodiment of this application.
[0133] It can be understood that by occluding the same positions of different-frame image samples, the context semantic information of the same positions between any two-frame image samples can be retained. By occluding different positions of different-frame image samples, the complete image content of the video sample can be retained as much as possible.
[0134] In a possible implementation, image feature extraction is performed on multiple occluded image sample blocks through a backbone network to obtain target image features, including:
[0135] Image feature extraction is performed on at least some of the unoccluded image sample blocks among the multiple image sample blocks through the backbone network to obtain target image features.
[0136] In the technical solution of this embodiment, image feature extraction is performed on at least some of the unoccluded image sample blocks among the multiple image sample blocks through the backbone network to obtain target image features. Then, the backbone network can only perform image feature extraction on at least some of the unoccluded image sample blocks, and no processing is required for the occluded part of the image sample blocks. Moreover, the occluded image sample blocks cannot extract effective information. Thus, without affecting training, the time required for image feature extraction can be reduced, thereby improving the training efficiency.
[0137] In a possible implementation, the backbone network includes at least one level of feature extraction sub-network. Each level of feature extraction sub-network includes a non-convolutional downsampling unit and a feature extraction unit based on an attention mechanism. Performing image feature extraction on multiple occluded image sample blocks through the backbone network includes:
[0138] For each level of feature extraction sub-network, downsampling is performed on the feature extraction sub-network through the non-convolutional downsampling unit of the feature extraction sub-network to obtain a downsampling result, and feature extraction is performed on the downsampling result through the feature extraction unit of the feature extraction sub-network to obtain a feature extraction result output by the feature extraction sub-network.
[0139] In this embodiment, the feature extraction sub-network can be part or all of the 3D Swin Transformer. 3D Swin Transformer is a deep neural network similar to CNN. Its core idea is to use Transformer modules to replace traditional convolutional modules to extract features in the input data. The non-convolutional downsampling can be patch merging, and the feature extraction unit based on the attention mechanism can be a swin transformer block.
[0140] In a possible implementation, the backbone network includes at least two cascaded feature extraction sub-networks. Image feature extraction is performed on multiple occluded image patches through the backbone network, and it further includes:
[0141] The input of the first-level feature extraction sub-network includes the output of the image occlusion network. The input of the nth-level feature extraction sub-network includes the output of the (n - 1)th-level feature extraction sub-network. The output of the nth-level feature extraction sub-network serves as the input of the (n + 1)th-level feature extraction sub-network. The output of the last-level feature extraction sub-network serves as the image features output by the image backbone network;
[0142] where 1 < n < m, and m is the total number of levels of the at least two cascaded feature extraction sub-networks.
[0143] In this embodiment, feature extraction can be performed through multiple levels of feature extraction sub-networks, thereby improving the effect of feature extraction and further enhancing the performance of the pre-trained model.
[0144] Please refer to Figure 8 , Figure 8 which is a schematic framework diagram of a feature extraction module provided by an embodiment of this application.
[0145] The detailed structure of the 3D Swin Transformer can be as Figure 8 shown. The entire network consists of multiple Swin Transformer modules, and each module contains multiple 3D convolutional layers and Transformer blocks. Among them, the 3D convolutional layers are used to extract local features, and the Transformer blocks are used to capture global context information. This structure enables the network to have both local perception and global perception capabilities and can perform information interaction between features at different scales.
[0146] Specifically, in each stage, downsampling is first performed through a Patch Merging layer (except for Stage 1). As shown in the figure below, assume that the input to Patch Merging is a single-channel feature map with a size of 4x4. Patch Merging divides each adjacent 2x2 pixels into a patch, and then stitches together the pixels at the same position (the same color) in each patch to obtain 4 feature maps. Then, these four feature maps are concatenated in the depth direction, and then passed through a LayerNorm layer. Finally, a fully connected layer performs a linear transformation in the depth direction of the feature map, doubling the depth of the feature map from C to C / 2. After passing through the Patch Merging layer, the height and width of the feature map are halved, and the depth is doubled.
[0147] In a possible implementation, the target image features are reconstructed into a target output image through an image reconstruction module in the model to be trained, including:
[0148] The target image features are processed by at least one of upsampling, linear mapping, or nonlinear mapping through an image reconstruction module in the model to be trained, so as to reconstruct the target image features into a target output image.
[0149] The technical solution of this embodiment converts the extracted video feature map into a pixel-level image to realize the reconstruction of the video. The reconstruction of the video is realized by using original pixel-level regression. Since the pixel values are continuous in the original space, the occluded pixel values can be most directly restored through regression prediction. However, usually, the video feature maps learned using the vision framework are downsampled, which results in the loss of some pixel information. In order to be able to perform regression of all pixel values, a reconstruction module needs to be designed to upsample the feature map or map it linearly or nonlinearly back to the size of the original video. This can maximize the retention of original pixel-level information, thereby improving the quality and accuracy of video reconstruction.
[0150] The following embodiments further illustrate the relevant content on how to confirm whether the training is completed based on any of the above embodiments.
[0151] In a possible implementation, the target image features include the first image features corresponding to each first image sample block in a plurality of first image sample blocks, the comparison image features include the second image features corresponding to the second image sample blocks corresponding to each first image sample block, and determining the first loss information between the target image features and the comparison image features includes:
[0152] For each first image sample block, determine first loss sub-information between the first image feature corresponding to the first image sample block and the second image feature corresponding to the second image sample corresponding to the first image sample block;
[0153] Based on the fusion of the first loss sub-information corresponding to each first image sample block, determine the first loss information between the target image feature and the comparison image feature.
[0154] In the technical solution of this embodiment, for each first image sample block, determine first loss sub-information between the first image feature corresponding to the first image sample block and the second image feature corresponding to the second image sample corresponding to the first image sample block;
[0155] Based on the fusion of the first loss sub-information corresponding to each first image sample block, determine the first loss information between the target image feature and the comparison image feature, which can improve the calculation accuracy of the feature loss information, and further improve the performance of the pre-trained model.
[0156] In a possible implementation manner, determining the second loss information between the target input image and the target output image includes:
[0157] For each first image pixel in the target input image, determine the second image pixel corresponding to the first image pixel in the target output image, and determine the second loss sub-information between the first image pixel and the corresponding second image pixel;
[0158] Based on the fusion of the second loss sub-information corresponding to each first image pixel in the target input image, determine the second loss information between the target input image and the target output image.
[0159] In the technical solution of this embodiment, the second image pixel corresponding to the first image pixel, and determine the second loss sub-information between the first image pixel and the corresponding second image pixel;
[0160] Based on the fusion of the second loss sub-information corresponding to each first image pixel in the target input image, determine the second loss information between the target input image and the target output image, which can improve the calculation accuracy of the pixel-level loss, and further improve the performance of the pre-trained model.
[0161] In the following embodiments, on the basis of any of the above embodiments, an example is given where the video sample comes from a news application. Specifically, when the terminal plays the video of the news application, the terminal obtains the video of the news application as the video sample, and then the terminal sends the video sample to the server.
[0162] The server extracts frames from the video sample to obtain at least two first image samples in the video sample, then divides the at least two first image samples into blocks to obtain a plurality of first image sample blocks, and then inputs the plurality of first image sample blocks into the model to be trained. At this time, the image occlusion network occludes the plurality of first image sample blocks to obtain the occluded plurality of first image sample blocks, and the occluded plurality of first image sample blocks include unoccluded image pixels; then the backbone network extracts features from the unoccluded first image sample blocks to obtain target image features, and then the reconstruction module reconstructs the extracted target image features to obtain a plurality of second image sample blocks. Then, second loss information is calculated based on the plurality of second image sample blocks and the plurality of first image sample blocks. At the same time, at least two second image samples can also be divided into blocks to obtain a plurality of second image sample blocks, and then the plurality of second image sample blocks are input into the model to be trained. At this time, the image occlusion network occludes the plurality of second image sample blocks to obtain the occluded plurality of second image sample blocks, and the occluded plurality of second image sample blocks include unoccluded image pixels; then the backbone network extracts features from the unoccluded second image sample blocks to obtain comparison image features, and then the first loss information between the target image features and the comparison image features is determined.
[0163] If it is determined based on the first loss information and the second loss information that the model to be trained has not finished training, a new video sample is obtained from the terminal for continued training. If it is determined based on the first loss information and the second loss information that the model to be trained has finished training, a pre-trained model is obtained.
[0164] For ease of understanding, based on any of the above embodiments, the following embodiments illustrate the technical solutions of the present application in combination with a specific example.
[0165] Please refer to Figure 9 , Figure 9 which is a schematic diagram of the framework of another model to be trained provided by the embodiments of the present application. As Figure 9 shown, among them, the encoder can be used as a feature extraction module, and the decoder can be used as an image reconstruction module.
[0166] The design of the self-supervised pre-training framework based on video proposed in the embodiments of the present application is mainly divided into the design of two modules: video feature reconstruction contrast learning and video pixel reconstruction contrast learning. The overall framework diagram is as Figure 9As shown in the figure. The whole model is divided into three parts: Encoder, Decoder, and Contrastive-loss. First, this framework inputs a video sample, and randomly extracts T frames as the input of the network through a multi-scale temporal input module: x ∈ C * T * H * W (T ∈ {4, 8, 16, 32}), where T is the time length, H and W are the width and height of the image respectively, and C represents the number of pixel values in the image block. As the input of the network, the sample x first undergoes a blocking operation, and the block size is 2 * 4 * 4. After blocking, the size is: Then, through the network Encoder, it is mapped to a high-dimensional space to obtain the feature q. N represents the number of image blocks.
[0167] To achieve pixel-level reconstruction loss supervision, the present invention uses the feature map x output by the backbone in the encoder feature map . The feature map x feature map Then it is reconstructed through the Decoder network, and finally the original video pixels y are obtained. To ensure the consistency between the reconstructed pixels and the original pixels, L2 loss is used for supervision. This method can not only improve the reconstruction quality, but also ensure a high degree of consistency between the reconstructed pixels and the original pixels, providing a better basis for subsequent tasks:
[0168] L image-recontruction loss = ‖y - x‖2;
[0169] where x, are the original video and the reconstructed video respectively.
[0170] To achieve feature-level reconstruction supervision, the present invention takes the following steps: First, input the original video pixels y into the encoder to obtain the reconstructed feature q y . Subsequently, the reconstructed feature is supervised by using L2 loss to achieve feature-level reconstruction:
[0171] L Feature-recontruction loss = ||q y - q||2;
[0172] Next, Encoder, Momentum encoder, and Decoder will be specifically introduced.
[0173] 1. Encoder
[0174] The Encoder is the main part of the self-supervised network. The parameters trained in this part will be used as the initial parameters for downstream tasks, playing a crucial role. It learns high-level semantic features in the dataset and encodes these features into low-dimensional vector representations, providing effective feature expressions for downstream tasks. The Encoder is divided into two modules: Maskblock and Backbone. The functions and detailed implementation methods of each part will be introduced in turn below:
[0175] Mask block. The function of this part is to cover a certain proportion of pixel information in the input feature x. There are two optional methods here: mask block and mask frame. First, for the mask block, it performs block masks on the input video frame x ∈ C*T*H*W in three dimensions. The size of the block is 2*4*4. First, the input video frame is divided into blocks: The size of the block will have a certain impact on the performance of self-supervision. First, randomly initialize the mask matrix x p ∈ [0,1]. The feature x after occlusion is m calculated as follows:
[0176] x m = x p *;
[0177] Backbone. Generally, traditional convolutional neural networks (CNNs) are widely used in image processing tasks. However, with the successful application of Transformer networks in the field of natural language processing, they have also begun to be more and more applied in the field of computer vision.
[0178] Due to the excellent performance of Transformer networks in the visual field, the present invention uses 3D Swin Transformer as the basic network for self-supervised experiments. 3D Swin Transformer is a deep neural network similar to CNN. Its core idea is to use Transformer modules to replace traditional convolutional modules to extract features from input data.
[0179] 3. Decoder
[0180] The role of the Decoder is to convert the video feature maps extracted by the Encoder into pixel-level images for video reconstruction. In contrast, the Encoder is responsible for learning latent feature representations, so it requires a deeper network structure and more channels. The Decoder, on the other hand, can adopt different network structures as long as it can output the same image as the Encoder input and complete the pixel-level reconstruction task. For example, structures such as 2-layer MLP, Transformer, etc. can all be used as the network structure of the Decoder.
[0181] Video reconstruction is achieved using raw pixel-level regression. Since pixel values are continuous in the original space, occlusion pixel values can be most directly restored through regression prediction. However, usually, the video feature maps learned using a vision framework are downsampled, which results in the loss of some pixel information. To perform regression on all pixel values, a Decoder needs to be designed to upsample the feature maps or map them linearly or non-linearly back to the size of the original video. This can maximize the retention of original pixel-level information, thereby improving the quality and accuracy of video reconstruction. Taking SwinTransformer as an example, if the input size of the video is x ∈ C * T * H * W, after chunking (the size of the chunk is 2 * 4 * 4), the input size becomes After passing through the Swin Transformer and the decoder, the output size is The present invention uses l2-loss to perform the reconstruction task on all pixels:
[0182] L Image-recontruction loss = ||x d - x p ||2.
[0183] The main application fields of the technical solution of this embodiment include, but are not limited to, video-based learning and recognition, including multi-label recognition, video classification, video retrieval, and video quality recognition, etc. It provides a pre-trained model adapted to various video-based downstream tasks.
[0184] As shown in Table 1, Table 1 is a schematic table of experimental results provided by an embodiment of the present application.
[0185] The experimental results show that the technical solution of this embodiment can effectively support existing video multi-label recognition projects and has been successfully applied to related services. In addition, in the video classification project, the technical solution of this embodiment has also achieved good results. Taking video tags as an example, a performance comparison was made between not using a pre-trained model and using the video self-supervised pre-training framework proposed by this technology. It was found that using the pre-trained model of this technology can significantly improve the performance of multi-label recognition and has good practical application prospects.
[0186] Table 1
[0187]
[0188] Please refer to Figure 10 , Figure 10 , which is a schematic structural diagram of a video-based self-supervised training device provided by an embodiment of the present application. As Figure 10 shown, the device 1000 may include an acquisition module 1010 and a training module 1020, where:
[0189] The acquisition module 1010 is used to acquire at least two first image samples in a video sample;
[0190] The training module 1020 is used to perform self-supervised training on a model to be trained based on at least two first image samples:
[0191] Feature extraction is performed on a target input image through a feature extraction module in the model to be trained to obtain target image features; and image reconstruction is performed on the target image features through an image reconstruction module in the model to be trained to obtain a target output image; and feature extraction is performed on the target output image through the feature extraction module to obtain comparison image features; wherein, the target input image is obtained based on at least two first image samples;
[0192] Determine the first loss information between the target image features and the comparison image features, and determine the second loss information between the target input image and the target output image;
[0193] If it is determined based on the first loss information and the second loss information that the model to be trained has not finished training, continue to perform self-supervised training on the model to be trained;
[0194] If it is determined based on the first loss information and the second loss information that the model to be trained has finished training, use the trained model to be trained as a pre-trained model to process the target video to obtain a target video processing result.
[0195] Optionally, when the training module 1020 obtains a target input image based on at least two first image samples, it can be used for at least one of the following:
[0196] Use at least two first image samples as the target input image;
[0197] Divide at least two first image samples into blocks to obtain a plurality of first image sample blocks, and use the plurality of first image sample blocks as the target input image.
[0198] Optionally, when the training module 1020 performs image reconstruction on the target image features through the image reconstruction module in the model to be trained to obtain a target output image, it can be used for:
[0199] The target image features are reconstructed at the pixel level through an image reconstruction module in the model to be trained, and a target output image matching the target input image is obtained;
[0200] Among them, if the target input image includes at least two first image samples, the target output image includes at least two second image samples; if the target input image includes multiple first image sample blocks, the target input image includes multiple second image sample blocks.
[0201] Optionally, the image to be extracted includes multiple image sample blocks, and the feature extraction module includes an image occlusion network and a backbone network. The feature extraction module extracts image features from the image to be extracted in the following way by the training module 1020:
[0202] The multiple image sample blocks are occluded through the image occlusion network to obtain the occluded multiple image sample blocks, and the occluded multiple image sample blocks include unoccluded image pixels;
[0203] The image features of the occluded multiple image sample blocks are extracted through the backbone network to obtain the image features output by the backbone network;
[0204] Among them, if the image to be extracted includes the target input image, the image features output by the image backbone network include the target image features; if the image to be extracted includes the target output image, the image features output by the image backbone network include the comparison image features.
[0205] Optionally, the multiple image sample blocks include at least two image sample blocks corresponding to each frame of image sample. When the training module 1020 occludes the multiple image sample blocks through the image occlusion network, it can be used for:
[0206] Occlude all the image sample blocks corresponding to at least one frame of image sample and retain all the image sample blocks corresponding to at least one frame of image sample;
[0207] Occlude some of the image sample blocks corresponding to each frame of image sample;
[0208] Randomly select some of the image sample blocks among the multiple image sample blocks for occlusion.
[0209] Optionally, when the training module 1020 extracts the image features of the occluded multiple image sample blocks through the backbone network to obtain the target image features, it can be used for:
[0210] Extract the image features of at least some of the unoccluded image sample blocks among the multiple image sample blocks through the backbone network to obtain the target image features.
[0211] Optionally, the backbone network includes at least one level of feature extraction sub-network, and each level of feature extraction sub-network includes a non-convolutional down-sampling unit and an attention mechanism-based feature extraction unit. When the training module 1020 performs image feature extraction on multiple occluded image sample blocks through the backbone network, it can be used for:
[0212] For each level of feature extraction sub-network, perform down-sampling on the feature extraction sub-network through the non-convolutional down-sampling unit of the feature extraction sub-network to obtain a down-sampling result, and perform feature extraction on the down-sampling result through the feature extraction unit of the feature extraction sub-network to obtain a feature extraction result output by the feature extraction sub-network.
[0213] Optionally, the backbone network includes at least two cascaded levels of feature extraction sub-networks. When the training module 1020 performs image feature extraction on multiple occluded image sample blocks through the backbone network, it can also be used for:
[0214] The input of the first-level feature extraction sub-network includes the output of the image occlusion network, the input of the nth-level feature extraction sub-network includes the output of the (n - 1)th-level feature extraction sub-network, the output of the nth-level feature extraction sub-network serves as the input of the (n + 1)th-level feature extraction sub-network, and the output of the last-level feature extraction sub-network serves as the image features output by the image backbone network;
[0215] where 1 < n < m, and m is the total number of at least two cascaded levels of feature extraction sub-networks.
[0216] Optionally, when the training module 1020 performs image reconstruction on the target image features through the image reconstruction module in the to-be-trained model to obtain the target output image, it can be used for:
[0217] Perform at least one of up-sampling, linear mapping, or non-linear mapping on the target image features through the image reconstruction module in the to-be-trained model to reconstruct the target image features into the target output image.
[0218] Optionally, the target image features include the first image features corresponding to each first image sample block in multiple first image sample blocks, and the comparison image features include the second image features corresponding to the second image sample blocks corresponding to each first image sample block. When the training module 1020 determines the first loss information between the target image features and the comparison image features, it can be used for:
[0219] For each first image sample block, determine the first loss sub-information between the first image feature corresponding to the first image sample block and the second image feature corresponding to the second image sample corresponding to the first image sample block;
[0220] Fuse based on the first loss sub-information corresponding to each first image sample block to determine the first loss information between the target image feature and the comparison image feature.
[0221] Optionally, when determining the second loss information between the target input image and the target output image, the training module 1020 can be used for:
[0222] For each first image pixel in the target input image, determine the second image pixel in the target output image corresponding to the first image pixel, and determine the second loss sub-information between the first image pixel and the corresponding second image pixel;
[0223] Fuse based on the second loss sub-information corresponding to each first image pixel in the target input image to determine the second loss information between the target input image and the target output image.
[0224] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the method according to the embodiments of the present application. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, and details are not described herein again.
[0225] An electronic device is provided in an embodiment of the present application, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the method according to any embodiment of the present application.
[0226] In an optional embodiment, an electronic device is provided, as Figure 11 shown, Figure 11 The electronic device 1100 shown includes: a processor 1101 and a memory 1103. Among them, the processor 1101 and the memory 1103 are connected, such as connected through a bus 1102. Optionally, the electronic device 1100 may further include a transceiver 1104, and the transceiver 1104 can be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 1104 is not limited to one, and the structure of the electronic device 1100 does not constitute a limitation to the embodiment of the present application.
[0227] The processor 1101 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of this application. The processor 1101 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0228] The bus 1102 may include a path for transmitting information between the above components. The bus 1102 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 1102 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0229] The memory 1103 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.
[0230] The memory 1103 is used to store the computer program for implementing the embodiments of the present application, and is controlled by the processor 1101 to execute. The processor 1101 is used to execute the computer program stored in the memory 1103 to implement the steps shown in the foregoing method embodiments.
[0231] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0232] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0233] It should be understood that although the flowcharts of the embodiments of the present application indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated in this document, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0234] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, other similar implementation means based on the technical idea of the present application also belong to the protection scope of the embodiments of the present application.
Claims
1. A video-based self-supervised training method, characterized in that, Including: Obtaining at least two first image samples in a video sample; Performing self-supervised training on a model to be trained based on the at least two first image samples: Performing feature extraction on a target input image through a feature extraction module in the model to be trained to obtain target image features; and performing image reconstruction on the target image features through an image reconstruction module in the model to be trained to obtain a target output image; and performing feature extraction on the target output image through the feature extraction module to obtain comparison image features; wherein the target input image is obtained based on the at least two first image samples; Determining first loss information between the target image features and the comparison image features, and determining second loss information between the target input image and the target output image; If it is determined based on the first loss information and the second loss information that the training of the model to be trained has not ended, updating the model parameters of the model to be trained and continuing the self-supervised training of the model to be trained; If it is determined based on the first loss information and the second loss information that the training of the model to be trained has ended, using the trained model to be trained as a pre-trained model to process a target video to obtain a target video processing result.
2. The method according to claim 1, wherein The method for obtaining the target input image based on the at least two first image samples includes at least one of the following: Using the at least two first image samples as the target input image; Dividing the at least two first image samples into blocks to obtain a plurality of first image sample blocks, and using the plurality of first image sample blocks as the target input image.
3. The method according to claim 2, wherein The step of performing image reconstruction on the target image features through the image reconstruction module in the model to be trained to obtain a target output image includes: Performing pixel-level image reconstruction on the target image features through the image reconstruction module in the model to be trained to obtain a target output image that matches the target input image; Wherein, if the target input image includes at least two first image samples, the target output image includes at least two second image samples; if the target input image includes a plurality of first image sample blocks, the target input image includes a plurality of second image sample blocks.
4. The method according to claim 2, wherein The image to be extracted includes a plurality of image sample blocks, the feature extraction module includes an image occlusion network and a backbone network, and the feature extraction module extracts image features from the image to be extracted in the following manner: Occluding the plurality of image sample blocks through the image occlusion network to obtain a plurality of occluded image sample blocks, and the plurality of occluded image sample blocks include unoccluded image pixels; Performing image feature extraction on the plurality of occluded image sample blocks through the backbone network to obtain the image features output by the backbone network; Among them, if the image to be extracted includes the target input image, the image features output by the image backbone network include the target image features; if the image to be extracted includes the target output image, the image features output by the image backbone network include the comparison image features.
5. The method according to claim 4, characterized in that, The multiple image sample blocks include at least two image sample blocks corresponding to each frame of image sample. The occlusion of the multiple image sample blocks by the image occlusion network includes at least one of the following: Occlude all the image sample blocks corresponding to at least one frame of image sample and retain all the image sample blocks corresponding to at least one frame of image sample; Occlude some of the image sample blocks corresponding to each frame of image sample; Randomly select some of the multiple image sample blocks for occlusion.
6. The method according to claim 5, wherein The extraction of image features from the occluded multiple image sample blocks by the backbone network to obtain target image features includes: Extract image features from at least some of the unoccluded image sample blocks among the multiple image sample blocks by the backbone network to obtain target image features.
7. The method according to claim 4, characterized in that The backbone network includes at least one level of feature extraction sub-network, and each level of feature extraction sub-network includes a non-convolutional downsampling unit and a feature extraction unit based on an attention mechanism. The extraction of image features from the occluded multiple image sample blocks by the backbone network includes: For each level of feature extraction sub-network, perform downsampling on the feature extraction sub-network through the non-convolutional downsampling unit of the feature extraction sub-network to obtain a downsampling result, and perform feature extraction on the downsampling result through the feature extraction unit of the feature extraction sub-network to obtain the feature extraction result output by the feature extraction sub-network.
8. The method according to claim 7, characterized in that, The backbone network includes at least two levels of cascaded feature extraction sub-networks. The extraction of image features from the occluded multiple image sample blocks by the backbone network further includes: The input of the first-level feature extraction sub-network includes the output of the image occlusion network, the input of the nth-level feature extraction sub-network includes the output of the (n - 1)th-level feature extraction sub-network, the output of the nth-level feature extraction sub-network is used as the input of the (n + 1)th-level feature extraction sub-network, and the output of the last-level feature extraction sub-network is used as the image features output by the image backbone network; Among them, 1 < n < m, and m is the total number of levels of the at least two levels of cascaded feature extraction sub-networks.
9. The method according to any one of claims 1 - 8, characterized in that, The reconstruction of the target image features by the image reconstruction module in the model to be trained to obtain the target output image includes: Perform at least one of upsampling, linear mapping, or non-linear mapping on the target image features through the image reconstruction module in the model to be trained to reconstruct the target image features into the target output image.
10. The method according to any one of claims 1-8, characterized in that, The target image features include first image features corresponding to each of the multiple first image sample blocks, the comparison image features include second image features corresponding to second image sample blocks corresponding to each of the first image sample blocks, and determining the first loss information between the target image features and the comparison image features includes: For each first image sample block, determining first loss sub-information between the first image feature corresponding to the first image sample block and the second image feature corresponding to the second image sample corresponding to the first image sample block; Based on the fusion of the first loss sub-information corresponding to each of the first image sample blocks, determining the first loss information between the target image features and the comparison image features.
11. The method according to any one of claims 1-8, characterized in that, Determining the second loss information between the target input image and the target output image includes: For each first image pixel in the target input image, determining a second image pixel corresponding to the first image pixel in the target output image, and determining second loss sub-information between the first image pixel and the corresponding second image pixel; Based on the fusion of the second loss sub-information corresponding to each of the first image pixels in the target input image, determining the second loss information between the target input image and the target output image.
12. A video-based self-supervised training device, characterized in that, Including: An acquisition module, configured to acquire at least two frames of first image samples in a video sample; A training module, configured to perform self-supervised training on a model to be trained based on the at least two frames of first image samples: Performing feature extraction on a target input image through a feature extraction module in the model to be trained to obtain target image features; and performing image reconstruction on the target image features through an image reconstruction module in the model to be trained to obtain a target output image; and performing feature extraction on the target output image through the feature extraction module to obtain comparison image features; wherein, the target input image is obtained based on the at least two frames of first image samples; Determining the first loss information between the target image features and the comparison image features, and determining the second loss information between the target input image and the target output image; If it is determined based on the first loss information and the second loss information that the model to be trained has not finished training, then continue to perform self-supervised training on the model to be trained; If it is determined based on the first loss information and the second loss information that the model to be trained has finished training, then use the trained model to be trained as a pre-trained model to process a target video to obtain a target video processing result.
13. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-11.
15. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-11.