Video Processing Model Training Method, Apparatus, and Device

Through the joint supervision method of visual semantics and motion transformation, the semantics and motion information of video frames are predicted using the de-entanglement decoder, which solves the problem of visual and motion decoupling in video processing model training, achieving more efficient training and better downstream task performance.

CN116310643BActive Publication Date: 2025-07-25BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
CN202310271487.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-17
Publication Date
2025-07-25
Estimated Expiration
2043-03-17

AI Technical Summary

Technical Problem

In the training of video processing models, it is difficult to effectively combine visual semantics and motion transformation for joint supervision, resulting in high training costs and insufficient downstream tasks.

Method used

Using a joint supervision method based on visual semantics and motion transformation, two de-entanglement decoders are used to predict semantic codebooks and inter-frame motion changes of mask video frames, and space-time feature extraction is enhanced through encoder learning, combining loss function optimization in pre-training and fine-tuning stages.

Benefits of technology

Reduces training costs and significantly improves the performance of video downstream tasks such as image classification, object detection, semantic segmentation and action recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310643B_ABST
    Figure CN116310643B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, and device for training a video processing model, relating to the field of artificial intelligence technologies, specifically to technical fields such as computer vision, augmented reality, virtual reality, deep learning, etc. A specific implementation of this method includes: obtaining a masked video frame, where the masked video frame includes a visible video block and a masked video block; inputting the masked video frame into an encoder to learn the features of the masked video frame; respectively inputting the features of the masked video frame into a visual decoder and a motion decoder to predict the visual codebook and hidden motion information of the masked video frame; calculating a loss based on the visual codebook and the hidden motion information; adjusting the parameters of the encoder, the visual decoder, and the motion decoder based on the loss to obtain a video processing model. This implementation provides a training method for a video processing model based on joint supervision of visual semantics and motion transformation, which not only reduces the training cost but also significantly improves the performance in video downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to technical fields such as computer vision, augmented reality, virtual reality, and deep learning. Background Art

[0002] Visual pre-training is an important research direction and application point in current deep learning. For large-scale visual pre-training schemes, first perform pre-training on a large-scale visual dataset using unsupervised learning, and then fine-tune on specific downstream tasks, which can achieve excellent results in a series of downstream tasks such as image classification, object detection, semantic segmentation, and action recognition. Summary of the Invention

[0003] Embodiments of the present disclosure propose a method, device, equipment, storage medium, and program product for training a video processing model.

[0004] In a first aspect, embodiments of the present disclosure propose a method for training a video processing model, including: obtaining a masked video frame, where the masked video frame includes a visible video block and a masked video block; inputting the masked video frame into an encoder to learn the features of the masked video frame; respectively inputting the features of the masked video frame into a visual decoder and a motion decoder to predict the visual codebook and hidden motion information of the masked video frame; calculating a loss based on the visual codebook and the hidden motion information; and adjusting the parameters of the encoder, the visual decoder, and the motion decoder based on the loss to obtain a video processing model.

[0005] In a second aspect, embodiments of the present disclosure propose a video processing method, including: obtaining a video to be processed for a target task, where the target task includes at least one of the following: image classification, object detection, semantic segmentation, and action recognition; inputting the video to be processed into the video processing model to obtain a processing result of the target task for the video to be processed, where the video processing model is trained by the method described in the first aspect.

[0006] In a third aspect, embodiments of the present disclosure propose a video processing model training device, including: a first acquisition module configured to obtain a masked video frame, where the masked video frame includes a visible video block and a masked video block; an encoding module configured to input the masked video frame into an encoder to learn the features of the masked video frame; a decoding module configured to respectively input the features of the masked video frame into a visual decoder and a motion decoder to predict the visual codebook and hidden motion information of the masked video frame; a first calculation module configured to calculate a loss based on the visual codebook and the hidden motion information; and a first training module configured to adjust the parameters of the encoder, the visual decoder, and the motion decoder based on the loss to obtain a video processing model.

[0007] Fourthly, an embodiment of the present disclosure provides a video processing apparatus, including: an acquisition module configured to acquire a video to be processed for a target task, where the target task includes at least one of the following: image classification, object detection, semantic segmentation, and action recognition; a processing module configured to input the video to be processed into a video processing model to obtain a processing result of the target task for the video to be processed, where the video processing model is trained by using the apparatus described in the third aspect.

[0008] Fifthly, an embodiment of the present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect or the second aspect.

[0009] Sixthly, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in any implementation manner of the first aspect or the second aspect.

[0010] Seventhly, an embodiment of the present disclosure provides a computer program product including a computer program, where the computer program, when executed by a processor, implements the method described in any implementation manner of the first aspect or the second aspect.

[0011] The embodiment of the present disclosure provides a method for training a video processing model based on joint supervision of visual semantics and motion transformation. Two disentangled decoders are used simultaneously to predict the semantic codebook and the inter-frame motion change in the masked video frames to decouple and reconstruct the visual and motion representations. Moreover, by jointly predicting the semantic codebook of the masked video frames and the inter-frame motion change, it is also possible to promote the encoder to obtain a strong spatio-temporal feature extraction ability, learn a more robust and general spatio-temporal video representation, and enable the encoder to be more efficiently migrated to video-related downstream tasks. This not only reduces the training cost but also significantly improves the performance in video downstream tasks.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. Description of the Drawings

[0013] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present disclosure will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0014] Figure 1It is a flowchart of an embodiment of a video processing model training method according to the present disclosure;

[0015] Figure 2 It is a flowchart of another embodiment of a video processing model training method according to the present disclosure;

[0016] Figure 3 It is a scenario diagram that can implement the video processing model training method of the embodiments of the present disclosure;

[0017] Figure 4 It is a flowchart of an embodiment of a video processing method according to the present disclosure;

[0018] Figure 5 It is a schematic structural diagram of an embodiment of a video processing model training apparatus according to the present disclosure;

[0019] Figure 6 It is a schematic structural diagram of an embodiment of a video processing apparatus according to the present disclosure;

[0020] Figure 7 It is a block diagram of an electronic device for implementing the video processing model training method or the video processing method of the embodiments of the present disclosure. Specific Embodiments

[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0022] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0023] Figure 1 Flow 100 of an embodiment of a video processing model training method according to the present disclosure is shown. The video processing model training method includes the following steps:

[0024] Step 101, obtaining a masked video frame.

[0025] In this embodiment, the execution subject of the video processing model training method can obtain a masked video frame.

[0026] Generally, to obtain video frames in a video, a masked video prediction framework is adopted. By masking some video blocks in the video frames, masked video frames can be obtained. Among them, the masked video frames can include visible video blocks and masked video blocks. Visible video blocks are the video blocks in the video frames that are not masked. Masked video blocks can be the video blocks in the video frames that are masked. In some embodiments, a preset proportion (such as 75%) of the video blocks in the video frames are randomly masked to obtain masked video frames.

[0027] Step 102: Input the masked video frames into an encoder to learn the features of the masked video frames.

[0028] In this embodiment, the above-mentioned execution entity can input the masked video frames into an encoder to learn the features of the masked video frames.

[0029] Among them, the encoder can compress the input into a latent space representation. Here, the encoder not only learns the features of the visible video blocks in the masked video frames, but also learns the features of the masked video blocks in the masked video frames. Moreover, the features learned by the encoder can also include the inherent spatio-temporal relationships in the video data.

[0030] Step 103: Input the features of the masked video frames into a visual decoder and a motion decoder respectively to predict the visual codebook and hidden motion information of the masked video frames.

[0031] In this embodiment, the above-mentioned execution entity can input the features of the masked video frames into a visual decoder to predict the visual codebook of the masked video frames, and input the features of the masked video frames into a motion decoder to predict the hidden motion information of the masked video frames.

[0032] Here, two disentangled decoders are used simultaneously to predict the reconstructed visual appearance of the video frames and the motion information in the video frames. Among them, the visual decoder can be used to predict the visual codebook of the video frames, and the motion decoder can be used to predict various motion information hidden in the video data, so as to provide supplementary semantic clues in the pre-training stage. Moreover, using two disentangled decoders can decouple the appearance and motion views at the object level, can effectively process redundant spatio-temporal data, and accelerate the convergence speed in the pre-training stage. In addition, through the two disentangled decoders, the encoder can also be forced to learn the inherent spatio-temporal relationships in the video data.

[0033] Step 104: Calculate the loss based on the visual codebook and the hidden motion information.

[0034] In this embodiment, the above-mentioned execution entity can calculate the loss based on the visual codebook and the hidden motion information.

[0035] Generally, for the unmasked original video frames corresponding to the masked video frames, the true visual codebook and motion information can be obtained. Based on the differences between the true visual codebook and the predicted visual codebook, and between the true motion information and the predicted motion information, the loss can be calculated. For example, by inputting the true visual codebook and the predicted visual codebook into the corresponding loss function, a loss can be calculated. By inputting the true motion information and the predicted motion information into the corresponding loss function, another loss can be obtained. By performing a weighted sum of the above two losses, the final total loss can be obtained.

[0036] Step 105: Based on the loss, adjust the parameters of the encoder, visual decoder, and motion decoder to obtain a video processing model.

[0037] In this embodiment, the above-mentioned execution subject can adjust the parameters of the encoder, visual decoder, and motion decoder based on the loss to obtain a video processing model.

[0038] Generally, based on the loss, adjust the parameters of the above encoder, visual decoder, and motion decoder until convergence, that is, complete the training of the video processing model.

[0039] In some embodiments, in order to improve the processing effect of the video processing model on video downstream tasks, after the video processing model is trained, a training sample set of the target task can also be obtained to continue training the video processing model. Specifically, use the sample video as the input and the target task processing result of the sample video as the output to continue training the video processing model. Among them, the training sample set can be a small labeled video data set of the target task, and each training sample can include the sample video and the target task processing result of the sample video. Use these labeled video data sets to fine-tune the parameters of the encoder, visual decoder, and motion decoder in the video processing model, so that the video processing model can achieve excellent results in the target task. Among them, the target task is the video downstream task, which can include but is not limited to image classification, object detection, semantic segmentation, action recognition, etc.

[0040] The embodiments of the present disclosure provide a method for training a video processing model based on joint supervision of visual semantics and motion transformation. Two disentangled decoders are used simultaneously to predict the semantic codebook and inter-frame motion changes in the masked video frames to decouple and reconstruct visual and motion representations. Moreover, by jointly predicting the semantic codebook and inter-frame motion changes of the masked video frames, it is also possible to promote the encoder to obtain a strong spatio-temporal feature extraction ability, learn a more robust and generalized spatio-temporal video representation, and make the encoder more efficiently transferable to video-related downstream tasks. This not only reduces the training cost but also significantly improves the performance in video downstream tasks.

[0041] Continue to refer to Figure 2, which shows a flow 200 of another embodiment of the video processing model training method according to the present disclosure. The video processing model training method includes the following steps:

[0042] Step 201, obtain a masked video frame.

[0043] In this embodiment, the specific operation of step 201 has been described in detail in step 101 of the embodiment shown in Figure 1 and will not be elaborated here.

[0044] Step 202, input the visible video blocks into the first encoder to learn the features of the visible video blocks.

[0045] In this embodiment, the execution subject of the video processing model training method may input the visible video blocks into the first encoder to learn the features of the visible video blocks. Among them, the first encoder can compress the input into a latent space representation. Here, the first encoder learns the features of the visible video blocks in the masked video frame.

[0046] Step 203, input the masked video blocks into the second encoder to learn the features of the masked video blocks.

[0047] In this embodiment, the above-mentioned execution subject may input the masked video blocks into the second encoder to learn the features of the masked video blocks. Among them, the second encoder can compress the input into a latent space representation. Here, the second encoder learns the features of the masked video blocks in the masked video frame.

[0048] Step 204, input the features of the visible video blocks into the latent variable regressor to obtain the predicted features of the masked video blocks.

[0049] In this embodiment, the above-mentioned execution subject may input the features of the visible video blocks into the latent variable regressor to obtain the predicted features of the masked video blocks.

[0050] To eliminate the difference between the appearance view and the motion view, a latent variable regressor is used to further map the latent representation output by the encoder. Specifically, the latent variable regressor can query from the embedding of the visible video blocks and predict the representation of the masked video blocks.

[0051] Among them, the latent variable regressor can be composed of stacked cross-attention modules. Each cross-attention module can use the latent variable representation of the visible video blocks as the key and value, use the learnable query mask (maskqueries) as the query, and predict the feature representation of the masked video blocks through cross-attention.

[0052] Step 205, calculate the alignment loss based on the predicted features of the masked video blocks and the features of the masked video blocks.

[0053] In this embodiment, the above-mentioned execution entity can calculate the alignment loss based on the prediction features and features of the masked video block, and perform constraints.

[0054] In the training stage, the true features of the masked video block learned by the second encoder are aligned with the predicted features of the masked video block predicted by the latent variable regressor, so as to complete the pretext task in the pre-training stage.

[0055] Among them, the alignment loss can be calculated by the following formula:

[0056]

[0057] Among them, M is the set of all masked video blocks in T frames, |M| is the number of tokens of the masked video block, r p is the output of the regressor, is the output of the masked video block passing through the encoder, can be used as a label to calculate the alignment loss with r p Calculate the alignment loss

[0058] Step 206: Input the predicted features of the masked video block into the visual decoder and the motion decoder respectively to predict the visual codebook and the hidden motion information.

[0059] In this embodiment, the above-mentioned execution entity can input the predicted features of the masked video block into the visual decoder to predict the visual codebook, and input the predicted features of the masked video block into the motion decoder to predict the hidden motion information.

[0060] Here, two disentangled decoders are used simultaneously to predict the reconstructed visual appearance of the video frame and the motion information in the video frame. Among them, the visual decoder can be used to predict the visual codebook of the video frame, and the motion decoder can be used to predict various motion information hidden in the video data, so as to provide supplementary semantic clues in the pre-training stage. Moreover, using two disentangled decoders can decouple the appearance and motion views at the target level, can effectively process redundant spatio-temporal data, and accelerate the convergence speed of the pre-training stage. In addition, through the two disentangled decoders, the encoder can also be forced to learn the inherent spatio-temporal relationship in the video data.

[0061] Step 207: Input the visual codebook into the pre-trained tokenizer to generate a discrete codebook.

[0062] In this embodiment, the above-mentioned execution entity can input the visual codebook into the pre-trained tokenizer to generate a discrete codebook.

[0063] For the visual decoder, an off-the-shelf tokenizer of a pre-trained discrete autoregressive encoder is used to generate a discrete codebook as the training target.

[0064] Step 208, calculate the cross-entropy loss based on the discrete codebook.

[0065] In this embodiment, the above-mentioned execution subject can calculate the cross-entropy loss based on the discrete codebook for constraint.

[0066] Among them, the cross-entropy loss can be calculated by the following formula:

[0067]

[0068] Among them, M is the set of all masked video blocks in T frames, |M| is the number of tokens of the masked video blocks, K is the number of categories of discrete representations generated by the tokenizer of the discrete autoregressive encoder, y m(i,j) is the label value of the i-th input video block corresponding to the j-th category, and p i,j is the predicted value of the i-th input video block corresponding to the j-th category.

[0069] Step 209, calculate the pixel difference corresponding to the masked video block based on the hidden motion information.

[0070] In this embodiment, the above-mentioned execution subject can calculate the pixel difference corresponding to the masked video block based on the hidden motion information.

[0071] For the motion decoder, calculate the pixel difference in the masked area of the input T frames, and use the pixel difference as the supervision information of the motion branch.

[0072] Step 210, calculate the mean squared error loss based on the pixel difference.

[0073] In this embodiment, the above-mentioned execution subject can calculate the mean squared error loss based on the pixel difference and optimize it through the mean squared error loss.

[0074] Among them, the mean squared error loss can be calculated by the following formula:

[0075]

[0076] Among them, M′ represents the set of all masked video blocks from the second frame to the T-th frame, D p is the output of the p-th masked video block corresponding to the motion decoder, is the pixel difference corresponding to the p-th masked video block.

[0077] Step 211: Calculate the sum of the alignment loss, cross-entropy loss, and mean squared error loss to obtain the total loss.

[0078] In this embodiment, the above-mentioned execution entity may calculate the sum of the alignment loss, cross-entropy loss, and mean squared error loss to obtain the total loss.

[0079] Among them, the total loss can be calculated by the following formula:

[0080]

[0081] Among them, is the cross-entropy loss, is the mean squared error loss, is the alignment loss.

[0082] Step 212: Adjust the parameters of the first encoder, second encoder, visual decoder, and motion decoder based on the total loss to obtain a video processing model.

[0083] In this embodiment, the above-mentioned execution entity may adjust the parameters of the first encoder, second encoder, visual decoder, and motion decoder based on the total loss until the training objective is completed to obtain a video processing model. Among them, the training objective may be that the performance of the video processing model composed of the first encoder, second encoder, visual decoder, and motion decoder reaches a preset performance.

[0084] Generally, adjust the parameters of the above first encoder, second encoder, visual decoder, and motion decoder based on the total loss until convergence, that is, complete the training of the video processing model.

[0085] From Figure 2 it can be seen that compared with the corresponding embodiment of Figure 1 , the process 200 of the training method of the video processing model in this embodiment highlights the training steps. Thus, the solution described in this embodiment aligns the true features of the masked video blocks learned by the second encoder with the predicted features of the masked video blocks predicted by the latent variable regressor, and can complete the pretext task in the pre-training stage. Based on the class label value of the masked video block and the class prediction value predicted by the tokenizer, calculate the cross-entropy loss as the supervision information for the visual decoder for constraint. Based on the output of the motion decoder and the pixel difference, calculate the mean squared error loss as the supervision information for the motion decoder for constraint. Combining the above three loss functions for training can improve the effect of the video processing model.

[0086] For ease of understanding, Figure 3 shows a scenario diagram that can implement the training method of the video processing model of the present disclosure.

[0087] Randomly mask the input video segment to obtain visible video blocks and masked video blocks. Input the visible video blocks into the first encoder to learn the features of the visible video blocks. At the same time, input the masked video blocks into the second encoder to learn the features of the masked video blocks.

[0088] To eliminate the differences between the appearance view and the motion view, input the features of the visible video blocks into the regressor, use the learnable query mask as the query, and predict the features of the masked video blocks through cross-attention. Calculate the alignment loss based on the output of the second encoder and the output of the regressor. Complete the pretext task in the pre-training stage in this way.

[0089] Input the predicted features of the masked video blocks into the visual decoder and the motion decoder simultaneously to predict the reconstructed visual appearance of the video and the motion information in the video. Among them, the motion decoder can predict various motion information hidden in the video data, and the visual decoder can predict the visual codebook of the video, so as to provide supplementary semantic clues in the pre-training stage.

[0090] For the visual decoder, use the tokenizer of the pre-trained discrete autoregressive encoder to generate a discrete codebook as the training target, and use the cross-entropy loss for constraint. For the motion decoder, calculate the pixel difference of the input T frames in the masked area. Use the pixel difference as the supervision information of the motion branch and optimize it through the mean squared error loss to optimize.

[0091] Finally, the total loss function in the pre-training stage is the sum of the alignment loss, the cross-entropy loss, and the mean squared error loss. Adjust the parameters based on the total loss to complete the training of the video processing model.

[0092] For further reference Figure 4 , which shows the process 400 of an embodiment of the video processing method according to the present disclosure. The video processing method includes the following steps:

[0093] Step 401, obtain the video to be processed for the target task.

[0094] In this embodiment, the execution subject of the video processing method can obtain the video to be processed for the target task. Among them, the target task may include, but is not limited to, image classification, object detection, semantic segmentation, action recognition, and so on.

[0095] Step 402, input the video to be processed into the video processing model to obtain the processing result of the target task of the video to be processed.

[0096] In this embodiment, the above-mentioned execution entity may input the video to be processed into a video processing model to obtain the target task processing result of the video to be processed. Among them, the video processing model may be trained using the method shown in Figure 1 or Figure 2 , which will not be elaborated here.

[0097] Generally, for the video frames in the video to be processed, the video frames to be processed are input into an encoder to learn the features of the video frames to be processed; the features of the video frames to be processed are input into a visual decoder and a motion decoder to predict the visual codebook and motion information of the video frames to be processed; based on the visual codebook and motion information of the video frames to be processed, target task processing is performed to obtain the target task processing result.

[0098] In some embodiments, in order to improve the processing effect of the video processing model on the target task, a small amount of labeled video data sets of the target task may be obtained, and the parameters of the video processing model are fine-tuned using these labeled video data sets. The video to be processed is input into the fine-tuned video processing model for processing, so that the video processing model can achieve excellent results in the target task.

[0099] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, an embodiment of a video processing model training device is provided in the present disclosure. This device embodiment corresponds to the method embodiment shown in Figure 1 , and this device can be specifically applied to various electronic devices.

[0100] As shown in Figure 5 , the video processing model training device 500 in this embodiment may include: a first acquisition module 501, an encoding module 502, a decoding module 503, a first calculation module 504, and a first training module 505. Among them, the first acquisition module 501 is configured to acquire masked video frames, where the masked video frames include visible video blocks and masked video blocks; the encoding module 502 is configured to input the masked video frames into an encoder to learn the features of the masked video frames; the decoding module 503 is configured to input the features of the masked video frames into a visual decoder and a motion decoder respectively to predict the visual codebook and hidden motion information of the masked video frames; the first calculation module 504 is configured to calculate a loss based on the visual codebook and hidden motion information; the first training module 505 is configured to adjust the parameters of the encoder, visual decoder, and motion decoder based on the loss to obtain a video processing model.

[0101] In this embodiment, in the video processing model training device 500: the specific processing of the first acquisition module 501, the encoding module 502, the decoding module 503, the first calculation module 504, and the first training module 505 and the technical effects brought by them can be respectively referred to Figure 1The relevant descriptions of steps 101-105 in the corresponding embodiments will not be elaborated here.

[0102] In some alternative implementation manners of this embodiment, the encoding module 502 includes: a first encoding sub-module configured to input a visible video block into a first encoder to learn the features of the visible video block; a regression sub-module configured to input the features of the visible video block into a latent variable regressor to obtain the predicted features of the masked video block.

[0103] In some alternative implementation manners of this embodiment, the decoding module 503 is further configured to: input the predicted features of the masked video block into a visual decoder and a motion decoder respectively to predict a visual codebook and hidden motion information.

[0104] In some alternative implementation manners of this embodiment, the first calculation module 504 is further configured to: input the visual codebook into a pre-trained tokenizer to generate a discrete codebook; calculate a cross-entropy loss based on the discrete codebook; calculate the pixel difference corresponding to the masked video block based on the hidden motion information; calculate a mean squared error loss based on the pixel difference.

[0105] In some alternative implementation manners of this embodiment, the encoding module 502 further includes: a second encoding sub-module configured to input the masked video block into a second encoder to learn the features of the masked video block.

[0106] In some alternative implementation manners of this embodiment, the video processing model training device 500 further includes: a second calculation module configured to calculate an alignment loss based on the predicted features of the masked video block and the features of the masked video block.

[0107] In some alternative implementation manners of this embodiment, the first training module 505 is further configured to: calculate the sum of the alignment loss, the cross-entropy loss, and the mean squared error loss to obtain a total loss; adjust the parameters of the first encoder, the second encoder, the visual decoder, and the motion decoder based on the total loss to obtain a video processing model.

[0108] In some alternative implementation manners of this embodiment, the latent variable regressor is composed of stacked cross-attention modules. The cross-attention module uses the latent variable representation of the visible video block as keys and values, uses a learnable query mask as the query, and predicts the features of the masked video block through cross-attention.

[0109] In some alternative implementation manners of this embodiment, the first acquisition module 501 is further configured to: acquire a video frame; perform random masking on a preset proportion of video blocks in the video frame to obtain a masked video frame.

[0110] In some alternative implementation manners of this embodiment, the video processing model training device 500 further includes: a second acquisition module configured to acquire a training sample set of a target task, where the training samples in the training sample set include sample videos and target task processing results of the sample videos, and the target task includes at least one of the following: image classification, object detection, semantic segmentation, and action recognition; a second training module configured to use the sample videos as inputs and the target task processing results of the sample videos as outputs to continue training the video processing model.

[0111] Further referring to Figure 6 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a video processing device. This device embodiment corresponds to Figure 4 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0112] As shown in Figure 6 , the video processing device 600 of this embodiment may include: an acquisition module 601 and a processing module 602. Among them, the acquisition module 601 is configured to acquire a video to be processed for a target task, where the target task includes at least one of the following: image classification, object detection, semantic segmentation, and action recognition; the processing module 602 is configured to input the video to be processed into a video processing model to obtain a target task processing result of the video to be processed, where the video processing model is trained by using the device shown in Figure 5 .

[0113] In this embodiment, in the video processing device 600: for the specific processing of the acquisition module 601 and the processing module 602 and the technical effects brought by them, reference may be respectively made to Figure 4 the relevant descriptions of steps 401-402 in the corresponding embodiment, which will not be elaborated here.

[0114] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0115] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0116] Figure 7FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0117] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0118] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0119] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the video processing model training method or the video processing method. For example, in some embodiments, the video processing model training method or the video processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the video processing model training method or the video processing method described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the video processing model training method or the video processing method in any other suitable manner (e.g., by means of firmware).

[0120] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed as an independent software package partially on the machine and partially on a remote machine, or executed entirely on a remote machine or server.

[0122] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0124] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0125] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0126] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution provided in this disclosure can be achieved, and this is not limited herein.

[0127] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for training a video processing model, comprising: Obtaining a masked video frame, wherein the masked video frame includes a visible video block and a masked video block; Inputting the masked video frame into an encoder to learn the features of the masked video frame; Inputting the features of the masked video frame into a visual decoder and a motion decoder respectively to predict the visual codebook and the hidden motion information of the masked video frame, wherein the visual decoder and the motion decoder are two disentangled encoders for decoupling the appearance and motion views and forcing the encoder to learn the inherent spatio-temporal relationships in the video data; Calculating a loss based on the visual codebook and the hidden motion information; Adjusting the parameters of the encoder, the visual decoder and the motion decoder based on the loss to obtain a video processing model.

2. The method according to claim 1, wherein The step of inputting the masked video frame into an encoder to learn the features of the masked video frame includes: Inputting the visible video block into a first encoder to learn the features of the visible video block; Inputting the features of the visible video block into a latent variable regressor to obtain the predicted features of the masked video block.

3. The method according to claim 2, wherein, The step of inputting the features of the masked video frame into a visual decoder and a motion decoder respectively to predict the visual codebook and the hidden motion information of the masked video frame includes: Inputting the predicted features of the masked video block into the visual decoder and the motion decoder respectively to predict the visual codebook and the hidden motion information.

4. The method according to claim 3, wherein The step of calculating a loss based on the visual codebook and the hidden motion information includes: Inputting the visual codebook into a pre-trained tokenizer to generate a discrete codebook; Calculating a cross-entropy loss based on the discrete codebook; Calculating the pixel difference corresponding to the masked video block based on the hidden motion information; Calculating a mean squared error loss based on the pixel difference.

5. The method according to claim 4, wherein, The step of inputting the masked video frame into an encoder to learn the features of the masked video frame further includes: Inputting the masked video block into a second encoder to learn the features of the masked video block.

6. The method according to claim 5, wherein The method further includes: Calculating an alignment loss based on the predicted features of the masked video block and the features of the masked video block.

7. The method according to claim 6, wherein, The step of adjusting the parameters of the encoder, the visual decoder and the motion decoder based on the loss to obtain a video processing model includes: Calculating the sum of the alignment loss, the cross-entropy loss and the mean squared error loss to obtain a total loss; Adjusting the parameters of the first encoder, the second encoder, the visual decoder and the motion decoder based on the total loss to obtain the video processing model.

8. The method according to claim 2, wherein The latent variable regressor is composed of stacked cross-attention modules, and the cross-attention module uses the latent variable representation of the visible video block as the key and value, takes a learnable query mask as the query, and predicts the features of the masked video block through cross-attention.

9. The method according to any one of claims 1-8, wherein, The step of obtaining a masked video frame includes: Obtaining a video frame; Randomly masking a preset proportion of video blocks in the video frame to obtain the masked video frame.

10. The method according to any one of claims 1-8, wherein, The method further includes: Obtain a training sample set for a target task, where the training samples in the training sample set include sample videos and the target task processing results of the sample videos, and the target task includes at least one of the following: image classification, object detection, semantic segmentation, and action recognition; Use the sample video as the input and the target task processing result of the sample video as the output to continue training the video processing model.

11. A video processing method, comprising: Obtain a video to be processed for a target task, where the target task includes at least one of the following: image classification, object detection, semantic segmentation, and action recognition; Input the video to be processed into a video processing model to obtain the target task processing result of the video to be processed, where the video processing model is trained using the method described in any one of claims 1-10.

12. A video processing model training device, comprising: A first acquisition module configured to acquire masked video frames, where the masked video frames include visible video blocks and masked video blocks; An encoding module configured to input the masked video frames into an encoder to learn the features of the masked video frames; A decoding module configured to input the features of the masked video frames into a visual decoder and a motion decoder respectively to predict the visual codebook and the hidden motion information of the masked video frames, where the visual decoder and the motion decoder are two disentangled encoders for decoupling the appearance and motion views and forcing the encoder to learn the inherent spatio-temporal relationships in the video data; A first calculation module configured to calculate a loss based on the visual codebook and the hidden motion information; A first training module configured to adjust the parameters of the encoder, the visual decoder, and the motion decoder based on the loss to obtain a video processing model.

13. The device according to claim 12, wherein The encoding module includes: A first encoding sub-module configured to input the visible video blocks into a first encoder to learn the features of the visible video blocks; A regression sub-module configured to input the features of the visible video blocks into a latent variable regressor to obtain the predicted features of the masked video blocks.

14. The device according to claim 13, wherein, The decoding module is further configured to: Input the predicted features of the masked video blocks into the visual decoder and the motion decoder respectively to predict the visual codebook and the hidden motion information.

15. The device according to claim 14, wherein, The first calculation module is further configured to: Input the visual codebook into a pre-trained tokenizer to generate a discrete codebook; Calculate the cross-entropy loss based on the discrete codebook; Calculate the pixel difference corresponding to the masked video blocks based on the hidden motion information; Calculate the mean square error loss based on the pixel difference.

16. The device according to claim 15, wherein, The encoding module further includes: A second encoding sub-module configured to input the masked video blocks into a second encoder to learn the features of the masked video blocks.

17. The device according to claim 16, wherein, The device further includes: A second calculation module configured to calculate an alignment loss based on the predicted features of the masked video blocks and the features of the masked video blocks.

18. The apparatus according to claim 17, wherein, The first training module is further configured to: Calculate the sum of the alignment loss, the cross-entropy loss, and the mean squared error loss to obtain the total loss; Based on the total loss, adjust the parameters of the first encoder, the second encoder, the visual decoder, and the motion decoder to obtain the video processing model.

19. The apparatus according to claim 13, wherein, The latent variable regressor consists of stacked cross-attention modules. The cross-attention module uses the latent variable representation of the visible video blocks as keys and values, takes a learnable query mask as the query, and predicts the features of the masked video blocks through cross-attention.

20. The apparatus according to any one of claims 12-19, wherein, The obtaining module is further configured to: Obtain video frames; Randomly mask a preset proportion of the video blocks in the video frames to obtain the masked video frames.

21. The apparatus according to any one of claims 12-19, wherein, The apparatus further includes: A second obtaining module, configured to obtain a training sample set of a target task, where the training samples in the training sample set include sample videos and the target task processing results of the sample videos, and the target task includes at least one of the following: image classification, object detection, semantic segmentation, and action recognition; A second training module, configured to use the sample videos as inputs and the target task processing results of the sample videos as outputs to further train the video processing model.

22. A video processing apparatus, comprising: An obtaining module, configured to obtain a video to be processed for a target task, where the target task includes at least one of the following: image classification, object detection, semantic segmentation, and action recognition; A processing module, configured to input the video to be processed into a video processing model to obtain the target task processing result of the video to be processed, where the video processing model is trained using the apparatus according to any one of claims 12-21.

23. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-10 or the method according to claim 11.

24. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method according to any one of claims 1-10 or the method according to claim 11.

25. A computer program product, comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-10 or the method according to claim 11.

Citation Information

Patent Citations

  • Video segmentation model training method and device, equipment and storage medium

    CN115496905A

Cited By

  • Video processing method and device, model training method and device, electronic equipment and storage medium

    CN119893137A