Model Training Method and Apparatus Based on Multimodal Correlation Data
By considering the modal loss of video and associated data during training and optimizing model parameters, the problem that existing models cannot effectively utilize modal association data is solved, thus improving video reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies fail to effectively utilize the correlation data of different modalities in videos when training video reconstruction models, resulting in the models being unable to reconstruct videos based on the correlation data of specific modalities.
By determining the modal correlation data of the training video, calculating the video loss function value and modal loss function value between the video and the correlation data, optimizing the model parameters until convergence, and constructing a video reconstruction prediction model.
The model was able to repair or reconstruct videos based on associated data, improving the accuracy and effectiveness of video reconstruction.
Smart Images

Figure CN115240099B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of algorithm model training technology, and in particular to a model training method and apparatus based on multimodal correlation data. Background Technology
[0002] With the development of algorithm technology, more and more companies are using algorithmic models to perform video-related data prediction tasks, such as video reconstruction. These tasks require algorithmic models to fully extract and learn video features. However, current technologies, when training such models, do not consider constructing training data based on the correlation data of different modalities corresponding to the video, nor do they consider simultaneously calculating the difference loss between different correlation data and the video during training. Therefore, the trained model cannot achieve the effect of video reconstruction based on the correlation data of a specific modality from a single frame. Clearly, current technologies have shortcomings that urgently need to be addressed. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a model training determination method and apparatus based on multimodal association data, which enables the model to learn the relationship between video and association data during training, thereby enabling the finally trained model to repair or even reconstruct video based on association data.
[0004] To address the aforementioned technical problems, the first aspect of this invention discloses a model training method based on multimodal correlation data, the method comprising:
[0005] Identify the training videos to be used to train the model;
[0006] Determine the associated data for at least one modality corresponding to the training video;
[0007] The training video and the associated data are input into the video reconstruction prediction model for training. During training, the video loss function value between the predicted video output by the video reconstruction prediction model and the input training video, as well as the modal loss function value between the predicted video and at least one of the input associated data, are calculated.
[0008] The model parameters of the video reconstruction prediction model are optimized based on the video loss function value and the modal loss function value until convergence, resulting in the trained video reconstruction prediction model.
[0009] As an optional implementation, in a first aspect of the invention, the modality includes at least one of an audio modality, a text modality, and an image modality; and / or, the associated data includes at least one of descriptive audio data, descriptive text data, and characterizing image data.
[0010] As an optional implementation, in a first aspect of the invention, the step of training the training video and the associated data into a video reconstruction prediction model, and calculating, during training, a video loss function value between the predicted video output by the video reconstruction prediction model and the input training video, and a modal loss function value between the predicted video and at least one of the input associated data, includes:
[0011] Perform frame extraction on the training video to obtain multiple training video frames corresponding to the training video.
[0012] The multiple training video frames and the associated data are input into the video reconstruction prediction model for training;
[0013] In the training, the video loss function values between the multiple predicted video frames output by the video reconstruction prediction model and the multiple training video frames input are calculated, as well as the modal loss function values between the predicted video and at least one of the input associated data.
[0014] As an optional implementation, in the first aspect of the present invention, the associated data is characterizing image data, and the step of determining the associated data of at least one modality corresponding to the training video includes:
[0015] The target representation frame image is determined from the plurality of training video frames, and the target representation frame image is copied to obtain a plurality of copied representation frame images;
[0016] The plurality of replicated representation frame images are determined as the representation image data corresponding to the training video.
[0017] As an optional implementation, in the first aspect of the present invention, the step of performing frame extraction on the training video to obtain a plurality of training video frames corresponding to the training video includes:
[0018] The frame extraction interval corresponding to the training video is determined based on the video parameters of the training video.
[0019] The training video is subjected to frame extraction operation according to the frame extraction interval to obtain multiple training video frames corresponding to the training video.
[0020] As an optional implementation, in a first aspect of the present invention, the video reconstruction prediction model includes a video reconstruction network and a modality reconstruction network; the video reconstruction network is used to receive the training video and the associated data and reconstruct the predicted video; the modality reconstruction network is used to extract the associated data features of the modality corresponding to the predicted video; the associated data features are used to compare with the associated data to calculate the modality loss function value; the modality reconstruction network is trained and converged through a training dataset including multiple training videos and corresponding training associated data of the modality;
[0021] And, the step of optimizing the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence, to obtain the trained video reconstruction prediction model, includes:
[0022] During training, the parameters of the modality reconstruction network are kept constant. The network parameters of the video reconstruction network are optimized based on the video loss function value and the modality loss function value until convergence, thus obtaining the trained video reconstruction network.
[0023] As an optional implementation, in the first aspect of the present invention, optimizing the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence includes:
[0024] Calculate the weighted sum of the video loss function value and the modal loss function value;
[0025] Based on the weighted sum, the model parameters of the video reconstruction prediction model are optimized until convergence.
[0026] As an optional implementation, in the first aspect of the present invention, the video loss function value is calculated as follows:
[0027] For any of the predicted video frames, calculate the frame loss function value between the predicted video frame and the corresponding training video frame;
[0028] Calculate the average of the frame loss function values for all the predicted video frames to obtain the loss function values between the multiple predicted video frames and the multiple input training video frames.
[0029] A second aspect of the present invention discloses a model training apparatus based on multimodal correlation data, the apparatus comprising:
[0030] The video determination module is used to determine the training videos to be used for training the model;
[0031] The association determination module is used to determine the association data of at least one modality corresponding to the training video;
[0032] The model training module is used to train the video reconstruction prediction model by inputting the training video and the associated data into the video reconstruction prediction model. During training, the module calculates the video loss function value between the predicted video output by the video reconstruction prediction model and the input training video, as well as the modal loss function value between the predicted video and at least one of the input associated data.
[0033] The model optimization module is used to optimize the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence, so as to obtain the trained video reconstruction prediction model.
[0034] As an optional implementation, in a second aspect of the invention, the modality includes at least one of an audio modality, a text modality, and an image modality; and / or, the associated data includes at least one of descriptive audio data, descriptive text data, and characterizing image data.
[0035] As an optional implementation, in a second aspect of the present invention, the model training module includes:
[0036] A frame extraction unit is used to perform frame extraction on the training video to obtain multiple training video frames corresponding to the training video.
[0037] The model training unit is used to input the multiple training video frames and the associated data into the video reconstruction prediction model for training.
[0038] The loss calculation unit is used to calculate, during the training, the video loss function value between a plurality of predicted video frames output by the video reconstruction prediction model and the plurality of training video frames input, as well as the modal loss function value between the predicted video and at least one of the associated data input.
[0039] As an optional implementation, in a second aspect of the present invention, the associated data is characterizing image data, and the specific method by which the association determination module determines the associated data of at least one modality corresponding to the training video includes:
[0040] The target representation frame image is determined from the plurality of training video frames, and the target representation frame image is copied to obtain a plurality of copied representation frame images;
[0041] The plurality of replicated representation frame images are determined as the representation image data corresponding to the training video.
[0042] As an optional implementation, in a second aspect of the present invention, the specific method by which the frame extraction unit performs frame extraction on the training video to obtain multiple training video frames corresponding to the training video includes:
[0043] The frame extraction interval corresponding to the training video is determined based on the video parameters of the training video.
[0044] The training video is subjected to frame extraction operation according to the frame extraction interval to obtain multiple training video frames corresponding to the training video.
[0045] As an optional implementation, in a second aspect of the present invention, the video reconstruction prediction model includes a video reconstruction network and a modality reconstruction network; the video reconstruction network is used to receive the training video and the associated data and reconstruct the predicted video; the modality reconstruction network is used to extract the associated data features of the modality corresponding to the predicted video; the associated data features are used to compare with the associated data to calculate the modality loss function value; the modality reconstruction network is trained and converged through a training dataset including multiple training videos and corresponding training associated data of the modality.
[0046] Furthermore, the specific method by which the model optimization module optimizes the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence, to obtain the trained video reconstruction prediction model, includes:
[0047] During training, the parameters of the modality reconstruction network are kept constant. The network parameters of the video reconstruction network are optimized based on the video loss function value and the modality loss function value until convergence, thus obtaining the trained video reconstruction network.
[0048] As an optional implementation, in a second aspect of the present invention, the specific method by which the model optimization module optimizes the model parameters of the video reconstruction prediction model until convergence based on the video loss function value and the modal loss function value includes:
[0049] Calculate the weighted sum of the video loss function value and the modal loss function value;
[0050] Based on the weighted sum, the model parameters of the video reconstruction prediction model are optimized until convergence.
[0051] As an optional implementation, in the second aspect of the present invention, the video loss function value is calculated as follows:
[0052] For any of the predicted video frames, calculate the frame loss function value between the predicted video frame and the corresponding training video frame;
[0053] Calculate the average of the frame loss function values for all the predicted video frames to obtain the loss function values between the multiple predicted video frames and the multiple input training video frames.
[0054] A third aspect of the present invention discloses another model training apparatus based on multimodal correlation data, the apparatus comprising:
[0055] Memory containing executable program code;
[0056] A processor coupled to the memory;
[0057] The processor calls the executable program code stored in the memory to execute some or all of the steps in the model training method based on multimodal association data disclosed in the first aspect of the present invention.
[0058] The fourth aspect of the present invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute some or all of the steps in the model training method based on multimodal association data disclosed in the first aspect of the present invention.
[0059] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0060] In this embodiment of the invention, a training video for training the model is determined; at least one modality-related data corresponding to the training video is determined; the training video and the related data are input into a video reconstruction prediction model for training; during training, a video loss function value between the predicted video output by the video reconstruction prediction model and the input training video, and a modal loss function value between the predicted video and at least one of the input related data are calculated; the model parameters of the video reconstruction prediction model are optimized based on the video loss function value and the modal loss function value until convergence, resulting in a trained video reconstruction prediction model. Therefore, this invention can train a model based on a training video and corresponding specific modality-related data, and considers the modal loss between the predicted video and the related data during training. This allows the model to learn the relationship between the video and the related data during training, enabling the finally trained model to repair or even reconstruct the video based on the related data. Attached Figure Description
[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0062] Figure 1 This is a flowchart illustrating a model training method based on multimodal correlation data disclosed in an embodiment of the present invention;
[0063] Figure 2 This is a flowchart illustrating another model training method based on multimodal association data disclosed in an embodiment of the present invention;
[0064] Figure 3 This is a schematic diagram of the structure of a model training device based on multimodal correlation data disclosed in an embodiment of the present invention;
[0065] Figure 4 This is a schematic diagram of another model training device based on multimodal correlation data disclosed in an embodiment of the present invention;
[0066] Figure 5 This is a schematic diagram of the structure of another model training device based on multimodal association data disclosed in an embodiment of the present invention. Detailed Implementation
[0067] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0068] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or end that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or ends.
[0069] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0070] This invention discloses a model training method and apparatus based on multimodal association data. It can train a model based on training videos and corresponding association data for specific modalities, and considers the modal loss between the predicted video and the association data during training. This allows the model to learn the relationship between the video and the association data during training, enabling the final trained model to repair or even reconstruct videos based on the association data. Detailed explanations follow.
[0071] Example 1
[0072] Please see Figure 1 , Figure 1 This is a flowchart illustrating a model training method based on multimodal correlation data disclosed in an embodiment of the present invention. Figure 1 The described method is applied in a video data processing device, which can be a corresponding processing terminal, processing equipment, or processing server. The server can be a local server or a cloud server; this embodiment of the invention does not impose any limitations. Figure 1 As shown, the model training method based on multimodal association data may include the following operations:
[0073] 101. Determine the training videos to be used to train the model.
[0074] Optionally, the video content of the training video can be determined based on the target processing video content of the model to be trained. It can also include multiple types of video content to improve the adaptive characteristics of the trained model. Optionally, the video content of the training video can be a continuous action or a continuous scene. Since the solution of this invention incorporates the assistance of correlated data, the camera angles or scenes in the video content of the training video can be switched.
[0075] 102. Determine the associated data for at least one modality corresponding to the training video.
[0076] Optionally, the modality may include at least one of an audio modality, a text modality, and an image modality. Correspondingly, the associated data may include at least one of descriptive audio data, descriptive text data, and image data.
[0077] Optionally, the associated data is related to the content of the training video and can be used to describe some features of the training video. It can be written by the operator based on the content of the training video, or it can be automatically predicted and generated by other video content algorithms. This invention does not limit it.
[0078] 103. Input the training video and associated data into the video reconstruction prediction model for training. During training, calculate the video loss function value between the predicted video output by the video reconstruction prediction model and the input training video, as well as the modal loss function value between the predicted video and at least one input associated data.
[0079] 104. Optimize the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence, and obtain the trained video reconstruction prediction model.
[0080] Optionally, the loss function value can be an L1 loss function value, an L2 loss function value, or other loss function values suitable for calculating the similarity between images or texts; this invention does not limit the specific loss function value.
[0081] Optionally, the gradient descent method can be used to continuously optimize the model parameters until the loss function value reaches its minimum, thereby enabling the model to converge and obtaining a trained video reconstruction prediction model.
[0082] As can be seen, the method described in the embodiments of the present invention can train the model based on the training video and the corresponding specific modal association data, and consider the modal loss between the predicted video and the association data during training, so that the model can learn the relationship between the video and the association data during training, and thus enable the finally trained model to repair or even reconstruct the video based on the association data.
[0083] As an optional implementation, the above steps, including training the video reconstruction prediction model by inputting the training video and associated data, and calculating the video loss function value between the predicted video output by the video reconstruction prediction model and the input training video, as well as the modal loss function value between the predicted video and at least one piece of input associated data, during training, include:
[0084] Perform frame extraction on the training video to obtain multiple training video frames corresponding to the training video.
[0085] Multiple training video frames and associated data are input into the video reconstruction prediction model for training;
[0086] During training, the video loss function values between multiple predicted video frames output by the video reconstruction prediction model and multiple training video frames input are calculated, as well as the modal loss function values between the predicted video and at least one associated data point input.
[0087] Optionally, the frame extraction operation can be performed manually or automatically by an algorithm according to a predetermined frame extraction interval; this invention does not impose any limitation on this.
[0088] With the above settings, frames can be extracted from the training video to obtain a small number of video frames that can be used to characterize the features of the training video as training data, which can effectively reduce the amount of data and workload.
[0089] As an optional implementation, the associated data is characterizing image data. The step described above, determining the associated data for at least one modality corresponding to the training video, includes:
[0090] The target representation frame image is determined from multiple training video frames, and the target representation frame image is copied to obtain multiple copied representation frame images;
[0091] Multiple replicated representation frame images are identified as the representation image data corresponding to the training video.
[0092] Optionally, the target representation frame image can be the first training video frame among multiple training video frames, i.e., the first frame image. This setting allows the model to learn the image relationship between the first frame image and the entire video. Alternatively, the target representation frame image can also be a video frame among multiple training video frames that plays a key representational role, such as a frame showing a landmark or representing a turning point in the plot. This can also be used to train the model to learn the image relationship between keyframe images and the entire video. Optionally, the number of duplicate representation frame images input into the model for each training iteration can be the same as the number of training video frames to facilitate subsequent loss calculation.
[0093] As an optional implementation, the video reconstruction prediction model includes a video reconstruction network and a modal reconstruction network. The video reconstruction network is used to receive training videos and associated data and reconstruct the prediction video, while the modal reconstruction network is used to extract the associated data features of the modality corresponding to the prediction video.
[0094] Specifically, the associated data features are used to compare with the associated data to calculate the modality loss function value. The modality reconstruction network is trained and converged through a training dataset that includes multiple training videos and corresponding modal training associated data. That is, it is trained before training the video reconstruction prediction model and can be directly used to extract the associated data features of a specific modality.
[0095] Optionally, the modal reconstruction network can be a text reconstruction network or an audio reconstruction network.
[0096] Optionally, the modality reconstruction network can be trained in advance using the original video frames and the corresponding modality association data. That is, the video frames are input into the modality reconstruction network, the SmoothL1 loss between the output feature vector and the vector of the association data is calculated, and the training of the modality reconstruction network is completed by optimizing this loss value until convergence.
[0097] Optionally, the video reconstruction network can be used to receive multiple replicated representation frame images and associated data and reconstruct multiple predicted video frames. Correspondingly, the video loss function value can be the video loss function value between the multiple predicted video frames output by the video reconstruction prediction model and the multiple training video frames input.
[0098] Optionally, the above steps, including optimizing the model parameters of the video reconstruction prediction model based on the video loss function value and the text loss function value until convergence, to obtain a trained video reconstruction prediction model, include:
[0099] During training, the parameters of the modal reconstruction network are kept constant. The network parameters of the video reconstruction network are optimized based on the video loss function value and the modal loss function value until convergence, resulting in a well-trained video reconstruction network.
[0100] Specifically, during training, the parameters of the modality reconstruction network are frozen, and only the network parameters of the video reconstruction network are optimized. Finally, the trained video reconstruction network can be used for video reconstruction.
[0101] Optionally, the modal reconstruction network includes an embedding module, a Transformer module, and a fully connected layer module.
[0102] Specifically, a modality reconstruction network can be constructed to maintain the semantic association between video frames and associated data. It consists of an embedding layer, two Transformer modules and a fully connected layer, which is used to receive the video frames reconstructed by the decoder of the video reconstruction network, reconstruct the feature vectors of specific modalities from the video frames through the modality reconstruction network, and optimize the distance between the feature vectors and the initial encoded feature vectors corresponding to the associated data through the modality loss function.
[0103] As an optional implementation, the video reconstruction network includes an encoder and a decoder, wherein the encoder is used to extract features, and the decoder is used to reconstruct the features. Optionally, the encoder includes a video embedding layer, a modal embedding layer, a feature fusion layer, and a first Transformer layer. The video embedding layer receives training video frames and processes them to obtain video features; the modal embedding layer receives associated data and processes it to obtain modal features; and the feature fusion layer fuses the video features and modal features to obtain training features, which are then input into the first Transformer layer.
[0104] The video embedding layer includes two-dimensional convolutional layers and / or three-dimensional convolutional layers.
[0105] Optionally, a two-dimensional convolutional patch embedding layer can be constructed as a video embedding layer, which is used to replace the high-dimensional original image features with a low-dimensional vector. Optionally, the patch embedding layer consists of a single convolutional layer with equal kernel size and stride, and the number of output channels can be selected as 768 (other options are also possible).
[0106] Optionally, a 3D convolutional video patch embedding layer can be constructed as the video embedding layer, i.e., a 3D PatchEmbedding layer. This 3D convolutional layer performs convolutions simultaneously in both spatial and temporal dimensions, thereby extracting additional correlation features between consecutive frames. Specifically, to ensure that each frame possesses correlation features, the stride of the convolution is not necessarily equal to the depth of the convolution. That is, the stride can be equal to 1. For example, if the convolution depth is 2 and the stride is 1, the first and second frames will obtain a feature vector through 3D convolution, and subsequently, the second and third frames will also obtain a feature vector, and so on. If the convolution depth is 2 and the stride is also 2, the first and second frames will convolve to obtain a feature vector, and then the convolutional features between the third and fourth frames will be directly calculated, without calculating the features between the second and third frames.
[0107] Specifically, this video embedding layer converts the input video frames into corresponding feature vectors. Taking a specific implementation of a two-dimensional convolutional image patch embedding layer as an example, consider 30 video frames, each with an image size of 224*224, resulting in dimensions (30, 3, 224, 224), where 3 represents the RGB three channels. The convolution kernel size is (16, 16), and the stride is 16. After conversion, it becomes (30, 768, 14, 14), which is then transformed into a feature vector (30, 768, 196). Subsequently, the order of the dimensions is reversed, and a class label is added at the initial position of dimension 196 for downstream tasks such as classification. The final feature vector has dimensions (30, 197, 768).
[0108] Optionally, the modal embedding layer may include a Tokenizer module and an Embedding module, wherein the Tokenizer module is used to convert the data units of the input associated data into token indices, and the Embedding module encodes the token indices into feature vectors.
[0109] Optionally, the feature fusion layer may include a first fully connected layer, a second fully connected layer, a feature fusion module, and a GELU activation layer. The input feature dimensions of the first and second fully connected layers must be equal to facilitate feature fusion. The first fully connected layer receives the output features from the video embedding layer and performs feature space transformation. The second fully connected layer receives the output features from the modality embedding layer and performs feature space transformation. The feature fusion module can use feature addition and averaging, concatenation, or other fusion methods to fuse the features of the first and second fully connected layers, and then transforms them through the GELU activation layer to obtain the final fused features.
[0110] In a specific scheme, the associated data is descriptive text data, which corresponds to text features. To better fuse video frame features and text features, two fully connected layers are constructed for feature space transformation. The feature vector dimension of the video frame is (30, 197, 768). Since no position markers are added, the feature vector dimension of the text features is (1, 196, 768) when using the maximum text length of 196. The two types of features are concatenated on the first dimension (starting from 0), and then the dimensional order of the concatenated features is changed, that is: (30, 197+196, 768) -> (30, 768, 197+196). Then, it is input into a fully connected layer, and the output dimension is (30, 768, 197). Finally, a GELU activation layer is used for non-linear transformation to obtain the fused features, and the dimension is transformed again to (30, 197, 768).
[0111] Optionally, the first Transformer layer includes multiple stacked Transformer modules, the structure of which can be referenced from the encoder structure of the VIT (Vision Transformer) network.
[0112] As can be seen, by implementing this optional implementation method, an encoder structure that can fully extract and fuse features from video frames and related data can be constructed, resulting in better prediction performance of the model.
[0113] As an optional implementation, the decoder includes a first fully connected layer, a second Transformer layer, and a second fully connected layer.
[0114] Specifically, the decoder first employs a fully connected layer to fuse and transform the encoder's features. This is followed by multiple Transformer modules, and finally, another fully connected layer is used to generate pixel-level video frames.
[0115] As can be seen, by implementing this optional implementation method, a decoder structure that can fully reconstruct the features of video frames can be constructed, thereby improving the prediction performance of the model.
[0116] As an optional implementation, the step described above, optimizing the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence, includes:
[0117] Calculate the weighted sum of the video loss function value and the modal loss function value;
[0118] The model parameters of the video reconstruction prediction model are optimized based on the weighted summation until convergence.
[0119] Optionally, the weights corresponding to the video loss function value and the modal loss function value can be adjusted by technicians based on experimental or empirical values to achieve the best representation effect.
[0120] Optionally, the video loss function value can be calculated as follows:
[0121] For any predicted video frame, calculate the frame loss function value between the predicted video frame and the corresponding training video frame;
[0122] Calculate the average of the frame loss function values for all predicted video frames to obtain the loss function values between multiple predicted video frames and multiple input training video frames.
[0123] In a specific implementation, the final video frame reconstruction process involves two parts in its training optimization loss. The first part is the image reconstruction loss between the decoded video frame and the original video frame, which can be achieved using SmoothL1 loss. The reconstruction loss between each decoded video frame and its corresponding real frame is calculated, and the average loss for each frame is taken as the final reconstruction loss. The second part is the modal association loss between the video and the associated data, i.e., the modal loss function. This is achieved by using the trained modal reconstruction module to re-express the feature vector of the modality corresponding to the associated data from the reconstructed video frame, and calculating the distance between this feature vector and the original feature vector of the associated data. This part of the loss is also optimized using SmoothL1. During this training process, the parameters of the modal reconstruction module are frozen and do not participate in the parameter update process.
[0124] Specifically, since the video frames input to the encoder are all identical (composed of stacked specific frame images), the decoder's task is actually to perform local image corrections based on the input association data within these existing video frames. If only the video frame reconstruction loss is used, during training, some frames may reconstruct well while others reconstruct poorly. Furthermore, the semantic information expressed by the reconstructed video frame sequence is chaotic because it only considers the similarity between the reconstructed and original frames, neglecting the overall continuity between the reconstructed frames. Therefore, a modal loss function is needed to re-express the modal association information from the reconstructed video frames. By narrowing the distance between the features of the reconstructed association data and the feature vectors of the original association data, the reconstructed video frame sequence can also express complete association information, thus maintaining the continuity of the reconstructed video frames.
[0125] Specifically, the two losses mentioned above are balanced by setting different weighting factors. The video loss function considers the similarity between each pair of video frames from the perspective of each frame, while the modal loss function considers the correlation as a whole. Since the parameters of the modal reconstruction network are fixed, the value of the modal loss function directly depends on the overall similarity between the reconstructed video frame sequence and the original video frame sequence. The closer the reconstructed video frame sequence is to the original video frame sequence, the closer the reconstructed modal features are to the modal features of the original associated data, and ultimately a model with better video reconstruction results can be trained.
[0126] Example 2
[0127] Please see Figure 2 , Figure 2 This is a flowchart illustrating another model training method based on multimodal correlation data disclosed in an embodiment of the present invention. Figure 2 The described method is applied in a video data processing device, which can be a corresponding processing terminal, processing equipment, or processing server. The server can be a local server or a cloud server; this embodiment of the invention does not impose any limitations. Figure 2 As shown, the model training method based on multimodal association data may include the following operations:
[0128] 201. Determine the training videos to be used to train the model.
[0129] 202. Determine the associated data for at least one modality corresponding to the training video.
[0130] 203. Determine the frame extraction interval corresponding to the training video based on the video parameters of the training video.
[0131] Optionally, video parameters may include image change parameters and / or video scene parameters, which can be used to characterize the complexity of the video content. When the complexity is high, the frame extraction interval should be shortened to obtain more video frames so that the video frames can fully reflect the video content. When the complexity is low, the frame extraction interval can be appropriately increased to obtain fewer video frames.
[0132] 204. Perform frame extraction on the training video according to the frame extraction interval to obtain multiple training video frames corresponding to the training video.
[0133] Optionally, the frame extraction interval described in this invention can be a time interval or a frame number interval.
[0134] Optionally, multiple frame extraction points can be determined based on chronological or frame-by-frame order, with the extraction interval as the interval. Then, the video frames corresponding to the extraction points in the training video are obtained to acquire multiple training video frames. It should be noted that the time interval or frame-by-frame interval between adjacent video frames in the multiple training video frames is not necessarily strictly the same as the aforementioned frame extraction interval. This is because when extracting frames according to chronological or frame-by-frame intervals, if the interval at the end is insufficient, the last frame may be directly determined as a training video frame.
[0135] Optionally, after obtaining multiple training video frames corresponding to the training video, the multiple training video frames can be bound to and saved with the identifier of the training video (such as storage path, video ID, etc.) for subsequent training.
[0136] 205. Input multiple training video frames and associated data into the video reconstruction prediction model for training.
[0137] 206. During training, calculate the video loss function values between multiple predicted video frames output by the video reconstruction prediction model and multiple training video frames input, as well as the modal loss function values between the predicted video and at least one associated data point input.
[0138] 207. Optimize the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence, and obtain the trained video reconstruction prediction model.
[0139] The specific technical details and explanations of technical terms for steps 201-202 and 205-207 above can be found in the description of steps 101-104 in Implementation 1 and the technical details of other steps with the same description, which will not be repeated here.
[0140] As can be seen, implementing the method described in the embodiments of the present invention can determine a reasonable frame extraction interval, thereby extracting appropriate video frames that can represent the video content, resulting in better prediction performance of the trained model.
[0141] As an optional implementation, the frame extraction interval includes multiple different frame extraction intervals. Accordingly, the step described above, performing frame extraction on the training video according to the frame extraction interval to obtain multiple training video frames corresponding to the training video, may include:
[0142] The training video is subjected to frame extraction operations at multiple different frame extraction intervals to obtain multiple training video frame groups corresponding to the training video.
[0143] Optionally, each group of training video frames is used as training data as a single input when training the video reconstruction prediction model.
[0144] Optionally, each group of training video frames may include multiple training video frames.
[0145] Specifically, to augment the dataset, the same training video may be sampled at various intervals, such as 5 frames per second and 2 frames per second, resulting in a difference in the speed of the video frame sequence. Optionally, frame sampling of the training video can be performed before the start of the entire training process. Early frame sampling can significantly reduce training time. If dynamic frame sampling intervals are used to sample short videos during the training phase, although this can greatly enrich the dataset and achieve data augmentation, the frame sampling speed is often slow, requiring re-sampling for each training iteration, which undoubtedly severely hinders the entire training process.
[0146] As an optional implementation, the step described above, which involves performing frame extraction on the training video according to the frame extraction interval to obtain multiple training video frames corresponding to the training video, includes:
[0147] The training video is subjected to frame extraction operation according to the frame extraction interval to obtain multiple candidate video frames corresponding to the training video.
[0148] For any two adjacent candidate video frames, calculate the image similarity between the two candidate video frames;
[0149] Determine whether the image similarity meets the preset similarity threshold condition;
[0150] If the judgment result is yes, then the two candidate video frames are determined as key video frames;
[0151] Based on all the key video frames in the multiple candidate video frames, determine the multiple training video frames corresponding to the training video.
[0152] Optionally, the image similarity can be a structural similarity parameter or a cosine similarity parameter. Optionally, the similarity threshold condition can be that the image similarity is greater than a certain similarity threshold. In this case, the two candidate video frames are not similar, and it can be considered that an image switch occurred between these two frames. These two frames are then retained as critical keyframes.
[0153] As an optional implementation, the step described above, determining the multiple training video frames corresponding to the training video based on all key video frames among the multiple candidate video frames, includes:
[0154] For the candidate video frames other than the key video frame among multiple candidate video frames, a frame extraction operation is performed according to the second frame extraction interval to obtain multiple extracted video frames; wherein, the second frame extraction interval is greater than the frame extraction interval.
[0155] All key video frames and extracted video frames are identified as multiple training video frames corresponding to the training video.
[0156] Optionally, the second frame extraction interval can be determined as follows:
[0157] Determine the total number of frames for all other candidate video frames;
[0158] The second frame skipping interval is determined based on the total number of frames and the preset frame-interval correspondence.
[0159] The second frame extraction interval is directly proportional to the total number of frames. That is, the more total frames there are, the more video content remains. Therefore, the larger the second frame extraction interval, the fewer frames are extracted. This is because for videos with a lot of content but few scene changes, it is not necessary to extract too many frames, and vice versa.
[0160] Optionally, determining multiple training video frames corresponding to the training video based on all key video frames among multiple candidate video frames may also include:
[0161] For multiple candidate video frames other than the key video frame, calculate the image similarity between at least two other candidate video frames.
[0162] Then, all candidate video frames whose image similarity meets the similarity threshold condition are retained to determine multiple training video frames corresponding to the training video.
[0163] Specifically, the video can first be frame-sampling at a relatively small interval. For example, if the video has a frame rate of 30, a 10-frame-per-second extraction interval would yield 50 frames for a 5-second video. Then, the similarity between any two consecutive frames within these 50 frames is calculated. When the similarity exceeds a set maximum threshold m, a scene transition is considered to have occurred between these two frames, and these two frames are retained as critical keyframes. After retaining all keyframes where scene transitions have occurred using this method, consecutive frames other than the keyframes are then subjected to equally spaced frame extraction. The second extraction interval can be determined based on the total length of the final sequence, or it can be adaptively selected by calculating the similarity between consecutive frames. For example, if the similarity between the first and second frames is less than a set minimum threshold n, indicating a small difference and high redundancy, only the first frame is retained, and the second frame is discarded. The third frame is then compared with the first frame; if the similarity is less than the set threshold n, the subsequent frame is discarded. This process continues until the subsequent frame exceeds the minimum threshold n or becomes a critical keyframe, at which point it is retained.
[0164] The keyframe selection method for scene transitions proposed above is intended to enable the model to learn the characteristic changes of scene transitions, so that the final reconstructed video can also exhibit certain scene transition effects or scene continuity.
[0165] As an optional implementation, the step of determining the frame extraction interval corresponding to the training video based on the video parameters of the training video includes:
[0166] Determine the parameters of the image changes in the training video;
[0167] The frame extraction interval corresponding to the training video is determined based on the image change parameters and preset parameter threshold conditions.
[0168] Optionally, the image change parameter can be an optical flow value parameter between different frames of the training video, such as the average, maximum, or weighted average of the optical flow motion between all adjacent frames. Correspondingly, the parameter threshold condition can be an optical flow value threshold condition.
[0169] Preferably, the optical flow value between each frame can be calculated, and the frame skipping interval can be determined by limiting the amount of optical flow movement between each frame. For example, a threshold for the amount of optical flow change can be set. If, after statistical analysis, it is found that the amount of optical flow change every K frames just exceeds this threshold, then the frame skipping interval can be set to K frames. If the amount of optical flow movement between each frame is large, a smaller frame skipping interval can be selected.
[0170] As can be seen, by implementing this optional implementation method, a more reasonable frame extraction interval can be determined based on the image change parameters, thereby enabling the video content of video frames to be determined reasonably and efficiently, thus improving the training efficiency of the model.
[0171] As an optional implementation, the step of determining the frame extraction interval corresponding to the training video based on the video parameters of the training video includes:
[0172] Determine the video scene parameters of the training video;
[0173] The frame extraction interval for the training video is determined based on the video scene parameters and the preset scene frame extraction correspondence.
[0174] Optionally, the video scene parameter can be the scene type of the training video. Optionally, the scene frame extraction correspondence is used to indicate the frame extraction interval corresponding to different scene types. For example, if a short video mainly shows changes in human movement, such as a layup, the movement is relatively fast, and the changes between each frame are relatively obvious; in this case, a smaller frame extraction interval can be used. If the content of the short video changes slowly or follows a certain pattern, such as a car slowly driving on a mountain road, a dashcam recording the road conditions ahead, and trees on both sides moving rhythmically towards the camera; this kind of regular or slow change can use a larger frame extraction interval. Specifically, the frame extraction interval varies depending on the different task scenarios.
[0175] Optionally, the number of different types of scenes appearing in the training video can be used to indicate the degree of scene change or the complexity of the video content. Optionally, the scene frame extraction correspondence is used to indicate the frame extraction interval corresponding to the number of different types of scenes. Generally speaking, the frame extraction interval is inversely proportional to the number of scenes. That is, the more scenes there are, the more changes there are in the video scenes, and the more complex the content is. In this case, the smaller the frame extraction interval, the more video frames are obtained, and vice versa.
[0176] As can be seen, by implementing this optional implementation method, a more reasonable frame extraction interval can be determined based on the video scene parameters, thereby enabling the video content of the video frame to be determined reasonably and efficiently, thus improving the training efficiency of the model.
[0177] As an optional implementation, in the above steps, after performing frame extraction on the training video according to the frame extraction interval to obtain multiple training video frames corresponding to the training video, the method further includes:
[0178] Determine whether the number of multiple training video frames exceeds a preset first frame number threshold;
[0179] If so, divide multiple training video frames into at least two training video frame groups with a number of video frames less than or equal to the first frame number threshold.
[0180] Each group of partitioned training video frames is used as training data as a single input when training the video reconstruction prediction model.
[0181] Specifically, after extracting frames from the video, the final frame count is calculated. To avoid excessive memory usage during training, a first frame count threshold, such as 30 frames, is set. If the result of extracting frames from a video is 54, the first 30 frames should be extracted as a sequence, and the last 24 frames should be extracted as a second sequence to segment the training data of a single input.
[0182] As can be seen, by implementing this optional implementation method, multiple training video frames can be divided into at least two training video frame groups with a number of video frames less than or equal to a first frame number threshold. This allows for the reasonable and efficient determination of training data for a single input, thereby reducing training costs and improving the training efficiency of the model.
[0183] As an optional implementation, in the above steps, after performing frame extraction on the training video according to the frame extraction interval to obtain multiple training video frames corresponding to the training video, the method further includes:
[0184] Determine whether the number of multiple training video frames is less than a preset second frame number threshold;
[0185] If so, extract video frames from the training video and fill them into multiple training video frames until the number of multiple training video frames equals the second frame number threshold.
[0186] Specifically, after extracting frames from the video, the final frame count is calculated. To avoid some videos having too few frames and failing to achieve the expected training effect, a second frame count threshold, such as 30 frames, is set. If the total number of frames in the video is less than 30, such as only 24 frames, 6 frames are randomly selected and copied and inserted in the original time order to complete the 30 frames.
[0187] As can be seen, by implementing this optional implementation method, video frames can be extracted from the training video and filled into multiple training video frames until the number of multiple training video frames is equal to the second frame number threshold. This can reasonably and efficiently complete the amount of training data input in a single instance, thereby improving the training efficiency and effectiveness of the model.
[0188] Example 3
[0189] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a model training device based on multimodal correlation data disclosed in an embodiment of the present invention. Figure 3 The described apparatus can be applied to corresponding video data processing devices, which can be corresponding processing terminals, processing equipment, or processing servers. The server can be a local server or a cloud server; this embodiment of the invention does not impose limitations. Figure 3 As shown, the device may include:
[0190] The video determination module 301 is used to determine the training video for training the model;
[0191] The association determination module 302 is used to determine the association data of at least one modality corresponding to the training video;
[0192] The model training module 303 is used to train the video reconstruction prediction model by inputting training videos and associated data. During training, it calculates the video loss function value between the predicted video output by the video reconstruction prediction model and the input training video, as well as the modal loss function value between the predicted video and at least one input associated data.
[0193] The model optimization module 304 is used to optimize the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence, so as to obtain the trained video reconstruction prediction model.
[0194] As an optional implementation, the modality includes at least one of an audio modality, a text modality, and an image modality; and / or, the associated data includes at least one of descriptive audio data, descriptive text data, and characterizing image data.
[0195] As an optional implementation method, such as Figure 4 As shown, the model training module 303 includes:
[0196] The frame extraction unit 3031 is used to perform frame extraction on the training video to obtain multiple training video frames corresponding to the training video.
[0197] The model training unit 3032 is used to train the video reconstruction prediction model by inputting multiple training video frames and associated data into the video reconstruction prediction model.
[0198] The loss calculation unit 3033 is used to calculate, during training, the video loss function value between multiple predicted video frames output by the video reconstruction prediction model and multiple training video frames input, as well as the modal loss function value between the predicted video and at least one associated data input.
[0199] As an optional implementation, the associated data is characterizing image data, and the association determination module 302 determines the specific method of the associated data for at least one modality corresponding to the training video, including:
[0200] The target representation frame image is determined from multiple training video frames, and the target representation frame image is copied to obtain multiple copied representation frame images;
[0201] Multiple replicated representation frame images are identified as the representation image data corresponding to the training video.
[0202] As an optional implementation, the frame extraction unit 3031 performs frame extraction on the training video to obtain multiple training video frames, including the following specific methods:
[0203] Determine the frame extraction interval corresponding to the training video based on the video parameters of the training video;
[0204] The training video is subjected to frame extraction at the extraction interval to obtain multiple training video frames corresponding to the training video.
[0205] As an optional implementation, the video reconstruction prediction model includes a video reconstruction network and a modality reconstruction network; the video reconstruction network is used to receive training videos and associated data and reconstruct the predicted video; the modality reconstruction network is used to extract the associated data features of the modality corresponding to the predicted video; the associated data features are used to compare with the associated data to calculate the modality loss function value; the modality reconstruction network is trained and converged through a training dataset including multiple training videos and corresponding modality training associated data;
[0206] Furthermore, the model optimization module 304 optimizes the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence, obtaining the specific method of the trained video reconstruction prediction model, including:
[0207] During training, the parameters of the modal reconstruction network are kept constant. The network parameters of the video reconstruction network are optimized based on the video loss function value and the modal loss function value until convergence, resulting in a well-trained video reconstruction network.
[0208] As an optional implementation, the specific method by which the model optimization module 304 optimizes the model parameters of the video reconstruction prediction model until convergence based on the video loss function value and the modal loss function value includes:
[0209] Calculate the weighted sum of the video loss function value and the modal loss function value;
[0210] The model parameters of the video reconstruction prediction model are optimized based on the weighted summation until convergence.
[0211] As an optional implementation method, the video loss function value is calculated as follows:
[0212] For any predicted video frame, calculate the frame loss function value between the predicted video frame and the corresponding training video frame;
[0213] Calculate the average of the frame loss function values for all predicted video frames to obtain the loss function values between multiple predicted video frames and multiple input training video frames.
[0214] Example 4
[0215] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of another model training device based on multimodal correlation data disclosed in an embodiment of the present invention. For example... Figure 5 As shown, the device may include:
[0216] Memory 401 storing executable program code;
[0217] Processor 402 coupled to memory 401;
[0218] The processor 402 calls the executable program code stored in the memory 401 to execute some or all of the steps in the model training method based on multimodal association data disclosed in Embodiment 1 or Embodiment 2 of the present invention.
[0219] Example 5
[0220] This invention discloses a computer storage medium storing computer instructions. When these computer instructions are invoked, they are used to execute some or all of the steps in the model training method based on multimodal association data disclosed in Embodiment 1 or Embodiment 2 of this invention.
[0221] The foregoing has described specific embodiments of this specification; other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than those shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily have to follow the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0222] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer-readable storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0223] The apparatus, device, non-volatile computer-readable storage medium and method provided in the embodiments of this specification are corresponding. Therefore, the apparatus, device and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, device and non-volatile computer storage medium will not be repeated here.
[0224] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0225] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0226] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0227] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.
[0228] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0229] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0230] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0231] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0232] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0233] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0234] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0235] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0236] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0237] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0238] Finally, it should be noted that the model training method and apparatus based on multimodal association data disclosed in the embodiments of the present invention are merely preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, not to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A model training method based on multimodal association data, characterized in that, The method includes: Identify the training videos to be used to train the model; Determine the associated data for at least one modality corresponding to the training video, wherein the modality includes at least one of audio modality, text modality, and image modality, and the associated data includes at least one of descriptive audio data, descriptive text data, and representative image data; The training video and the associated data are input into a video reconstruction prediction model for training. During training, the video loss function value between the predicted video output by the video reconstruction prediction model and the input training video, and the modal loss function value between the predicted video and at least one of the input associated data are calculated. The video reconstruction prediction model includes a video reconstruction network and a modal reconstruction network. The video reconstruction network is used to receive the training video and the associated data and reconstruct the predicted video. The modal reconstruction network is used to extract the associated data features of the modality corresponding to the predicted video. The associated data features are compared with the associated data to calculate the modal loss function value. The model parameters of the video reconstruction prediction model are optimized based on the video loss function value and the modal loss function value until convergence, resulting in the trained video reconstruction prediction model.
2. The model training method based on multimodal association data according to claim 1, characterized in that, The step of training the video reconstruction prediction model by inputting the training video and the associated data, and calculating the video loss function value between the predicted video output by the video reconstruction prediction model and the input training video, and the modal loss function value between the predicted video and at least one of the input associated data, includes: Perform frame extraction on the training video to obtain multiple training video frames corresponding to the training video. The multiple training video frames and the associated data are input into the video reconstruction prediction model for training; In the training, the video loss function values between the multiple predicted video frames output by the video reconstruction prediction model and the multiple training video frames input are calculated, as well as the modal loss function values between the predicted video and at least one of the input associated data.
3. The model training method based on multimodal association data according to claim 2, characterized in that, The associated data is image data, and the associated data for determining at least one modality corresponding to the training video includes: The target representation frame image is determined from the plurality of training video frames, and the target representation frame image is copied to obtain a plurality of copied representation frame images; The plurality of replicated representation frame images are determined as the representation image data corresponding to the training video.
4. The model training method based on multimodal association data according to claim 1, characterized in that, The step of performing frame extraction on the training video to obtain multiple training video frames corresponding to the training video includes: The frame extraction interval corresponding to the training video is determined based on the video parameters of the training video. The training video is subjected to frame extraction operation according to the frame extraction interval to obtain multiple training video frames corresponding to the training video.
5. The model training method based on multimodal association data according to claim 1, characterized in that, The modality reconstruction network is trained and converged using a training dataset that includes multiple training videos and corresponding training association data of the modalities. And, the step of optimizing the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence, to obtain the trained video reconstruction prediction model, includes: During training, the parameters of the modality reconstruction network are kept constant. The network parameters of the video reconstruction network are optimized based on the video loss function value and the modality loss function value until convergence, thus obtaining the trained video reconstruction network.
6. The model training method based on multimodal association data according to claim 1, characterized in that, The step of optimizing the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence includes: Calculate the weighted sum of the video loss function value and the modal loss function value; Based on the weighted sum, the model parameters of the video reconstruction prediction model are optimized until convergence.
7. The model training method based on multimodal association data according to claim 6, characterized in that, The video loss function value is calculated as follows: For any of the predicted video frames, calculate the frame loss function value between the predicted video frame and the corresponding training video frame; Calculate the average of the frame loss function values for all the predicted video frames to obtain the loss function values between the multiple predicted video frames and the multiple input training video frames.
8. A model training device based on multimodal correlation data, characterized in that, The device includes: The video determination module is used to determine the training videos to be used for training the model; The association determination module is used to determine the association data of at least one modality corresponding to the training video. The modality includes at least one of audio modality, text modality and image modality. The association data includes at least one of descriptive audio data, descriptive text data and characterizing image data. The model training module is used to train a video reconstruction prediction model by inputting the training video and the associated data into the model. During training, it calculates the video loss function value between the predicted video output by the model and the input training video, and the modal loss function value between the predicted video and at least one piece of the input associated data. The video reconstruction prediction model includes a video reconstruction network and a modal reconstruction network. The video reconstruction network receives the training video and the associated data and reconstructs the predicted video. The modal reconstruction network extracts the associated data features corresponding to the modality of the predicted video. The associated data features are compared with the associated data to calculate the modal loss function value. The model optimization module is used to optimize the model parameters of the video reconstruction prediction model based on the video loss function value and the modal loss function value until convergence, so as to obtain the trained video reconstruction prediction model.
9. A model training device based on multimodal correlation data, characterized in that, The device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the model training method based on multimodal association data as described in any one of claims 1-7.
Citation Information
Patent Citations
Multi-modal feature fusion text-guided image restoration method
CN111340122A
Multi-modal feature extraction model training method and device, and electronic equipment
CN113486833A