Method and apparatus for generating video data, non-volatile storage medium
By using deep learning models to analyze and repair image blocks in video data and generate target video data, the problem of low amplification efficiency in the prior art is solved, and the effect of providing sufficient deep learning model training samples is achieved.
Patent Information
- Application Number
- CN202411274986.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-09-11
AI Technical Summary
The prior art cannot efficiently amplify video data, resulting in the inability to provide sufficient training samples for deep learning models.
By acquiring the initial video data, the original image is analyzed using the first deep learning model, the attention map is generated and the target attention map is determined, which is converted into a masked image block, and the second deep learning model is used for repair processing. Finally, the repaired image block is fused with the original image to generate the target video data.
It realizes efficient amplification of video data, provides sufficient training samples for deep learning models, and solves the problem of low video data amplification efficiency.
Smart Images

Figure CN119233045B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video data enhancement, and in particular, to a method and apparatus for generating video data, and a non-volatile storage medium. Background Art
[0002] Deep learning provides an effective optimization framework for automatically and dynamically extracting intrinsic patterns from continuous observable processes. Deep learning relies on large-scale observable data to capture implicit patterns as substitutes for physical laws. During the training of deep learning, the following problems exist: 1. High-quality / resolution video data is relatively scarce, and the training cost using such data is extremely high; the sensors on the earth are extremely unevenly distributed, and many regions cannot effectively utilize this data due to data scarcity. Although some efforts have been made to solve this problem, such as transfer learning and active learning, data-driven methods still lack interpretability, resulting in a lack of generalization ability during the transfer process and poor performance in some extreme scenarios, such as tracking cyclones and sensing turbulence. This remains a common challenge faced by deep learning models. 2. Designs customized for specific tasks endow the model with special capabilities and high performance. However, complex designs make it difficult for the model to generalize. Currently, common video data augmentation methods include cropping, flipping, image brightness adjustment, etc. The current methods have certain limitations.
[0003] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] Embodiments of this application provide a method and apparatus for generating video data, and a non-volatile storage medium, to at least solve the technical problem that the related art cannot efficiently augment video data, resulting in the inability to provide sufficient training samples for deep learning models.
[0005] According to an aspect of the embodiments of this application, a method for generating video data is provided, including: obtaining initial video data; analyzing the original images in the initial video data using a first deep learning model to obtain a plurality of image blocks output by the deep learning model and the attention maps respectively corresponding to the plurality of image blocks, where the attention map includes: the attention score corresponding to the image block; determining a target attention map in the attention map whose attention score satisfies a first preset condition; converting the target image block corresponding to the target attention map into a masked image block; using a second deep learning model to perform a repair process on the masked image block to obtain a target image block; performing a fusion process on the target image block and the original image to obtain a target image, and updating the target image to the initial video data to obtain target video data.
[0006] Optionally, the second deep learning model includes: a U-Net model, where the U-Net model includes: a second variational autoencoder and a second variational decoder; the second deep learning model is obtained by the following method: extracting n frames of first images other than the original image from the initial video data, performing feature extraction on the first images to obtain first image features; obtaining text information for describing the first images, performing feature extraction on the text information to obtain text features; using the second variational autoencoder to compress the first images from the image space to the latent space to obtain second image features, where the dimension of the latent space is smaller than the dimension of the image space; determining the attention weights between the text features and the second image features, and according to the attention weights, performing weighted summation on the first image features to obtain third image features; using the second variational decoder to restore the third image features from the latent space to the image space to obtain fourth image features; determining a first loss function according to the error between the first image features and the fourth image features, and obtaining the model parameters in the U-Net model when the first loss function satisfies a second preset condition; using the model weight parameters in the U-Net model to replace the model weight parameters in the second deep learning model to obtain the trained second deep learning model.
[0007] Optionally, using the second deep learning model to perform repair processing on the masked image block to obtain a target image block includes: using the second variational autoencoder to compress the masked image block from the image space to the latent space to obtain a first feature; performing element-wise multiplication processing on the pixel values of the masked image block and the original image to obtain a second image; adding noise to the first feature within the target region corresponding to the second image to obtain a second feature; performing denoising processing on the second feature for n iterations: adding the second feature in the i-th iteration to the predicted noise in the (i + 1)-th iteration determined according to the second deep learning model to obtain the second feature in the (i + 1)-th iteration, where n is a positive integer greater than 1, and i is a positive integer not greater than n; using the second variational decoder to restore the second feature obtained in the n-th iteration from the latent space to the image space to obtain a target image block.
[0008] Optionally, the first deep learning model includes: a first variational autoencoder and a first variational decoder. The first deep learning model is trained by the following method: Obtain training video data, where the training video data includes: training images; Segment the training images into multiple training image patches of a preset size; Convert each training image patch into a first vector respectively to obtain a plurality of first vectors, where the first vector is a one-dimensional vector; Determine the position encoding information corresponding to each training image patch respectively to obtain a plurality of position encoding information; Input the plurality of first vectors and the plurality of position encoding information into the first variational autoencoder to obtain the target vector output by the first variational autoencoder; Determine the loss function according to the error between the target vector and the vector corresponding to the training image patch, and complete the training of the deep learning model when the loss function meets the second preset condition.
[0009] Optionally, the first variational autoencoder includes: a plurality of identical encoder layers, and each encoder layer includes: a multi-head self-attention mechanism and a feed-forward neural network.
[0010] Optionally, the attention map is a row vector; Determining the target attention map whose attention score meets the first preset condition in the attention map includes: determining the value corresponding to each row vector to obtain a plurality of values; Determining the row vector with the smallest value among the plurality of values as the target attention map.
[0011] Optionally, before converting the target image patch corresponding to the target attention map into a masked image patch, the method further includes: determining and storing the position information of the target image patch in the original image; Performing a fusion process on the target image patch and the original image to obtain a target image, including: based on the position information, replacing the target image patch in the original image with the target image patch to obtain the target image.
[0012] According to another aspect of the embodiments of the present application, there is also provided a video data generation device, including: an acquisition module, configured to acquire initial video data; a first determination module, configured to analyze the original images in the initial video data by using a first deep learning model to obtain a plurality of image patches output by the deep learning model and the attention maps respectively corresponding to the plurality of image patches, where the attention map includes: the attention score corresponding to the image patch; a second determination module, configured to determine a target attention map whose attention score meets the first preset condition in the attention map; a conversion module, configured to convert the target image patch corresponding to the target attention map into a masked image patch; a filling module, configured to perform a repair process on the masked image patch by using a second deep learning model to obtain a target image patch; a third determination module, configured to perform a fusion process on the target image patch and the original image to obtain a target image, and update the target image to the initial video data to obtain target video data.
[0013] According to another aspect of the embodiments of the present application, a non-volatile storage medium is further provided. The storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the above method for generating video data.
[0014] According to another aspect of the embodiments of the present application, an electronic device is further provided, including: a memory and a processor. The processor is used to run a program stored in the memory, and when the program runs, it executes the above method for generating video data.
[0015] According to another aspect of the embodiments of the present application, a computer program is further provided, and when the computer program is executed by a processor, it implements the above method for generating video data.
[0016] According to another aspect of the embodiments of the present application, a computer program product is further provided. The computer program product includes a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above method for generating video data.
[0017] In the embodiments of the present application, the following method is adopted: obtaining initial video data; analyzing the original images in the initial video data by using a first deep learning model to obtain multiple image blocks output by the deep learning model and the attention maps respectively corresponding to the multiple image blocks, where the attention maps include: the attention scores corresponding to the image blocks; determining target attention maps in the attention maps whose attention scores meet a first preset condition; converting the target image blocks corresponding to the target attention maps into mask image blocks; using a second deep learning model to perform repair processing on the mask image blocks to obtain target image blocks; performing fusion processing on the target image blocks and the original images to obtain target images, and updating the target images to the initial video data to obtain target video data. In this way, the purpose of efficiently amplifying video data is achieved, thereby realizing the technical effect of providing sufficient training samples for the deep learning model, and further solving the technical problem that due to the inability of related technologies to efficiently amplify video data, sufficient training samples for the deep learning model cannot be provided. Description of the Drawings
[0018] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0019] Figure 1 is a flowchart of a method for generating video data according to an embodiment of the present application;
[0020] Figure 2It is a flowchart of another method for generating video data according to an embodiment of the present application;
[0021] Figure 3 It is a flowchart of a method for regenerating the region with the lowest attention score according to an embodiment of the present application;
[0022] Figure 4 It is a structural diagram of a video data generation device according to an embodiment of the present application;
[0023] Figure 5 It is a hardware structure block diagram of a computer terminal for a method of generating video data according to an embodiment of the present application. Detailed implementation manners
[0024] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0026] According to an embodiment of the present application, a method embodiment of a method for generating video data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0027] Figure 1 It is a flowchart of a method for generating video data according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:
[0028] Step S101, obtain initial video data.
[0029] For example, the initial video data is video data in the KTH format. Among them, the initial video data includes multiple original images.
[0030] Step S102: Analyze the original images in the initial video data using the first deep learning model to obtain multiple image patches output by the deep learning model and the attention maps corresponding to the multiple image patches respectively. Among them, the attention map includes: the attention score corresponding to the image patch.
[0031] The deep learning model is, for example, a ViT (Vision Transformer) model. Among them, the attention mechanism in the ViT model is used to quantify the degree of association between various parts of the input data, helping the model learn the hierarchical structure and feature relationships existing in the input image. Its core idea is to calculate the correlation between each image patch and other image patches, so as to determine the importance weight of each image patch.
[0032] For example, the attention map corresponding to each image patch can be determined by the following method: First, divide the input image into several image patches of a fixed size, and convert each image patch into a low-dimensional vector representation (embedding vector) through a linear mapping. Then, in the multi-head self-attention layer, perform attention calculation on these embedding vectors. Specifically, for each head, obtain an attention matrix by calculating the query vector, key vector, and value vector. The query vector is used to determine the position to be focused on, the key vector is used to match with the query vectors at other positions, and the value vector is the feature representation of the corresponding position. Finally, splice and linearly transform the attention matrices of multiple heads to obtain the final attention output. The above output contains the attention scores of each image patch to other image patches, that is, the attention map.
[0033] It can be understood that the attention map can be presented in the form of a two-dimensional matrix. The rows and columns of the matrix respectively correspond to different image patches of the input image. The element value in the matrix represents the attention weight between the corresponding two image patches. The larger the weight, the more the model pays attention to the association between these two image patches when processing information. For example, the position with a larger element value indicates that the corresponding two image patches have a strong correlation in the feature extraction and classification process.
[0034] Step S103: Determine the target attention map in the attention map whose attention score meets the first preset condition.
[0035] Methods for attention scores include summation, averaging, taking the maximum value, etc. For example, all pixel values in the attention map can be added as the attention score, or the average value of the attention map can be calculated as the score.
[0036] For example, the attention map with the maximum attention score is determined as the above-mentioned target attention map. Specifically: 1. Create a variable to store the current maximum attention score and the corresponding attention map. The maximum score can be initialized to a relatively small value so that it can be updated correctly during the comparison process. 2. Traverse all the calculated attention maps and their corresponding scores. For each attention map, compare its score with the current maximum score. 3. If the score of the current attention map is greater than the current maximum score, update the maximum score and the corresponding attention map. Update the maximum score to the score of the current attention map and save the corresponding attention map. Repeat steps 2 and 3: continue to traverse the remaining attention maps until all attention maps are processed.
[0037] Step S104: Convert the target image patch corresponding to the target attention map into a mask image patch.
[0038] Among them, the mask image patch is an array or matrix, and its size is the same as the original image. When using the mask, only the corresponding pixels with a value of 1 (or True) in the mask array are processed, while the pixels with a value of 0 (or False) are ignored.
[0039] Step S105: Use the second deep learning model to perform repair processing on the mask image patch to obtain the target image patch.
[0040] The second deep learning model is, for example, an image inpainting model based on a convolutional neural network. This model includes: U-Net, where U-Net consists of an encoder and a decoder. The encoder extracts multi-scale features of the image, the decoder restores the extracted features to an image, and skip connections fuse the features of different levels of the encoder with the features of the decoder. U-Net can effectively utilize the context information and multi-scale features of the image to accurately repair the missing part.
[0041] Step S106: Perform fusion processing on the target image patch and the original image to obtain the target image, and update the target image to the initial video data to obtain the target video data.
[0042] According to the above steps, the method includes: obtaining initial video data; analyzing the original images in the initial video data using a first deep learning model to obtain multiple image patches output by the deep learning model and the attention maps respectively corresponding to the multiple image patches, where the attention map includes: the attention score corresponding to the image patch; determining a target attention map in the attention maps whose attention score meets a first preset condition; converting the target image patch corresponding to the target attention map into a masked image patch; using a second deep learning model to perform repair processing on the masked image patch to obtain the target image patch; performing fusion processing on the target image patch and the original image to obtain a target image, and updating the target image to the initial video data to obtain target video data, thereby achieving the purpose of efficiently augmenting video data and realizing the technical effect of providing sufficient training samples for the deep learning model.
[0043] The following gives an exemplary illustration and explanation of Figure 1 the steps shown.
[0044] According to some optional embodiments of the present application, the second deep learning model includes: a U-Net model, where the U-Net model includes: a second variational autoencoder and a second variational decoder. The second deep learning model is obtained by training through the following method:
[0045] Extracting n frames of first images other than the original images from the initial video data, performing feature extraction on the first images to obtain first image features; obtaining text information for describing the first images, performing feature extraction on the text information to obtain text features; using the second variational autoencoder to compress the first images from the image space to the latent space to obtain second image features, where the dimension of the latent space is smaller than the dimension of the image space; determining the attention weights between the text features and the second image features, and performing weighted summation on the first image features according to the attention weights to obtain third image features; using the second variational decoder to restore the third image features from the latent space to the image space to obtain fourth image features; determining a first loss function according to the error between the first image features and the fourth image features, and obtaining the model parameters in the U-Net model when the first loss function meets a second preset condition; replacing the model weight parameters in the second deep learning model with the model weight parameters in the U-Net model to obtain the trained second deep learning model.
[0046] The second variational autoencoder compresses the first images into a low-dimensional latent space and extracts the key features of the images. For example, for a natural scenery image, the second variational autoencoder can extract features such as color, texture, and object shape and represent them as a latent vector.
[0047] When training the second variational autoencoder, the parameters of the encoder and decoder are optimized by minimizing the reconstruction loss and the regularization term. The reconstruction loss ensures that the decoder can recover an image similar to the original image from the latent vector, and the regularization term constrains the distribution of the latent space to have good properties.
[0048] The Contrastive Language-Image Pretraining (CLIP) model is used to extract features from the text. Among them, CLIP is a model pre-trained on large-scale image and text datasets, which can effectively convert text descriptions into high-dimensional feature vectors. For example, for the text description "beautiful sunset", CLIP can extract features related to the sunset such as colors, atmosphere, objects, etc., and represent them as a text feature vector.
[0049] Furthermore, a text-image attention module is constructed to control the diffusion of image features in the latent space. The text-image attention module can adjust the distribution of image features in the latent space according to the correlation between the text feature vector and the image latent vector. For example, if the text description emphasizes a specific object or color in the image, the text-image attention module can enhance the image features related to that object or color while suppressing other irrelevant features.
[0050] Preferably, the text-image attention module can be implemented using techniques such as multi-head attention mechanism and attention pooling. By calculating the attention weights between the text features and the image features, the text information is fused into the image features, thereby guiding the image generation process.
[0051] Furthermore, the image latent vector processed by the text-image attention module is input into the second variational auto-decoder. The second variational auto-decoder gradually recovers the pixel values of the image according to the feature information in the latent vector, and generates an image that conforms to the text description.
[0052] For example, during the training process of the second deep learning model, the optimizer used is AdamW, the initial learning rate is 0.00001, the loss function is the mean squared error loss, the gradients of the VAE and CLIP modules are set to False, the gradient of the U-Net is set to True, and the number of training epochs is 100.
[0053] Finally, the weights of the trained U-Net are replaced into the weights of the Inpanting model to complete the weight update of the Inpanting model. By updating the U-Net weights to obtain the learning of the Inpanting model for new data images, better high-quality images can be generated.
[0054] Further, the masked image patch is processed by a second deep learning model to obtain a target image patch, which can be achieved by the following method: The masked image patch is compressed from the image space to the latent space by a second variational autoencoder to obtain a first feature; the masked image patch is multiplied element by element with the pixel values of the original image to obtain a second image; noise is added to the first feature within the target region corresponding to the second image to obtain a second feature; the second feature is subjected to n iterations of denoising processing: the second feature in the i-th iteration is added to the predicted noise in the (i + 1)-th iteration determined according to the second deep learning model to obtain the second feature in the (i + 1)-th iteration, where n is a positive integer greater than 1, and i is a positive integer not greater than n; the second feature obtained in the n-th iteration is restored from the latent space to the image space by a second variational auto-decoder to obtain a target image patch.
[0055] That is to say, first, image features are extracted by a second variational autoencoder and mapped into the VAE latent space. Then, noise is added to the image features based on the DDIM scheduler. Then, a cyclic denoising process is performed. In the cyclic denoising process, the masked region is multiplied by the original image, so that the diffusion generation is limited within the masked interval. In each iteration, the U-Net model predicts the next-step noise and adds it to the current feature map to complete the denoising process. Finally, the image is restored by a second variational auto-decoder. Through the above steps, the regenerated image data is different from the original image, and the original image can be specifically changed to achieve the purpose of image augmentation.
[0056] According to some other optional embodiments of the present application, the first deep learning model includes: a first variational autoencoder and a first variational auto-decoder, and the first deep learning model is obtained by training through the following method: training video data is obtained, where the training video data includes: training images; the training images are segmented into multiple training image patches of a preset size; each training image patch is respectively converted into a first vector to obtain a plurality of first vectors, where the first vector is a one-dimensional vector; the position encoding information corresponding to each training image patch is respectively determined to obtain a plurality of position encoding information; the plurality of first vectors and the plurality of position encoding information are input into the first variational autoencoder to obtain a target vector output by the first variational autoencoder; according to the error between the target vector and the vector corresponding to the training image patch, a loss function is determined, and when the loss function meets the second preset condition, the training of the deep learning model is completed.
[0057] Specifically, first, the input image is segmented into image patches of a fixed size, where the size of each image patch is 16X16 pixels. The image is divided into 64 image patches, and each image patch is flattened into a one-dimensional vector with a dimension of 256, and then mapped to a fixed dimension of 256 through a linear layer. To retain the position information of the image patches, a position encoding is added to each image patch embedding. The position encoding is a vector with the same dimension as the image patch embedding and is used to represent the position of the image patch in the original image. The sum of the image patch embedding and the position encoding serves as the input to the first variational autoencoder, which is stacked by multiple identical layers, and each layer includes a multi-head self-attention mechanism and a feed-forward neural network. The first variational auto-decoder restores the output sequence of the first variational autoencoder to image patches, calculates the loss between the input image and the generated image, and thus autoregressive training can be performed.
[0058] Preferably, the size of the training image is (128, 128), the height and width of each cut image patch are 16, the number of channels is 1, there are two multi-head attention mechanisms, the Adam optimizer is used, the initial learning rate is 0.001, the number of training times is 100, the learning rate adjustment strategy is the cosine annealing strategy, the loss function is the mean squared error loss function, and the optimal weights are saved.
[0059] Preferably, the first variational autoencoder includes: multiple identical encoder layers, and each encoder layer includes: a multi-head self-attention mechanism and a feed-forward neural network.
[0060] Preferably, the attention map is a row vector. To determine the target attention map whose attention scores satisfy the first preset condition in the attention map, it can be achieved by the following method: determine the numerical values corresponding to each row vector to obtain a plurality of numerical values; determine the row vector with the smallest numerical value among the plurality of numerical values as the target attention map.
[0061] Specifically, first, for each row vector, determine a method to calculate its corresponding value. For example, the norm of the row vector (such as L1 norm, L2 norm, etc.), the sum of the elements in the row vector, the average value of the elements in the row vector, etc. can be calculated. If calculating the L1 norm, that is, summing the absolute values of all elements in the row vector; if calculating the sum, directly add all elements in the row vector; if calculating the average value, add all elements in the row vector and then divide by the length of the row vector. Secondly, create a variable to store the current minimum value and the corresponding row vector. The minimum value can be initialized to a relatively large value so that it can be correctly updated during the comparison process. Thirdly, traverse all row vectors and their corresponding values. For each row vector, compare its corresponding value with the current minimum value. If the value corresponding to the current row vector is less than the current minimum value, update the minimum value and the corresponding target row vector. Update the minimum value to the value of the current row vector and save the corresponding row vector. Finally, continue to traverse the remaining row vectors until all row vectors are processed.
[0062] In some alternative embodiments of the present application, before converting the target image block corresponding to the target attention map into a mask image block, the following steps may also be performed: Determine the position information of the target image block in the original image and store the above position information in the memory. Further, perform a fusion process on the target image block and the original image to obtain a target image, which can be achieved by the following method: Search for the previously stored position information in the memory, and based on the found position information, use the target image block to replace the target image block in the original image to obtain the target image.
[0063] Figure 2 is a flowchart of another method for generating video data according to an embodiment of the present application, as Figure 2 shown, the method includes the following steps:
[0064] Step S201, obtain video data and process the video data into an npy format file.
[0065] Specifically, use KTH video data. Each video data has 20 frames of video. A total of 500 video data are used. The dataset is processed into an npy file to make a dataset. 350 videos are taken as the training set and 150 videos are taken as the validation set.
[0066] Step S202: Extract the video data into separate images, and then construct a ViT model for autoregressive training to implement an Encode-Decode structure. Specifically, first divide the input image into N image patches (patches) of a fixed size, where the size of each image patch is P×P pixels. Each image patch is flattened into a one-dimensional vector and mapped to a fixed dimension D through a linear layer. To preserve the position information of the image patches, a position encoding is added to each image patch embedding. The position encoding is a vector with the same dimension as the image patch embedding, used to represent the position of the image patch in the original image. The sum of the image patch embedding and the position encoding serves as the input to the Transformer encoder, which is stacked by multiple identical layers, and each layer includes a multi-head self-attention mechanism and a feed-forward neural network. Decode restores the vectors to image patches in the output sequence of the Transformer encoder, calculates the loss between the input image and the generated image, and completes the autoregressive training.
[0067] Specifically, first divide the input image into image patches of a fixed size. The size of each image patch is 16X16 pixels. The image is divided into 64 image patches, each of which is flattened into a one-dimensional vector with a dimension of 256 and mapped to a fixed dimension of 256 through a linear layer. To preserve the position information of the image patches, a position encoding is added to each image patch embedding. Among them, the position encoding is a vector with the same dimension as the image patch embedding, used to represent the position of the image patch in the original image. The sum of the image patch embedding and the position encoding serves as the input to the first variational autoencoder, which is stacked by multiple identical layers, and each layer includes a multi-head self-attention mechanism and a feed-forward neural network. The first variational auto-decoder restores the output sequence of the first variational autoencoder to image patches, calculates the loss between the input image and the generated image, and thus can perform autoregressive training.
[0068] Preferably, the size of the training image is (128, 128), the height and width of each cut image patch are 16, the number of channels is 1, there are two multi-head attention mechanisms, the Adam optimizer is used, the initial learning rate is 0.001, the number of training times is 100, the learning rate adjustment strategy is the cosine annealing strategy, the loss function is the mean square error loss function, and the optimal weights are saved.
[0069] Through the above steps, important region features are learned during the training process to complete the autoregressive task. The parameters of the trained model are fixed. When the image is input into the network again, important region features will be automatically extracted to complete the reconstruction task, and low-attention features are obtained through the calculated attention feature map.
[0070] Furthermore, re-enter the training set into the model. During each calculation of the image, extract the attention map under the first module. The size of the attention image is 64×64, which represents the influence degree of each image patch on other image patches. By calculating the sum of each row of the feature map and then calculating the minimum value of all rows, the smallest region among the 64 image patches can be obtained. Calculate the coordinates of the specific image patch based on which block it is, and then generate a Mask image occlusion. The same Mask image is used for each video. It is used for subsequent diffusion to generate the image of this region. By inversely calculating the features of the unimportant region in the way of extracting image features through the Vit architecture, generating the mask image of this region is a very effective perturbation augmentation method.
[0071] Specifically, in step S203, when using the trained ViT model to regenerate the image, each time an image is input, the Transformer structure will generate a corresponding attention map. By extracting these attention maps, the region with the lowest attention score in the image can be calculated. It can be understood that these regions represent the parts with lower attention from the model and may need further repair. Generate a mask Mask image for each image according to the attention map. These mask images identify the regions with the worst attention scores. Extract 5 frames of images from each video. Use the same text description for the entire fine-tuned image data. Freeze the text module during the training process and generate images through the text. Calculate the loss value between the generated image and the original image. The purpose is to let the U-Net model learn the overall image feature distribution and fine-tune and update the weight model of U-Net by fine-tuning the text-to-image model.
[0072] Specifically, extract images from the KTH video data again. Extract 2 frames of images from each video to create a dataset for fine-tuning the diffusion model. Use the training method of text-to-image to update the weights of U-Net. Use the same text description for all images. Build a Stable Diffusion model. Extract features of the image and text respectively using different encoders. Use the VAE module to extract features of the image and use CLIP to extract features of the text. Control the diffusion of image features in the latent space through the image-text attention module. Finally, use the VAE module for image restoration. The optimizer used for training is AdamW, and the initial learning rate is 0.00001. The loss function is the mean-variance loss. The gradients of the VAE and CLIP modules are set to False, and the gradient of the U-Net is set to True. Train for 100 batches. Replace the saved U-Net weights with the U-Net in the Inpanting weights to complete the weight update of Inpanting. By updating the U-Net weights to obtain the learning of the Inpanting model for the new data image, better high-quality images can be generated.
[0073] Step S204: Replace the U-Net weights of the image inpainting diffusion model, and finally regenerate the region with the lowest attention score through the image inpainting diffusion model to complete the augmentation of video data.
[0074] Figure 3 It is a flowchart of a method for regenerating the region with the lowest attention score according to an embodiment of the present application. As Figure 3 shown, the method includes the following steps: First, extract image features through the second variational autoencoder and map them into the VAE latent space. Then, add noise to the image features based on the DDIM scheduler. Next, perform a cyclic denoising process. During the cyclic denoising process, multiply the masked region by the original image so that the diffusion generation is limited within the masked interval. In each iteration, the U-Net model predicts the next-step noise and adds it to the current feature map to complete the denoising process. Finally, restore the image through the second variational auto-decoder. Through the above steps, the regenerated image data is different from the original image, and the original image can be specifically changed to achieve the purpose of image augmentation.
[0075] In summary, in this embodiment, the model of the Transformer architecture is autoregressively trained. After the training is completed, a weight is obtained. The model infers and calculates the image again, extracts the attention map of the first block of the Transformer, calculates the states in each attention map, and obtains the region of the image patch with the weakest attention by summing each row of the attention map. Through this method, the image patches not focused on by the network can be obtained, and new images can be generated by regenerating the image patches through diffusion. Retraining the network can effectively improve the model effect and generalization. In addition, by training text-to-image to learn the video distribution characteristics, the global feature distribution of the video can be captured. Using the U-Net weights to update the image inpainting model can complete the filling from the overall distribution when generating image patches, and will be as different as possible from the original image patches. However, the training method of the image inpainting module is autoregressive image patches, which will cause the newly generated image patches to be as similar as possible to the original image.
[0076] It should be noted that the content of evaluating and assessing the above-trained recognition model is as follows: The evaluation metrics used are MAE and MSE, which are metrics for measuring the difference between the predicted value and the true value of the image. The smaller the MAE, the more accurate the model and the better the training effect. Verify whether the model effect can be improved by using different proportions of data on a model. This proportion refers to the size of the training set used in the entire training set. Generally, the effect is better when the data set is smaller. Among them, Table 1 shows the improvement percentages of different proportions of data, and Table 2 is used to reflect the effectiveness of using multiple models on the data set.
[0077] Model Data Ratio MAE Improvement Ratio ViT 10% 57% ViT 20% 25% ViT 50% 9% ViT 70% 11% ViT 100% 6%
[0078] Table 1
[0079] Model MAE Improvement Ratio Model ViT 12.4% ViT Mmvp 7.1% Mmvp MAU 10.0% MAU SimVP 15.4% SimVP ConvLSTM 13.4% ConvLSTM
[0080] Table 2
[0081] Figure 4 is a structural diagram of a video data generation device according to an embodiment of the present application. As Figure 4 shown, the device includes:
[0082] An acquisition module 41, configured to acquire initial video data.
[0083] A first determination module 42, configured to analyze the original images in the initial video data by using a first deep learning model to obtain a plurality of image blocks output by the deep learning model and attention maps respectively corresponding to the plurality of image blocks, where the attention maps include: attention scores corresponding to the image blocks.
[0084] A second determination module 43, configured to determine a target attention map whose attention score satisfies a first preset condition in the attention maps.
[0085] A conversion module 44, configured to convert the target image block corresponding to the target attention map into a mask image block.
[0086] A filling module 45, configured to perform a repair process on the mask image block by using a second deep learning model to obtain a target image block.
[0087] A third determination module 46, configured to perform a fusion process on the target image block and the original image to obtain a target image, and update the target image to the initial video data to obtain target video data.
[0088] Optionally, the second deep learning model includes: a U-Net model, where the U-Net model includes: a second variational autoencoder and a second variational decoder. The video data generation device further includes: a first training module for training the second deep learning model by the following method: extracting n frames of first images other than the original image from the initial video data, performing feature extraction on the first images to obtain first image features; obtaining text information for describing the first images, performing feature extraction on the text information to obtain text features; using the second variational autoencoder to compress the first images from the image space to the latent space to obtain second image features, where the dimension of the latent space is smaller than the dimension of the image space; determining the attention weights between the text features and the second image features, and performing weighted summation on the first image features according to the attention weights to obtain third image features; using the second variational decoder to restore the third image features from the latent space to the image space to obtain fourth image features; determining a first loss function according to the error between the first image features and the fourth image features, and obtaining the model parameters in the U-Net model when the first loss function satisfies the second preset condition; using the model weight parameters in the U-Net model to replace the model weight parameters in the second deep learning model to obtain the trained second deep learning model.
[0089] Optionally, the padding module 45 is further configured to perform the following steps: using the second variational autoencoder to compress the masked image patch from the image space to the latent space to obtain a first feature; performing element-wise multiplication on the masked image patch and the pixel values of the original image to obtain a second image; adding noise to the first feature within the target region corresponding to the second image to obtain a second feature; performing n iterations of denoising on the second feature: adding the second feature in the i-th iteration to the predicted noise in the (i + 1)-th iteration determined according to the second deep learning model to obtain the second feature in the (i + 1)-th iteration, where n is a positive integer greater than 1 and i is a positive integer not greater than n; using the second variational decoder to restore the second feature obtained in the n-th iteration from the latent space to the image space to obtain the target image patch.
[0090] Optionally, the first deep learning model includes: a first variational autoencoder and a first variational decoder. The video data generation device further includes: a second training module for training the first deep learning model by the following method: obtaining training video data, where the training video data includes: training images; segmenting the training images into multiple training image blocks of a preset size; respectively converting each training image block into a first vector to obtain a plurality of first vectors, where the first vector is a one-dimensional vector; respectively determining the position encoding information corresponding to each training image block to obtain a plurality of position encoding information; inputting the plurality of first vectors and the plurality of position encoding information into the first variational autoencoder to obtain a target vector output by the first variational autoencoder; determining a loss function according to the error between the target vector and the vector corresponding to the training image block, and completing the training of the deep learning model when the loss function satisfies a second preset condition.
[0091] Optionally, the first variational autoencoder includes: a plurality of identical encoder layers, and each encoder layer includes: a multi-head self-attention mechanism and a feed-forward neural network.
[0092] Optionally, the attention map is a row vector; the second determination module 43 is further configured to perform the following steps: determining the numerical value corresponding to each row vector to obtain a plurality of numerical values; and determining the row vector with the smallest numerical value among the plurality of numerical values as the target attention map.
[0093] Optionally, the video data generation device further includes: a fourth determination module for performing the following steps before converting the target image block corresponding to the target attention map into a masked image block: determining and storing the position information of the target image block in the original image. The third determination module 46 is further configured to, based on the position information, replace the target image block in the original image with the target image block to obtain a target image.
[0094] It should be noted that the above Figure 4 each module may be a program module (for example, a set of program instructions for implementing a specific function), or a hardware module. For the latter, it may be presented in the following forms, but not limited to: the presentation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.
[0095] It should be noted that Figure 4 The preferred implementation manners of the embodiments shown may be referred to Figure 1 the relevant descriptions of the embodiments shown, and will not be elaborated here.
[0096] Figure 5 shows a hardware structure block diagram of a computer terminal for implementing a video data generation method. As Figure 5As shown, the computer terminal 50 may include one or more processors 502 (shown as 502a, 502b, ……, 502n in the figure) (the processor 502 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 504 for storing data, and a transmission module 506 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 5 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the computer terminal 50 may further include more or fewer components than Figure 5 shown in the figure, or have a different configuration from Figure 5 that shown in the figure.
[0097] It should be noted that the above one or more processors 502 and / or other data processing circuits are generally referred to as "data processing circuits" herein. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 50. As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0098] The memory 504 may be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for generating video data in the embodiments of the present application. The processor 502 executes various functional applications and data processing by running the software programs and modules stored in the memory 504, that is, implements the above method for generating video data. The memory 504 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 504 may further include a memory remotely provided with respect to the processor 502, and these remote memories may be connected to the computer terminal 50 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0099] The transmission module 506 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the computer terminal 50. In one example, the transmission module 506 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 506 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0100] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 50.
[0101] It should be noted here that in some alternative embodiments, the above Figure 5 illustrated computer terminal may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 5 is only an example of a specific specific instance and is intended to illustrate the types of components that may exist in the above computer terminal.
[0102] It should be noted that Figure 5 the illustrated computer terminal is used to execute Figure 1 the method for generating the video data shown, so the relevant explanations in the above method for executing the command also apply to this electronic device, which will not be elaborated here.
[0103] The embodiment of the present application also provides a non-volatile storage medium. The non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the above method for generating video data.
[0104] The program that the non-volatile storage medium executes the following functions: obtaining initial video data; analyzing the original images in the initial video data using a first deep learning model to obtain multiple image blocks output by the deep learning model and the attention maps respectively corresponding to the multiple image blocks, wherein the attention map includes: the attention score corresponding to the image block; determining a target attention map in the attention map whose attention score meets a first preset condition; converting the target image block corresponding to the target attention map into a masked image block; using a second deep learning model to perform a repair process on the masked image block to obtain a target image block; fusing the target image block with the original image to obtain a target image, and updating the target image to the initial video data to obtain target video data.
[0105] An embodiment of the present application further provides an electronic device, including: a memory and a processor, where the processor is configured to run a program stored in the memory, and when the program runs, it executes the above method for generating video data.
[0106] The processor is configured to run a program that executes the following functions: obtaining initial video data; analyzing the original images in the initial video data using a first deep learning model to obtain multiple image blocks output by the deep learning model and the attention maps respectively corresponding to the multiple image blocks, where the attention map includes: the attention score corresponding to the image block; determining a target attention map in the attention map whose attention score meets a first preset condition; converting the target image block corresponding to the target attention map into a mask image block; using a second deep learning model to perform a repair process on the mask image block to obtain a target image block; fusing the target image block with the original image to obtain a target image, and updating the target image to the initial video data to obtain target video data.
[0107] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0108] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0109] In the above embodiments of the present application, the information collected is information and data authorized by the user or fully authorized by all parties, and the processing of the relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with relevant laws, regulations, and standards, takes necessary protection measures, does not violate public order and good customs, and provides a corresponding operation entry for the user to select authorization or rejection.
[0110] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0111] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0112] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0113] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs that can store program codes.
[0114] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for generating video data, characterized in that: include: Get initial video data; Analyze the original image in the initial video data using a first deep learning model to obtain a plurality of image blocks output by the deep learning model and attention maps corresponding to the plurality of image blocks, wherein the attention maps include: attention scores corresponding to the image blocks; Determining, in the attention map, a target attention map in which the attention score satisfies a first preset condition; Convert the target image block corresponding to the target attention map into a mask image block; Performing a repair process on the mask image block using a second deep learning model to obtain a target image block, wherein the second deep learning model includes: a U-Net model, and the U-Net model includes: a second variational autoencoder and a second variational autodecoder; The second deep learning model is obtained by training in the following method: extracting n frames of the first image other than the original image from the initial video data, performing feature extraction on the first image, and obtaining first image features; obtaining text information used to describe the first image, performing feature extraction on the text information, and obtaining text features; using a second variational autoencoder to compress the first image from an image space to a latent space, and obtaining second image features, wherein the dimension of the latent space is smaller than the dimension of the image space; determining an attention weight between the text features and the second image features, and performing weighted summation on the first image features according to the attention weights, and obtaining third image features; using the second variational autodecoder to restore the third image features from the latent space to the image space, and obtaining fourth image features; determining a first loss function according to an error between the first image features and the fourth image features, and obtaining model parameters in the U-Net model when the first loss function satisfies a second preset condition; replacing model weight parameters in the second deep learning model with model weight parameters in the U-Net model, and obtaining a second deep learning model that has completed training; The target image block is fused with the original image to obtain a target image, and the target image is updated to the initial video data to obtain target video data.
2. The method according to claim 1, characterized in that Performing a repair process on the mask image block using a second deep learning model to obtain a target image block, including: compressing the mask image block from the image space to the latent space using the second variational autoencoder to obtain a first feature; Performing element-by-element multiplication of the mask image block and the pixel value of the original image to obtain a second image; In a target area corresponding to the second image, adding noise to the first feature to obtain a second feature; Performing n-iteration denoising processing on the second feature: adding the second feature in the i-th iteration to the predicted noise in the i+1-th iteration determined according to the second deep learning model to obtain the second feature in the i+1-th iteration, where n is a positive integer greater than 1, and i is a positive integer not greater than n; The second feature obtained in the nth iteration is restored from the latent space to the image space using the second variational self-decoder to obtain the target image block.
3. The method according to claim 1, characterized in that The first deep learning model includes: a first variational autoencoder and a first variational autodecoder, and the first deep learning model is trained by the following method: Acquire training video data, wherein the training video data includes: training images; Dividing the training image into a plurality of training image blocks of preset sizes; Convert each of the training image blocks into a first vector respectively to obtain a plurality of the first vectors, wherein the first vector is a one-dimensional vector; Respectively determine the position coding information corresponding to each of the training image blocks to obtain a plurality of the position coding information; Inputting a plurality of the first vectors and a plurality of the position encoding information into the first variational autoencoder to obtain a target vector output by the first variational autoencoder; A loss function is determined based on an error between the target vector and the vector corresponding to the training image block, and the training of the deep learning model is completed when the loss function satisfies a second preset condition.
4. The method according to claim 3, characterized in that The first variational autoencoder includes: a plurality of identical encoder layers, each of the encoder layers includes: a multi-head self-attention mechanism and a feedforward neural network.
5. The method according to claim 1 or 3, characterized in that: The attention map is a row vector; Determining, in the attention map, a target attention map in which the attention score satisfies a first preset condition, comprising: Determine a value corresponding to each of the row vectors to obtain multiple values; The row vector with the smallest value among the multiple values is determined as the target attention map.
6. The method according to claim 1, characterized in that Before converting the target image block corresponding to the target attention map into a mask image block, the method further includes: Determine and store the position information of the target image block in the original image; The target image block is fused with the original image to obtain a target image, including: Based on the position information, the target image block in the original image is replaced with the target image block to obtain the target image.
7. A device for generating video data, characterized in that: include: An acquisition module, used for acquiring initial video data; A first determination module is configured to analyze the original image in the initial video data using a first deep learning model to obtain a plurality of image blocks output by the deep learning model and attention maps corresponding to the plurality of image blocks, wherein the attention maps include: attention scores corresponding to the image blocks; A second determination module is used to determine, in the attention map, a target attention map in which the attention score satisfies a first preset condition; A conversion module, used to convert the target image block corresponding to the target attention map into a mask image block; A filling module, used to perform a repair process on the mask image block using a second deep learning model to obtain a target image block, wherein the second deep learning model includes: a U-Net model, and the U-Net model includes: a second variational autoencoder and a second variational autodecoder; The second deep learning model is obtained by training in the following method: extracting n frames of the first image other than the original image from the initial video data, performing feature extraction on the first image, and obtaining first image features; obtaining text information used to describe the first image, performing feature extraction on the text information, and obtaining text features; using a second variational autoencoder to compress the first image from an image space to a latent space, and obtaining second image features, wherein the dimension of the latent space is smaller than the dimension of the image space; determining an attention weight between the text features and the second image features, and performing weighted summation on the first image features according to the attention weights, and obtaining third image features; using the second variational autodecoder to restore the third image features from the latent space to the image space, and obtaining fourth image features; determining a first loss function according to an error between the first image features and the fourth image features, and obtaining model parameters in the U-Net model when the first loss function satisfies a second preset condition; replacing model weight parameters in the second deep learning model with model weight parameters in the U-Net model, and obtaining a second deep learning model that has completed training; The third determination module is used to fuse the target image block with the original image to obtain a target image, and update the target image into the initial video data to obtain target video data.
8. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute the method for generating video data according to any one of claims 1 to 6.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the method for generating video data according to any one of claims 1 to 6 when running.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating video data according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Target detection network training method, target detection method and related device
CN117274768A