Video generation model training method, video generation method and device
By extracting the pose and background image sequences from the sample fitting video and training the video generation model, the problem of time and space inconsistency and background in the 2D virtual fitting solution was solved, and a stable fitting video with background and clothes was generated, which improved the video quality.
Patent Information
- Application Number
- CN202210236968.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-03-11
AI Technical Summary
The fitting videos generated by the existing 2D virtual fitting scheme have problems such as time and space inconsistency and inconsistency between the background area and the clothing area, resulting in poor quality of the resulting fitting videos.
By extracting the pose image sequence and background image sequence from the sample fitting video and using the video generation model for training, the generated fitting video combines the deformed clothes image, the pose image and the background image sequence to learn the continuous relationship between pose and background, and generates a fitting video that is consistent in time and space.
It improves the picture stability of the fitting video and the coordination between the background area and the clothing area, and improves the overall quality of the fitting video.
Smart Images

Figure CN114638375B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to a method for training a video generation model, a video generation method and a device. Background Art
[0002] With the development of online e-commerce platforms, virtual try-on technology that simulates clothes selected by users on people can enhance the shopping experience of users. Since 3D virtual try-on solutions require a large amount of computing resources, 2D virtual try-on solutions are the mainstream research direction in this field.
[0003] Currently, the try-on module in each 2D virtual try-on solution generates a virtual try-on video in a frame-by-frame manner. That is, after obtaining a person video uploaded by a user, each time the deformed clothing image and a video frame of the person video are input into the try-on module, and then the try-on module outputs a synthesized try-on image for this video frame. Then, the synthesized try-on images output by the try-on module for each video frame are spliced into a try-on video. The spliced try-on video has the problem of spatio-temporal inconsistency, that is, the space in the try-on video is not continuous in time, the generated try-on video is jittery, and the background area and the clothing area in the generated try-on video are not coordinated, resulting in poor quality of the generated try-on video. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a method for training a video generation model, a video generation method and a device to improve the quality of the generated try-on video. The specific technical solutions are as follows:
[0005] In a first aspect, the embodiments of the present application provide a method for training a video generation model, the method including:
[0006] Obtain a sample try-on video and a sample clothing image, wherein the sample try-on video is a video of a person wearing the sample clothing;
[0007] Extract a pose image sequence and a background image sequence from the sample try-on video, wherein the pose image sequence contains the pose information of the person in each video frame of the sample try-on video, and the background image sequence includes the images in each video frame of the sample try-on video except for the area to be tried on;
[0008] Perform deformation processing on the sample clothing image according to the pose information of the person in each video frame of the sample try-on video to obtain a deformed clothing image sequence;
[0009] Input the deformed clothing image sequence, the pose image sequence, and the background image sequence into a video generation model to obtain a synthesized virtual try-on video output by the video generation model;
[0010] Calculate the loss function value based on the synthesized virtual try-on video and the sample virtual try-on video, and adjust the parameters of the video generation model based on the loss function value until the video generation model converges, and then determine that the training of the video generation model is completed.
[0011] In a possible implementation, the video generation model includes: three encoders, multiple MPDT modules based on multi-scale image patches, and one decoder. Each MPDT module includes three input terminals. The output terminals of different encoders are connected to different input terminals of the starting MPDT module in the multiple MPDT modules, and the output terminal of the ending MPDT module in the multiple MPDT modules is connected to the input terminal of the decoder; the step of inputting the deformed clothing image sequence, the pose image sequence, and the background image sequence into the video generation model to obtain the synthesized virtual try-on video output by the video generation model includes:
[0012] Input the deformed clothing image sequence, the pose image sequence, and the background image sequence into an encoder respectively to obtain deformed clothing sequence features, pose sequence features, and background sequence features;
[0013] Iteratively process the deformed clothing sequence features, the pose sequence features, and the background sequence features through the multiple MPDT modules to obtain a first fusion feature value;
[0014] Input the first fusion feature value into the decoder to obtain an image sequence output by the decoder;
[0015] Generate the synthesized virtual try-on video based on the image sequence output by the decoder.
[0016] In a possible implementation, the image sequence output by the decoder includes an initial virtual try-on image sequence and a clothing region mask sequence; the step of generating the synthesized virtual try-on video based on the image sequence output by the decoder includes:
[0017] Fuse the initial virtual try-on image sequence, the clothing region mask sequence, and the deformed clothing image sequence to obtain a fused image sequence;
[0018] Fuse the fused image sequence, the background image mask sequence, and the background image sequence to obtain the synthesized virtual try-on video.
[0019] In a possible implementation, the MPDT module includes a first multi-head attention module, a second multi-head attention module, and a third multi-head attention module; the iterative processing of the deformed clothing sequence features, the pose sequence features, and the background sequence features through the multiple MPDT modules to obtain a first fusion feature value includes:
[0020] Input the deformed clothing sequence features and the pose attention feature set into the first multi-head attention module of the starting MPDT module to obtain deformed clothing attention features, where the pose attention feature set includes the features obtained by each head of the second multi-head attention module in the starting MPDT module processing the pose sequence features;
[0021] Input the background sequence features and the pose attention feature set into the third multi-head attention module of the starting MPDT module to obtain background attention features;
[0022] After splicing the deformed clothing attention features and the background attention features, perform a convolution operation to obtain a second fusion feature value;
[0023] For each MPDT module, perform a residual connection between the output and the input of the first multi-head attention module of this MPDT module as the input of the first multi-head attention module of the next MPDT module; perform a residual connection between the second fusion feature value output by this MPDT module and the input of the second multi-head attention module of this MPDT module as the input of the second multi-head attention module of the next MPDT module; perform a residual connection between the output and the input of the third multi-head attention module of this MPDT module as the input of the third multi-head attention module of the next MPDT module, until the first fusion feature value output by the ending MPDT module is obtained.
[0024] In a possible implementation, the inputting the deformed clothing sequence features and the pose attention feature set into the first multi-head attention module of the starting MPDT module to obtain deformed clothing attention features includes:
[0025] Input the deformed clothing sequence features into the first multi-head attention module to obtain first key-value pairs obtained by each head of the first multi-head attention module processing the deformed clothing sequence features;
[0026] Input the pose attention feature set into the first multi-head attention module, so that the first multi-head attention module obtains the deformed clothing attention features based on the pose attention feature set and multiple first key-value pairs;
[0027] Inputting the background sequence feature and the pose attention feature set into the third multi-head attention module of the starting MPDT module to obtain the background attention feature includes:
[0028] Inputting the background sequence feature into the third multi-head attention module to obtain second key-value pairs processed by each head of the third multi-head attention module for the deformed clothing sequence feature;
[0029] Inputting the pose attention feature set into the third multi-head attention module, so that the third multi-head attention module obtains the background attention feature based on the pose attention feature set and multiple second key-value pairs.
[0030] In a possible implementation manner, the fusion of the initial fitting image sequence, the clothing region mask sequence, and the deformed clothing image sequence to obtain a fused image sequence includes:
[0031] Performing fusion calculation according to the following formula:
[0032] I masked1 T = M C1 T ⊙ C1 T + (1 - M C1 T ) ⊙ I R1 T
[0033] where C1 T represents the deformed clothing image sequence, M C1 T represents the clothing region mask sequence, I R1 T represents the initial fitting image sequence, I masked1 T represents the fused image sequence.
[0034] The fusion of the fused image sequence, the background image mask sequence, and the background image sequence to obtain the synthesized fitting video includes:
[0035] Performing fusion calculation according to the following formula:
[0036] I T 1 = (1 - M a1 T ) ⊙ I masked1 T + M a1 T ⊙ A1 T
[0037] where Ma1 T Denote the background image mask sequence, A1 T Denote the background image sequence, I T 1 denotes the synthesized virtual fitting video.
[0038] In a second aspect, an embodiment of the present application provides a video generation method, the method comprising:
[0039] Obtain a video to be processed and a target clothing image, where the video to be processed is a video containing a person;
[0040] Extract a pose image sequence and a background image sequence from the video to be processed, where the pose image sequence contains the pose information of the person in each video frame of the video to be processed, and the background image sequence includes the images in each video frame of the video to be processed except for the area to be tried on;
[0041] Perform a deformation process on the target clothing image according to the pose information of the person in each video frame of the video to be processed, to obtain a deformed clothing image sequence;
[0042] Input the deformed clothing image sequence, the pose image sequence, the background image sequence and the input video into a video generation model, to obtain a target synthesized virtual fitting video, where the video generation model is a video generation model trained by the method described in the first aspect above.
[0043] In a third aspect, an embodiment of the present application provides a video generation model training device, the device comprising:
[0044] A first acquisition module, configured to acquire a sample virtual fitting video and a sample clothing image, where the sample virtual fitting video is a video of a person wearing a sample clothing;
[0045] A first extraction module, configured to extract a pose image sequence and a background image sequence from the sample virtual fitting video, where the pose image sequence contains the pose information of the person in each video frame of the sample virtual fitting video, and the background image sequence includes the images in each video frame of the sample virtual fitting video except for the area to be tried on;
[0046] A first deformation module, configured to perform a deformation process on the sample clothing image according to the pose information of the person in each video frame of the sample virtual fitting video, to obtain a deformed clothing image sequence;
[0047] A first generation module, configured to input the deformed clothing image sequence, the pose image sequence and the background image sequence into a video generation model, to obtain the synthesized virtual fitting video output by the video generation model;
[0048] A calculation module, configured to calculate a loss function value according to the synthesized virtual fitting video and the sample virtual fitting video, and adjust the video generation model parameters based on the loss function value until the video generation model converges, and then determine that the training of the video generation model is completed.
[0049] Fourthly, an embodiment of the present application provides a video generation device, which includes:
[0050] A second acquisition module, configured to acquire a video to be processed and a target clothing image, where the video to be processed is a video including a person;
[0051] A second extraction module, configured to extract a pose image sequence and a background image sequence from the video to be processed, where the pose image sequence includes pose information of a person in each video frame of the video to be processed, and the background image sequence includes images other than the area to be tried on in each video frame of the video to be processed;
[0052] A second deformation module, configured to deform the target clothing image according to the pose information of a person in each video frame of the video to be processed to obtain a deformed clothing image sequence;
[0053] A second generation module, configured to input the deformed clothing image sequence, the pose image sequence, the background image sequence into a video generation model to obtain a target synthesized virtual fitting video, where the video generation model is a video generation model trained by the device described in the third aspect above
[0054] Fifthly, an embodiment of the present application provides an electronic device, which includes a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory communicate with each other through the communication bus;
[0055] The memory is used to store a computer program;
[0056] The processor is configured to implement the method steps described in the first aspect or the second aspect when executing the program stored in the memory.
[0057] Sixthly, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements any of the above-mentioned video generation model training methods or video generation methods.
[0058] Seventhly, an embodiment of the present application provides a computer program product including instructions, and when it runs on a computer, it causes the computer to execute any of the above-mentioned video generation model training methods or video generation methods.
[0059] Adopting the solution provided by the embodiment of the present application, compared with the prior art in which the Try-on module generates a try-on image for each video frame of the person video and then stitches all the generated try-on images into a try-on video, the embodiment of the present application can extract the pose image sequence and the background image sequence from the sample try-on video. Furthermore, the video generation model in the embodiment of the present application is trained based on the deformed clothing image sequence, the pose image sequence, and the background image sequence. Since the continuous pose image sequence and background image sequence extracted from the sample try-on video are used when training the video generation model, rather than only processing one video frame of the sample try-on video, during the process of the video generation model generating the try-on video, it can learn the relationship between consecutive pose images in the pose image sequence and the relationship between consecutive background images in the background image sequence. Thus, the try-on video generated based on this video generation model has spatio-temporal consistency, and the picture of the try-on video is stable without jitter. Moreover, since the input during training the video generation model includes the background image sequence, the video generation model combines the background image sequence when generating the try-on video, making the background area and the clothing area in the generated try-on video more coordinated, which can improve the picture quality of the try-on video. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.
[0061] Figure 1 It is a flowchart of a method for training a video generation model provided by an embodiment of the present application;
[0062] Figure 2a It is an exemplary schematic diagram of a first key point information provided by an embodiment of the present application;
[0063] Figure 2b It is an exemplary schematic diagram of a second key point information provided by an embodiment of the present application;
[0064] Figure 2c It is a schematic diagram of a background image sequence provided by an embodiment of the present application;
[0065] Figure 3 It is a flowchart of another method for training a video generation model provided by an embodiment of the present application;
[0066] Figure 4 It is a flowchart of another method for training a video generation model provided by an embodiment of the present application;
[0067] Figure 5 It is a flowchart of a video generation method provided by an embodiment of the present application;
[0068] Figure 6 Schematic diagram of the video synthesis model structure provided by an embodiment of the present application;
[0069] Figure 7 Schematic diagram of the structure of a video generation model training device provided by an embodiment of the present application;
[0070] Figure 8 Schematic diagram of the structure of a video generation device provided by an embodiment of the present application;
[0071] Figure 9 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0072] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.
[0073] As Figure 1 shown, an embodiment of the present application provides a video generation model training method, which can be applied to an electronic device. The electronic device can be a device such as a smart phone, a tablet computer, a desktop computer, a server, etc. The method includes:
[0074] S101. Obtain a sample try-on video and a sample clothing image.
[0075] Among them, the sample try-on video is a video of a person wearing the sample clothing.
[0076] S102. Extract a pose image sequence and a background image sequence from the sample try-on video.
[0077] Among them, the pose image sequence contains the pose information of the person in each video frame of the sample try-on video, and the background image sequence includes the images in each video frame of the sample try-on video except for the area to be tried on.
[0078] S103. Perform deformation processing on the sample clothing image according to the pose information of the person in each video frame of the sample try-on video to obtain a deformed clothing image sequence.
[0079] S104. Input the deformed clothing image sequence, the pose image sequence, and the background image sequence into the video generation model to obtain a synthesized try-on video output by the video generation model.
[0080] S105. Calculate the loss function value according to the synthesized try-on video and the sample try-on video, and adjust the video generation model parameters based on the loss function value until the video generation model converges, and then determine that the training of the video generation model is completed.
[0081] According to the embodiment of the present application, compared with the prior art method in which the Try-on module generates a fitting image for each video frame of the character video, and then splices all the generated fitting images into a fitting video, the embodiment of the present application can extract a posture image sequence and a background image sequence from the sample fitting video, and then the video generation model in the embodiment of the present application is obtained by training based on the deformed clothing image sequence, the posture image sequence and the background image sequence. Because the continuous posture image sequence and the background image sequence extracted from the sample fitting video are used when training the video generation model, not only one video frame of the sample fitting video is processed, so the relationship between the continuous posture images in the posture image sequence and the relationship between the continuous background images in the background image sequence can be learned during the generation of the fitting video by the video generation model, so that the fitting video generated based on the video generation model has spatiotemporal consistency, and the picture of the fitting video is stable and does not shake. And because the input when training the video generation model includes the background image sequence, the video generation model combines the background image sequence when generating the fitting video, so that the background area and the clothing area in the generated fitting video are more coordinated, which can improve the picture quality of the fitting video.
[0082] With respect to the above S101, the sample clothing image is an image of the sample clothing in a flat state.
[0083] With respect to the above S102, in one implementation, a human posture recognition algorithm can be used to perform human posture recognition on a human figure in each video frame of a sample fitting video to obtain a sequence of human posture images. The posture information in the posture image includes first key point information and second key point information. The first key point information is information representing key points of each joint of a human body, and the second key point information is information representing the shape of each part of a human body. The first key point information can be obtained by an Openpose algorithm, such as Figure 2a As shown, Figure 2a is an exemplary schematic diagram of the first key point information, and the second key point information can be obtained by the Densepose algorithm, such as Figure 2b As shown, Figure 2b is an exemplary schematic diagram of the second key point information; wherein, the Openpose algorithm is a human posture recognition algorithm used to identify the key points of each joint of the human body in the image, and the Densepose algorithm is a human posture recognition algorithm that converts a 2D human body image into a 3D human body image.
[0084] The background image sequence is obtained by processing the to-be-fitted area of the portrait in each video frame of the sample fitting video to the same pixel value, such as Figure 2c As shown, Figure 2c is a schematic diagram of the background image sequence. Figure 2cThe upper garment area of the portrait included therein is processed to the same pixel value, that is, the background image includes the background except the portrait area, the part of the portrait area except the upper garment area, and the occluder of the upper garment area.
[0085] For the above S103, in one implementation, according to the pose information of the person in each video frame of the sample try-on video, through CP-VTON, Parser Free Appearence Flow Network (PF-AFN), Adaptive Content Generation and Preserving Network (ACGPN), VITON-HD and other try-on methods, the deformation (warp) module is used to deform the sample clothing image to obtain a deformed clothing sequence. Among them, CP-VTON is a feature-preserving virtual try-on network (CP-VTON) proposed at the International Conference on Computer Vision in Europe, and VITON-HD is an image-based virtual try-on network.
[0086] For example, the deformation module can deform the sample clothing image by the Thin Plate Spline (TPS) method or the appearance-flow-based warp method.
[0087] Among them, the deformed clothing in each frame of the deformed clothing sequence fits the pose of the person in the sample try-on video.
[0088] For the above S105, in the embodiments of the present application, the loss function value is calculated according to the synthesized try-on video and the sample try-on video, and based on the loss function, it is determined whether the video generation model converges. If the video generation model converges, it is determined that the training of the video generation model is completed; if the video generation model does not converge, the parameters of the video generation model are adjusted according to the loss function value, and S101 is returned to obtain the next sample try-on video and sample clothing image until the video generation model converges, and it is determined that the training of the video generation model is completed.
[0089] Among them, the loss function in the embodiments of the present application can be:
[0090]
[0091] Among them, λ1, λ2, λ3, λ4 are hyperparameters, represents the loss function of the synthesized try-on video and the sample try-on video in all regions, represents the loss function of the synthesized try-on video and the sample try-on video in the try-on region, L TPGAN represents the adversarial loss, Lperc Represents a perceptual loss function.
[0092] In another embodiment of the present application, the above video generation model includes: three encoders, one decoder, and multiple multi-scale patch-based dual-stream transformer (MPDT) modules. Each MPDT module includes three input terminals. The output terminals of different encoders are connected to different input terminals of the starting MPDT module among the multiple MPDT modules. The output terminal of the ending MPDT module among the multiple MPDT modules is connected to the input terminal of the decoder. On this basis, as Figure 3 shown, the above S104 can be specifically implemented as:
[0093] S1041. Input the deformed clothing image sequence, the pose image sequence, and the background image sequence into one encoder respectively to obtain the deformed clothing sequence feature, the pose sequence feature, and the background sequence feature.
[0094] Among them, the deformed clothing image sequence can be input into the first encoder to obtain the deformed clothing sequence feature output by the first encoder.
[0095] Input the pose image sequence into the second encoder to obtain the pose sequence feature output by the second encoder.
[0096] Input the background image sequence into the third encoder to obtain the background sequence feature output by the third encoder.
[0097] S1042. Iteratively process the deformed clothing sequence feature, the pose sequence feature, and the background sequence feature through multiple MPDT modules to obtain the first fusion feature value.
[0098] In the embodiment of the present application, in the video generation model, multiple MPDT modules are connected in sequence. The first MPDT module among the multiple MPDT modules is the starting MPDT module, and the last MPDT module is the ending MPDT module. Each MPDT module includes three input terminals. The output terminals of different encoders are connected to different input terminals of the starting MPDT module among the multiple MPDT modules. The three output terminals of the starting MPDT module are respectively connected to the three input terminals of the second MPDT module. The three output terminals of the second MPDT module are respectively connected to the three input terminals of the third MPDT module, and so on.
[0099] After inputting the deformed clothing sequence feature, the pose sequence feature, and the background sequence feature into the starting MPDT module, they will be iteratively processed through multiple MPDT modules in sequence to obtain the first fusion feature value output by the ending MPDT module among the multiple MPDT modules.
[0100] Among them, except for the end MPDT module, the three outputs of each MPDT module will be respectively subjected to residual connections with the three inputs of this MPDT module, serving as the three inputs of the next MPDT module, thereby realizing iterative processing of multiple MPDT modules. The specific residual connection method will be introduced below.
[0101] S1043. Input the first fusion eigenvalue into the decoder to obtain the image sequence output by the decoder.
[0102] In the embodiment of the present application, the output end of the end MPDT module is connected to the input end of the decoder, and the decoder can decode the first fusion eigenvalue output by the end MPDT module to obtain the image sequence.
[0103] S1044. Generate a synthetic fitting video based on the image sequence output by the decoder.
[0104] By adopting this method, the deformed clothing image sequence, the pose image sequence, and the background image sequence can be respectively processed into the deformed clothing sequence feature, the pose sequence feature, and the background sequence feature through three encoders. Furthermore, these features can be iteratively processed by multiple MPDT modules, which enables multiple MPDT modules to learn the relationships between consecutive pose images in the pose image sequence and the relationships between consecutive background images in the background image sequence. As a result, the first fusion eigenvalue output by multiple MPDT modules can fully reflect the deformed clothing sequence feature, the pose sequence feature, and the background sequence feature. Therefore, based on the image sequence decoded from this first fusion eigenvalue, a synthetic fitting video with a coordinated background and clothing area can be generated, and the fitting video has spatio-temporal consistency and a stable and non-shaking picture.
[0105] In another embodiment of the present application, the MPDT module includes a first multi-head attention module, a second multi-head attention module, and a third multi-head attention module. Among them, the three input ends of the MPDT module are respectively the input ends of the first multi-head attention module, the second multi-head attention module, and the third multi-head attention module. The output end of the first encoder is connected to the input end of the first multi-head attention module of the starting MPDT module, the output end of the second encoder is connected to the input end of the second multi-head attention module of the starting MPDT module, and the output end of the third encoder is connected to the input end of the third multi-head attention module of the starting MPDT module. On this basis, the above S1042 can be implemented as:
[0106] Step 1. Input the deformed clothing sequence feature and the pose attention feature set into the first multi-head attention module of the starting MPDT module to obtain the deformed clothing attention feature.
[0107] Among them, the pose attention feature set includes the features obtained by each head of the second multi-head attention module in the starting MPDT module processing the pose sequence features.
[0108] In one implementation, each head in the second multi-head attention module intercepts the pose sequence features of image patches of different sizes, and then performs convolution processing on the extracted pose sequence features through a 1×1 convolutional kernel to obtain pose attention features. For example, the second multi-head attention module can have 4 heads. Among them, the first head intercepts the pose sequence features of image patches with a size of 64×48, the second head intercepts the pose sequence features of image patches with a size of 32×24, the third head intercepts the pose sequence features of image patches with a size of 16×12, and the fourth head intercepts the pose sequence features of image patches with a size of 8×6. Each of the 4 heads outputs a pose attention feature, and the pose attention feature set contains 4 pose attention features. The convolution formula in each head is:
[0109] q = conv q (p)
[0110] where p is the pose sequence feature, q is the pose attention feature, and conv q () represents the convolution function.
[0111] The above step one can be specifically implemented as:
[0112] Input the deformed clothing sequence features into the first multi-head attention module to obtain the first key-value pairs obtained by each head of the first multi-head attention module processing the deformed clothing sequence features; input the pose attention feature set into the first multi-head attention module so that the first multi-head attention module obtains the deformed clothing attention features based on the pose attention feature set and multiple first key-value pairs.
[0113] In one implementation, each head in the first multi-head attention module intercepts the deformed clothing sequence features of image patches of different sizes, and then performs convolution processing on the intercepted clothing sequence features through two 1×1 convolutional kernels to obtain the first key-value pairs. For example, the first multi-head attention module can have 4 heads. Among them, the first head intercepts the deformed clothing sequence features of image patches with a size of 64×48, the second head intercepts the deformed clothing sequence features of image patches with a size of 32×24, the third head intercepts the deformed clothing sequence features of image patches with a size of 16×12, and the fourth head intercepts the deformed clothing sequence features of image patches with a size of 8×6. Each head obtains a first key-value pair, and the convolution formula is:
[0114] (K C , V C ) = (conv K (C), conv V (C))
[0115] Among them, (K C , V C ) represents the first key-value pair, conv K () and conv V () represent the convolution function, and C represents the deformed clothing attention feature.
[0116] Among them, the first multi-head attention module and the second multi-head attention module contain the same number and size of heads. For example, both the first multi-head attention module and the second multi-head attention module contain 4 heads. The cropping size of the first head in the first multi-head attention module is the same as that of the first head in the second multi-head attention module. The cropping size of the second head in the first multi-head attention module is the same as that of the second head in the second multi-head attention module. The cropping size of the third head in the first multi-head attention module is the same as that of the third head in the second multi-head attention module. The cropping size of the fourth head in the first multi-head attention module is the same as that of the fourth head in the second multi-head attention module.
[0117] Within the heads with the same cropping size, perform attention operations on the above-mentioned feature q and the first key-value pair (K C , V C ) to obtain the deformed clothing attention result output by this head.
[0118] Specifically, the deformed clothing attention result in each head is calculated through the following formula:
[0119]
[0120]
[0121] Among them, i represents the i-th head, represents the pose attention feature of the i-th head, softmax j () represents the normalization exponential function, r1*r2 represents the cropping size of the i-th head, cn represents the number of channels, and represent the key-value pair of the j-th image patch, Ω C represents the deformed clothing area, represents the attention weight, represents the deformed clothing attention result within the i-th head.
[0122] Input the in each head into a 3*3 convolution kernel for fusion to obtain the deformed clothing attention feature ATT C .
[0123] Step 2: Input the background sequence features and the pose attention feature set into the third multi-head attention module of the starting MPDT module to obtain background attention features.
[0124] Specifically, Step 2 can be implemented as follows: Input the background sequence features into the third multi-head attention module to obtain the second key-value pairs obtained by each head of the third multi-head attention module processing the deformed clothing sequence features; input the pose attention feature set into the third multi-head attention module, so that the third multi-head attention module obtains background attention features based on the pose attention feature set and multiple second key-value pairs.
[0125] In one implementation, each head in the third multi-head attention module intercepts the deformed clothing sequence features of image patches of different sizes, and then performs convolution processing on the intercepted background sequence features through two 1*1 convolutional kernels to obtain the second key-value pairs. For example, the first multi-head attention module can have 4 heads. Among them, the first head intercepts the deformed clothing sequence features of an image patch with a size of 64*48, the second head intercepts the deformed clothing sequence features of an image patch with a size of 32*24, the third head intercepts the deformed clothing sequence features of an image patch with a size of 16*12, and the fourth head intercepts the deformed clothing sequence features of an image patch with a size of 8*6. Each head obtains a second key-value pair, and the convolution formula is:
[0126] (K A ,V A )=(conv K (A),conv V (A))
[0127] where, (K A ,V A ) represents the second key-value pair, A represents the background sequence features, conv V () and conv K () represent the convolution function.
[0128] Among them, the way the third multi-head attention module generates background attention features is the same as the way the first multi-head attention module generates deformed clothing attention features, and reference can be made to the way the first multi-head attention module generates deformed clothing attention features above.
[0129] Step 3: After splicing the deformed clothing attention features and the background attention features, perform convolution operation to obtain the second fusion feature value.
[0130] In one implementation, the spliced deformed clothing attention features and background attention features can be convolved and fused through a 1*1 convolutional kernel. The specific fusion formula is:
[0131]
[0132] Among them, o represents the second fusion feature, and W1 and b1 are the 1*1 convolution parameters.
[0133] Step 4: For each MPDT module, perform a residual connection between the output and the input of the first multi-head attention module of this MPDT module as the input of the first multi-head attention module of the next MPDT module; perform a residual connection between the second fusion feature value output by this MPDT module and the input of the second multi-head attention module of this MPDT module as the input of the second multi-head attention module of the next MPDT module; perform a residual connection between the output and the input of the third multi-head attention module of this MPDT module as the input of the third multi-head attention module of the next MPDT module, until the first fusion feature value output by the end MPDT module is obtained.
[0134] For example, for the starting MPDT module, after performing a residual connection between the deformed clothing attention feature output by the first multi-head attention module in the starting MPDT module and the deformed clothing sequence feature output by the encoder, input it into the first multi-head attention module of the next MPDT module; after performing a residual connection between the second fusion feature value output by the starting MPDT module and the pose sequence feature output by the encoder, input it into the second multi-head attention module of the next MPDT module; after performing a residual connection between the background attention feature output by the third multi-head attention module in the starting MPDT module and the background sequence feature output by the encoder, input it into the third multi-head attention module of the next MPDT module.
[0135] When each subsequent level of MPDT module iterates, the input of the second multi-head attention of this MPDT module is the result of a residual connection between the second fusion feature value output by the previous-level MPDT module and the input of the second multi-head attention module in the previous-level MPDT module.
[0136] The input of the first multi-head attention module of the current MPDT module is the result of a residual connection between the output of the first multi-head attention module in the previous-level MPDT module and the input of the first multi-head attention module in the previous-level MPDT module.
[0137] The input of the third multi-head attention module in the current MPDT module is the result of a residual connection between the output of the third multi-head attention module in the previous-level MPDT module and the input of the third multi-head attention module in the previous-level MPDT module.
[0138] By adopting the embodiment of the present application, the three multi-head attention modules included in the MPDT module perform convolution operations on the deformed clothing sequence features, posture sequence features and background sequence features from different sizes, respectively, to obtain multiple first key-value pairs, posture attention feature sets and multiple second key-value pairs, so that the second multi-head attention module can learn the relationship between continuous posture images in the posture image sequence, and the third multi-head attention module can learn the relationship between continuous background images in the background image sequence. Then, the first multi-head attention module can use the posture attention feature set to perform attention operations on multiple first key-value pairs to obtain deformed clothing attention features, and the third multi-head attention module can use the posture attention feature set to perform attention operations on multiple second key-value pairs to obtain background attention features, so that the MPDT module can fuse the deformed clothing attention features and the background attention features to generate a second fused feature value. The second fused feature value integrates multiple image sequence features, improves the stability of the picture in the final generated fitting video, and also performs attention operations on the background sequence features, so that the background area in the final generated fitting video is coordinated, thereby improving the picture quality of the fitting video.
[0139] Multiple MPDT modules are used for iterative processing. The features output by each MPDT module can be residually connected with the input features, so that the image sequence features will not be lost during the iterative processing of multiple MPDT modules, further improving the picture quality of the fitting video generated based on the video generation model.
[0140] In another embodiment of the present application, the image sequence output by the decoder includes an initial fitting image sequence and a clothing region mask sequence, such as Figure 4 As shown, the above S1044 can be specifically implemented as follows:
[0141] S10441. Fusing the initial fitting image sequence, the clothing region mask sequence and the deformed clothing image sequence to obtain a fused image sequence.
[0142] The fusion calculation is performed according to the following formula:
[0143] I masked1 T =M C1 T ⊙C1 T +(1-M C1 T )⊙I R1 T
[0144] Among them, C1 T represents the deformed clothing image sequence, M C1 T represents the clothing region mask sequence, I R1T represents the initial try-on image sequence, I masked1 T represents the fused image sequence.
[0145] S10442. Fuse the fused image sequence, the background image mask sequence, and the background image sequence to obtain a synthesized try-on video.
[0146] Among them, the background image mask sequence is a sequence obtained by magnifying the clothing area in the human body segmentation map of each video frame in the sample try-on video after human body parsing.
[0147] Perform fusion calculation according to the following formula:
[0148] I T 1 = (1 - M a1 T ) ⊙ I masked1 T + M a1 T ⊙ A1 T
[0149] Among them, M a1 T represents the background image mask sequence, and A1 T represents the background image sequence, and I T 1 represents the synthesized try-on video.
[0150] By adopting the embodiment of the present application, by fusing the initial try-on image sequence and the clothing area mask sequence decoded by the decoder with the deformed clothing image sequence and the background image sequence, the stability of the video frame sequence of the finally obtained synthesized try-on video can be further improved.
[0151] Corresponding to the above embodiment, the embodiment of the present application also provides a method for generating a virtual try-on video, as Figure 5 shown, the method includes:
[0152] S501. Obtain a video to be processed and a target clothing image.
[0153] Among them, the video to be processed is a video containing a person, and the target clothing image is an image of the clothing to be tried on selected by the user in a flat state.
[0154] S502. Extract a pose image sequence and a background image sequence from the video to be processed.
[0155] Among them, the pose image sequence contains the pose information of the person in each video frame of the video to be processed, and the background image sequence includes the images in each video frame of the video to be processed except for the area to be tried on.
[0156] The method for extracting the pose image sequence and the background image sequence from the video to be processed is the same as the method for extracting the pose image sequence and the background image sequence from the sample try-on video in the above embodiment, and the relevant description in the above embodiment can be referred to.
[0157] S503. Perform a deformation process on the target clothing image according to the pose information of the person in each video frame of the video to be processed, and obtain a deformed clothing image sequence.
[0158] Among them, the method for performing a deformation process on the target clothing image is the same as the method for deforming the sample clothing image in the above embodiment, and the relevant description above can be referred to.
[0159] S504. Input the deformed clothing image sequence, the pose image sequence, the background image sequence, and the input video into the model, and obtain the target synthesized try-on video.
[0160] Among them, the video generation model is the video generation model trained by the above video generation model training method.
[0161] Adopting the embodiment of the present application, compared with the method in the prior art where the Try-on module generates a try-on image for each video frame of the person video and then stitches all the generated try-on images into a try-on video, the embodiment of the present application can extract the deformed clothing image sequence, the pose image sequence, and the background image sequence from the video to be processed and the target clothing image, and input the deformed clothing image sequence, the pose image sequence, and the background image sequence into the video generation model, so that the video generation model can generate a try-on video based on the deformed clothing image sequence, the pose image sequence, and the background image sequence. Because the video generation model uses the continuous pose image sequence and background image sequence extracted from the video to be processed when generating the try-on video, rather than only processing one video frame in the video to be processed, during the process of the video generation model generating the try-on video, it can learn the relationship between consecutive pose images in the pose image sequence and can also learn the relationship between consecutive background images in the background image sequence. Therefore, the try-on video generated based on this video generation model has spatio-temporal consistency, and the picture of the try-on video is stable without jitter. And because the input during the training of the video generation model includes the background image sequence, the video model combines the background image sequence features when generating the try-on video, making the background area and the clothing area in the generated try-on video more coordinated, and can improve the picture quality of the try-on video.
[0162] Figure 6 This is the structural diagram of the video synthesis model provided by the embodiment of the present application. The following will be described in conjunction with Figure 6 for illustration.
[0163] The deformed clothing image sequence The person pose image sequence and the background image sequence A1 T They are respectively input into an encoder to obtain the deformed clothing image feature C, the human pose image feature P, and the background image feature A.
[0164] The features C, P, and A are respectively input into three multi-head attention modules. Each head of the three multi-head attention modules intercepts the features C, P, and A of image patches of different sizes for convolution, and respectively obtains the deformed clothing image key-value pairs (K C , V C ), obtains the query value Q, and the background image key-value pairs (K BG , V BG ).
[0165] Figure 5 Each multi-head attention module has four heads. The first head intercepts image patches with sizes r1 = 64 and r2 = 48, the second head intercepts image patches with sizes r1 = 32 and r2 = 24, the third head intercepts image patches with sizes r1 = 16 and r2 = 12, and the fourth head intercepts image patches with sizes r1 = 8 and r2 = 6.
[0166] Within the heads of the same size, the Q value is used to perform attention operations on the deformed clothing image key-value pairs (K C , V C ) to obtain the attention result of this head. The attention results of each head are fused using a 3*3 convolution kernel to obtain the deformed clothing image attention result ATT C .
[0167] Within the heads of the same size, the Q value is used to perform attention operations on the background image key-value pairs (K BG , V BG ) to obtain the attention result of this head. The attention results of each head are fused using a 3*3 convolution kernel to obtain the background image attention result ATT A .
[0168] ATT C and ATT A are concatenated and then fused through a 1*1 convolution to obtain the eigenvalue o.
[0169] The eigenvalue o and the feature P output by the encoder are subjected to residual connection and used as the input P of the next-level MPDT Block module; ATT C and the eigenvalue C are subjected to residual connection and used as the input C of the next-level MPDT Block module; ATT A and the eigenvalue A are subjected to residual connection and used as the input A of the next-level MPDT Block module for iterative processing to obtain the finally output fused eigenvalue O.
[0170] Among them, when each level of MPDT Block module iterates, the eigenvalue C input to the current MPDT Block module is the ATT of the input feature C of the previous level and the output of the previous level C The result of the residual connection.
[0171] The eigenvalue P input to the current MPDT Block module is the result of the residual connection of the eigenvalue P of the input of the previous level and the eigenvalue o of the output of the previous level.
[0172] The eigenvalue A input to the current MPDT Block module is the ATT of the input eigenvalue A of the previous level and the output of the previous level A The result of the residual connection.
[0173] Input the finally output fused eigenvalue O into the decoder to obtain the synthetic clothing sequence I r1 T And the clothing area mask sequence M C1 T .
[0174] Fuse the synthetic clothing sequence I r1 T First with the clothing area mask sequence M C1 T And the deformed clothing image sequence To obtain the intermediate fusion sequence I masked1 T
[0175] Then fuse the intermediate fusion sequence I masked1 T With the background image mask sequence M a1 T And the background image sequence A1 T To obtain the virtual try-on video sequence I T 1.
[0176] Corresponding to the above method embodiment, the embodiment of the present application also provides a video generation model training device, as Figure 7 Shown, the device includes:
[0177] The first acquisition module 701 is used to acquire a sample try-on video and a sample clothing image, where the sample try-on video is a video of a person wearing a sample clothing;
[0178] The first extraction module 702 is used to extract a pose image sequence and a background image sequence from the sample try-on video, where the pose image sequence contains the pose information of the person in each video frame of the sample try-on video, and the background image sequence includes the image except the area to be tried on in each video frame of the sample try-on video;
[0179] The first deformation module 703 is configured to perform deformation processing on the sample clothing image according to the pose information of the person in each video frame of the sample fitting video, so as to obtain a sequence of deformed clothing images;
[0180] The first generation module 704 is configured to input the sequence of deformed clothing images, the sequence of pose images, and the sequence of background images into the video generation model, so as to obtain the synthesized fitting video output by the video generation model;
[0181] The calculation module 705 is configured to calculate the value of the loss function according to the synthesized fitting video and the sample fitting video, and adjust the parameters of the video generation model based on the value of the loss function until the video generation model converges, and then determine that the training of the video generation model is completed.
[0182] In another embodiment of the present application, the video generation model includes: three encoders, multiple MPDT modules based on multi-scale image patches, and a decoder. Each MPDT module includes three input ends. The output ends of different encoders are connected to different input ends of the starting MPDT module among the multiple MPDT modules. The output end of the ending MPDT module among the multiple MPDT modules is connected to the input end of the decoder; the first generation module 704 is specifically configured to:
[0183] Input the sequence of deformed clothing images, the sequence of pose images, and the sequence of background images into an encoder respectively, so as to obtain the deformed clothing sequence feature, the pose sequence feature, and the background sequence feature;
[0184] Iteratively process the deformed clothing sequence feature, the pose sequence feature, and the background sequence feature through multiple MPDT modules to obtain the first fusion feature value;
[0185] Input the first fusion feature value into the decoder to obtain the image sequence output by the decoder;
[0186] Generate the synthesized fitting video based on the image sequence output by the decoder.
[0187] In another embodiment of the present application, the image sequence output by the decoder includes the initial fitting image sequence and the clothing region mask sequence; the first generation module 704 is specifically configured to:
[0188] Fuse the initial fitting image sequence, the clothing region mask sequence, and the sequence of deformed clothing images to obtain a fused image sequence;
[0189] Fuse the fused image sequence, the background image mask sequence, and the sequence of background images to obtain the synthesized fitting video.
[0190] In another embodiment of the present application, the MPDT module includes a first multi-head attention module, a second multi-head attention module, and a third multi-head attention module; the first generation module 704 is specifically configured to:
[0191] Input the deformed clothing sequence features and the pose attention feature set into the first multi-head attention module of the starting MPDT module to obtain the deformed clothing attention features. The pose attention feature set includes the features obtained by each head of the second multi-head attention module in the starting MPDT module processing the pose sequence features;
[0192] Input the background sequence features and the pose attention feature set into the third multi-head attention module of the starting MPDT module to obtain the background attention features;
[0193] After splicing the deformed clothing attention features and the background attention features, perform a convolution operation to obtain the second fusion feature value;
[0194] For each MPDT module, perform a residual connection between the output and the input of the first multi-head attention module of this MPDT module as the input of the first multi-head attention module of the next MPDT module; perform a residual connection between the second fusion feature value output by this MPDT module and the input of the second multi-head attention module of this MPDT module as the input of the second multi-head attention module of the next MPDT module; perform a residual connection between the output and the input of the third multi-head attention module of this MPDT module as the input of the third multi-head attention module of the next MPDT module, until the first fusion feature value output by the ending MPDT module is obtained.
[0195] In another embodiment of the present application, the first generation module 704 is specifically configured to:
[0196] Input the deformed clothing sequence features into the first multi-head attention module to obtain the first key-value pairs obtained by each head of the first multi-head attention module processing the deformed clothing sequence features;
[0197] Input the pose attention feature set into the first multi-head attention module, so that the first multi-head attention module obtains the deformed clothing attention features based on the pose attention feature set and multiple first key-value pairs;
[0198] The first generation module 704 is specifically configured to:
[0199] Input the background sequence features into the third multi-head attention module to obtain the second key-value pairs obtained by each head of the third multi-head attention module processing the deformed clothing sequence features;
[0200] Input the pose attention feature set into the third multi-head attention module, so that the third multi-head attention module obtains the background attention features based on the pose attention feature set and multiple second key-value pairs.
[0201] In another embodiment of the present application, the first generation module 704 is specifically configured to:
[0202] The fusion calculation is performed according to the following formula:
[0203] I masked1 T = M C1 T ⊙ C1 T +(1 - M C1 T ) ⊙ I R1 T
[0204] The first generation module 704 is specifically used for:
[0205] The fusion calculation is performed according to the following formula:
[0206] I T 1 = (1 - M a1 T ) ⊙ I masked1 T + M a1 T ⊙ A1 T
[0207] Among them, C1 T represents the deformed clothing image sequence, M C1 T represents the clothing area mask sequence, I R1 T represents the initial try-on clothing image sequence, I masked1 T represents the fused image sequence, M a1 T represents the background image mask sequence, A1 T represents the background image sequence, I T 1 represents the synthesized try-on clothing video.
[0208] This application embodiment also provides a video generation device, as Figure 8 shown, the device includes:
[0209] The second acquisition module 801 is used to acquire the video to be processed and the target clothing image, where the video to be processed is a video containing a person;
[0210] The second extraction module 802 is used to extract the pose image sequence and the background image sequence from the video to be processed, where the pose image sequence contains the pose information of the person in each video frame of the video to be processed, and the background image sequence includes the images in each video frame of the video to be processed except for the area to be tried on;
[0211] A second deformation module 803, configured to perform deformation processing on the target clothing image according to the pose information of the person in each video frame of the video to be processed, so as to obtain a sequence of deformed clothing images;
[0212] A second generation module 804, configured to generate a target synthesized virtual try-on video by inputting the sequence of deformed clothing images, the sequence of pose images, the sequence of background images, and the input video into a video generation model, where the video generation model is a video generation model trained by the above-mentioned video generation model training device.
[0213] An embodiment of the present application further provides an electronic device, such as Figure 9 shown, including a processor 901, a communication interface 902, a memory 903, and a communication bus 904. Among them, the processor 901, the communication interface 902, and the memory 903 complete communication with each other through the communication bus 904.
[0214] The memory 903 is used to store a computer program;
[0215] The processor 901 is configured to implement the steps in the above-mentioned video generation model training method or video generation method when executing the program stored in the memory 903.
[0216] The communication bus mentioned in the above-mentioned electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0217] The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0218] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0219] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU for short), a Network Processor (NP for short), etc.; it can also be a Digital Signal Processor (DSP for short), an Application Specific Integrated Circuit (ASIC for short), a Field-Programmable Gate Array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0220] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the video generation model training methods or video generation methods in the above embodiments.
[0221] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, and when it runs on a computer, it causes the computer to execute any one of the video generation model training methods or video generation methods in the above embodiments.
[0222] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a Solid State Disk (SSD)).
[0223] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0224] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.
[0225] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A method for training a video generation model, characterized in that, The method includes: Obtaining a sample try-on video and a sample clothing image, where the sample try-on video is a video of a person wearing the sample clothing; Extracting a pose image sequence and a background image sequence from the sample try-on video, where the pose image sequence contains the pose information of the person in each video frame of the sample try-on video, and the background image sequence includes the images in each video frame of the sample try-on video except for the area to be tried on; Performing a deformation process on the sample clothing image according to the pose information of the person in each video frame of the sample try-on video to obtain a deformed clothing image sequence; Inputting the deformed clothing image sequence, the pose image sequence, and the background image sequence into a video generation model to obtain a synthesized try-on video output by the video generation model; Calculating a loss function value based on the synthesized try-on video and the sample try-on video, and adjusting the parameters of the video generation model based on the loss function value until the video generation model converges, and determining that the training of the video generation model is completed.
2. The method according to claim 1, wherein The video generation model includes: three encoders, multiple MPDT modules based on multi-scale image blocks, and a decoder. Each MPDT module includes three input terminals. The output terminals of different encoders are connected to different input terminals of the starting MPDT module among the multiple MPDT modules. The output terminal of the ending MPDT module among the multiple MPDT modules is connected to the input terminal of the decoder. The MPDT module includes a first multi-head attention module, a second multi-head attention module, and a third multi-head attention module; The step of inputting the deformed clothing image sequence, the pose image sequence, and the background image sequence into the video generation model to obtain a synthesized try-on video output by the video generation model includes: Inputting the deformed clothing image sequence, the pose image sequence, and the background image sequence into an encoder respectively to obtain a deformed clothing sequence feature, a pose sequence feature, and a background sequence feature; Inputting the deformed clothing sequence feature and the pose attention feature set into the first multi-head attention module of the starting MPDT module to obtain a deformed clothing attention feature, where the pose attention feature set includes the features obtained by each head of the second multi-head attention module in the starting MPDT module processing the pose sequence feature; Inputting the background sequence feature and the pose attention feature set into the third multi-head attention module of the starting MPDT module to obtain a background attention feature; After splicing the deformed clothing attention feature and the background attention feature, performing a convolution operation to obtain a second fusion feature value; For each MPDT module, perform a residual connection between the output and the input of the first multi-head attention module of this MPDT module as the input of the first multi-head attention module of the next MPDT module; perform a residual connection between the second fused feature value output by this MPDT module and the input of the second multi-head attention module of this MPDT module as the input of the second multi-head attention module of the next MPDT module; perform a residual connection between the output and the input of the third multi-head attention module of this MPDT module as the input of the third multi-head attention module of the next MPDT module until the first fused feature value output by the end MPDT module is obtained. Input the first fused feature value into the decoder to obtain the image sequence output by the decoder. Generate the synthetic try-on video based on the image sequence output by the decoder.
3. The method according to claim 2, wherein The image sequence output by the decoder includes an initial try-on image sequence and a clothing region mask sequence; the generating of the synthetic try-on video based on the image sequence output by the decoder includes: Fuse the initial try-on image sequence, the clothing region mask sequence, and the deformed clothing image sequence to obtain a fused image sequence. Fuse the fused image sequence, the background image mask sequence, and the background image sequence to obtain the synthetic try-on video.
4. The method according to claim 2, characterized in that The inputting of the deformed clothing sequence feature and the pose attention feature set into the first multi-head attention module of the starting MPDT module to obtain the deformed clothing attention feature includes: Input the deformed clothing sequence feature into the first multi-head attention module to obtain the first key-value pairs processed by each head of the first multi-head attention module for the deformed clothing sequence feature. Input the pose attention feature set into the first multi-head attention module so that the first multi-head attention module obtains the deformed clothing attention feature based on the pose attention feature set and multiple first key-value pairs. The inputting of the background sequence feature and the pose attention feature set into the third multi-head attention module of the starting MPDT module to obtain the background attention feature includes: Input the background sequence feature into the third multi-head attention module to obtain the second key-value pairs processed by each head of the third multi-head attention module for the deformed clothing sequence feature. Input the pose attention feature set into the third multi-head attention module so that the third multi-head attention module obtains the background attention feature based on the pose attention feature set and multiple second key-value pairs.
5. The method according to claim 3, characterized in that, The fusing of the initial try-on image sequence, the clothing region mask sequence, and the deformed clothing image sequence to obtain a fused image sequence includes: Perform fusion calculation according to the following formula: I masked1 T = M C1 T ⊙ C1 T +(1 - M C1 T )⊙ I R1 T Among them, C1 T represents the deformed clothing image sequence, M C1 T represents the clothing area mask sequence, I R1 T represents the initial try-on image sequence, I masked1 T represents the fused image sequence; The fusing of the fused image sequence, the background image mask sequence, and the background image sequence to obtain the synthetic try-on video includes: Perform fusion calculation according to the following formula: I T 1 = (1 - M a1 T ) ⊙ I masked1 T + M a1 T ⊙ A1 T Among them, M a1 T represents the background image mask sequence, and A1 T represents the background image sequence, and I T 1 represents the synthesized virtual fitting video.
6. A video generation method, characterized in that, The method includes: Obtain a video to be processed and a target clothing image, where the video to be processed is a video containing a person. Extract a pose image sequence and a background image sequence from the video to be processed, where the pose image sequence contains the pose information of the person in each video frame of the video to be processed, and the background image sequence includes the images in each video frame of the video to be processed except for the area to be tried on clothes; Perform a deformation process on the target clothing image according to the pose information of the person in each video frame of the video to be processed to obtain a deformed clothing image sequence; Input the deformed clothing image sequence, the pose image sequence, the background image sequence and the input video into a video generation model to obtain a target synthesized try-on video, where the video generation model is a video generation model trained by the method according to any one of claims 1-5.
7. A video generation model training device, characterized in that The device includes: A first acquisition module, configured to acquire a sample try-on video and a sample clothing image, where the sample try-on video is a video of a person wearing a sample clothing; A first extraction module, configured to extract a pose image sequence and a background image sequence from the sample try-on video, where the pose image sequence contains the pose information of the person in each video frame of the sample try-on video, and the background image sequence includes the images in each video frame of the sample try-on video except for the area to be tried on clothes; A first deformation module, configured to perform a deformation process on the sample clothing image according to the pose information of the person in each video frame of the sample try-on video to obtain a deformed clothing image sequence; A first generation module, configured to input the deformed clothing image sequence, the pose image sequence and the background image sequence into a video generation model to obtain a synthesized try-on video output by the video generation model; A calculation module, configured to calculate a loss function value according to the synthesized try-on video and the sample try-on video, and adjust the video generation model parameters based on the loss function value until the video generation model converges, and then determine that the training of the video generation model is completed.
8. A video generation device, characterized in that, The device includes: A second acquisition module, configured to acquire a video to be processed and a target clothing image, where the video to be processed is a video containing a person; A second extraction module, configured to extract a pose image sequence and a background image sequence from the video to be processed, where the pose image sequence contains the pose information of the person in each video frame of the video to be processed, and the background image sequence includes the images in each video frame of the video to be processed except for the area to be tried on clothes; A second deformation module, configured to perform a deformation process on the target clothing image according to the pose information of the person in each video frame of the video to be processed to obtain a deformed clothing image sequence; A second generation module, configured to input the deformed clothing image sequence, the pose image sequence, the background image sequence and the input video into a video generation model to obtain a target synthesized try-on video, where the video generation model is a video generation model trained by the device according to claim 7.
9. An electronic device, characterized in that, Includes a processor, a communication interface, a memory and a communication bus, where the processor, the communication interface and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; A processor, when executing a program stored in a memory, implements the method steps described in any one of claims 1-5 or 6.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method steps described in any one of claims 1-5 or 6 are implemented.