A method, device, equipment and medium for generating virtual fitting videos

By using optical flow information to correct the first pixel coordinate map in the virtual fitting video, the jitter problem caused by unstable timing of the clothing area in the virtual fitting video is solved, and stable virtual fitting video generation is achieved.

CN114638754BActive Publication Date: 2025-06-03BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210236556.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-11
Publication Date
2025-06-03
Estimated Expiration
2042-03-11

AI Technical Summary

Technical Problem

The clothes area in the existing virtual fitting video has a problem of unstable timing, resulting in jitter.

Method used

By acquiring the video to be fitted and the first pixel coordinate diagram sequence, the first pixel coordinate diagram of each video frame is corrected using optical flow information to generate a stable sequence of deformed clothes images, thereby generating a virtual fitted video that is not shaken.

Benefits of technology

The picture of the clothing area in the virtual fitting video is stable and no longer shakes, improving the quality of the virtual fitting experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114638754B_ABST
    Figure CN114638754B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method, apparatus, device, and medium for generating a virtual fitting video, which relates to the field of artificial intelligence technology. The solution includes: obtaining a video to be tried on and a first pixel coordinate map sequence, and for each video frame in the video to be tried on, obtaining the optical flow information between the previous video frame and the current video frame of the current video frame; for each video frame in the video to be tried on, correcting the first pixel coordinate map corresponding to the current video frame based on the optical flow information corresponding to the current video frame to obtain a second pixel coordinate map corresponding to the current video frame; generating a deformed clothing image based on the second pixel coordinate map corresponding to each video frame in the video to be tried on to obtain a sequence of deformed clothing images; and generating a virtual fitting video based on the sequence of deformed clothing images and the video to be tried on. The stability of the picture in the clothing area in the virtual fitting video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and medium for generating virtual fitting videos. Background Art

[0002] With the development of online e-commerce platforms, virtual fitting technology that simulates the clothes selected by users on a person can enhance the shopping experience of users. Before simulating the clothes selected by users on the characters in the original video, it is necessary to deform the clothes selected by users so that the deformed clothes are consistent with the postures and shapes of the characters in the original video.

[0003] However, currently, each generated deformed clothing image is obtained based on a single video frame in the original video. After superimposing each deformed clothing image on a video frame in the original video respectively, there is a problem of temporal instability in the clothing area of the generated virtual fitting video, and the clothing area in the virtual fitting video will jitter. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a method, apparatus, device, and medium for generating virtual fitting videos to achieve a stable picture in the clothing area of the virtual fitting video without jitter. The specific technical solutions are as follows:

[0005] In a first aspect, the embodiments of this application provide a method for generating a virtual fitting video, the method including:

[0006] Obtain a video to be fitted and a first sequence of pixel coordinate maps, where each first pixel coordinate map in the first sequence of pixel coordinate maps is used to represent the mapping relationship between the area to be fitted in a video frame of the video to be fitted and the pixel points of the clothing image to be tried on;

[0007] For each video frame in the video to be fitted, obtain the optical flow information between the previous video frame and the current video frame of this video frame;

[0008] For each video frame in the video to be fitted, correct the first pixel coordinate map corresponding to this video frame based on the optical flow information corresponding to this video frame to obtain a second pixel coordinate map corresponding to this video frame;

[0009] Generate a deformed clothing image based on the second pixel coordinate map corresponding to each video frame in the video to be fitted to obtain a sequence of deformed clothing images;

[0010] Generate a virtual fitting video based on the sequence of deformed clothing images and the video to be fitted.

[0011] In a possible implementation, for each video frame in the to-be-tried-on video, correcting the first pixel coordinate map corresponding to the video frame based on the optical flow information corresponding to the video frame to obtain a second pixel coordinate map corresponding to the video frame, including:

[0012] Converting the first pixel coordinate map corresponding to the video frame into a first vector;

[0013] Performing a correction calculation on the first vector based on the optical flow information corresponding to the video frame to obtain a second vector;

[0014] Converting the second vector into a second pixel coordinate map.

[0015] In a possible implementation, the performing a correction calculation on the first vector based on the optical flow information corresponding to the video frame to obtain a second vector includes:

[0016] Converting the first pixel coordinate map corresponding to the previous video frame of the video frame into a third vector;

[0017] Calculating a predicted vector corresponding to the video frame according to the third vector and the optical flow information between the video frame and the previous video frame;

[0018] Performing a correction calculation on the first vector based on the predicted vector to obtain the second vector.

[0019] In a possible implementation, the calculating a predicted vector corresponding to the video frame according to the third vector and the optical flow information between the video frame and the previous video frame includes:

[0020] Calculating the predicted vector according to the following formula:

[0021] f′ t =ω t-1 (f t-1 )

[0022] where ω t-1 () is an optical flow prediction function obtained according to the optical flow information from the (t - 1)-th frame to the t-th frame in the to-be-tried-on video, and f t-1 represents the vector obtained by converting the first pixel coordinate map corresponding to the (t - 1)-th video frame.

[0023] In a possible implementation, the performing a correction calculation on the first vector based on the predicted vector to obtain the second vector includes:

[0024] Calculating the difference between the first vector and the predicted vector;

[0025] When the difference is less than or equal to a preset threshold, using the predicted vector to perform a correction calculation on the first vector to obtain the second vector;

[0026] When the difference is greater than the preset threshold, the first vector is used as the second vector.

[0027] In a possible implementation, the first pixel coordinate map sequence is obtained through the following steps:

[0028] For each video frame in the to-be-tried-on video, obtain the first key point image, the second key point image, and the background region image of this video frame; both the first key point image and the second key point image are used to represent the human pose in the to-be-tried-on video, and the background region image includes the background region in the to-be-tried-on video except the to-be-tried-on region, and the occluder image on the to-be-tried-on region;

[0029] For each video frame in the to-be-tried-on video, input the first key point image, the second key point image, the background region image, and the to-be-tried-on clothes image of this video frame into the deformation warp module to obtain the first pixel coordinate map of this video frame.

[0030] In a possible implementation, generating the virtual try-on video based on the deformed clothes image sequence and the to-be-tried-on video includes:

[0031] Input the deformed clothes image sequence, the first key point image sequence, the second key point image sequence, and the background region image sequence of the to-be-tried-on video into the try-on Try-on module to obtain the virtual try-on video.

[0032] In a second aspect, an embodiment of the present application provides a virtual try-on video generation device, and the device includes:

[0033] A first acquisition module, configured to acquire a to-be-tried-on video and a first pixel coordinate map sequence, and each first pixel coordinate map in the first pixel coordinate map sequence is used to represent: the mapping relationship between the to-be-tried-on region of a video frame in the to-be-tried-on video and the pixel points of the to-be-tried-on clothes image;

[0034] A second acquisition module, configured to, for each video frame in the to-be-tried-on video, acquire the optical flow information between the previous video frame and this video frame of this video frame;

[0035] A correction module, configured to, for each video frame in the to-be-tried-on video, correct the first pixel coordinate map corresponding to this video frame based on the optical flow information corresponding to this video frame to obtain the second pixel coordinate map corresponding to this video frame;

[0036] A first generation module, configured to generate a deformed clothes image based on the second pixel coordinate map corresponding to each video frame in the to-be-tried-on video, to obtain a deformed clothes image sequence;

[0037] A second generation module, configured to generate a virtual fitting video based on the deformed clothing image sequence and the video to be tried on.

[0038] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0039] The memory is used to store a computer program;

[0040] The processor is configured to implement the method steps described in the first aspect above when executing the program stored on the memory.

[0041] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the virtual fitting video generation method described in any one of the above is implemented.

[0042] In a fifth aspect, an embodiment of the present application provides a computer program product containing instructions, which when running on a computer, causes the computer to execute the virtual fitting video generation method described in any one of the above.

[0043] By adopting the above technical solutions, the optical flow information between adjacent frames in the video to be tried on can be obtained, and the optical flow information is used to correct the first pixel coordinate map corresponding to each video frame of the video to be tried on. Since the optical flow information is the motion relationship of pixel points between adjacent frames in the video to be tried on, after correcting the first pixel coordinate map of each video frame using the motion relationship of pixel points between adjacent frames, the motion relationship of pixel points between adjacent first pixel coordinate maps in the first pixel coordinate map sequence can be made to conform to the motion relationship of pixel points between adjacent video frames in the video to be tried on, which also makes the time sequence of the deformed clothing image sequence generated according to the first pixel coordinate map sequence stable. Furthermore, the picture of the clothing area in the finally generated virtual fitting video is stable and does not shake. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art.

[0045] Figure 1 It is a schematic structural diagram of a virtual fitting device provided by an embodiment of the present application;

[0046] Figure 2 It is a flowchart of a virtual fitting video generation method provided by an embodiment of the present application;

[0047] Figure 3aAn exemplary schematic diagram of a first key point image provided by an embodiment of the present application;

[0048] Figure 3b An exemplary schematic diagram of a second key point image provided by an embodiment of the present application;

[0049] Figure 4 A flowchart of another virtual fitting video generation method provided by an embodiment of the present application;

[0050] Figure 5 A structural schematic diagram of a virtual fitting video generation device provided by an embodiment of the present application;

[0051] Figure 6 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0052] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.

[0053] The embodiment of the present application provides a virtual fitting video generation method, which is applied to a virtual fitting device, as Figure 1 shown. The system includes a Warp module 101, a correction module 102, and a Try-on module 103.

[0054] Among them, the Warp module 101 is used to generate a pixel coordinate map corresponding to each video frame included in the to-be-tried-on clothing video based on the to-be-tried-on clothing image and the to-be-tried-on clothing video. Specifically, the Warp module 101 can learn the point-to-point mapping relationship between the pixel points in the to-be-tried-on clothing image and the human to-be-tried-on area of each video frame through an appearance-flow-based warp method, and then generate a pixel coordinate map for representing the above mapping relationship for each video frame.

[0055] The correction module 102 is used to correct the pixel coordinate map corresponding to each video frame generated by the Warp module 101. Furthermore, the Warp module 101 is also used to deform the to-be-tried-on clothing image based on the corrected pixel coordinate map corresponding to each video frame to obtain a deformed clothing image sequence.

[0056] The Try-on module 103 is used to generate a virtual fitting video according to the deformed clothing image sequence generated by the Warp module 101 and the to-be-tried-on clothing video.

[0057] Based on Figure 1 the device shown, the virtual fitting video generation method provided by the embodiment of the present application will be introduced in detail below.

[0058] An embodiment of the present application provides a method for generating a virtual fitting video, which can be applied to an electronic device. The electronic device can be a device such as a smart phone, a tablet computer, a desktop computer, a server, etc. For example, Figure 2 as shown, the method includes:

[0059] S201. Obtain a video to be tried on and a first sequence of pixel coordinate maps.

[0060] Each first pixel coordinate map in the first sequence of pixel coordinate maps is used to represent the mapping relationship between the area to be tried on in a video frame in the video to be tried on and the pixel points of the image of the clothes to be tried on.

[0061] S202. For each video frame in the video to be tried on, obtain the optical flow information between the previous video frame and the current video frame of this video frame.

[0062] S203. For each video frame in the video to be tried on, correct the first pixel coordinate map corresponding to this video frame based on the optical flow information corresponding to this video frame, and obtain the second pixel coordinate map corresponding to this video frame.

[0063] S204. Based on the second pixel coordinate map corresponding to each video frame in the video to be tried on, generate a deformed clothes image, and obtain a sequence of deformed clothes images.

[0064] S205. Generate a virtual fitting video based on the sequence of deformed clothes images and the video to be tried on.

[0065] By using the embodiment of the present application, the optical flow information between adjacent frames in the video to be tried on can be obtained, and the first pixel coordinate map corresponding to each video frame in the video to be tried on is corrected by using the optical flow information. Since the optical flow information is the motion relationship between pixel points between adjacent frames in the video to be tried on, after correcting the first pixel coordinate map of each video frame by using the motion relationship between pixel points between adjacent frames, the motion relationship between pixel points between adjacent first pixel coordinate maps in the first sequence of pixel coordinate maps can be made to conform to the motion relationship between pixel points between adjacent video frames in the video to be tried on, which also makes the sequence of deformed clothes images generated according to the first sequence of pixel coordinate maps stable in time series. Furthermore, the picture of the clothes area in the finally generated virtual fitting video is stable and does not shake.

[0066] Regarding the above S201, the video to be tried on is a video uploaded by a user and containing a person who wants to try on clothes, and the image of the clothes to be tried on is an image of the clothes to be virtually tried on by the user in a flat state.

[0067] For example, in various software with virtual fitting functions, users can upload a video containing a human figure and select the clothes to be tried on. Among them, the video to be tried on is the video uploaded by the user, and the image of the clothes to be tried on is the image of the clothes selected by the user for trying on.

[0068] The first pixel coordinate map is a two-dimensional map with a size of H*W, that is, a two-dimensional map with a height and width of H and W respectively. There are a total of H*W points in this two-dimensional map, and each point has a two-dimensional coordinate, and each two-dimensional coordinate represents the position of a pixel point on the clothes to be tried on.

[0069] The first pixel coordinate map sequence is obtained through the following steps:

[0070] Step 1: For each video frame in the video to be tried on, obtain the first key point image, the second key point image, and the background region image of this video frame.

[0071] Among them, both the first key point image and the second key point image are used to represent the human pose in this video frame of the video to be tried on. The background region image includes the background region other than the region to be tried on in the video to be tried on, and the image of the occluder on the region to be tried on.

[0072] In one implementation, the human pose in each video frame of the video to be tried on can be recognized through a human pose recognition algorithm to obtain the corresponding first key point map and the second key point image for each video frame. The first key point image is a pose image containing human key points, and the second key point image is a densepose image containing the shapes of each part of the human body. The first key point image can be obtained through the Openpose algorithm, as Figure 3a shown, Figure 3a is an exemplary schematic diagram of the first key point image. The second key point image can be obtained through the Densepose algorithm, as Figure 3b shown, Figure 3b is an exemplary schematic diagram of the second key point image; among them, the Openpose algorithm is a human pose recognition algorithm used to recognize the key points of each joint of the human body in an image, and the Densepose algorithm is a human pose recognition algorithm that converts a 2D human image into a 3D human image.

[0073] The background region image is obtained by processing the region to be tried on of the human figure in each video frame of the video to be tried on into the same pixel value. It can be understood that in the case where there is no occluder in the video frame of the video to be tried on, the background region image does not include the occluder image either.

[0074] Step 2: For each video frame in the to-be-tried-on video, input the first key-point image, the second key-point image, the background region image, and the to-be-tried-on clothing image of this video frame into the warp module to obtain the first pixel coordinate map.

[0075] In the embodiments of the present application, the first pixel coordinate map corresponding to each video frame in the to-be-tried-on video can be obtained, and these first pixel coordinate maps can form a first pixel coordinate map sequence.

[0076] Among them, the warp module can be a warp module based on a Thin Plate Spline (TPS) model or an Appearance Flow Warping Module (AFWM) model.

[0077] Regarding the above S202, among them, the optical flow information represents the motion relationship of pixel points in the to-be-tried-on video between frames.

[0078] In one implementation, the to-be-tried-on video can be input into an optical flow prediction network to obtain the optical flow information between adjacent video frames in the to-be-tried-on video. For example, the optical flow prediction network can be FlowNet2.0.

[0079] Regarding the above S203, the optical flow information obtained in the embodiments of the present application is used to correct the pixel coordinates in each frame of the pixel coordinate map from the second frame to the last frame in the first pixel coordinate map sequence, and the pixel coordinates in the first pixel coordinate map of the first frame in the first pixel coordinate map sequence are not corrected.

[0080] Regarding the above S204, in one implementation, the pixel points at the corresponding coordinate positions can be found from the to-be-tried-on clothing image according to the pixel coordinates of each point in the second pixel coordinate map and filled into the second pixel coordinate map to obtain the deformed clothing image.

[0081] For example, if the height and width of the two-dimensional map are 6 and 3 respectively, there are a total of 6 * 3 = 18 points on the two-dimensional map. Assuming that one of the points is point a, and the two-dimensional coordinates of point a are (m, n), then when generating the anti-occlusion deformed clothing, the pixel point at the coordinate position (m, n) on the anti-occlusion to-be-tried-on clothing image needs to be filled into point a in the two-dimensional map.

[0082] Regarding the above S205, in one implementation, the deformed clothing image sequence, the first key-point image sequence, the second key-point image sequence, and the background region image sequence of the to-be-tried-on video can be input into the Try-on module to obtain the virtual try-on video.

[0083] Among them, the first key point image sequence is a sequence composed of first key point images corresponding to each video frame of the to-be-tried-on video.

[0084] The second key point image sequence is a sequence composed of second key point images corresponding to each video frame of the to-be-tried-on video.

[0085] The background region image sequence is a sequence composed of background region images corresponding to each video frame of the to-be-tried-on video.

[0086] Among them, the try-on module can be the Try-on module in try-on solutions such as CP-VTON, Parser Free AppearenceFlow Network (PF-AFN), Adaptive Content Generation and Preserving Network (ACGPN), or VITON-HD. CP-VTON is a feature-preserving virtual try-on network (CP-VTON) proposed at the International Conference on Computer Vision in Europe, and VITON-HD is an image-based virtual try-on network.

[0087] In another embodiment of the present application, the first pixel coordinate map can be converted into a vector, the converted vector is corrected, and the corrected vector is converted into a second pixel coordinate map. As Figure 4 shown, the above S203 can be implemented as:

[0088] S2031. Convert the first pixel coordinate map corresponding to this video frame into a first vector.

[0089] Among them, the first vector is a one-dimensional vector.

[0090] In one implementation, the first pixel coordinate map can be converted into a one-dimensional vector through a reshape function.

[0091] S2032. Perform a correction calculation on the first vector based on the optical flow information corresponding to this video frame to obtain a second vector.

[0092] S2032 can be specifically implemented as:

[0093] Step 1. Convert the first pixel coordinate map corresponding to the previous video frame of this video frame into a third vector.

[0094] Step 2. Calculate the predicted vector corresponding to this video frame according to the third vector and the optical flow information between this video frame and the previous video frame.

[0095] In the embodiment of the present application, the predicted vector can be calculated according to the following formula:

[0096] f' t = ω t-1 (f t-1 )

[0097] where ω t-1 () is the optical flow prediction function obtained from the optical flow information of the (t - 1)-th frame to the t-th frame in the to-be-tried-on video, and f t-1 represents the vector obtained by converting the first pixel coordinate map corresponding to the (t - 1)-th frame video frame.

[0098] Step 3: Perform a correction calculation on the first vector based on the prediction vector to obtain a second vector.

[0099] In the embodiments of the present application, the difference between the first vector and the prediction vector can be calculated; when the difference is less than or equal to a preset threshold, the prediction vector is used to perform a correction calculation on the first vector to obtain a second vector; when the difference is greater than the preset threshold, the first vector is used as the second vector.

[0100] Specifically, the second vector can be calculated according to the following optical flow correction formula

[0101]

[0102] δ t = ‖f t - ω t-1 (f t-1 )‖

[0103] where f t represents the first vector of the first pixel coordinate map corresponding to the t-th frame video frame, ω t-1 () is the optical flow prediction function obtained from the optical flow information of the (t - 1)-th frame to the t-th frame in the to-be-tried-on video, δ t represents the difference between the first vector of the first pixel coordinate map corresponding to the t-th frame video frame and the prediction vector corresponding to the t-th frame video frame, ε represents the preset threshold of the difference δ t , and Ω represents the overlapping area between the clothing area in the t-th frame video frame of the to-be-tried-on video and the deformed clothing image of the first pixel coordinate map.

[0104] It should be noted that when performing correction through this formula, the prediction vector ω t-1 () corresponding to the t-th frame video frame can be obtained through predictive calculation using the optical flow prediction function ω t-1 () between the (t - 1)-th frame video frame and the t-th frame video frame and the first vector f t-1 (f t-1 ).

[0105] Through the formula δ t = ‖f t - ωt-1 (f t-1 ) ‖ Calculate the difference between the first vector of the first pixel coordinate map corresponding to the t-th video frame and the predicted vector corresponding to the t-th video frame.

[0106] When the difference δ t is less than or equal to the preset threshold ε, use the predicted vector to correct the first vector of the first pixel coordinate map.

[0107] When the difference δ t is greater than the preset threshold ε, do not correct the first vector of the first pixel coordinate map.

[0108] Among them, the preset threshold ε can be set according to actual needs. For example, it can be set to 0.05.

[0109] When using this formula, only correct the pixel coordinates of the points in the first pixel coordinate map that belong to the overlapping region Ω, and do not correct the pixel coordinates of the points in the first pixel coordinate map that do not belong to the region outside the overlapping region.

[0110] In one implementation, the overlapping region Ω can be obtained through the region where the deformed clothing image generated from the first pixel coordinate map overlaps with the clothing region of the person in the video frame of the to-be-tried-on video corresponding to the first pixel coordinate map.

[0111] S2033. Convert the second vector into a second pixel coordinate map.

[0112] In one implementation, the second vector can be converted into a second pixel coordinate map through the reshape function.

[0113] Adopting the embodiments of the present application, by converting the first pixel coordinate image into a first vector, then using the optical flow correction formula to perform correction calculation on the first vector to obtain a second vector, and further, the second vector can be converted into a second pixel coordinate map, which realizes the correction of the first pixel coordinate map according to the optical flow information. At the same time, the pixel coordinates can be accurately adjusted for the overlapping region through the optical flow correction formula, so that the picture of the clothing region in the finally generated virtual try-on video is stable and does not shake.

[0114] Corresponding to the above method embodiments, the embodiments of the present application also provide a virtual try-on video generation device, as Figure 5 shown, the device includes:

[0115] The first acquisition module 501 is used to acquire the to-be-tried-on video and the first pixel coordinate map sequence. Each first pixel coordinate map in the first pixel coordinate map sequence is used to represent: the mapping relationship between the to-be-tried-on region of a video frame in the to-be-tried-on video and the pixel points of the to-be-tried-on clothing image;

[0116] The second acquisition module 502 is configured to obtain the optical flow information between the previous video frame and the current video frame for each video frame in the to-be-tried-on video;

[0117] The correction module 503 is configured to correct the first pixel coordinate map corresponding to each video frame in the to-be-tried-on video based on the optical flow information corresponding to the video frame, so as to obtain the second pixel coordinate map corresponding to the video frame;

[0118] The first generation module 504 is configured to generate a deformed clothing image based on the second pixel coordinate map corresponding to each video frame in the to-be-tried-on video, so as to obtain a deformed clothing image sequence;

[0119] The second generation module 505 is configured to generate a virtual try-on video based on the deformed clothing image sequence and the to-be-tried-on video.

[0120] In another embodiment of the present application, the correction module 503 is specifically configured to:

[0121] Convert the first pixel coordinate map corresponding to the video frame into a first vector;

[0122] Perform correction calculation on the first vector based on the optical flow information corresponding to the video frame to obtain a second vector;

[0123] Convert the second vector into a second pixel coordinate map.

[0124] In another embodiment of the present application, the correction module 503 is specifically configured to:

[0125] Convert the first pixel coordinate map corresponding to the previous video frame of the video frame into a third vector;

[0126] Calculate the predicted vector corresponding to the video frame according to the third vector and the optical flow information between the video frame and the previous video frame;

[0127] Perform correction calculation on the first vector based on the predicted vector to obtain a second vector.

[0128] In another embodiment of the present application, the correction module 503 is specifically configured to:

[0129] Calculate the predicted vector according to the following formula:

[0130] f′ t =ω t-1 (f t-1 )

[0131] where ω t-1 () is an optical flow prediction function obtained according to the optical flow information from the (t-1)-th frame to the t-th frame in the to-be-tried-on video, and f t-1 represents the vector obtained by converting the first pixel coordinate map corresponding to the (t-1)-th frame video frame.

[0132] In another embodiment of the present application, the correction module 503 is specifically configured to:

[0133] Calculate the difference between the first vector and the predicted vector;

[0134] When the difference is less than or equal to a preset threshold, use the predicted vector to perform a correction calculation on the first vector to obtain a second vector;

[0135] When the difference is greater than the preset threshold, use the first vector as the second vector.

[0136] In another embodiment of the present application, the first acquisition module 501 is further configured to:

[0137] For each video frame in the to-be-tried-on video, obtain the first key point image, the second key point image, and the background region image of the video frame; both the first key point image and the second key point image are used to represent the human posture in the to-be-tried-on video, and the background region image includes the background region other than the to-be-tried-on region in the to-be-tried-on video, and the occlusion image on the to-be-tried-on region.

[0138] For each video frame in the to-be-tried-on video, input the first key point image, the second key point image, the background region image, and the to-be-tried-on clothing image of the video frame into the deformation warp module to obtain the first pixel coordinate map of the video frame.

[0139] In another embodiment of the present application, the second generation module 505 is specifically configured to:

[0140] Input the deformed clothing image sequence, the first key point image sequence, the second key point image sequence, and the background region image sequence of the to-be-tried-on video into the Try-on module to obtain a virtual try-on video.

[0141] The embodiment of the present application further provides an electronic device, as Figure 6 shown, including a processor 601, a communication interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.

[0142] The memory 603 is used to store a computer program;

[0143] When the processor 601 is used to execute the program stored in the memory 603, it implements the steps of the virtual try-on video generation method in the above embodiment.

[0144] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0145] The communication interface is used for communication between the above electronic device and other devices.

[0146] The memory can include a Random Access Memory (RAM), and can also include a non-volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.

[0147] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0148] In another embodiment provided by this application, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the virtual fitting video generation method described in any one of the above embodiments.

[0149] In another embodiment provided by this application, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute the virtual fitting video generation method described in any one of the above embodiments.

[0150] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0151] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.

[0152] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.

[0153] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.

Claims

1. A method for generating a virtual fitting video, characterized in that, the method includes: Obtain a video to be tried on and a first sequence of pixel coordinate maps, where each first pixel coordinate map in the first sequence of pixel coordinate maps is used to represent: the mapping relationship between the area to be tried on in a video frame of the video to be tried on and the pixel points of the clothing image to be worn; For each video frame in the video to be tried on, obtain the optical flow information between the previous video frame and the current video frame of this video frame; For each video frame in the video to be tried on, correct the first pixel coordinate map corresponding to this video frame based on the optical flow information corresponding to this video frame to obtain the second pixel coordinate map corresponding to this video frame; Based on the second pixel coordinate map corresponding to each video frame in the video to be tried on, generate a deformed clothing image to obtain a sequence of deformed clothing images; Generate a virtual fitting video based on the sequence of deformed clothing images and the video to be tried on; The step of, for each video frame in the video to be tried on, correcting the first pixel coordinate map corresponding to this video frame based on the optical flow information corresponding to this video frame to obtain the second pixel coordinate map corresponding to this video frame includes: Convert the first pixel coordinate map corresponding to this video frame into a first vector; Perform a correction calculation on the first vector based on the optical flow information corresponding to this video frame to obtain a second vector; Convert the second vector into a second pixel coordinate map; The step of performing a correction calculation on the first vector based on the optical flow information corresponding to this video frame to obtain a second vector includes: Convert the first pixel coordinate map corresponding to the previous video frame of this video frame into a third vector; According to the third vector and the optical flow information between this video frame and the previous video frame, calculate the predicted vector corresponding to this video frame; Perform a correction calculation on the first vector based on the predicted vector to obtain the second vector.

2. The method according to claim 1, characterized in that, the step of calculating the predicted vector corresponding to this video frame according to the third vector and the optical flow information between this video frame and the previous video frame includes: Calculate the predicted vector according to the following formula: f t ′ = ω t-1 (f t-1 ) Among them, ω t-1 () is an optical flow prediction function obtained from the optical flow information from the (t - 1)-th frame to the t-th frame in the to-be-tried-on video, and f t-1 represents the vector obtained by converting the first pixel coordinate map corresponding to the (t - 1)-th frame video frame.

3. The method according to claim 2, characterized in that, the step of performing a correction calculation on the first vector based on the predicted vector to obtain the second vector includes: Calculate the difference between the first vector and the predicted vector; When the difference is less than or equal to a preset threshold, use the predicted vector to perform a correction calculation on the first vector to obtain the second vector; When the difference is greater than the preset threshold, use the first vector as the second vector.

4. The method according to any one of claims 1-3, characterized in that, the first sequence of pixel coordinate maps is obtained through the following steps: For each video frame in the video to be tried on, obtain the first key point image, the second key point image and the background region image of this video frame; both the first key point image and the second key point image are used to represent the human pose in the video to be tried on, and the background region image includes the background region in the video to be tried on except the area to be tried on, and the image of the occluder on the area to be tried on; For each video frame in the to-be-tried-on video, input the first key-point image, the second key-point image, the background region image, and the to-be-tried-on clothing image of this video frame into the deformation warp module to obtain the first pixel coordinate map of this video frame.

5. The method according to claim 4, wherein, generating the virtual try-on video based on the deformed clothing image sequence and the to-be-tried-on video includes: inputting the deformed clothing image sequence, the first key-point image sequence, the second key-point image sequence, and the background region image sequence of the to-be-tried-on video into the try-on module to obtain the virtual try-on video.

6. A virtual try-on video generation device, wherein, the device includes: a first acquisition module, configured to acquire a to-be-tried-on video and a first pixel coordinate map sequence, and each first pixel coordinate map in the first pixel coordinate map sequence is used to represent: the mapping relationship between the to-be-tried-on area of a video frame in the to-be-tried-on video and the pixel points of the to-be-tried-on clothing image; a second acquisition module, configured to, for each video frame in the to-be-tried-on video, acquire the optical flow information between the previous video frame of this video frame and this video frame; a correction module, configured to, for each video frame in the to-be-tried-on video, correct the first pixel coordinate map corresponding to this video frame based on the optical flow information corresponding to this video frame to obtain the second pixel coordinate map corresponding to this video frame; a first generation module, configured to generate a deformed clothing image based on the second pixel coordinate map corresponding to each video frame in the to-be-tried-on video to obtain a deformed clothing image sequence; a second generation module, configured to generate a virtual try-on video based on the deformed clothing image sequence and the to-be-tried-on video; The correction module is specifically configured to: convert the first pixel coordinate map corresponding to this video frame into a first vector; perform a correction calculation on the first vector based on the optical flow information corresponding to this video frame to obtain a second vector; convert the second vector into a second pixel coordinate map; The correction module is specifically configured to: convert the first pixel coordinate map corresponding to the previous video frame of this video frame into a third vector; calculate the predicted vector corresponding to this video frame according to the third vector and the optical flow information between this video frame and the previous video frame; perform a correction calculation on the first vector based on the predicted vector to obtain a second vector.

7. An electronic device, wherein, it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete mutual communication through the communication bus; the memory is used for storing a computer program; the processor is configured to, when executing the program stored on the memory, implement the method steps of any one of claims 1-5.

8. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method steps of any one of claims 1-5.

Citation Information

Patent Citations

  • Virtual hair style modeling method of images and videos

    CN103606186A

  • Video virtual try-on method and device based on mixed optical flow

    CN111275518A