Method, apparatus, device, storage medium and program product for virtually trying on a garment
By optimizing virtual clothing try-on technology using STMVN and DAFVN networks, and by utilizing multi-scale feature maps, optical flow fields, attention maps, and historical frame information, the problems of clothing misalignment and artifacts were solved, thereby improving the coherence and realism of virtual try-on videos.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2026-03-24
AI Technical Summary
Traditional video-based virtual try-on clothing technology is prone to significant misalignment and distortion of patterns and text on clothing during the pose transition process for generating prediction results, resulting in a lack of realism in the generated videos.
We employ the Spatiotemporal Memory Video Virtual Try-On Network (STMVN) and the Deformable Attention Flow-Based Try-On Network (DAFVN). By optimizing multi-scale feature maps, optical flow fields, and attention maps, and combining historical frame information for weighted summation, we optimize the try-on results for each human image and improve the coherence of the video.
In complex poses and with large deformations, it improves the problems of clothing misalignment and artifacts, and enhances the coherence and realism of the synthesized video.
Smart Images

Figure CN119417574B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer equipment, storage medium, and computer program product for virtual try-on of clothing. Background Technology
[0002] With the development of artificial intelligence technology, computer vision is increasingly entering people's lives. From automatic navigation to facial recognition payment, from intelligent object recognition to one-click image cutout, these applications of computer vision technology are influencing and changing people's lifestyles in all aspects.
[0003] In the e-commerce sector, people can buy clothing without leaving home. The emergence of virtual try-on technology allows people to try on clothes in a virtual space while shopping online.
[0004] Some virtual try-on technologies can be video-based, directly replacing clothing within the video. Users can upload a video containing themselves and an image of the clothing they want to try on; this technology can then generate a composite video of them trying on the clothing.
[0005] Traditional video-based virtual try-on clothing technology suffers from a lack of guidance information during the pose transition process to generate prediction results. This can lead to significant misalignment and distortion of patterns and text on the clothing, resulting in a lack of realism in the final video. Summary of the Invention
[0006] Therefore, it is necessary to provide a method, apparatus, computer equipment, storage medium, and computer program product for virtual clothing try-on in response to the above-mentioned technical problems.
[0007] This application provides a method for virtual clothing try-on, the method comprising:
[0008] Extract multiple frames of human body images from a video;
[0009] Based on the clothing images and each frame of the human body image, preliminary fitting results are obtained for each frame of the human body image.
[0010] The initial fitting results of each human body image are optimized to obtain the optimized fitting results of each human body image.
[0011] Specifically, when obtaining the preliminary fitting results of the current frame human body image, multiple scales of current frame human body feature maps and clothing feature maps are obtained based on the current frame human body image and clothing image. The optical flow field and attention map are optimized according to the current frame human body feature maps and clothing feature maps from small scale to large scale to obtain the target optical flow field and target attention map. The target optical flow field and target attention map are used as guiding information to process the current frame human body image and clothing image to obtain the preliminary fitting results of the current frame human body image.
[0012] Specifically, when optimizing the preliminary try-on results of the current frame human body image, the preliminary try-on results of the current frame human body image, the preliminary try-on results of historical frame human body images, and the optimized try-on results of historical frame human body images are obtained; the current frame key feature map is obtained based on the preliminary try-on results of the current frame human body image; the historical frame key feature map is obtained based on the preliminary try-on results of historical frame human body images; the historical frame value feature map is obtained based on the optimized try-on results of historical frame human body images; and the historical frame value feature map is weighted and summed based on the similarity between the current frame key feature map and the historical frame key feature map to obtain the optimized try-on result of the current frame human body image.
[0013] In one embodiment, the optical flow field and attention map are optimized based on the current frame's human feature map and clothing feature map from small scale to large scale to obtain the target optical flow field and target attention map, including:
[0014] Based on the current frame human feature map from small scale to large scale, obtain the self-optical flow field and self-attention map corresponding to each scale, and take the self-optical flow field and self-attention map corresponding to the last scale as the target self-optical flow field and target self-attention map.
[0015] Specifically, when obtaining the self-optical flow field and self-attention map corresponding to the current scale, if the current scale is not the minimum scale, the current frame human feature map of the current scale is obtained and fused with the current frame human feature map of the previous scale to obtain the current frame human fusion feature map. Based on the self-optical flow field and self-attention map corresponding to the previous scale, the current frame human fusion feature map is transformed to obtain the current frame human transformation feature map. Based on the current frame human transformation feature map, the self-optical flow field and self-attention map corresponding to the current scale are obtained.
[0016] Based on the clothing feature maps from small scale to large scale, the cross optical flow field and cross attention map corresponding to each scale are obtained. The cross optical flow field and cross attention map corresponding to the last scale are taken as the target cross optical flow field and target cross attention map.
[0017] Specifically, when obtaining the cross optical flow field and cross attention map corresponding to the current scale, if the current scale is not the minimum scale, the clothing feature map of the current scale is obtained and fused with the clothing feature map of the previous scale to obtain the clothing fusion feature map. Based on the cross optical flow field and cross attention map corresponding to the previous scale, the clothing fusion feature map is transformed to obtain the clothing transformation feature map. Based on the clothing transformation feature map, the cross optical flow field and cross attention map corresponding to the current scale are obtained; the previous scale is smaller than the current scale.
[0018] In one embodiment, the method further includes:
[0019] When obtaining the self-optical flow field and self-attention map corresponding to the current scale, if the current scale is not the minimum scale, the self-optical flow field and self-attention map corresponding to the current scale are obtained based on the human body feature map of the current frame at the current scale.
[0020] In one embodiment, the method further includes:
[0021] When obtaining the cross-optical flow field and cross-attention map corresponding to the current scale, if the current scale is the smallest scale, the cross-optical flow field and cross-attention map corresponding to the current scale are obtained based on the clothing feature map of the current scale.
[0022] In one embodiment, the target optical flow field and target attention map are used as guiding information to process the current frame human body image and clothing image to obtain a preliminary fitting result of the current frame human body image, including:
[0023] The current frame's human body image and clothing image are shallow encoded separately to obtain the current frame's shallow encoded human body feature map and shallow encoded clothing feature map;
[0024] Based on the target self-optical flow field, the target self-attention map and the current frame shallow-coded human body feature map, deformable attention distortion is performed to obtain the current frame target human body clothing feature map;
[0025] Based on the target cross optical flow field, the target cross attention map and the shallow-coded clothing feature map, deformable attention distortion is performed to obtain the target clothing feature map;
[0026] Merge the target human body clothing feature map and the target clothing feature map in the current frame to obtain the merged feature map of the current frame;
[0027] The current frame's merged feature map is shallow-decoded to obtain preliminary fitting results for the current frame's human body image.
[0028] In one embodiment, based on the similarity between the current frame key feature map and the historical frame key feature maps, a weighted sum of the historical frame value feature maps is performed to obtain the optimized fitting result of the current frame human body image, including:
[0029] Based on the relative similarity between the current frame key feature map and each historical frame key feature map, the relative weights assigned to each historical frame key feature map are determined, thus obtaining the weights of each historical frame key feature map; the greater the similarity, the greater the weight.
[0030] Based on the weights of the feature maps of each historical frame, the feature maps of each historical frame are weighted and summed to obtain the optimized fitting result of the human body image in the current frame.
[0031] This application provides a device for virtual clothing try-on, the device comprising:
[0032] The human body image acquisition module is used to acquire multiple frames of human body images from videos;
[0033] The initial fitting processing module is used to obtain preliminary fitting results for each frame of the human body image based on the clothing image and each frame of the human body image.
[0034] The optimized try-on processing module is used to optimize the preliminary try-on results of each frame of human body image to obtain the optimized try-on results of each frame of human body image.
[0035] When obtaining the preliminary fitting results of the current frame human body image, the initial fitting processing module is further configured to: obtain multiple scales of current frame human body feature maps and clothing feature maps based on the current frame human body image and clothing image; optimize the optical flow field and attention map according to the current frame human body feature maps and clothing feature maps from small scale to large scale to obtain the target optical flow field and target attention map; and use the target optical flow field and target attention map as guiding information to process the current frame human body image and clothing image to obtain the preliminary fitting results of the current frame human body image.
[0036] Specifically, when optimizing the preliminary fitting results of the current frame human body image, the optimized fitting processing module is further configured to: obtain the preliminary fitting results of the current frame human body image, the preliminary fitting results of historical frame human body images, and the optimized fitting results of historical frame human body images; obtain the current frame key feature map based on the preliminary fitting results of the current frame human body image; obtain the historical frame key feature map based on the preliminary fitting results of historical frame human body images; obtain the historical frame value feature map based on the optimized fitting results of historical frame human body images; and perform a weighted summation of the historical frame value feature map based on the similarity between the current frame key feature map and the historical frame key feature map to obtain the optimized fitting result of the current frame human body image.
[0037] This application provides a computer device, including a memory and a processor, wherein the memory stores a computer program and the processor executes the above-described method.
[0038] This application provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor using the methods described above.
[0039] This application provides a computer program product having a computer program stored thereon, the computer program being executed by a processor using the above-described method.
[0040] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for virtual clothing try-on acquire multiple frames of human body images from a video; based on the clothing images and each frame of human body images, obtain preliminary try-on results for each frame of human body images; optimize the preliminary try-on results for each frame of human body images to obtain optimized try-on results for each frame of human body images; when obtaining the preliminary try-on results for the current frame of human body images, based on the current frame of human body images and clothing images, obtain current frame human body feature maps and clothing feature maps at multiple scales; optimize the optical flow field and attention map according to the current frame human body feature maps and clothing feature maps from small scale to large scale to obtain target optical flow field and target attention map; use the target optical flow field and target attention map as guiding information to process the current frame of human body images and clothing images to obtain the preliminary try-on results for the current frame of human body images. This improves the issues of clothing misalignment and obvious artifacts in try-on images under complex poses and large deformations. When optimizing the initial try-on results of the current frame's human body image, the process involves acquiring the initial try-on results of the current frame's human body image, the initial try-on results of historical frames' human body images, and the optimized try-on results of historical frames' human body images. Based on the initial try-on results of the current frame's human body image, a key feature map of the current frame is obtained; based on the initial try-on results of historical frames' human body images, a key feature map of historical frames is obtained; based on the similarity between the key feature map of the current frame and the key feature maps of historical frames, a weighted sum of the historical frame's key feature maps is performed to obtain the optimized try-on result of the current frame's human body image. Combined with the feature information of historical frames, the initial try-on result of the current frame is optimized, improving the coherence of the synthesized video. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart illustrating a method for virtually trying on clothing in one embodiment;
[0043] Figure 2 This is a schematic diagram of the STMVN processing architecture in one embodiment;
[0044] Figure 3 This is a schematic diagram of the processing architecture of DAFVN in one embodiment;
[0045] Figure 4 This is a structural block diagram of a device for virtual clothing try-on in one embodiment;
[0046] Figure 5This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0048] The virtual clothing try-on method provided in this application can be executed by a computer device, and the method includes... Figure 1 The steps shown are as follows:
[0049] Step S101: Obtain multiple frames of human images from the video.
[0050] Users can upload videos containing themselves, and computer devices can extract multiple frames of human images from the videos.
[0051] After obtaining multiple frames of human body images, they can be processed by the Space-Time Memory Video Virtual Try-On Network constructed in this application. First, a preliminary try-on result for each frame of the human body image is obtained, followed by an optimized try-on result for each frame. The optimized try-on results for each frame form a synthesized video. The full English name of the Space-Time Memory Video Virtual Try-On Network is: Space-Time Memory Video Virtual Try-On Network, abbreviated as STMVN. The model architecture diagram of STMVN is shown below. Figure 2 As shown.
[0052] Step S102: Based on the clothing image and each frame of the human body image, obtain the preliminary fitting results for each frame of the human body image.
[0053] Specifically, when obtaining the preliminary fitting results of the current frame human body image, multiple scales of current frame human body feature maps and clothing feature maps are obtained based on the current frame human body image and clothing image. The optical flow field and attention map are optimized according to the current frame human body feature map and clothing feature map from small scale to large scale to obtain the target optical flow field and target attention map. The target optical flow field and target attention map are used as guiding information to process the current frame human body image and clothing image to obtain the preliminary fitting results of the current frame human body image.
[0054] The clothing images include clothing that the user wants to try on; this clothing can be referred to as the target clothing.
[0055] Taking the current frame as frame i as an example, the clothing image and the human body image in frame i can be input into STMVN. STMVN includes a human body keypoint detection network and a human body parsing network.
[0056] By processing the i-th frame of human image using a human keypoint detection network, the keypoint information of the i-th frame can be obtained. The keypoint information of the i-th frame is then converted into the i-th frame human pose map, which can be a 3-channel RGB human pose map.
[0057] By processing the i-th frame of the human body image using a human body parsing network, we can obtain the i-th frame's human body parsing information. This i-th frame's parsing image is then converted into an i-th frame's clothing mask image. Specifically, the regions of the human body in the i-th frame's parsing image that need to be covered by clothing are determined, and the pixels of these regions are set to black pixels to obtain the image with the clothing removed, i.e., the i-th frame's clothing mask image. If the target clothing is a top, the regions of the human body that need to be covered by the clothing must include at least the left arm, right arm, and the area from the neck to the lower limbs.
[0058] Next, STMVN can stitch together the i-th frame human pose image and the i-th frame human clothing mask image, and the stitched result can be called the i-th frame human stitched image. STMVN also includes a DeformableAttention Flows Virtual Try-On Network (DAFVN). STMVN can input the i-th frame human stitched image and clothing image into DAFVN to obtain the preliminary try-on result of the i-th frame human image, in order to solve the problem of clothing misalignment and artifacts under complex poses and large deformations.
[0059] After the i-th frame human clothing mask image and the i-th frame human clothing mask image are stitched together to form the i-th frame human mosaic image, they are sent to DAFVN. The i-th frame human mosaic image and clothing image can guide the generation of optical flow field and attention map, and provide a reference for predicting the distortion degree of the target clothing and the pixel prediction of the human torso.
[0060] DAFVN stitches the i-th frame of the human body image with the clothing image to obtain the preliminary fitting results of the i-th frame of the human body image. The specific process is as follows:
[0061] DAFVN can include two pyramid feature extractors, called the source pyramid feature extractor and the reference pyramid feature extractor. Inputting the clothing image into the source pyramid feature extractor yields clothing feature maps at multiple scales. Inputting the i-th frame of the human body mosaic image into the reference pyramid feature extractor yields i-th frame human body feature maps at multiple scales. The optical flow field and attention map are optimized based on the i-th frame human body and clothing feature maps from small to large scales to obtain the target optical flow field and target attention map. Then, using the target optical flow field and target attention map as guiding information, the i-th frame human body image and clothing image are processed to obtain the preliminary fitting results for the i-th frame human body image.
[0062] Using the above method, preliminary fitting results can be obtained for each frame of human body image, and then proceed to step S103.
[0063] Step S103: Optimize the preliminary fitting results of each human body image to obtain the optimized fitting results of each human body image.
[0064] Specifically, when optimizing the preliminary try-on results of the current frame human body image, the preliminary try-on results of the current frame human body image, the preliminary try-on results of historical frame human body images, and the optimized try-on results of historical frame human body images are obtained; the current frame key feature map is obtained based on the preliminary try-on results of the current frame human body image; the historical frame key feature map is obtained based on the preliminary try-on results of historical frame human body images; the historical frame value feature map is obtained based on the optimized try-on results of historical frame human body images; and the historical frame value feature map is weighted and summed based on the similarity between the current frame key feature map and the historical frame key feature map to obtain the optimized try-on result of the current frame human body image.
[0065] Taking the current frame as frame i as an example, we can obtain the preliminary fitting results of the human body image in frame i, the preliminary fitting results of human body images in several historical frames before frame i, and the optimized fitting results of human body images in several historical frames before frame i.
[0066] STMVN can also include a Space-Time Correspondence Try-On Network (STCN) based on a memory model.
[0067] DAFVN performs well in solving the problems of misalignment and artifacts under complex poses and large deformations. However, DAFVN does not utilize the spatiotemporal information in the video, resulting in insufficient video continuity. Therefore, this application uses STCN to correct the mispredicted pixels in the preliminary fitting results of the current frame by utilizing information from historical frames, reducing the misalignment of the predicted clothing and human body between frames, thereby making the synthesized video more coherent. Moreover, the STCN of this application can optimize the preliminary fitting results more concisely and effectively.
[0068] STCN comprises several key encoders, which can be identical. The preliminary fitting result of the i-th frame of human images is input into one of the key encoders, and the result is called the i-th frame key feature map. The preliminary fitting results of several historical frames of human images are input into each key encoder, respectively, to obtain the key feature map for each historical frame. Each historical frame key feature map can be stored in memory.
[0069] STCN can also include several value encoders, which can be identical. The optimized fitting results of several historical frame human images are input into each value encoder to obtain a feature map for each historical frame. Each historical frame feature map can be stored in memory.
[0070] Next, STCN can calculate the similarity between the current frame key feature map and each historical frame key feature map, determine the weight of the corresponding historical frame value feature map based on the similarity, and then perform a weighted summation on each historical frame value feature map according to the weight. The weighted summation result is called the weighted value feature map of the i-th frame.
[0071] STCN may also include a decoder, into which the weighted feature map of the i-th frame is input, and the result is used as the optimized fitting result of the i-th frame human body image.
[0072] In the aforementioned method for virtual clothing try-on, multiple frames of human body images are acquired from a video; based on the clothing images and each frame of human body images, a preliminary try-on result for each frame of human body images is obtained; the preliminary try-on result for each frame of human body images is optimized to obtain an optimized try-on result for each frame of human body images; when obtaining the preliminary try-on result for the current frame of human body images, multiple scales of current frame human body feature maps and clothing feature maps are obtained based on the current frame human body images and clothing images; the optical flow field and attention map are optimized according to the current frame human body feature maps and clothing feature maps from small scale to large scale to obtain a target optical flow field and target attention map; using the target optical flow field and target attention map as guiding information, the current frame human body images and clothing images are processed to obtain the preliminary try-on result for the current frame human body images, enabling complex poses and large deformations to be effectively simulated. The issues of clothing misalignment and obvious artifacts in the try-on images under certain conditions are improved. When optimizing the preliminary try-on results of the current frame human body image, the preliminary try-on results of the current frame human body image, the preliminary try-on results of the historical frame human body images, and the optimized try-on results of the historical frame human body images are obtained. The key feature map of the current frame human body image is obtained based on the preliminary try-on results of the current frame human body image; the key feature map of the historical frame human body image is obtained based on the preliminary try-on results of the historical frame human body images; the value feature map of the historical frame human body image is obtained based on the optimized try-on results of the historical frame human body images; based on the similarity between the key feature map of the current frame and the key feature map of the historical frame, the value feature map of the historical frame is weighted and summed to obtain the optimized try-on result of the current frame human body image. Combined with the feature information of the historical frames, the initial try-on result of the current frame is optimized, improving the coherence of the synthesized video.
[0073] In one embodiment, the optical flow field and attention map are optimized based on the current frame's human feature map and clothing feature map from small scale to large scale to obtain the target optical flow field and target attention map, including:
[0074] Based on the current frame human feature map from small scale to large scale, obtain the self-optical flow field and self-attention map corresponding to each scale, and take the self-optical flow field and self-attention map corresponding to the last scale as the target self-optical flow field and target self-attention map.
[0075] Specifically, when obtaining the self-optical flow field and self-attention map corresponding to the current scale, if the current scale is not the minimum scale, the current frame human feature map of the current scale is obtained and fused with the current frame human feature map of the previous scale to obtain the current frame human fusion feature map. Based on the self-optical flow field and self-attention map corresponding to the previous scale, the current frame human fusion feature map is transformed to obtain the current frame human transformation feature map. Based on the current frame human transformation feature map, the self-optical flow field and self-attention map corresponding to the current scale are obtained.
[0076] Based on the clothing feature maps from small scale to large scale, the cross optical flow field and cross attention map corresponding to each scale are obtained. The cross optical flow field and cross attention map corresponding to the last scale are taken as the target cross optical flow field and target cross attention map.
[0077] Specifically, when obtaining the cross optical flow field and cross attention map corresponding to the current scale, if the current scale is not the minimum scale, the clothing feature map of the current scale is obtained and fused with the clothing feature map of the previous scale to obtain the clothing fusion feature map. Based on the cross optical flow field and cross attention map corresponding to the previous scale, the clothing fusion feature map is transformed to obtain the clothing transformation feature map. Based on the clothing transformation feature map, the cross optical flow field and cross attention map corresponding to the current scale are obtained; the previous scale is smaller than the current scale.
[0078] Specifically, the method provided in this application also includes:
[0079] When obtaining the self-optical flow field and self-attention map corresponding to the current scale, if the current scale is not the minimum scale, the self-optical flow field and self-attention map corresponding to the current scale are obtained based on the human body feature map of the current frame at the current scale.
[0080] Specifically, the method provided in this application also includes:
[0081] When obtaining the cross-optical flow field and cross-attention map corresponding to the current scale, if the current scale is the smallest scale, the cross-optical flow field and cross-attention map corresponding to the current scale are obtained based on the clothing feature map of the current scale.
[0082] Combination Figure 3 Let's take the current frame as the i-th frame and the number of scales as an example.
[0083] Figure 3 This is a schematic diagram of the DAFVN processing framework; see reference. Figure 3 DAFVN can sample multiple streams simultaneously and extract feature-level and pixel-level information from different semantic regions through a deformation attention mechanism. It can simultaneously perform clothing distortion and human body synthesis. The preliminary fitting results generated by DAFVN have high quality, which is beneficial for the next stage of optimization.
[0084] DAFVN includes a reference pyramid feature extractor. You can input the i-th frame human pose image and the i-th frame human clothing mask image into the reference pyramid feature extractor to obtain large-scale, medium-scale, and small-scale i-th frame human feature images.
[0085] DAFVN includes a Deformable Attention Flows Network (DAFN). A small-scale human feature map of the i-th frame can be input into the DAFN to obtain the corresponding optical flow field at that small scale. and self-attention map Next, the i-th frame human feature map at the medium scale is fused with the i-th frame human feature map at the small scale to obtain the i-th frame human feature map fused at the medium and small scales.
[0086] DAFVN also includes a Deformable Attention Warp (DAWarp). This unit can fuse the human body feature map of the i-th frame at medium to small scales, and the self-optical flow field. and self-attention map Input into DAWarp to enable self-optical flow field in DAWarp and self-attention map The i-th frame human body fusion feature map is transformed from mesoscale to small scale, and the result is called the i-th frame human body transformed feature map from mesoscale to small scale.
[0087] Then, the i-th frame human body transformation feature map from mesoscale to small scale is input into DAFN to obtain the self-optical flow field corresponding to the mesoscale. and self-attention map Next, the large-scale i-th frame human feature map is fused with the medium-scale i-th frame human feature map to obtain the large-scale-medium-scale i-th frame human feature map.
[0088] The large-scale-to-medium-scale human body fusion feature map of the i-th frame and the self-light flow field are combined. and self-attention map Input into DAWarp to enable self-optical flow field in DAWarp and self-attention map The i-th frame human body fusion feature map is transformed from large-scale to medium-scale, and the result is called the i-th frame human body transformed feature map from large-scale to medium-scale.
[0089] Then, the large-scale to mesoscale human body transformation feature map of the i-th frame is input into the DAFN to obtain the self-optical flow field corresponding to the mesoscale. and self-attention map Self-optical flow field Optical flow field that is one type of target optical flow field; self-attention map Attention maps are one type of attention maps.
[0090] DAFVN also includes a source pyramid feature extractor, which can take clothing images as input and obtain large-scale, medium-scale, and small-scale clothing feature maps.
[0091] Small-scale clothing feature maps can be input into DAFN to obtain the corresponding cross-optical flow field at the small scale. and cross attention map Next, the mesoscale clothing feature map is fused with the small-scale clothing feature map to obtain a mesoscale-small-scale fused clothing feature map.
[0092] It can fuse mesoscale and small-scale clothing feature maps and cross-optical flow fields. and cross attention map Input into DAWarp to achieve cross-optical flow in DAWarp and cross attention map The result of converting the mesoscale to smallscale clothing fusion feature map is called the mesoscale to smallscale clothing conversion feature map.
[0093] Then, the mesoscale-to-small-scale clothing transformation feature map is input into DAFN to obtain the cross-optical flow field corresponding to the mesoscale. and cross attention map Next, the large-scale clothing feature map is fused with the medium-scale clothing feature map to obtain a large-scale-medium-scale clothing fused feature map.
[0094] Integrating large-scale and meso-scale clothing feature maps and cross-optical flow fields and cross attention map Input into DAWarp to achieve cross-optical flow in DAWarp and cross attention map The result of converting the large-scale to medium-scale clothing fusion feature map is called the large-scale to medium-scale clothing conversion feature map.
[0095] Then, the large-scale to mesoscale clothing transformation feature map is input into DAFN to obtain the cross-optical flow field corresponding to the mesoscale. and cross attention map Cross-optical flow field Optical flow field that is one of the target optical flow fields; cross-attention map Attention maps are a type of target attention map.
[0096] In one embodiment, the target optical flow field and target attention map are used as guiding information to process the current frame's human body image and clothing image to obtain preliminary fitting results for the current frame's human body image, including:
[0097] The current frame's human body image and clothing image are shallow-encoded separately to obtain shallow-encoded human body feature maps and shallow-encoded clothing feature maps. Based on the target's self-optical flow field, target self-attention map, and the current frame's shallow-encoded human body feature map, deformable attention distortion is applied to obtain the current frame's target human body and clothing feature map. Based on the target's cross-optical flow field, target cross-attention map, and shallow-encoded clothing feature map, deformable attention distortion is applied to obtain the target clothing feature map. The current frame's target human body and clothing feature maps and the target clothing feature map are merged to obtain the current frame's merged feature map. The current frame's merged feature map is shallow-decoded to obtain the preliminary fitting results of the current frame's human body image.
[0098] Taking the current frame as the i-th frame, the target optical flow field includes the self-optical flow field. and cross optical flow field Furthermore, the target attention map includes the self-attention map. and cross attention map For example.
[0099] DAFVN includes a shallow encoder and a shallow decoder.
[0100] By inputting the human clothing mask image of the i-th frame into the shallow encoder, the shallow encoded human feature image of the i-th frame can be obtained; the shallow encoded human feature image of the i-th frame and the self-optical flow field and self-attention map Input DAWarp to perform deformable attention distortion, and the result is called the target human clothing feature map of the i-th frame.
[0101] Inputting a clothing image into a shallow encoder yields a shallow-coded clothing feature map. This shallow-coded clothing feature map, along with a cross-optical flow field, is then processed. and cross attention map Input DAWarp to perform deformable attention distortion, and the result is called the target clothing feature map.
[0102] The target human clothing feature map and the target clothing feature map of the i-th frame are merged to obtain the merged feature map of the i-th frame; the merged feature map of the i-th frame is input into the shallow decoder, and the result is used as the preliminary fitting result of the i-th frame human image.
[0103] In one embodiment, based on the similarity between the current frame key feature map and the historical frame key feature maps, a weighted sum of the historical frame value feature maps is performed to obtain the optimized fitting result of the current frame human body image, including:
[0104] Based on the relative similarity between the current frame key feature map and each historical frame key feature map, the relative weights assigned to each historical frame key feature map are determined, resulting in the weights of each historical frame key feature map; the greater the similarity, the greater the weight; based on the weights of each historical frame key feature map, the historical frame key feature maps are weighted and summed to obtain the optimized fitting result of the current frame human body image.
[0105] Taking the current frame as the i-th frame as an example.
[0106] STCN includes several identical key encoders; STCN may also include several identical value encoders; STCN may also include decoders.
[0107] The preliminary fitting results of the i-th frame human image are input into a one-key encoder, and the result is called the i-th frame key feature map; the dimension of the i-th frame key feature map is... ; It is the number of channels in the key feature map of the i-th frame. and Here, represents the height and width of the key feature map in frame i, respectively. The preliminary fitting results from several historical frames of human images are input into each key encoder to obtain the key feature map for each historical frame; the dimension of each historical frame key feature map is... ; It is the number of channels in the historical frame key feature map. and These represent the height and width of the historical frame key feature map, respectively. The optimized try-on results of several historical frame human images are input into each value encoder to obtain the value feature map for each historical frame; the dimension of each historical frame feature map is... ; It is the number of channels in the historical frame feature map. and These represent the height and width of the historical frame key feature map, respectively.
[0108] STCN can calculate the similarity between the key feature map of frame i and the key feature maps of each historical frame, obtaining several similarity scores. Based on the relative magnitudes of these similarity scores, the relative weights assigned to the corresponding historical frame value feature maps are determined. For example, if the similarity between the key feature map of frame i and the key feature map of frame (i-1) is greater than the similarity between the key feature map of frame i and the key feature map of frame (i-2), then the weight assigned to the value feature map of frame (i-1) is greater than the weight assigned to the value feature map of frame (i-2). Thus, the weights of each historical frame value feature map can be obtained. Based on these weights, the value feature maps of each historical frame are weighted and summed. The weighted sum is called the weighted feature map of frame i. The weighted feature map of frame i is input into the decoder, and the result is used as the optimized fitting result for the human body image of frame i.
[0109] To construct the aforementioned STMVN, training can be divided into three stages. The first stage trains the DAFVN, the second stage trains the initial STCN using a static method, and the third stage trains the final STCN using a dynamic method. A training set can be constructed, consisting of 661 training videos, totaling 159,170 frames.
[0110] The training sets used in each stage of training are described below.
[0111] Specifically, the first phase of training aims to obtain a DAFVN with the ability to predict preliminary fitting results. For example, all images in the training set can be used as the training set for the first phase of training, with a batch size of 4, a training epoch of 20, an initial learning rate of 5×10^(-5), which is reduced to 0.2 times the original value every 3 epochs, and the Adam optimizer can be used.
[0112] The second phase of static training aims to obtain an initial STCN with basic image generation capabilities. For example, 661 training videos from the training set can be acquired, with 30 frames sampled from each video, resulting in 19,830 images. At the beginning of each epoch, sampling is performed again. For each of the 19,830 sampled images, each image is copied five times, resulting in 19,830 five-frame videos, which are used as the training set for static training. The batch size for the second phase of training is set to 1, the epoch is 90, and the initial learning rate is 2 × 10^(-4), which is reduced to 0.2 times the original rate every 30 epochs.
[0113] The third stage of dynamic training aims to obtain the final STCN capable of generating realistic costume change videos. For example, 661 training videos can be acquired from the training set, with each video sampling 5 frames. The interval between adjacent sampled frames in a training video does not exceed k frames, resulting in 661 five-frame videos per epoch. The sampling interval k changes with the number of training epochs, initially set to 5, increasing by 5 every 600 epochs, and then resetting to 5 at the last 600 epochs. At the beginning of each epoch, the 661 training videos in the training set are resampled, ensuring that the 661 five-frame videos trained in each epoch are newly generated and not duplicated from previous ones. The batch size for the third stage of training is 1, the number of epochs is 3000, and the initial learning rate is 2×10^(-4), which is reduced to 0.2 times the original rate every 1000 epochs.
[0114] The following describes the loss functions used in each stage of training.
[0115] In the first stage of training DAFVN, the loss function used can be obtained by weighting three loss functions: L1 loss function, perceptual loss function, and style loss function.
[0116] The L1 loss function is obtained by calculating the L1 distance between the pre-test image and the actual test image. The L1 loss function helps to optimize the details of the generated image. The calculation formula is shown in equation (1). Indicates the L1 distance. This indicates a pre-test wear pattern. This shows a real try-on photo.
[0117] (1)
[0118] The perceptual loss function is obtained by calculating the L1 distance between five feature maps of different scales extracted from the pre-test image and the real test image. The VGG-19 network can be used to extract five feature maps of different scales. The perceptual loss function helps to improve the structural realism of the generated image, and its calculation formula is shown in Equation (2).
[0119] This represents the j-th scale feature map extracted by the VGG-19 network.
[0120] (2)
[0121] The style loss function maintains style consistency between the pre-test image and the real test image by calculating the L1 distance between the Gram feature matrices of feature maps at different scales extracted from the pre-test image and the real test image. The Gram matrix can capture the texture and color information of the image, so using the style loss function can make the generated image more realistic and natural. Specifically, the formula for calculating the style loss function is shown in equation (3). Let represent the Gram matrix of the feature map at the j-th scale.
[0122] (3)
[0123] The loss function used when training DAFVN is shown in equation (4). , and These are the weights for the L1 loss function, the perceptual loss function, and the style loss function, respectively.
[0124] (4)
[0125] During the second and third training phases, the goal of the loss function is to correct mispredicted pixels and improve the coherence of the synthesized video. Therefore, this application introduces a stream consistency loss function. This is used to estimate the coherence loss between frames. The optical flow field can be divided into the forward optical flow field. and backward optical flow field The forward optical flow field represents the optical flow from the previous frame to the next frame, and the backward optical flow field represents the optical flow from the next frame to the previous frame. This can be expressed using equation (5). and Alignment The distortion operation is performed using bilinear interpolation. Similarly, it can be calculated using equation (6). .
[0126] (5)
[0127] (6)
[0128] If the frames are continuous, then, and They are equal in size but opposite in direction, and similarly... and Since they are equal in magnitude and opposite in direction, as shown in equation (7), the flow consistency loss function can be defined as:
[0129] (7)
[0130] in This indicates whether pixel p belongs to a human body area that needs to be covered by clothing; if so, then... The value is 1 if the value is less than 1, and 0 otherwise. To save computation, the flow consistency loss function in this application can only calculate the human body region to be covered by clothing. N represents the total number of pixels within the human body region to be covered by clothing. The better the frame coherence, the better. The closer the value is to 0.
[0131] In addition to the stream consistency loss function, the loss functions in the second and third stages also need to evaluate the generation quality of the video images themselves, applying the L1 loss function of equation (1). The perceptual loss function of equation (2) and the L2 loss function as shown in equation (8). .
[0132] (8)
[0133] To enhance the realism of the synthesized videos, this application also uses a multi-scale discriminator based on generative adversarial networks to calculate the generative adversarial loss function. The multi-scale discriminator in this application uses three discriminators with the same network structure but different discrimination scales, denoted as D1, D2, and D3. D1 accepts the generated image at the original scale, D2 accepts the generated image downsampled to half the scale, and D3 accepts the generated image downsampled to a quarter scale. D3 has the largest receptive field, enabling it to perceive the structure of the generated image and improve the global consistency of the synthesis. D1 has the finest scale, which can guide the generator to generate finer details. The calculation formula is shown in equation (9). Wherein This indicates a real video. This indicates the initial virtual try-on video input from STCN.
[0134] (9)
[0135] The loss function used in the second and third stages of training is shown in (10). , , , and As weight.
[0136] (10)
[0137] After the three stages of training are completed, the DAFVN obtained from the first stage of training and the final STCN obtained from the third stage of training can be obtained to form the DAFVN, which can be used for inference scenarios.
[0138] In inference scenarios, when generating virtual try-on videos of clothing using the DAFVN provided in this application, the text and patterns of the clothing can be well preserved. This is because the DAFVN uses extracted multi-level feature maps to perform cascaded updates and optimizations on two optical flow fields and attention maps, which can preserve both high-level structural information and low-level detail information of the clothing. The solution provided in this application can well preserve the texture and material of the clothing, partly because DAFVN preserves low-level detail information, and partly because STCN has the ability to correct mispredicted pixels, thus recovering the texture and material features of the clothing. The virtual try-on videos of clothing generated by the solution provided in this application have a significantly improved quality compared to traditional methods.
[0139] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0140] Based on the same inventive concept, this application also provides a virtual clothing try-on apparatus for implementing the virtual clothing try-on method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations of one or more virtual clothing try-on apparatus embodiments provided below can be found in the limitations of the virtual clothing try-on method described above, and will not be repeated here.
[0141] In one embodiment, such as Figure 4 As shown, a device for virtual clothing try-on is provided, comprising:
[0142] The human body image acquisition module 401 is used to acquire multiple frames of human body images from the video.
[0143] The initial fitting processing module 402 is used to obtain the preliminary fitting results of each frame of human body image based on the clothing image and each frame of human body image.
[0144] The optimized try-on processing module 403 is used to optimize the preliminary try-on results of each frame of human body image to obtain the optimized try-on results of each frame of human body image.
[0145] When obtaining the preliminary fitting results of the current frame human body image, the initial fitting processing module 402 is further configured to: obtain multiple scales of current frame human body feature maps and clothing feature maps based on the current frame human body image and clothing image; optimize the optical flow field and attention map according to the current frame human body feature maps and clothing feature maps from small scale to large scale to obtain the target optical flow field and target attention map; and process the current frame human body image and clothing image with the target optical flow field and target attention map as guiding information to obtain the preliminary fitting results of the current frame human body image.
[0146] Specifically, when optimizing the preliminary fitting results of the current frame human body image, the optimized fitting processing module 403 is further configured to: acquire the preliminary fitting results of the current frame human body image, the preliminary fitting results of historical frame human body images, and the optimized fitting results of historical frame human body images; obtain the current frame key feature map based on the preliminary fitting results of the current frame human body image; obtain the historical frame key feature map based on the preliminary fitting results of historical frame human body images; obtain the historical frame value feature map based on the optimized fitting results of historical frame human body images; and perform a weighted summation of the historical frame value feature map based on the similarity between the current frame key feature map and the historical frame key feature map to obtain the optimized fitting result of the current frame human body image.
[0147] In one embodiment, the initial fitting processing module 402 is further configured to:
[0148] Based on the current frame human feature map from small scale to large scale, obtain the self-optical flow field and self-attention map corresponding to each scale, and take the self-optical flow field and self-attention map corresponding to the last scale as the target self-optical flow field and target self-attention map; based on the clothing feature map from small scale to large scale, obtain the cross optical flow field and cross attention map corresponding to each scale, and take the cross optical flow field and cross attention map corresponding to the last scale as the target cross optical flow field and target cross attention map;
[0149] Specifically, when obtaining the self-optical flow field and self-attention map corresponding to the current scale, if the current scale is not the minimum scale, the current frame human feature map of the current scale is obtained and fused with the current frame human feature map of the previous scale to obtain the current frame human fusion feature map. Based on the self-optical flow field and self-attention map corresponding to the previous scale, the current frame human fusion feature map is transformed to obtain the current frame human transformation feature map. Based on the current frame human transformation feature map, the self-optical flow field and self-attention map corresponding to the current scale are obtained.
[0150] Specifically, when obtaining the cross optical flow field and cross attention map corresponding to the current scale, if the current scale is not the minimum scale, the clothing feature map of the current scale is obtained and fused with the clothing feature map of the previous scale to obtain the clothing fusion feature map. Based on the cross optical flow field and cross attention map corresponding to the previous scale, the clothing fusion feature map is transformed to obtain the clothing transformation feature map. Based on the clothing transformation feature map, the cross optical flow field and cross attention map corresponding to the current scale are obtained; the previous scale is smaller than the current scale.
[0151] In one embodiment, the initial fitting processing module 402 is further configured to:
[0152] When obtaining the self-optical flow field and self-attention map corresponding to the current scale, if the current scale is not the minimum scale, the self-optical flow field and self-attention map corresponding to the current scale are obtained based on the human body feature map of the current frame at the current scale.
[0153] In one embodiment, the initial fitting processing module 402 is further configured to:
[0154] When obtaining the cross-optical flow field and cross-attention map corresponding to the current scale, if the current scale is the smallest scale, the cross-optical flow field and cross-attention map corresponding to the current scale are obtained based on the clothing feature map of the current scale.
[0155] In one embodiment, the initial fitting processing module 402 is further configured to:
[0156] The current frame's human body image and clothing image are shallow-encoded separately to obtain shallow-encoded human body feature maps and shallow-encoded clothing feature maps. Based on the target's self-optical flow field, target self-attention map, and the current frame's shallow-encoded human body feature map, deformable attention distortion is performed to obtain the current frame's target human body and clothing feature map. Based on the target's cross-optical flow field, target cross-attention map, and the shallow-encoded clothing feature map, deformable attention distortion is performed to obtain the target clothing feature map. The current frame's target human body and clothing feature maps and target clothing feature maps are merged to obtain the current frame's merged feature map. The current frame's merged feature map is shallow-decoded to obtain the preliminary fitting results of the current frame's human body image.
[0157] In one embodiment, the optimized fitting process module 403 is further configured to:
[0158] Based on the relative similarity between the current frame key feature map and each historical frame key feature map, the relative weights assigned to each historical frame key feature map are determined, resulting in the weights of each historical frame key feature map; the greater the similarity, the greater the weight; based on the weights of each historical frame key feature map, the historical frame key feature maps are weighted and summed to obtain the optimized fitting result of the current frame human body image.
[0159] The various modules in the aforementioned virtual clothing try-on device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0160] In one exemplary embodiment, a computer device is provided, the internal structure of which can be as shown in the figure. Figure 5As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores the data involved in the aforementioned methods. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for virtually trying on clothing.
[0161] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0162] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in the various method embodiments described above.
[0163] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the various method embodiments described above.
[0164] In one embodiment, a computer program product is provided having a computer program stored thereon, the computer program being executed by a processor of the steps described in the various method embodiments above.
[0165] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0166] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0167] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0168] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for virtual try-on of clothing, characterized in that, The method includes: Extract multiple frames of human body images from a video; Based on the clothing images and each frame of the human body image, preliminary fitting results are obtained for each frame of the human body image. The initial fitting results of each human body image are optimized to obtain the optimized fitting results of each human body image. Specifically, when obtaining the preliminary fitting results of the current frame human body image, multiple scales of current frame human body feature maps and clothing feature maps are obtained based on the current frame human body image and clothing image. The optical flow field and attention map are optimized according to the current frame human body feature maps and clothing feature maps from small scale to large scale to obtain the target optical flow field and target attention map. The target optical flow field and target attention map are used as guiding information to process the current frame human body image and clothing image to obtain the preliminary fitting results of the current frame human body image. Specifically, when optimizing the preliminary try-on results of the current frame human body image, the preliminary try-on results of the current frame human body image, the preliminary try-on results of historical frame human body images, and the optimized try-on results of historical frame human body images are obtained; the current frame key feature map is obtained based on the preliminary try-on results of the current frame human body image; the historical frame key feature map is obtained based on the preliminary try-on results of historical frame human body images; the historical frame value feature map is obtained based on the optimized try-on results of historical frame human body images; and the historical frame value feature map is weighted and summed based on the similarity between the current frame key feature map and the historical frame key feature map to obtain the optimized try-on result of the current frame human body image. The step of optimizing the optical flow field and attention map based on the current frame's human feature map and clothing feature map from small to large scales to obtain the target optical flow field and target attention map includes: Based on the current frame human feature map from small scale to large scale, obtain the self-optical flow field and self-attention map corresponding to each scale, and take the self-optical flow field and self-attention map corresponding to the last scale as the target self-optical flow field and target self-attention map. Specifically, when obtaining the self-optical flow field and self-attention map corresponding to the current scale, if the current scale is not the minimum scale, the current frame human feature map of the current scale is obtained and fused with the current frame human feature map of the previous scale to obtain the current frame human fusion feature map. Based on the self-optical flow field and self-attention map corresponding to the previous scale, the current frame human fusion feature map is transformed to obtain the current frame human transformation feature map. Based on the current frame human transformation feature map, the self-optical flow field and self-attention map corresponding to the current scale are obtained. Based on the clothing feature maps from small scale to large scale, the cross optical flow field and cross attention map corresponding to each scale are obtained. The cross optical flow field and cross attention map corresponding to the last scale are taken as the target cross optical flow field and target cross attention map. Specifically, when obtaining the cross optical flow field and cross attention map corresponding to the current scale, if the current scale is not the minimum scale, the clothing feature map of the current scale is obtained and fused with the clothing feature map of the previous scale to obtain the clothing fusion feature map. Based on the cross optical flow field and cross attention map corresponding to the previous scale, the clothing fusion feature map is transformed to obtain the clothing transformation feature map. Based on the clothing transformation feature map, the cross optical flow field and cross attention map corresponding to the current scale are obtained; the previous scale is smaller than the current scale.
2. The method according to claim 1, characterized in that, The method further includes: When obtaining the self-optical flow field and self-attention map corresponding to the current scale, if the current scale is not the minimum scale, the self-optical flow field and self-attention map corresponding to the current scale are obtained based on the human body feature map of the current frame at the current scale.
3. The method according to claim 1, characterized in that, The method further includes: When obtaining the cross-optical flow field and cross-attention map corresponding to the current scale, if the current scale is the smallest scale, the cross-optical flow field and cross-attention map corresponding to the current scale are obtained based on the clothing feature map of the current scale.
4. The method according to claim 1, characterized in that, Using the target optical flow field and target attention map as guiding information, the current frame human body image and clothing image are processed to obtain preliminary fitting results for the current frame human body image, including: The current frame's human body image and clothing image are shallow encoded separately to obtain the current frame's shallow encoded human body feature map and shallow encoded clothing feature map; Based on the target self-optical flow field, the target self-attention map and the current frame shallow-coded human body feature map, deformable attention distortion is performed to obtain the current frame target human body clothing feature map; Based on the target cross optical flow field, the target cross attention map and the shallow-coded clothing feature map, deformable attention distortion is performed to obtain the target clothing feature map; Merge the target human body clothing feature map and the target clothing feature map in the current frame to obtain the merged feature map of the current frame; The current frame's merged feature map is shallow-decoded to obtain preliminary fitting results for the current frame's human body image.
5. The method according to any one of claims 1 to 4, characterized in that, Based on the similarity between the current frame key feature map and the historical frame key feature maps, a weighted sum is performed on the historical frame value feature maps to obtain the optimized fitting result of the current frame human image, including: Based on the relative similarity between the current frame key feature map and each historical frame key feature map, the relative weights assigned to each historical frame key feature map are determined, thus obtaining the weights of each historical frame key feature map; the greater the similarity, the greater the weight. Based on the weights of the feature maps of each historical frame, the feature maps of each historical frame are weighted and summed to obtain the optimized fitting result of the human body image in the current frame.
6. A device for virtual clothing try-on, characterized in that, The device includes: The human body image acquisition module is used to acquire multiple frames of human body images from videos; The initial fitting processing module is used to obtain preliminary fitting results for each frame of the human body image based on the clothing image and each frame of the human body image. The optimized try-on processing module is used to optimize the preliminary try-on results of each frame of human body image to obtain the optimized try-on results of each frame of human body image. When obtaining the preliminary fitting results of the current frame human body image, the initial fitting processing module is further configured to: obtain multiple scales of current frame human body feature maps and clothing feature maps based on the current frame human body image and clothing image; optimize the optical flow field and attention map according to the current frame human body feature maps and clothing feature maps from small scale to large scale to obtain the target optical flow field and target attention map; and use the target optical flow field and target attention map as guiding information to process the current frame human body image and clothing image to obtain the preliminary fitting results of the current frame human body image. Specifically, when optimizing the preliminary fitting results of the current frame human body image, the optimized fitting processing module is further configured to: obtain the preliminary fitting results of the current frame human body image, the preliminary fitting results of historical frame human body images, and the optimized fitting results of historical frame human body images; obtain the current frame key feature map based on the preliminary fitting results of the current frame human body image; obtain the historical frame key feature map based on the preliminary fitting results of historical frame human body images; obtain the historical frame value feature map based on the optimized fitting results of historical frame human body images; and perform a weighted summation of the historical frame value feature map based on the similarity between the current frame key feature map and the historical frame key feature map to obtain the optimized fitting result of the current frame human body image. The initial fitting processing module is further configured to: Based on the current frame human feature map from small scale to large scale, obtain the self-optical flow field and self-attention map corresponding to each scale, and take the self-optical flow field and self-attention map corresponding to the last scale as the target self-optical flow field and target self-attention map. Specifically, when obtaining the self-optical flow field and self-attention map corresponding to the current scale, if the current scale is not the minimum scale, the current frame human feature map of the current scale is obtained and fused with the current frame human feature map of the previous scale to obtain the current frame human fusion feature map. Based on the self-optical flow field and self-attention map corresponding to the previous scale, the current frame human fusion feature map is transformed to obtain the current frame human transformation feature map. Based on the current frame human transformation feature map, the self-optical flow field and self-attention map corresponding to the current scale are obtained. Based on the clothing feature maps from small scale to large scale, the cross optical flow field and cross attention map corresponding to each scale are obtained. The cross optical flow field and cross attention map corresponding to the last scale are taken as the target cross optical flow field and target cross attention map. Specifically, when obtaining the cross optical flow field and cross attention map corresponding to the current scale, if the current scale is not the minimum scale, the clothing feature map of the current scale is obtained and fused with the clothing feature map of the previous scale to obtain the clothing fusion feature map. Based on the cross optical flow field and cross attention map corresponding to the previous scale, the clothing fusion feature map is transformed to obtain the clothing transformation feature map. Based on the clothing transformation feature map, the cross optical flow field and cross attention map corresponding to the current scale are obtained; the previous scale is smaller than the current scale.
7. The apparatus according to claim 6, characterized in that, The initial fitting processing module is also used for: When obtaining the self-optical flow field and self-attention map corresponding to the current scale, if the current scale is not the minimum scale, the self-optical flow field and self-attention map corresponding to the current scale are obtained based on the human body feature map of the current frame at the current scale.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Virtual fitting video generation method and device, equipment and medium
CN114638754A
Appearance flow estimation method fusing double attention mechanism for garment deformation
CN118229838A