End-to-end analysis-free virtual reloading method, system and device based on double optical streams

By introducing dual-optical flow mechanism and task adaptive processing in the virtual dressing method, end-to-end parsing-free virtual dressing is achieved, solving the shortcomings of the existing methods in terms of training time and effect, and significantly improving the quality and efficiency of the dressing are carried out.

CN120088435APending Publication Date: 2025-06-03GUANGDONG OCEAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510121262.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing virtual dressing methods, based on analysis and unresolved methods, have their own shortcomings, and cannot effectively meet the increasingly high virtual dressing requirements, especially in terms of training time and dressing effect.

Method used

The end-to-end non-analytical virtual dressing method based on dual-optical flow is adopted. By obtaining the clothing pictures to be worn and the mannequin diagram, feature map encoding and task adaptive processing are performed. Combining the dual-optical flow decoder and feature map distortion module, the splicing and distortion of the clothing and human body feature maps is realized to generate virtual dressing results.

Benefits of technology

The training time of the virtual dressing method without analytics is greatly reduced, and the dressing effect is better than the original dressing method through the dual-optical flow mechanism, solving the problems of missing details and large calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088435A_ABST
    Figure CN120088435A_ABST
Patent Text Reader

Abstract

The invention discloses an end-to-end analysis-free virtual reloading method, system and device based on double optical streams, and belongs to the technical field of virtual reloading. The method comprises the following steps: S1, obtaining and coding a to-be-worn clothes graph and a human body model graph to obtain corresponding feature graphs; s2, performing task adaptive processing on the clothing feature maps and the human body feature maps to obtain two groups of clothing feature maps and two groups of human body feature maps; s3, splicing the first group of clothing feature maps and the first group of human body feature maps, and performing dual-optical flow decoding processing to obtain an optical flow field; s4, distortion of the garment feature maps is carried out through the optical flow field and the second set of garment feature maps, and distortion data are obtained; and S5, carrying out reloading decoding by using the distortion data and the second group of human body feature maps to obtain a virtual reloading result. According to the method, the training time of non-analytic virtual reloading can be greatly shortened, and a reloading effect better than that of an original non-analytic virtual reloading method is obtained under a double-optical-flow mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of virtual dressing, and particularly relates to an end-to-end virtual dressing method, system and device based on dual optical flow. Background Art

[0002] In order to reduce the return cost of online retailers and enable shoppers to have the same experience as offline shopping, image virtual dressing technology has been favored by many researchers in recent years.

[0003] Currently, the mainstream virtual dressing methods are generally divided into two types: analytical virtual dressing methods and non-analytical virtual dressing methods.

[0004] Figure 1 Shows that the virtual dressing method based on the parsing graph usually inputs the relevant information of the human body (segmentation images of each part, pose point maps, depth information maps), the model image, and the image to be dressed into the clothing distortion model to generate a distorted clothing image or distorted clothing features before generating. Then, the distorted result is input into the dressing network together with the relevant information of the human body and the model image to complete the final dressing.

[0005] Figure 2 Shows the operation process of the non-analytical virtual dressing method. This method uses the knowledge distillation technology based on an analytical virtual dressing model to obtain a student model, which only needs the model image and the image to be dressed to complete the dressing, and the dressing process is more concise.

[0006] Although the training process of the analytical virtual dressing method is simpler than that of the non-analytical method, it needs to introduce many additional human body information, such as segmentation maps of each part, pose maps, depth information maps, and these relevant information need to be obtained using additional deep learning models. Eventually, the user experience of using electronic dressing products is greatly reduced due to the long dressing time. Moreover, due to the extreme dependence of the analytical virtual dressing method on human body information, once the parsing information is inaccurate, the virtual dressing result will be difficult to satisfy (see Figure 3 the left side of the black dotted line).

[0007] Although the non-analytical virtual dressing method avoids the phenomenon of damaged dressing results caused by bad parsing results (such as Figure 3 the right side of the black dotted line), and improves the user experience of the product. However, it increases the overall training process of the virtual dressing network because of using the knowledge distillation method, that is, on the basis of the analytical virtual dressing method, a student model that only needs to input the image to be dressed and the model image to complete the dressing needs to be obtained using knowledge distillation.

[0008] It can be seen that both the parsing-based virtual clothing change method and the non-parsing virtual clothing change method have their own deficiencies and cannot well meet the increasingly high requirements for virtual clothing change. Summary of the Invention

[0009] The present invention aims to solve at least one of the technical problems in the above related technologies to some extent.

[0010] To this end, the object of the present invention is to provide an end-to-end non-parsing virtual clothing change method, system and device based on dual optical flow, which can greatly reduce the training time of the non-parsing virtual clothing change method, and with the help of the dual optical flow mechanism, obtain better results than the original non-parsing virtual clothing change method.

[0011] In order to solve the above technical problems, the present invention is implemented as follows:

[0012] The embodiment of the present invention provides an end-to-end non-parsing virtual clothing change method based on dual optical flow, and the method includes:

[0013] S1. Obtain the clothing image to be worn and the human body model image, and encode them respectively to obtain the clothing feature map and the human body feature map;

[0014] S2. Perform task adaptive processing on the clothing feature map to obtain a first group of clothing feature maps and a second group of clothing feature maps; perform task adaptive processing on the human body feature map to obtain a first group of human body feature maps and a second group of human body feature maps;

[0015] S3. Splice the first group of clothing feature maps and the first group of human body feature maps, and perform dual optical flow decoding processing to obtain an optical flow field for distorting the shape of the clothing feature map;

[0016] S4. Use the optical flow field and the second group of clothing feature maps to distort the clothing feature map to obtain distortion data;

[0017] S5. Use the distortion data and the second group of human body feature maps to perform clothing change decoding processing to obtain the result of virtual clothing change.

[0018] In addition, according to the end-to-end non-parsing virtual clothing change method based on dual optical flow of the present invention, the following additional technical features may also be provided:

[0019] In some of the embodiments, the distortion of the clothing feature map is performed by point sampling.

[0020] In some of the embodiments, the task adaptive processing is a feature map reweighting processing based on the Hadamard product.

[0021] In some of the embodiments, the steps of the task adaptive processing in step S2 include:

[0022] Obtain the encoder feature map F;

[0023] Process the encoded feature map F with a dense block layer to obtain the first feature map F d ;

[0024] Successively process the first feature map F d with a Reshape layer, two fully connected layers, and a sigmoid activation function to obtain the global weight map G Fd ;

[0025] Use the global weight map G Fd to perform weighted processing on the first feature map F d to obtain the final weighted feature map F out .

[0026] In some of these embodiments, the content of performing weighted processing on the first feature map F d includes:

[0027] Multiply the global weight map G Fd and the first feature map F d to obtain the first weighted feature map F w d ;

[0028] Solve the Hadamard product of the first feature map F d and the first weighted feature map F w d and use a 1x1 convolutional layer to fuse the channel features of multiple multiplexed layers to obtain the final weighted feature map F out .

[0029] In some of these embodiments, the feature map dimension of the weighted feature map F out is the same as the feature map dimension of the encoder feature map F.

[0030] In some of these embodiments, the dual optical flow decoding process in step S3 is implemented using a pyramidal optical flow decoder. The pyramidal optical flow decoder includes four stacked modules. The subsequent module receives the optical flow passed from the previous module, and the size of the optical flow estimated by each module gradually increases from small to large.

[0031] In some of these embodiments, the processing content of each module in the pyramidal optical flow decoder includes:

[0032] Obtain the data to be processed;

[0033] Process the data to be processed using a bilinear interpolation upsampling operator to expand the spatial dimension of the optical flow to twice the original;

[0034] Perform convolutional calculation for style editing on the upsampled data to obtain global optical flow;

[0035] Use the global optical flow as a sampling grid to warp the clothing feature map of this layer, obtaining the warped clothing feature map of this layer;

[0036] Concatenate the warped clothing feature map with the human body feature map, and then perform ordinary convolutional calculation to obtain local optical flow;

[0037] Adjust the global optical flow with the local optical flow to obtain the output optical flow of the current pyramid layer

[0038] An embodiment of the present invention also provides an end-to-end non-analytical virtual clothing changing system based on dual optical flow, and the system includes:

[0039] A data acquisition module for acquiring the to-be-worn clothing map and the human body model map;

[0040] An encoder module, including a clothing encoder and a human body encoder, which are respectively used for encoding the to-be-worn clothing map and the human body model map to obtain the corresponding clothing feature map and human body feature map;

[0041] A task adaptation module, including a clothing task adaptation module and a human body task adaptation module, which are respectively used for performing task adaptation processing on the clothing feature map and the human body feature map to obtain the first group of clothing feature maps and the second group of clothing feature maps, the first group of human body feature maps and the second group of human body feature maps;

[0042] A splicing module for splicing the first group of clothing feature maps and the first group of human body feature maps;

[0043] A dual optical flow decoder for performing dual optical flow decoding processing on the data spliced by the splicing module to obtain an optical flow field for distorting the form of the clothing feature map;

[0044] A feature map distortion module for distorting the clothing feature map according to the optical flow field to obtain distorted data;

[0045] A clothing changing decoder for performing clothing changing decoding processing according to the distorted data and the second group of human body feature maps to obtain the result of virtual clothing changing.

[0046] An embodiment of the present invention also provides an end-to-end non-analytical virtual clothing changing device based on dual optical flow, including a processor and a memory. A software program is stored on the memory, and when the processor runs the software program stored on the memory, it can implement the steps of the end-to-end non-analytical virtual clothing changing method described in any one of the above.

[0047] Compared with the prior art, the present invention has at least the following beneficial effects:

[0048] In the embodiment of the present invention, the provided end-to-end virtual clothing change method based on dual optical flow can obtain a better visual clothing change result through the clothing distortion correction mechanism of dual optical flow; specifically, local optical flow is used to adjust the global optical flow, solving the problem of missing details that occurs when only using global optical flow for virtual clothing change, and avoiding the neglect of some details due to too strong global vision;

[0049] In the embodiment of the present invention, the provided end-to-end virtual clothing change method based on dual optical flow, through the module of feature map reweighting based on the Hadamard product, that is, the task adaptation module, can bring the clothing change performance of the model close to the level of using multiple encoders, but only increases a small amount of computational complexity, far superior to the way of using multiple encoders.

[0050] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flowchart of an analytical-based virtual clothing change method in the prior art;

[0052] Figure 2 is the process of a non-analytical virtual clothing change method in the prior art;

[0053] Figure 3 is a comparison diagram of clothing change results between the analytical-based virtual clothing change method and the non-analytical virtual clothing change method disclosed in the present invention;

[0054] Figure 4 is a framework diagram of an end-to-end virtual clothing change network disclosed in an embodiment of the present invention;

[0055] Figure 5 is a flowchart of a feature map reweighting module based on the Hadamard product disclosed in an embodiment of the present invention;

[0056] Figure 6 is a schematic diagram of the structure of a dual optical flow decoder disclosed in an embodiment of the present invention;

[0057] Figure 7 is an effect diagram of virtual clothing change after the cooperation of a dual optical flow encoder disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] Next, the embodiments of the present invention will be described in detail through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0060] In the existing virtual dressing methods, two models, namely, a clothing distortion model and a dressing model generation model, are required to operate together to complete the dressing task. Therefore, in the training stage, it is often necessary to first train the clothing distortion model and then jointly train the clothing distortion model and the generation model, which is quite time-consuming. In order to greatly reduce the training time of the non-analytical virtual dressing method, a single end-to-end network with a clothing distortion module and a generation module is designed to complete virtual dressing. For the overall architecture of the end-to-end virtual dressing network designed by the present invention, please refer to Figure 4 As shown, the overall framework includes two encoders and two decoders. Input the to-be-worn clothing image and the human model image into the virtual dressing model. They are respectively encoded by specific encoders to obtain the clothing feature map and the human body feature map, and then two sets of feature maps are obtained through their respective task adaptation modules. The two sets of feature maps respectively generated from the clothing feature map and the human body feature map both include the feature map applicable to the generation task and the feature map applicable to the optical flow estimation task. One set of feature maps is used as input data and enters the optical flow decoder. Before input, the clothing feature map and the human body feature map will first perform feature map splicing along the channel, and then a set of optical flow fields is decoded and output by the optical flow decoder. It will be used to distort the shape of the clothing feature map, that is, to perform feature map distortion together with the other set of features of the clothing feature map Figure 1 ; The distortion method is carried out by point sampling. After the feature map is distorted, the data and the other set of features of the human body feature map Figure 1 are decoded by the dressing decoder together to generate the final dressing result. Experimentally, it doesn't make much difference which set of the two sets of feature maps is used as the input of the optical flow encoder and which set of feature maps is used as the input of the encoder, because the numerical ranges of the weight sampling of the two are very close. If the mean value is 0.1, the range is [-2.9, 3.1], and if the mean value is 0.01, the range is [-3.0, 3.0].

[0061] The weight initialization methods of the two task adaptation modules are somewhat different, and they respectively follow Gaussian distribution sampling with two different mean values, one is 0.1 and the other is 0.01.

[0062] In the above embodiments, the optical flow decoding task and the clothing change generation task are quite different, and different human body parts need to be concerned so as to reasonably decouple the features suitable for their respective tasks and bring better clothing change effects, as shown in the following experimental chart 1:

[0063] Table 1

[0064]

[0065] It can be found from Table 1 that the clothing change performance of the model using multiple encoders is better than that of the model sharing encoders. FID and SSIM in the table are image quality evaluation indicators. The smaller the FID, the closer the generated image is to the image in the validation set, and the larger the SSIM, the higher the similarity of the clothing change image to the image in the validation set in terms of local details. Although using more encoders can bring performance improvement, it also greatly increases the computational complexity of the virtual clothing change model, which is not conducive to practical applications. To solve the above difficulties, the present invention designs a module for reweighting feature maps based on the Hadamard product, that is, Figure 4 the task adaptive module in Figure 5 . The results show that it can bring the clothing change performance of the model close to the level of using multiple encoders, but only increases a small amount of computational complexity, that is, greatly reduces the computational complexity compared with using multiple encoders. The structures and working processes of the task adaptive modules corresponding to the clothing feature map and the human body feature map are the same. The working process of the task adaptive module is as shown in

[0066] Specifically: d First, the task adaptive module receives the encoder feature map F, and then enters the Dense Bolck layer (dense block layer) for processing to obtain the first feature map F d ; The Dense Block layer can adopt existing modules, which have a unique feature map reuse mechanism. While enriching the diversity of feature map content, it does not require each layer to learn all new features separately, and the feature content is more coherent. Then, the first feature map F Fd is input into the Reshape layer and two fully connected layers, and then processed through a sigmoid activation function to obtain the global weight map G Fd . Finally, the feature map is weighted. The content of the weighted processing includes: multiplying the global weight map G Fd and the first feature map F d to obtain the first weighted feature map F w d . Then, the Hadamard product of the first feature map F d and the first weighted feature map F w d is solved, and a 1x1 convolutional layer is used to fuse the channel features multiplexed in multiple layers to obtain the second weighted feature map, that is, the final weighted feature map Fout At this time, the dimension of the feature map is BxCxHxW.

[0067] In the above embodiments of the present invention, the dual optical flow decoder is the main virtual clothing change interaction part of the present invention. Through the clothing distortion correction mechanism of the dual optical flow, a clothing change result with better visual effect can be obtained. The overall module structure is as Figure 6 shown. The dual optical flow refers to performing two optical flow estimations. The first optical flow is to perform the first distortion on the feature map, and the second optical flow is used to correct the first optical flow, and then the corrected optical flow is used to correct the feature map again. The dual optical flow decoder is pyramid-shaped, and the size of the estimated optical flow will gradually increase from small to large, and the magnification factor for each increase is 2 (for example, 32x32, 64x64, 128x128, 256x256). A total of four layers of modules are stacked. Each layer of the module will receive the optical flow passed from the previous layer. Then the total magnification factor is 16 times. After the optical flow f t-1 is input to this layer of the module, it will first pass through the upsampling operator u p of bilinear interpolation. The upsampling will expand the spatial dimension of the optical flow to 2 times the original; then, using the expanded optical flow as the sampling grid, the clothing feature map of this layer is distorted, and then the feature map is input to the customized convolution using style editing for operation. The process of style editing is as follows:

[0068] w' ijk = s i ·w ijk

[0069] where w and w' represent the weights of the original convolution and the convolution weights after style editing respectively, and s i is the adjustment factor corresponding to the i-th feature map. i represents the i-th input feature map, j refers to the j-th output feature map, and k represents the window size of the convolution.

[0070] After the editing is completed, an operation of demodulate is also required to eliminate the influence of δ in the output feature map statistics (to ensure the stability of gradient conduction). Specifically, assume that the input activation values are independent and identically distributed random variables with unit standard deviation. If after style editing and convolution, the standard deviation of the output activation values will be corrected to:

[0071]

[0072] Therefore, when performing the inverse editing operation on it, only need to multiply each feature map j by 1 / σ j to achieve. Or this operation can be directly integrated into the convolution weights:

[0073]

[0074] Among them, ε is an extremely small constant used to ensure the stability of data operations; w i ' j ' k refers to the demodulated convolutional weight. Thus, the convolution with the ability of style editing is obtained. Because it utilizes the adjustment factor of the style vector with a global receptive field, the optical flow generated after its interaction with the input feature map has a global receptive field and is mainly responsible for distorting the overall clothing. The convolution calculation of style editing is part of the optical flow decoder, corresponding to Figure 6 conv_m. After the input feature map undergoes the convolution calculation of style editing, the global optical flow f ci is obtained, and then the global optical flow f ci is used as the sampling grid to distort the clothing feature map C i . The distorted clothing feature map is concatenated with the human body feature map p i , and then calculated with a normal convolution to generate the optical flow f ri . This local optical flow f ri is used to adjust the global optical flow f ci , and finally the optical flow f i output by the current pyramid layer is obtained. Each calculation of the optical flow requires concatenating the clothing feature map and the human body feature map and then calculating, because in this way, the convolution module can adjust this part of the optical flow during convolution and calculation comparison; although the optical flow can be generated without referring to the human body feature map, the effect is not good.

[0075] The steps of adjusting the global optical flow with the local optical flow include:

[0076] 1) Generate coordinate points according to the magnitude of the global optical flow. For example, (0,0), (1,0), (0,1), (1,1), and the array dimension is [B, H, W, 2], where B refers to the training batch size, H refers to the length of the global optical flow matrix, and W refers to the width of the global optical flow matrix;

[0077] 2) Add the coordinate values and the optical flow S = O + offset, where S is the output new coordinate value, O is the coordinate point matrix, and offset is the optical flow;

[0078] 3) Normalize the values stored in the new coordinate value S to [-1, 1]. The method is:

[0079] S_out = S / size_value / 2 - 1;

[0080] 4) Finally, perform point sampling on the global optical flow according to the newly generated coordinate value S_out to obtain the adjusted optical flow.

[0081] Figure 2The concept of optical flow transfer in the prior art is part of an analytical virtual dressing method. In the analytical virtual dressing method, the optical flow obtained from the analytical virtual dressing method is used as a supervision signal of a teacher model to the student model. Corresponding to the present invention, specifically, it is the supervision of the optical flow finally output by the present invention.

[0082] As Figure 7 shown, it is the effect of virtual dressing after the cooperation of the dual optical flow encoders designed by the present invention. If only the global optical flow (that is, f ci ) obtained previously is used to distort the clothing and then generate the dressed image, it can be found that it is very difficult to align the cuffs of the clothes with the shoulders of the human body. This may be because the global vision is too strong, resulting in the omission of some details when generating the optical flow. By adding the local optical flow in the present invention and using the dual optical flow to distort the clothing, the above problems are immediately solved.

[0083] Based on the above improvements, the model effect of the present invention can be seen in Table 2 as follows.

[0084] Table 2

[0085]

[0086] Compared with other existing virtual dressing models, the model of the present invention has a better dressing effect and greatly reduces the training time of the analytical image method by 75%.

[0087] For the parts not detailed in the present invention, reference can be made to the prior art in this field or the well-known techniques of those skilled in the art, and no further elaboration will be made here.

[0088] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them belong to the protection scope of the present invention.

Claims

1. An end-to-end non-analytic virtual dressing method based on dual optical flows, characterized in that: The method comprises: S1. Obtain a picture of clothing to be worn and a picture of a human model, and encode them respectively to obtain a clothing feature picture and a human feature picture; S2, performing task adaptive processing on the clothing feature graph to obtain a first group of clothing feature graphs and a second group of clothing feature graphs; performing task adaptive processing on the human body feature graph to obtain a first group of human body feature graphs and a second group of human body feature graphs; S3, splicing the first group of clothing feature maps and the first group of human body feature maps, and performing dual optical flow decoding processing to obtain an optical flow field for distorting the shape of the clothing feature map; S4, using the optical flow field and the second set of clothing feature maps to distort the clothing feature map to obtain distorted data; S5. Perform a dressing-changing decoding process using the distorted data and the second group of human body feature maps to obtain a virtual dressing-changing result.

2. The end-to-end non-analytic virtual dressing method based on dual optical flows according to claim 1 is characterized in that: The clothing feature map distortion is performed using point sampling.

3. The end-to-end non-analytic virtual dressing method based on dual optical flows according to claim 1 is characterized in that: The task adaptive processing is a feature map re-weighting processing based on the Hammar product.

4. The end-to-end non-analytic virtual dressing method based on dual optical flows according to any one of claims 1 to 3, characterized in that: The steps of task adaptive processing in step S2 include: Get the encoder feature map F; The encoded feature map F is processed with a dense block layer to obtain the first feature map F d ; The first feature map F is activated by using the Reshape layer, two fully connected layers and a sigmoid activation function. d Processing is performed to obtain the global weight graph G Fd ; With the global weight graph G Fd For the first feature map F d After weighted processing, the final weighted feature map F out .

5. The end-to-end non-analytic virtual dressing method based on dual optical flows according to claim 4 is characterized in that: For the first feature map F d The weighted content includes: The global weight graph G Fd and the first feature map F d Multiply them to get the first weighted feature map F w d ; Solve the first characteristic graph F d And the first weighted feature map F w d Hammar product, and use 1x1 convolution layer to fuse the multi-layer multiplexed channel features to obtain the final weighted feature map F out .

6. The end-to-end non-analytic virtual dressing method based on dual optical flows according to claim 4 is characterized in that: Weighted feature map F out The feature map dimension of is the same as the feature map dimension of the encoder feature map F.

7. The end-to-end non-analytic virtual dressing method based on dual optical flows according to claim 1, characterized in that: The dual optical flow decoding process in step S3 is implemented by a pyramid optical flow decoder. The pyramid optical flow decoder includes four stacked modules. The modules in the next layer receive the optical flow transmitted by the modules in the previous layer. The size of the optical flow estimated by each layer of modules gradually increases from small to large.

8. The end-to-end non-analytic virtual dressing method based on dual optical flows according to claim 7 is characterized in that: The processing content of each layer module in the pyramid optical flow decoder includes: Get the data to be processed; The upsampling operator of bilinear interpolation is used to process the data to be processed, and the spatial dimension of the optical flow is expanded to twice the original; Perform style-edited convolution on the upsampled data to obtain the global optical flow; The global optical flow is used as a sampling grid to distort the clothing feature map of the layer, so as to obtain the distorted clothing feature map of the layer; The distorted clothing feature map is spliced ​​with the human body feature map, and then ordinary convolution calculation is performed to obtain the local optical flow; The global optical flow is adjusted with the local optical flow to obtain the output optical flow of the current pyramid layer.

9. An end-to-end non-analytic virtual dressing system based on dual optical flows, characterized in that: The system comprises: A data acquisition module, used to acquire images of clothing to be worn and images of human models; The encoder module includes a clothing encoder and a human body encoder, which are used to encode the image of the clothing to be worn and the image of the human model, respectively, to obtain the corresponding clothing feature image and human body feature image; The task adaptation module includes a clothing task adaptation module and a human task adaptation module, which are used to perform task adaptation processing on clothing feature maps and human feature maps respectively, and obtain a first group of clothing feature maps and a second group of clothing feature maps, a first group of human feature maps and a second group of human feature maps respectively; A splicing module, used for splicing the first set of clothing feature images and the first set of human body feature images; A dual optical flow decoder is used to perform dual optical flow decoding on the data spliced ​​by the splicing module to obtain an optical flow field for distorting the shape of the clothing feature map; A feature map distortion module is used to distort the second set of clothing feature maps according to the optical flow field to obtain distorted data; The dress-changing decoder is used to perform dress-changing decoding processing according to the distorted data and the second set of human body feature maps to obtain a virtual dress-changing result.

10. An end-to-end non-analyzing virtual dressing device based on dual optical flows, comprising a processor and a memory, wherein a software program is stored in the memory, characterized in that: When the processor runs the software program stored in the memory, it can implement the steps of the end-to-end non-parsed virtual dressing method based on dual optical flow as described in any one of claims 1 to 8.