Image processing method and device, electronic device, and storage medium
By generating a deformed image of the target clothing adapted to the human body and fusing it with the target person, the problems of low quality and poor real-time performance of virtual dressing images caused by relying on human body analysis in existing technologies are solved, and high-quality and real-time virtual dressing effects are achieved.
Patent Information
- Application Number
- CN202110141360.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-27
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2041-01-29
AI Technical Summary
Existing virtual dressing technology is highly dependent on human body analysis results, resulting in low quality of virtual dressing images and weak real-time performance. The human body analysis process is time-consuming and cannot achieve real-time, high-quality virtual dressing.
By obtaining the appearance flow features of the target clothing that adapts to the target person's body deformation, a deformed image is generated and fused with the target person to achieve virtual clothing change without relying on the body analysis results.
High-quality virtual dressing image generation is achieved, the real-time performance of virtual dressing is improved, and the problems of low image quality and weak real-time performance caused by relying on human body analysis results are avoided.
Smart Images

Figure CN113570685B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and more specifically, to an image processing method and device, an electronic device, and a computer-readable storage medium. Background Art
[0002] Virtual dressing technology involves fusing images of the human body with images of clothing to create an image of the user wearing the clothing. This allows users to see how the clothing will look without having to wear the actual clothing. Virtual dressing technology is widely used in online shopping, clothing displays, fashion design, and virtual try-ons for offline shopping.
[0003] Existing virtual dress-up techniques rely on human body analysis results from human images. An ideal virtual dress-up dataset should contain images of a designated person wearing any clothing, images containing the target clothing, and images of the designated person wearing the target clothing. However, because images of the same person wearing two different clothing items while maintaining identical postures are difficult to obtain, currently used virtual dress-up datasets only contain images of the designated person wearing the target clothing. Human body analysis results are then used to remove the designated person's target clothing area, and then the human image is reconstructed using images containing the target clothing.
[0004] This demonstrates that existing virtual dress-up technologies are highly dependent on human body analysis results. Inaccurate human body analysis results in a virtual dress-up image where the target person and the target clothing do not match. Furthermore, in practical applications, human body analysis is time-consuming, making it impossible to achieve real-time virtual dress-up results. Summary of the Invention
[0005] To solve the above technical problems, embodiments of the present application provide an image processing method and device, an electronic device, and a computer-readable storage medium.
[0006] According to one aspect of an embodiment of the present application, an image processing method is provided, comprising: acquiring a first image containing a target person and a second image containing a target garment; generating, based on image features corresponding to the first image and image features corresponding to the second image, appearance flow features for characterizing the deformation of the target garment adapted to the human body of the target person, and generating, based on the appearance flow features, a deformed image of the target garment adapted to the human body; and generating, based on the appearance flow features, a virtual dressing image by fusing the deformed image and the first image, in which the target person wears the target garment adapted to the human body.
[0007] According to one aspect of an embodiment of the present application, an image processing device is provided, comprising: an image acquisition module configured to acquire a first image containing a target person and a second image containing a target garment; an information generation module configured to generate, based on image features corresponding to the first image and image features corresponding to the second image, appearance flow features for characterizing the deformation of the target garment adapted to the human body of the target person, and to generate, based on the appearance flow features, a deformed image of the target garment adapted to the human body; and a virtual dressing module configured to generate a virtual dressing image based on the fusion of the deformed image and the first image, in which the target person wears the target garment adapted to the human body.
[0008] According to one aspect of an embodiment of the present application, an electronic device is provided, including a processor and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the image processing method described above is implemented.
[0009] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a computer, the computer executes the image processing method described above.
[0010] According to one aspect of an embodiment of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image processing methods provided in the various optional embodiments described above.
[0011] In the technical solution provided in the embodiments of the present application, there is no need to rely on the results of human body analysis to perform virtual dressing. Instead, the deformation of the target clothing adapted to the human body is generated by obtaining the appearance flow characteristics of the deformation produced by the target clothing adapting to the human body of the target person. Finally, the deformed target clothing and the target person are fused to obtain a dressing image, thereby solving various problems caused by relying on human body analysis results for virtual dressing in the existing technology implementation and achieving high-quality virtual dressing.
[0012] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, serving to explain the principles of the present application. It is obvious that the drawings described below are merely some embodiments of the present application, and a person of ordinary skill in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0014] Figure 1 is a schematic diagram of an implementation environment involved in this application;
[0015] Figure 2 1 is a structural diagram of a virtual dress-up student model according to an embodiment of the present application;
[0016] Figure 3 yes Figure 2 The structure diagram of the first clothing deformation sub-model 11 in one embodiment is shown;
[0017] Figure 4 yes Figure 3 Schematic diagram of the process of appearance flow feature prediction performed by the "FN-2" module in the second image feature layer;
[0018] Figure 5 is a flowchart of an image processing method shown in another embodiment of the present application;
[0019] Figure 6 1 is a schematic diagram of a training process for a virtual dress-up student model according to an embodiment of the present application;
[0020] Figure 7 is a block diagram of an image processing device according to an embodiment of the present application;
[0021] Figure 8 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0022] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0023] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0024] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0025] It should also be noted that the term "plurality" used in this application refers to two or more than two. The terms "first," "second," and the like may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are used solely to distinguish one concept from another. For example, a first image may be referred to as a second image, and similarly, a second image may be referred to as a first image, without departing from the scope of this application.
[0026] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0027] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies include natural language processing and machine learning.
[0028] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.
[0029] Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision techniques such as using cameras and computers to replace the human eye in identifying, tracking, and measuring objects. This involves further processing the images, transforming them into images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, virtual reality, augmented reality, simultaneous localization and mapping, and common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0030] The following describes the image processing method provided in the embodiment of the present application based on artificial intelligence technology and computer vision technology.
[0031] The embodiment of the present application provides an image processing method, the execution subject is a computer device, which can fuse human body images and clothing images. In one possible implementation, the computer device is a terminal, which can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, car computer, etc. In another possible implementation, the computer device is a server, which can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, wherein multiple servers can form a blockchain, and the server is a node on the blockchain. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms.
[0032] The image processing method provided in the embodiments of the present application can be applied to any scenario where a human body image and an image of clothing are to be merged. For example, in the scenario of a virtual dress-up during online shopping, if a user wants to see the effect of wearing a certain item of clothing, they only need to provide an image of the user's body and an image of the clothing item. Using the method provided in the embodiments of the present application, the human body image and the clothing image are processed to obtain an image of the user wearing the clothing item, thus enabling online virtual dress-up without the user having to wear the actual clothing.
[0033] In addition, the image processing method provided in the embodiment of the present application can also be applied to scenarios such as clothing design, clothing display, or virtual try-on for offline shopping to provide real-time virtual dressing functions, which are not listed here one by one.
[0034] See also Figure 1 , Figure 1 This is a flowchart of an image processing method according to one embodiment of the present application. The image processing method includes at least steps S110 to S150, which can be implemented as a virtual dress-up student model. The virtual dress-up student model is an artificial intelligence model that can achieve virtual dress-up of a target person without relying on human body analysis results. This model not only generates high-quality virtual dress-up images but also improves the real-time performance of virtual dress-up.
[0035] The following for Figure 1 The image processing method shown is described in detail:
[0036] Step S110 , obtaining a first image containing a target person and a second image containing a target garment.
[0037] The target person mentioned in this embodiment refers to the person to be virtually dressed up, and the target clothing refers to the clothing that the target task wants to wear.
[0038] For example, in a scenario involving virtual dressing during online shopping, the target person is the user currently shopping online. The first image is a body image provided by the user, while the second image can be a picture of the target clothing loaded on the shopping platform. It should be noted that the target person in the first image and the target clothing in the second image can be determined based on the actual application scenario and are not limited here.
[0039] Step S130 , generating appearance flow features for characterizing the deformation of the target clothing adapted to the target person's body based on the image features corresponding to the first image and the image features corresponding to the second image, and generating a deformed image of the target clothing adapted to the body based on the appearance flow features.
[0040] First, the image features corresponding to the first image are obtained by performing image feature extraction on the first image, and the image features corresponding to the second image are obtained by performing image feature extraction on the second image. For example, in some embodiments, the first image can be input into a first image feature extraction model, and the second image can be input into a second image feature extraction model. Both the first image feature extraction model and the second image feature extraction model are configured with image feature extraction algorithms, thereby obtaining the image features output by the first image feature extraction model for the first image, and obtaining the image features output by the second image feature extraction model for the second image.
[0041] The image features corresponding to the first image output by the first image feature extraction model and the image features corresponding to the second image output by the second image feature extraction model can be multi-layer image features, which refer to multiple feature maps obtained in sequence during the process of image feature extraction on the first image and the second image.
[0042] Exemplarily, the first image feature extraction model and the second image feature extraction model can be pyramid feature extraction models, in which a feature pyramid network (FPN) is configured. The feature graph pyramid output by the feature pyramid network is also the multi-layer image feature corresponding to the image. For example, in some embodiments, the bottom-up part of the pyramid feature extraction model can be used to extract image features from the first image and the second image. The bottom-up part is understood as using a convolutional network to extract image features. As the convolution deepens, the image spatial resolution is reduced and spatial information is lost, but high-level semantic information is enriched, thereby obtaining multi-layer image features with feature graph sizes sorted from large to small.
[0043] The appearance flow feature refers to a two-dimensional coordinate vector, which is usually used to indicate which pixel of the source image can be used to reconstruct a specified pixel of the target image. In order to achieve high-quality virtual dressing, this embodiment needs to establish an accurate and dense correspondence between the target person's body and the target clothing, so that the target clothing can be deformed to adapt to the body. Therefore, in this embodiment, the source image refers to the second image, specifically the target clothing area in the second image, and the target image to be reconstructed refers to the image after the target clothing is deformed to adapt to the target person's body in the first image.
[0044] It can be seen that the appearance flow feature can characterize the deformation of the target clothing adapted to the body of the target person in the first image. According to the obtained appearance flow feature, the deformed image of the target clothing adapted to the body can be generated.
[0045] When the image features corresponding to the first image are multi-layer image features output by the first image feature extraction model, and the image features corresponding to the second image are multi-layer image features output by the second image feature extraction model, the appearance flow features can be extracted layer by layer based on the multi-layer image features output by the first image feature extraction model and the second image feature extraction model, and the appearance flow features extracted from the last image feature layer are used as the final appearance flow features.
[0046] For example, at the first image feature layer, appearance flow features representing the deformation of the target garment as it adapts to the target person's body can be extracted based on the image features output by the first and second image feature extraction models. At each subsequent image feature layer, the appearance flow features output by the previous image feature layer are optimized based on the image features output by the first and second image feature extraction models to obtain the appearance flow features corresponding to the current image feature layer.
[0047] During the process of extracting appearance flow features layer by layer based on the multi-layer image features output by the first and second image feature extraction models, the appearance flow features can also be extracted based on a preset second-order smoothness constraint. The second-order smoothness constraint is a constraint imposed on the linear correspondence between adjacent appearance flows to further preserve features such as target clothing patterns and stripes, thereby improving the image quality of the generated deformed image of the target clothing adapted to the target person's body.
[0048] Step S150 , generating a virtual dress-changing image based on the deformed image of the target clothing adapted to the target person's body and the first image. In the virtual dress-changing image, the target person wears the target clothing adapted to the human body.
[0049] The virtual dressing image is generated by fusing the deformed image of the target clothing adapted to the target person's body with the first image. This can be achieved by using an image fusion algorithm suitable for virtual dressing, such as the Res-UNet algorithm, which is not limited in this embodiment.
[0050] From the above, it can be seen that the technical solution provided in the embodiment does not need to rely on the human body analysis results to perform virtual dressing. Instead, it generates a deformation of the target clothing adapted to the human body by obtaining the appearance flow characteristics of the deformation caused by the target clothing adapting to the human body of the target person. Finally, the deformed target clothing and the target person are fused to obtain a dressing image. This avoids the problems of low quality of virtual dressing images and weak real-time performance of virtual dressing caused by relying on human body analysis results for virtual dressing, and achieves high-quality virtual dressing.
[0051] See also Figure 2 , Figure 2 This is a structural diagram of a virtual dress-up student model according to an embodiment of the present application. The exemplary virtual dress-up student model 10 includes a first clothing deformation sub-model 11 and a first dress-up generation sub-model 12, wherein the first clothing deformation sub-model 11 can perform Figure 1 In step S130 of the embodiment shown, the first outfit generation sub-model 12 may execute Figure 1 Step S150 in the illustrated embodiment.
[0052] like Figure 2 As shown, by inputting a first image containing a target person and a second image containing a target garment into the virtual dressing-up student model 10, the virtual dressing-up student model 10 can output a corresponding virtual dressing-up image. In the output virtual dressing-up image, the target person wears the target garment that is adapted to the human body.
[0053] In addition to the first image and the second image, the virtual dress-up student model 10 does not require any other additional input signals, and there is no need to input the human body analysis result of the target person contained in the first image into the virtual dress-up student model 10 .
[0054] Figure 3 yes Figure 2 The first clothing deformation sub-model 11 is a schematic structural diagram in one embodiment. Figure 3 As shown, the first clothing deformation sub-model 11 includes a first image feature extraction model, a second image feature extraction model and an appearance flow feature prediction model.
[0055] The first image feature extraction model is used to extract the image features corresponding to the first image, and the second image feature extraction model is used to extract the image features corresponding to the second image. Figure 3 As shown, the first image feature extraction model extracts image features from the first image and sequentially obtains the multi-layer image features shown in c1 to c3, and the second image feature extraction model extracts image features from the second image and sequentially obtains the multi-layer image features shown in p1 to p3.
[0056] It should be noted that Figure 3 The multi-layer image features shown are only examples. The number of layers of image features corresponding to the input image extracted by the first image feature extraction model and the second image feature extraction model can be set according to actual needs, and this embodiment does not limit this.
[0057] The appearance flow feature prediction model is used to extract appearance flow features layer by layer based on the multi-layer image features output by the first image feature extraction model and the second image feature extraction model, and the appearance flow features extracted from the last image feature layer are used as the final appearance flow features. For example, Figure 3The "FN-1" module is used to predict appearance flow features at the first image feature layer, the "FN-2" module is used to predict appearance flow features at the second image feature layer, and the "FN-3" module is used to predict appearance flow features at the third image feature layer. In other words, the appearance flow feature prediction model is a progressive appearance flow feature prediction model.
[0058] like Figure 3 As shown, the appearance flow feature prediction model extracts appearance flow features representing the deformation of the target garment as it adapts to the target person's body based on the image features output by the first and second image feature extraction models at the first image feature layer. At each subsequent image feature layer, the appearance flow feature prediction model optimizes the appearance flow features output by the previous image feature layer based on the image features output by the first and second image feature extraction models to obtain appearance flow features corresponding to the current image feature layer.
[0059] Through such a progressive processing method, as the multi-layer image features are continuously deepened with the convolution, the image spatial resolution gradually decreases and the spatial information is gradually lost, but the high-level semantic information is enriched, so that the feature information contained in the appearance flow features obtained layer by layer in the appearance flow feature prediction model becomes more and more rich and accurate. For example, Figure 3 The appearance flow features f1 to f3 shown in FIG. 1 gradually contain richer feature information and gradually adapt to the human body of the target person.
[0060] It can be seen from this that the appearance flow features obtained by the appearance flow feature prediction model in the last image feature layer can very accurately reflect the deformation caused by the target clothing adapting to the target person's body, and the deformed image corresponding to the target clothing generated based on the appearance flow features obtained by the appearance flow feature prediction model in the last image feature layer can establish an accurate and close correspondence with the target person's body, so that subsequent fusion can be carried out based on the accurate deformation of the target clothing and the target person's body to obtain a high-quality virtual dressing image.
[0061] Figure 4 yes Figure 3 The flow chart of appearance flow feature prediction performed by the “FN-2” module in the second image feature layer is shown in FIG. Figure 4As shown, first, the appearance flow feature f1 output by the previous image feature layer is upsampled to obtain the upsampled feature f1', and then the image feature c2 of the second image output by the current feature layer is deformed according to the upsampled feature f1' to obtain the first deformed feature c2'. Next, based on the image feature p2 of the first image output by the current image feature layer, the first deformed feature c2' is corrected to obtain the corrected feature r2, and the corrected feature r2 is convolved to obtain the first convolution feature f2'. Next, based on the feature f2'' obtained by splicing the first convolution feature f2'' and the up-sampled feature f1', a second deformation process is performed on the image feature c2 of the second image output by the current image feature layer to obtain the second deformed feature p2 c2'. The second deformed feature is the splicing of the image feature p2 of the first image output by the current image feature layer and another feature c2'. Finally, a second convolution calculation is performed on the second deformed feature p2 c2', and the calculated second convolution feature f2' is spliced with the first convolution feature f2', so that the appearance flow feature f2 corresponding to the current image feature layer can be obtained.
[0062] As can be seen above, upsampling the appearance flow features output by the previous image feature layer helps improve the resolution of the appearance flow features of the current image feature layer. Subsequent two deformation processes and two convolution calculations further refine the feature information contained in the upsampled features. This is equivalent to adding spatial information to the appearance flow features output by the previous image feature layer. This optimizes the appearance flow features output by the previous image feature layer, resulting in appearance flow features that further reflect the deformation of the target clothing to the target person's body.
[0063] It is also mentioned that, in some embodiments, the appearance flow feature prediction model, in the process of extracting appearance flow features layer by layer based on the multi-layer image features output by the first image feature extraction model and the second image feature extraction model, also extracts appearance flow features based on the second-order smoothness constraint conditions preset for the linear correspondence between adjacent appearance flows, so as to further retain the patterns, stripes and other features of the target clothing.
[0064] Figure 5 FIG. 1 is a flowchart of an image processing method according to another embodiment of the present application. Figure 5 As shown, this method Figure 1 The embodiment shown further includes steps S210 to S250, which are described in detail as follows:
[0065] Step S210, calling the virtual dressing assistant model, inputting the human body analysis result corresponding to the character image containing the specified character and the first clothing image containing the clothing to be changed into the virtual dressing assistant model, and obtaining the assistant image output by the virtual dressing assistant model. In the assistant image, the specified character wears a deformed image of the human body adapted to the specified character.
[0066] First of all, this embodiment discloses Figure 2 The training process of the virtual dress-up student model shown in the figure is as follows. During the training phase of the virtual dress-up student model, a virtual dress-up teaching assistant model is called upon for auxiliary training. Specifically, the virtual dress-up teaching assistant model is an artificial intelligence model that relies on human body analysis results. By inputting the human body analysis results corresponding to a person image containing a specified person and a first clothing image containing the clothing to be changed into, the virtual dress-up teaching assistant model can output the corresponding teaching assistant image. In the teaching assistant image, the specified person wears a deformed image of the human body adapted to the specified person.
[0067] In this embodiment, the virtual dress-up dataset is composed of an image containing a designated person's image, a first clothing image containing the clothing to be changed, and a second clothing image containing the original clothing worn by the designated person. The number of person images, first clothing images, and second clothing images can be multiple, and different designated persons can be included in different person images, which is not a limitation in this embodiment.
[0068] In step S230, a second clothing image including the original clothing worn by the designated person in the character image and the teaching assistant image are input into the virtual dressing-up student model to be trained, and a student image output by the virtual dressing-up student model to be trained is obtained. In the student image, the designated person wears a deformed image of the human body adapted to the designated person in the teaching assistant image.
[0069] Since the virtual dressing student model does not rely on the human body analysis results to achieve virtual dressing, and the features extracted by the virtual dressing assistant model based on the human body analysis results will contain richer semantic information and feature expressions, this embodiment uses the virtual dressing assistant model to guide the training of the virtual dressing student model.
[0070] In other words, this embodiment uses knowledge distillation to train the virtual dress-up student model. Knowledge distillation involves using the intrinsic information of a teacher network to train the student network. In this embodiment, the teacher network is a virtual dress-up assistant model, and the intrinsic information of the teacher network refers to the feature expressions and semantic information extracted by the virtual dress-up assistant model based on the human body analysis results.
[0071] The trained virtual dressing student model has fully learned the accurate and dense correspondence between the human body and clothing. Therefore, in practical applications, there is no need to obtain the human body analysis results of the target person. The virtual dressing student model can still output high-quality virtual dressing images based on the first image containing the target person and the second image containing the target clothing as input.
[0072] Specifically, this embodiment inputs the teaching assistant image output by the virtual dress-up teaching assistant model as teaching assistant knowledge into the virtual dress-up student model to be trained, and inputs a second clothing image containing the original clothing worn by the designated character in the character image into the virtual dress-up student model to be trained, so that the virtual dress-up student model to be trained outputs a student image. In the student image, the designated character wears a deformed image of the human body adapted to the designated character in the teaching assistant image.
[0073] In step S250 , the character image is used as the teacher image, and the parameters of the virtual dress-up student model to be trained are updated according to the image loss information between the student image and the teacher image.
[0074] This embodiment uses a person image as a teacher image to supervise the training process of the virtual dress-up student model. This means that the virtual dress-up student model can be directly supervised by the teacher image during training, which helps improve the performance of the virtual dress-up student model. The resulting trained virtual dress-up student model, in practical applications, is no longer dependent on human body analysis results and can output high-quality virtual dress-up images based on the input first and second images.
[0075] The image loss information between the student image and the teacher image can be obtained by calculating the loss function value of the student image and the teacher image. Exemplarily, the image loss value of the student image relative to the teacher image can be obtained. The image loss value can include at least one of the pixel distance loss function value, the perceptual loss function value, and the adversarial loss function value. Then, a sum operation is performed on the at least one image loss value to obtain the image loss sum value of the student image relative to the teacher image. Finally, the image loss sum value is used as the image loss information between the student image and the teacher image to update the parameters of the virtual dress-up student model to be trained, thereby completing the training of the virtual dress-up student model.
[0076] By training the virtual dressing student model to be trained multiple times to gradually improve the model performance of the virtual dressing student model, when the image loss information between the student image and the teacher image is less than or equal to the preset image loss threshold, it means that the virtual dressing student model has achieved better model performance, and the training process of the virtual dressing student model can be ended.
[0077] It should also be mentioned that the human body analysis results can include information such as human body key points, human body posture heat maps, dense posture estimation, etc. In most cases, the virtual dressing assistant model can extract richer semantic information based on the human body analysis results, and the predicted appearance flow features will also be more accurate. Therefore, the image quality of the assistant image output by the virtual dressing assistant model should be higher than the student image output by the virtual dressing student model.
[0078] If the human body parsing results input into the virtual dressing assistant model are inaccurate, the virtual dressing assistant model will provide completely wrong guidance to the virtual dressing student model during the training process of the virtual dressing student model. Therefore, it is necessary to set up an adjustable knowledge distillation mechanism to ensure that only accurate teaching assistant images can be used to train the virtual dressing student model.
[0079] Specifically, before step S250, the image quality difference between the teaching assistant image and the student image is obtained. If this image quality difference is positive, it indicates that the image quality of the teaching assistant image is greater than that of the student image, and step S250 is then executed to train the virtual dress-up student model based on this teaching assistant image. If this image quality difference is negative or zero, it indicates that the image quality of the teaching assistant image is not greater than that of the student image, and the human body analysis result input into the virtual dress-up teaching assistant model may be completely incorrect. Therefore, step S250 is terminated and the next round of virtual dress-up student model training begins.
[0080] Figure 6 FIG. 1 is a diagram showing a training process of a virtual dress-up student model according to an embodiment of the present application. Figure 6 As shown, the virtual dressing assistant model 20 is used as an auxiliary model for training the virtual dressing student model 10. The virtual dressing assistant model 20 outputs a corresponding teaching assistant image based on the first clothing image input and the human body analysis results obtained by performing human body analysis on the character image (i.e., the teacher image). The teaching assistant image and the second clothing image output by the virtual dressing assistant model 20 are then input into the virtual dressing student model 10 to obtain the student image output by the virtual dressing student model 10. Based on the image loss information between the student image and the teacher image, the parameters of the virtual dressing student model 10 can be updated.
[0081] The virtual dressing assistant model 20 includes a second clothing deformation sub-model 21 and a second clothing generation sub-model 22. By calling the second clothing deformation sub-model 21, a deformed image of the human body of the designated person with the clothing to be changed can be generated based on the human body analysis results and the image features of the first clothing image. Figure 3 and Figure 4By calling the second clothing-changing generation sub-model 22, the teaching assistant image can be generated by fusing the deformed image corresponding to the clothing to be changed output by the second clothing deformation sub-model and the image regions other than the region where the original clothing is worn in the character image.
[0082] In another embodiment, by calling the second clothing generation sub-model 22, the area of the character image containing the original clothing worn by the specified character can be cleared according to the human body analysis results to obtain other image areas in the character image except the area wearing the original clothing.
[0083] It should be noted that the first clothing deformation sub-model contained in the virtual dress-up student model and the second clothing deformation sub-model contained in the virtual dress-up teaching assistant model can have the same network structure, for example, Figure 3 The network structure shown. The first dressing generation sub-model contained in the virtual dressing student model and the second dressing generation sub-model contained in the virtual dressing assistant model can also have the same network structure. For example, the first dressing generation sub-model and the second dressing generation sub-model can be composed of an encoder-decoder network and a residual network. The residual network is used to normalize the upper network to which it is connected, thereby facilitating parameter optimization during the model training process.
[0084] From the above, it can be seen that this application uses a novel "teacher-teaching assistant-student" knowledge distillation mechanism to train a virtual dressing student model that does not rely on human body analysis results. The virtual dressing student model is supervised by the teacher image during the training process, so that the virtual dressing student model finally trained does not need to rely on human body analysis results to generate highly realistic virtual dressing results, thereby achieving high-quality virtual dressing without relying on human body analysis results.
[0085] Figure 7 FIG. 1 is a block diagram of an image processing device according to an embodiment of the present application. Figure 7 As shown, in an exemplary embodiment, the image processing device includes:
[0086] The image acquisition module 310 is configured to acquire a first image containing a target person and a second image containing a target garment; the information generation module 330 is configured to generate, based on image features corresponding to the first image and image features corresponding to the second image, appearance flow features for characterizing the deformation of the target garment adapted to the human body of the target person, and generate a deformed image of the target garment adapted to the human body based on the appearance flow features; the virtual dressing module 350 is configured to generate a virtual dressing image based on the fusion of the deformed image and the first image, in which the target person wears the target garment adapted to the human body.
[0087] In another exemplary embodiment, the information generation module 330 includes:
[0088] The multi-layer image feature acquisition unit is configured to use the first image as the input signal of the first image feature extraction model, and the second image as the input signal of the second image feature extraction model, and extract the multi-layer image features corresponding to the input signals through the first image feature extraction model and the second image feature extraction model; the appearance flow feature extraction unit is configured to extract the appearance flow features layer by layer based on the multi-layer image features output by the first image feature extraction model and the second image feature extraction model, and use the appearance flow features extracted for the last image feature layer as the final appearance flow features generated.
[0089] In another exemplary embodiment, the appearance flow feature extraction unit includes:
[0090] The first feature extraction subunit is configured to extract, at the first image feature layer, appearance flow features used to characterize the deformation of the target clothing adapted to the human body of the target person based on the image features output by the first image feature extraction model and the second image feature extraction model; the second feature extraction subunit is configured to optimize the appearance flow features outputted for the previous image feature layer based on the image features outputted by the first image feature extraction model and the second image feature extraction model at each image feature layer subsequent to the first image feature layer, so as to obtain appearance flow features corresponding to the current image feature layer.
[0091] In another exemplary embodiment, the second feature extraction subunit includes:
[0092] The first deformation processing subunit is configured to perform upsampling processing on the appearance flow features output by the previous image feature layer to obtain upsampled features, and perform a first deformation processing on the image features of the second image output by the current image feature layer to obtain first deformed features; the correction processing subunit is configured to perform correction processing on the first deformed features based on the image features of the first image output by the current image feature layer, and perform a first convolution calculation on the corrected features obtained by the correction processing to obtain first convolution features; the second deformation processing subunit is configured to perform a second deformation processing on the image features of the second image output by the current image feature layer based on the features obtained by splicing the first convolution features and the upsampled features to obtain second deformed features; the appearance flow feature acquisition subunit is configured to perform a second convolution calculation on the second deformed features, and splice the calculated second convolution features with the first convolution features to obtain appearance flow features corresponding to the current image feature layer.
[0093] In another exemplary embodiment, the information generation module 330 further includes:
[0094] The second-order smoothness constraint unit is configured to extract appearance flow features layer by layer based on the multi-layer image features output by the first image feature extraction model and the second image feature extraction model, and further extract appearance flow features based on a second-order smoothness constraint condition preset for the linear correspondence between adjacent appearance flows.
[0095] In another exemplary embodiment, the information generation module 330 is configured as the first clothing deformation sub-model contained in the virtual dressing-up student model, and the virtual dressing-up module 350 is configured as the first dressing-up generation sub-model contained in the virtual dressing-up student model.
[0096] In another exemplary embodiment, the image processing apparatus further includes:
[0097] The teaching assistant image acquisition module is configured to call the virtual dressing-up teaching assistant model, input the human body analysis result corresponding to the character image containing the specified character and the first clothing image containing the clothing to be changed into the virtual dressing-up teaching assistant model, and obtain the teaching assistant image output by the virtual dressing-up teaching assistant model, in which the specified character wears a deformed image of the human body adapted to the specified character; the student image acquisition module is configured to input the second clothing image containing the original clothing worn by the specified character in the character image and the teaching assistant image into the virtual dressing-up student model to be trained, and obtain the student image output by the virtual dressing-up student model to be trained, in which the specified character wears a deformed image of the human body adapted to the specified character in the teaching assistant image; the parameter updating module is configured to use the character image as the teacher image, and update the parameters of the virtual dressing-up student model to be trained according to the image loss information between the student image and the teacher image.
[0098] In another exemplary embodiment, the image processing apparatus further includes:
[0099] The image quality difference acquisition module is configured to obtain the image quality difference between the teaching assistant image and the student image. If the image quality difference is positive, the character image is used as the teacher image, and the parameters of the virtual dressing student model to be trained are updated based on the image loss information between the student image and the teacher image.
[0100] In another exemplary embodiment, the teaching assistant image acquisition module includes:
[0101] The second clothing deformation sub-model calling unit is configured to call the second clothing deformation sub-model contained in the virtual dressing assistant model, and generate a deformed image of the human body of the clothing to be replaced that is adapted to the specified person based on the human body analysis results and the image features of the first clothing image; the second dressing generation sub-model calling unit is configured to call the second dressing generation sub-model contained in the virtual dressing model, and generate the assistant image based on the deformed image corresponding to the clothing to be replaced output by the second clothing deformation sub-model and other image areas in the character image except the area wearing the original clothing.
[0102] In another exemplary embodiment, the teaching assistant image acquisition module further includes:
[0103] The image area information acquisition unit is configured to call the second dressing generation sub-model contained in the virtual dressing model, and based on the human body analysis results, clear the area of the character image containing the original clothing worn by the specified character to obtain other image areas in the character image except the area wearing the original clothing.
[0104] In another exemplary embodiment, the parameter updating module includes:
[0105] An image loss value acquisition unit is configured to obtain an image loss value of a student image relative to a teacher image, where the image loss value includes at least one of a pixel distance loss function value, a perceptual loss function value, and an adversarial loss function value; a loss value summation unit is configured to perform a sum operation on at least one image loss value to obtain an image loss sum value of the student image relative to the teacher image; and a model parameter updating unit is configured to use the image loss sum value as image loss information between the student image and the teacher image to perform parameter updates on a virtual dress-up student model to be trained.
[0106] In another exemplary embodiment, the first outfit generation sub-model is composed of an encoder-decoder network and a residual network, and the residual network is used to normalize the connected upper layer network.
[0107] It should be noted that the apparatus provided in the above embodiment and the method provided in the above embodiment belong to the same concept, wherein the specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here.
[0108] An embodiment of the present application further provides an electronic device, including a processor and a memory, wherein the memory stores computer-readable instructions, which implement the image processing method as described above when executed by the processor.
[0109] Figure 8 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown.
[0110] It should be noted that Figure 8 The computer system 1600 of the electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.
[0111] like Figure 8 As shown, computer system 1600 includes a central processing unit (CPU) 1601, which can perform various appropriate actions and processes according to programs stored in read-only memory (ROM) 1602 or programs loaded from storage portion 1608 into random access memory (RAM) 1603, such as executing the methods described in the above embodiments. Various programs and data required for system operation are also stored in RAM 1603. CPU 1601, ROM 1602, and RAM 1603 are connected to each other via bus 1604. Input / output (I / O) interface 1605 is also connected to bus 1604.
[0112] The following components are connected to the I / O interface 1605: an input section 1606 including a keyboard, a mouse, and the like; an output section 1607 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 1608 including a hard disk; and a communication section 1609 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1609 performs communication processing via a network such as the Internet. A drive 1610 is also connected to the I / O interface 1605 as needed. Removable media 1611, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1610 as needed, so that computer programs read from the removable media can be installed in the storage section 1608 as needed.
[0113] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1609, and / or installed from a removable medium 1611. When the computer program is executed by the central processing unit (CPU) 1601, the various functions defined in the system of the present application are executed.
[0114] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0116] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0117] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image processing method described above. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device.
[0118] Another aspect of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image processing method provided in each of the above embodiments.
[0119] The above content is only a preferred exemplary embodiment of the present application and is not intended to limit the implementation scheme of the present application. Ordinary technicians in this field can easily make corresponding changes or modifications based on the main ideas and spirit of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection required by the claims.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a first image containing a target person and a second image containing a target garment; The first clothing deformation sub-model included in the virtual dress-up student model generates, based on image features corresponding to the first image and image features corresponding to the second image, appearance flow features for representing deformation of the target clothing adapted to the target person's body, and generates a deformed image of the target clothing adapted to the person's body based on the appearance flow features; generating a virtual dress-changing image by fusing the deformed image and the first image using a first dress-changing generation sub-model included in the virtual dress-changing student model, wherein the target person wears target clothing adapted to the human body in the virtual dress-changing image; The process of training the virtual dress-up student model includes: Invoking a virtual dressing-up teaching assistant model, inputting a human body analysis result corresponding to a character image containing a designated character and a first clothing image containing clothing to be changed into the virtual dressing-up teaching assistant model, and obtaining a teaching assistant image output by the virtual dressing-up teaching assistant model, wherein the designated character is wearing a deformed image adapted to the body of the designated character; Inputting a second clothing image containing the original clothing worn by the designated person in the person image and the teaching assistant image into a virtual dress-changing student model to be trained, thereby obtaining a student image output by the virtual dress-changing student model to be trained, wherein the designated person wears a deformed image of a human body adapted to the designated person in the teaching assistant image; The character image is used as a teacher image, and parameters of the virtual dress-up student model to be trained are updated according to image loss information between the student image and the teacher image.
2. The method according to claim 1, characterized in that Generating, based on the image features corresponding to the first image and the image features corresponding to the second image, appearance flow features for characterizing deformation of the target clothing caused by adapting to the body of the target person, includes: Using the first image as an input signal of a first image feature extraction model, and using the second image as an input signal of a second image feature extraction model, extracting multi-layer image features corresponding to the input signals through the first image feature extraction model and the second image feature extraction model; According to the multi-layer image features output by the first image feature extraction model and the second image feature extraction model, appearance flow features are extracted layer by layer, and the appearance flow features extracted from the last image feature layer are used as the final appearance flow features.
3. The method according to claim 2, characterized in that The extracting of appearance flow features layer by layer based on the multi-layer image features output by the first image feature extraction model and the second image feature extraction model includes: Extracting, at a first image feature layer, appearance flow features for characterizing deformation of the target clothing resulting from adaptation to the target person's body based on image features output by the first image feature extraction model and the second image feature extraction model; Each image feature layer after the first image feature layer optimizes the appearance flow features output by the previous image feature layer based on the image features output by the first image feature extraction model and the second image feature extraction model to obtain the appearance flow features corresponding to the current image feature layer.
4. The method according to claim 3, characterized in that Each image feature layer subsequent to the first image feature layer optimizes the appearance flow features output by the previous image feature layer according to the image features output by the first image feature extraction model and the second image feature extraction model, including: Performing upsampling processing on the appearance flow features output by the previous image feature layer to obtain upsampled features, and performing a first deformation processing on the image features of the second image output by the current image feature layer to obtain a first deformed feature; Performing correction processing on the first deformed feature based on the image feature of the first image output by the current image feature layer, and performing a first convolution calculation on the corrected feature obtained by the correction processing to obtain a first convolution feature; performing a second deformation process on the image features of the second image output by the current image feature layer according to the features obtained by concatenating the first convolution features and the up-sampled features to obtain second deformed features; A second convolution calculation is performed on the second deformed feature, and the second convolution feature obtained by the calculation is concatenated with the first convolution feature to obtain an appearance flow feature corresponding to the current image feature layer.
5. The method according to claim 2, characterized in that In the process of extracting appearance flow features layer by layer based on the multi-layer image features output by the first image feature extraction model and the second image feature extraction model, the appearance flow features are also extracted based on a second-order smoothness constraint condition preset for the linear correspondence between adjacent appearance flows.
6. The method according to claim 1, characterized in that Before updating the parameters of the virtual dress-up student model to be trained based on image loss information between the student image and the teacher image, the method further includes: Obtaining an image quality difference between the teaching assistant image and the student image; If the image quality difference is a positive value, the step of using the character image as the teacher image and updating the parameters of the virtual dress-up student model to be trained according to the image loss information between the student image and the teacher image is executed.
7. The method according to claim 1, characterized in that The calling of the virtual dressing-up teaching assistant model, inputting a human body analysis result corresponding to a human image containing a human figure and a first clothing image containing clothing to be changed into the virtual dressing-up teaching assistant model, and obtaining a teaching assistant image output by the virtual dressing-up teaching assistant model, includes: Invoking a second clothing deformation sub-model included in the virtual dressing assistant model to generate a deformed image of the clothing to be changed adapted to the body of the designated person based on the human body analysis result and image features of the first clothing image; The second dressing generation sub-model contained in the virtual dressing model is called, and the teaching assistant image is generated by fusing the deformed image corresponding to the clothing to be changed output by the second clothing deformation sub-model and the other image areas in the character image except the area wearing the original clothing.
8. The method according to claim 7, characterized in that The method further comprises: The second dressing generation sub-model contained in the virtual dressing model is called, and based on the human body analysis result, the area of the character image containing the original clothing worn by the specified character is cleared to obtain other image areas in the character image except the area wearing the original clothing.
9. The method according to claim 1, characterized in that The updating of parameters of the virtual dress-changing student model to be trained according to the image loss information between the student image and the teacher image includes: Obtaining an image loss value of the student image relative to the teacher image, where the image loss value includes at least one of a pixel distance loss function value, a perceptual loss function value, and an adversarial loss function value; performing a sum operation on the at least one image loss value to obtain a sum of image losses of the student image relative to the teacher image; The image loss sum value is used as the image loss information between the student image and the teacher image, and the parameters of the virtual dress-up student model to be trained are updated.
10. The method according to claim 1, characterized in that The first dress-changing generation sub-model consists of an encoder-decoder network and a residual network, and the residual network is used to normalize the connected upper-layer network.
11. An image processing device, characterized in that: The device comprises: An image acquisition module configured to acquire a first image containing a target person and a second image containing a target garment; an information generation module configured to generate, based on image features corresponding to the first image and image features corresponding to the second image, appearance flow features for characterizing the deformation of the target garment adapted to the human body of the target person, and to generate, based on the appearance flow features, a deformed image of the target garment adapted to the human body; the information generation module is further configured to generate a first garment deformation sub-model contained in the virtual dress-up student model; a virtual dressing module configured to generate a virtual dressing image based on the fusion of the deformed image and the first image, wherein the target person wears target clothing adapted to the human body in the virtual dressing image; the virtual dressing module is further configured to generate a first dressing sub-model contained in the virtual dressing student model; a teaching assistant image acquisition module configured to call a virtual dressing-up teaching assistant model, input a human body analysis result corresponding to a character image containing a designated character, and a first clothing image containing clothing to be changed into, into the virtual dressing-up teaching assistant model, and obtain a teaching assistant image output by the virtual dressing-up teaching assistant model, wherein the designated character is wearing a deformed image adapted to the designated character's body; a student image acquisition module configured to input a second clothing image containing the original clothing worn by the designated person in the person image and the teaching assistant image into a virtual dress-changing student model to be trained, and obtain a student image output by the virtual dress-changing student model to be trained, wherein the designated person in the student image is wearing a deformed image of a human body adapted to the designated person in the teaching assistant image; The parameter updating module is configured to use the character image as a teacher image and update the parameters of the virtual dress-up student model to be trained according to the image loss information between the student image and the teacher image.
12. An electronic device, characterized in that: include: a memory storing computer-readable instructions; The processor reads the computer-readable instructions stored in the memory to execute the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the method according to any one of claims 1 to 10.
14. A computer program product, characterized in that comprising computer instructions stored in a computer-readable storage medium; A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method according to any one of claims 1 to 10.