Image processing method, device, storage medium, equipment and product

By using feature replacement and makeup transfer models, the problem of poor makeup detail transfer in existing face-swapping technologies has been solved, achieving a more realistic and natural face-swapping image effect.

CN115147261BActive Publication Date: 2025-11-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210537076.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2025-11-25
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

In existing face-swapping technologies, the transfer of makeup details is not good, resulting in blurry makeup and poor image quality in the swapped images.

Method used

Image processing methods are used to transfer the identity features of the target object and the makeup features of the template object to the initial feature-replaced image through a feature transfer model. The makeup transfer model is then used to preserve the makeup features of the template object, thereby improving the realism of the image.

Benefits of technology

It achieves complete preservation of the makeup features of the template object in face-swapped images, improving the realism and quality of the images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147261B_ABST
    Figure CN115147261B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an image processing method, device, storage medium, equipment and product. The image processing method obtains a target object image and a template object image, performs feature replacement processing on the template object image according to the target object image to obtain an initial feature replacement image, preliminarily replaces an object identity feature in the template object, and then migrates a makeup feature in the template object to a virtual object contained in the initial feature replacement image obtained through the feature replacement processing, so as to ensure that the virtual object after makeup migration retains complete makeup features in the template object, and make the virtual object in the obtained target feature replacement image more real and natural.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to an image processing method and device, a computer readable storage medium, an electronic device and a computer program product. BACKGROUND

[0002] With the development of artificial intelligence and computer technology, image synthesis technology such as face swapping technology has emerged. Face swapping refers to replacing a face region in a source image with a face region in a template image to change the identity feature of a target image. Face swapping technology has many application scenarios, for example, it can be applied to film and television character production, game character design, virtual image and privacy protection scenarios.

[0003] At present, common face swapping methods include: using three-dimensional (3-Dimension, 3D) modeling technology to perform three-dimensional face reconstruction on a source image and a template image to obtain a new face three-dimensional model, and generate a face swapping target image; or, for a specified face swapping object, a large number of face images containing the face swapping object are obtained, a neural network model is trained, and the trained model is used for face swapping.

[0004] However, the face swapping target image obtained by using the method of the prior art is not real and natural, and the face swapping effect is poor. SUMMARY

[0005] To solve the above technical problems, embodiments of the present application provide an image processing method, device, computer readable storage medium, electronic device and computer program product.

[0006] According to an aspect of an embodiment of the present application, an image processing method is provided, including: obtaining a target object image containing a target object, and obtaining a template object image containing a template object; performing feature replacement processing on the template object image according to the target object image to obtain an initial feature replacement image; wherein the initial feature replacement image contains a virtual object, and the virtual object has an object identity feature of the target object and an object additional attribute feature of the template object; migrating a makeup feature of the template object to the virtual object contained in the initial feature replacement image to obtain a target feature replacement image.

[0007] According to an aspect of some embodiments of the present application, an image processing apparatus is provided, comprising: an image acquisition module configured to acquire a target object image containing a target object, and acquire a template object image containing a template object; a feature replacement module configured to perform feature replacement processing on the template object image according to the target object image to obtain an initial feature replacement image; wherein the initial feature replacement image contains a virtual object, and the virtual object has object identity features of the target object and object additional attribute features of the template object; a makeup transfer module configured to transfer makeup features of the template object to the virtual object contained in the initial feature replacement image to obtain a target feature replacement image.

[0008] According to an aspect of some embodiments of the present application, a computer readable storage medium having computer readable instructions stored thereon is provided, when the computer readable instructions are executed by a processor of a computer, the computer is caused to perform the image processing method as described above.

[0009] According to an aspect of some embodiments of the present application, an electronic device is provided, comprising: a processor; and a memory for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the image processing method as described above.

[0010] According to an aspect of some embodiments of the present application, a computer program product is also provided, the computer program product comprising computer instructions for implementing the image processing method as described above when executed by a processor.

[0011] In the technical solutions provided in the embodiments of the present application, the target object image is acquired, and the template object image is acquired, then the feature replacement processing is performed on the template object image according to the target object image to obtain the initial feature replacement image, so as to preliminarily replace the object identity features in the template object, and then the makeup features in the template object are transferred to the virtual object contained in the initial feature replacement image obtained by the feature replacement processing, so as to ensure that the virtual object after the makeup transfer retains the complete makeup features in the template object, and make the virtual object in the obtained target feature replacement image more real and natural.

[0012] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0013] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments consistent with the application and, together with the description, further serve to explain the principles of the application. It is to be understood that the drawings are only schematic, and that they do not purport to be to scale with respect to one another. The embodiments will be described with reference to the drawings in which:

[0014] Figure 1 is a comparison chart of image processing effects in the prior art and the image processing effects of the embodiments involved in the present application;

[0015] Figure 2 is a schematic diagram of the implementation environment involved in the present application;

[0016] Figure 3 is a flowchart of an image processing method shown by an exemplary embodiment of the present application;

[0017] Figure 4 is a schematic diagram of obtaining a target object image and a template object image shown by an exemplary embodiment of the present application;

[0018] Figure 5 is a schematic diagram of an image processing method shown by an exemplary embodiment of the present application;

[0019] Figure 6 is a flowchart of an image processing method shown by another exemplary embodiment of the present application;

[0020] Figure 7 is a schematic diagram of a makeup transfer model shown by an exemplary embodiment of the present application;

[0021] Figure 8 is a schematic diagram of a decoding network layer shown by an exemplary embodiment of the present application;

[0022] Figure 9 is a schematic diagram of obtaining a training sample triple shown by an exemplary embodiment of the present application;

[0023] Figure 10 is a schematic diagram of adding makeup materials to an initial template object sample image according to a makeup material library shown by an exemplary embodiment of the present application;

[0024] Figure 11 is a schematic diagram of a training process of a makeup transfer model shown by an exemplary embodiment of the present application;

[0025] Figure 12 is a flowchart of an image processing method shown by another exemplary embodiment of the present application;

[0026] Figure 13is a block diagram of an image processing apparatus according to an example embodiment of the present application;

[0027] Figure 14 A structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. DETAILED DESCRIPTION

[0028] The example embodiments will be described in detail with reference to the drawings, of which examples are shown. In the following description, the same numbers are used to denote the same or similar elements unless otherwise described. The embodiments described in the following example embodiments do not represent all the implementations in accordance with this application. Rather, they are merely examples in accordance with some aspects of this application as detailed in the appended claims.

[0029] The block diagrams shown in the drawings are merely functional entities, and do not necessarily have to correspond to physically independent entities. That is, the functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0030] The flowcharts shown in the drawings are merely illustrative, and do not necessarily include all the contents and operations / steps, nor do they have to be executed in the order described. For example, some operations / steps can be further divided, and some operations / steps can be combined or partially combined, so the actual execution order can be changed depending on the actual situation.

[0031] In the present application, "a plurality of" means two or more. The association relationship of "and / or" describes the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally means that the associated objects before and after are in an "or" relationship.

[0032] The following briefly introduces the technologies that can be used in the embodiments of the present application.

[0033] Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0034] Artificial intelligence technology is a comprehensive discipline involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0035] Computer vision (CV) is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to identify and measure targets, and further to do image processing, so that the computer processing becomes more suitable for human eye observation or image transmission to instrument detection. As a scientific discipline, computer vision researches related theories and technologies, trying to establish artificial intelligence systems that can obtain information from images or multidimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. It also includes common face recognition, fingerprint recognition and other biometric identification technologies.

[0036] Machine learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. Its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning.

[0037] With the research and progress of artificial intelligence technology, artificial intelligence technology has been researched and applied in many fields, such as smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned vehicles, autonomous vehicles, drones, robots, smart medical care, smart customer service, etc. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0038] In the prior art, when performing feature replacement processing on objects in an image, such as replacing a face in an image, the makeup details cannot be well transferred, and the resolution of the transferred image is low. For example Figure 1As shown, when the makeup features of the template object contained in the template object image are migrated to the virtual object contained in the initial feature replacement image using the prior art, the makeup migration image a obtained has problems such as makeup blur and poor image quality.

[0039] Therefore, in order to completely retain the makeup features in the template object image in the image after face replacement when the template object in the template object image has makeup features, and improve the quality of the image after face replacement, an embodiment of the present application provides an image processing method, device, computer readable storage medium, electronic device and computer program product to obtain a face replacement image as shown in makeup migration image b, which completely retains the makeup features in the template object image, so that the image after face replacement is more realistic and natural, and the image quality is improved. Figure 1

[0040] The image processing method provided by the embodiment of the present application will be described below. Figure 2 is a schematic diagram of an implementation environment of the image processing method in the present application. As shown in Figure 2 The implementation environment includes a terminal 210 and a server 220. The terminal 210 and the server 220 can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.

[0041] The terminal 210 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 210 can generally refer to one of a plurality of terminals, and the embodiment is only exemplified by the terminal 210. Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminals can be only one, or the above terminals can be dozens or hundreds, or more, and at this time, the implementation environment of the above image processing method also includes other terminals. The number and type of the terminal are not limited in the embodiment of the present application.

[0042] The server 220 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms, etc. The server 220 is used to provide background services for the application program running on the terminal 210.

[0043] ​Optionally, the wireless or wired networks described above use standard communications technologies and / or protocols. The networks typically carry Internet traffic, but can also include, without limitation, any combination of local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), mobile, wired or wireless networks, proprietary networks, or virtual private networks. In some embodiments, technologies and / or formats including, without limitation, Hyper Text Markup Language (HTML), Extensible Markup Language (XML), and others are used to represent data exchanged over the networks. In addition, conventional encryption technologies, such as the Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Networks (VPNs), Internet Protocol Security (IPsec), and others, can be used to encrypt all or some of the links between the nodes. In other embodiments, custom and / or proprietary data communications technologies and / or formats can be used instead of, or in addition to, the ones described above.

[0044] Optionally, the server 220 undertakes the primary image processing work, and the terminal 210 undertakes the secondary image processing work; or the server 220 undertakes the secondary image processing work, and the terminal 210 undertakes the primary image processing work; or the server 220 or the terminal 210 can undertake the image processing work individually.

[0045] Illustratively, the terminal 210 sends an image processing instruction to the server 220, and the server 220 receives the image processing instruction sent by the terminal 210 and acquires a target object image and a template object image; the server 220 performs feature replacement processing on the template object image according to the target object image to obtain an initial feature replacement image; and the server 220 migrates the makeup features of the template object to a virtual object contained in the initial feature replacement image to obtain a target feature replacement image.

[0046] Referring to Figure 3 , Figure 3 is a flowchart of an image processing method according to an example embodiment of the present application. The image processing method can be applied to Figure 2The illustrated implementation environment, and by the server 220 in the implementation environment. It should be understood that the method can also be applied to other exemplary implementation environments, and by the devices in other implementation environments, the embodiment does not limit the implementation environment to which the method is applied.

[0047] The image processing method proposed in the embodiment of the application will be described in detail below with the server as a specific execution subject.

[0048] As Figure 3 As shown in an exemplary embodiment, the image processing method at least includes steps S310 to S330, which are described in detail as follows:

[0049] Step S310, obtaining a target object image containing a target object, and obtaining a template object image containing a template object.

[0050] It should be noted that the target object image refers to the image of the target object; the template object image refers to the image of the template object.

[0051] In the embodiment of the application, the way of obtaining the target object image and the template object image can be different according to the specific scene. For example, the target object image or the template object image can be pre-stored in the memory of the computer device, so that the target object image or the template object image is obtained, that is, the target object image or the template object image stored in the memory of the computer device is obtained; the user can also directly input the target object image or the template object image, and when the computer device needs to obtain the target object image or the template object image, the image input operation of the user is received to obtain the target object image or the template object image; the computer device can also be connected with an image acquisition device, and the target object image or the template object image in the current field of view is collected in real time through the image acquisition device, or the video frame corresponding to the pre-stored video frame sequence is obtained, and the pre-stored video frame is taken as the target object image or the template object image. The computer device can also obtain the target object image and the template object image through other ways, which are not limited by the application.

[0052] It should be noted that in the specific embodiments of the application, the data related to the target object image, the template object image and the like are involved, and when the above embodiments of the application are applied to specific products or technologies, the permission or consent of the user needs to be obtained, and the collection, use and processing of the related data need to comply with the relevant laws, regulations and standards of the country and region.

[0053] The target object image and the template object image can be real images or synthesized images. For example, the target object image is a real person image captured by an image acquisition device, and the template object image is a virtual person image synthesized by a face image synthesis software; or the template object image is a real person image captured by an image acquisition device, and the target object image is a virtual person image synthesized by a face image synthesis software; or the template object image and the target object image are both real person images captured by an image acquisition device; or the template object image and the target object image are both virtual person images synthesized by a face image synthesis software. The embodiments are not limited in this regard.

[0054] The target object image or the template object image can be a user-specified image or a randomly selected image. For example, the server obtains filtering condition information of the template object image input by the user, and the filtering condition information includes, but is not limited to, information such as the gender and age of the template object in the template object image and the similarity between the target object and the template object. The filtering condition information is used to match a plurality of candidate face images, to obtain a matching value corresponding to each candidate face image, and to select the candidate face image with the highest matching value as the template object image. It can be understood that the target object image can also be filtered by the filtering condition information. The filtering condition information is used to obtain a target object image or a template object image that meets the user's demand, so as to improve the user experience.

[0055] For example, refer to Figure 4 , Figure 4 The schematic diagram for obtaining the target object image and the template object image is provided for the exemplary embodiments of the present application. As shown in Figure 4 , the server provides an image acquisition interface to the terminal, the image acquisition interface is provided with a target object image acquisition component and a template object image acquisition component, the user can trigger the image upload buttons corresponding to the target object image acquisition component and the template object image acquisition component to upload the target object image and the template object image, and then the user can trigger the submit button to send the target object image in the target object image acquisition component and the template object image in the template object image acquisition component to the server.

[0056] Optionally, one template object can correspond to one or more target objects. For example, if the target object image contains multiple face images, the user is prompted to select a target object, and one or more face images are selected as target objects according to the user's target object selection operation. For example, the user triggers the image upload button of the target object image acquisition component to upload a target object image to the server. The server performs face region recognition on the target object image, and the obtained face region recognition result includes face region A1, face region A2, and face region A3. The server returns the face region recognition result to the terminal, so that the terminal displays face region A1, face region A2, and face region A3 in the target object image to the user. The terminal monitors that the user has performed a target object selection operation on face region A2 and face region A3, and sends corresponding target object selection information to the server. The server selects the face images corresponding to face region A2 and face region A3 as target objects according to the target object selection information. It can be understood that when one template object corresponds to multiple target objects, feature replacement processing needs to be performed on each target object according to the template object, and a target feature replacement image corresponding to each target object is obtained.

[0057] Optionally, one target object can correspond to one or more template objects. For example, if the template object image contains multiple face images, the user is prompted to select a template object, and one or more face images are selected as template objects according to the user's template object selection operation. For example, the user triggers the image upload button of the template object image acquisition component to upload a template object image to the server. The server performs face region recognition on the template object image, and the obtained face region recognition result includes face region B1, face region B2, and face region B3. The server returns the face region recognition result to the terminal, so that the terminal displays face region B1, face region B2, and face region B3 in the template object image to the user. The terminal monitors that the user has performed a template object selection operation on face region B2 and face region B3, and sends corresponding template object selection information to the server. The server selects the face images corresponding to face region B2 and face region B3 as template objects according to the template object selection information. It can be understood that when one target object corresponds to multiple template objects, multiple feature replacement processes need to be performed on the target object according to each template object, and a target feature replacement image corresponding to each template object is obtained.

[0058] In the embodiments of the present application, the image processing method can further include a step of pre-processing the obtained target object image or template object image. Illustratively, pre-processing the obtained target object image or template object image can include noise removal, brightness enhancement, etc. of the target object image or template object image. The noise removal of the target object image or template object image can filter out the noise and color spots in the image to be processed by using a noise reduction algorithm. The brightness enhancement of the target object image or template object image can adjust the RGB color distribution, replace the brightness extraction algorithm, perform sharpening processing, enhance the contrast, and enhance the edge, etc. The pre-processing of the obtained target object image or template object image can avoid errors in subsequent processing caused by defects in the target object image or template object image itself.

[0059] In step S320, the template object image is processed according to the target object image to obtain an initial feature replacement image. The initial feature replacement image contains a virtual object, and the virtual object has the object identity feature of the target object and the object additional attribute feature of the template object.

[0060] It should be noted that the image features of a person include the object identity feature and the object additional attribute feature. The object identity feature refers to the key features in the face of a person, such as the eyebrows, eyes, ears, nose, mouth, etc. that affect the appearance of the person. The object additional attribute feature refers to other features in the face of a person, such as the hairstyle, the posture of the person, the expression of the person, the decorations, etc. that do not affect the appearance of the person.

[0061] The purpose of the feature replacement processing is to replace the object identity feature of the template object with the object identity feature of the target object to obtain a virtual object containing the object identity feature of the target object and the object additional attribute feature of the template object.

[0062] In the embodiments of the present application, the feature replacement model trained can be used to perform feature replacement processing on the target object image and the template object image. For example, a feature replacement model based on a generative adversarial network (GAN) algorithm can replace the object identity features in the template object with the object identity features in the target object. Generally, the GAN algorithm can use an identity encoder to extract the object identity features of the target object, which can include the shape of the eyes, the distance between the mouth and the eyes, the bending program of the mouth, and the like. At the same time, an attribute extractor is used to extract the object additional attribute features of the template object, such as the pose, contour, facial expression, hairstyle, scene lighting, and the like of the face. Then, the object identity features of the target object and the object additional attribute features of the template object are input into the feature replacement model, the feature replacement model replaces the object identity features in the template object with the object identity features of the target object, and retains the object additional attribute features of the template object, to obtain an initial feature replacement image output by the feature replacement model. The feature replacement processing can also be performed on the target object image and the template object image by an image fusion method. For example, the target object region and the template object region are obtained, the object identity features in the target object region are affine transformed into the template object region, and the affine transformed object identity features of the target object and the object identity features of the template object are fused, such as Poisson fusion, alpha fusion, and the like, to obtain an initial feature replacement image. It can be understood that the specific implementation of face replacement can be flexibly selected according to the actual application scenario, and the present application does not limit this.

[0063] In step S330, the makeup features of the template object are migrated to the virtual object contained in the initial feature replacement image, to obtain a target feature replacement image.

[0064] It should be noted that the makeup features refer to the external appearance performance formed by a person through a certain dressing modification, for example, a person uses cosmetics and tools to render, draw, arrange, enhance the three-dimensional impression, adjust the shape and color, and the like of the eyebrows, eyes, ears, nose, mouth, and the like of the face, so as to achieve the purpose of beautifying the visual experience.

[0065] After obtaining the initial feature-replaced image, the virtual object in the image possesses the object identity features of the target object and the object-attached attribute features of the template object. However, when the template object contains makeup features, since makeup features are generally attached to object identity features (e.g., eye makeup features are attached to the eye features in the object identity features), replacing the object identity features of the template object with those of the target object will cause the target object's object identity features to overwrite the makeup features in the template object. Therefore, it is difficult to retain the makeup features of the template object in the resulting virtual object. Since makeup features do not affect a person's appearance, it is necessary to retain the makeup features of the template object in the virtual object to improve the accuracy of feature replacement processing.

[0066] In this embodiment, a trained makeup transfer model can be used to transfer makeup features of a template object to a virtual object contained in an initial feature-replacement image. For example, the makeup transfer model extracts makeup features from the template object and virtual object features from the virtual object. Makeup features may include eye makeup features, lip makeup features, cheek makeup features, etc. The makeup transfer model calculates based on the makeup features and virtual object features, transferring the makeup features from the template object to the virtual object to obtain the target feature-replacement image output by the makeup transfer model. Alternatively, image fusion methods can be used to transfer makeup features from the template object to the virtual object contained in the initial feature-replacement image. For example, regions containing makeup features in the template object are extracted to obtain makeup features for multiple parts, such as makeup features for the upper eyelashes, lower eyelashes, and lips. Then, the makeup features of each part are aligned with the corresponding parts of the virtual object using keypoints to transfer the makeup features from the template object to the virtual object, obtaining the target feature-replacement image. It is understood that the specific implementation method of makeup transfer can be flexibly selected according to the actual application scenario, and this application does not impose any limitations on this.

[0067] like Figure 5 As shown, after performing feature replacement processing on the template object image based on the target object image, an initial feature-replaced image is obtained. Then, makeup transfer is performed on the initial feature-replaced image based on the template object image to obtain the target feature-replaced image. This target feature-replaced image can better preserve the makeup features in the template object image while reflecting the object identity features of the target object in the target object image, making the obtained target feature-replaced image have a more realistic face-swapping effect.

[0068] Since the makeup features generally depend on the object identity features, when the object identity features of the template object are replaced by the object identity features of the target object, the object identity features of the target object will cover the makeup features in the template object, so that it is difficult to retain the makeup features in the template object in the resulting virtual object, and thus the face replacement effect is reduced. Therefore, in order to improve the face replacement effect, the image processing method provided in the embodiments of the present application performs feature replacement processing on the template object image according to the target object image, and then migrates the makeup features in the template object to the virtual object contained in the initial feature replacement image obtained by the feature replacement processing, so as to ensure that the virtual object after makeup migration retains the complete makeup features in the template object, so that the resulting target feature replacement image is more accurate. In addition, the virtual object before makeup migration and the template object have the same object additional attribute features except the makeup features, that is, the virtual object before makeup migration and the template object have high similarity, so as to facilitate the subsequent makeup migration process, and improve the accuracy of makeup migration.

[0069] Referring to Figure 6 , Figure 6 is a flowchart of an image processing method according to another exemplary embodiment. As shown in Figure 6 , in an exemplary embodiment, the process of migrating the makeup features of the template object to the virtual object contained in the initial feature replacement image to obtain the target feature replacement image in step S330 can include the following steps:

[0070] In step S331, the template object image is input into the first encoding network contained in the trained makeup migration model to obtain the makeup features of the template object output by the first encoding network.

[0071] In the embodiments of the present application, the first encoding network contained in the makeup migration model is used to extract features of the template object to obtain the makeup features output by the first encoding network. The makeup features refer to the results obtained by vectorizing the input template object, and the vectorization refers to representing the template object input into the first encoding network by using feature vectors.

[0072] In step S332, the initial feature replacement image is input into the second encoding network and the multi-layer perceptron contained in the makeup migration model to obtain the virtual object features of the virtual object output by the second encoding network and the style features of the virtual object output by the multi-layer perceptron.

[0073] In the embodiments of the present application, the second encoding network included in the makeup transfer model performs feature extraction on the virtual object in the initial feature replacement image to obtain virtual object features output by the second encoding network. The virtual object features refer to the results obtained by vectorizing the input virtual object. At the same time, the multilayer perceptron included in the makeup transfer model performs style extraction on the virtual object in the initial feature replacement image to obtain style features of the virtual object output by the multilayer perceptron. For example, as shown in FIG. 8, the virtual object is identified to obtain face part elements such as eye part elements and mouth part elements included in the face region of the virtual object, so as to extract part element style features corresponding to each face part element, and then obtain the style features of the virtual object according to the part element style features. Figure 5

[0074] Optionally, in the embodiments of the present application, the multilayer perceptron (MLP) includes but is not limited to an artificial neural network (ANN), and in addition to an input layer and an output layer, at least one hidden layer is further included, that is, the multilayer perceptron is at least a three-layer structure, and full connection is performed between layers. The input layer is the bottommost layer of the multilayer perceptron, the middle layer is the hidden layer, and the last is the output layer, which maps multiple input data sets to a single output data set. The multilayer perceptron is used to convert the corresponding latent factors of the virtual object into an intermediate latent space to obtain the style features of the virtual object, so that the style features include multiple independent features, so that the decoding network can more easily perform rendering, and at the same time, avoid the combination of features that do not exist in the training data.

[0075] In step S333, the makeup features of the template object, the virtual object features of the virtual object, and the style features of the virtual object are input into the decoding network included in the makeup transfer model to obtain a target feature replacement image output by the decoding network.

[0076] The decoding network is used to fuse the information included in the style features of the virtual object with the makeup features of the template object and the virtual object features of the virtual object, and finally generate a target feature replacement image after makeup transfer.

[0077] In the prior art, when performing face replacement on an image, the makeup details cannot be well transferred, and the resolution of the transferred image is low, the makeup is blurred, and the image quality is poor.

[0078] ​Therefore, in order to improve the clarity of the image after makeup transfer and retain more makeup details, this application adopts the above-mentioned makeup transfer model, which uses a multilayer perceptron to perform nonlinear mapping on the latent factors corresponding to the virtual object to obtain a more error-correcting intermediate latent space. The intermediate latent space is then used as style information to act on the spatial data, thereby making the target feature replacement image after makeup transfer clearer and retaining more refined makeup features of the template object.

[0079] Optional, please refer to Figure 7 , Figure 7 This is a schematic diagram of a makeup transfer model provided for an exemplary embodiment of this application. Figure 7 As shown, the makeup transfer model includes a first encoding network, a second encoding network, a multilayer perceptron, and a decoding network. The first and second encoding networks each include n encoding network layers, and the decoding network includes n decoding network layers. The output of each encoding network layer in the first and second encoding networks serves as the input to the next encoding network layer and the corresponding decoding network layer. Here, n is an integer greater than 2.

[0080] like Figure 7 As shown, the feature map extracted by each encoding network layer is output to the next encoding network layer for processing, and also needs to be output to the corresponding decoding network layer of the decoding network for processing. It should be noted that the corresponding decoding network layer here refers to a decoding network layer whose size matches that of the currently output feature map. For example, if the size of the currently output feature map is 32*32*512, then the corresponding decoding network layer in the decoding network is a decoding network layer capable of processing feature maps of size 32*32*512.

[0081] Furthermore, after each decoding network layer of the decoding network obtains the makeup features from the input of the first encoding network layer, the virtual object features from the input of the second encoding network layer, and the style features from the input of the multilayer perceptron, it decodes and synthesizes the makeup features, virtual object features, and style features, and outputs the decoding result to the next decoding network layer. This process continues until the last decoding network layer outputs the target features to replace the image.

[0082] In some embodiments, the step S333 inputs the makeup features of the template object, the virtual object features of the virtual object, and the style features of the virtual object into a decoding network included in the makeup transfer model to obtain a target feature replacement image output by the decoding network, and the process includes: adjusting convolution weights of the decoding network according to a preset scaling ratio to obtain scaled convolution weights; normalizing the scaled convolution weights to obtain a new decoding network; and performing decoding calculation on the makeup features of the template object, the virtual object features of the virtual object, and the style features of the virtual object according to the new decoding network to obtain the target feature replacement image output by the decoding network.

[0083] For example, please refer to Figure 8 , Figure 8 The schematic diagram of the decoding network layer provided for the exemplary embodiments of the present application is shown in FIG. 2. As shown in FIG. 2, the decoding network layer is composed of an affine transformation template A, a modulation template Mod-Demod, an up-sampling module Upsample, a plurality of convolution modules Conv, etc. Figure 8

[0084] Among them, the learnable affine transformation template A can be composed of a fully connected layer; the up-sampling module Upsample can use deconvolution for up-sampling operation, represents the makeup features output by the encoding network layer in the first encoding network, represents the virtual object features output by the encoding network layer in the second encoding network, and w represents the style features of the virtual object output by the multi-layer perception, represents the output of the decoding network layer of the previous layer, and respectively pass through the newly added convolution network to obtain

[0085] Figure 8 represents the output of the decoding network layer of the current layer, and can be represented as:

[0086]

[0087] Among them, and have the same spatial resolution, and Concat represents concatenating the three features.

[0088] Mod in the modulation template is used to adjust the convolution weights, and the specific calculation method is shown in the following formula:

[0089] w′ ijk = s i · w ijk

[0090] Among them, w′ ijk represents the scaled convolution weights, and w represents the original convolution weights.​​ijk denotes the convolution weight before scaling, s i denotes the preset scaling ratio of the i-th style feature of the input, j denotes the j-th decoding network layer, and k is a convolution kernel.

[0091] Further, the Demod in the modulation template demodulates the weight of the convolution layer, normalizes the scaled convolution weight, aims to restore the output to the unit standard deviation, and obtains the weight of the new convolution layer. For specific calculation methods, refer to the following formula:

[0092]

[0093] wherein w" ijk denotes the weight of the new convolution layer, and the constant ∈ is added to avoid the denominator being 0.

[0094] In some embodiments, the training process of the makeup transfer model includes: obtaining a training sample triple, the training sample triple including a template object sample image, an initial feature replacement sample image, and a target feature replacement sample image; inputting the template object sample image and the initial feature replacement sample image into the makeup transfer model to obtain a predicted image output by the makeup transfer model; correcting network parameters of the makeup transfer model according to the predicted image and the target feature replacement sample image to obtain a trained makeup transfer model.

[0095] It should be noted that the template object sample image refers to a face image with makeup features, the initial feature replacement sample image refers to a face image with the additional attribute features of the object in the template object sample image, and the makeup features between the initial feature replacement sample image and the template object sample image are inconsistent, and the target feature replacement sample image refers to a face image with the additional attribute features of the object in the template object sample image and the makeup features. Among them, the training sample carries a sample label, and the sample label is used to indicate the real class information to which the training sample belongs.

[0096] In the embodiments of the present application, by inputting the template object sample image and the initial feature replacement sample image into the makeup transfer model, a predicted image output by the makeup transfer model is obtained, and then the target feature replacement sample image in the training sample triple is taken as the target output of the makeup transfer model, the difference between the predicted image and the target feature replacement sample image is calculated to determine whether the makeup transfer model training end condition is met. Wherein, the makeup transfer model training end condition includes any one of the following: the number of training times reaches the number threshold; the loss function converges; the loss function is less than the loss function threshold. The number threshold and the loss function threshold are set according to experience or flexibly adjusted according to the application scene, and the embodiments of the present application do not limit this.

[0097] In the embodiment of the present application, the purpose of the makeup transfer model is to transfer the complete makeup features contained in the template object image to the initial feature replacement image to obtain a target feature replacement image containing more makeup details. Since the initial template object image and the initial feature replacement image have the same object additional attribute features, the following method is used to obtain training sample triplets that are more suitable for the application scenario, which can specifically include the following steps: obtaining an initial template object sample image and a target object sample image; performing feature replacement processing on the initial template object sample image according to the target object sample image to obtain an initial feature replacement sample image; performing makeup material adding processing on the initial feature replacement sample image and the initial template object sample image according to a makeup material library to obtain a target feature replacement sample image and a template object sample image; and obtaining a training sample triplet according to the template object sample image, the initial feature replacement sample image, and the target feature replacement sample image.

[0098] It should be noted that the initial template object sample image and the target object sample image do not contain makeup features, and the object identity features of the initial template object sample image and the target object sample image are different; the makeup material library pre-stores a plurality of makeup materials, such as eye makeup materials, eyebrow makeup materials, and the like.

[0099] In the embodiment of the present application, the initial template object sample image is processed by feature replacement according to the target object sample image to obtain an initial feature replacement sample image. The specific implementation steps of the feature replacement processing can be referred to step S320 in Figure 3 . Details are not described here. By performing feature replacement processing on the initial template object sample image and the target object sample image that do not contain makeup features, the interference of makeup features on the feature replacement processing process is avoided, and more accurate training data can be obtained to improve the training effect of the subsequent makeup transfer model.

[0100] For example, please refer to Figure 9 , Figure 9 for a schematic diagram of obtaining a training sample triplet. As shown in Figure 9 , the initial template object sample image and the target object sample image are input into the feature replacement model for feature replacement processing to obtain an initial feature replacement sample image output by the feature replacement model. Then, the initial template object sample image and the initial feature replacement sample image are processed by makeup material adding processing according to the makeup materials in the makeup material library, so as to obtain a template object sample image corresponding to the initial template object sample image and a target feature replacement sample image corresponding to the initial feature replacement sample image by performing key point detection and makeup material adding on the initial template object sample image and the initial feature replacement sample image. Then, a training sample triplet is obtained according to the template object sample image, the initial feature replacement sample image, and the target feature replacement sample image.

[0101] As shown in Figure 10 , Figure 10 is a schematic diagram of makeup material adding processing on an initial template object sample image according to a makeup material library. As shown in Figure 10 , the makeup material obtained through the makeup material library includes eyelash material and eye shadow material, then key point detection is performed on the initial template object sample image, and alignment is performed according to the detected key points and the key points corresponding to the eyelash material and the eye shadow material, so as to add the eyelash material and the eye shadow material to the initial template object sample image, and thus obtain a template object sample image, in which the makeup features include the above-mentioned eyelash material and eye shadow material. It can be understood that the specific implementation of the initial feature replacement sample image and the initial template object sample image in the makeup material adding processing is consistent, and will not be repeated here.

[0102] Through the above-mentioned method of obtaining training sample triplets, a large number of training sample triplets can be obtained, and the quality of the training sample triplets is guaranteed, so as to obtain a more accurate makeup transfer model.

[0103] In some embodiments, in the above-mentioned exemplary embodiments, the network parameters of the makeup transfer model are corrected according to the predicted image and the target feature replacement sample image, including: inputting the predicted image and the target feature replacement sample image into the discriminative network to obtain a discriminative result output by the discriminative network; wherein the discriminative result is used to represent an image output as a predicted target in the predicted image and the target feature replacement sample image; taking the target feature replacement sample image as an actual target output, and calculating a loss function value according to the predicted target output and the actual target output; and correcting the network parameters of the makeup transfer model according to the loss function value.

[0104] It should be noted that the discriminative network is used to classify the input image.

[0105] As shown in Figure 11 , Figure 11 is a schematic diagram of the training process of the makeup transfer model. As shown in Figure 11 , the template object sample image is input into the first encoding network, and the initial feature replacement sample image is input into the second encoding network and the multi-layer perceptron, then the outputs of the first encoding network, the second encoding network and the multi-layer perceptron are input into the decoding network, so that the decoding network outputs the predicted image. Further, the predicted image and the target feature replacement sample image are input into the discriminative network to obtain a discriminative result output by the discriminative network, and then an adversarial loss function is obtained according to the discriminative result, which can be expressed as the following formula:

[0106]

[0107] wherein GT represents the target feature replacement sample image, input represents the initial feature replacement sample image, refer represents the template object sample image, min G max D represents the minimum maximum function of G and D, D(GT) represents the discrimination of the real target output, the closer the discrimination result of D(GT) is to 1, the better, G(input, refer) represents the predicted image output by the makeup transfer model, the closer the discrimination result D(G(input, refer)) of G(input, refer) is to 0, the better.

[0108] Further, a reconstruction loss function is obtained according to the predicted image and the target feature replacement sample image, and the reconstruction loss function can be represented by the following formula:

[0109] L rec = |G(input) - GT|1+ |LPIPS(G(input)) - LPIPS(GT)|1

[0110] wherein G(input) represents the predicted image output by the makeup transfer model, GT represents the target feature replacement sample image, and LPIPS is a perceptual loss function.

[0111] Then, the obtained loss function is as follows:

[0112] L = L GAN + L rec

[0113] The server obtains a loss function value according to the loss function, to determine whether the makeup transfer model reaches a training completion condition according to the loss function value, when the training completion condition is not reached, the loss function value is used to update the model parameters in the makeup transfer model in reverse, and the training step of the makeup transfer model is iterated until the training completion condition of the makeup transfer model is reached, and the trained makeup transfer model is obtained.

[0114] Wherein, when training the makeup transfer model, the first encoding network and the second encoding network can be trained by using the already trained decoding network and the discriminator network, and different learning rates are set, the learning rate is used to indicate the model parameter correction according to the loss function value, for example, the higher the learning rate, the more model parameters are considered when the model parameter correction is performed according to the loss function value, and the lower the learning rate, the less model parameters are considered when the model parameter correction is performed according to the loss function value. For example, the learning rates of the first encoding network, the second encoding network, the decoding network and the discriminator network are 100:100:10:1.

[0115] In some embodiments, please refer toFigure 12 , Figure 12 is a flow chart of an image processing method according to another exemplary embodiment. As shown in FIG. 32, the process of performing feature replacement processing on the template object image according to the target object image to obtain an initial feature replacement image in step S320 can include the following steps. Figure 12

[0116] In step S321, the target object image and the template object image are input into the trained feature replacement model to obtain the object identity feature of the target object in the target object image and the object additional attribute feature of the template object in the template object image output by the feature replacement model.

[0117] In the present exemplary embodiment, the target object image is input into the face key attribute extraction network of the feature replacement model, and the template object image is input into the face additional attribute extraction network of the feature replacement model. The face key attribute extraction network and the face additional attribute extraction network include a plurality of convolutional layers, so as to perform face key attribute extraction on the target object based on the plurality of convolutional layers in the face key attribute extraction network to obtain the object identity feature of the target object, and perform face additional attribute extraction on the template object based on the plurality of convolutional layers in the face additional attribute extraction network to obtain the object additional attribute feature of the template object.

[0118] For example, the feature map output by the previous convolutional layer in the face key attribute extraction network is used to input the next convolutional layer. Each convolutional layer can correspond to different resolution parameters, and each convolutional layer can generate a feature map corresponding to width, height and depth, wherein the width, height and depth are related to the size of the convolution kernel and the number of convolution kernels on the convolutional layer. Each convolutional layer is matched with a corresponding fully connected layer. Each fully connected layer projects and embeds the feature map output by the corresponding convolutional layer into the corresponding feature space to obtain the feature vector corresponding to each convolutional layer, and takes the feature vector output by the last convolutional layer as the object identity feature.

[0119] In some embodiments, inputting the target object image and the template object image into the trained feature replacement model to obtain the object identity feature of the target object in the target object image and the object additional attribute feature of the template object in the template object image output by the feature replacement model includes: identifying the target object included in the target object image to obtain a corresponding to-be-replaced region of the target object; extracting at least two sub-identity feature information from the to-be-replaced region, wherein the at least two sub-identity feature information is used to represent the object identity information at different positions in the to-be-replaced region; and fusing the at least two sub-identity feature information to obtain the object identity feature of the target object in the target object image.

[0120] ​Optionally, the face sub-regions include, but are not limited to, a face part type dimension, such as a face part type including, but not limited to, an ear type, an eye type, a mouth type, and the like.

[0121] In the embodiment, the target object is identified to obtain a to-be-replaced region corresponding to the target object, so as to extract face sub-region images of at least two face sub-regions from the to-be-replaced region, such as an eye image corresponding to an eye sub-region and an ear image corresponding to an ear sub-region. Then, the encoding features corresponding to each face sub-region image are obtained to obtain a plurality of sub-identity feature information, such as an eye encoding feature and an ear encoding feature. Further, the plurality of sub-identity feature information is spliced to obtain an object identity feature of the target object.

[0122] By respectively extracting features of the face sub-regions contained in the to-be-replaced region, the direct coupling relationship between the face image of the target object and the face image of the template object is removed, and the face changing effect is improved, so that the subsequent virtual object is more realistic.

[0123] In step S322, the object identity feature of the target object and the object additional attribute feature of the template object are fused to obtain a virtual object.

[0124] For example, the object identity feature of the target object and the object additional attribute feature of the template object are input into an image generation network in the feature replacement model, and after the object identity feature and the object additional attribute feature are sequentially fused by a plurality of cascaded feature fusion network layers in the image generation network, a virtual object is obtained, which includes the object identity feature of the target object and the object additional attribute feature of the template object.

[0125] When the face key attribute extraction network and the face additional attribute extraction network include a plurality of convolution layers, the feature map output by each convolution layer is taken as the input of the next convolution layer and the corresponding feature fusion network layer. The corresponding feature fusion network layer here refers to a feature fusion network layer matching the size of the current output feature map, that is, the input of each feature fusion network layer includes the output of the previous layer feature fusion network layer, the output of the corresponding convolution layer in the face key attribute extraction network, and the output of the corresponding convolution layer in the face additional attribute extraction network. The plurality of cascaded feature fusion network layers sequentially fuse the object identity features and the object additional attribute features at different levels to obtain a final virtual object, which includes not only the object identity feature of the target object, but also the object additional attribute feature of the template object.

[0126] In step S323, an initial feature replacement image is obtained according to the virtual object.

[0127] It can be understood that the initial feature replacement image includes a person content and a background content, the person content of the initial feature replacement image contains a virtual object, and the background content of the initial feature replacement image can be consistent with the background content of the target object image, can be consistent with the background content of the template object image, or can be other background content such as other background content specified by a user. The present application does not limit this.

[0128] Optionally, in the embodiment of the present application, the image processing method described above includes but is not limited to being applied in a video scene. For example, in a video conference scene, after detecting that a terminal where a conference object is located triggers a face replacement request, a user-specified target object image and a template object image carried in the face replacement request are obtained, and then the template object image is processed according to the target object image to obtain an initial feature replacement image. The template object image can be a video frame in a conference video obtained by the terminal where the conference object is located. Further, in order to improve the authenticity of the face replacement image, the makeup features of the template object in the template object image are obtained, which include but are not limited to nose makeup features, eye makeup features, mouth makeup features, etc., and the makeup features are migrated to the virtual object contained in the initial feature replacement image to obtain a target feature replacement image. Then, the image displayed on the current conference picture of the terminal where the conference object is located is processed according to the target feature replacement image determined above, so as to display the target feature replacement image, thereby realizing the continuous display of the target feature replacement image after face replacement in the video conference process of the conference object.

[0129] Optionally, in the embodiment of the present application, the image processing method described above includes but is not limited to being applied in a video production scene. For example, a video short includes multiple person objects, and a face replacement request containing a target object image and a template object image is generated according to a selection operation performed by a user on the multiple person objects, wherein at least one of the target object image and the template object image is a person object contained in the video short. Then, the template object image is processed according to the target object image to obtain an initial feature replacement image, and the initial feature replacement image contains a virtual object. Further, the makeup features of the template object in the template object image are obtained, which include but are not limited to nose makeup features, eye makeup features, mouth makeup features, etc., and the makeup features are migrated to the virtual object contained in the initial feature replacement image to obtain a target feature replacement image. The target feature replacement image contains a virtual object with complete makeup features of the template object. The virtual object in the target feature replacement image is used to process all video frames containing the target object in the video short to replace the target object in these video frames with the virtual object in the target feature replacement image, thereby obtaining a processed video short.

[0130] It can be seen that in the technical scheme provided in the embodiment of the present application, the target object image and the template object image are obtained, the template object image is processed for feature replacement according to the target object image, and an initial feature replacement image is obtained to preliminarily replace the object identity feature in the template object. Then, the makeup feature in the template object is migrated to the virtual object contained in the initial feature replacement image obtained by the feature replacement processing, so as to ensure that the virtual object after makeup migration retains the complete makeup feature in the template object, and the obtained target feature replacement image is more accurate.

[0131] Figure 13 is a block diagram of an image processing apparatus shown in an example embodiment of the present application. The image processing apparatus can be applied to Figure 1 the implementation environment shown in the figure. The image processing apparatus can also be applicable to other example implementation environments and be specifically configured in other devices, and the embodiment does not limit the implementation environment to which the apparatus is applicable.

[0132] As shown in Figure 13 the example image processing apparatus 1300 includes an image acquisition module 1310, a feature replacement module 1320, and a makeup migration module 1330. Specifically:

[0133] The image acquisition module 1310 is configured to acquire a target object image containing a target object, and acquire a template object image containing a template object.

[0134] The feature replacement module 1320 is configured to process the template object image for feature replacement according to the target object image to obtain an initial feature replacement image; wherein the initial feature replacement image contains a virtual object, and the virtual object has an object identity feature of the target object and an object additional attribute feature of the template object.

[0135] The makeup migration module 1330 is configured to migrate the makeup feature of the template object to the virtual object contained in the initial feature replacement image to obtain a target feature replacement image.

[0136] In the example image processing apparatus, in order to improve the face replacement effect, after the template object image is processed for feature replacement according to the target object image, the makeup feature in the template object is migrated to the virtual object contained in the initial feature replacement image obtained by the feature replacement processing, so as to ensure that the virtual object after makeup migration retains the complete makeup feature in the template object, and the obtained target feature replacement image is more accurate. Moreover, the virtual object before makeup migration and the template object have the same object additional attribute features except the makeup feature, that is, the virtual object before makeup migration and the template object have a high similarity, and thus the subsequent makeup migration process is facilitated, so as to improve the accuracy of makeup migration.

[0137] On the basis of the above exemplary embodiments, the makeup transfer module 1330 further includes a first encoding module, a second encoding module, and a decoding module. Specifically:

[0138] The first encoding module is configured to input the template object image into the first encoding network included in the trained makeup transfer model to obtain the makeup feature of the template object output by the first encoding network.

[0139] The second encoding module is configured to input the initial feature replacement image into the second encoding network and the multi-layer perceptron included in the makeup transfer model to obtain the virtual object feature of the virtual object output by the second encoding network and the style feature of the virtual object output by the multi-layer perceptron.

[0140] The decoding module is configured to input the makeup feature of the template object, the virtual object feature of the virtual object, and the style feature of the virtual object into the decoding network included in the makeup transfer model to obtain the target feature replacement image output by the decoding network.

[0141] In the exemplary image processing device, the multi-layer perceptron included in the makeup transfer model extracts the style of the virtual object in the initial feature replacement image to obtain the style feature of the virtual object output by the multi-layer perceptron, i.e., converts the latent factor corresponding to the virtual object into a more error-corrected intermediate latent space, which is the style feature of the virtual object. The style feature contains multiple independent features, so that the decoding network can more easily perform rendering, while avoiding feature combinations that do not exist in the training data, making the obtained target feature replacement image after makeup transfer clearer and retaining more detailed makeup features of the template object.

[0142] On the basis of the above exemplary embodiments, the decoding module further includes a scaling calculation module, a normalization calculation module, and a decoding subunit. Specifically:

[0143] The scaling calculation module is configured to adjust the convolution weights of the decoding network according to a preset scaling ratio to obtain scaled convolution weights.

[0144] The normalization calculation module is configured to normalize the scaled convolution weights to obtain a new decoding network.

[0145] The decoding subunit is configured to perform decoding calculation on the makeup feature of the template object, the virtual object feature of the virtual object, and the style feature of the virtual object according to the new decoding network to obtain the target feature replacement image output by the decoding network.

[0146] In the exemplary image processing device, the convolution weights in the decoding network are updated to better untangle the input features, thereby further improving the quality of the generated target feature replacement image.

[0147] On the basis of the above exemplary embodiments, the first encoding network and the second encoding network each comprise a plurality of encoding network layers, and the decoding network comprises a plurality of decoding network layers; the output of each encoding network layer in the first encoding network and the second encoding network is taken as the input of the corresponding next encoding network layer and the decoding network layer.

[0148] In the exemplary image processing apparatus, the first encoding network and the second encoding network each comprise a plurality of encoding network layers, and the initial feature replacement image is sequentially subjected to feature extraction to obtain the virtual object features of the virtual object, and the template object image is sequentially subjected to feature extraction to obtain the makeup features of the template object, so that the facial features at multiple levels can be fully extracted, and the input feature map is sequentially subjected to decoding processing by the plurality of decoding network layers comprised by the decoding network, so that the makeup features of the template object are more completely retained on the target feature replacement image output by the decoding network, and the quality of the target feature replacement image is improved.

[0149] On the basis of the above exemplary embodiments, the image processing apparatus 1300 further comprises a training sample acquisition module, a predicted image acquisition module, and a network parameter correction module, specifically:

[0150] The training sample acquisition module is configured to acquire a training sample triple, and the training sample triple comprises a template object sample image, an initial feature replacement sample image, and a target feature replacement sample image.

[0151] The predicted image acquisition module is configured to input the template object sample image and the initial feature replacement sample image into the makeup transfer model to obtain a predicted image output by the makeup transfer model.

[0152] The network parameter correction module is configured to correct the network parameters of the makeup transfer model according to the predicted image and the target feature replacement sample image to obtain a trained makeup transfer model.

[0153] In the exemplary image processing apparatus, the makeup transfer model is trained according to the training sample triple, the target feature replacement sample image is taken as the target output, and the network parameters of the makeup transfer model are updated according to the difference between the predicted image and the target feature replacement sample image to obtain a trained makeup transfer model, so that the makeup transfer model learns the ability of makeup transfer.

[0154] On the basis of the above exemplary embodiments, the training sample acquisition module comprises an image acquisition unit, a face replacement unit, a makeup material adding unit, and a training sample confirmation unit. Specifically:

[0155] The image acquisition unit is configured to acquire an initial template object sample image and a target object sample image.

[0156] The face changing unit is configured to perform feature replacement processing on the initial template object sample image according to the target object sample image, to obtain an initial feature replacement sample image.

[0157] The makeup material adding unit is configured to perform makeup material adding processing on the initial feature replacement sample image and the initial template object sample image according to a makeup material library, to obtain a target feature replacement sample image and a template object sample image.

[0158] The training sample confirming unit is configured to obtain a training sample triple according to the template object sample image, the initial feature replacement sample image, and the target feature replacement sample image.

[0159] In the exemplary image processing device, a large number of training sample triples can be obtained by performing feature replacement processing on the initial template object sample image according to the target object sample image, and performing makeup material adding processing on the initial feature replacement sample image and the initial template object sample image, and the quality of the training sample triples is ensured, so that a more accurate makeup transfer model can be obtained.

[0160] On the basis of the above exemplary embodiments, the network parameter correction module includes a discrimination module, a loss function value calculation module, and a parameter correction module. Specifically:

[0161] The discrimination module is configured to input the predicted image and the target feature replacement sample image into a discrimination network to obtain a discrimination result output by the discrimination network; wherein the discrimination result is used to represent an image that is a predicted target output in the predicted image and the target feature replacement sample image.

[0162] The loss function value calculation module is configured to take the target feature replacement sample image as an actual target output, and calculate a loss function value according to the predicted target output and the actual target output.

[0163] The parameter correction module is configured to correct the network parameters of the makeup transfer model according to the loss function value.

[0164] In the exemplary image processing device, the authenticity of the input predicted image and the target feature replacement sample image is judged by the discrimination network to verify whether the predicted image can be judged by the discrimination network as an actual target output, and then the loss function value is calculated based on the discrimination structure, and the network parameters of the makeup transfer model are adjusted in reverse according to the loss function value, so that the obtained makeup transfer model is more accurate.

[0165] On the basis of the above exemplary embodiments, the feature replacement module 1320 includes a feature extraction module, a feature fusion module, and an initial feature replacement image acquisition module. Specifically:

[0166] The feature extraction module is configured to input the target object image and the template object image into the trained feature replacement model to obtain object identity features of the target object in the target object image and object additional attribute features of the template object in the template object image output by the feature replacement model.

[0167] The feature fusion module is configured to perform fusion processing on the object identity features of the target object and the object additional attribute features of the template object to obtain the virtual object.

[0168] The initial feature replacement image acquisition module is configured to obtain an initial feature replacement image according to the virtual object.

[0169] In the exemplary image processing device, the object identity features are obtained by performing feature extraction on the target object in the target object image, the object additional attribute features are obtained by performing feature extraction on the template object in the template object image, and then the object identity features and the object additional attribute features are fused to obtain the virtual object, so that the virtual object contains the object identity features of the target object and the object additional attribute features of the template object.

[0170] Based on the above exemplary embodiment, the feature extraction module includes a region-to-be-replaced identification module, a sub-region information extraction module, and an information fusion module. Specifically,

[0171] The region-to-be-replaced identification module is configured to identify the target object included in the target object image to obtain a region-to-be-replaced corresponding to the target object.

[0172] The sub-region information extraction module is configured to extract at least two sub-identity feature information from the region-to-be-replaced, wherein the at least two sub-identity feature information is used to represent object identity information at different positions in the region-to-be-replaced.

[0173] The information fusion module is configured to fuse the at least two sub-identity feature information to obtain the object identity features of the target object in the target object image.

[0174] In the exemplary image processing device, the target object is divided into multiple sub-regions, feature extraction is performed on the multiple sub-regions respectively to obtain multiple sub-identity feature information, and then the multiple sub-identity feature information is fused to obtain the object identity features of the target object, so that the direct coupling relationship between the face image of the target object and the face image of the template object is removed, and the face replacement effect is improved, and the virtual object obtained subsequently is more realistic.

[0175] It should be noted that the image processing apparatus provided in the above embodiments and the image processing method provided in the above embodiments belong to the same concept, and the specific manner in which each module and unit performs operations has been described in the method embodiments in detail, which will not be repeated here. The image processing apparatus provided in the above embodiments can be divided into different functional modules according to the needs in actual application, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above, which is not limited here.

[0176] Embodiments of the present application also provide an electronic device, comprising: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the electronic device implements the image processing method provided in each of the above embodiments.

[0177] Figure 14 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. It should be noted that, Figure 14 The computer system 1400 of the electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present application.

[0178] As Figure 14 shown, the computer system 1400 includes a central processing unit (CPU) 1401, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1402 or programs loaded from a storage portion 1408 into a random access memory (RAM) 1403, such as performing the methods described in the above embodiments. In the RAM 1403, various programs and data required for system operation are also stored. The CPU 1401, the ROM 1402, and the RAM 1403 are connected to each other through a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.

[0179] The following components are connected to the I / O interface 1405: an input portion 1406 including input devices such as a keyboard and mouse; an output portion 1407 including output devices such as a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), and a speaker; a storage portion 1408 including a hard disk; and a communication portion 1409 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication portion 1409 performs communication processing via a network such as the Internet. A drive 1410 is also connected to the I / O interface 1405 as necessary. A removable media 1411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 1410 as necessary, so that a computer program read therefrom is installed in the storage portion 1408 as necessary.

[0180] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present application. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing a computer program for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication portion 1409, and / or installed from the removable media 1411. When the computer program is executed by the Central Processing Unit (CPU) 1401, various functions defined in the system of the present application are executed.

[0181] It should be noted that the computer readable medium shown in the embodiments of the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable signal medium can include a data signal propagating in the baseband or as a carrier wave part of a signal propagating in the baseband, in which the computer readable computer program is carried. Such a propagating data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit programs for use by or in connection with an instruction execution system, device or apparatus. The computer program contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0182] The units described in the embodiments of the present application can be implemented in software, or can be implemented in hardware, and the described units can also be arranged in a processor. In some cases, the names of these units do not constitute a limitation on the units themselves.

[0183] Another aspect of the present application also provides a computer readable storage medium having a computer program stored thereon, which is executed by a processor to implement the image processing method as described above. The computer readable storage medium can be included in the electronic device described in the above embodiments, or can exist separately and not be assembled into the electronic device.

[0184] Another aspect of the present application also provides a computer program product or computer program, which includes computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the computer device execute the image processing method provided in each of the above embodiments.

[0185] The above merely provides preferred exemplary embodiments of the present application, and is not intended to limit the implementation of the present application. Based on the main concept and spirit of the present application, the person skilled in the art can easily make corresponding changes or modifications, and the protection scope of the present application should be subject to the protection scope required by the claims.

Claims

1. An image processing method, characterized by, The method comprises: obtaining a target object image containing a target object, and obtaining a template object image containing a template object; performing feature replacement processing on the template object image according to the target object image to obtain an initial feature replacement image; wherein the initial feature replacement image contains a virtual object, and the virtual object has object identity features of the target object and object additional attribute features of the template object; inputting the template object image into a first encoding network included in a makeup transfer model trained to obtain makeup features of the template object output by the first encoding network; inputting the initial feature replacement image into a second encoding network and a multi-layer perceptron included in the makeup transfer model to obtain virtual object features of the virtual object output by the second encoding network and style features of the virtual object output by the multi-layer perceptron; inputting the makeup features of the template object, the virtual object features of the virtual object, and the style features of the virtual object into a decoding network included in the makeup transfer model to obtain a target feature replacement image output by the decoding network.

2. The method of claim 1, wherein, The inputting of the makeup features of the template object, the virtual object features of the virtual object, and the style features of the virtual object into the decoding network included in the makeup transfer model to obtain the target feature replacement image output by the decoding network comprises: adjusting convolution weights of the decoding network according to a preset scaling ratio to obtain scaled convolution weights; normalizing the scaled convolution weights to obtain a new decoding network; performing decoding calculation on the makeup features of the template object, the virtual object features of the virtual object, and the style features of the virtual object according to the new decoding network to obtain the target feature replacement image output by the decoding network.

3. The method of claim 1, wherein, The first encoding network and the second encoding network each comprise a plurality of encoding network layers, and the decoding network comprises a plurality of decoding network layers; an output of each encoding network layer in the first encoding network and the second encoding network is input into a corresponding next encoding network layer and a decoding network layer.

4. The method of claim 1, wherein, The training process of the makeup transfer model comprises: obtaining a training sample triple comprising a template object sample image, an initial feature replacement sample image, and a target feature replacement sample image; inputting the template object sample image and the initial feature replacement sample image into the makeup transfer model to obtain a predicted image output by the makeup transfer model; correcting network parameters of the makeup transfer model according to the predicted image and the target feature replacement sample image to obtain the makeup transfer model trained.

5. The method of claim 4, wherein, The obtaining of the training sample triple comprises: obtaining an initial template object sample image and a target object sample image; performing feature replacement processing on the initial template object sample image according to the target object sample image to obtain an initial feature replacement sample image; According to the makeup material library, makeup material adding processing is performed on the initial feature replacement sample image and the initial template object sample image to obtain a target feature replacement sample image and a template object sample image; The training sample triplets are obtained according to the template object sample image, the initial feature replacement sample image and the target feature replacement sample image.

6. The method of claim 4, wherein, The network parameters of the makeup transfer model are corrected according to the predicted image and the target feature replacement sample image, and the method comprises the steps of: The predicted image and the target feature replacement sample image are input into a discriminant network to obtain a discriminant result output by the discriminant network; wherein the discriminant result is used to represent an image output by the predicted target in the predicted image and the target feature replacement sample image; The target feature replacement sample image is taken as an actual target output, and a loss function value is calculated according to the predicted target output and the actual target output; The network parameters of the makeup transfer model are corrected according to the loss function value.

7. The method according to any one of claims 1 to 6, characterized in that, The feature replacement processing is performed on the template object image according to the target object image to obtain an initial feature replacement image, and the method comprises the steps of: The target object image and the template object image are input into the trained feature replacement model to obtain object identity features of a target object in the target object image and object additional attribute features of a template object in the template object image output by the feature replacement model; The object identity features of the target object and the object additional attribute features of the template object are fused to obtain a virtual object; The initial feature replacement image is obtained according to the virtual object.

8. The method of claim 7, wherein, The target object image and the template object image are input into the trained feature replacement model to obtain object identity features of a target object in the target object image and object additional attribute features of a template object in the template object image output by the feature replacement model, and the method comprises the steps of: The target object contained in the target object image is identified to obtain a replacement area corresponding to the target object; At least two sub-identity feature information is extracted from the replacement area, wherein the at least two sub-identity feature information is used to represent object identity information at different positions in the replacement area; The at least two sub-identity feature information is fused to obtain object identity features of a target object in the target object image.

9. An image processing apparatus characterized by comprising: The device comprises: An image acquisition module configured to acquire a target object image containing a target object and acquire a template object image containing a template object; A feature replacement module configured to perform feature replacement processing on the template object image according to the target object image to obtain an initial feature replacement image; wherein the initial feature replacement image contains a virtual object, and the virtual object has object identity features of the target object and object additional attribute features of the template object; A first encoding module configured to input the template object image into a first encoding network contained in a trained makeup transfer model to obtain makeup features of the template object output by the first encoding network; a second encoding module, configured to input the initial feature replacement image into a second encoding network and a multi-layer perception included in the makeup transfer model, to obtain virtual object features of the virtual object output by the second encoding network and style features of the virtual object output by the multi-layer perception; a decoding module, configured to input the makeup features of the template object, the virtual object features of the virtual object, and the style features of the virtual object into a decoding network included in the makeup transfer model, to obtain a target feature replacement image output by the decoding network.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by a processor to implement the image processing method in any one of claims 1 to 8.

11. An electronic device, comprising: comprise: a processor; and a memory for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the image processing method in any one of claims 1 to 8.

12. A computer program product, characterised in that, The computer program product comprises computer instructions for implementing the image processing method in any one of claims 1 to 8 when executed by a processor.