Image processing method, device, electronic device, and computer-readable storage medium
By acquiring the features of the pending image and reference image for face reconstruction, generating rendered images and depth images, the problems of face deformation and clarity are solved, and the decoupling control of head posture and expression are achieved.
Patent Information
- Application Number
- CN202110633640.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-07
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2041-06-07
AI Technical Summary
The existing face driving method is based on facial key point recognition, resulting in face deformation and clarity problems in different face images, and it is impossible to achieve decoupling control of head posture and expression.
Reconstruct the face by obtaining the shape, texture, expression and posture features of the pending and reference images, generate rendered images and depth maps, and use these features for image synthesis to generate target images.
Accurate image synthesis between different face images is realized, face deformation problem is solved, and head posture and expression decoupling control is realized.
Smart Images

Figure CN113221847B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to an image processing method, device, electronic device and computer-readable storage medium. Background Art
[0002] With the continuous development of society, electronic devices such as mobile phones and tablet computers have been widely used in learning, entertainment, work, etc., playing an increasingly important role. These electronic devices are equipped with cameras, which can be used for applications such as taking pictures, recording videos or live broadcasting.
[0003] In applications such as live streaming, AR (Augmented Reality), and emoji creation, face-driven technology can identify the current user's facial state and drive another face to express that facial state. However, existing face-driven methods are based on facial key point recognition. This has high image requirements and requires that the source image and the driven image be the same face. Otherwise, facial deformation will occur, and it cannot achieve decoupled control of facial expression and posture. Summary of the Invention
[0004] In view of this, the object of the present invention is to provide an image processing method, device, electronic device and computer-readable storage medium to solve the problem of facial deformation and inability to achieve decoupling control of facial expressions and postures in the prior art.
[0005] In order to achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows:
[0006] In a first aspect, the present invention provides an image processing method, comprising: acquiring an image to be processed and a reference image; the face to be processed in the image to be processed is the same as or different from the reference face in the reference image; performing face reconstruction based on the shape features and texture features of the face to be processed, and the expression features and / or posture features of the reference face to obtain a rendering image and a depth map; generating a target image based on the image to be processed, the rendering image and the depth map; wherein the target image has the face to be processed; and the face to be processed has the expression features and / or posture features of the reference face.
[0007] In a second aspect, the present invention provides an image processing device, comprising: an acquisition module for acquiring an image to be processed and a reference image; the face to be processed in the image to be processed is the same as or different from the reference face in the reference image; a processing module for performing face reconstruction based on the shape features and texture features of the face to be processed, and the expression features and / or posture features of the reference face to obtain a rendering image and a depth map; the processing module is also used to generate a target image based on the image to be processed, the rendering image and the depth map; the target image has the face to be processed; the face to be processed has the expression features and / or the posture features.
[0008] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the image processing method described in the first aspect.
[0009] In a fourth aspect, the present invention provides a computer-readable storage medium having machine-executable instructions stored thereon, wherein the machine-executable instructions, when executed by a processor, implement the image processing method described in the first aspect.
[0010] The present invention provides an image processing method, device, electronic device and computer-readable storage medium, which includes: obtaining an image to be processed and a reference image; the face to be processed in the image to be processed is the same as or different from the reference face in the reference image; performing face reconstruction based on the shape features and texture features of the face to be processed, and the expression features and / or posture features of the reference face to obtain a rendering image and a depth map; generating a target image based on the image to be processed, the rendering image and the depth map; wherein the target image has the face to be processed; and the face to be processed has the expression features and / or the posture features.
[0011] The difference from the prior art is that the existing face driving method is a driving method based on facial key point recognition. Once the driving image and the source image are not the same face, it will cause the problem of facial deformation. At the same time, this face driving method cannot achieve the decoupling control of head posture and expression. The embodiment of the present invention provides a face driving method that no longer relies on face key technology for face driving, but instead uses the shape features and texture features of the face to be processed, as well as the expression features and / or posture features of the reference face to reconstruct the face and obtain a rendering image and a depth map. It can be seen that the obtained rendering image and depth map contain the shape features of the face to be processed, the expression features and / or posture features of the reference face, which can achieve the effect of face driving and the effect of decoupling control of head posture and expression. At the same time, the depth map also contains the depth information of the face to be processed, and the rendering image also contains the texture information of the face to be processed. Based on this information, the accuracy of the face to be processed in the generated target image can be guaranteed, solving the face deformation problem in the prior art.
[0012] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 A schematic flow chart of an image processing method provided by an embodiment of the present invention;
[0015] Figure 2 A schematic diagram of a scenario provided by an embodiment of the present invention;
[0016] Figure 3 A schematic flowchart of an implementation method of step S105 provided in an embodiment of the present invention;
[0017] Figure 4 A schematic flowchart of an implementation of step S106 provided in an embodiment of the present invention;
[0018] Figure 5 A schematic diagram of another scenario provided by an embodiment of the present invention;
[0019] Figure 6 A schematic diagram of a model training provided by an embodiment of the present invention;
[0020] Figure 7A and Figure 7B is a schematic diagram of a user interface provided by an embodiment of the present invention;
[0021] Figure 8 A functional module diagram of an image processing device provided by an embodiment of the present invention;
[0022] Figure 9 A structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The following will be combined with the accompanying drawings to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0024] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but is merely intended to represent selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.
[0025] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0026] Face-driven technology is increasingly popular in applications such as live streaming, AR (Augmented Reality), and emoji creation. Face-driven technology can identify the current user's facial state and then drive another face to express that facial state. Specifically, given two images, one is a source image (containing the face to be processed) and the other is a driving image (containing a reference face). These two images are processed, and the output image contains the face to be processed, but the face to be processed has the same head pose, expression, and driving image as the reference face.
[0027] The existing face driving method is a driving method based on the recognition of facial key points. For example, the relevant technology uses the input of facial key points and source images into the neural network model for face driving. However, different faces differ in shape (including face), facial feature positions, etc. The key point-based generation scheme uses key points as input and marks the facial feature positions through key points. Then, once the driving image and the source image are not the same face, the facial shape and facial feature distance represented by the key points extracted from the driving image are different from the key points of the face in the source image, which will cause the face shape of the final generated image to be the same as the face shape in the driving image, but the facial skin color is the same as the skin color of the face in the source image, which will cause the face deformation and low clarity. At the same time, this face driving method cannot achieve decoupled control of head posture and expression.
[0028] It is understandable that the existing face driving method based on facial key point technology has high requirements on images. The source image and the driving image must be the same face. Otherwise, there will be problems with facial clarity and facial deformation, and it is impossible to achieve decoupled control of facial expressions and postures.
[0029] In order to solve the above technical problems, an embodiment of the present invention provides a face driving method, that is, no longer relying on key face technologies for face driving, but using the shape features, texture features of the source image and the expression features and / or posture features of the driving image to reconstruct the face, obtain the reconstructed face rendering image and face depth map, and then perform image synthesis based on the face rendering image and the face depth map to obtain the target image. Since more face information, such as texture information, depth information, etc., is provided in the image synthesis process, the problems of face deformation and poor clarity in the existing technology can be solved. Furthermore, due to the expression features and / or posture features used in the process of obtaining the rendering image and the depth map, the effect of decoupling control of head posture and expression can be achieved.
[0030] To facilitate understanding of the above technical effects, the image processing method provided by the embodiment of the present invention will be introduced below with reference to relevant drawings.
[0031] See Figure 1 , Figure 1 A schematic flowchart of an image processing method provided in an embodiment of the present invention may include:
[0032] S104: Acquire the image to be processed and the reference image.
[0033] It is understood that in embodiments of the present invention, the face to be processed in the processed image and the reference face in the reference image may be the same or different. That is, if the face to be processed and the reference face are the same, then the face in the target image subsequently obtained by embodiments of the present invention will naturally not be deformed. Even if the face to be processed and the reference face are different, the deformation problem can still be resolved through the subsequent processing steps of embodiments of the present invention.
[0034] In some possible embodiments, the image to be processed may be one image, which may be obtained by obtaining a pre-stored image or an image captured by an acquisition device; the reference image may be a single image or several consecutive frames of images. Therefore, the reference image may be obtained by obtaining a pre-stored image or an image captured by an acquisition device, or all frame images of the video.
[0035] S105 , performing face reconstruction based on the shape features and texture features of the face to be processed, and the expression features and / or posture features of the reference face, to obtain a rendering image and a depth map.
[0036] It is understood that in embodiments of the present invention, facial reconstruction can be performed based on the facial features of the image to be processed and the reference image, thereby obtaining a rendering image and a depth map of the reconstructed face. In order to ensure that the target image subsequently obtained includes the face to be processed, the face in the rendering image and the depth map generated in step S105 includes the shape features of the face to be processed and the expression features and / or posture features of the reference face.
[0037] It can also be understood that the above-mentioned rendering refers to the planar face image obtained by rendering the three-dimensional face model, and the depth map is a 256x256 image, in which the pixel value of each pixel is between 0-255. The size of the pixel value represents the distance between the face position represented by the pixel and the screen, that is, the depth of the face. The closer the pixel value is to 255, the closer the pixel point is to the screen, and the closer the pixel value is to 0, the farther the pixel point is from the screen. For example, with regard to the position of the nose and eyes on the face, the nose is more prominent and the distance from the screen is smaller. Therefore, the pixel value corresponding to the pixel point of the nose is larger than the pixel value corresponding to the eye point.
[0038] In an embodiment of the present invention, the applicant has discovered through research that when a person turns his head, the value of his texture features will change due to the influence of lighting. Therefore, in order to eliminate the influence of lighting on texture features, the generated face can be fine-tuned based on additional depth information to make the generated face more realistic.
[0039] S106: Generate a target image based on the image to be processed, the rendering image, and the depth map.
[0040] The target image has a face to be processed, and the face to be processed has expression features and / or posture features of a reference face.
[0041] It can be understood that in order to make the face to be processed in the target image closer to the face to be processed in the image to be processed, the features in the image to be processed can be referenced in the process of obtaining the target image. At the same time, since the rendering image and the depth image should contain the shape features of the face to be processed, the expression features and / or posture features of the reference face, the generated target image can have the face to be processed, and the face to be processed has the expression features and / or posture features of the reference face.
[0042] An image processing method provided by an embodiment of the present invention differs from the prior art in that the prior art face driving method is a driving method based on facial key point recognition. If the driving image and the source image are not the same face, the face will be deformed. At the same time, this face driving method cannot achieve decoupled control of head posture and expression. The embodiment of the present invention provides a face driving method that no longer relies on key face technology for face driving. Instead, it uses the shape and texture features of the face to be processed, as well as the expression features and / or posture features of a reference face to reconstruct the face, obtaining a rendering image and a depth map. It can be seen that the obtained rendering image and depth map contain the shape features of the face to be processed, the expression features and / or posture features of the reference face, and can achieve the effect of face driving while also achieving the effect of decoupling control of head posture and expression. At the same time, the depth map also contains the depth information of the face to be processed, and the rendering image also contains the texture information of the face to be processed. Based on this information, the accuracy of the face to be processed in the generated target image can be guaranteed, solving the face deformation problem in the prior art.
[0043] To understand the above effects, see Figure 2 , Figure 2 A schematic diagram of a scenario provided by an embodiment of the present invention. Figure 2 It can be seen that the embodiment of the present invention can obtain a rendering image and a depth map of a face based on the image to be processed and the reference image. It can be seen that the face in the rendering image and the depth map is already the face to be processed, and then the target image is obtained based on the obtained rendering image, depth map and image to be processed. The face in the obtained target image is the face to be processed, and the face to be processed has the expression features and posture features of the reference face.
[0044] In some possible embodiments, in order to ensure that the obtained rendering image and depth map contain the facial features of the face to be processed and the expression features and / or posture features of the reference face, and to ensure that the problem of face deformation does not occur, a method for obtaining the rendering image and depth map is given below. Figure 3 , Figure 3 This is a schematic flowchart of an implementation of step S105 provided in an embodiment of the present invention. Step S105 may include the following sub-steps:
[0045] Sub-step S105-1, performing face reconstruction on the image to be processed to obtain shape features and texture features.
[0046] It is understood that in order to ensure that the final target image contains the face to be processed, the shape features can be used in the subsequent face reconstruction process to ensure that the face in the rendering image and depth image is the face to be processed. At the same time, in order to prevent facial deformation, the texture features of the face to be processed can also be obtained. The texture features mentioned here are the regular distribution of grayscale values caused by the repeated arrangement of objects in the image. Such features are the texture features of the image. In other words, for different faces, the grayscale values presented in the image have different regular distributions. Therefore, obtaining texture features can ensure that the face obtained in the subsequent reconstruction is nearly identical to the face to be processed.
[0047] In some possible embodiments, if the face to be processed also has expression features and / or posture features, then performing face reconstruction on the face to be processed can also obtain the expression features and / or posture features of the face to be processed.
[0048] Sub-step S105-2: reconstructing the face of the reference image to obtain expression features and / or posture features.
[0049] It can be understood that in order to make the processed face in the final generated target image have the expression features and / or posture features of the reference face, the expression features and / or posture features of the reference face should be considered in the process of obtaining the rendering image and depth map, so as to express the state of the reference face through the processed face.
[0050] Sub-step S105-3: inputting the expression features and / or posture features, as well as the shape features and texture features into a preset parameter model to obtain a three-dimensional face model.
[0051] In an embodiment of the present invention, the preset parameter model can be used to reconstruct the face based on the acquired expression features and / or posture features, as well as the shape features and texture features to obtain a three-dimensional face model.
[0052] In one possible implementation, expression features, shape features, and texture features can be input into a preset parameter model to obtain a three-dimensional face model. In this way, the obtained three-dimensional face has these three features. In another possible implementation, posture, shape features, and texture features can also be input into a preset parameter model to obtain a three-dimensional face with these three features. In another possible implementation, expression features, posture features, shape features, and texture features can also be input into a preset parameter model to obtain a three-dimensional face with these four features. In this way, the effect of decoupling control of expression features and posture features can be achieved.
[0053] Sub-step S105-4: obtaining a rendering image and a depth map according to the three-dimensional face model.
[0054] Through the above process, the rendering image and depth map can be obtained. It can be seen that since the three-dimensional face model is obtained based on the shape features, texture features of the face to be processed and the expression features and / or posture features of the reference face, the three-dimensional face should have the shape and texture of the face to be processed, and the expression and / or posture of the reference face.
[0055] In some possible embodiments, in order to finally obtain the target image, an implementation method is given below. Figure 4 , Figure 4 This is a schematic flowchart of an implementation of step S106 provided in an embodiment of the present invention. Step S106 may include the following sub-steps:
[0056] Sub-step S106-1: input the image to be processed, the rendering image and the depth map into the pre-trained image generation model.
[0057] Sub-step S106-2: extracting semantic features of the image to be processed, the rendering image, and the depth image through the image generation model, and generating a target image based on the obtained semantic features.
[0058] It is understood that the above-mentioned image generation model can be, but is not limited to, a Unet neural network. Since the rendering image contains the texture features of the face to be processed, the texture features provide the Unet neural network with more facial information. Since the Unet neural network has jump links, the face rendering image can be directly output through the jump links, which can prevent the generated face from being deformed. Furthermore, adding a depth map to the input of the Unet neural network can improve the problem of facial texture values being changed due to lighting issues. Because when reconstructing a face using a preset parameter model, only the position of the texture features is moved, and the texture value is not changed. However, when a person turns their head, the texture value will change due to the influence of lighting. Inputting the depth map into the Unet neural network allows the Unet neural network to obtain not only the planar information of the face (the face rendering image) but also the depth information of the face. In this way, when learning, the Unet neural network can fine-tune the generated face based on this additional depth information, making the generated face more realistic.
[0059] In order to facilitate the understanding of the above implementation process, the following Figure 2 Based on this, another scenario diagram is given, see Figure 5 , Figure 5 Another scenario schematic diagram provided for an embodiment of the present invention shows that, based on face reconstruction, the shape features and texture features of the face to be processed, the expression features and posture features of the reference face can be obtained respectively, and then these features are input into a preset parameter model to obtain a three-dimensional face model, and then the three-dimensional face model is rendered to obtain a rendering image and a depth map. It can be seen that the shape of the face in the rendering image and the depth map is the shape of the face to be processed, and the expression and posture are the expression and posture of the reference face. Then, the image to be processed, the rendering image and the depth map are simultaneously input into the image generation model. The image generation model can propose semantic features of the face in the model to be processed, for example, extract the shape features and texture features in the rendering image and the depth features extracted from the depth map, and can also extract other features in the image to be processed, and combine these features for image generation. The final generated image can have the shape, texture of the face to be processed and the expression and posture of the reference face.
[0060] In some possible embodiments, please combine Figure 6 , Figure 6 A model training diagram provided in an embodiment of the present invention is provided. The image generation model in an embodiment of the present invention is trained in the following manner:
[0061] Step 1: obtain a training sample image set; the training sample image set includes a training source image and a training reference image; the training source image and the training reference image have the same face.
[0062] It is understandable that the training source image and the training reference image are the same person, because this is the only way to construct the image loss. When the training source image and the training reference image are different, the generated image and the training reference image have the same expression but are faces of different people.
[0063] Step 2: Train the initial face-driven model based on the training sample image set.
[0064] Step 3: If the loss function value of the image generation model is within a preset threshold range, the trained image generation model is obtained.
[0065] It can be understood that during training, the generated image can be further reconstructed using face reconstruction to obtain the texture features, shape features, expression features and posture features of the generated image, and these features are respectively combined with the shape features, texture features of the training original image and the expression features and posture features of the training reference image to construct the shape feature loss function, texture feature loss function, expression feature loss function and posture feature loss function, so that more supervision information is obtained during training and the effect is improved. At the same time, the generated image and the training source image can be used to construct the image perception loss function and the image loss function, and the shape feature loss function, texture feature loss function, expression feature loss function, posture feature loss function, image perception loss function and image loss function are combined to jointly construct the loss function of the image generation model. In one implementation method, the final constructed loss function can be expressed as: loss function = 10*image perception loss function + 10*image loss function + shape feature loss function + expression feature function + posture loss function + texture loss function.
[0066] In some possible embodiments, the number of training times may be, but is not limited to, 150 epochs, with 40,000 iterations per epoch. When the loss value of the loss function drops from 80 to 6 or 7, the trained image generation model can be obtained.
[0067] In some possible embodiments, in order to facilitate user operation, a method for obtaining the image to be processed and the reference image is also given below. Figure 7A and Figure 7B , Figure 7A and Figure 7B is a schematic diagram of a user interface provided by an embodiment of the present invention.
[0068] like Figure 7AAs shown, in one implementation, the image to be processed and the reference image are both still images. The method for obtaining the image to be processed and the reference image can be: in response to a user operation, a face-driven interface is displayed; the face-driven interface has a source image entry area and a reference image entry area. It can be understood that the source image is the image to be processed in this embodiment of the present invention. When an entry operation instruction is received in the source image entry area, the obtained image is used as the image to be processed; when an entry operation instruction is received in the source image entry area, the obtained image is used as the reference image.
[0069] Continue to see Figure 7B In another implementation, the image to be processed is a still image, and the reference image is each frame of a video. The method for obtaining the image to be processed and the reference image may be: in response to a user operation, a face-driven interface is displayed; the face-driven interface has a source image input area and a reference image input area; when an input operation instruction is received in the source image input area, the obtained image is used as the reference image. When a selection operation instruction is received in the reference image input area, the obtained frame image of the video file is used as the reference image.
[0070] It should be noted that in the process of obtaining the image to be processed and the reference image, the source image entry area and the reference image entry area may not be in the same display interface. For example, there may be a source image entry area in one display interface, and when the user's entry operation is received, it triggers the display of another interface, and there may be a reference image entry area on the other interface.
[0071] Continue to see Figure 7A and Figure 7B The interface also has an icon to guide the user to generate an image, such as a "start" icon. When a generation operation instruction is received on the face-driven display interface, the generated target image is displayed in the preview area.
[0072] In order to execute the corresponding steps in the above embodiments and various possible methods, an implementation method of an image processing device is given below. Figure 8 , Figure 8 This is a functional block diagram of an image processing device provided by an embodiment of the present invention. It should be noted that the basic principles and technical effects of the image processing device provided by this embodiment are the same as those of the above embodiments. For the sake of simplicity, any parts not mentioned in this embodiment can be referred to the corresponding contents of the above embodiments. The image processing device 20 includes:
[0073] An acquisition module 21 is configured to acquire an image to be processed and a reference image; the face to be processed in the image to be processed is the same as or different from the reference face in the reference image;
[0074] The processing module 22 is used to perform face reconstruction based on the shape features and texture features of the face to be processed, as well as the expression features and / or posture features of the reference face, to obtain a rendering image and a depth map; and to generate a target image based on the image to be processed, the rendering image, and the depth map; wherein the target image has the face to be processed; and the face to be processed has the expression features and / or posture features of the reference face.
[0075] Optionally, both the rendering image and the depth map have the expression features and / or posture features and the shape features of the face to be processed; the rendering image also has the texture features of the face to be processed; the depth map also has the depth features of the face to be processed; the processing module 22 is specifically used to perform face reconstruction on the image to be processed to obtain the shape features and the texture features; perform face reconstruction on the reference image to obtain the expression features and / or the posture features; input the expression features and / or the posture features, as well as the shape features and the texture features into a preset parameter model to obtain a three-dimensional face model; and obtain the rendering image and the depth map based on the three-dimensional face model.
[0076] Optionally, the processing module 22 is further specifically used to input the image to be processed, the rendering image and the depth map into a pre-trained image generation model; extract the semantic features of the image to be processed, the rendering image and the depth map respectively through the image generation model, and generate the target image based on the obtained semantic features.
[0077] Optionally, the image generation model is trained in the following manner: obtaining a training sample image set; the training sample image set includes a training source image and a training reference image; the training source image and the training reference image have the same face; based on the training sample image set, training the initial face-driven model; if the loss function value of the image generation model is within a preset threshold range, obtaining the trained image generation model.
[0078] Optionally, the acquisition module 21 is specifically configured to: display a face-driven interface in response to a user operation; the face-driven interface has a source image entry area and a reference image entry area; when an entry operation instruction is received in the source image entry area, use the obtained image as the image to be processed; when an entry operation instruction is received in the source image entry area, use the obtained image as the reference image; or, when a selection operation instruction is received in the source image entry area, use the obtained frame image of the video file as the reference image;
[0079] Optionally, the acquisition module 21 is also used to display the image to be processed and the reference image in the preview area of the face drive display interface; when a generation operation instruction is received on the face drive display interface, the generated target image is displayed in the preview area.
[0080] An embodiment of the present invention further provides an electronic device, such as Figure 9 , Figure 9 This is a block diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 80 includes a communication interface 81, a processor 82 and a memory 83. The processor 82, the memory 83 and the communication interface 81 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory 83 can be used to store software programs and modules, such as program instructions / modules corresponding to the image processing method provided in an embodiment of the present invention. The processor 82 executes various functional applications and data processing by executing the software programs and modules stored in the memory 83. The communication interface 81 can be used for signaling or data communication with other node devices. In the present invention, the electronic device 80 can have multiple communication interfaces 81.
[0081] Among them, the memory 83 can be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), etc.
[0082] The processor 82 can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0083] Optionally, the above modules can be stored in the form of software or firmware. Figure 9 The memory shown in FIG. 1 or solidified in the operating system (OS) of the electronic device, and can be Figure 9 Meanwhile, the data and program codes required to execute the above modules may be stored in the memory.
[0084] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or actions, or can be implemented using a combination of dedicated hardware and computer instructions.
[0085] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0086] If the functions are implemented as software modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0087] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire an image to be processed and a reference image; wherein the face to be processed in the image to be processed is the same as or different from the reference face in the reference image; Reconstructing the face based on the shape features and texture features of the face to be processed and the expression features and / or posture features of the reference face to obtain a rendering image and a depth map; wherein, when reconstructing the face using a preset parameter model, only the position of the texture features is moved, and the texture value is not changed; Generate a target image according to the image to be processed, the rendering image and the depth map, including: extracting semantic features of the image to be processed, the rendering image and the depth map respectively through the image generation model, and generating the target image according to the obtained semantic features; the image generation model is trained in the following manner: obtain a training sample image set; the training sample image set includes a training source image and a training reference image; the training source image and the training reference image have the same human face; train the initial image generation model according to the training sample image set; if the loss function value of the image generation model is within a preset threshold range, obtain the trained image generation model; wherein the loss function is constructed by combining a shape feature loss function, a texture feature loss function, an expression feature loss function, a posture feature loss function, an image perception loss function and an image loss function; The target image has the human face to be processed; the human face to be processed has the expression feature and / or the posture feature.
2. The image processing method according to claim 1, wherein: The rendering image and the depth image both have the expression features and / or posture features and the shape features; the rendering image also has the texture features; and the depth image also has the depth features of the face to be processed; The step of performing face reconstruction based on the shape features and texture features of the face to be processed and the expression features and / or posture features of the reference face to obtain a rendering image and a depth map includes: Performing face reconstruction on the image to be processed to obtain the shape feature and the texture feature; Performing facial reconstruction on the reference image to obtain the expression features and / or the posture features; Inputting the expression features and / or the posture features, as well as the shape features and the texture features into the preset parameter model to obtain a three-dimensional face model; The rendering image and the depth map are obtained according to the three-dimensional face model.
3. The image processing method according to claim 1, wherein: The step of generating a target image according to the image to be processed, the rendering image and the depth map further includes: The image to be processed, the rendering image, and the depth map are input into the pre-trained image generation model.
4. The image processing method according to claim 1, wherein: The step of obtaining the image to be processed and the reference image includes: In response to user operations, a face drive interface is displayed; the face drive interface has a source image entry area and a reference image entry area; When receiving an input operation instruction of the source image input area, using the obtained image as the image to be processed; When an input operation instruction is received in the reference image input area, the obtained image is used as the reference image; or when a selection operation instruction is received in the reference image input area, the obtained frame image of the video file is used as the reference image.
5. The image processing method according to claim 4, characterized in that The method further comprises: Displaying the image to be processed and the reference image in a preview area of the face drive interface; When a generation operation instruction is received on the face drive interface, the generated target image is displayed in the preview area.
6. An image processing device, characterized in that include: An acquisition module is used to acquire an image to be processed and a reference image; the face to be processed in the image to be processed is the same as or different from the reference face in the reference image; a processing module, configured to perform face reconstruction based on the shape and texture features of the face to be processed, and the expression features and / or posture features of the reference face, to obtain a rendering image and a depth map; wherein, when performing face reconstruction using a preset parameter model, only the positions of the texture features are moved, and the texture values are not changed; The processing module is also used to generate a target image based on the image to be processed, the rendering image and the depth map, including: extracting the semantic features of the image to be processed, the rendering image and the depth map respectively through the image generation model, and generating the target image based on the obtained semantic features; the image generation model is trained in the following manner: obtaining a training sample image set; the training sample image set includes a training source image and a training reference image; the training source image and the training reference image have the same face; according to the training sample image set, the initial image generation model is trained; if the loss function value of the image generation model is within a preset threshold range, the trained image generation model is obtained; wherein the loss function is constructed in combination with a shape feature loss function, a texture feature loss function, an expression feature loss function, a posture feature loss function, an image perception loss function and an image loss function; wherein the target image has the face to be processed; the face to be processed has the expression feature and / or the posture feature.
7. The image processing device according to claim 6, wherein: The rendering image and the depth image both have the expression features and / or posture features and the shape features of the face to be processed; the rendering image also has the texture features; the depth image also has the depth features of the face to be processed; the processing module is specifically configured to: Performing face reconstruction on the image to be processed to obtain the shape feature and the texture feature; Performing facial reconstruction on the reference image to obtain the expression features and / or the posture features; Inputting the expression features and / or the posture features, as well as the shape features and the texture features into the preset parameter model to obtain a three-dimensional face model; The rendering image and the depth map are obtained according to the face model.
8. An electronic device, characterized in that: The image processing method comprises a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the image processing method according to any one of claims 1 to 5.
9. A computer-readable storage medium having machine-executable instructions stored thereon, characterized in that: When the machine executable instructions are executed by a processor, the image processing method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Face reconstruction method, apparatus and device based on reconstruction network, medium and product
CN108776983A
Three-dimensional face reconstruction method and device, electronic equipment and storage medium
CN112819947A