A face driving method, device, electronic device and storage medium

By adjusting the face key points in the video image and generating the target mask image and superimposing it, inputting it to the image generation model, the problem of poor image integrity during the face driving process is solved, and high-quality face driving effect is achieved.

CN114626979BActive Publication Date: 2025-05-27GUANGZHOU HUYA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210272249.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2025-05-27
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

The prior art is difficult to ensure the integrity of the image during the face driving process, resulting in poor driving effect and prone to image rupture.

Method used

By determining the image to be processed and the video to be processed, multi-frame video images of the video are obtained, and the face key points of the image are adjusted based on the face key points of the video image, the target mask image is generated and the image is superimposed, and input to the image generation model to generate a face-driven video.

Benefits of technology

Fill in the missing information during the face driving process, ensure the integrity of the face during the driving process, avoid the problem of image rupture, and improve the driving effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114626979B_ABST
    Figure CN114626979B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a face driving method, apparatus, electronic device, and storage medium. The method includes: determining an image to be processed and a video to be processed, obtaining multiple video images of the video to be processed, respectively adjusting the face key points of the image to be processed based on the face key points of each video image to obtain multiple first images, where each first image includes first face key points. For each first image, superimposing the target mask image corresponding to the first face key points in the first image on the first image to obtain multiple second images, inputting each second image into a trained image generation model, outputting multiple target images, and generating a face driving video based on each target image. This application can fill in the missing information during the face driving process, ensure the integrity of the face during the driving process, and avoid the problem of image rupture when driving the face.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more particularly, to a face driving method, device, electronic device, and storage medium. Background Art

[0002] Face driving of an image refers to driving the face in the image so that the face contained in the image changes.

[0003] Currently, it is impossible to generate content that the original image does not carry. For example, if the original image is an image with a closed mouth, driving it to open the mouth will result in a cracking and tearing phenomenon, leading to a poor driving effect. Summary of the Invention

[0004] The purpose of the present invention is to provide a face driving method, device, electronic device, and storage medium, which can ensure the integrity of the driven image, thereby improving the driving effect.

[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present application are as follows:

[0006] In a first aspect, an embodiment of the present application provides a face driving method, and the method includes:

[0007] Determine an image to be processed and a video to be processed, wherein the portrait in the image to be processed is driven based on the video to be processed;

[0008] Obtain multiple video images of the video to be processed;

[0009] Adjust the face key points of the image to be processed based on the face key points of each video image respectively to obtain multiple first images, wherein each first image includes first face key points;

[0010] For each first image, superimpose the target mask image corresponding to the first face key point in the first image on the first image to obtain multiple second images;

[0011] Input each second image into a trained image generation model to output multiple target images;

[0012] Generate a face driving video based on each target image.

[0013] In an alternative embodiment, the step of adjusting the face key points of the image to be processed based on the face key points of each video image respectively to obtain multiple first images includes:

[0014] For each video image, determine the second face key points in the video image and the first coordinates of each second face key point;

[0015] Determine the third face key points of the image to be processed and the second coordinates of each of the third face key points;

[0016] Adjust the corresponding second coordinates based on the first coordinates to obtain multiple adjusted first images.

[0017] In an alternative embodiment, the step of, for each of the first images, superimposing the target mask image corresponding to the first face key point in the first image on the first image to obtain multiple second images includes:

[0018] For each of the first images, determine the first face key points in each of the first images;

[0019] Construct multiple first triangular mesh images based on the first face key points in each of the first images;

[0020] Obtain the first mask image of the standard face;

[0021] Obtain the fourth face key points of the face in the first mask image, and construct a second triangular mesh image based on the fourth face key points;

[0022] Respectively adjust the second triangular mesh image based on each of the first triangular mesh images to obtain multiple updated target mask images;

[0023] Superimpose each of the target mask images on the corresponding first image to obtain multiple second images.

[0024] In an alternative embodiment, the step of respectively adjusting the second triangular mesh image based on each of the first triangular mesh images to obtain multiple updated target mask images includes:

[0025] For each of the first triangular network images, determine each first triangular mesh in the first triangular mesh image;

[0026] For each of the first triangular meshes, determine the corresponding second triangular mesh, where the second triangular mesh is a triangular mesh in the second triangular mesh image;

[0027] Based on the first triangular mesh, perform an affine transformation on the second triangular mesh to obtain multiple updated target mask images.

[0028] In an alternative embodiment, the method further includes:

[0029] Train the image generation model;

[0030] This step includes:

[0031] Obtain the second mask image of the standard face and the face key points in the second mask image;

[0032] Based on the face key points in the second mask image, construct a third triangular mesh image;

[0033] Determine the image to be trained;

[0034] Determine the face key points in the image to be trained, and construct a fourth triangular mesh image based on the face key points of the image to be trained;

[0035] Based on each triangular mesh in the fourth triangular mesh image, adjust each triangular mesh in the third triangular mesh image to obtain a third mask image;

[0036] Overlay the third mask image with the image to be trained to obtain a target training image;

[0037] Train an image generation model based on the target training image, where the input of the image generation model is the target training image and the output is the training image.

[0038] In an alternative embodiment, the step of overlaying the target mask image corresponding to the first face key point in the first image with the first image for each of the first images to obtain a plurality of second images includes:

[0039] For each of the first images, overlay the target mask image corresponding to the first face key point in the first image with the first image;

[0040] Perform a feathering operation on the target mask image overlaid in the first image to obtain a plurality of second images.

[0041] In a second aspect, an embodiment of the present application provides a face driving device, the device includes: a determination module and a processing module;

[0042] The determination module is configured to determine an image to be processed and a video to be processed, wherein the portrait in the image to be processed is driven based on the video to be processed;

[0043] Obtain multiple video images of the video to be processed;

[0044] The processing module is configured to respectively adjust the face key points of the image to be processed based on the face key points of each of the video images to obtain a plurality of first images, wherein each of the first images includes first face key points;

[0045] For each of the first images, superimpose the target mask image corresponding to the first face key points in the first image on the first image to obtain a plurality of second images;

[0046] Input each of the second images into a trained image generation model to output a plurality of target images;

[0047] Generate a face-driven video based on each of the target images.

[0048] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the face driving method are implemented.

[0049] In a fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the face driving method are implemented.

[0050] The present application has the following beneficial effects:

[0051] In the present application, by determining the image and video to be processed, obtaining multiple video images of the video to be processed, and respectively adjusting the face key points of the image to be processed based on the face key points of each video image, a plurality of first images are obtained. Each of the first images includes first face key points. For each first image, superimpose the target mask image corresponding to the first face key points in the first image on the first image to obtain a plurality of second images. Input each of the second images into a trained image generation model to output a plurality of target images, and generate a face-driven video based on each of the target images. The present application can fill in the missing information in the face driving process, ensure the integrity of the face during the driving process, and avoid the problem of image rupture when driving the face. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a schematic block diagram of the electronic device provided by the embodiment of the present invention;

[0054] Figure 2 It is one of the schematic flowcharts of a face driving method provided by the embodiment of the present invention;

[0055] Figure 3Schematic diagram of the second image provided by the embodiment of the present invention;

[0056] Figure 4 Schematic diagram II of the process of a face driving method provided by the embodiment of the present invention;

[0057] Figure 5 Schematic diagram III of the process of a face driving method provided by the embodiment of the present invention;

[0058] Figure 6 Schematic diagram IV of the process of a face driving method provided by the embodiment of the present invention;

[0059] Figure 7 Schematic diagram of the third triangular mesh image provided by the embodiment of the present invention;

[0060] Figure 8 Schematic diagram of the fourth triangular mesh image provided by the embodiment of the present invention;

[0061] Figure 9 Structural block diagram of the face driving device provided by the embodiment of the present invention. Detailed implementation manners

[0062] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0063] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0064] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0065] In the description of the present invention, it should be noted that if terms such as "upper", "lower", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings or the orientation or positional relationship in which the product of the invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0066] In addition, terms such as "first" and "second" are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0067] In the description of the present application, it should also be noted that unless otherwise clearly specified and defined, the terms "set", "installed", "connected", and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0068] Through a large number of studies by the inventors, it is found that currently, it is impossible to generate content that the original picture does not carry. For example, if the original image is an image of a closed mouth, directly driving the position relationship of some local areas in the human face to achieve the driving of the human face. For example, when driving a human face image in a closed mouth state to an open mouth state, the key point positions of the lip part in the human face are changed, so as to achieve the driving from the closed mouth state to the open mouth state. At this time, when driving the open mouth, a phenomenon of rupture and tearing will occur, resulting in a poor driving effect.

[0069] In view of the discovery of the above problems, the present embodiment provides a human face driving method, device, electronic device and storage medium, which can determine a to-be-processed image and a to-be-processed video, obtain multiple frames of video images of the to-be-processed video, respectively adjust the human face key points of the to-be-processed image based on the human face key points of each video image to obtain multiple first images. For each first image, the target mask image corresponding to the first human face key point in the first image is superimposed on the first image to obtain multiple second images. The second images are input into a trained image generation model to output multiple target images, and based on each target image, a human face driving video is generated. The present application can fill in the missing information in the human face driving process, ensure the integrity of the human face during the driving process, and avoid the problem of image rupture when driving the human face. The solution provided by the present embodiment will be elaborated in detail below.

[0070] The present embodiment provides an electronic device that can drive a human face. In a possible implementation manner, the electronic device may be a user terminal. For example, the electronic device may be, but is not limited to, a server, a smart phone, a personal computer (PC), a tablet computer, a personal digital assistant (PDA), a mobile internet device (MID), etc.

[0071] Please refer toFigure 1 , Figure 1 1 is a schematic diagram of the structure of the electronic device 100 provided in the embodiment of the present application. The electronic device 100 may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0072] The electronic device 100 includes a face driving device 110 , a memory 120 and a processor 130 .

[0073] The memory 120 and the processor 130 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines. The face drive device 110 includes at least one software function module that can be stored in the memory 120 in the form of software or firmware or solidified in the operating system (OS) of the electronic device 100. The processor 130 is used to execute the executable modules stored in the memory 120, such as the software function modules and computer programs included in the face drive device 110.

[0074] The memory 120 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. The memory 120 is used to store a program, and the processor 130 executes the program after receiving an execution instruction.

[0075] Please refer to Figure 2 , Figure 2 For application Figure 1 A flow chart of a face driving method of the electronic device 100 is shown, and the method including each step is described in detail below.

[0076] Step 201: Determine the image and video to be processed.

[0077] Among them, the portrait of the image to be processed is driven and processed based on the video to be processed.

[0078] Step 202: Obtain multiple video images of the video to be processed.

[0079] Step 203: Adjust the face key points of the image to be processed based on the face key points of each video image respectively to obtain multiple first images.

[0080] Among them, each first image contains first face key points.

[0081] Step 204: For each first image, superimpose the target mask image corresponding to the first face key point in the first image on the first image to obtain multiple second images.

[0082] Step 205: Input each second image into the trained image generation model to output multiple target images.

[0083] Step 206: Generate a face-driven video based on each target image.

[0084] It should be noted that the image to be processed and the video to be processed can be an image and a video containing the same person, or the person contained in the image to be processed and the video to be processed can be different people. For example: the person contained in the image to be processed is person A, and the person in the video to be processed is person B.

[0085] The video to be processed can be a video captured in real time through the camera of an electronic device, or a video cached in the electronic device.

[0086] Determine multiple video images containing people in the video to be processed. It should be noted that the content captured in the video to be processed may include images of people and non-people. Therefore, in order to achieve face driving, it is necessary to determine multiple video images containing people in the video to be processed, so as to facilitate subsequent driving of the face in the image to be processed.

[0087] Since the face contained in the image to be processed is a static face image, in order to drive the face image in the image to be processed from a static image to a dynamic image, it is necessary to adjust the face key points in the image to be processed based on the face key points in multiple video images containing people in the video to be processed.

[0088] It should be noted that the face key points can include lip key points and eye key points. When it is necessary to drive the mouth of the person in the image to be processed, obtain the lip key points of the face in each video image and the lip key points in the image to be processed, and adjust the lip key points in the image to be processed based on the lip key points in the video image to obtain multiple adjusted images to be processed, that is, multiple first images. Each first image contains first face key points, that is, lip key points.

[0089] When it is necessary to drive the eyes of the person in the image to be processed, obtain the eye key points of the face in each video image and the eye key points in the image to be processed, and adjust the eye key points in the image to be processed based on the eye key points in the video image, to obtain multiple images to be processed after adjustment, that is, multiple first images. Each first image contains eye key points.

[0090] Of course, in addition to being able to adjust the eye key points or lip key points of the image to be processed based on the video image alone, it is also possible to simultaneously adjust the lip key points and eye key points in the image to be processed based on the eye key points and lip key points in the video image, so as to realize driving the mouth and eyes of the face in the image to be processed simultaneously.

[0091] In one example, it is also possible to first determine the video image of the target person in the video to be processed, and then adjust the face key points in the image to be processed based on the face key points of the video image of the target person, to obtain multiple first images.

[0092] Based on each first image, obtain the target mask image corresponding to each first image. After superimposing the target mask image on the corresponding first image, obtain the first image with a mask, that is, the second image. As Figure 3 shown, it is the second image after superimposition. The target mask image is actually an image with the same shape as the area to be driven in the first image and with a mosaic. For example: if the lips in the first image are in an open state, the target mask image is a mosaic image that is the same as the open area of the lips in the first image.

[0093] Exemplarily, for each first image, superimpose the target mask image corresponding to the first face key point in the first image on the first image; perform a feathering operation on the target mask image superimposed on the first image, to obtain multiple second images.

[0094] Feathering is a way to transition the superimposed target mask image towards the first image. The larger the feathering value, the wider the blurring range, that is, the softer the color gradient. The smaller the feathering value, the narrower the blurring range. It can be adjusted according to the actual situation. Edge feathering makes the final effect transition more natural.

[0095] Input each second image into an image generation model, and output a target image after filling the target mask. Among them, the face of the target image is the same as the face of the image to be processed. The only difference is that when the face in the image to be processed is in a closed - mouth state, the target image is in an open - mouth state, and when the face in the image to be processed is in a closed - eye state, the target image is in an open - eye state. Based on multiple target images, generate a face driving video of the image to be processed.

[0096] In this application, by determining the image to be processed and the video to be processed, multiple video images of the video to be processed are obtained. Based on the face key points of each video image, the face key points of the image to be processed are adjusted respectively to obtain multiple first images. Among them, each first image contains first face key points. For each first image, the target mask image corresponding to the first face key point in the first image is superimposed on the first image to obtain multiple second images. Each second image is input into a trained image generation model to output multiple target images. Based on each target image, a face-driven video is generated. This application can fill in the missing information in the face driving process, ensure the integrity of the face during the driving process, and avoid the problem of image rupture when driving the face.

[0097] Regarding step 203 above, in another embodiment of this application, as Figure 4 shown, a face driving method is provided, which specifically includes the following steps:

[0098] Step 203-1: For each video image, determine the second face key points in the video image and the first coordinates of each second face key point.

[0099] Step 203-2: Determine the third face key points of the image to be processed and the second coordinates of each third face key point.

[0100] Step 203-3: Adjust the corresponding second coordinates based on the first coordinates to obtain multiple adjusted first images.

[0101] Exemplarily, when it is necessary to drive the lip face key points of the image to be processed, for each video image, through face recognition technology, the lip key points in each video image are determined, and the first coordinates of each lip key point are determined. Also through face recognition technology, the third face key points in the image to be processed are determined, that is, the lip key points and the second coordinates of each lip key point in the image to be processed.

[0102] It should be noted that for each second face key point in each video image, each second face key point corresponds to a third face key point in the image to be processed, that is, each first coordinate of each second face key point corresponds to a second coordinate. Therefore, by adjusting the second coordinates based on the first coordinates and moving the second coordinates towards the first coordinates, the adjusted first image is obtained.

[0103] Regarding step 204 above, in another embodiment of this application, as Figure 5 shown, a face driving method is provided, which specifically includes the following steps:

[0104] Step 204-1: For each first image, determine the first face key points in each first image.

[0105] Step 204-2: Based on the first face key points in each first image, construct multiple first triangular mesh images.

[0106] Step 204-3: Obtain the first mask image of the standard face.

[0107] Step 204-4: Obtain the fourth face key points of the face in the first mask image, and construct a second triangular mesh image based on the fourth face key points.

[0108] Step 204-5: Adjust the second triangular mesh image based on each first triangular mesh image respectively to obtain multiple updated target mask images.

[0109] Step 204-6: Overlay each target mask image with the corresponding first image to obtain multiple second images.

[0110] Based on the face recognition technology, determine the first face key points in each first image, construct a first triangular mesh image based on the first face key points, obtain the first mask image of the standard face, and based on the face recognition technology, determine the fourth face key points in the first mask image. When the first face key points of the first image are the lip edge contour points, the face key points of the first mask image are also the lip edge contour points, and construct a second triangular mesh image based on the face key points of the first mask.

[0111] Exemplarily, when any first triangular mesh image contains triangular meshes numbered 1, 2, and 3, determine the triangular meshes corresponding to the triangular mesh numbered 1, the triangular meshes corresponding to the triangular mesh numbered 2, and the triangular meshes corresponding to the triangular mesh numbered 3 in the second triangular mesh image, adjust the corresponding triangular meshes in the second triangular mesh image based on the triangular mesh numbered 1, adjust the corresponding triangular meshes in the second triangular mesh image based on the triangular mesh numbered 2, and adjust the corresponding triangular meshes in the second triangular mesh image based on the triangular mesh numbered 3. Thus, the adjusted target mask image is obtained. Among them, there is a corresponding relationship between the first triangular mesh image and the target mask image, that is, different first triangular mesh images correspond to different target mask images.

[0112] Exemplarily, for each first triangular network image, determine each first triangular mesh in the first triangular mesh image; for each first triangular mesh, determine the corresponding second triangular mesh, where the second triangular mesh is a triangular mesh in the second triangular mesh image; based on the first triangular mesh, perform an affine transformation on the second triangular mesh to obtain multiple updated target mask images.

[0113] An affine transformation is geometrically defined as an affine transformation or affine mapping between two vector spaces, consisting of a non-singular linear transformation (a transformation using a linear function) followed by a translation transformation.

[0114] Finally, after superimposing each first image with the corresponding target mask image respectively, multiple second images are obtained.

[0115] Superimpose each target mask image with the corresponding first image respectively, so as to fill the closed mouth in the image to be processed with the target mask image, provide an input image for the subsequent input of the image generation model, and thus ensure the generation of the driving image.

[0116] For the training of the image generation model, in another embodiment of the present application, as Figure 6 shown, a face driving method is provided, which specifically includes the following steps:

[0117] Step 301: Obtain the second mask image of the standard face and the face key points in the second mask image.

[0118] Step 302: Based on the face key points in the second mask image, construct a third triangular mesh image.

[0119] Step 303: Determine the image to be trained.

[0120] Step 304: Determine the face key points in the image to be trained, and construct a fourth triangular mesh image based on the face key points of the image to be trained.

[0121] Step 305: Based on each triangular mesh in the fourth triangular mesh image, adjust each triangular mesh in the third triangular mesh image to obtain a third mask image.

[0122] Step 306: Superimpose the third mask image with the image to be trained to obtain a target training image.

[0123] Step 307: Train the image generation model based on the target training image.

[0124] Among them, the input of the image generation model is the target training image, and the output is the training image.

[0125] Obtain the second mask image of the standard face and the image to be trained. Respectively, based on the face recognition technology, determine the face key points in the second mask image and the face key points in the image to be trained, and construct the third triangular mesh image and the fourth triangular mesh image. As Figure 7 shown, it is the third triangular mesh image constructed based on the face key points in the second mask image. As Figure 8 shown, it is the fourth triangular mesh image constructed based on the face key points of the image to be trained.

[0126] Based on each triangular mesh in the fourth triangular mesh image, adjust each triangular mesh in the third triangular mesh image to obtain the third mask image.

[0127] For the process of adjusting the corresponding triangular mesh in the third triangular mesh image by the triangular mesh of the specific fourth triangular mesh image, it is the same as the method of adjusting the second triangular mesh image based on each first triangular mesh image, which will not be elaborated here.

[0128] It should be noted that the image to be trained can be multiple images randomly extracted from a person's video. When there are multiple images to be trained, multiple third mask images are obtained. And there is a corresponding relationship between the image to be trained and the third mask image, that is, different images to be trained correspond to different third mask images.

[0129] Overlay each third mask image with the corresponding image to be trained to obtain multiple target training images. Use the multiple target training images as the training set to train the image generation model, so that the input of the trained image generation model is the target training image carrying the mask image, and the output is the training image, that is, the image after filling the mask.

[0130] Please refer to Figure 9 In addition, this embodiment of the present application also provides a face driving device 110 applied to Figure 1 the electronic device 100 described above. The face driving device 110 includes:

[0131] A determination module 111 and a processing module 112;

[0132] The determination module 111 is used to determine the image to be processed and the video to be processed. Among them, the portrait of the image to be processed is driven based on the video to be processed;

[0133] Obtain multiple frame video images of the video to be processed;

[0134] The processing module 112 is configured to adjust the face key points of the image to be processed based on the face key points of each of the video images respectively, so as to obtain multiple first images, wherein each of the first images includes first face key points;

[0135] For each of the first images, superimpose the target mask image corresponding to the first face key point in the first image on the first image to obtain multiple second images;

[0136] Input each of the second images into the trained image generation model to output multiple target images;

[0137] Generate a face-driven video based on each of the target images.

[0138] Optionally, the processing module 112 is further configured to:

[0139] For each of the video images, determine the second face key points in the video image and the first coordinates of each of the second face key points;

[0140] Determine the third face key points of the image to be processed and the second coordinates of each of the third face key points;

[0141] Adjust the corresponding second coordinates based on the first coordinates to obtain multiple adjusted first images.

[0142] In summary, in this application, by determining the image to be processed and the video to be processed, obtaining multiple video images of the video to be processed, and adjusting the face key points of the image to be processed based on the face key points of each video image respectively, multiple first images are obtained, wherein each first image includes first face key points. For each first image, superimpose the target mask image corresponding to the first face key point in the first image on the first image to obtain multiple second images. Input each of the second images into the trained image generation model to output multiple target images, and generate a face-driven video based on each of the target images. This application can fill in the missing information in the face driving process, ensure the integrity of the face during the driving process, and avoid the problem of image rupture when driving the face.

[0143] This application further provides an electronic device 100, which includes a processor 130 and a memory 120. The memory 120 stores computer-executable instructions, and when the computer-executable instructions are executed by the processor 130, the face driving method is implemented.

[0144] This application embodiment further provides a storage medium, which stores a computer program, and when the computer program is executed by the processor 130, the face driving method is implemented.

[0145] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0146] In addition, each functional module in various embodiments of this application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0147] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0148] As described above, these are only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily conceive of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A face driving method, characterized in that, the method includes: Determine an image to be processed and a video to be processed, wherein, based on the video to be processed, the portrait in the image to be processed is driven; Obtain multiple video images of the video to be processed; For each of the video images, determine the second face key points in the video image and the first coordinates of each of the second face key points; Determine the third face key points in the image to be processed and the second coordinates of each of the third face key points; Adjust the corresponding second coordinates based on the first coordinates to obtain multiple adjusted first images, wherein each of the first images includes first face key points; For each of the first images, superimpose the target mask image corresponding to the first face key points in the first image on the first image to obtain multiple second images, wherein the target mask image is an image with the same shape as the area to be driven in the first image and with mosaics; Input each of the second images into a trained image generation model to output multiple target images; Training the image generation model includes: obtaining a second mask image of a standard face and the face key points in the second mask image; constructing a third triangular mesh image based on the face key points in the second mask image; determining an image to be trained, wherein the image to be trained is multiple images randomly extracted from a person video; determining the face key points in the image to be trained and constructing a fourth triangular mesh image based on the face key points in the image to be trained; adjusting each triangular mesh in the third triangular mesh image based on each triangular mesh in the fourth triangular mesh image to obtain a third mask image; superimposing the third mask image on the image to be trained to obtain a target training image; training the image generation model based on the target training image, wherein the input of the image generation model is the target training image and the output is a training image; Generate a face driving video based on each of the target images.

2. The method according to claim 1, characterized in that, the step of, for each of the first images, superimposing the target mask image corresponding to the first face key points in the first image on the first image to obtain multiple second images, includes: For each of the first images, determine the first face key points in each of the first images; Construct multiple first triangular mesh images based on the first face key points in each of the first images; Obtain a first mask image of a standard face; Obtain the fourth face key points of the face in the first mask image and construct a second triangular mesh image based on the fourth face key points; Adjust the second triangular mesh image based on each of the first triangular mesh images respectively to obtain multiple updated target mask images; Superimpose each of the target mask images on the corresponding first image to obtain multiple second images.

3. The method according to claim 2, characterized in that, The step of respectively adjusting the second triangular mesh image based on each of the first triangular mesh images to obtain multiple updated target mask images includes: For each first triangular network image, determine each first triangular mesh in the first triangular mesh image; For each of the first triangular meshes, determine a corresponding second triangular mesh in the second triangular mesh image, where the second triangular mesh is a triangular mesh in the second triangular mesh image; Based on the first triangular mesh, perform an affine transformation on the second triangular mesh to obtain multiple updated target mask images.

4. The method according to claim 1, wherein, The step of superimposing the target mask image corresponding to the first face key point in the first image on the first image for each of the first images to obtain multiple second images includes: For each of the first images, superimpose the target mask image corresponding to the first face key point in the first image on the first image; Perform a feathering operation on the target mask image superimposed on the first image to obtain multiple second images.

5. A face driving device, wherein, The device includes: a determination module and a processing module; The determination module is configured to determine an image to be processed and a video to be processed, wherein the portrait in the image to be processed is driven based on the video to be processed; Obtain multiple frame video images of the video to be processed; The processing module is configured to, for each of the video images, determine second face key points in the video image and first coordinates of each of the second face key points; Determine third face key points in the image to be processed and second coordinates of each of the third face key points; Based on the first coordinates, adjust the corresponding second coordinates to obtain multiple adjusted first images, where each of the first images includes first face key points; For each of the first images, superimpose the target mask image corresponding to the first face key point in the first image on the first image to obtain multiple second images, where the target mask image is an image with the same shape as the area to be driven in the first image and with mosaics; Input each of the second images into a trained image generation model to output multiple target images; Training the image generation model includes: obtaining a second mask image of a standard face and the facial key points in the second mask image; constructing a third triangular mesh image based on the facial key points in the second mask image; determining the images to be trained, where the images to be trained are multiple images randomly selected from a person video; determining the facial key points in the images to be trained and constructing a fourth triangular mesh image based on the facial key points of the images to be trained; adjusting each triangular mesh in the third triangular mesh image based on each triangular mesh in the fourth triangular mesh image to obtain a third mask image; superimposing the third mask image on the images to be trained to obtain target training images; training the image generation model based on the target training images, where the input of the image generation model is the target training images and the output is the training images; Generating a face-driven video based on each of the target images.

6. An electronic device, characterized in that it includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method according to any one of claims 1-4.

7. A storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Human face three-dimensional reconstruction and human face replacing video editing system and method based on deep learning

    CN107067429A

  • Video generating method and apparatus, electronic device, and storage medium

    WO2021232690A1