Face image processing method, apparatus, device, and storage medium
By acquiring optical flow information through adversarial training and using a generator and discriminator to generate videos with synchronized facial expressions, the problem of the limited processing effect of image processing models in existing technologies is solved, and diverse and realistic facial image processing effects are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2021-08-10
- Publication Date
- 2026-05-22
AI Technical Summary
Existing image processing models can only replace specific two faces, resulting in limited processing effects and failing to provide diverse and realistic facial image processing results.
By acquiring optical flow information from first and second facial sample images of the same target object, and using a generator and discriminator for adversarial training, a video with synchronized facial expressions can be generated, achieving diverse facial image processing.
The generated videos are highly realistic, and the image processing model can drive any face to change expressions, achieving diverse image processing effects.
Smart Images

Figure CN115708120B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for facial image processing. Background Technology
[0002] Currently, for entertainment purposes, internet applications offer features such as face-swapping. To achieve this, an image processing model is typically trained using a large amount of facial image data of both the target and original faces. The trained model can then replace the original face in an image with the target face. However, because this image processing model can only replace two specific faces and cannot provide other processing effects, the results are limited. Therefore, there is an urgent need for a more diverse and realistic facial image processing method. Summary of the Invention
[0003] This application provides a facial image processing method, apparatus, device, and storage medium. The technical solution provided by this application results in videos with high realism and diverse image processing effects from the image processing model. The technical solution is as follows:
[0004] On the one hand, a facial image processing method is provided, which includes:
[0005] Acquire a first face sample image and a second face sample image of the same target object, and acquire first optical flow information, which is used to represent the offset of multiple key points in the first face sample image and the second face sample image;
[0006] The optical flow information prediction model of the image processing model is used to obtain the second optical flow information based on the first optical flow information. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image.
[0007] Based on the first face sample image and the second optical flow information, a predicted image of the first face sample image is generated by the generator of the image processing model.
[0008] The discriminator of the image processing model distinguishes between the predicted image and the second face sample image, obtaining a first discrimination result and a second discrimination result. The first discrimination result is used to indicate whether the predicted image is a real image; the second discrimination result is used to indicate whether the second face sample image is a real image.
[0009] Based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information, the image processing model is trained, and the image processing model is used to process the input image.
[0010] On the one hand, a facial image processing method is provided, which includes:
[0011] Acquire a facial image and a first video, wherein the facial image is a facial image of a first object, and the first video includes multiple facial images of a second object, and the multiple facial images have expression changes;
[0012] The face image and the first video are processed by an image processing model to obtain a second video. The second video includes multiple face images of the first object, and the expression changes of the multiple face images of the first object are the same as the expression changes of the multiple face images in the first video.
[0013] The image processing model is obtained through adversarial training using a first face sample image, a second face sample image, and second optical flow information of the same target object. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image. The second optical flow information is determined based on the first optical flow information, which is used to represent the offset of multiple key points in the first face sample image and the second face sample image.
[0014] On one hand, a facial image processing apparatus is provided, the apparatus comprising:
[0015] The first optical flow acquisition module is used to acquire a first face sample image and a second face sample image of the same target object, and to acquire first optical flow information, which is used to represent the offset of multiple key points in the first face sample image and the second face sample image.
[0016] The second optical flow acquisition module is used to predict the model based on the first optical flow information through the optical flow information of the image processing model, and to acquire the second optical flow information. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image.
[0017] The generation module is used to generate a predicted image of the first face sample image based on the first face sample image and the second optical flow information through the generator of the image processing model.
[0018] The discrimination module is used to discriminate the predicted image and the second face sample image through the discriminator of the image processing model to obtain a first discrimination result and a second discrimination result. The first discrimination result is used to indicate whether the predicted image is a real image; the second discrimination result is used to indicate whether the second face sample image is a real image.
[0019] The model training module is used to train the image processing model based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information. The image processing model is used to process the input image.
[0020] In one possible implementation, the scaling unit is used for:
[0021] Based on the scale difference between the first intermediate feature map and the first face sample image, the scale of the second optical flow information is reduced to obtain the third optical flow information.
[0022] In one possible implementation, the second processing unit is configured to:
[0023] Based on the scale difference between the first intermediate feature map and the first face sample image, the scale of the second optical flow information is reduced, and the offset of each pixel in the scaled second optical flow information is reduced.
[0024] On one hand, a facial image processing apparatus is provided, the apparatus comprising:
[0025] The acquisition module is used to acquire a facial image and a first video. The facial image is a facial image of a first object, and the first video includes multiple facial images of a second object, and the multiple facial images have expression changes.
[0026] The image processing module is used to process the face image and the first video through an image processing model to obtain a second video. The second video includes multiple face images of the first object, and the expression changes of the multiple face images of the first object are the same as the expression changes of the multiple face images in the first video.
[0027] The image processing model is obtained through adversarial training using a first face sample image, a second face sample image, and second optical flow information of the same target object. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image. The second optical flow information is determined based on the first optical flow information, which is used to represent the offset of multiple key points in the first face sample image and the second face sample image.
[0028] On one hand, a computer device is provided, which includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, the computer program being loaded and executed by the one or more processors to implement the facial image processing method.
[0029] On the one hand, a computer-readable storage medium is provided, which stores at least one computer program that is loaded and executed by a processor to implement the facial image processing method.
[0030] On one hand, a computer program product or computer program is provided, which includes program code stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the above-described facial image processing method.
[0031] The technical solution provided in this application determines the optical flow information that represents the pixel offset in the face image based on key points in the first and second face sample images. Based on the optical flow information and the input sample images, adversarial training is achieved, enabling the image processing model to learn the facial expressions in the video to drive facial expression changes. The generated video is highly realistic, and the image processing model can be used to drive any face, achieving diversified image processing effects. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 This is a schematic diagram of the implementation environment of a facial image processing method provided in an embodiment of this application;
[0034] Figure 2 This is a schematic diagram of the training structure of a facial image processing model provided in an embodiment of this application;
[0035] Figure 3 This is a flowchart of a facial image processing method provided in an embodiment of this application;
[0036] Figure 4 This is a flowchart of a facial image processing method provided in an embodiment of this application;
[0037] Figure 5 This is a flowchart of a facial image processing method provided in an embodiment of this application;
[0038] Figure 6 This is a flowchart of a facial image processing method provided in an embodiment of this application;
[0039] Figure 7 This is a schematic diagram illustrating a process for acquiring second optical flow information provided in an embodiment of this application;
[0040] Figure 8 This is a schematic diagram illustrating the generation process of a second intermediate feature map provided in an embodiment of this application;
[0041] Figure 9 This is a schematic diagram illustrating the generation result of a face image provided in an embodiment of this application;
[0042] Figure 10 This is a schematic diagram of the structure of a facial image processing device provided in an embodiment of this application;
[0043] Figure 11 This is a schematic diagram of the structure of a facial image processing device provided in an embodiment of this application;
[0044] Figure 12 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application;
[0045] Figure 13 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0047] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.
[0048] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the celebrity facial images and user-captured facial expression videos involved in this application were obtained with full authorization.
[0049] In this application, the term "at least one" means one or more, and "multiple" means two or more, for example, multiple face images means two or more face images.
[0050] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0051] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0052] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing, tracking, and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content, behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0053] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.
[0054] Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and cryptographic algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying platform, a platform product service layer, and an application service layer.
[0055] Portrait-driven: Given a face image to be driven and a driving video containing a series of expressions and poses, the purpose of portrait-driven is to generate a video that makes the face in the face image to be driven perform the expressions in the driving video.
[0056] The generator is trained using a Generative Adversarial Network (GAN), which consists of a generator and a discriminator. The discriminator takes either real samples or the generator's output as input and aims to distinguish the generator's output from real samples as closely as possible. The generator, on the other hand, tries to deceive the discriminator. The two networks compete against each other, constantly adjusting their parameters until the generator can produce highly realistic images.
[0057] Figure 1 This is a schematic diagram illustrating the implementation environment of a facial image processing method provided in this application embodiment. See also... Figure 1 The implementation environment may include terminal 110 and server 120.
[0058] Optionally, terminal 110 can be a tablet computer, laptop computer, desktop computer, etc., but is not limited to these. The terminal 110 runs an application that supports image processing to process images input by the user or images captured by the user.
[0059] Optionally, server 120 may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.
[0060] The terminal 110 can communicate with the server 120 to use the image processing functions provided by the server 120. For example, the terminal 110 can upload images to the server 120 for processing, and the server 120 can return the image processing results to the terminal 110. It should be noted that the facial image processing method provided in this application embodiment can be executed by either the terminal or the server, and this application embodiment does not limit this.
[0061] Optionally, the aforementioned terminal 110 and server 120 can serve as nodes on a blockchain system to store image processing-related data.
[0062] After introducing the implementation environment of the embodiments of this application, the application scenarios of the embodiments of this application will be introduced below in conjunction with the above implementation environment. It should be noted that in the following description, the facial image processing method provided by the embodiments of this application can be applied to face-driven scenarios. That is, through the facial image processing method provided by the embodiments of this application, when the terminal 110 obtains a face image to be driven and a driving video, a new video can be generated after the above facial image processing process. The face in the new video is the face in the original face image, and the expression of the face changes with the change of the facial expression in the driving video. For example, a platform can provide some celebrity face images, and users can upload their own expression videos on the platform. The platform can generate a dynamic video of the celebrity's face using the celebrity face image and expression video.
[0063] Furthermore, the facial image processing method provided in this application embodiment can also be applied to other facial image processing scenarios, such as animation production scenarios, and this application embodiment does not limit this application.
[0064] In this embodiment, a computer device can implement the facial image processing method provided in this application through an image processing model. The following describes the method in conjunction with... Figure 2 This section provides a brief explanation of the training structure of the image processing model.
[0065] See Figure 2 ,Should Figure 2The diagram illustrates a training structure for an image processing model, which includes a generator 201, a keypoint detector 202, an optical flow information prediction model 203, a discriminator 204, and a loss calculation unit 205. The keypoint detector 202 detects keypoints in a first face sample image F1 and a second face sample image F2. Based on the detected keypoints, the computer device obtains first optical flow information between the first face sample image F1 and the second face sample image F2. The optical flow information prediction model 203 makes predictions based on the first optical flow information to obtain the first optical flow information. The generator 201 processes the acquired first face sample image F1 based on the second optical flow information to obtain a predicted image F3. The discriminator 204 determines whether the input image is generated by the generator or a real image. The loss calculation unit 205 calculates the value of a loss function based on the discriminator 204's judgment result, the predicted image F3, the second face sample image F2, the first optical flow information, and the second optical flow information. Based on the loss function value, the network parameters of the image processing model are updated for the next training iteration. After training, the image processing model including the generator 201, the keypoint detector 202, and the optical flow prediction model 203 can be released as the trained image processing model.
[0066] The overall training process is described below, which includes at least two parts: training the discriminator and training the generator and optical flow prediction model. When training the discriminator, the network parameters of the generator and optical flow prediction model are kept constant. Based on sample images and model processing results, the discriminator's network parameters are adjusted. After adjustment to meet certain conditions, the discriminator's network parameters are kept constant again, and the generator and optical flow prediction model's network parameters are adjusted again based on sample images and model processing results. After adjustment to meet certain conditions, the discriminator is trained again. This alternating training process enables the image processing model to learn the ability to drive the input image based on the input video. Figure 3 This is a flowchart of a facial image processing method provided in an embodiment of this application. Taking a computer device as the execution subject as an example, see [link to flowchart]. Figure 3 The method includes:
[0067] 301. A computer device acquires a first face sample image and a second face sample image of the same target object, and acquires first optical flow information, which is used to represent the offset of multiple key points in the first face sample image and the second face sample image.
[0068] The first and second face sample images are two face sample images of the same target object in the same video. The target object can be a person, an animal, or a virtual character, etc.
[0069] During any iteration, the computer device acquires a pair of sample images from the sample image set, that is, acquires a first face sample image and a second face sample image. In some embodiments, the first face sample image appears before the second face sample image in the driving video.
[0070] In some embodiments, the process of acquiring the sample image set includes: a computer device acquiring a driving video, which is a dynamic video of a target object, wherein the images in the driving video contain the face of the target object, and the facial expressions of the target object are different in multiple frames of the driving video, for example, the facial expressions of the target object in the driving video change over time. After acquiring the driving video, the computer device extracts multiple frames from the driving video and adds the extracted multiple frames to the sample image set for model training.
[0071] 302. The computer device uses the optical flow information prediction model of the image processing model to obtain the second optical flow information based on the first optical flow information. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image.
[0072] The first optical flow information is the optical flow information of key points in the face sample image, and the second optical flow information is the optical flow information of all pixels in the face sample image obtained through prediction. The optical flow information prediction model can predict the optical flow information of multiple pixels in the face sample image based on fewer pixels, that is, the offset of multiple pixels.
[0073] 303. The computer device generates a predicted image of the first face sample image based on the first face sample image and the second optical flow information through the generator of the image processing model.
[0074] Since the second optical flow information is the offset of all pixels in the first face sample image, the image after the offset of the pixels in the first face sample image can be predicted based on the second optical flow information. The training goal of the generator is to make the generated predicted image have the same expression as the second face sample image.
[0075] 304. The computer device uses the discriminator of the image processing model to distinguish between the predicted image and the second face sample image, and obtains a first discrimination result and a second discrimination result. The first discrimination result is used to indicate whether the predicted image is a real image, and the second discrimination result is used to indicate whether the second face sample image is a real image.
[0076] 305. Based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information, the computer device trains the image processing model, which is used to process the input image.
[0077] In this embodiment, the total loss function is calculated by a loss calculation unit, and the network parameters of the image processing model are updated based on the total loss function value. The updated image processing model is then trained in the next iteration.
[0078] The technical solution provided in this application determines the optical flow information that represents the pixel offset in the face image based on key points in the first and second face sample images. Based on the optical flow information and the input sample images, adversarial training is achieved, enabling the image processing model to learn the facial expressions in the video to drive facial expression changes. The generated video is highly realistic, and the image processing model can be used to drive any face, achieving diversified image processing effects.
[0079] The training process involves multiple iterations. Below, we will use only one iteration of this training process as an example to illustrate this facial image processing method. Figure 4 This is a flowchart of a facial image processing method provided in an embodiment of this application. Taking a computer device as the execution subject as an example, see [link to flowchart]. Figure 4 The methods include:
[0080] 401. During the i-th iteration, the computer device acquires a first face sample image and a second face sample image of the same target object, and executes steps 402 and 403. The first face sample image and the second face sample image include faces, where i is a positive integer.
[0081] Step 401 is the same as step 301, and will not be repeated here.
[0082] 402. The computer device inputs the first face sample image into the generator of the image processing model. The generator extracts features from the first face sample image to obtain a first intermediate feature map, and then executes step 407.
[0083] In this embodiment, the generator includes at least one convolutional layer. The generator convolves the first face sample image through this at least one convolutional layer. If the generator includes one convolutional layer, the first face sample image is convolved through this convolutional layer to obtain a first intermediate feature map. If the generator includes two or more convolutional layers, for any one convolutional layer, the convolutional layer convolves the input image or feature map and inputs the convolution result into the next convolutional layer, and so on, with the last convolutional layer outputting the first intermediate feature map. In some embodiments, the generator further includes a pooling layer to pool the feature map output by the convolutional layers to obtain the first intermediate feature map.
[0084] 403. The computer device inputs the first face sample image and the second face sample image into a key point detector, and the key point detector detects the input first face sample image and the second face sample image to obtain the first position of multiple key points in the first face sample image and the second position of the second face sample image.
[0085] In some embodiments, the keypoint detector detects the locations of keypoints in a face sample image. The keypoint detector pre-stores semantic features of multiple keypoints, and based on these semantic features, determines the locations of the multiple keypoints in a first face sample image and a second face sample image, respectively; that is, a first location and a second location. The semantic features represent the characteristics of the keypoint, such as which facial feature it belongs to, its approximate location, and its relationship with surrounding pixels.
[0086] The following explanation uses the detection process of a key point as an example. The semantic feature of this key point is that its grayscale value is significantly higher than that of the surrounding pixels. These surrounding pixels refer to the pixels within a 3×3 matrix region centered on the key point. The grayscale value matrix of this facial sample image is as follows: Based on the semantic features of the key point, the grayscale matrix is traversed to find the pixel that best matches the semantic features. In the example above, the pixel corresponding to the second row and second column of the grayscale matrix is the key point. This process is only an example of key point detection, and the embodiments of this application are not limited thereto.
[0087] 404. The computer device generates first optical flow information based on the first position and the second position, the first optical flow information being used to represent the offset of multiple key points in the first face sample image and the second face sample image.
[0088] For each keypoint in the first facial sample image, a second position of that keypoint in the second facial sample image is determined. The second position of the keypoint is subtracted from its first position to obtain the offset of the keypoint. This offset indicates the direction and amount of the offset. The first and second positions are represented using coordinates in the same coordinate system, and the offset is represented in vector form.
[0089] In some embodiments, the first optical flow information is expressed as a matrix. When the offset is represented in vector form, the matrix includes multiple vectors (which can also be viewed as coordinates), each vector corresponding to a key point. The vector is used to represent the offset direction and offset amount of the key point.
[0090] 405. The computer device inputs the first optical flow information into the optical flow information prediction model, and obtains the second optical flow information based on the first optical flow information through the optical flow information prediction model. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image.
[0091] In this embodiment of the application, obtaining second optical flow information based on the first optical flow information through the optical flow information prediction model includes: processing the first optical flow information through the optical flow information prediction model of the image processing model to obtain the second optical flow information, wherein the scale of the second optical flow information is the same as the scale of the first face sample image.
[0092] The following is based on Figure 7 The process of generating this second optical flow information is explained below: Figure 7 The data includes a first face sample image F1, a second face sample image F2, first optical flow information 701, an optical flow information prediction model 203, and second optical flow information 702. The arrows in the first optical flow information 701 represent the offsets of key points. The first optical flow information 701 is input into the optical flow information prediction model 203, which outputs the second optical flow information 702 based on it. The arrows in the first optical flow information 701 represent the offsets of individual pixels. Observation shows that the number of pixels in the second optical flow information 702 is significantly greater than the number of pixels in the first optical flow information 701; that is, this process can be understood as an estimation process for a dense running field.
[0093] 406. The computer device scales the second optical flow information to obtain third optical flow information with the same scale as the first intermediate feature map.
[0094] The second optical flow information includes the optical flow information of all pixels in the first face sample image. Therefore, the second optical flow information has the same scale as the first face sample image. The first intermediate feature map is obtained by feature extraction from the first face sample image. Therefore, the scale of the first intermediate feature map is different from that of the first face sample image. So, the second optical flow information should be processed first to obtain the third optical flow information with the same scale as the first intermediate feature map. Then, the first intermediate feature map is processed based on the third optical flow information.
[0095] In some embodiments, processing the second optical flow information to obtain the third optical flow information includes: scaling down the second optical flow information according to the scale difference between the first intermediate feature map and the first face sample image to obtain the third optical flow information. In some embodiments, a computer device determines the scale difference between the first intermediate feature map and the first face sample image, for example, determining a scale ratio, and proportionally scaling down the second optical flow information based on the scale difference. This proportional scaling down refers to proportionally reducing the offset of each pixel in the second optical flow information to obtain the third optical flow information. During the above scaling down process, the offset direction of the pixels remains unchanged.
[0096] For example, if the second optical flow information is a 6×6 matrix and the first intermediate feature map is a 3×3 matrix, then for a pixel in the second optical flow information with an offset of (-6, 10) and a scale ratio of 2, after being scaled down proportionally, the offset of that pixel is (-3, 5).
[0097] 407. The computer device inputs the third optical flow information into the generator, and the generator offsets the pixels in the first intermediate feature map based on the third optical flow information to obtain a second intermediate feature map. The second intermediate feature map is then upsampled to obtain a predicted image of the first face sample image.
[0098] In some embodiments, a pixel P in the first intermediate feature map i For example, based on the third optical flow information, the process of offsetting the pixel in the first intermediate feature map includes: the pixel P i The position coordinates of the first intermediate feature map are (x i y i The offset of this pixel is (m) i n i ), for pixel P i The pixel is offset, and the position of the offset pixel is (x... i +m i y i +n iThen, all pixels of the first intermediate feature map are offset as described above to obtain the second intermediate feature map.
[0099] Figure 8 For example, a pixel P in the first intermediate feature map. i For example, this is a schematic diagram of the process of offsetting pixels in the first intermediate feature map based on the third optical flow information.
[0100] In some embodiments, the generator includes at least one transposed convolutional layer. Accordingly, upsampling the second intermediate feature map to obtain a predicted image of the first face sample image includes: the generator performing a transposed convolution on the first face sample image through the at least one transposed convolutional layer to obtain the predicted image of the first face sample image. If the generator includes one transposed convolutional layer, the second intermediate feature map is transposed through this layer to obtain the predicted image. If the generator includes two or more transposed convolutional layers, for any one transposed convolutional layer, the transposed convolutional layer performs a transposed convolution on the input feature map and inputs the transposed convolution result into the next level transposed convolutional layer, and so on, with the last transposed convolutional layer outputting the predicted image. In some embodiments, the transposed convolutional layer can also be an interpolation layer, that is, performing bilinear interpolation on the second intermediate feature image to obtain the predicted image. In some embodiments, the generator further includes a feature concatenation layer and a convolutional layer. The feature concatenation layer is used to concatenate features from the output of the transposed convolutional layer, and the concatenated feature result is input into the convolutional layer. The convolutional layer convolves the input feature concatenation result to obtain the predicted image.
[0101] In some embodiments, the generator employs a U-Net architecture, which is divided into two parts: an encoder and a decoder. The encoder is used to obtain a second intermediate feature map, and the decoder is used to obtain a predicted image.
[0102] The following will use a generator with a U-shaped network architecture as an example to illustrate the image processing flow of the generator: The first face sample image is input into the generator. The encoder of the generator convolves the first face sample image to obtain the first intermediate feature map. Based on the third optical flow information and the first intermediate feature map, the second intermediate feature map is generated. Then, the decoder performs transpose convolution, cropping, feature stitching and other operations on the second intermediate feature map to output the predicted image.
[0103] 408. The computer device uses the discriminator of the image processing model to distinguish between the predicted image and the second face sample image, and obtains a first discrimination result and a second discrimination result. The first discrimination result is used to indicate whether the predicted image is a real image, and the second discrimination result is used to indicate whether the second face sample image is a real image.
[0104] In this embodiment of the application, the discriminator is used to determine whether the input image is an image generated by the generator or a real image.
[0105] In some embodiments, the discrimination result output by the discriminator is represented by a score, which ranges from (0, 1). The higher the score, the more realistic the input image; the lower the score, the less realistic the input image, i.e., the more likely it is an image generated by the generator. Here, the realistic image refers to an image that has not been processed by the generator on the computer device.
[0106] 409. Based on the first discrimination result and the second discrimination result, the computer device obtains the first function value of the first branch function in the total loss function, the first branch function being used to represent the discrimination accuracy of the discriminator for the input image.
[0107] In adversarial training, the discriminator is used to detect the realism of the images generated by the generator, enabling the generator to produce images similar to real images. The first branch function is constructed based on the discriminator's accuracy in predicting the image, its accuracy in predicting the reference image, and the generator's accuracy.
[0108] In this embodiment, the first branch function is expressed by the following formula (1):
[0109] (1)
[0110] in, This represents the first function value of the first branch function. This represents the first facial sample image. This represents the second facial sample image. This indicates the predicted image. This indicates the result of the first discrimination. This represents the second discrimination result, where log represents the logarithmic function used to calculate the discrimination result, and E represents the expected value of the discrimination result, which reflects the magnitude of the average value of the discrimination result. In the training process of a Generative Adversarial Network (GAN), the discriminator is trained first, followed by the generator. The goal of training the discriminator is to maximize the value of a first function. A larger first function value indicates higher accuracy in the discriminator's judgment of the input results, meaning a stronger ability to distinguish between real and generated images. Conversely, the goal of training the generator is to minimize the first function value. A smaller first function value indicates that the generated image is closer to the real image.
[0111] 410. The computer device obtains a second function value of the second branch function in the total loss function based on the predicted image and the second face sample image, the second branch function being used to represent the difference between the predicted image and the second face sample image.
[0112] To determine the accuracy of the predicted image generated by the generator, a second branch function is constructed based on the difference between the predicted image and a second face sample image used as a reference image. The smaller the value of the second function of the second branch function, the smaller the difference between the predicted image and the second face sample image, meaning the stronger the generator's ability to generate an image similar to the real image.
[0113] In this embodiment, the second branch function is expressed by the following formula (2):
[0114] (2)
[0115] in, This represents the result of calculating the perceptual similarity of the second facial sample. This represents the result of calculating the perceptual similarity of the predicted image.
[0116] 411. The computer device determines the third function value of the third branch function in the total loss function based on the first optical flow information and the second optical flow information. The third branch function is used to represent the prediction accuracy of the multiple key points.
[0117] The optical flow prediction model predicts the optical flow information of multiple pixels in the first facial sample image based on the first optical flow information, i.e., the second optical flow information. Since the second optical flow information contains the optical flow information of the multiple key points, the accuracy of the prediction result of the optical flow prediction model can be determined by judging the difference between the multiple key points in the second optical flow information and the first optical flow information. Therefore, the third branch function is constructed based on the first optical flow information and the second optical flow information. The smaller the value of the third function, the smaller the difference between the prediction result and the first optical flow information, i.e., the stronger the prediction ability of the optical flow prediction model.
[0118] In this embodiment, the third branch function is expressed by the following formula (3):
[0119] (3)
[0120] Where n represents the number of key points in the first optical flow information, and n is a positive integer. This represents the position of the key point, where i is a positive integer less than n. This represents the optical flow information of the i-th key point in the first optical flow information. This represents the optical flow information of the i-th key point in the second optical flow information.
[0121] 412. The computer device performs a weighted summation of the first function value, the second function value, and the third function value to obtain the function value of the total loss function.
[0122] The training process of this image processing model includes training the generator, discriminator, and optical flow information prediction model. The three branches of the total loss function can reflect the training status of the above three parts. Based on the total loss function, the network parameters of the image processing model are updated.
[0123] In this embodiment, the total loss function is expressed by the following formula (4):
[0124] (4)
[0125] Where α represents the weight of the second branch function, and β represents the weight of the third branch function. In some embodiments, the specific values of α and β are... .
[0126] 413. If the total loss function value of the i-th iteration or the current iteration meets the training stopping condition, then training is stopped, and the computer device determines the image processing model used in the i-th iteration as the trained image processing model.
[0127] The conditions for stopping training include: the total loss function value converges or the number of iterations reaches a threshold. This application does not limit these conditions.
[0128] 414. If the function value of the total loss function in the i-th iteration or the current iteration does not meet the stopping training condition, then update the network parameters of the image processing model and perform the (i+1)-th iteration training based on the updated image processing model.
[0129] In the above training process, taking the current i+1th iteration as an example of training the discriminator, if the function value of the total loss function in the i-th iteration or the current iteration does not meet the stop training condition, and the first function value of the first branch function does not meet the first condition, the network parameters of the generator and the optical flow information prediction model remain unchanged, and the network parameters of the discriminator in the image processing model are updated. Based on the updated discriminator's network parameters, the (i+1)th iteration of training is performed until the first function value of the first branch function obtained in the j-th iteration satisfies the first condition. Then, the training object is switched, and training begins again from the (j+1)th iteration. The discriminator's network parameters remain unchanged, while the network parameters of the generator and the optical flow prediction model are updated. If the first function value of the first branch function obtained in the (j+1)-th iteration does not satisfy the second condition, training continues for the generator and the optical flow prediction model. If the first function value of the first branch function obtained in the k-th iteration satisfies the second condition, the training object is switched again, and training continues from the (k+1)-th iteration. This process of switching training objects is repeated multiple times to achieve adversarial training. Training stops when the total loss function value or the current iteration satisfies the training stopping condition. Here, i, j, and k are positive integers, and i... <j<k。
[0130] In some embodiments, the image processing model further includes an image enhancement model. The computer device inputs the predicted image into the image enhancement model, which processes the predicted image to obtain an enhanced image with a higher resolution than the predicted image. Based on the enhanced image and the predicted image, the computer device can obtain a second function value of the second branch function of the total loss function. Based on the second function value, the first function value, and the third function value, the total loss function value is obtained. The image processing model is then trained based on this total loss function value. By training with images enhanced by the image enhancement model, the image processing model can output high-quality images, such as high-resolution images, thus improving its processing capabilities.
[0131] The technical solution provided in this application determines the optical flow information that represents the pixel offset in the face image based on key points in the first and second face sample images. Based on the optical flow information and the input sample images, adversarial training is achieved, enabling the image processing model to learn the facial expressions in the video to drive facial expression changes. The generated video is highly realistic, and the image processing model can be used to drive any face, achieving diversified image processing effects.
[0132] Figure 5This is a flowchart of a facial image processing method provided in an embodiment of this application. Taking a computer device as the execution subject as an example, see [link to flowchart]. Figure 5 The method includes:
[0133] 501. A computer device acquires a facial image and a first video, wherein the facial image is a facial image of a first object, and the first video includes multiple facial images of a second object, and the multiple facial images have expression changes.
[0134] The first video is a driving video, which is used to drive the facial image.
[0135] 502. The computer device processes the facial image and the first video using an image processing model to obtain a second video. The second video includes multiple facial images of the first object, and the expression changes of the multiple facial images of the first object are the same as the expression changes of the multiple facial images in the first video.
[0136] The image processing model is obtained through adversarial training using a first face sample image, a second face sample image, and second optical flow information of the same target object. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image. The second optical flow information is determined based on the first optical flow information, which is used to represent the offset of multiple key points in the first face sample image and the second face sample image.
[0137] The image processing model processes the facial image as follows: obtaining a first facial image from a first video based on the facial image; obtaining a second facial image based on the first facial image; the keypoint detector of the image processing model obtains multiple keypoints of the first facial image and the second facial image; obtaining first optical flow information based on the multiple keypoints; the optical flow information prediction model of the image processing model obtains second optical flow information based on the first optical flow information; and the generator of the image processing model generates a predicted image based on the second optical flow information and the facial image.
[0138] The aforementioned computer device processes the face image to be processed based on multiple face images in the first video to obtain multiple predicted images, and generates a second video based on the multiple predicted images.
[0139] In some embodiments, the obtained multiple predicted images are image-enhanced, and a second video is generated based on the image-enhanced predicted images.
[0140] The processing procedure of the above image processing model is the same as that of the key point detector, optical flow information prediction model and generator in the above embodiments, and will not be repeated here.
[0141] After briefly introducing the image processing process, the following section will use a single image processing step as an example to illustrate this facial image processing method. Figure 6 This is a flowchart of a facial image processing method provided in an embodiment of this application. Taking a computer device as the execution subject as an example, see [link to flowchart]. Figure 6 The methods include:
[0142] 601. The computer device acquires a facial image and a first video, and executes steps 602 and 603.
[0143] 602. The computer device inputs the face image into the generator of the image processing model. The generator extracts features from the face image and generates an intermediate feature map. Step 609 is then executed.
[0144] Step 602 is the same as step 402, and will not be described in detail here.
[0145] 603. The computer device determines a first face image in the first video based on the face image, and the facial expression in the first face image matches the facial expression in the face image.
[0146] In some embodiments, a keypoint detector is used to detect keypoints in the face image and multiple images in the first video to obtain keypoints in the face image and multiple images in the first video. Based on the keypoints of the face image, they are matched one by one with the keypoints of the multiple images in the first video to find the frame image in the first video with the highest similarity to the keypoints of the face image, which is then used as the first face image that matches the face image.
[0147] 604. The computer device obtains a second face image from the first video based on a first face image in the first video, the second face image being located after the first face image and corresponding to the same target object.
[0148] In this embodiment of the application, a second face image whose timestamp or image number is located after the first face image is obtained from the first video, based on the timestamp or image number of the first face image.
[0149] 605. The computer device inputs the first face image and the second face image into a key point detector. The key point detector detects the input first face image and the second face image to obtain the first position of multiple key points in the first face image and the second position of multiple key points in the second face sample image.
[0150] Step 605 is similar to step 403 and will not be described in detail here. It should be noted that in some embodiments, the key points of the first face image have already been obtained, so the key point detector can only detect the second face image.
[0151] 606. The computer device generates first optical flow information based on the first position and the second position, the first optical flow information being used to represent the offset of multiple key points in the first face image and the second face image.
[0152] Step 606 is similar to step 404, and will not be repeated here.
[0153] 607. The computer device inputs the first optical flow information into the optical flow information prediction model, and obtains the second optical flow information based on the first optical flow information through the optical flow information prediction model. The second optical flow information is used to represent the offset between multiple pixels in the first face image and multiple pixels in the second face image.
[0154] Step 607 is similar to step 405, and will not be repeated here.
[0155] 608. The computer device scales the second optical flow information to obtain third optical flow information with the same scale as the first intermediate feature map.
[0156] Step 608 is similar to step 406, and will not be repeated here.
[0157] 609. The computer device inputs the third optical flow information into the generator, and the generator offsets the pixels in the intermediate feature map based on the third optical flow information, and upsamples the offset intermediate feature map to obtain the predicted image of the face image.
[0158] Step 609 is similar to step 407, and will not be repeated here.
[0159] 610. If the second face image is the last frame of the first video, the computer device generates the second video based on the generated prediction image.
[0160] In some embodiments, after acquiring any one of the predicted images, the computer device performs image enhancement on the predicted image, and when generating the second video, it uses multiple image-enhanced predicted images to generate the video, thereby improving the video quality.
[0161] By using the aforementioned image processing model to drive facial images, the facial images can exhibit dynamic facial expression changes consistent with the driven image, thereby achieving the goal of diversifying image processing effects and making the facial expressions more realistic.
[0162] See Figure 9 , Figure 9The generation effect of the facial image is demonstrated. Assuming that the first video contains a total of 4 frames, image 901 is used as the first facial image, and images 902, 903 and 904 are used as the second facial images in sequence. Image 905 is driven to obtain images 906, 907 and 908. The expressions of images 906, 907 and 908 are consistent with those of images 902, 903 and 904, respectively. The expression changes shown in the second video generated based on images 906, 907 and 908 are also consistent with the expression changes in images 902, 903 and 904.
[0163] 611. If the second face image is not the last frame of the first video, the computer device updates the second face image and repeats steps 603 to 609 above.
[0164] Figure 10 This is a schematic diagram of a facial image processing device provided in an embodiment of this application. See also... Figure 10 The device includes: a first optical flow acquisition module 1001, a second optical flow acquisition module 1002, a generation module 1003, a discrimination module 1004, and a model training module 1005.
[0165] The first optical flow acquisition module 1001 is used to acquire a first face sample image and a second face sample image of the same target object, and to acquire first optical flow information, which is used to represent the offset of multiple key points in the first face sample image and the second face sample image.
[0166] The second optical flow acquisition module 1002 is used to obtain second optical flow information based on the first optical flow information by using the optical flow information prediction model of the image processing model. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image.
[0167] The generation module 1003 is used to generate a predicted image of the first face sample image based on the first face sample image and the second optical flow information through the generator of the image processing model.
[0168] The discrimination module 1004 is used to discriminate the predicted image and the second face sample image through the discriminator of the image processing model to obtain a first discrimination result and a second discrimination result. The first discrimination result is used to indicate whether the predicted image is a real image; the second discrimination result is used to indicate whether the second face sample image is a real image.
[0169] The model training module 1005 is used to train the image processing model based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information. The image processing model is used to process the input image.
[0170] In one possible implementation, the generation module 1003 includes:
[0171] The scaling unit is used to scale the second optical flow information to obtain third optical flow information with the same scale as the first intermediate feature map;
[0172] The prediction image generation unit is used to extract features from the first face sample image through the generator to obtain a first intermediate feature map, offset the pixels in the first intermediate feature map based on the third optical flow information to obtain a second intermediate feature map, and upsample the second intermediate feature map to obtain a prediction image of the first face sample image.
[0173] In one possible implementation, the scaling unit is used for:
[0174] Based on the scale difference between the first intermediate feature map and the first face sample image, the scale of the second optical flow information is reduced to obtain the third optical flow information.
[0175] In one possible implementation, the model training module 1005 includes:
[0176] The first determining unit is used to determine the function value of the total loss function based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information.
[0177] The update unit is used to update the network parameters of the image processing model when the function value or the current iteration does not meet the training stopping condition, and to perform the next iteration of training based on the updated image processing model;
[0178] The second determining unit is used to determine the image processing model corresponding to the current iteration as the trained image processing model when the function value or the current iteration meets the training stopping condition.
[0179] In one possible implementation, the second processing unit is configured to:
[0180] Based on the scale difference between the first intermediate feature map and the first face sample image, the scale of the second optical flow information is reduced, and the offset of each pixel in the scaled second optical flow information is reduced.
[0181] In one possible implementation, the first determining unit is configured to:
[0182] Based on the first discrimination result and the second discrimination result, the first function value of the first branch function in the total loss function is obtained. The first branch function is used to represent the discrimination accuracy of the discriminator for the input image.
[0183] Based on the predicted image and the second face sample image, the second function value of the second branch function in the total loss function is obtained. The second branch function is used to represent the difference between the predicted image and the second face sample image.
[0184] Based on the first optical flow information and the second optical flow information, the third function value of the third branch function in the total loss function is determined. The third branch function is used to represent the prediction accuracy of the multiple key points.
[0185] The first function value, the second function value, and the third function value are weighted and summed to obtain the function value of the total loss function.
[0186] In one possible implementation, the updating unit is used for:
[0187] If the first function value of the first branch function does not satisfy the first condition, keep the network parameters of the generator and the optical flow information prediction model unchanged, update the network parameters of the discriminator in the image processing model, and perform the next iteration training based on the updated discriminator.
[0188] If the first function value of the first branch function satisfies the first condition, the network parameters of the discriminator remain unchanged, the network parameters of the generator and the optical flow information prediction model in the image processing model are updated, and the next iteration training is performed based on the updated generator and optical flow information prediction model.
[0189] It should be noted that the facial image processing device provided in the above embodiments is only illustrated by the division of the above functional modules when processing facial images. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the facial image processing device and the facial image processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0190] Figure 11 This is a schematic diagram of a facial image processing device provided in an embodiment of this application. See also... Figure 11 The device includes: an acquisition module 1101 and an image processing module 1102.
[0191] The acquisition module 1101 is used to acquire a face image and a first video. The face image is a face image of a first object, and the first video includes multiple face images of a second object, and the multiple face images have expression changes.
[0192] Image processing module 1102 is used to process the face image and the first video through an image processing model to obtain a second video. The second video includes multiple face images of the first object, and the expression changes of the multiple face images of the first object are the same as the expression changes of the multiple face images in the first video.
[0193] The image processing model is obtained through adversarial training using second optical flow information between a first face sample image, a second face sample image, and face sample images of the same target object. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image. The second optical flow information is determined based on the first optical flow information, which is used to represent the offset of multiple key points in the first face sample image and the second face sample image.
[0194] It should be noted that the facial image processing device provided in the above embodiments is only illustrated by the division of the above functional modules when processing facial images. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the facial image processing device and the facial image processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0195] This application provides a computer device for performing the above-described method. This computer device can be implemented as a terminal or a server. The structure of the terminal will be described below:
[0196] Figure 12 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal 1200 can be a tablet computer, a laptop computer, or a desktop computer. The terminal 1200 may also be referred to as a terminal, a portable terminal, a laptop terminal, a desktop terminal, or other names.
[0197] Typically, terminal 1200 includes one or more processors 1201 and one or more memories 1202.
[0198] Processor 1201 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1201 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 1201 may also include a main processor and a coprocessor. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1201 may integrate a Graphics Processing Unit (GPU) for rendering and drawing content required for the display screen. In some embodiments, processor 1201 may also include an AI processor for handling computational operations related to machine learning.
[0199] The memory 1202 may include one or more computer-readable storage media, which may be non-transitory. The memory 1202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1202 are used to store at least one computer program, which is executed by the processor 1201 to implement the facial image processing method provided in the method embodiments of this application.
[0200] In some embodiments, the terminal 1200 may also optionally include a peripheral device interface 1203 and at least one peripheral device. The processor 1201, memory 1202, and peripheral device interface 1203 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1203 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1204, a display screen 1205, a camera assembly 1206, an audio circuit 1207, a positioning assembly 1208, and a power supply 1209.
[0201] Peripheral interface 1203 can be used to connect at least one input / output (I / O) related peripheral device to processor 1201 and memory 1202. In some embodiments, processor 1201, memory 1202 and peripheral interface 1203 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1201, memory 1202 and peripheral interface 1203 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0202] The radio frequency (RF) circuit 1204 is used to receive and transmit radio frequency (RF) signals, also known as electromagnetic signals. The RF circuit 1204 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1204 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1204 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc.
[0203] Display screen 1205 is used to display a user interface (UI). This UI may include graphics, text, icons, video, and any combination thereof. When display screen 1205 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1201 for processing. In this case, display screen 1205 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard.
[0204] The camera assembly 1206 is used to capture images or videos. Optionally, the camera assembly 1206 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal.
[0205] The audio circuit 1207 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input to the processor 1201 for processing, or input to the radio frequency circuit 1204 to realize voice communication.
[0206] The positioning component 1208 is used to locate the current geographical location of the terminal 1200 in order to enable navigation or location-based services (LBS).
[0207] The power supply 1209 is used to supply power to the various components in the terminal 1200. The power supply 1209 can be AC power, DC power, a disposable battery, or a rechargeable battery.
[0208] In some embodiments, the terminal 1200 further includes one or more sensors 1210. The one or more sensors 1210 include, but are not limited to: an accelerometer 1211, a gyroscope 1212, a pressure sensor 1213, a fingerprint sensor 1214, an optical sensor 1215, and a proximity sensor 1216.
[0209] Accelerometer 1211 can detect the magnitude of acceleration on the three coordinate axes of a coordinate system established with terminal 1200.
[0210] The gyroscope sensor 1212 can detect the orientation and rotation angle of the terminal 1200. The gyroscope sensor 1212 can work in conjunction with the accelerometer sensor 1211 to collect the user's 3D movements on the terminal 1200.
[0211] The pressure sensor 1213 can be installed on the side bezel of the terminal 1200 and / or on the lower layer of the display screen 1205. When the pressure sensor 1213 is installed on the side bezel of the terminal 1200, it can detect the user's grip signal on the terminal 1200, and the processor 1201 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1213. When the pressure sensor 1213 is installed on the lower layer of the display screen 1205, the processor 1201 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1205.
[0212] The fingerprint sensor 1214 is used to collect the user's fingerprint. The processor 1201 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 1214, or the fingerprint sensor 1214 identifies the user's identity based on the collected fingerprint.
[0213] The optical sensor 1215 is used to collect ambient light intensity. In one embodiment, the processor 1201 can control the display brightness of the display screen 1205 based on the ambient light intensity collected by the optical sensor 1215.
[0214] The proximity sensor 1216 is used to detect the distance between the user and the front of the terminal 1200.
[0215] Those skilled in the art will understand that Figure 12 The structure shown does not constitute a limitation on terminal 1200 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0216] The aforementioned computer equipment can also be implemented as a server. The structure of a server is described below:
[0217] Figure 13 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1300 can vary significantly due to different configurations or performance. It may include one or more processors 1301 and one or more memories 1302. The one or more memories 1302 store at least one computer program, which is loaded and executed by the one or more processors 1301 to implement the methods provided in the various method embodiments described above. Of course, the server 1300 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1300 may also include other components for implementing device functions, which will not be elaborated upon here.
[0218] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program that can be executed by a processor to perform the facial image processing method described above. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, or an optical data storage device, etc.
[0219] In an exemplary embodiment, a computer program product or computer program is also provided, which includes program code stored in a computer-readable storage medium. The processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the above-described facial image processing method.
[0220] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.
[0221] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0222] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A facial image processing method, characterized in that, The method includes: Acquire a first face sample image and a second face sample image of the same target object, and acquire first optical flow information. The first optical flow information is used to represent the offset of multiple key points in the first face sample image and the second face sample image. The optical flow information prediction model of the image processing model is used to obtain the second optical flow information based on the first optical flow information. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image. Based on the first face sample image and the second optical flow information, a predicted image of the first face sample image is generated by the generator of the image processing model. The discriminator of the image processing model is used to distinguish between the predicted image and the second face sample image to obtain a first discrimination result and a second discrimination result. The first discrimination result is used to indicate whether the predicted image is a real image, and the second discrimination result is used to indicate whether the second face sample image is a real image. Based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information, the image processing model is trained, and the image processing model is used to process the input image.
2. The method according to claim 1, characterized in that, The step of generating a predicted image of the first face sample image based on the first face sample image and the second optical flow information, using the generator of the image processing model, includes: The second optical flow information is scaled to obtain third optical flow information with the same scale as the first intermediate feature map; The generator extracts features from the first face sample image to obtain a first intermediate feature map. Based on the third optical flow information, the pixels in the first intermediate feature map are offset to obtain a second intermediate feature map. The second intermediate feature map is upsampled to obtain a predicted image of the first face sample image.
3. The method according to claim 2, characterized in that, The scaling of the second optical flow information to obtain third optical flow information with the same scale as the first intermediate feature map includes: Based on the scale difference between the first intermediate feature map and the first face sample image, the second optical flow information is scaled down to obtain the third optical flow information.
4. The method according to claim 1, characterized in that, The image processing model is trained based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information. The image processing model is used to process the input image, including: Based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information, the function value of the total loss function is determined; If the function value or the current iteration does not meet the training stopping condition, the network parameters of the image processing model are updated, and the next iteration of training is performed based on the updated image processing model. If the function value or the current iteration satisfies the training stopping condition, the image processing model corresponding to the current iteration is determined as the image processing model that has been trained.
5. The method according to claim 4, characterized in that, The step of determining the value of the total loss function based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information includes: Based on the first discrimination result and the second discrimination result, the first function value of the first branch function in the total loss function is obtained. The first branch function is used to represent the discrimination accuracy of the discriminator for the input image. Based on the predicted image and the second face sample image, the second function value of the second branch function in the total loss function is obtained, and the second branch function is used to represent the difference between the predicted image and the second face sample image; Based on the first optical flow information and the second optical flow information, the third function value of the third branch function in the total loss function is determined, and the third branch function is used to represent the prediction accuracy of the multiple key points; The first function value, the second function value, and the third function value are weighted and summed to obtain the function value of the total loss function.
6. The method according to claim 5, characterized in that, The step of updating the network parameters of the image processing model when the function value or the current iteration does not meet the training stopping condition, and then performing the next iteration of training based on the updated image processing model, includes: If the first function value of the first branch function does not meet the first condition, keep the network parameters of the generator and the optical flow information prediction model unchanged, update the network parameters of the discriminator in the image processing model, and perform the next iteration training based on the updated discriminator. If the first function value of the first branch function satisfies the first condition, keep the network parameters of the discriminator unchanged, update the network parameters of the generator and the optical flow information prediction model in the image processing model, and perform the next iteration training based on the updated generator and optical flow information prediction model.
7. A facial image processing method, characterized in that, The method includes: Acquire a facial image and a first video, wherein the facial image is a facial image of a first object, and the first video includes multiple facial images of a second object, and the multiple facial images have expression changes; The face image and the first video are processed by an image processing model to obtain a second video. The second video includes multiple face images of the first object, and the expression changes of the multiple face images of the first object are the same as the expression changes of the multiple face images in the first video. The image processing model is obtained through adversarial training using a first face sample image, a second face sample image, and second optical flow information of the same target object. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image. The second optical flow information is determined based on the first optical flow information, which is used to represent the offset of multiple key points in the first face sample image and the second face sample image.
8. A facial image processing device, characterized in that, The device includes: The first optical flow acquisition module is used to acquire a first face sample image and a second face sample image of the same target object, and to acquire first optical flow information. The first optical flow information is used to represent the offset of multiple key points in the first face sample image and the second face sample image. The second optical flow acquisition module is used to predict the model through the optical flow information of the image processing model, and to acquire the second optical flow information based on the first optical flow information. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image. The generation module is used to generate a predicted image of the first face sample image based on the first face sample image and the second optical flow information, through the generator of the image processing model. The discrimination module is used to discriminate the predicted image and the second face sample image using the discriminator of the image processing model, and obtain a first discrimination result and a second discrimination result. The first discrimination result is used to indicate whether the predicted image is a real image; the second discrimination result is used to indicate whether the second face sample image is a real image. The model training module is used to train the image processing model based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information. The image processing model is used to process the input image.
9. The apparatus according to claim 8, characterized in that, The generation module includes: The scaling unit is used to scale the second optical flow information to obtain third optical flow information with the same scale as the first intermediate feature map; The prediction image generation unit is used to extract features from the first face sample image through the generator to obtain a first intermediate feature map, offset the pixels in the first intermediate feature map based on the third optical flow information to obtain a second intermediate feature map, and upsample the second intermediate feature map to obtain a prediction image of the first face sample image.
10. The apparatus according to claim 9, characterized in that, The model training module includes: The first determining unit is used to determine the function value of the total loss function based on the first discrimination result, the second discrimination result, the predicted image, the second face sample image, the first optical flow information, and the second optical flow information; The update unit is used to update the network parameters of the image processing model when the function value or the current iteration does not meet the training stopping condition, and to perform the next iteration training based on the updated image processing model. The second determining unit is used to determine the image processing model corresponding to the current iteration process as the trained image processing model when the function value or the current iteration satisfies the training stopping condition.
11. The apparatus according to claim 10, characterized in that, The first determining unit is configured to: Based on the first discrimination result and the second discrimination result, the first function value of the first branch function in the total loss function is obtained. The first branch function is used to represent the discrimination accuracy of the discriminator for the input image. Based on the predicted image and the second face sample image, the second function value of the second branch function in the total loss function is obtained, and the second branch function is used to represent the difference between the predicted image and the second face sample image; Based on the first optical flow information and the second optical flow information, the third function value of the third branch function in the total loss function is determined, and the third branch function is used to represent the prediction accuracy of the multiple key points; The first function value, the second function value, and the third function value are weighted and summed to obtain the function value of the total loss function.
12. The apparatus according to claim 11, characterized in that, The update unit is used for: If the first function value of the first branch function does not meet the first condition, keep the network parameters of the generator and the optical flow information prediction model unchanged, update the network parameters of the discriminator in the image processing model, and perform the next iteration training based on the updated discriminator. If the first function value of the first branch function satisfies the first condition, keep the network parameters of the discriminator unchanged, update the network parameters of the generator and the optical flow information prediction model in the image processing model, and perform the next iteration training based on the updated generator and optical flow information prediction model.
13. A facial image processing device, characterized in that, The device includes: The acquisition module is used to acquire a facial image and a first video, wherein the facial image is a facial image of a first object, and the first video includes multiple facial images of a second object, and the multiple facial images have expression changes. An image processing module is used to process the facial image and the first video through an image processing model to obtain a second video. The second video includes multiple facial images of the first object, and the expression changes of the multiple facial images of the first object are the same as the expression changes of the multiple facial images in the first video. The image processing model is obtained through adversarial training using a first face sample image, a second face sample image, and second optical flow information of the same target object. The second optical flow information is used to represent the offset between multiple pixels in the first face sample image and multiple pixels in the second face sample image. The second optical flow information is determined based on the first optical flow information, which is used to represent the offset of multiple key points in the first face sample image and the second face sample image.
14. A computer device, characterized in that, The computer device includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, the computer program being loaded and executed by the one or more processors to implement the facial image processing method as described in any one of claims 1 to 7.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the face image processing method as described in any one of claims 1 to 7.