Character expression synthesis method and device, electronic equipment and storage medium

By coding, convolution and diffusion processing of the target face sample image, the expression synthesis model is optimized, and the problem of randomness and unpredictability of the diffusion model in character expression synthesis is solved, achieving more accurate and real expression synthesis.

CN120339453APending Publication Date: 2025-07-18PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510430668.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing diffusion model has strong randomness and unpredictability in character expression synthesis, which affects the accuracy of expression synthesis.

Method used

By obtaining the target face sample image, expressive extraction and encoding processing, using two-dimensional convolution and zero convolution processing, combined with diffusion processing, the original expression synthesis model is optimized to generate the target expression synthesis model, and expressive synthesis of the reference character images to improve accuracy.

Benefits of technology

Improve the accuracy and authenticity of character expression synthesis, ensuring that the generated expression image is highly similar to the reference expression image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339453A_ABST
    Figure CN120339453A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a character expression synthesis method and device, electronic equipment and a storage medium, belongs to the technical field of artificial intelligence, and is suitable for financial science and technology scenes. The method comprises the steps of obtaining a target face sample image and performing expression extraction to obtain a sample expression image; performing coding processing, two-dimensional convolution processing, zero convolution processing and diffusion processing on the sample expression image through the original expression synthesis model to obtain a prediction sample face image; performing model optimization on the original expression synthesis model based on the prediction sample face image and the sample expression image to obtain a target expression synthesis model; performing expression extraction on the reference figure image to obtain a reference expression image; performing expression synthesis on the reference expression image through a target expression synthesis model to obtain a target face image; wherein the target face image has the expression of the reference expression image. According to the embodiment of the invention, the accuracy of character expression synthesis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and is applicable to the fintech scenario. In particular, it relates to a method and device for synthesizing human expressions, an electronic device, and a storage medium. Background Art

[0002] Human expression synthesis is an artificial intelligence technology aimed at integrating the expressions of reference persons into candidate pictures to achieve expression transfer. Expression synthesis can be applied to various application scenarios. For example, in the fintech scenario, it can generate human images for business personnel.

[0003] Currently, diffusion models are mainly used to achieve human expression synthesis. However, due to the strong randomness of the human images synthesized by diffusion models, the synthesized expression images may have large differences and unpredictability, thus affecting the accuracy of human expression synthesis.

[0004] Therefore, how to improve the accuracy of human expression synthesis has become an urgent technical problem to be solved. Summary of the Invention

[0005] The main purpose of the embodiments of this application is to propose a method and device for synthesizing human expressions, an electronic device, and a storage medium, aiming to improve the accuracy of human expression synthesis.

[0006] To achieve the above purpose, the first aspect of the embodiments of this application proposes a method for synthesizing human expressions, and the method includes:

[0007] Obtain a target face sample image;

[0008] Extract the expression from the target face sample image to obtain a sample expression image;

[0009] Encode the sample expression image through a preset original expression synthesis model to obtain a sample decoding feature;

[0010] Perform two-dimensional convolution processing on the sample decoding feature through the original expression synthesis model to obtain a sample two-dimensional convolution feature;

[0011] Perform zero convolution processing on the sample two-dimensional convolution feature through the original expression synthesis model to obtain a sample zero convolution feature;

[0012] Perform diffusion processing on the sample zero convolution feature through the original expression synthesis model to obtain a predicted sample face image;

[0013] Optimize the original expression synthesis model based on the predicted sample face image and the sample expression image to obtain a target expression synthesis model;

[0014] Extract the expression from the pre-acquired reference person image to obtain a reference expression image;

[0015] Perform expression synthesis on the reference expression image through the target expression synthesis model to obtain a target face image; wherein, the target face image has the expression of the reference expression image.

[0016] In some embodiments, the model optimization of the original expression synthesis model based on the predicted sample face image and the sample expression image to obtain the target expression synthesis model includes:

[0017] Calculate the loss of the predicted sample face image and the sample expression image to obtain expression synthesis loss data;

[0018] Adjust the parameters of the original expression synthesis model based on the expression synthesis loss data to obtain the target expression synthesis model.

[0019] In some embodiments, the calculating the loss of the predicted sample face image and the sample expression image to obtain expression synthesis loss data includes:

[0020] Extract the expression from the predicted sample face image to obtain predicted expression image data;

[0021] Extract the expression features from the predicted expression image data to obtain predicted expression features;

[0022] Extract the expression features from the sample expression image to obtain sample expression features;

[0023] Calculate the similarity loss between the predicted expression features and the sample expression features to obtain the expression synthesis loss data.

[0024] In some embodiments, the obtaining the target face sample image includes:

[0025] Obtain the original face sample data;

[0026] Perform face detection on the original face sample data to obtain an initial face sample image;

[0027] Perform size screening on the initial face sample image to obtain the target face sample image.

[0028] In some embodiments, the performing size screening on the initial face sample image to obtain the target face sample image includes:

[0029] Perform size analysis on the initial face sample image to obtain sample face size data;

[0030] Filter out invalid sample images from the initial face sample images according to the sample face size data; wherein, the sample face size data of the invalid sample images is greater than a preset maximum threshold of the face region size or less than a preset minimum threshold of the face region size;

[0031] Remove the invalid sample images from the initial face sample images to obtain the target face sample images.

[0032] In some embodiments, the performing face detection on the original face sample data to obtain initial face sample images includes:

[0033] Perform face recognition on the original face sample data to obtain sample face region data and the number of sample faces;

[0034] Perform face number screening on the original face sample data based on the number of sample faces to obtain candidate face sample data;

[0035] Perform face cropping on the candidate face sample data based on the sample face region data to obtain the initial face sample images.

[0036] In some embodiments, the obtaining of the original face sample data includes:

[0037] Obtain basic face sample data; wherein, the basic face sample data includes basic sample images and sample labels, and the sample labels are used to represent whether the basic sample images are valid images or invalid images;

[0038] Perform label screening on the basic sample images based on the sample labels to obtain training sample images; wherein, the training sample images are the valid images;

[0039] Perform resolution screening on the training sample images to obtain the original face sample data.

[0040] To achieve the above object, a second aspect of the embodiments of the present application provides a character expression synthesis device, and the device includes:

[0041] A sample data acquisition module, configured to acquire target face sample images;

[0042] A sample expression extraction module, configured to perform expression extraction on the target face sample images to obtain sample expression images;

[0043] A sample encoding processing module, configured to perform encoding processing on the sample expression images through a preset original expression synthesis model to obtain sample decoding features;

[0044] A sample two-dimensional convolution module for performing two-dimensional convolution processing on the sample decoded features through the original expression synthesis model to obtain sample two-dimensional convolution features;

[0045] A sample zero convolution module for performing zero convolution processing on the sample two-dimensional convolution features through the original expression synthesis model to obtain sample zero convolution features;

[0046] A sample diffusion processing module for performing diffusion processing on the sample zero convolution features through the original expression synthesis model to obtain a predicted sample face image;

[0047] A model optimization module for optimizing the original expression synthesis model based on the predicted sample face image and the sample expression image to obtain a target expression synthesis model;

[0048] A target expression extraction module for extracting an expression from a pre-acquired reference person image to obtain a reference expression image;

[0049] A target expression synthesis module for synthesizing an expression on the reference expression image through the target expression synthesis model to obtain a target face image; wherein, the target face image has the expression of the reference expression image.

[0050] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0051] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0052] The method and device for synthesizing human facial expressions, electronic device, and storage medium proposed in this application obtain a target human face sample image, extract the facial expression from the target human face sample image to obtain a sample facial expression image. Subsequently, the sample facial expression image is subjected to encoding processing, two-dimensional convolutional processing, zero convolutional processing, and diffusion processing through a preset original facial expression synthesis model to obtain a predicted sample human face image. Further, based on the predicted sample human face image and the sample facial expression image, the original facial expression synthesis model is optimized to obtain a target facial expression synthesis model, which can accurately learn facial expression features and perform human facial expression synthesis. Finally, the facial expression of a pre-obtained reference human image is extracted to obtain a reference facial expression image, and the reference facial expression image is subjected to facial expression synthesis through the target facial expression synthesis model to obtain a target human face image. Among them, the target human face image has the facial expression of the reference facial expression image, improving the accuracy and authenticity of human facial expression synthesis. Description of the Drawings

[0053] Figure 1 is a flowchart of the method for synthesizing human facial expressions provided in an embodiment of this application;

[0054] Figure 2 is Figure 1 a flowchart of step S101 in

[0055] Figure 3 is Figure 2 a flowchart of step S201 in

[0056] Figure 4 is Figure 2 a flowchart of step S202 in

[0057] Figure 5 is Figure 2 a flowchart of step S203 in

[0058] Figure 6 is Figure 1 a flowchart of step S107 in

[0059] Figure 7 is Figure 6 a flowchart of step S601 in

[0060] Figure 8 is a schematic structural diagram of the device for synthesizing human facial expressions provided in an embodiment of this application;

[0061] Figure 9 is a schematic hardware structure diagram of the electronic device provided in an embodiment of this application. Detailed Embodiments

[0062] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0063] It should be noted that although functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0065] First, some terms involved in the present application are analyzed:

[0066] Artificial intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also refers to the theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0067] Synthesis of human facial expressions. The synthesis of human facial expressions is an artificial intelligence technology that synthesizes a human face image with a specific expression by analyzing the feature information of the human face, including facial geometric features and texture features, such as the geometric features of the nose and eyes and facial muscle features. The technology of synthesizing human facial expressions usually involves the precise modeling of human face images, using knowledge of facial anatomy to find the movement rules of facial muscles, and decomposing the expression area into several constituent units. By changing the parameters of these expression units, a rich variety of facial expressions can be synthesized, thereby enhancing the realism and emotional expression ability of virtual images. The technology of synthesizing human facial expressions has broad application prospects in the fields of film and television special effects, virtual image production, interactive games, etc. In the fintech scenario, corresponding human images can be generated based on the facial expression characteristics of business personnel for activities such as business promotion.

[0068] The ControlNet model is a neural network technology used to enhance image generation. The ControlNet model achieves more refined control over the content of the generated image by adding an independent control module to the original diffusion model, thereby allowing users to guide the generation process according to specific requirements. The ControlNet model can significantly improve the controllability and quality of image generation while maintaining the stability of the original diffusion model.

[0069] Currently, the diffusion model is mainly used to synthesize human facial expressions. However, due to the strong randomness of the human face images synthesized by the diffusion model, the synthesized expression images may have large differences and unpredictability, thus affecting the accuracy of synthesizing human facial expressions.

[0070] Based on this, the embodiments of this application provide a method and device for synthesizing human facial expressions, an electronic device, and a storage medium, aiming to improve the accuracy of synthesizing human facial expressions.

[0071] The method and device for synthesizing human facial expressions, the electronic device, and the storage medium provided by the embodiments of this application are specifically described through the following embodiments. First, the method for synthesizing human facial expressions in the embodiments of this application is described.

[0072] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0073] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0074] The method for synthesizing human facial expressions provided by the embodiments of this application relates to the field of artificial intelligence technology. The method for synthesizing human facial expressions provided by the embodiments of this application can be applied to a terminal, or to a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the method for synthesizing human facial expressions, etc., but is not limited to the above forms.

[0075] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0076] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0077] Figure 1 is an optional flowchart of the method for synthesizing human facial expressions provided by the embodiments of the present application. Figure 1 The method in may include but is not limited to steps S101 to S109.

[0078] Step S101, obtain a target face sample image;

[0079] Step S102, perform expression extraction on the target face sample image to obtain a sample expression image;

[0080] Step S103, perform encoding processing on the sample expression image through a preset original expression synthesis model to obtain a sample decoding feature;

[0081] Step S104, perform two-dimensional convolution processing on the sample decoding feature through the original expression synthesis model to obtain a sample two-dimensional convolution feature;

[0082] Step S105, perform zero convolution processing on the sample two-dimensional convolution feature through the original expression synthesis model to obtain a sample zero convolution feature;

[0083] Step S106, perform diffusion processing on the sample zero convolution feature through the original expression synthesis model to obtain a predicted sample face image;

[0084] Step S107, optimize the original expression synthesis model based on the predicted sample face image and the sample expression image to obtain a target expression synthesis model;

[0085] Step S108, perform expression extraction on a pre-obtained reference person image to obtain a reference expression image;

[0086] Step S109, perform expression synthesis on the reference expression image through the target expression synthesis model to obtain a target face image; wherein, the target face image has the expression of the reference expression image.

[0087] Steps S101 to S109 illustrated in the embodiments of the present application, by obtaining a target face sample image and performing expression extraction on the target face sample image to obtain a sample expression image; subsequently, encoding processing, two-dimensional convolution processing, zero convolution processing, and diffusion processing are performed on the sample expression image through a preset original expression synthesis model to obtain a predicted sample face image. Further, based on the predicted sample face image and the sample expression image, the original expression synthesis model is optimized to obtain a target expression synthesis model, which can accurately learn expression features and perform human expression synthesis; finally, expression extraction is performed on a pre-obtained reference person image to obtain a reference expression image, and the reference expression image is subjected to expression synthesis through the target expression synthesis model to obtain a target face image; wherein, the target face image has the expression of the reference expression image, improving the accuracy and authenticity of human expression synthesis.

[0088] Please refer to Figure 2 , in some embodiments, step S101 may include but is not limited to steps S201 to S203:

[0089] Step S201, obtain original face sample data;

[0090] Step S202, perform face detection on the original face sample data to obtain an initial face sample image;

[0091] Step S203, perform size screening on the initial face sample image to obtain a target face sample image.

[0092] Steps S201 to S203 illustrated in the embodiments of the present application, by obtaining original face sample data and performing face detection to obtain an initial face sample image; further performing size screening on the initial face sample image to obtain a target face sample image, can effectively remove non-conforming or low-quality samples, reduce data noise, facilitate subsequent accurate face expression synthesis, and improve the accuracy and efficiency of face expression synthesis.

[0093] Please refer to Figure 3 , in some embodiments, step S201 may include but is not limited to steps S301 to S303:

[0094] Step S301, obtain basic face sample data; wherein, the basic face sample data includes a basic sample image and a sample label, and the sample label is used to characterize whether the basic sample image is a valid image or an invalid image;

[0095] Step S302, perform label screening on the basic sample image based on the sample label to obtain a training sample image; wherein, the training sample image is a valid image;

[0096] Step S303: Screen the resolution of the training sample images to obtain the original face sample data.

[0097] Steps S301 to S303 illustrated in the embodiments of the present application obtain the basic face sample data; wherein, the basic face sample data includes basic sample images and sample labels, and the sample labels are used to represent whether the basic sample images are valid images or invalid images. Based on the sample labels, the valid images are screened out from the basic sample images to obtain the training sample images. Finally, the resolution of the training sample images is screened to obtain the original face sample data, which can remove the invalid images, reduce noise interference, ensure the clarity of the images, help the model learn more accurate expression features, and thus improve the training efficiency and accuracy of the human expression synthesis model.

[0098] In step S301 of some embodiments, the basic face sample data is a pre-collected sample data set, which can be obtained from Internet channels or from business scenarios, such as the Laion-5B data set, the images of authorized business personnel, etc. Among them, the basic face sample data includes basic sample images and sample labels, and the sample labels are used to represent whether the basic sample images are valid images or invalid images.

[0099] It should be noted that the valid images and invalid images are specific to the specific application scenario. For example, in the fintech scenario, generating a human image for business personnel for business promotion belongs to a commercial scenario. Some images in the basic face sample data do not meet the requirements of the fintech scenario (such as anime characters, AI pictures, etc.), or the image is not for commercial use, so these images need to be excluded.

[0100] Therefore, for each basic sample image, screening is performed based on the sample label of the basic sample image, and the valid images are screened out from it to improve the quality of the training sample images for human expression synthesis, and avoid potential copyright disputes and privacy leakage risks, ensuring the legality and compliance of data use.

[0101] Furthermore, in step S303 of some embodiments, screening the resolution of the training sample images based on a preset resolution threshold can ensure the clarity of the input image data during model training, and improve the model's recognition ability of expression features and the quality and accuracy of synthesized human expressions.

[0102] Specifically, the preset resolution threshold needs to be set according to the actual application scenario. For example, the resolution threshold is set to 2048*2048, 1024*1024, 512*512, etc., which is not limited thereto.

[0103] Please refer to Figure 4In some embodiments, step S302 may include but is not limited to steps S401 to S403:

[0104] Step S401, performing face recognition on the original face sample data to obtain sample face region data and the number of sample faces;

[0105] Step S402, screening the original face sample data based on the number of sample faces to obtain candidate face sample data;

[0106] Step S403, performing face capture on the candidate face sample data based on the sample face region data to obtain an initial face sample image.

[0107] In steps S401 to S403 shown in the embodiment of the present application, by performing face recognition on the original face sample data, sample face area data and the number of sample faces are obtained; based on the number of sample faces, images with only one face are screened out from the original face sample data to obtain candidate face sample data; finally, face capture is performed on the candidate face sample data based on the sample face area data to obtain an initial face sample image, thereby accurately locating and capturing the face image in the face area, which can simplify the data complexity of the model processing and help enhance the model's recognition and learning capabilities for expression features.

[0108] In step S401 of some embodiments, face recognition is performed on the original face sample data through a pre-trained face recognition model to obtain sample face region data and the number of sample faces. The face recognition model may be a model such as MTCNN, FaceNet, DeepFace, ArcFace, etc., which needs to be selected in combination with the actual application scenario, but is not limited thereto.

[0109] It is understandable that if the original face sample data contains two or more faces, the target expression cannot be determined when the model learns the expression features, resulting in the expression features learned by the model being inaccurate and lacking in specificity, which in turn affects the accuracy and reliability of face recognition or expression analysis.

[0110] Therefore, in step S402 of some embodiments, it is necessary to filter out the data with the number of sample faces being 1 in the initial face sample data to obtain candidate face sample data.

[0111] Finally, face extraction is performed on the candidate face sample data based on the sample face area data, which can remove the rest of the background information in the image except the face, reduce the impact on expression feature learning, and improve the effect of subsequent facial expression synthesis tasks.

[0112] See also Figure 5, in some embodiments, step S303 may further include but is not limited to steps S501 to S503:

[0113] Step S501, perform size analysis on the initial face sample image to obtain sample face size data;

[0114] Step S502, screen out invalid sample images from the initial face sample image according to the sample face size data; wherein, the sample face size data of the invalid sample images is greater than the maximum threshold of the preset face area size or less than the minimum threshold of the preset face area size;

[0115] Step S503, remove the invalid sample images from the initial face sample image to obtain the target face sample image.

[0116] Steps S501 to S503 illustrated in the embodiments of the present application, by performing size analysis on the initial face sample image, screening out invalid sample images (i.e., face images with sizes greater than the maximum threshold of the face area size or less than the minimum threshold of the face area size) according to the preset face area size threshold, and then removing the invalid sample images from the initial face sample image, can ensure that the finally obtained target face sample image has a reasonable size range that meets the requirements of the model data, which helps to improve the training effect of the subsequent face expression synthesis model.

[0117] In some embodiments, the maximum threshold of the preset face area size and the minimum threshold of the preset face area size need to be set according to the input of the model. For example, the maximum threshold of the face area size is set to 512*512, and the minimum threshold of the face area size is set to 128*128, which is not limited thereto.

[0118] In some embodiments, after step S503, the method for synthesizing the facial expression of the person further includes:

[0119] Obtain the number of samples of the target face sample image;

[0120] If the number of samples is less than the preset sample number threshold, repeat steps S201 to S203.

[0121] Specifically, the specific implementation manner is basically the same as the specific implementation manner of the above steps S201 to S203, and will not be elaborated herein.

[0122] In some embodiments, the sample number threshold needs to be set according to the actual application scenario. For example, 200,000 sample images, or 500,000 sample images, which is not limited thereto.

[0123] In step S102 of some embodiments, the expression of the target face sample image can be extracted through a preset expression extraction model. Specifically, the expression extraction model can be models such as FaceNet, ArcFace, etc., and is not limited thereto.

[0124] It can be understood that the sample expression image is equivalent to extracting the expression image from the face, which helps to remove redundant information in the image, thereby improving the model's learning ability of expression features.

[0125] In some embodiments, after step S102 and before step S103, the method for synthesizing human expressions further includes constructing an original expression synthesis model. Among them, the original expression synthesis model is based on the ControlNet model. By using the Micron-Bert model as the expression image encoder for the encoder structure in the ControlNet model, the classification layer and linear layer in the Micron-Bert model have been deleted.

[0126] It can be understood that changing the encoder structure in the ControlNet model to the Micron-Bert model can effectively reduce the computational complexity of the model during human expression synthesis and reduce the consumption of computing resources.

[0127] Furthermore, the output result of the original expression image encoder is input into an original two-dimensional convolutional sub-model (3-layer CONV2D layer) for feature dimension adjustment; then, the output result of the original two-dimensional convolutional layer is input into the original zero convolutional sub-model (the Zeroconvolution layer in the ControlNet model); finally, the result of the original zero convolutional layer is input into the original diffusion sub-model (stable diffusion model) in the ControlNet model, thereby realizing the synthesis of human expressions and improving the accuracy of human expression synthesis.

[0128] In addition, it can also overcome the limitation that the ControlNet model only supports controlling the generation result according to information such as posture and shape, but cannot generate a human face according to expression information.

[0129] It should be noted that during the training process of the original expression synthesis model, the parameters of the expression image encoder are frozen. During the training process, the parameters of the original two-dimensional convolutional sub-model, the original zero convolutional sub-model, and the original diffusion sub-model are mainly adjusted to improve the model's learning ability of expression features, thereby improving the accuracy of human expression synthesis.

[0130] In step S103 of some embodiments, the sample expression image is encoded through the expression image encoder in the original expression synthesis model, and feature representations are learned from the sample expression image to obtain sample decoded features, and the sample encoded features have stronger expressive power.

[0131] In step S104 of some embodiments, since the original two-dimensional convolutional sub-model includes multiple layers of two-dimensional convolutional layers, the sample decoded features are gradually subjected to feature extraction and dimensional transformation through the multiple layers of two-dimensional convolutional layers to obtain the sample two-dimensional convolutional features, which can enhance the representation ability of the encoded features and improve the model's ability to capture expression details.

[0132] In step S105 of some embodiments, the sample two-dimensional convolutional features are subjected to zero convolution processing through the original zero convolution sub-model to obtain the sample zero convolution features, which is beneficial to reducing the computational overhead.

[0133] In step S106 of some embodiments, noise is gradually added to and removed from the sample zero convolution features through the original diffusion sub-model, and finally a predicted sample face image is generated. The expression of the predicted sample face image has a certain similarity to the expression of the sample expression image.

[0134] It should be noted that through the model training process shown in the above steps S103 to S106, the expression features of the target face sample data can be gradually refined and transformed. Finally, while retaining the expression features, a high-quality predicted sample face image is generated, which helps to improve the detail richness and authenticity of the human expression image, and enhance the adaptability and generalization ability of the model to different expression feature changes.

[0135] Please refer to Figure 6 , in some embodiments, step S107 includes but is not limited to steps S601 to S602:

[0136] Step S601, calculate the loss between the predicted sample face image and the sample expression image to obtain the expression synthesis loss data;

[0137] Step S602, adjust the parameters of the original expression synthesis model based on the expression synthesis loss data to obtain the target expression synthesis model.

[0138] Steps S601 to S602 shown in the embodiments of the present application can continuously optimize the model by calculating the loss between the predicted sample face image and the sample expression image and adjusting the parameters of the original expression synthesis model accordingly, so that the generated expression image is closer to the real expression, thereby improving the accuracy and naturalness of expression synthesis.

[0139] Please refer to Figure 7 , in some embodiments, step S601 may include but is not limited to steps S701 to S704:

[0140] Step S701, extract the expression from the predicted sample face image to obtain the predicted expression image data;

[0141] Step S702: Extract facial expression features from the predicted facial expression image data to obtain predicted facial expression features;

[0142] Step S703: Extract facial expression features from the sample facial expression image to obtain sample facial expression features;

[0143] Step S704: Calculate the similarity loss between the predicted facial expression features and the sample facial expression features to obtain facial expression synthesis loss data.

[0144] Steps S701 to S704 shown in the embodiments of the present application, by extracting facial expressions from the predicted sample facial images and further extracting facial expression features to obtain predicted facial expression features; extracting facial expression features from the sample facial expression images to obtain sample facial expression features, and then calculating the similarity loss according to the predicted facial expression features and the sample facial expression features, can accurately quantify the difference between the predicted sample facial images and the sample facial expression images, thereby enhancing the realism and accuracy of the synthesis of any facial expressions and ensuring that the finally generated facial expression images can highly restore the facial expression features in the sample facial expression images.

[0145] In step S701 of some embodiments, the predicted sample facial image can be subjected to facial expression extraction through a preset facial expression extraction model. Specifically, the facial expression extraction model can be models such as FaceNet and ArcFace, and is not limited thereto.

[0146] In step S702 of some embodiments, the predicted facial expression image data can be subjected to facial expression feature extraction through a pre-trained facial expression feature extraction model. Specifically, the facial expression feature extraction model can be algorithms and models such as the LBP algorithm, ResNet model, ASM model, and AAM model. In addition, the facial expression feature extraction model can also be the Mi cron-Bert model, and is not limited thereto.

[0147] In step S703 of some embodiments, the sample facial expression image is subjected to facial expression feature extraction through a facial expression feature extraction model to obtain sample facial expression features. The specific implementation manner is basically the same as that of the above step S702 and will not be repeated.

[0148] In some embodiments, step S704 may include but is not limited to the following steps:

[0149] Perform feature dimension transformation on the predicted facial expression features to obtain the first facial expression features.

[0150] Perform feature dimension transformation on the sample facial expression features to obtain the second facial expression features; wherein, the first facial expression features and the second facial expression features have the same data dimension;

[0151] Calculate the similarity loss between the first facial expression features and the second facial expression features to obtain facial expression synthesis loss data.

[0152] Specifically, the dimensions of the predicted expression features and the sample expression features can be transformed to a preset feature dimension through methods such as dimensionality reduction or regularization. Among them, the preset feature dimension needs to be set according to the design of the model structure, computing resources, and application scenarios. For example, the preset feature dimension can be 1024 dimensions, 512 dimensions, 256 dimensions, etc., which is not limited thereto.

[0153] By calculating the similarity between the predicted expression features and the sample expression features, feature similarity data (with a value range of 0 to 1) is obtained, and 1 is subtracted from the feature similarity data to obtain expression synthesis loss data. The goal of model training is to minimize the expression synthesis loss data, that is, to improve the similarity between the predicted expression features and the sample expression features.

[0154] In step S602 of some embodiments, methods such as backpropagation, gradient descent, and momentum update can be used to adjust the parameters of the original expression synthesis model based on the expression synthesis loss data to obtain the target expression synthesis model. Specifically, the parameters of the original two-dimensional convolutional submodel, the original zero convolutional submodel, and the original diffusion submodel are adjusted to obtain the target two-dimensional convolutional submodel, the target zero convolutional submodel, and the target diffusion submodel. That is, the target expression synthesis model includes an expression image encoder, a target two-dimensional convolutional submodel, a target zero convolutional submodel, and a target diffusion submodel.

[0155] In step S108 of some embodiments, an expression can be extracted from a pre-obtained reference person image through an expression extraction model. Specifically, the expression extraction model can be models such as FaceNet, ArcFace, etc., which is not limited thereto.

[0156] In some embodiments, step S109 may include but is not limited to the following steps:

[0157] The reference expression image is encoded by the expression image encoder in the target expression synthesis model to learn the feature representation from the reference expression image and obtain the target decoded feature;

[0158] In step S104 of some embodiments, since the target two-dimensional convolutional submodel includes multiple layers of two-dimensional convolutional layers, the target decoded feature is gradually subjected to feature extraction and dimensionality transformation through the multiple layers of two-dimensional convolutional layers to obtain the target two-dimensional convolutional feature;

[0159] The target two-dimensional convolutional feature is subjected to zero convolution processing through the target zero convolutional submodel to obtain the target zero convolutional feature;

[0160] Noise is gradually added to and removed from the target zero convolutional feature through the target diffusion submodel, and finally a predicted target face image is generated, and the expression of the predicted target face image is the same as that of the reference expression image.

[0161] It can be understood that after training the target expression synthesis model, expression features can be learned from the reference face images of any person, and based on the expression features, person expression synthesis can be performed, which can be directly applied to multiple scenarios. For example, in the fintech scenario, corresponding person images can be generated based on the expression features of business personnel for activities such as business promotion, enabling personalized promotion and increasing customer stickiness.

[0162] Please refer to Figure 8 , this embodiment of the present application also provides a person expression synthesis device that can implement the above person expression synthesis method. The device includes:

[0163] A sample data acquisition module 801, configured to acquire a target face sample image;

[0164] A sample expression extraction module 802, configured to extract an expression from the target face sample image to obtain a sample expression image;

[0165] A sample encoding processing module 803, configured to perform encoding processing on the sample expression image through a preset original expression synthesis model to obtain a sample decoding feature;

[0166] A sample two-dimensional convolution module 804, configured to perform two-dimensional convolution processing on the sample decoding feature through the original expression synthesis model to obtain a sample two-dimensional convolution feature;

[0167] A sample zero convolution module 805, configured to perform zero convolution processing on the sample two-dimensional convolution feature through the original expression synthesis model to obtain a sample zero convolution feature;

[0168] A sample diffusion processing module 806, configured to perform diffusion processing on the sample zero convolution feature through the original expression synthesis model to obtain a predicted sample face image;

[0169] A model optimization module 807, configured to optimize the original expression synthesis model based on the predicted sample face image and the sample expression image to obtain a target expression synthesis model;

[0170] A target expression extraction module 808, configured to extract an expression from a pre-acquired reference person image to obtain a reference expression image;

[0171] A target expression synthesis module 809, configured to perform expression synthesis on the reference expression image through the target expression synthesis model to obtain a target face image; wherein, the target face image has the expression of the reference expression image.

[0172] The specific implementation manner of this person expression synthesis device is basically the same as the specific embodiment of the above person expression synthesis method, and will not be elaborated herein.

[0173] The embodiments of the present application also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned human expression synthesis method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0174] Please refer to Figure 9 , Figure 9 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0175] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0176] A memory 902, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902, and the processor 901 is used to call and execute the human expression synthesis method of the embodiments of the present application;

[0177] An input / output interface 903, which is used to implement information input and output;

[0178] A communication interface 904, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);

[0179] A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0180] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.

[0181] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned human expression synthesis method is implemented.

[0182] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0183] The method and apparatus for synthesizing human facial expressions, electronic device, and storage medium provided in the embodiments of the present application obtain a target face sample image, extract an expression from the target face sample image to obtain a sample expression image, and then perform encoding processing, two-dimensional convolution processing, zero convolution processing, and diffusion processing on the sample expression image through a preset original expression synthesis model to obtain a predicted sample face image. Further, the original expression synthesis model is optimized based on the predicted sample face image and the sample expression image to obtain a target expression synthesis model, which can accurately learn expression features and perform human facial expression synthesis. Finally, an expression is extracted from a pre-obtained reference human image to obtain a reference expression image, and the reference expression image is subjected to expression synthesis through the target expression synthesis model to obtain a target face image, where the target face image has the expression of the reference expression image, improving the accuracy and authenticity of human facial expression synthesis.

[0184] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0185] Those skilled in the art can understand that the technical solutions shown in the figures do not limit the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.

[0186] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0187] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.

[0188] As used in the description of the present application and the above drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0189] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one)" or a similar expression thereof refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c may be single or multiple.

[0190] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other may be through some interfaces, and the indirect coupling or communication connection of devices or units may be in an electrical, mechanical, or other form.

[0191] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0192] In addition, the functional units in each embodiment of this application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0193] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0194] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, but this does not limit the scope of rights of the embodiments of this application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.

Claims

1. A method for synthesizing human facial expressions, characterized in that, The method includes: Obtain a target face sample image; Extract an expression from the target face sample image to obtain a sample expression image; Perform encoding processing on the sample expression image through a preset original expression synthesis model to obtain sample decoded features; Perform two-dimensional convolution processing on the sample decoded features through the original expression synthesis model to obtain sample two-dimensional convolution features; Perform zero convolution processing on the sample two-dimensional convolution features through the original expression synthesis model to obtain sample zero convolution features; Perform diffusion processing on the sample zero convolution features through the original expression synthesis model to obtain a predicted sample face image; Optimize the original expression synthesis model based on the predicted sample face image and the sample expression image to obtain a target expression synthesis model; Extract an expression from a pre-obtained reference person image to obtain a reference expression image; Perform expression synthesis on the reference expression image through the target expression synthesis model to obtain a target face image; wherein, the target face image has the expression of the reference expression image.

2. The method according to claim 1, wherein The optimizing the original expression synthesis model based on the predicted sample face image and the sample expression image to obtain a target expression synthesis model includes: Calculate a loss between the predicted sample face image and the sample expression image to obtain expression synthesis loss data; Adjust the parameters of the original expression synthesis model based on the expression synthesis loss data to obtain the target expression synthesis model.

3. The method according to claim 2, wherein The calculating a loss between the predicted sample face image and the sample expression image to obtain expression synthesis loss data includes: Extract an expression from the predicted sample face image to obtain predicted expression image data; Extract expression features from the predicted expression image data to obtain predicted expression features; Extract expression features from the sample expression image to obtain sample expression features; Calculate a similarity loss between the predicted expression features and the sample expression features to obtain the expression synthesis loss data.

4. The method according to any one of claims 1 to 3, characterized in that, The obtaining a target face sample image includes: Obtain original face sample data; Perform face detection on the original face sample data to obtain an initial face sample image; Perform size screening on the initial face sample image to obtain the target face sample image.

5. The method according to claim 4, characterized in that, The performing size screening on the initial face sample image to obtain the target face sample image includes: Perform size analysis on the initial face sample image to obtain sample face size data; Screen out invalid sample images from the initial face sample image according to the sample face size data; wherein, the sample face size data of the invalid sample images is greater than a preset maximum threshold of the face region size or less than a preset minimum threshold of the face region size; Remove the invalid sample images from the initial face sample image to obtain the target face sample image.

6. The method according to claim 4, characterized in that The performing face detection on the original face sample data to obtain an initial face sample image includes: Perform face recognition on the original face sample data to obtain sample face region data and the number of sample faces; Perform face quantity screening on the original face sample data based on the quantity of the sample faces to obtain candidate face sample data; Perform face cropping on the candidate face sample data based on the sample face region data to obtain the initial face sample image.

7. The method according to claim 4, wherein The obtaining of the original face sample data includes: Obtain basic face sample data; wherein, the basic face sample data includes a basic sample image and a sample label, and the sample label is used to represent whether the basic sample image is a valid image or an invalid image; Perform label screening on the basic sample image based on the sample label to obtain a training sample image; wherein, the training sample image is the valid image; Perform resolution screening on the training sample image to obtain the original face sample data.

8. A device for synthesizing human expressions, characterized in that, The device includes: A sample data acquisition module, configured to acquire a target face sample image; A sample expression extraction module, configured to perform expression extraction on the target face sample image to obtain a sample expression image; A sample encoding processing module, configured to perform encoding processing on the sample expression image through a preset original expression synthesis model to obtain a sample decoding feature; A sample two-dimensional convolution module, configured to perform two-dimensional convolution processing on the sample decoding feature through the original expression synthesis model to obtain a sample two-dimensional convolution feature; A sample zero convolution module, configured to perform zero convolution processing on the sample two-dimensional convolution feature through the original expression synthesis model to obtain a sample zero convolution feature; A sample diffusion processing module, configured to perform diffusion processing on the sample zero convolution feature through the original expression synthesis model to obtain a predicted sample face image; A model optimization module, configured to perform model optimization on the original expression synthesis model based on the predicted sample face image and the sample expression image to obtain a target expression synthesis model; A target expression extraction module, configured to perform expression extraction on a pre-acquired reference person image to obtain a reference expression image; A target expression synthesis module, configured to perform expression synthesis on the reference expression image through the target expression synthesis model to obtain a target face image; wherein, the target face image has the expression of the reference expression image.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.