Expression processing method, model training method and related devices

By constructing a target processing model of the decoupling training framework, the problem of low accuracy in determining the basis coefficient of facial expressions in the existing technology is solved, high-quality facial expression capture and migration is achieved, and the expression expression effect and realism of virtual images are improved.

CN120070158APending Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311634741.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When determining the basis coefficient of facial expressions in the prior art, the model prediction accuracy is poor, resulting in large errors in the basis of expressions, affecting the expression expression effect and sense of reality of virtual images.

Method used

By constructing a target processing model of the decoupled training framework, the expression basis coefficients, sample images and reconstruction images of expression parameters are used as training data, the coded feature vectors of facial expressions are extracted, and the expression basis coefficients are mapped to obtain the expression basis coefficients, which are used to adjust the virtual image.

Benefits of technology

It improves the expression effect and realism of virtual images, realizes high-quality facial expression capture and migration, and enhances the expressiveness of virtual images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070158A_ABST
    Figure CN120070158A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an expression processing method, a model training method and a related device, at least relates to technologies such as artificial intelligence and the like, can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like, and improves the expression expression effect of a target image. The expression processing method comprises the following steps: acquiring a first processing image comprising a facial expression; and extracting a coding feature vector of the facial expression through the target processing model, mapping the coding feature vector of the facial expression, and performing facial adjustment on the target image based on the expression base coefficient of the facial expression to obtain a target image image containing the facial expression. The target processing model is a machine learning model obtained by training the initial processing model by using an expression base coefficient of the expression parameter, an expression label, a target sample image and a reconstructed sample image, the target sample image comprises the expression parameter of the first sample image and a supplementary parameter of the second sample image, and the supplementary parameter does not comprise the expression parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and particularly to a method for expression processing, a method for model training, and related devices. Background Art

[0002] Facial model parameterization is an important task for constructing and driving virtual avatars, and it is particularly applicable to scenarios such as special effects movies and computer games. In facial model parameterization, the technology of driving virtual avatars by determining facial expression basis coefficients can realistically create various expression animations of virtual avatars, thereby greatly enhancing the realism of virtual avatars, improving the expressiveness, and bringing users a more immersive experience. In the work of driving virtual avatars, an important task is to accurately obtain the facial expression basis coefficients of the virtual avatars.

[0003] However, in the traditional solutions for determining expression basis coefficients, it is usually dependent on algorithms such as emotion driven monocular face capture and animation (EMOCA) and 3D dense face alignment (3DDFA). It is necessary to use PCA bases in three-dimensional videos to form a facial expression basis model, and perform weakly supervised training by means of information such as facial contour feature points and emotion classification that cannot accurately display expressions. The model prediction accuracy is poor, resulting in a large expression basis error, affecting the expression performance effect of virtual avatars, and reducing the realism and expressiveness of virtual avatars. Summary of the Invention

[0004] The embodiments of the present application provide a method for expression processing, a method for model training, and related devices, which are used to accurately determine the expression basis coefficients in an image, so as to improve the expression performance effect, realism, and expressiveness of the target image.

[0005] In a first aspect, an embodiment of the present application provides a method for facial expression processing. The method includes: obtaining a first processed image containing a facial expression; extracting an encoded feature vector of the facial expression based on a target processing model, where the target processing model is a machine learning model obtained by training an initial processing model with the expression basis coefficients of the expression parameters, the expression labels of the first sample images, target sample images, and reconstructed sample images in the first sample images, the target sample images include the expression parameters and supplementary parameters of the second sample images, the supplementary parameters do not include the expression parameters, the expression parameters are used to indicate the expression situation in the first sample images, and the reconstructed sample images are obtained based on the supplementary parameters and the expression basis coefficients of the expression parameters; performing a mapping process on the encoded feature vector of the facial expression based on the target processing model to obtain the expression basis coefficients of the facial expression, where the expression basis coefficients of the facial expression are used to indicate the weights of one or more expression bases in the facial expression; and performing a facial adjustment on a target image based on the expression basis coefficients of the facial expression to obtain a target image containing the facial expression.

[0006] In a second aspect, an embodiment of the present application provides a method for model training. The method includes: obtaining a first sample image containing expression parameters and a second sample image containing supplementary parameters, where the supplementary parameters do not include the expression parameters, and the expression parameters are used to indicate the expression situation in the first sample image; extracting an encoded feature vector of the expression parameters and determining the expression basis coefficients of the expression parameters based on the encoded feature vector of the expression parameters, where the expression basis coefficients of the expression parameters are used to indicate the weights of one or more expression bases in the expression of the first sample image; extracting an encoded feature vector of the supplementary parameters and determining a reconstructed sample image based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters; and training an initial processing model with the expression basis coefficients of the expression parameters, the expression labels of the first sample images, target sample images, and the reconstructed sample images to obtain a target processing model, where the target processing model is used to process a first processed image to obtain the expression basis coefficients of the facial expression in the first processed image, and the expression basis coefficients of the facial expression are used to perform a facial adjustment on a target image to obtain a target image containing the facial expression, and the target sample images include the expression parameters and the supplementary parameters.

[0007] In a third aspect, an embodiment of the present application provides an expression processing device. The expression processing device includes an acquisition module and a processing module. Among them, the acquisition module is used to acquire a first processed image including a facial expression. The processing module is used to extract an encoded feature vector of the facial expression based on a target processing model. The target processing model is a machine learning model obtained by training an initial processing model with the expression basis coefficients of the expression parameters in the first sample image, the expression labels of the first sample image, the target sample image, and the reconstructed sample image. The target sample image includes the expression parameters and supplementary parameters of the second sample image. The supplementary parameters do not include the expression parameters. The expression parameters are used to indicate the expression situation in the first sample image. The reconstructed sample image is obtained based on the supplementary parameters and the expression basis coefficients of the expression parameters. The processing module is used to perform a mapping process on the encoded feature vector of the facial expression to obtain the expression basis coefficients of the facial expression. The expression basis coefficients of the facial expression are used to indicate the weights of one or more expression bases in the facial expression. The processing module is used to perform a facial adjustment on the target image based on the expression basis coefficients of the facial expression to obtain a target image including the facial expression.

[0008] In some alternative embodiments, the processing module is configured to: perform a weighted process based on the expression basis coefficients of the facial expression and the corresponding expression bases to determine the expression results corresponding to the expression bases; perform a facial adjustment on the target image based on the expression results of all the expression bases.

[0009] In some other alternative embodiments, the acquisition module is further configured to acquire a second processed image including facial attributes before performing a facial adjustment on the target image based on the expression basis coefficients of the facial expression to obtain a target image including the facial expression. The facial attributes do not include facial expressions. The processing module is configured to: use the second processed image as the input of the target processing model to obtain the facial coefficients of the facial attributes. The processing module is configured to: perform a facial adjustment on the target image based on the expression basis coefficients of the facial expression and the facial coefficients of the facial attributes to obtain a target image including the facial expression and the facial attributes.

[0010] In some other alternative embodiments, the processing module is configured to: based on the target processing model, the encoded feature vector of the facial attributes of the second processed image; perform a mapping process on the encoded feature vector of the facial attributes based on the target processing model to obtain the facial coefficients of the facial attributes.

[0011] Fourth aspect, an embodiment of the present application provides a model training device. The model training device includes an acquisition unit and a processing unit. Among them, the acquisition unit is used to acquire a first sample image containing expression parameters and a second sample image containing supplementary parameters, where the supplementary parameters do not include the expression parameters, and the expression parameters are used to indicate the expression situation in the first sample image. The processing unit is used to extract the encoded feature vector of the expression parameters and determine the expression basis coefficients of the expression parameters based on the encoded feature vector of the expression parameters. The expression basis coefficients of the expression parameters are used to indicate the weights of one or more expression bases in the expression of the first sample image. The processing unit is used to extract the encoded feature vector of the supplementary parameters and determine a reconstructed sample image based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters. The processing unit is used to train an initial processing model based on the expression basis coefficients of the expression parameters, the expression label of the first sample image, the target sample image, and the reconstructed sample image to obtain a target processing model. The target processing model is used to process a first processed image to obtain the expression basis coefficients of the facial expression in the first processed image. The expression basis coefficients of the facial expression are used to perform facial adjustment on a target image to obtain a target image containing the facial expression. The target sample image includes the expression parameters and the supplementary parameters.

[0012] In some optional embodiments, the acquisition unit is further used to acquire a third sample image containing facial parameters before determining the reconstructed sample image based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters. The facial parameters do not include the expression parameters, and the supplementary parameters also do not include the facial parameters. The processing unit is used to: extract the encoded feature vector of the facial parameters and determine the facial coefficients of the facial attributes based on the encoded feature vector of the facial parameters; determine the reconstructed sample image based on the encoded feature vector of the supplementary parameters, the expression basis coefficients of the expression parameters, and the facial coefficients of the facial attributes.

[0013] In some other optional embodiments, the target processing model is further used to process a second processed image to obtain the facial coefficients of the facial attributes in the second processed image. The facial coefficients of the facial attributes are also used to perform facial adjustment on the target image. The adjusted target image also includes the facial attributes of the second processed image. The facial attributes do not include the facial expression.

[0014] In some other optional embodiments, the processing unit is used to perform a mapping process on the encoded feature vector of the expression parameters based on a preset regression model in the initial processing model to obtain the expression basis coefficients of the expression parameters.

[0015] In some other alternative embodiments, the processing unit is configured to: perform a fusion process on the expression basis coefficients of the expression parameters and the encoded feature vector of the supplementary parameters to obtain a target fusion feature; decode the target fusion feature to obtain a reconstructed sample image;

[0016] In some other alternative embodiments, the processing unit is configured to: calculate the difference between the expression basis coefficients of the expression parameters and the expression label of the first sample image to obtain a first loss value; calculate the image difference between the target sample image and the reconstructed sample image to obtain a second loss value; update the model parameters of the initial processing model based on the first loss value and the second loss value to obtain a target processing model.

[0017] In some other alternative embodiments, the processing unit is configured to: calculate a first similarity distance between the expression basis coefficients of the expression parameters and the expression label of the first sample image; determine a first loss value based on the first similarity distance.

[0018] In some other alternative embodiments, the obtaining unit is further configured to: before calculating the image difference between the target face image and the reconstructed sample image to obtain a second loss value, obtain the pixel value information of the target sample image and the pixel value information of the reconstructed sample image. The processing unit is configured to: extract first pixel information from the pixel value information of the target sample image, where the first pixel information is used to indicate the facial contour condition in the target sample image; calculate the pixel difference between the pixel value information of the reconstructed sample image and the pixel value information of the target sample image to obtain image difference information; determine the second loss value based on the first pixel information and the image difference information.

[0019] In some other alternative embodiments, the processing unit is further configured to: before updating the model parameters of the initial processing model based on the first loss value and the second loss value to obtain a target processing model, extract the encoded feature vector of the expression parameters of the target sample image, and perform a mapping process on the encoded feature vector of the expression parameters of the target sample image to obtain the expression basis coefficients of the expression parameters in the target sample image; calculate the difference between the expression basis coefficients of the expression parameters in the first sample image and the expression basis coefficients of the expression parameters in the target sample image to obtain a third loss value; update the model parameters of the initial processing model based on the first loss value, the second loss value, and the third loss value to obtain a target processing model.

[0020] In some other alternative embodiments, the processing unit is configured to: calculate a second similarity distance between the expression basis coefficients of the expression parameters and the expression basis coefficients of the target sample image; and determine a third loss value based on the second similarity distance.

[0021] A third aspect of the embodiments of the present application provides an expression processing device, including: a memory, an input / output (I / O) interface, and a memory. The memory is used to store program instructions. The processor is configured to execute the program instructions in the memory to perform the expression processing method corresponding to the embodiments of the first aspect above; or, execute the model training method corresponding to the embodiments of the second aspect above.

[0022] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions are run on a computer, the computer is caused to execute the expression processing method corresponding to the embodiments of the first aspect above; or, execute the model training method corresponding to the embodiments of the second aspect above.

[0023] A fifth aspect of the embodiments of the present application provides a computer program product containing instructions. When the computer program product is run on a computer or a processor, the computer or the processor is caused to execute the expression processing method corresponding to the embodiments of the first aspect above; or, execute the model training method corresponding to the embodiments of the second aspect above.

[0024] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0025] In the embodiments of the present application, the expression basis coefficients of expression parameters, the expression labels of the first sample images, the target sample images, and the reconstructed sample images are used as training data to train an initial processing model, so as to obtain a target processing model. The mentioned target sample images include the expression parameters of the first sample images and the supplementary parameters of the second sample images, and the supplementary parameters do not include expression parameters. That is to say, the target sample images are understood to be rendered through expression parameters and supplementary parameters. The mentioned reconstructed sample images are obtained based on the supplementary parameters and the expression basis coefficients of the expression parameters. In this way, after obtaining the first processed image containing facial expressions, the target processing model can be used to extract the encoded feature vector of the first processed image, and the encoded feature vector of the facial expression is mapped to obtain the expression basis coefficients of the facial expression. The expression basis coefficients of the facial expression can indicate the weights of one or more expression bases in the facial expression. In this way, the target image is further adjusted based on the expression basis coefficients of the facial expression, so as to obtain the target image containing the facial expression. Through the above method, in the embodiments of the present application, supplementary parameters that do not include expression parameters are constructed to construct a decoupled training framework (i.e., the target processing model) between the expression parameters and the supplementary parameters. In this way, in the subsequent process of using the target processing model, the facial expressions in the images can be focused on, high-quality image-based facial expression capture can be achieved, and the captured high-quality expression basis coefficients can be transferred to the target image to realize the facial adjustment of the target image, so as to obtain the target image containing the corresponding facial expression in the image, greatly improving the expression effect, realism, and expressiveness of the target image. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained without creative efforts based on these drawings.

[0027] Figure 1 FIG. shows a schematic diagram of an application scenario provided by an embodiment of the present application;

[0028] Figure 2 FIG. shows an alternative schematic diagram of the system framework provided by an embodiment of the present application;

[0029] Figure 3 FIG. shows another alternative schematic diagram of the system framework provided by an embodiment of the present application;

[0030] Figure 4 FIG. shows a schematic flowchart of a method for model training provided by an embodiment of the present application;

[0031] Figure 5 Shows a schematic diagram of a scene of image rendering provided by an embodiment of the present application;

[0032] Figure 6 Shows another schematic flowchart of the method for model training provided by an embodiment of the present application;

[0033] Figure 7 Shows another schematic diagram of a scene of image rendering provided by an embodiment of the present application;

[0034] Figure 8 Shows the schematic flowchart of the method for expression processing provided by an embodiment of the present application;

[0035] Figure 9 Shows another schematic flowchart of the method for expression processing provided by an embodiment of the present application;

[0036] Figure 10 Shows a schematic diagram of a functional module of an expression processing device provided by an embodiment of the present application;

[0037] Figure 11 Shows a schematic diagram of a functional module of a model training device provided by an embodiment of the present application;

[0038] Figure 12 Shows the schematic hardware structure of an expression processing device provided by an embodiment of the present application. Detailed implementation manners

[0039] The embodiments of the present application provide a method for expression processing, a method for model training, and related devices, which are used to accurately determine the expression basis coefficients in an image, so as to improve the expression effect, authenticity, and expressiveness of the target image.

[0040] It can be understood that in the specific implementation manners of the present application, relevant data such as user information is involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0041] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0042] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.

[0043] An important task in constructing a virtual avatar is to parameterize the facial model, which is usually applicable to scenarios such as special effects movies, self-media live broadcasts, and computer games. During the process of facial model parameterization, it is usually necessary to determine the expression basis coefficients of the facial model of the virtual avatar. However, in the traditional scheme for determining the expression basis, the PCA basis in 3D videos is usually used to form the facial expression basis model, and weak supervision training is carried out by means of information such as facial contour feature points and emotion classification that cannot accurately display expressions, resulting in poor model prediction accuracy, large expression basis errors, and affecting the expression performance, realism and expressiveness of the virtual avatar. Moreover, the expression basis determined by this traditional scheme for determining the expression basis cannot be reasonably transferred to any virtual avatar, virtual character and other virtual images.

[0044] Therefore, to solve the above-mentioned technical problems, the embodiments of the present application provide a method for facial expression processing. Correspondingly, the embodiments of the present application also provide a method for model training. Through the method for model training, a target processing model that decouples facial expression parameters and supplementary parameters can be trained. That is to say, in the present application, by constructing supplementary parameters that do not include facial expression parameters, a decoupling training framework (i.e., the target processing model) between facial expression parameters and supplementary parameters is constructed, so that during the subsequent use of the target processing model, the facial expressions in the image can be focused on. In this way, during the process of performing the facial expression processing method in the subsequent model usage stage, first, a first processed image including facial expressions is obtained, and the first processed image is used as the input of the target processing model to extract the encoded feature vector of the facial expression through the target processing model, and the encoded feature vector of the facial expression is mapped and processed to determine the expression basis coefficient of the facial expression. Thus, the target image is adjusted facially according to the expression basis coefficient of the facial expression to obtain a target image including facial expressions. In other words, the embodiments of the present application can achieve high-quality image-based facial expression capture by means of the trained target processing model, and can transfer the captured high-quality expression basis coefficients to the target image, so as to obtain a target image including facial expressions, improving the expression effect, realism, and expressiveness of the facial expressions of the target image.

[0045] Optionally, the expression basis coefficient mentioned in the embodiments of the present application can be understood as the weight coefficient of the expression basis. For example, it can include but is not limited to the weight coefficient of the facial action coding system (FACS) expression basis. Taking the weight coefficient of the FACS expression basis as an example, since the FACS expression basis has expression meanings, such as "left eye stare" and "right eyebrow raise", then by adjusting the target image facially with the weight coefficient of the FACS expression basis, the target image finally obtained can include the facial expression corresponding to the expression basis corresponding to the weight coefficient of the FACS expression basis. That is to say, the facial expression corresponding to the expression basis corresponding to the weight coefficient of the FACS expression basis can be transferred to any virtual image such as a virtual person or a virtual character, or transferred to any real image such as a real person or a real character.

[0046] With the research and progress of artificial intelligence (AI) technology, AI technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, AI technology will be applied in more fields and play an increasingly important role. Therefore, the methods for expression processing and model training provided in the embodiments of this application can be implemented based on artificial intelligence.

[0047] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech technology, natural language processing technology, and machine learning / deep learning.

[0048] In the embodiments of this application, the artificial intelligence technologies mainly involved include the above-mentioned computational vision (image), machine learning / deep learning, etc. For example, it may involve video processing, video semantic understanding, face recognition, etc. in computer vision (CV). Target recognition, target detection and localization, etc. are included in video semantic understanding; face 3D reconstruction, face detection, etc. are included in face recognition.

[0049] The method for expression processing provided in this application can be applied to an expression processing device with data processing capabilities. Exemplarily, the method for model training provided in this application can also be applied to the above-mentioned expression processing device. For example, the described expression processing device may include, but is not limited to, a terminal device, or a server, etc. Among them, the terminal device may include, but is not limited to, a smart phone, a desktop computer, a laptop computer, a tablet computer, a smart speaker, a vehicle-mounted device, a smart watch, a wearable smart device, a smart voice interaction device, a smart home appliance, an aircraft, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, etc. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), as well as big data and artificial intelligence platforms, etc. This application does not make specific limitations. In addition, the mentioned terminal device and server can be directly or indirectly connected through wired communication or wireless communication, etc. This application does not make specific limitations.

[0050] The above-mentioned expression processing device can be capable of implementing the above-mentioned computer vision technology. As a scientific discipline, the described computer vision technology studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pretrained models in the visual field such as swin-transformer, ViT, V-MOE, MAE, etc. can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology generally includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0051] In this application, technologies such as image processing and video processing in computer vision technology can be used to implement the recognition and processing of objects such as videos or images. In the embodiments of this application, the expression processing device can, by implementing the above-mentioned computer vision technology, implement a machine learning model obtained by training an initial processing model based on the expression basis coefficients of expression parameters, the expression labels of the first sample images, the target sample images, and the reconstructed sample images; and based on the target processing model, implement processing such as extraction and mapping of encoded feature vectors of the processed images to obtain the expression basis coefficients of facial expressions and other functions.

[0052] In addition, the expression processing device can also have machine learning capabilities. Machine learning is a multi-disciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as neural networks. In the artificial intelligence models adopted in the methods for expression processing and model training provided in the embodiments of this application, it mainly involves the application of neural networks, and the expression basis coefficients of facial expressions in the processed images are determined through neural networks.

[0053] As a schematic description, Figure 1 shows a schematic diagram of the application scenario provided by the embodiments of this application. As Figure 1 shown, at least object A and object B are included in this application scenario. Among them, by extracting the facial expression of object A, the corresponding expression basis coefficients are determined, such as the weight coefficient of "rolling the eyes" of both eyes. Further, the expression basis coefficients of the facial expression of object A are transferred to the face of object B so that object B can have the expression and other characteristics of object A, such as object B can also show the expression of "rolling the eyes" of both eyes.

[0054] Exemplarily, Figure 2 shows an optional schematic diagram of the system framework provided by the embodiments of this application. As Figure 2 shown, in this system framework, it includes a model training stage and a model usage stage.

[0055] In the model training stage, it is necessary to obtain a first sample image and a second sample image. Moreover, the first sample image includes expression parameters, and the second sample image includes supplementary parameters. The described supplementary parameters do not include the expression parameters. In other words, the first sample image needs to provide the expression parameters of the face in the first sample image; while in the second sample image, other parameters except the expression parameters in the second sample image need to be provided, that is, the supplementary parameters do not need to include the expression parameters. As a schematic description, the supplementary parameters can include but are not limited to the angle of the face, the shape of the face, the facial identity (ID), etc., which are not specifically defined in this application. Through the facial identity, it is possible to identify which user object the face belongs to.

[0056] For example, taking Figure 2 the first sample image shown as a face image (such as face image A) and the second sample image as another face image (such as face image B) as an example, the face image A needs to provide the expression parameters of the face image A, such as "rolling eyes". While in the face image B, it is not necessary to provide the relevant expressions of the face image B (such as "the right corner of the mouth turns up"), but other facial parameters except the expression parameters in the face image B need to be provided, such as providing supplementary parameters such as the facial angle, the facial shape, or the facial identity.

[0057] After obtaining the first sample image and the second sample image, the encoded feature vectors of the expression parameters and the encoded feature vectors of the supplementary parameters are respectively extracted through an encoding module, etc., and the expression basis coefficients of the expression parameters are determined based on the encoded feature vectors of the expression parameters. The mentioned expression basis coefficients of the expression parameters can indicate the weights of one or more expression bases in the expression parameters. For example, the weight of the "right eyebrow raise" expression base in the expression parameters is 0.5, etc., which is not specifically defined.

[0058] It should be noted that the encoding module for extracting the encoded feature vectors of the expression parameters and the encoding module for extracting the encoded feature vectors of the supplementary parameters can be encoding modules with the same model structure or encoding modules with different model structures, which are not specifically defined in this application.

[0059] Further, based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters, a reconstructed sample image is determined through, for example, a decoding model. In addition, it is also necessary to render the target sample image, that is, the target sample image includes the expression parameters of the first sample image and the supplementary parameters of the second sample image. Thus, the initial processing model is trained based on the expression basis coefficients of the expression parameters, the expression labels of the first sample image, the target sample image, and the reconstructed sample image to obtain the target processing model. Optionally, the target processing model can be trained by calculating the loss value between the expression basis coefficients of the expression parameters and the expression labels of the first sample image, and calculating the loss value between the target sample image and the reconstructed sample image.

[0060] In this way, in the model usage stage, after obtaining the first processed image containing facial expressions, the target processing model trained in the above model training stage is used to process the first processed image to obtain the expression basis coefficients of the facial expressions in the first processed image. In this way, the target image is further adjusted based on the expression basis coefficients of the facial expressions to obtain the target image with facial expressions.

[0061] It should be noted that the target image mentioned in the embodiments of the present application may include, but is not limited to, virtual images or real images, etc., which are not limited in the present application. The virtual images mentioned may include, but are not limited to, virtual humans, virtual characters, etc.; the real images described may include, but are not limited to, real people, or other real characters, etc., which are not limited in the present application.

[0062] Optionally, Figure 3 Another optional schematic diagram of the system framework provided by the embodiments of the present application is shown. As Figure 3 shown, on the basis of the system framework shown above, it is also necessary to obtain a third sample image. For this third sample image, it is required to include facial parameters. The facial parameters described do not include expression parameters, and in the second sample image, the second sample image does not need to include the facial parameters accordingly. Figure 2 For example,

[0063] For example, Figure 3Taking the first sample image shown as a face image (e.g., face image A), the second sample image as another face image (e.g., face image B), and the third sample image as another face image (e.g., face image C) as an example, the face image A needs to provide the expression parameters of the face image A, such as "rolling eyes". In the face image C, the facial identifier of the face image C (i.e., the facial ID of the face image C) can be provided. While in the face image B, the relevant expressions such as the expression of the face image B (e.g., "mouth open") do not need to be provided, nor the facial identifier of the face image B, but other facial parameters except the expression parameters and the facial identifier in the face image B are provided, such as supplementary parameters like facial angle and facial shape.

[0064] Similarly, after obtaining the third sample image, the encoded feature vector of the facial parameters needs to be extracted through an encoding module, etc., and the facial coefficient of the facial parameters in the third sample image is determined based on the encoded feature vector of the facial parameters. Through the facial coefficient of the facial parameters, the weight of the facial parameters can be indicated.

[0065] It should be noted that the encoding module for extracting the encoded feature vector of the facial parameters, the encoding module for extracting the encoded feature vector of the facial parameters, and the encoding module for extracting the encoded feature vector of the supplementary parameters can be encoding modules with the same model structure or encoding modules with different model structures, which are not specifically limited in this application.

[0066] Furthermore, through, for example, a decoding model, etc., a reconstructed sample image is determined based on the encoded feature vector of the supplementary parameters, the expression basis coefficient of the expression parameters, and the facial coefficient of the facial parameters in the third sample image. In addition, it is also necessary to render a target sample image based on the expression parameters of the first sample image, the supplementary parameters of the second sample image, and the facial parameters of the third sample image. Thus, the initial processing model is trained based on the expression basis coefficient of the expression parameters, the expression label of the first sample image, the target sample image, and the reconstructed sample image to obtain the target processing model.

[0067] In this way, in the model usage stage, not only the first processed image containing facial expressions needs to be obtained, but also the second processed image of the facial attributes needs to be obtained. It should be noted that the facial attributes mentioned here do not include facial expressions. Further, with the help of the above Figure 3 The target processing model trained is used to process the first processed image to obtain the expression basis coefficient of the facial expression in the first processed image. Similarly, with the help of the above Figure 3The trained model processes the second processed image to obtain the facial coefficients of the facial attributes in the second processed image. In this way, based on the expression basis coefficients of the facial expression and the facial coefficients of the facial attributes, the target image is adjusted facially, so as to obtain the target image including the facial expression and the facial attributes. That is to say, the adjusted target image includes the facial expression of the first processed image and the facial attributes in the second processed image.

[0068] It should be noted that the facial coefficients mentioned in the embodiments of the present application can also be understood as the weight coefficients of the facial attributes. For example, when the facial attribute is the facial shape, the facial coefficient can also be understood as the weight coefficient of the facial shape, and can also be called the face pinching basis coefficient. Or, when the facial attribute is the facial angle, the facial coefficient can also be understood as the weight coefficient of the facial angle, etc., which will not be elaborated here.

[0069] In other words, a decoupled training framework between the expression parameters and the supplementary parameters is constructed (i.e., the Figure 2 shown target processing model), or a decoupled training framework between the expression parameters, the facial parameters, and the supplementary parameters is constructed (i.e., the Figure 3 shown target processing model), so that in the subsequent process of using the target processing model, the facial expression in the image can be focused on, or other facial parameters can also be concerned. In this way, after constructing the target processing model, high-quality facial expression capture based on images can be realized by means of the target processing model, and the captured high-quality expression basis coefficients, or other facial coefficients can be migrated to the target image to obtain the target image including the facial expression, or the target image including the facial expression and the facial attributes, so as to improve the expression performance of the target image, as well as the realism and expressiveness of the target image.

[0070] It should be noted that the training images provided above, such as but not limited to the first sample image, the second sample image, the third sample image, etc., can be obtained by taking pictures of the face of an object, or can also be images intercepted from the face of an object in a video, etc., which are not specifically limited in the embodiments of the present application.

[0071] Similarly, the processed images provided above, such as but not limited to the first processed image and the second processed image, can also be obtained by taking pictures of the face of an object, or can also be images intercepted from the face of an object in a video, etc., which are not specifically limited in the embodiments of the present application.

[0072] In addition, regarding the first sample image, the second sample image, and the third sample image mentioned above, they may include different facial images of the same object, or may also include facial images of different objects. Specifically, no limitation is made in this application. The differences among the first sample image, the second sample image, and the third sample image are as follows: the first sample image includes expression parameters, the third sample image includes facial parameters (excluding expression parameters), and the second sample image contains other supplementary parameters except for expression parameters and facial parameters. Specifically, no specific limitation is made on the human object in these three sample images in this application.

[0073] Regarding the first processed image and the second processed image mentioned above, they may include different facial images of the same object, or may also include facial images of different objects. The difference between the first processed image and the second processed image is as follows: the first processed image includes facial expressions, and the second processed image includes facial attributes other than facial expressions. Specifically, no specific limitation is made on the human object in these two processed images in this application.

[0074] In addition, the facial shape mentioned above can also be understood as the facial contour. Taking a human face as an example, the facial shape can also be called the face shape. According to morphological classification, the face shape can be divided into the following ten types: (1) round face shape; (2) oval face shape; (3) ovoid face shape; (4) inverted ovoid face shape; (5) square face shape; (6) rectangular face shape; (7) trapezoidal face shape; (8) inverted trapezoidal face shape; (9) rhomboid face shape; (10) pentagonal face shape. According to glyph classification, the face shape can be divided into the following eight types: (1) national character face shape; (2) eye character face shape; (3) field character face shape; (4) oil character face shape; (5) vertical middle character face shape; (6) armor character face shape; (7) use character face shape; (8) wind character face shape. The contours in the embodiments of this application include but are not limited to the above types, and no specific limitation is made here.

[0075] Exemplarily, the method for expression processing and the method for model training based on artificial intelligence provided in the embodiments of this application can both be applied to various application scenarios for adjusting the face of a target image. For example, it can be applied in news broadcasts, weather forecasts, game commentaries, and game scenarios where it is allowed to construct game characters with the same face shape as the user's own; it can also be used in scenarios where virtual images are used to undertake personalized services, such as one-on-one services for personal face-to-face psychological doctors, virtual assistants, etc.; or, it can also be applicable to scenarios such as self-media live broadcasts. In these scenarios, by using the method provided in the embodiments of this application, the expression basis coefficients can be determined, and then a face model can be constructed based on the expression basis coefficients, so as to realize the adjustment of the face of the target image based on the face model and the expression basis coefficients.

[0076] As a schematic description, facial adjustment of the target image may include facial adjustment of the target image in the video or facial adjustment of the target image in a single image, which is not specifically limited in this application.

[0077] For example, taking a video as an example, facial adjustment of the target image in the video based on the expression basis coefficients of facial expressions can make the target image in each frame of the video include facial expressions, so as to obtain a target image with facial expressions in each frame. In this way, a driving video can also be generated based on the obtained target image with facial expressions.

[0078] Or, taking a single image as an example, facial adjustment of the target image in the single image based on the expression basis coefficients of facial expressions can make the target image in the single image include facial expressions, so as to obtain a target image with facial expressions.

[0079] Exemplarily, the model training method and the expression processing method provided in this application can also be applied to scenarios such as cloud technology, artificial intelligence, intelligent transportation, assisted driving, big data, etc., which are not specifically limited in this application.

[0080] As a schematic description, due to the execution of the above-described expression processing method, it is necessary to rely on the target processing model trained by the method of model training in the early stage. Therefore, first, from the perspective of embodiments, the model training method provided in the embodiments of this application will be described in detail. Exemplarily, Figure 4 shows a schematic flowchart of the model training method provided in the embodiments of this application. As Figure 4 shown, the model training method at least includes the following steps:

[0081] 401. Obtain a first sample image containing expression parameters and a second sample image containing supplementary parameters. The supplementary parameters do not include expression parameters, and the expression parameters are used to indicate the expression situation in the first sample image.

[0082] In this example, the first sample image may include but is not limited to a face image, an animal face image, or other facial images, etc. In this application, only a face image is taken as an example for illustration. The described second sample image may also include but is not limited to a face image, an animal face image, a cartoon character image, or other facial images, etc. In the embodiments of this application, only a face image is also taken as an example for illustration.

[0083] For the first sample image and the second sample image, the difference is that the first sample image includes expression parameters, and the second sample image includes supplementary parameters. The mentioned supplementary parameters do not include expression parameters. For example, the supplementary parameters may include but are not limited to facial shape, facial angle, light, background, etc., and specific reference can be made to the foregoingFigure 2 Or Figure 3 For the supplementary parameters shown in Figure 3 , they are understood here and will not be elaborated.

[0084] In addition, through the expression parameters, the expression situation in the first sample image can be obtained. For example, "the left corner of the mouth is upturned" or "the eyes are widened", etc., which are not limited here.

[0085] 402. Extract the encoded feature vector of the expression parameter, and determine the expression basis coefficient of the expression parameter based on the encoded feature vector of the expression parameter.

[0086] In this example, after obtaining the first sample image, the expression parameter of the first sample image can be encoded through an encoding module such as an expression encoding module in the initial processing model to extract the encoded feature vector of the expression parameter. For example, one or more convolutional layers in the encoding module can be used to perform convolutional processing on the expression parameter of the first sample image to encode and obtain the encoded feature vector of the expression parameter in the first sample image.

[0087] After extracting the encoded feature vector of the expression parameter, the expression basis coefficient of the expression parameter can be determined based on the encoded feature vector of the expression parameter. For example, the encoded feature vector of the expression parameter can be used as the input of a preset regression model in the initial processing model, so as to perform mapping processing on the encoded feature vector of the expression parameter through the preset regression model to obtain the expression basis coefficient of the expression parameter. The expression basis coefficient of the described expression parameter can sometimes also be called the expression weight coefficient, which can be used to reflect the weight of one or more expression bases in the expression parameter in the first sample image; or, it can also be understood as the proportion of the weight of one or more expression bases in the expression parameter.

[0088] It should be noted that the encoding module and the preset regression model mentioned here can be coupled to the initial processing model and used as encoding sub-modules, regression sub-modules, etc. in the initial processing model, which are not specifically limited. Or, the encoding module and the preset regression model mentioned here can also exist independently of the initial processing model in actual applications, which are not specifically limited.

[0089] 403. Extract the encoded feature vector of the supplementary parameter, and determine the reconstructed sample image based on the encoded feature vector of the supplementary parameter and the expression basis coefficient of the expression parameter.

[0090] In this example, after obtaining the second sample image, the supplementary parameter of the second sample image can also be encoded through an encoding module such as an encoding module to obtain the encoded feature vector of the supplementary parameter. For example, one or more convolutional layers in the encoding module can be used to perform convolutional processing on the supplementary parameter of the second sample image to encode and obtain the encoded feature vector of the supplementary parameter in the second sample image.

[0091] Thus, after determining the expression basis coefficients of the expression parameters in step 402 and the encoded feature vectors of the supplementary parameters in step 403, image reconstruction can be performed based on the encoded feature vectors of the supplementary parameters and the expression basis coefficients of the expression parameters. As a schematic description, the expression basis coefficients of the expression parameters and the encoded feature vectors of the supplementary parameters are fused to obtain a target fusion feature. In this way, after obtaining the target fusion feature, the target fusion feature is decoded through, for example, a decoder in the initial processing model to obtain a reconstructed sample image. For example, one or more deconvolution layers in the decoder can be used to perform deconvolution calculations on the target fusion feature to decode and obtain the reconstructed sample image.

[0092] It should be noted that the decoder mentioned here can be coupled to the initial processing model and serve as a decoding sub-model in the initial processing model, etc., without specific limitation. Or, the decoder mentioned here can also exist independently of the initial processing model in actual applications, without specific limitation.

[0093] 404. Train the initial processing model based on the expression basis coefficients of the expression parameters, the expression label of the first sample image, the target sample image, and the reconstructed sample image to obtain a target processing model, where the target sample image includes expression parameters and supplementary parameters.

[0094] In this example, in order to verify whether the image of the reconstructed sample image generated in step 403 is accurate and determine whether there is a large deviation from the actual image in the subsequent steps, before training the model in this application, it is also necessary to render a target sample image based on the first sample image and the second sample image. In this way, the target sample image is used as the image label of the previously generated reconstructed sample image, and then the difference between the two is calculated.

[0095] As a schematic description, the expression parameters of the first sample image and the supplementary parameters of the second sample image can be rendered to obtain a target sample image that includes both expression parameters and supplementary parameters. Or, the target sample image and the first sample image include the same expression parameters, and the target sample image and the second sample image include the same supplementary parameters except for the expression.

[0096] For example, Figure 5 shows a schematic diagram of a scene of image rendering provided by an embodiment of the present application. As Figure 5As shown, the scenario at least includes a first sample image and a second sample image. Taking the first sample image as face image A and the second sample image as face image B as an example, by rendering the expression parameters of face image A and the supplementary parameters of face image B, the target sample image, i.e., face image Q, is constructed and rendered.

[0097] In the above Figure 5 , face image Q has exactly the same facial expression as face image A, and face image Q has exactly the same supplementary parameters of other parts of the face except the facial expression as face image B, such as the facial identifier of face image B, the angle, light, focal length or background of face image B, etc.

[0098] In this way, after rendering the target sample image, the expression basis coefficients of the expression parameters, the expression labels of the first sample image, the target sample image, and the reconstructed sample image can be used as training data. Furthermore, the expression basis coefficients of the expression parameters, the expression labels of the first sample image, the target sample image, and the reconstructed sample image are used as the input of the initial processing model to train the initial processing model, so as to obtain the target processing model through training.

[0099] As a schematic description, since it is desired that the output of the deep neural network is as close as possible to the value that is truly desired to be predicted, the weight vectors of each layer of the neural network can be updated by comparing the predicted value of the current network with the truly desired target value and then according to the difference between the two (of course, there is usually an initialization process before the first update, that is, parameters are preconfigured for each layer in the deep neural network). For example, if the predicted value of the network is high, the weight vector is adjusted to make it predict lower, and continuous adjustment is made until the neural network can predict the truly desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function, and they are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0100] Therefore, during the specific training process, a loss function can be added synchronously to improve the learning ability of the processing model. Exemplarily, the difference between the expression basis coefficients of the expression parameters and the expression labels of the first sample image can be calculated to obtain a first loss value. As a schematic description, specifically, the first similarity distance between the expression basis coefficients of the expression parameters and the expression labels of the first sample image can be calculated, and then the first similarity distance can be determined as the first loss value. The described first similarity distance can include but is not limited to the Euclidean distance, cosine similarity distance, etc., and no specific limitation is made.

[0101] Taking the Euclidean distance as an example, the second norm of the expression basis coefficients of the expression parameters and the expression labels of the first sample image can be solved to obtain the corresponding Euclidean distance, and then the corresponding first loss value can be obtained. For example, the above-mentioned first loss value satisfies where represents the expression basis coefficient of the i-th expression parameter, represents the expression label of the i-th first sample image, and i is an integer greater than or equal to 1.

[0102] Similarly, it is also necessary to calculate the image difference between the target sample image and the reconstructed sample image to obtain a second loss value. As a schematic description, during the process of calculating the second loss value, the pixel value information of the target sample image and the pixel value information of the reconstructed sample image can be obtained, and the pixel difference between the pixel value information of the target sample image and the pixel value information of the reconstructed sample image can be calculated to obtain the image pixel difference. In addition, it is also necessary to extract first pixel information from the pixel value information of the target sample image. Through this first pixel information, the facial contour situation in the target sample image can be reflected. In this way, after calculating the image pixel difference and the first pixel information, the second loss value is calculated based on the image pixel difference and the first pixel information.

[0103] For example, a matrix can be used to describe the corresponding image pixel values, etc. For example, using matrix I rec to represent the pixel value information of the reconstructed sample image, using matrix I c to represent the pixel value information of the target sample image, and matrix M c to represent the first pixel information. At this time, the calculated second loss value satisfies L rec = ||M c ⊙ (I rec - I c ) || 1 . Among them, I rec - I c represents the image pixel difference, and the operator ⊙ represents the Hadamard product.

[0104] After calculating the first loss value and the second loss value, update the model parameters of the initial processing model based on the first loss value and the second loss value. For example, calculate the sum of the loss values between the first loss value and the second loss value, such as L reg +λ rec ×L rec ,and then through this L reg +λ rec ×L rec Adjust the model parameters of the initial processing model to obtain the target processing model.

[0105] It should be noted that λ rec is an adjustable weight coefficient, and can take values such as 0.1, 0.2, etc. Specifically, it is not limited in the embodiments of this application.

[0106] In some alternative embodiments, to ensure that the expression of the first sample image is consistent with the expression of the rendered target sample image, then in the process of calculating the loss value for adjusting the model parameters of the initial processing model, the expression consistency can also be taken into consideration. Specifically, it can be understood with reference to the following method, that is:

[0107] Before updating the model parameters of the initial processing model based on the first loss value and the second loss value mentioned above, the encoded feature vector of the expression parameters of the target sample image can also be extracted, and the encoded feature vector of the expression parameters of the target sample image is mapped to obtain the expression basis coefficient of the expression parameters in the target sample image. It should be noted that how to map to obtain the expression basis coefficient of the expression parameters in the target sample image can be specifically understood with reference to the process of determining the expression basis coefficient of the expression parameters in the first sample image in step 402 above, and will not be elaborated here.

[0108] In this way, calculate the difference between the expression basis coefficient of the expression parameters and the expression basis coefficient of the target sample image to obtain the third loss value. As a schematic description, the second similarity distance between the expression basis coefficient of the expression parameters in the first sample image and the expression basis coefficient of the expression parameters in the target sample image can be calculated, and this second similarity distance can be used as the third loss value. The described second similarity distance can include but is not limited to Euclidean distance, cosine similarity distance, etc., and is not specifically limited.

[0109] Taking the Euclidean distance as an example, the second-order norm of the expression basis coefficient of the expression parameters in the first sample image and the expression basis coefficient of the expression parameters in the target sample image can be solved to obtain the corresponding Euclidean distance, and then the third loss value can be obtained. For example, the third loss value mentioned above satisfies Among them, represents the expression basis coefficient of the expression parameters in the i-th first sample image, Indicates the expression basis coefficients representing the expression parameters in the target sample image.

[0110] In this way, by taking the third loss value into account during model training, the model parameters of the initial processing model can be updated based on the first loss value, the second loss value, and the third loss value. As an exemplary description, specifically, the sum of the loss values between the first loss value, the second loss value, and the third loss value can be calculated, for example, L reg + λ rec × L rec + λ cons × L cons , and then through this L reg + λ rec × L rec + λ cons × L cons Adjust the model parameters of the initial processing model to obtain the target processing model. It should be noted that λ cons is an adjustable weight coefficient, for example, it can take values such as 0.1, 0.2, 0.3, etc., and is not specifically limited in the embodiments of this application.

[0111] Updating the initial processing model by incorporating the third loss value greatly improves the model's ability to capture accurate facial expressions.

[0112] The above Figure 4 During the process of training the model, the model is mainly trained by decoupling the expression parameters and the supplementary parameters. In some other alternative embodiments, the expression parameters, other facial parameters, and supplementary parameters can also be decoupled to train the model, so that the obtained target processing model can also be applicable to determining the target-driven information including both the expression basis coefficients and other facial coefficients.

[0113] Exemplarily, based on the system framework shown above Figure 3 and the method embodiments shown Figure 4 , Figure 6 Another schematic flowchart of the model training method provided by the embodiments of this application is shown. As Figure 6 shown, the model training method at least includes the following steps:

[0114] 601. Obtain a first sample image containing expression parameters and a second sample image containing supplementary parameters, where the supplementary parameters do not include expression parameters, and the expression parameters are used to indicate the expression situation in the first sample image.

[0115] In this example, the content of the first sample image, the second sample image, the expression parameters, the supplementary parameters, etc. described here can all be understood with reference to the content described in step 401 above Figure 4 , and will not be elaborated here.

[0116] 602. Extract the encoded feature vector of the expression parameters, and determine the expression basis coefficients of the expression parameters based on the encoded feature vector of the expression parameters.

[0117] In this example, the encoded feature vector of the expression parameters and the expression basis coefficients of the expression parameters described can also be understood with reference to the content described in step 402 above, and will not be elaborated here. Figure 4 In step 402 above, and will not be elaborated here.

[0118] 603. Obtain a third sample image containing facial parameters, where the facial parameters do not include expression parameters, and the supplementary parameters do not include facial parameters either.

[0119] In this example, in addition to obtaining the first sample image and the second sample image, a third sample image can also be obtained. It should be noted that in the third sample image, facial parameters need to be included, and these facial parameters do not include expression parameters. Additionally, in this embodiment, the supplementary parameters of the second sample image neither include expression parameters nor facial parameters.

[0120] For example, assuming that the parameters related to the face include four parameters: facial expression, facial shape, facial identification, and facial angle. At this time, the facial expression parameter can be used as the expression parameter in the first sample image. If the facial parameter in the third sample image is the facial identification, then the supplementary parameters of the second sample image at this time include the facial shape and the facial angle. Or, if the facial parameter in the third sample image is the facial angle, then the supplementary parameters of the second sample image at this time include the facial shape and the facial identification. Specifically, it is not limited in this application.

[0121] It should be noted that for steps 601 and 603 mentioned above, their execution order is not limited. For example, step 601 can be executed first, and then step 603; or step 603 can be executed first, and then step 601; or, steps 601 and 603 can be executed synchronously.

[0122] 604. Extract the encoded feature vector of the facial parameters, and determine the facial coefficients of the facial parameters based on the encoded feature vector of the facial parameters.

[0123] In this example, after obtaining the third sample image, the facial parameters in the third sample image can be encoded to extract the encoded feature vector of the facial parameters. For example, one or more convolutional layers in the encoding module can be used to perform convolutional processing on the facial parameters of the third sample image to encode and obtain the encoded feature vector of the facial parameters. Specifically, it can also be understood with reference to the process of extracting the encoded feature vector of the expression parameters in the first sample image in step 402 above, and will not be elaborated here. Figure 4 In step 402 above, and will not be elaborated here.

[0124] In this way, after extracting the encoded feature vector of the facial parameters, the facial coefficient of the facial parameters in the third sample image can be determined based on the encoded feature vector of the facial parameters. For example, the encoded feature vector of the facial parameters can be used as the input of a preset regression model in the initial processing model, so as to perform mapping processing on the encoded feature vector of the facial parameters through the preset regression model, and obtain the facial coefficient of the facial parameters in the third sample image.

[0125] For example, when the facial parameter is the facial shape, etc., the facial coefficient described at this time can sometimes also be called the face pinching basis coefficient, which can be used to reflect the weight of the facial shape in the image; or, it can also be understood as the weight of the facial shape in the obtained target image.

[0126] 605. Extract the encoded feature vector of the supplementary parameters.

[0127] In this example, after obtaining the second sample image, the supplementary parameters in the second sample image can be encoded to extract the encoded feature vector of the supplementary parameters. Specifically, reference can also be made to the process of extracting the encoded feature vector of the supplementary parameters of the second sample image in step 403 above for understanding, and details are not described here. Figure 4 It should be noted that for steps 602, 604, and 605 mentioned above, their execution order is not limited. For example, step 602 can be executed first, then step 604, and finally step 605; or, step 604 can be executed first, then step 602, and finally step 605; or, steps 602, 604, and 605 can be executed synchronously, etc. In practical applications, there are also other execution orders between steps 602, 604, and 605, and no detailed examples are given here.

[0128]

[0129] 606. Determine the reconstructed sample image based on the encoded feature vector of the supplementary parameters, the expression basis coefficient of the expression parameters, and the facial coefficient of the facial parameters.

[0130] In this example, after extracting the encoded feature vector of the supplementary parameters, the expression basis coefficient of the expression parameters, and the facial coefficient of the facial parameters, the reconstructed sample image can be determined based on the encoded feature vector of the supplementary parameters, the expression basis coefficient of the expression parameters, and the facial coefficient.

[0131] Figure 3 As a schematic description, feature fusion processing can be performed on the encoded feature vector of the supplementary parameters, the expression basis coefficient of the expression parameters, and the facial coefficient of the facial parameters, and then the fused feature vector can be decoded to obtain the reconstructed sample image. Specifically, reference can be made to the above Figure 3Understand it based on the schematic diagram of the system framework shown, which will not be elaborated here.

[0132] 607. Train the initial processing model based on the expression basis coefficients of the expression parameters, the expression labels of the first sample image, the target sample image, and the reconstructed sample image to obtain the target processing model. The target sample image includes expression parameters, supplementary parameters, and facial parameters.

[0133] In this example, before training the model, it is also necessary to render the target sample image based on the first sample image, the second sample image, and the third sample image. As a schematic description, the expression parameters of the first sample image, the supplementary parameters of the second sample image, and the facial parameters of the third sample image can be rendered to obtain a target sample image that includes both expression parameters, facial parameters, and supplementary parameters. Or rather, the target sample image and the first sample image include the same expression parameters, the target sample image and the third sample image include the same facial parameters, and the target sample image and the second sample image include the same supplementary parameters other than the expression parameters and facial parameters.

[0134] For example, Figure 7 shows another schematic diagram of the image rendering provided by the embodiments of the present application. As Figure 7 shown, this scenario at least includes the first sample image, the second sample image, and the third sample image. Taking the first sample image as the face image A, the second sample image as the face image B, and the third sample image as the face image C as an example, by rendering the expression parameters of the face image A, the facial identifiers of the face image C, and the supplementary parameters of the face image B, the target sample image, that is, the face image Q, is constructed and rendered.

[0135] In the above Figure 7 , the face image Q has exactly the same facial expression as the face image A, the face image Q has exactly the same facial identifiers as the face image C, and the face image Q has exactly the same supplementary parameters of other parts of the face except for the facial expression and facial identifiers as the face image B, such as the angle, light, focal length, or background of the face B.

[0136] In this way, after rendering the target sample image, the expression basis coefficients of the expression parameters, the expression labels of the first sample image, the target sample image, and the reconstructed sample image determined in step 606 can be used as training data. Furthermore, the expression basis coefficients of the expression parameters, the expression labels of the first sample image, the target sample image, and the reconstructed sample image are used as the input of the initial processing model to train the initial processing model, thereby training the target processing model.

[0137] How to train the target processing model here can refer to the above-mentioned Figure 4 The sum of the first loss value and the second loss value calculated in step 404, or the sum of the first loss value, the second loss value, and the third loss value is used to update the model parameters of the initial processing model. Specifically, it can refer to the content of step 403 in the above-mentioned Figure 4 for understanding, and the steps are not elaborated here.

[0138] The difference from the above-mentioned target processing model obtained through Figure 4 training is that: the target processing model obtained through Figure 6 training here can not only be used to determine the expression basis coefficients of the facial expressions in the first processed image, but also determine the facial coefficients of other facial attributes in the second processed image. The other facial attributes described do not include the facial expression. In this way, based on Figure 6 the expression basis coefficients of the facial expressions and the facial coefficients of the facial attributes determined by the shown target processing model, the face of the target image is adjusted, which not only enables the generated target image to include the facial expressions corresponding to the expression basis coefficients, but also includes the facial attributes corresponding to the facial coefficients. Through the above method, the applicable scenarios are greatly enriched, meeting the facial adjustment needs of different users for different target images.

[0139] For example, assuming that on the basis of considering the facial expression, the facial parameter of the facial identifier used to identify the relevant identity of person A is also considered. At this time, the adjusted target image can show the facial expression of person A corresponding to the facial identifier. Or, assuming that the facial angle (for example, the face is tilted up 30°) is also considered. At this time, the adjusted target image can show the facial expression with the face tilted up 30°.

[0140] It should be noted that the above Figure 6 only shows the example of incorporating the facial parameters of the third sample image for illustration. In practical applications, training images with other different facial parameters can also be obtained, such as the fourth training image with a facial angle, or the fifth training image with a facial shape, etc. At this time, one or more of the fourth training image and the fifth training image can also be incorporated into the model training to train a target processing model that meets different conditions, which is not specifically limited in the embodiments of this application.

[0141] Through the above Figure 4 method, a decoupled training framework between the expression parameters and the supplementary parameters is constructed (that is, the target processing model shown in Figure 4 ); or through the above Figure 6 method, a decoupled training framework between the expression parameters, the facial parameters, and the supplementary parameters is constructed (that is, Figure 6The shown target processing model). During the subsequent use of the target processing model, it is possible to focus on the facial expressions in the image, or also focus on other facial parameters, greatly improving the ability to capture facial expressions, or facial expressions and facial attributes, etc. In addition, through this method, the generalization, multi-frame stability, and accuracy of objects in different images are all strongly improved in performance. In this way, after constructing the target processing model, high-quality image-based facial expression capture can be achieved by means of the target processing model, and the captured high-quality expression basis coefficients can be migrated to the target image to achieve facial adjustment of the target image, thereby obtaining a target image containing the corresponding facial expression in the image; or, it is also possible to use the expression basis coefficients of the captured facial expressions and the facial coefficients of the facial attributes to complete the facial adjustment of the target image, thereby obtaining the facial expression and other facial attributes in the image. Through the above method, the expression performance effect of the target image is greatly improved, as well as the realism and expressiveness of the target image. In addition, compared with existing products in the industry, the target processing model trained by this application can, in the scenario of monocular facial image input, achieve device independence and improve the effect of expression capture.

[0142] The above Figure 4 and Figure 6 Mainly from the perspective of method embodiments, the method for model training provided by the embodiments of this application is described. After training the target processing model based on Figure 4 the described method, in the method for performing expression processing, the expression basis coefficients can be determined by means of the target processing model. Exemplarily, on the basis of the Figure 4 shown target processing model, Figure 8 a schematic flowchart of the method for expression processing provided by the embodiments of this application is shown. As Figure 8 shown, the method for expression processing provided by this application at least includes the following steps:

[0143] 801. Obtain a first processed image containing a facial expression.

[0144] In this example, the first processed image may include, but is not limited to, a human face image, an animal face image, or other facial images, etc., which are not specifically limited in this application. It should be noted that the first processed image includes a facial expression.

[0145] 802. Extract the encoded feature vector of the facial expression based on the target processing model, where the target processing model is a machine learning model trained by using the expression basis coefficients of the expression parameters in the first sample image, the expression labels of the first sample image, the target sample image, and the reconstructed sample image to train the initial processing model.

[0146] In this example, the target processing model described herein is a machine learning model obtained by training an initial processing model using the expression basis coefficients of the expression parameters in the first sample image, the expression labels of the first sample image, the target sample image, and the reconstructed sample image as training data, and using the expression basis coefficients of the facial expression in the first processed image as the training target. For the specific training process, reference can be made to the content described in the foregoing Figure 4 and understood accordingly, which will not be elaborated here.

[0147] It should be noted that for the content such as the target sample image and the reconstructed sample image mentioned here, specific reference can be made to the content shown in the foregoing Figure 4 and understood accordingly, which will not be elaborated here.

[0148] After obtaining the first processed image containing the facial expression in step 801, the first processed image can be used as the input of the target processing model to extract the encoded feature vector of the facial expression in the first processed image through the target processing model.

[0149] As a schematic description, specifically, through the encoding module in the target processing model, the facial expression in the first processed image is encoded to obtain the encoded feature vector of the facial expression in the first processed image. More specifically, through the convolutional layer in the encoding module of the target processing model, convolutional calculation processing is performed on the facial expression in the first processed image to extract the corresponding encoded feature vector of the facial expression.

[0150] 803. Perform a mapping process on the encoded feature vector of the facial expression based on the target processing model to obtain the expression basis coefficients of the facial expression, and the expression basis coefficients of the facial expression are used to indicate the weights of one or more expression bases in the facial expression.

[0151] In this example, after extracting the encoded feature vector of the facial expression, the encoded feature vector of the target facial expression is used as the input of the target processing model to perform a mapping process on the encoded feature vector of the facial expression through the target processing model to obtain the expression basis coefficients of the facial expression. Through the expression basis coefficients of the facial expression, the weights of one or more expression bases in the facial expression can be reflected.

[0152] For example, if the facial expression includes the expression base of "right eyebrow raise", and if the calculated expression basis coefficient of this expression base of "right eyebrow raise" is 0.03 at this time, then in the subsequent process of driving the virtual image, the right eyebrow of the virtual image needs to be driven according to this weight of 0.03.

[0153] 804. Based on the expression basis coefficients of the facial expression, perform facial adjustment on the target image to obtain a target image with the facial expression.

[0154] In this example, after obtaining the expression basis coefficients of the facial expression, the face of the target image is adjusted based on the expression basis coefficients of the facial expression, so as to obtain a target image including the facial expression.

[0155] As a schematic description, weighted processing can be performed based on the expression basis coefficients of the facial expression and the corresponding expression bases to determine the performance results of the corresponding expression bases, and then the face of the target image can be adjusted according to the performance results of all the expression bases.

[0156] For example, assume that the expression basis coefficients of the facial expression include 0.2, 0.15, and 0.3. Among them, 0.2 is used to indicate the weight of expression basis 1, 0.15 is used to represent the weight of expression basis 2, and 0.3 is used to represent the weight of expression basis 3. At this time, 0.2, 0.15, and 0.3 can be used to weight expression basis 1, expression basis 2, and expression basis 3 respectively, so as to obtain the performance results of expression basis 1, the performance results of expression basis 2, and the performance results of expression basis 3. In this way, the performance results of expression basis 1, the performance results of expression basis 2, and the performance results of expression basis 3 are superimposed to complete the facial adjustment process of the target image. That is to say, the expression bases involved in the facial expression are superimposed according to the expression basis coefficients for adjusting the face of the target image.

[0157] Optionally, taking the target image in the video scene as an example, adjusting the face of the target image in the video based on the expression basis coefficients of the facial expression can make the target image in each frame of the video include the facial expression, so as to obtain a target image including the facial expression for each frame. In this way, a driving video can also be generated based on the obtained target image including the facial expression.

[0158] Or, taking the target image in a single-image scene as an example, adjusting the face of the target image in the single image based on the expression basis coefficients of the facial expression can make the target image in the single image include the facial expression, so as to obtain a target image including the facial expression.

[0159] Through the above method, it is possible to focus on the facial expression in the image, achieve high-quality image-based facial expression capture, and transfer the captured high-quality expression basis coefficients to the target image to complete the facial adjustment of the target image, so that the obtained target image includes the facial expression, improving the expression effect, realism, and expressiveness of the target image.

[0160] Optionally, based on the Figure 6 previously trained target processing model, Figure 9 Another flowchart of the expression processing method provided by the embodiment of the present application is shown. As Figure 9As shown in the figure, the method for facial expression processing provided by this application at least includes the following steps:

[0161] 901. Obtain a first processed image containing a facial expression.

[0162] In this example, the first processed image may include, but is not limited to, a human face image, an animal face image, or other facial images, etc., which are not specifically limited in this application. It should be noted that the first processed image includes a facial expression.

[0163] 902. Obtain a second processed image containing facial attributes, where the facial attributes do not include facial expressions.

[0164] In this example, the second processed image may include, but is not limited to, a human face image, an animal face image, or other facial images, etc., which are not specifically limited in this application.

[0165] It should be noted that the difference between the second processed image and the first processed image is that: the first processed image includes a facial expression, and the second processed image includes other facial attributes, and these facial attributes do not include facial expressions. For example, the facial attributes may include, but are not limited to, facial identifiers, facial shapes, or facial angles, etc., which are not specifically limited.

[0166] 903. Extract the encoded feature vector of the facial expression based on the target processing model. The target processing model is a machine learning model obtained by training an initial processing model using the expression basis coefficients of the expression parameters in the first sample image, the expression labels of the first sample image, the target sample image, and the reconstructed sample image as training data.

[0167] In this example, the target processing model described here is a machine learning model obtained by training an initial processing model using the expression basis coefficients of the expression parameters in the first sample image, the expression labels of the first sample image, the target sample image, and the reconstructed sample image as training data, and using the expression basis coefficients of the facial expression in the first processed image and the facial coefficients of the facial attributes in the second processed image as training targets. The specific training process can be understood by referring to the content described in the foregoing Figure 6 and will not be elaborated here.

[0168] It should be noted that the content such as the target sample image and the reconstructed sample image mentioned here can be specifically understood by referring to the content shown in the foregoing Figure 6 and will not be elaborated here.

[0169] After obtaining the first processed image containing the facial expression in step 901, the first processed image can be used as the input of the target processing model to extract the encoded feature vector of the facial expression in the first processed image through the target processing model.

[0170] As a schematic description, specifically, the encoding module in the target processing model encodes the facial expression in the first processed image to obtain the encoded feature vector of the facial expression in the first processed image. More specifically, through the convolutional layer in the encoding module of the target processing model, convolutional calculation is performed on the facial expression in the first processed image to extract the encoded feature vector of the corresponding facial expression.

[0171] 904. Map the encoded feature vector of the facial expression based on the target processing model to obtain the expression basis coefficients of the facial expression, and the expression basis coefficients of the facial expression are used to indicate the weights of one or more expression bases in the facial expression.

[0172] In this example, after the encoded feature vector of the facial expression is extracted, the encoded feature vector of the target facial expression is used as the input of the target processing model to perform mapping processing on the encoded feature vector of the facial expression through the target processing model, so as to obtain the expression basis coefficients of the facial expression. Through the expression basis coefficients of the facial expression, the weights of one or more expression bases in the facial expression can be reflected.

[0173] 905. Use the second processed image as the input of the target processing model to obtain the facial coefficients of the facial attributes.

[0174] In this example, after the second processed image is obtained, the previously Figure 6 trained target processing model can also be used to process the second processed image to obtain the corresponding facial coefficients of the facial attributes.

[0175] As a schematic description, the encoded feature vector of the facial attributes of the second processed image can be extracted through the target processing model. For example, the encoding module can encode the facial attributes in the second processed image to obtain the encoded feature vector of the facial attributes in the second processed image. In this way, based on the target processing model, mapping processing is performed on the encoded feature vector of the facial attributes to obtain the facial coefficients of the facial attributes. It should be noted that through the facial coefficients of the facial attributes, the weights of the corresponding facial attributes can be known.

[0176] It should be noted that the execution order of steps 901 and 902 mentioned above is not limited. For example, step 901 can be executed first and then step 902; or, step 902 can be executed first and then step 901; or, steps 901 and 902 can be executed synchronously. Similarly, the execution order of steps 903 to 904 and step 905 mentioned above is not limited. For example, steps 903 to 904 can be executed first and then step 905; or, step 905 can be executed first and then steps 903 to 904; or, steps 903 to 904 and step 905 can be executed synchronously.

[0177] 906. Based on the expression basis coefficients of the facial expression and the facial coefficients of the facial attributes, perform facial adjustment on the target image to obtain a target image containing the facial expression and the facial attributes.

[0178] In this example, after obtaining the expression basis coefficients of the facial expression through step 904 and the facial coefficients of the facial attributes through step 905, perform facial adjustment on the target image based on the expression basis coefficients of the facial expression and the facial coefficients of the facial attributes to obtain a target image containing the facial expression and the facial attributes. That is to say, in the adjusted target image, it includes both the facial expression of the first processed image and the facial attributes in the second processed image.

[0179] For example, assume that the expression basis coefficients of the facial expression determined from the first processed image include 0.2, 0.15, and 0.3. Among them, 0.2 is used to indicate the weight of expression basis 1, 0.15 is used to represent the weight of expression basis 2, and 0.3 is used to represent the weight of expression basis 3. In addition, the facial coefficient of the facial attribute (such as the facial angle "lower the head 25° to the lower left") determined from the second processed image is 0.35.

[0180] At this time, 0.2, 0.15, and 0.3 can be used to weight expression basis 1, expression basis 2, and expression basis 3 respectively, and the facial coefficient 0.35 is used to weight the facial angle "lower the head 25° to the lower left", so as to superimpose the weighted expression basis 1, the weighted expression basis 2, the weighted expression basis 3, and the weighted facial angle "lower the head 25° to the lower left" to complete the facial adjustment process of the target image, thereby obtaining the target image. At this time, the obtained target image contains the facial expression superimposed by the weighted expression basis 1, the weighted expression basis 2, and the weighted expression basis 3, and also includes the facial angle of "lower the head 25° to the lower left".

[0181] Optionally, taking the target image in the video scene as an example of the target image, based on the expression basis coefficients of facial expressions and the facial coefficients of facial attributes, performing facial adjustment on the target image in the video can make the target image in each frame of the video include facial expressions and facial attributes, so as to obtain a target image of each frame including facial expressions and facial attributes. In this way, a driving video can also be generated based on the obtained target image including facial expressions and facial attributes.

[0182] Alternatively, taking the target image in a single-image scene as an example of the target image, based on the expression basis coefficients of facial expressions and the facial coefficients of facial attributes, performing facial adjustment on the target image in the single image can make the target image in the single image include facial expressions and facial attributes, so as to obtain a target image including facial expressions and facial attributes.

[0183] Through the above method, not only the facial expressions in the image are focused on, but also other facial attributes in the image are concerned, realizing high-quality capture of facial expressions and other facial attributes based on images, and the facial adjustment of the target image can be completed based on the captured high-quality expression basis coefficients and facial coefficients, so that the target image includes the corresponding facial expressions and other facial attributes in the image. This not only improves the expression effect, realism and expressiveness of the virtual image, but also can adaptively consider other facial attributes according to user needs, expanding the applicable range of the target image and greatly enriching the usage scenarios.

[0184] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of the method. It can be understood that in order to implement the above functions, the corresponding hardware structures and / or software modules for executing each function are included. Those skilled in the art should easily realize that, combining the modules and algorithm steps of each example described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0185] The embodiments of the present application can divide the device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0186] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0187] A detailed description of the facial expression processing device in the embodiments of the present application will be given below. Figure 10 It is a schematic diagram of an embodiment of the facial expression processing device provided in the embodiments of the present application. As Figure 10 shown, the facial expression processing device may include an acquisition module 1001 and a processing module 1002.

[0188] Among them, the acquisition module 1001 is used to acquire a first processed image containing a facial expression. Specifically, reference can be made to the content shown in step 801 in the foregoing Figure 8 for understanding; or, reference can be made to the content shown in step 901 in the foregoing Figure 9 for understanding, and details will not be elaborated here.

[0189] The processing module 1002 is used to extract an encoded feature vector of the facial expression based on a target processing model. The target processing model is a machine learning model obtained by training an initial processing model with the expression basis coefficients of the expression parameters in the first sample image, the expression labels of the first sample image, the target sample image, and the reconstructed sample image as training data. The target sample image includes expression parameters and supplementary parameters of the second sample image. The supplementary parameters do not include expression parameters. The expression parameters are used to indicate the expression situation in the first sample image. The reconstructed sample image is obtained based on the supplementary parameters and the expression basis coefficients of the expression parameters. Specifically, reference can be made to the content shown in step 802 in the foregoing Figure 8 for understanding; or, reference can be made to the content shown in step 903 in the foregoing Figure 9 for understanding, and details will not be elaborated here.

[0190] The processing module 1002 is used to perform a mapping process on the encoded feature vector of the facial expression to obtain the expression basis coefficients of the facial expression. The expression basis coefficients of the facial expression are used to indicate the weights of one or more expression bases in the facial expression. Specifically, reference can be made to the content shown in step 803 in the foregoing Figure 8 for understanding; or, reference can be made to the content shown in step 903 in the foregoing Figure 9 for understanding, and details will not be elaborated here.

[0191] The processing module 1002 is configured to perform facial adjustment on the target image based on the expression basis coefficients of the facial expressions to obtain a target image with facial expressions. Specifically, reference may be made to the content shown in step 804 described above Figure 8 for understanding; or, reference may be made to the content shown in step 904 described above Figure 9 for understanding, which will not be elaborated here.

[0192] In some alternative embodiments, the processing module 1002 is configured to: perform weighted processing on the expression basis coefficients of the facial expressions and the corresponding expression bases to determine the performance results of the corresponding expression bases; perform facial adjustment on the target image based on the performance results of all expression bases.

[0193] In some other alternative embodiments, the acquisition module 1001 is further configured to obtain a second processed image including facial attributes before performing facial adjustment on the target image based on the expression basis coefficients of the facial expressions to obtain a target image with the facial expressions, where the facial attributes do not include facial expressions. The processing module 1002 is configured to use the second processed image as the input of the target processing model to obtain the facial coefficients of the facial attributes. The processing module 1002 is configured to: perform facial adjustment on the target image based on the expression basis coefficients of the facial expressions and the facial coefficients of the facial attributes to obtain a target image with the facial expressions and the facial attributes. Specifically, reference may be made to the content shown in step 905 to step 906 described above Figure 9 for understanding, which will not be elaborated here.

[0194] In some other alternative embodiments, the processing module 1002 is configured to: based on the encoded feature vector of the facial attributes of the second processed image of the target processing model; perform mapping processing on the encoded feature vector of the facial attributes by the target processing model to obtain the facial coefficients of the facial attributes.

[0195] The expression processing device in the embodiments of the present application is described above from the perspective of modular functional entities. Next, the model training device in the embodiments of the present application will be described from the perspective of modular functional entities. Figure 11 It is a schematic diagram of an embodiment of the model training device provided in the embodiments of the present application. As Figure 11 shown, the model training device may include an acquisition unit 1101 and a processing unit 1102.

[0196] Among them, the acquisition unit 1101 is configured to obtain a first sample image including expression parameters and a second sample image including supplementary parameters, where the supplementary parameters do not include expression parameters, and the expression parameters are used to indicate the expression situation in the first sample image. Specifically, reference may be made to the content described in step 401 above Figure 4 for understanding; or, reference may be made to the content described in step 401 above Figure 6For the content shown in step 601, it is understood here and will not be elaborated.

[0197] The processing unit 1102 is configured to extract the encoded feature vector of the expression parameters, and determine the expression basis coefficients of the expression parameters based on the encoded feature vector of the expression parameters. The expression basis coefficients of the expression parameters are used to indicate the weights of one or more expression bases in the expression of the first sample image. Specifically, reference can be made to the content described in step 402 above; or, reference can be made to the content shown in step 602 above. It will not be elaborated here. Figure 4 For understanding, reference can be made to the content described in step 402 above; or, reference can be made to the content shown in step 602 above. It will not be elaborated here. Figure 6 For the content shown in step 601, it is understood here and will not be elaborated.

[0198] The processing unit 1102 is configured to extract the encoded feature vector of the supplementary parameters, and determine the reconstructed sample image based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters. Specifically, reference can be made to the content described in step 403 above; or, reference can be made to the content shown in steps 605 to 606 above. It will not be elaborated here. Figure 4 For understanding, reference can be made to the content described in step 403 above; or, reference can be made to the content shown in steps 605 to 606 above. It will not be elaborated here. Figure 6 For the content shown in steps 605 to 606, it is understood here and will not be elaborated.

[0199] The processing unit 1102 is configured to train the initial processing model based on the expression basis coefficients of the expression parameters, the expression label of the first sample image, the target sample image, and the reconstructed sample image to obtain a target processing model. The target processing model is used to process the first processed image to obtain the expression basis coefficients of the facial expression in the first processed image. The expression basis coefficients of the facial expression are used to perform facial adjustment on the target image to obtain a target image with the facial expression. The target sample image includes expression parameters and supplementary parameters. Specifically, reference can be made to the content described in step 404 above; or, reference can be made to the content shown in step 607 above. It will not be elaborated here. Figure 4 For understanding, reference can be made to the content described in step 404 above; or, reference can be made to the content shown in step 607 above. It will not be elaborated here. Figure 6 For the content shown in step 607, it is understood here and will not be elaborated.

[0200] In some optional embodiments, the obtaining unit 1101 is further configured to obtain a third sample image including facial parameters before determining the reconstructed sample image based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters. The facial parameters do not include expression parameters, and the supplementary parameters do not include facial parameters. The processing unit 1102 is configured to: extract the encoded feature vector of the facial parameters, and determine the facial coefficients of the facial parameters based on the encoded feature vector of the facial parameters; determine the reconstructed sample image based on the encoded feature vector of the supplementary parameters, the expression basis coefficients of the expression parameters, and the facial coefficients of the facial parameters. Specifically, reference can be made to the content shown in steps 603 to 604 above for understanding. Figure 6 For understanding, reference can be made to the content shown in steps 603 to 604 above.

[0201] In some other alternative embodiments, the target processing model is further configured to process the second processed image to obtain the facial coefficients of the facial attributes in the second processed image, and the facial coefficients of the facial attributes are further used to perform facial adjustment on the target image, and the adjusted target image further includes the facial attributes of the second processed image, where the facial attributes do not include facial expressions.

[0202] In some other alternative embodiments, the processing unit 1102 is configured to perform a mapping process on the encoded feature vector of the expression parameters based on a preset regression model to obtain the expression basis coefficients of the expression parameters.

[0203] In some other alternative embodiments, the processing unit 1102 is configured to: perform a fusion process on the expression basis coefficients of the expression parameters and the encoded feature vector of the supplementary parameters to obtain a target fusion feature; decode the target fusion feature to obtain a reconstructed sample image;

[0204] In some other alternative embodiments, the processing unit 1102 is configured to: calculate the difference between the expression basis coefficients of the expression parameters and the expression labels of the first sample image to obtain a first loss value; calculate the image difference between the target sample image and the reconstructed sample image to obtain a second loss value; update the model parameters of the initial processing model based on the first loss value and the second loss value to obtain a target processing model.

[0205] In some other alternative embodiments, the processing unit 1102 is configured to: calculate a first similarity distance between the expression basis coefficients of the expression parameters and the expression labels of the first sample image; determine the first loss value based on the first similarity distance.

[0206] In some other alternative embodiments, the obtaining unit 1101 is further configured to: before calculating the image difference between the target face image and the reconstructed sample image to obtain a second loss value, obtain the pixel value information of the target sample image and the pixel value information of the reconstructed sample image. The processing unit 1102 is configured to: extract first pixel information from the pixel value information of the target sample image, where the first pixel information is used to indicate the facial contour situation in the target sample image; calculate the pixel difference between the pixel value information of the reconstructed sample image and the pixel value information of the target sample image to obtain image difference information; determine the second loss value based on the first pixel information and the image difference information.

[0207] In some other alternative embodiments, the processing unit 1102 is further configured to: before updating the model parameters of the initial processing model based on the first loss value and the second loss value to obtain the target processing model, extract the encoded feature vector of the expression parameters of the target sample image, and perform mapping processing on the encoded feature vector of the expression parameters of the target sample image to obtain the expression basis coefficients of the expression parameters in the target sample image; calculate the difference between the expression basis coefficients of the expression parameters in the first sample image and the expression basis coefficients of the expression parameters in the target sample image to obtain a third loss value; and update the model parameters of the initial processing model based on the first loss value, the second loss value, and the third loss value to obtain the target processing model.

[0208] In some other alternative embodiments, the processing unit 1102 is configured to: calculate the second similarity distance between the expression basis coefficients of the expression parameters and the expression basis coefficients of the target sample image; and determine the third loss value based on the second similarity distance.

[0209] The above describes the expression processing device and the model training device in the embodiments of the present application from the perspective of modular functional entities. Next, the expression processing device in the embodiments of the present application will be described from the perspective of hardware processing. Figure 12 FIG. is a schematic structural diagram of the expression processing device provided in the embodiments of the present application. This expression processing device may vary greatly due to configuration or performance differences. For example, it may include, but is not limited to Figure 10 the expression processing device shown in Figure 11 the model training device shown in

[0210] As Figure 12 shown, this expression processing device 300 may vary greatly due to configuration or performance differences. It may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the expression processing device. Further, the central processor 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the expression processing device 300. Exemplarily, the central processor 322 is used to execute the computer-executable instructions stored in the storage media 330, so as to implement the model training method or the expression processing method provided in the above embodiments of the present application.

[0211] The expression processing device 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM, and so on.

[0212] Exemplarily, Figure 12 the central processing unit 322 in may cause the expression processing device to execute the methods in the corresponding method embodiments by invoking the computer-executable instructions stored in the memory 332. Figures 4 to 9 the corresponding method embodiments.

[0213] Specifically, Figure 10 the functions / implementation processes of the processing module 1002 in and Figure 11 the processing unit 1102 in can be implemented by the central processing unit 322 in Figure 12 invoking the computer-executable instructions stored in the memory 332. Figure 10 the acquisition module 1001 in and Figure 11 the acquisition unit 1101 in the functions / implementation processes of can be implemented by Figure 12 the input / output interface 358 in.

[0214] The steps performed by the expression processing device in the foregoing embodiments may be based on the Figure 12 structure of the expression processing device shown.

[0215] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.

[0216] An embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.

[0217] In the foregoing embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0218] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0219] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0220] In addition, each functional unit in various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0221] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0222] The above embodiments can be implemented in whole or in part through software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product.

[0223] A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, a process or function according to the embodiments of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as an SSD), etc.

[0224] The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for facial expression processing, characterized in that, it includes: Obtain a first processed image containing a facial expression; Extract an encoded feature vector of the facial expression based on a target processing model, where the target processing model is a machine learning model obtained by training an initial processing model with the expression basis coefficients of the expression parameters in a first sample image, the expression labels of the first sample image, a target sample image, and a reconstructed sample image. The target sample image includes the expression parameters and supplementary parameters of a second sample image, and the supplementary parameters do not include the expression parameters. The expression parameters are used to indicate the expression situation in the first sample image, and the reconstructed sample image is obtained based on the supplementary parameters and the expression basis coefficients of the expression parameters; Perform a mapping process on the encoded feature vector of the facial expression based on the target processing model to obtain the expression basis coefficients of the facial expression, and the expression basis coefficients of the facial expression are used to indicate the weights of one or more expression bases in the facial expression; Based on the expression basis coefficients of the facial expression, perform facial adjustment on a target image to obtain a target image containing the facial expression.

2. The method according to claim 1, characterized in that, Performing facial adjustment on a target image based on the expression basis coefficients of the facial expression includes: Perform a weighted process on the expression basis coefficients of the facial expression and the corresponding expression bases to determine the expression results corresponding to the expression bases; Perform facial adjustment on the target image based on the expression results of all the expression bases.

3. The method according to any one of claims 1 to 2, characterized in that, Before performing facial adjustment on a target image based on the expression basis coefficients of the facial expression to obtain a target image containing the facial expression, the method further includes: Obtain a second processed image containing facial attributes, where the facial attributes do not include the facial expression; Use the second processed image as the input of the target processing model to obtain the facial coefficients of the facial attributes; Performing facial adjustment on a target image based on the expression basis coefficients of the facial expression to obtain a target image containing the facial expression includes: Perform facial adjustment on the target image based on the expression basis coefficients of the facial expression and the facial coefficients of the facial attributes to obtain a target image containing the facial expression and the facial attributes.

4. The method according to claim 3, characterized in that, Using the second processed image as the input of the target processing model to obtain the facial coefficients of the facial attributes includes: Extract an encoded feature vector of the facial attributes of the second processed image based on the target processing model; Perform a mapping process on the encoded feature vector of the facial attributes based on the target processing model to obtain the facial coefficients of the facial attributes.

5. A method for model training, characterized in that, it includes: Obtain a first sample image containing expression parameters and a second sample image containing supplementary parameters, where the supplementary parameters do not include the expression parameters, and the expression parameters are used to indicate the expression situation in the first sample image; Extract the encoded feature vector of the expression parameters, and determine the expression basis coefficients of the expression parameters based on the encoded feature vector of the expression parameters, where the expression basis coefficients of the expression parameters are used to indicate the weights of one or more expression bases in the expression of the first sample image; Extract the encoded feature vector of the supplementary parameters, and determine the reconstructed sample image based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters; Train the initial processing model based on the expression basis coefficients of the expression parameters, the expression label of the first sample image, the target sample image, and the reconstructed sample image to obtain a target processing model, where the target processing model is used to process the first processed image to obtain the expression basis coefficients of the facial expression in the first processed image, and the expression basis coefficients of the facial expression are used to perform facial adjustment on the target image to obtain a target image including the facial expression, and the target sample image includes the expression parameters and the supplementary parameters.

6. The method according to claim 5, wherein, Before determining the reconstructed sample image based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters, the method further includes: Obtain a third sample image including facial parameters, where the facial parameters do not include the expression parameters, and the supplementary parameters also do not include the facial parameters; Extract the encoded feature vector of the facial parameters, and determine the facial coefficients of the facial parameters based on the encoded feature vector of the facial parameters; Determining the reconstructed sample image based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters includes: Determine the reconstructed sample image based on the encoded feature vector of the supplementary parameters, the expression basis coefficients of the expression parameters, and the facial coefficients of the facial parameters.

7. The method according to claim 6, wherein, The target processing model is further used to process the second processed image to obtain the facial coefficients of the facial attributes in the second processed image, and the facial coefficients of the facial attributes are also used to perform facial adjustment on the target image, and the adjusted target image also includes the facial attributes of the second processed image, and the facial attributes do not include the facial expression.

8. The method according to any one of claims 5 to 7, wherein, Determining the expression basis coefficients of the expression parameters based on the encoded feature vector of the expression parameters includes: Perform mapping processing on the encoded feature vector of the expression parameters based on a preset regression model in the initial processing model to obtain the expression basis coefficients of the expression parameters.

9. The method according to any one of claims 5 to 7, wherein, Determining the reconstructed sample image based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters includes: Perform fusion processing on the expression basis coefficients of the expression parameters and the encoded feature vector of the supplementary parameters to obtain a target fusion feature; Decode the target fusion feature to obtain the reconstructed sample image.

10. The method according to any one of claims 5 to 7, wherein, Training an initial processing model based on the expression basis coefficients of the expression parameters, the expression labels of the first sample image, the target sample image, and the reconstructed sample image to obtain a target processing model, including: Calculating the difference between the expression basis coefficients of the expression parameters and the expression labels of the first sample image to obtain a first loss value; Calculating the image difference between the target sample image and the reconstructed sample image to obtain a second loss value; Updating the model parameters of the initial processing model based on the first loss value and the second loss value to obtain a target processing model.

11. The method according to claim 10, wherein, Calculating the difference between the expression basis coefficients of the expression parameters and the expression labels of the first sample image to obtain a first loss value, including: Calculating a first similarity distance between the expression basis coefficients of the expression parameters and the expression labels of the first sample image; Determining a first loss value based on the first similarity distance.

12. The method according to claim 10, wherein, Before calculating the image difference between the target sample image and the reconstructed sample image to obtain a second loss value, the method further includes: Obtaining the pixel value information of the target sample image and the pixel value information of the reconstructed sample image; Calculating the image difference between the target sample image and the reconstructed sample image to obtain a second loss value, including: Extracting first pixel information from the pixel value information of the target sample image, where the first pixel information is used to indicate the facial contour situation in the target sample image; Calculating the pixel difference between the pixel value information of the reconstructed sample image and the pixel value information of the target sample image to obtain image difference information; Determining the second loss value based on the first pixel information and the image difference information.

13. The method according to claim 10, wherein, Before updating the model parameters of the initial processing model based on the first loss value and the second loss value to obtain a target processing model, the method further includes: Extracting the encoded feature vector of the expression parameters of the target sample image and performing a mapping process on the encoded feature vector of the expression parameters of the target sample image to obtain the expression basis coefficients of the expression parameters of the target sample image; Calculating the difference between the expression basis coefficients of the expression parameters in the first sample image and the expression basis coefficients of the expression parameters of the target sample image to obtain a third loss value; Updating the model parameters of the initial processing model based on the first loss value and the second loss value to obtain a target processing model, including: Updating the model parameters of the initial processing model based on the first loss value, the second loss value, and the third loss value to obtain a target processing model.

14. The method according to claim 13, wherein, Calculating the difference between the expression basis coefficients of the expression parameters in the first sample image and the expression basis coefficients of the expression parameters of the target sample image to obtain a third loss value, including: Calculate a second similarity distance between the expression basis coefficients of the expression parameters in the first sample image and the expression basis coefficients of the expression parameters in the target sample image; Determine a third loss value based on the second similarity distance.

15. An expression processing device, Characterized in that it includes: An acquisition module, configured to acquire a first processed image including a facial expression; The processing module is configured to extract an encoded feature vector of the facial expression based on a target processing model, where the target processing model is a machine learning model obtained by training an initial processing model using the expression basis coefficients of the expression parameters in the first sample image, the expression labels of the first sample image, the target sample image, and the reconstructed sample image. The target sample image includes the expression parameters and supplementary parameters of the second sample image, and the supplementary parameters do not include the expression parameters. The expression parameters are used to indicate the expression situation in the first sample image, and the reconstructed sample image is obtained based on the supplementary parameters and the expression basis coefficients of the expression parameters; The processing module is configured to perform a mapping process on the encoded feature vector of the facial expression to obtain the expression basis coefficients of the facial expression, and the expression basis coefficients of the facial expression are used to indicate the weights of one or more expression bases in the facial expression; The processing module is configured to perform a facial adjustment on the target image based on the expression basis coefficients of the facial expression to obtain a target image including the facial expression.

16. A model training device, Characterized in that it includes: An acquisition unit, configured to acquire a first sample image including expression parameters and a second sample image including supplementary parameters, where the supplementary parameters do not include the expression parameters, and the expression parameters are used to indicate the expression situation in the first sample image; A processing unit, configured to extract an encoded feature vector of the expression parameters and determine the expression basis coefficients of the expression parameters based on the encoded feature vector of the expression parameters. The expression basis coefficients of the expression parameters are used to indicate the weights of one or more expression bases in the expression of the first sample image; The processing unit is configured to extract an encoded feature vector of the supplementary parameters and determine a reconstructed sample image based on the encoded feature vector of the supplementary parameters and the expression basis coefficients of the expression parameters; The processing unit is configured to train an initial processing model based on the expression basis coefficients of the expression parameters, the expression labels of the first sample image, the target sample image, and the reconstructed sample image to obtain a target processing model. The target processing model is used to process a first processed image to obtain the expression basis coefficients of the facial expression in the first processed image. The expression basis coefficients of the facial expression are used to perform a facial adjustment on the target image to obtain a target image including the facial expression. The target sample image includes the expression parameters and the supplementary parameters.

17. An expression processing device, Characterized in that it includes: An input / output interface, a processor, and a memory, and program instructions are stored in the memory; The processor is configured to execute program instructions stored in the memory and perform the method according to any one of claims 1 to 14.

18. A computer-readable storage medium, characterized in that the computer-readable storage medium comprises instructions which, when run on a computer device, cause the computer device to perform the method according to any one of claims 1 to 14.

19. A computer program product, characterized in that the computer program product comprises instructions which, when run on a computer device, cause the computer device to perform the method according to any one of claims 1 to 14.