Method executed by electronic equipment and electronic equipment
By referencing high-quality facial images of the same user and utilizing personalized facial features for image restoration, the problem of image quality degradation in extreme scenarios of electronic devices is solved, achieving high-fidelity and natural texture image restoration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING SAMSUNG TELECOM R&D CENT
- Filing Date
- 2024-10-23
- Publication Date
- 2026-04-24
AI Technical Summary
When electronic devices take photos at high zoom levels, in low light conditions, or when objects are moving, the image quality degrades drastically. Existing technologies struggle to effectively restore image quality, particularly in terms of fidelity and texture naturalness.
By referencing high-quality facial images of the same user, personalized facial features are obtained, and these features are used for image restoration, avoiding artifact generation and ensuring texture fidelity.
It achieves high-fidelity image restoration in extreme scenarios, avoids artifacts, and the restored image is clearer with natural textures.
Smart Images

Figure CN121921207A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and more specifically, to a method executed by an electronic device, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the development of electronic technology, the imaging capabilities of electronic devices (such as mobile phones) are becoming increasingly powerful. However, due to the limitations of physical hardware size, the image quality of electronic devices degrades sharply in certain scenarios, such as high-magnification zoom photography, insufficient lighting, or moving objects, failing to meet users' requirements for photo quality.
[0003] Electronic device manufacturers have invested heavily in research and development to improve image quality, for example, by designing image quality enhancement algorithms based on artificial intelligence (AI). However, in some scenarios that cause severe image quality degradation, there is still significant room for improvement in image quality. Summary of the Invention
[0004] According to a first aspect of this disclosure, a method performed by an electronic device is provided, which may include: obtaining a second face image of a user based on a first face image of the user, wherein the second face image represents a face image that does not include personalized facial features of the user; obtaining personalized facial features of the user based on the first face image, the second face image, and a third face image of the user; and performing image restoration on the third face image based on the personalized facial features to obtain a fourth face image corresponding to the third face image.
[0005] Optionally, the second face image is generated based on a pre-established face component library for each face region in the first face image.
[0006] Optionally, obtaining the user's personalized facial features based on the first face image, the second face image, and the user's third face image includes: obtaining the user's first facial features based on the first face image, the second face image, and the third face image; and obtaining the user's personalized facial features based on the first facial features and the second facial features extracted from the third face image.
[0007] Optionally, the first facial feature includes at least one of the three-dimensional geometric structure features and texture features of the face; the second facial feature includes at least one of the brightness features, color features, expression features and pose features of the face.
[0008] Optionally, obtaining a second face image of the user based on a first face image includes: transforming the first face image into a first geometric structure feature corresponding to a set direction; determining the features of each part of the face corresponding to the first geometric structure feature based on a pre-established face component library; and obtaining the second face image based on the features of each part.
[0009] Optionally, for each part of the face corresponding to the first geometric structure feature, the features of each part are determined based on a pre-established face component library, including: for each part, determining n cluster centers of each part from the face component library, where n is a positive integer greater than or equal to 1; and obtaining the features of each part using a first attention network based on each part and the n cluster centers of each part.
[0010] Optionally, based on each part and n cluster centers of each part, the features of each part are obtained using a first attention network, including: for each part, using the features of the part as a first query vector of the first attention network; extracting a first key vector and a first value vector of the first attention network from each of the n cluster centers of the part; obtaining n intermediate features using the first attention network based on the first query vector and the first key vector and the first value vector corresponding to the n cluster centers respectively; and obtaining the features of the part by feature fusion of the n intermediate features.
[0011] Optionally, obtaining the user's first facial feature based on the first face image, the second face image, and the third face image includes: transforming the first face image into a first geometric structure feature corresponding to a set direction; obtaining a face image with texture features based on the first geometric structure feature and texture features in the first face image; obtaining the user's enhanced features based on the face image and the second face image; and obtaining the first facial feature based on the enhanced features and texture features and three-dimensional geometric structure features in the third face image.
[0012] Optionally, the set direction is the forward direction of the face.
[0013] Optionally, obtaining the first facial feature based on the enhanced features and the texture and three-dimensional geometric features in the third face image includes: transforming the three-dimensional geometric features in the enhanced features or the three-dimensional geometric features in the first face image into a second geometric feature corresponding to the pose in the third face image, based on the pose features in the third face image and the pose features in the first face image; performing texture mapping based on the texture features in the enhanced features and the second geometric features to obtain the mapped texture features; and obtaining the first facial feature based on the second geometric features, the mapped texture features, and the three-dimensional geometric features and texture features in the third face image.
[0014] Optionally, obtaining the first facial feature based on the second geometric structure feature, the mapped texture feature, and the three-dimensional geometric structure feature and texture feature in the third face image includes: obtaining a second query vector of the second attention network based on the three-dimensional geometric structure feature and texture feature in the third face image; obtaining a second key vector and a second value vector of the second attention network based on the second geometric structure feature and the mapped texture feature; and obtaining the first facial feature using the second attention network based on the second query vector, the second key vector, and the second value vector.
[0015] Optionally, obtaining the user's personalized facial features based on the first facial feature and the second facial feature extracted from the third facial image includes: determining a weight applied to the first facial feature; obtaining a weighted first facial feature based on the weight and the first facial feature; and obtaining the user's personalized facial features based on the weighted first facial feature and the second facial feature.
[0016] Optionally, determining the weights applied to the first facial feature includes: determining the weights based on at least one of geometric consistency information between the third face image and the first face image and weight control information input by the user.
[0017] Optionally, determining the weight based on at least one of the geometric structure consistency information between the third face image and the first face image and the weight control information input by the user includes: transforming the three-dimensional geometric structure features in the enhancement features or the three-dimensional geometric structure features in the first face image into a second geometric structure feature corresponding to the pose in the third face image based on the pose features in the third face image and the pose features in the first face image; and obtaining the geometric structure consistency information based on the second geometric structure feature and the three-dimensional geometric structure features in the third face image.
[0018] Optionally, obtaining the geometric structure consistency information based on the second geometric structure features and the three-dimensional geometric structure features in the third face image includes: performing normalization processing on the second geometric structure features and the three-dimensional geometric structure features in the third face image respectively; and obtaining the geometric structure consistency information using a fully connected network based on the normalized features.
[0019] Optionally, image restoration of the third face image based on the personalized facial features to obtain a fourth face image corresponding to the third face image includes: obtaining an artifact image corresponding to the third face image; performing degradation processing on the artifact image to obtain a degraded artifact image as a negative sample for a diffusion model used for image restoration; performing a diffusion process on the third face image using the diffusion model based on the negative sample, the first face image, and the personalized facial features to obtain a first reconstruction feature corresponding to the negative sample and a second reconstruction feature corresponding to the first face image; obtaining a reconstruction feature of the third face image based on the first reconstruction feature and the second reconstruction feature; and obtaining a fourth face image corresponding to the third face image based on the reconstruction feature.
[0020] Optionally, obtaining the artifact image corresponding to the third face image includes: determining the region in the third face image where artifacts exist based on the third face image and a pre-established standard face database; performing an artifact generation operation on the region to obtain the artifact image.
[0021] Optionally, based on the third face image and a pre-established standard face database, determining the regions in the third face image where artifacts exist includes: aligning the facial feature points in the third face image with the facial feature points in the standard face database; comparing the semantic information of the face regions in the third face image with the semantic information of the corresponding face regions in the standard face database; and identifying face regions with inconsistent semantic information as regions where artifacts exist.
[0022] Optionally, performing an artifact generation operation on the region includes: selecting at least one singular operator corresponding to the region from a pre-established library of singular operator operations; and performing an artifact generation operation on the region using the at least one singular operator.
[0023] Optionally, the artifact image is degraded by: downsampling the artifact image to obtain a downsampled artifact image; and upsampling the downsampled artifact image to obtain the degraded artifact image.
[0024] Optionally, obtaining a fourth face image corresponding to the third face image includes: obtaining reconstructed features of the third face image and decoding the reconstructed features to obtain decoded features; generating an adaptive convolution kernel corresponding to each pixel region based on texture features in different pixel regions of the personalized facial features; using the generated adaptive convolution kernel to perform texture correction on the decoded features to obtain texture-corrected decoded features; and obtaining the fourth face image based on the texture-corrected decoded features.
[0025] Optionally, obtaining the fourth face image based on the texture-corrected decoding features includes: obtaining the user's global facial features based on the reconstructed features; modulating the texture-corrected decoding features based on the global facial features to obtain modulated decoding features; and generating the fourth face image based on the modulated decoding features.
[0026] Optionally, the global facial features include a first global facial feature and a second global facial feature. Modulating the texture-corrected decoding features based on the global facial features to obtain modulated decoding features includes: multiplying the texture-corrected decoding features with the first global facial feature to obtain a first feature; and adding the first feature with the second global facial feature to obtain the modulated decoding features.
[0027] According to a fourth aspect of this disclosure, a method performed by an electronic device is provided, comprising: obtaining an artifact image corresponding to a third face image of a user; performing degradation processing on the artifact image to obtain negative samples associated with the third face image; obtaining positive samples associated with the third face image based on a first face image of the user; and performing image restoration on the third face image based on the negative samples and the positive samples to obtain a fourth face image corresponding to the third face image.
[0028] Optionally, obtaining the artifact image corresponding to the user's third face image includes: determining the region in the third face image where artifacts exist based on the third face image and a pre-established standard face database; performing an artifact generation operation on the region to obtain the artifact image.
[0029] Optionally, based on the third face image and a pre-established standard face database, determining the regions in the third face image where artifacts exist includes: aligning the facial feature points in the third face image with the facial feature points in the standard face database; comparing the semantic information of the face regions in the third face image with the semantic information of the corresponding face regions in the standard face database; and identifying face regions with inconsistent semantic information as regions where artifacts exist.
[0030] Optionally, performing an artifact generation operation on the region includes: selecting at least one singular operator corresponding to the region from a pre-established library of singular operator operations; and performing an artifact generation operation on the region using the at least one singular operator.
[0031] Optionally, the artifact image is degraded by: downsampling the artifact image to obtain a downsampled artifact image; and upsampling the downsampled artifact image to obtain the degraded artifact image.
[0032] Optionally, image restoration is performed on the third face image based on the negative samples and the positive samples to obtain a fourth face image corresponding to the third face image, including: performing a diffusion process on the third face image using a diffusion model based on the negative samples and the positive samples to obtain a first reconstruction feature corresponding to the negative samples and a second reconstruction feature corresponding to the positive samples; obtaining a reconstruction feature of the third face image based on the first reconstruction feature and the second reconstruction feature; and obtaining a fourth face image corresponding to the third face image based on the reconstruction feature.
[0033] According to a third aspect of this disclosure, a method performed by an electronic device includes: obtaining reconstructed features of a third face image based on a user's third face image; obtaining decoded features of the third face image based on the reconstructed features; performing texture correction on the decoded features based on the user's personalized facial features to obtain texture-corrected decoded features; and obtaining a fourth face image corresponding to the third face image based on the texture-corrected decoded features.
[0034] Optionally, the personalized facial features are obtained through the following operations: obtaining a second facial image based on the user's first facial image, wherein the second facial image represents a facial image that does not include the personalized facial features; and obtaining the personalized facial features based on the first facial image, the second facial image, and the third facial image.
[0035] Optionally, the second face image is generated based on a pre-established face component library for each face region in the first face image.
[0036] Optionally, texture correction is performed on the decoded features based on the user's personalized facial features to obtain texture-corrected decoded features, including: generating an adaptive convolution kernel corresponding to each pixel region based on the texture features in different pixel regions of the personalized facial features; and using the generated adaptive convolution kernel to perform texture correction on the decoded features to obtain texture-corrected decoded features.
[0037] Optionally, obtaining a fourth face image corresponding to the third face image based on the texture-corrected decoding features includes: obtaining the user's global facial features based on the reconstructed features; modulating the texture-corrected decoding features based on the global facial features to obtain modulated decoding features; and generating the fourth face image based on the modulated decoding features.
[0038] Optionally, the global facial features include a first global facial feature and a second global facial feature. Modulating the texture-corrected decoding features based on the global facial features to obtain modulated decoding features includes: multiplying the texture-corrected decoding features with the first global facial feature to obtain a first feature; and adding the first feature with the second global facial feature to obtain the modulated decoding features.
[0039] According to a fourth aspect of this disclosure, an electronic device is provided, which may include: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform a method of an exemplary embodiment of this disclosure.
[0040] According to a fifth aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program or instructions, which, when executed by at least one processor, cause the at least one processor to perform a method of an exemplary embodiment of this disclosure.
[0041] According to a sixth aspect of this disclosure, a computer program product including a computer program is provided, wherein when the computer program is executed by a processor, a method for implementing exemplary embodiments of this disclosure is provided.
[0042] According to embodiments of this disclosure, personalized facial features for face restoration are obtained by referencing high-quality facial images of the same user, and image restoration is performed using the personalized facial features to obtain clearer facial images.
[0043] According to embodiments of this disclosure, by performing image restoration based on negative samples used for image restoration, artifacts are avoided during the image restoration process, thereby obtaining clearer face images.
[0044] According to embodiments of this disclosure, texture correction is performed on the decoded features of the face image to be restored based on the user's personalized facial features, thereby ensuring the texture fidelity of the restored face image.
[0045] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0046] The accompanying drawings, which are incorporated in and form part of this specification, illustrate exemplary embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0047] Figure 1 This is a flowchart of a method performed by an electronic device according to an exemplary embodiment of the present disclosure.
[0048] Figure 2 This is a schematic diagram of generating a second face image according to an exemplary embodiment of the present disclosure.
[0049] Figure 3 This is a schematic diagram illustrating the generation of enhanced features according to exemplary embodiments of the present disclosure.
[0050] Figure 4 This is a schematic diagram of obtaining a first facial feature according to an exemplary embodiment of the present disclosure.
[0051] Figure 5 This is a schematic diagram illustrating how a user's facial features are obtained based on a third face image and a first face image, according to an exemplary embodiment of the present disclosure.
[0052] Figure 6 This is a schematic diagram illustrating the generation of negative samples according to exemplary embodiments of the present disclosure.
[0053] Figure 7 This is a schematic diagram of a face generator according to an exemplary embodiment of the present disclosure.
[0054] Figure 8 This is a schematic diagram illustrating the modulation of decoding features according to an exemplary embodiment of the present disclosure.
[0055] Figure 9 This is an overall architecture diagram of face image recovery according to exemplary embodiments of the present disclosure.
[0056] Figure 10 This is a schematic diagram of a face image recovery process according to an exemplary embodiment of the present disclosure.
[0057] Figure 11 This is a flowchart of a method performed by an electronic device according to another exemplary embodiment of the present disclosure;
[0058] Figure 12 This is a flowchart of a method performed by an electronic device according to yet another exemplary embodiment of the present disclosure;
[0059] Figure 13 This is a block diagram illustrating an electronic device according to exemplary embodiments of the present disclosure.
[0060] Figure 14 A schematic diagram of the structure of an electronic device to which embodiments of the present disclosure apply is shown. Detailed Implementation
[0061] The following description, with reference to the accompanying drawings, is provided to aid in a thorough understanding of the various embodiments of this disclosure as defined by the claims and their equivalents. This description includes various specific details to aid understanding but should be considered exemplary only. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of this disclosure. Furthermore, for clarity and brevity, descriptions of well-known functions and structures may be omitted.
[0062] The terms and wording used in the following description and claims are not limited to their dictionary meanings, but are merely used by the inventors to enable a clear and consistent understanding of this disclosure. Therefore, it will be apparent to those skilled in the art that the following description of various embodiments of this disclosure is for illustrative purposes only and not for limiting the purpose of this disclosure as defined in the appended claims and their equivalents.
[0063] It should be understood that the singular forms of “a,” “an,” and “the” can also include plural references unless the context clearly indicates otherwise. Thus, for example, the reference to “component surface” includes referring to one or more such surfaces. When we say that an element is “connected” or “coupled” to another element, the element can be directly connected or coupled to the other element, or it can mean that the connection between the element and the other element is established through an intermediate element. Furthermore, the use of “connected” or “coupled” herein can include wireless connections or wireless couplings.
[0064] The terms “comprising” or “may include” refer to the presence of a corresponding disclosed function, operation, or component that may be used in the various embodiments of this disclosure, rather than limiting the presence of one or more additional functions, operations, or features. Furthermore, the terms “comprising” or “having” may be interpreted as indicating certain characteristics, numbers, steps, operations, constituent elements, components, or combinations thereof, but should not be construed as excluding the possibility of the presence of one or more other characteristics, numbers, steps, operations, constituent elements, components, or combinations thereof.
[0065] The term "or" as used in the various embodiments of this disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not explicitly defined, the multiple items may refer to one, more, or all of the multiple items. For example, the description "parameter A includes A1, A2, A3" can be implemented as parameter A includes A1 or A2 or A3, or it can be implemented as parameter A includes at least two of the three items A1, A2, and A3.
[0066] Unless otherwise defined, all terms used in this disclosure (including technical or scientific terms) have the same meaning as understood by one of those skilled in the art to which this disclosure pertains. Common terms as defined in dictionaries are to be interpreted as having a meaning consistent with the context in the relevant technical field and should not be interpreted ideally or overly formally, unless expressly defined in this disclosure.
[0067] At least some of the functions of the device or electronic device provided in this disclosure embodiment can be implemented by an AI model, such as implementing at least one module of a plurality of modules of the device or electronic device by an AI model. AI-related functions can be executed by non-volatile memory, volatile memory, and a processor.
[0068] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as central processing unit (CPU), application processor (AP), etc., or pure graphics processing unit, such as graphics processing unit (GPU), vision processing unit (VPU), and / or AI-specific processors, such as neural processing unit (NPU).
[0069] The one or more processors control the processing of input data based on predefined operating rules or artificial intelligence (AI) models stored in non-volatile and volatile memory. These predefined operating rules or AI models are provided through training or learning.
[0070] Here, "providing through learning" refers to obtaining predefined operating rules or an AI model with desired characteristics by applying a learning algorithm to multiple learning datasets. This learning can be performed within the device or electronic device itself, in which the AI is executed according to the embodiment, and / or can be implemented via a separate server / system.
[0071] AI models can contain multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network computations by calculating the input data of that layer (such as the computation results of the previous layer and / or the input data of the AI model) and the multiple weight values of the current layer. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q-networks.
[0072] A learning algorithm is a method of training a predetermined target device (e.g., a robot) using multiple learning data sets to enable, allow, or control the target device to make determinations or predictions. Examples of such learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0073] The methods provided in this disclosure may relate to one or more fields in the technical fields of speech, language, image, video, or data intelligence.
[0074] Optionally, in the context of speech or language, in the method performed by an electronic device according to this disclosure, a speech signal as an analog signal may be received via a speech input device (e.g., a microphone), and the speech portion may be converted into computer-readable text using an Automatic Speech Recognition (ASR) model. The user's utterance intent can be obtained by interpreting the converted text using a Natural Language Understanding (NLU) model. The ASR model or NLU model may be an artificial intelligence model. The artificial intelligence model may be processed by a dedicated artificial intelligence processor designed in a hardware architecture specified for processing the artificial intelligence model. Language understanding is a technique for recognizing and applying / processing human language / text, including, for example, natural language processing, machine translation, dialogue systems, question answering, or speech recognition / synthesis.
[0075] Optionally, when dealing with the field of images or videos, in the method performed by an electronic device according to this disclosure, output data can be obtained by using image data as input data for an artificial intelligence model. The methods of this disclosure can relate to the field of visual understanding in artificial intelligence technology, which is a technology for recognizing and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.
[0076] Optionally, in the field of data intelligence processing, in the method performed by an electronic device according to this disclosure, during the reasoning or prediction phase, an artificial intelligence model can be used to perform prediction by using real-time input data. The processor of the electronic device can perform preprocessing operations on the data to transform it into a form suitable for use as input to the artificial intelligence model. Reasoning and prediction are techniques for making logical inferences and predictions by determining information, including, for example, knowledge-based reasoning, optimization prediction, preference-based planning, or recommendation.
[0077] In this application, the artificial intelligence model can be obtained through training. Here, "obtained through training" means obtaining a predefined operational rule or artificial intelligence model configured to perform desired features (or objectives) by training a basic artificial intelligence model with multiple training data using a training algorithm. The artificial intelligence model may include multiple neural network layers. Each of the multiple neural network layers includes multiple weight values, and neural network computation is performed by calculating the results of the previous layer and the multiple weight values.
[0078] Face restoration (also known as face redrawing or face hallucination) is an important branch of image quality enhancement technology. It is a technique specifically designed to improve the image quality of the human face area. This improvement includes, but is not limited to, the restoration and generation of details (e.g., interpolating finer textures to blurry areas and generating semantically reasonable and natural textures to areas without texture), reducing image noise, and eliminating blur to make the image clearer.
[0079] Currently, solutions for face restoration generally fall into two main categories: one is based on generative adversarial networks (GANs). However, due to the limited capacity of generator models (CNN-based or Transformer-based), they still fail to achieve satisfactory results in certain extreme scenarios. The other is based on AI-driven cloud models, which use multiple frames of images with different exposure times as input to the model, aiming to utilize supplementary information between frames. This method, however, still fails to achieve satisfactory results in some scenarios, such as long-distance zoom (10X-100X). This is because the sensor captures very little information due to the distance, and increasing the exposure time does not provide significant benefits.
[0080] Generally, for severely degraded face restoration, human visual assessment involves two basic aspects: (1) whether the face before and after restoration is the same person, i.e., the fidelity of the restoration; and (2) whether the restored texture is natural and free of artifacts. However, existing solutions have problems in both of these aspects.
[0081] For example, due to physical limitations of the hardware, in some situations (such as high-magnification zoom), the sensor can only capture relatively little effective image information and lacks identifiable texture structure information. Even so, in existing schemes, the degradation removal process may further eliminate some useful structural and texture information, leading to a further reduction in effective identity (ID) features, thereby reducing the fidelity of the restored face image.
[0082] For example, facial images restored using existing methods may exhibit feature artifacts. Feature artifacts refer to unrealistic, unreasonable, and unnatural textures generated by neural networks, such as hair on the face. The causes of feature artifacts include: (1) if degradation removal is not ideal, some inappropriate textures may remain. These inappropriate textures may introduce unexpected artifacts during the texture generation stage; (2) when using a diffusion model, the restored information may exhibit unnecessary artifacts due to randomness.
[0083] Diffusion-based techniques, as a general-purpose technique, are widely used in various visual tasks. However, when using diffusion techniques to improve extremely degraded face images, at least one of the following problems may occur: (1) the restored face image lacks fidelity; (2) artifacts may be generated. Extreme degradation mainly refers to the severe loss of effective information in face images, such as face images taken with high-magnification zoom, where information loss is caused by factors such as hardware limitations, object motion, or insufficient lighting.
[0084] Based on this, this disclosure proposes a high-fidelity scheme for recovering extremely degraded, low-quality face images. The technical solutions of this disclosure and their effects are explained below through descriptions of several optional embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other; identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0085] Figure 1 This is a flowchart of a method performed by an electronic device according to an exemplary embodiment of the present disclosure.
[0086] Reference Figure 1 In step S101, a second face image of the user is obtained based on the user's first face image. For example, the second face image represents a face image that does not include the user's personalized facial features. Personalized facial features can be understood as distinctive and identifiable facial features that can represent a user.
[0087] The first face image can be a clear, high-quality face image belonging to the same user as the face image to be recovered (which may be referred to as the third face image). For example, the third face image could be an extremely degraded face image taken using high-magnification zoom. The first face image can be a high-quality reference face image of the same user.
[0088] In this disclosure, the second face image may also be referred to as a neutral face image, a standard face image, or an average face. The second face image may be generated based on a pre-established face part library. The face part library may be a database including at least one cluster center for each part of the face. Each cluster center may be obtained from face images of multiple sample users using a clustering algorithm.
[0089] For example, for each face region in the first face image, at least one cluster center corresponding to the face region is selected from the face component library, and the features of the face region are obtained based on the at least one cluster center (e.g., by cascading or averaging), and then the second face image is obtained based on the features of each face region.
[0090] According to the embodiment, the first face image can be transformed into a first geometric structure feature corresponding to a set direction. For each part of the face corresponding to the first geometric structure feature, the features of each part are determined based on a pre-established face component library. Based on the features of each part, a second face image is obtained.
[0091] For example, based on the pose features in the first face image, the three-dimensional geometric structure features in the first face image can be transformed into first geometric structure features corresponding to the forward face direction. For each part of the forward face corresponding to the first geometric structure features, n cluster centers for each part are determined from a pre-established face component library, where n is a positive integer greater than or equal to 1. Based on each part and the n cluster centers of each part, a first attention network is used to obtain the features of each part. Based on the features of each part, a second face image is obtained.
[0092] For each part of the positive face corresponding to the first geometric structure feature, the feature of that part can be used as the first query vector of the first attention network. The first key vector and the first value vector of the first attention network are extracted from each of the n cluster centers of that part. Based on the first query vector and the first key vector and the first value vector corresponding to the n cluster centers respectively, the first attention network is used to obtain n intermediate features. By performing feature fusion on the n intermediate features, the feature of that part is obtained.
[0093] The above explanation uses the setting of the face direction as an example, but the setting direction of this disclosure is not limited to this.
[0094] Figure 2 This is a schematic diagram illustrating the generation of a second face image according to an exemplary embodiment of the present disclosure. The second face image can be obtained using a neutral face generator (also referred to as a standard face generator or average face generator) of the present disclosure. The neutral face generator may be formed by an attention network and includes a library of face parts, which is a pre-built library of standard face parts for each part of a face, such as the left and right eyes, nose, mouth, and skin.
[0095] For example, for a frontal face relative to the first face image, the nearest n cluster centers are found for each part of the frontal face. The cluster centers of each part are then weighted and fused to obtain the neutral face features for each part. The following description uses the eyes in a frontal face as an example.
[0096] First, based on the pose features in the first face image, the three-dimensional geometric structure features in the first face image are transformed into first geometric structure features corresponding to the forward face direction, thus obtaining each part of the forward face corresponding to the first geometric structure features. For example... Figure 2 As shown, for the eye region, n nearest cluster centers are found from the face component library. A first key vector and a first value vector can be extracted from each cluster center, and the eye region features in the first geometric structure features are used as the first query vector. For example, a first key vector K1 and a first value vector V1 can be extracted from cluster center 1, and the eye region features in the first geometric structure features are used as the first query vector Q. A softmax function operation is performed based on Q and K1, and the result is multiplied by V1 to obtain intermediate features for cluster center 1. For each cluster center, the above operation is performed to obtain n intermediate features. Then, the n intermediate features are fused based on similarity to obtain neutral face features for the eye region. The above operations are merely exemplary, and this disclosure is not limited thereto.
[0097] By correcting the first face image to a standard pose (such as a forward pose), the user's facial features can be extracted more effectively. Furthermore, considering the diversity of neutral facial features corresponding to different parts of the face, multiple similar neutral facial features for each part are identified from a facial component library. These similar neutral facial features are then fused based on similarity to obtain the corresponding neutral facial features for each part, thereby achieving accurate extraction of features from each part of the user's face.
[0098] In step S102, the user's personalized facial features are obtained based on the user's first face image, second face image, and third face image.
[0099] A user's facial features may include at least one of the following: brightness features, color features, expression features, pose features, three-dimensional geometric structure features, and texture features.
[0100] For example, facial brightness and color features are separate features, and expression features, pose features, three-dimensional geometric features, and texture features do not include color attributes.
[0101] Facial features are features used to represent facial expressions (such as smiling or frowning) and do not include spatial geometric properties.
[0102] Pose features are used to represent facial features in three-dimensional space. They are typically described using three angles: pitch, yaw, and roll.
[0103] Three-dimensional (3D) geometric features are features used to represent the geometric structure of the face, such as the facial contours, the height of the bridge of the nose, and the prominence of the cheekbones.
[0104] Texture features can include coarse texture features, fine texture features, and personalized features. Coarse texture features can include boundary texture features of facial components, such as the shape of the corners of the eyes and the lip line. Fine texture features can include fine-grained texture features, such as texture details like hair and pores. Personalized features can include unique markers such as moles, wrinkles, and scars.
[0105] These features can constitute a user's identity, making the user visually recognizable and distinguishable from other users.
[0106] As an example, based on a first face image, a second face image, and a third face image, a user's first facial feature is obtained. Based on the first facial feature and a second facial feature extracted from the third face image, the user's personalized facial features are obtained. For example, the first facial feature may include at least one of the three-dimensional geometric structure features and texture features of the face. The second facial feature may include at least one of the brightness features, color features, expression features, and pose features of the face.
[0107] As an example, a first face image, a second face image, and a third face image can be input into a neural network to obtain the user's first facial features. Alternatively, the pose features, 3D geometric features, and texture features of the faces contained in the second face image, the third face image, and the first face image can be input into the neural network to obtain the user's first facial features.
[0108] A neural network-based facial feature extractor can be used to perform the feature extraction process. For example, a facial feature extractor can be used to extract at least one of the following facial features from a third face image: brightness features, color features, expression features, pose features, 3D geometric features, and texture features. A facial feature extractor can also be used to extract at least one of the following facial features from a first face image: brightness features, color features, expression features, pose features, 3D geometric features, and texture features. The facial feature extractor used to extract facial features from the third face image can be the same as or different from the facial feature extractor used to extract facial features from the first face image. In different cases, the facial feature extractor used to extract facial features from the third face image and the facial feature extractor used to extract facial features from the first face image can share weights.
[0109] As an example, facial features can be extracted from a third face image and a first face image, respectively, and personalized facial features of the user can be obtained by processing the extracted facial features (such as fusion).
[0110] According to embodiments of this disclosure, the brightness features, color features, expression features, and pose features of a face can be kept consistent with a third face image, and the three-dimensional geometric features and texture features of the face can be enhanced by referring to a first face image. This ensures accurate identification of the user's identity during feature enhancement and avoids altering their appearance.
[0111] According to embodiments of this disclosure, a first face image can be transformed into a first geometric structure feature corresponding to a set direction. Based on the first geometric structure feature and the texture feature in the first face image, a face image with texture features is obtained. Based on the face image and a second face image, the user's enhanced features are obtained. Based on the enhanced features and the texture feature and three-dimensional geometric structure feature in a third face image, a first facial feature is obtained.
[0112] As an example, based on the pose features in the first face image, the three-dimensional geometric features in the first face image can be transformed into first geometric features corresponding to the forward face direction. Based on the first geometric features and the texture features in the first face image, a forward face image with texture features is obtained. Based on the forward face image and the second face image, the user's enhanced features are obtained. Based on the enhanced features and the texture features and three-dimensional geometric features in the third face image, the user's first facial features are obtained. Here, the enhanced features may include at least one of the 3D geometric features and texture features, or may include personalized features from the texture features.
[0113] Figure 3This is a schematic diagram illustrating the generation of enhanced features according to exemplary embodiments of the present disclosure. Feature enhancement can be performed by the neutral face generator and facial feature calibrator of the present disclosure. The facial feature calibrator can be formed by a neural network. For example, the facial feature calibrator may include a facial feature deformer and a facial texture mapping network.
[0114] Reference Figure 3 A facial feature deformer can transform the 3D geometric features of a first face image into first geometric features corresponding to the forward face direction based on the pose features in the first face image. A facial texture mapping network can obtain a forward face image with texture features based on the first geometric features and the texture features in the first face image. For example, the facial feature deformer can transform the 3D geometric features in the first face image to the forward face direction. The facial texture mapping network can attach the texture features from the first face image to the transformed 3D geometry, obtaining detailed 3D facial features with texture features along the forward face direction.
[0115] After obtaining the second face image from the neutral face generator, enhanced features for the user can be obtained based on the positive face image and the second face image. Here, the enhanced features may include at least one of 3D geometric features and texture features, or may include personalized features within the texture features. Figure 3 The network structure described is merely exemplary, and this disclosure is not limited thereto.
[0116] According to embodiments of this disclosure, based on pose features in a third face image and pose features in a first face image, the three-dimensional geometric structure features in the enhancement features or the three-dimensional geometric structure features in the first face image can be transformed into a second geometric structure feature corresponding to the pose in the third face image. Texture mapping is performed based on the texture features in the enhancement features and the second geometric structure features to obtain the mapped texture features. Based on the second geometric structure features, the mapped texture features, and the three-dimensional geometric structure features and texture features in the third face image, the user's first facial features are obtained.
[0117] For example, the first facial features can be obtained by inputting the second geometric structure features, the mapped texture features, and the three-dimensional geometric structure features and texture features in the third face image into a neural network.
[0118] For example, based on the three-dimensional geometric structure features and texture features in the third face image, the second query vector of the second attention network can be obtained. Based on the second geometric structure features and the mapped texture features, the second key vector and the second value vector of the second attention network can be obtained respectively. Based on the second query vector, the second key vector and the second value vector, the first facial features can be obtained using the second attention network.
[0119] In addition, in order to measure the extent to which the first face image is used when recovering the third face image (i.e. how much information from the first face image is needed to recover the third face image), the weights applied to the first facial features can be determined. Based on these weights and the first facial features, a weighted first facial feature is obtained. Based on the weighted first facial feature and the second facial feature, the user's personalized facial features are obtained.
[0120] As an example, the weights applied to the first facial feature can be determined based on at least one of geometric structural consistency information between the third face image and the first face image and weight control information input by the user.
[0121] Based on the pose features in the third face image and the pose features in the first face image, the three-dimensional geometric structure features in the enhancement features or the three-dimensional geometric structure features in the first face image can be transformed into a second geometric structure feature corresponding to the pose in the third face image. Based on the second geometric structure feature and the three-dimensional geometric structure feature in the third face image, geometric structure consistency information can be obtained.
[0122] For example, normalization processing can be performed on the second geometric structure features and the three-dimensional geometric structure features in the third face image, respectively. Based on the normalized features, a fully connected network can be used to obtain geometric structure consistency information.
[0123] Figure 4 This is a schematic diagram illustrating the acquisition of a first facial feature according to an exemplary embodiment of the present disclosure. The first facial feature can be obtained by a high-fidelity ID feature extractor of the present disclosure. The high-fidelity ID feature extractor can be implemented by a neural network.
[0124] Reference Figure 4 The facial feature deformer in the high-fidelity ID feature extractor can transform the 3D geometric features in the enhancement features or the 3D geometric features in the first face image into a second geometric feature corresponding to the pose in the third face image, based on the pose features in the third face image and the pose features in the first face image. For example, a transformation matrix can be calculated based on the third face image and the first face image, and the 3D geometric features in the first face image can be transformed into a second geometric feature corresponding to the pose in the third face image, so that the pose of the third face image matches that of the first face image.
[0125] The facial texture mapping network in the high-fidelity ID feature extractor can perform texture mapping based on texture features and second geometric structure features in the enhanced features to obtain the mapped texture features.
[0126] The mapped texture features and normalized second geometric structure features can be used for subsequent feature extraction. During subsequent feature extraction, a second attention network can be used to perform the extraction. Normalization ensures that the features are represented in a standardized metric space.
[0127] For example, based on the 3D geometric structure features and texture features in a third face image, a second query vector for the second attention network can be obtained. Based on the normalized second geometric structure features (or the second geometric structure features themselves) and the mapped texture features, a second key vector and a second value vector for the second attention network can be obtained, respectively. Based on the second query vector, the second key vector, and the second value vector, the second attention network can be used to obtain the first facial features. Figure 4 In this context, θ, φ, and ρ can represent preprocessing operations on the input features to obtain the feature matrix of the input features. T This represents a transformation of the feature matrix. It can be based on the second query vector Q and the second key vector K. T By performing a multiplication operation and then multiplying the result with the second value vector V, the first facial feature can be obtained.
[0128] Normalization can be performed on the second geometric structural features and the three-dimensional geometric structural features in the third face image, respectively. Based on the normalized features, a fully connected network is used to obtain geometric structural consistency information. Geometric structural consistency information can also be referred to as the confidence level for the third face image.
[0129] Furthermore, the weights applied to the first facial feature can be determined based on at least one of the geometric consistency information between the third and first facial images and the weight control information input by the user. For example, the value of the geometric consistency information can be multiplied by the weight value input by the user to obtain the final weight. The final weight is then applied to the first facial feature to obtain the final first facial feature. For example, when the third facial image is unclear and the user expects to retain fewer attributes from the third facial image, the user can input a higher weight value for the first facial feature. Figure 4 The network structure shown is merely exemplary, and this disclosure is not limited thereto.
[0130] Figure 5 This is a schematic diagram illustrating the acquisition of a user's personalized facial features based on a third face image and a first face image, according to exemplary embodiments of the present disclosure. The extraction of personalized facial features can be achieved by an efficient ID feature extractor. For example, the efficient ID feature extractor may be in the form of a neural network and may include a facial feature extractor, a feature enhancer, and a high-fidelity ID feature extractor.
[0131] Reference Figure 5The facial feature extractor can be used to extract at least one of the following features from the third face image and the first face image: brightness feature 1, color feature 2, expression feature 3, pose feature 4, three-dimensional geometric structure feature 5, and texture feature 6. Figure 5 The two facial feature extractors in the algorithm can be the same or different. If the two facial feature extractors are not used, they can share weights.
[0132] Brightness features, color features, expression features, and pose features extracted from third-party face images can be directly used as the user's second facial features.
[0133] Feature enhancers can be used to enhance the 3D geometric and texture features in a first face image. A feature enhancer may include a facial feature calibrator and a neutral face generator. For example, pose features, 3D geometric features, and texture features from the first face image can be input into the feature enhancer, which then outputs the pose features of the first face image along with the enhanced 3D geometric and texture features. The process of generating enhanced features can be referred to... Figure 3 The description will not be repeated here. Figure 5 The diagram illustrates a feature enhancer that can output pose features, 3D geometric features, and texture features, but this disclosure is not limited thereto. The feature enhancer can output 3D geometric features and texture features, and then the 3D geometric features and texture features output by the feature enhancer, along with pose features from a first face image, are input into a high-fidelity ID feature extractor.
[0134] The high-fidelity ID feature extractor can perform adaptive consistency calculation based on pose features, 3D geometric structure features, and texture features in a third face image, as well as the output of the feature enhancer. This process obtains geometric structure consistency information. Then, based on the 3D geometric structure features and texture features in the third face image, as well as the normalized second geometric structure features and mapped texture features obtained during the adaptive consistency calculation, the first facial features can be obtained.
[0135] Considering the varying degrees of relevance users place on the referenced first facial image, users can adjust the weights of the first facial features. The obtained first facial features can be weighted based on the user-input weights combined with geometric consistency information. Finally, personalized facial features for the user are obtained based on the weighted first facial features and second facial features from a third facial image.
[0136] In step S103, the third face image is restored based on the obtained personalized facial features to obtain a fourth face image corresponding to the third face image.
[0137] As an example, a diffusion model can be used to perform image restoration. A diffusion model can include an encoder for feature encoding, a neural network for performing the diffusion process, and a decoder for feature decoding. For instance, the encoder of the diffusion model can encode features of a third face image to obtain encoded features of the third face image. Then, based on the obtained personalized facial features of the user, a neural network (such as UNet) in the diffusion model is used to reconstruct features from the encoded features to obtain reconstructed features. The decoder of the diffusion model then decodes the reconstructed features to obtain a fourth face image corresponding to the third face image, i.e., the restored face image.
[0138] According to embodiments of this disclosure, to avoid generating artifacts during the diffusion process, an artifact image corresponding to the third face image can be obtained; the artifact image is degraded to obtain a degraded artifact image as a negative sample of the diffusion model; based on the negative sample, the first face image (as a positive sample of the diffusion model), and the user's personalized facial features, a neural network in the diffusion model is used to perform a diffusion process on the third face image (i.e., the encoded features of the third face image) to obtain a first reconstructed feature corresponding to the negative sample and a second reconstructed feature corresponding to the first face image. Based on the first and second reconstructed features, the reconstructed features of the third face image are obtained. Based on the reconstructed features, a fourth face image corresponding to the third face image is obtained.
[0139] When generating artifact images, the regions containing artifacts in the third face image can be determined based on the third face image and a pre-established standard face database. Then, artifact generation operations are performed on these regions to obtain the artifact image.
[0140] For example, facial feature points in a third-party face image can be aligned with facial feature points in a standard face database. After alignment, the semantic information of the face region in the third-party face image can be compared with the semantic information of the corresponding face region in the standard face database. Face regions with inconsistent semantic information are identified as regions containing artifacts. Then, an artifact generation operation is performed on this region to obtain an artifact image. For example, at least one singular operator corresponding to this region can be selected from a pre-established singular operator operation library, and the artifact generation operation is performed on this region using the selected singular operator.
[0141] During degradation processing, the artifact image can be downsampled to obtain a downsampled artifact image, and then the downsampled artifact image can be upsampled to obtain a degraded artifact image.
[0142] Figure 6This is a schematic diagram illustrating the generation of negative samples according to an exemplary embodiment of the present disclosure. Negative samples can be generated by a neural network-based negative sample generator. The negative sample generator can detect regions in a third face image where artifacts may occur, generate semantically relevant artifact images for these regions, degrade the generated artifact images, and use the degraded artifact images as negative samples for a diffusion model to reconstruct the third face image.
[0143] Reference Figure 6 Negative sample generators can include face generators, simulated zoom degraders, and facial feature extractors, all of which can be formed by neural networks.
[0144] The weird face generator can be used to generate images of strange (or deformed) faces. It aligns facial landmarks in a third-party face image with those in a standard face database. After alignment, the semantic information of the face regions in the third-party image is compared with the semantic information of the corresponding face regions in the standard face database to find potential semantic regions that may have characteristic artifacts. By performing artifact generation operations on these potential semantic regions, the weird face image is generated accordingly.
[0145] Figure 7 This is a schematic diagram of a face generator according to an exemplary embodiment of the present disclosure.
[0146] like Figure 7 As shown in the example of the weird face generator, the weird face generator may include a feature point detector, a standard face database, and a library of strange operators. The feature point detector may, for example, be formed by a neural network.
[0147] Feature point detectors can be used to capture the overall facial structure of third-party face images. Due to the robustness of feature points, semantic segmentation is not required. Feature point detectors can detect regions where facial features are incomplete or distorted due to blurring, as well as areas with inaccurate semantic information.
[0148] A standard face database can be pre-built using a series of standard faces and stored in the weird face generator. In the weird face generator, facial feature points of a third-party face image are compared with facial feature points in the standard face database to extract areas where facial features are incomplete or distorted due to blurring, as well as areas with inaccurate semantic information.
[0149] To address potential artifacts in different facial regions, a library of singular operators can be pre-built. For example, for the facial skin region, which is prone to generating non-skin texture features, such as hair on the face, a local random reproduction operator can be applied to this region. As another example, for regions primarily consisting of facial parts, which are prone to feature artifacts such as local deformations (e.g., distorted corners of the eyes or mouth), a local distortion operator can be applied to this region. Furthermore, the singular operator library may also include global operators for overall facial deformation, local displacement operators for boundary occlusion (e.g., a hat obscuring the face), local random copy operators for hair on the face, and local missing operators for missing objects on the face (e.g., glasses). The above examples are merely illustrative, and this disclosure is not limited thereto.
[0150] Based on the detected regions where characteristic artifacts may occur, at least one operator can be randomly selected from the singular operator library and applied to the region to perform artifact generation operations, thereby obtaining strange face images.
[0151] Return to reference Figure 6 Simulated zoom degradation can simulate various forms of degradation depending on different situations. For example, the simulated zoom degradation can dynamically downsample the generated deformed face image based on the zoom ratio (zoom magnification) of the image, and then upsample the downsampled image through a super-resolution reconstruction (SR) network and a Local Implicit Image Function (LIIF) to simulate zoom operation.
[0152] A facial feature extractor can extract features from degraded artifact images. The extracted features (negative features) can be used for inverse supervision of feature reconstruction, thereby avoiding potential feature artifacts.
[0153] like Figure 6 As shown, features extracted from the degraded artifact image can be used as negative samples. Alternatively, the degraded artifact image can be used as a negative sample, in which case the negative sample generator may not include a facial feature extractor. Figure 6 The structures shown are merely exemplary, and this disclosure is not limited thereto.
[0154] According to embodiments of this disclosure, considering the potential differences between the texture features reconstructed by the diffusion model and the actual texture features, this disclosure can dynamically adjust the convolution kernel of the adaptive convolution using extracted personalized facial features, allowing different convolution kernels to be applied to different facial regions. In other words, the adjusted convolution kernel can have different feature correction capabilities.
[0155] As an example, reconstructed features of a third-party face image can be obtained, and decoded features can be obtained. For instance, the features of the third-party face image can be reconstructed using a neural network in a diffusion model that performs the diffusion process, obtaining reconstructed features, which are then decoded by the decoder of the diffusion model to obtain decoded features.
[0156] Next, based on the texture features in different pixel regions of the user's personalized facial features, an adaptive convolution kernel corresponding to each pixel region can be generated. The generated adaptive convolution kernel is used to perform texture correction on the decoded features to obtain texture-corrected decoded features. Based on the texture-corrected decoded features, a fourth face image corresponding to the third face image is obtained.
[0157] In the diffusion model, the diffusion process takes place in a deep latent space and is primarily responsible for generating macroscopic semantics, such as indicating the positions of eyes and mouths. The decoding process is mainly responsible for generating fine textures, such as hair. If the decoding process is not properly guided, the generated fine textures may be inconsistent with the real texture. Therefore, to ensure the consistency and semantic integrity of the decoded features after texture correction, this disclosure can modulate the corrected decoded features by extracting global information from the reconstructed features rebuilt by the diffusion model.
[0158] As an example, the reconstructed features of the basis diffusion model obtain the user's global facial features. Based on these global facial features, the texture-corrected decoded features are modulated to obtain modulated decoded features. A fourth face image is then generated based on these modulated decoded features. For instance, the global facial features may include a first global facial feature and a second global facial feature. When applying the global facial features, the texture-corrected decoded features can be multiplied with the first global facial feature to obtain a first feature. Then, the first feature can be added with the second global facial feature to obtain the modulated decoded features.
[0159] Figure 8 This is a schematic diagram illustrating the modulation of decoded features according to an exemplary embodiment of the present disclosure. The modulation operation can be performed by a texture fidelity maintainer of the present disclosure. The texture fidelity maintainer can be implemented by a neural network and may include a texture corrector and a semantic consistency and feature integrity maintainer. Figure 8 The upper part can represent the operation of the texture corrector. Figure 8 The lower half of the expression can represent the operation of the semantic consistency and feature integrity maintainer.
[0160] By extracting texture features from facial features, a correction kernel (i.e., a convolution kernel) can be generated for different pixel regions. This kernel is then applied to a learnable convolution, generating a set of adaptive convolutions for different pixel regions. These adaptive convolutions adapt to the feature patterns of different pixel regions. For example, the texture features of wrinkled regions are different from those of smooth regions, and the corresponding adaptive convolution kernels are also different. Figure 8 As shown, the reconstructed features can first be reshaped and resized to obtain the reshaped and resized features f. Using each pixel region (such as pixel region K) in feature f, a corresponding convolution kernel (such as w) can be generated. Applying the corresponding convolution kernel to the same pixel region in the decoded feature V can obtain the texture-corrected features of that pixel region.
[0161] The generated adaptive convolutional kernels can be applied to different regions of the reconstructed features output by the decoder to achieve texture correction in different regions. For example, texture-corrected decoded features V' can be obtained by applying a corresponding convolutional kernel to each pixel region.
[0162] The first global feature vector (i.e., the first global facial feature) γ and the second global feature vector (the second global facial feature) β can be extracted from the reconstructed features, respectively. For example... Figure 8 As shown, the reconstructed features can first be reshaped and resized to obtain the reshaped and resized features. Then, a convolution operation can be performed on the features to obtain β, and then a convolution operation can be performed on β to obtain γ.
[0163] γ can be applied as a product to the corrected decoded features (or the features obtained after normalizing the corrected decoded features) to obtain the first feature, and β can be applied as an additive expression to the first feature. Global feature modulation ensures the integrity and semantic consistency of the output features, further avoiding the generation of artifacts. Figure 8 The examples shown are merely illustrative, and this disclosure is not limited thereto.
[0164] According to embodiments of this disclosure, based on the diffusion model, this disclosure utilizes a high-quality reference face image and introduces two components—an efficient ID feature extractor and a texture fidelity preserver—to address the issue of insufficient fidelity. These two components can be specifically designed to ensure fidelity in terms of identity and texture, respectively. Furthermore, an adaptive negative sample generator is introduced to guide the diffusion process and avoid generating artifacts.
[0165] Figure 9 This is an overall architecture diagram of face image recovery according to exemplary embodiments of the present disclosure.
[0166] Reference Figure 9 Highly discriminative ID features (such as first facial features) can be extracted using an efficient ID feature extractor. For example, an efficient ID feature extractor can perform the above-mentioned steps. Figure 5 The described operation. Features extracted by the efficient ID feature extractor can be applied to the diffusion process in the latent space, and the extracted texture features can dynamically correct intermediate diffusion results. For example, a texture fidelity preserver can perform the above-mentioned... Figure 8 The described operations. Furthermore, the constructed negative samples can also influence the diffusion process in the latent space, guiding the diffusion model to avoid generating artifacts. For example, the negative sample generator can perform the operations described above. Figure 6 The operation described is then applied to the diffusion process of the diffusion model.
[0167] Figure 10 This is a schematic diagram of a face image recovery process according to an exemplary embodiment of the present disclosure.
[0168] Reference Figure 10 Based on a third face image and a first face image, the efficient ID feature extractor disclosed herein can be used to obtain a user's personalized facial features. The efficient ID feature extractor can be described as described above. Figure 5 The description of the operation to be performed.
[0169] Negative samples for the diffusion model can be obtained using the negative sample generator disclosed herein, based on a third-party face image. The negative sample generator can be described in accordance with the above reference. Figure 6 The description of the operation is as follows. Furthermore, the first face image can be used as a positive sample for the diffusion model.
[0170] When running the diffusion model, the third-person face image can first be downsampled. Encoded features of the third-person face image are obtained through a low-quality adapter and encoder. Based on the encoded features, negative samples, and personalized facial features from an efficient ID feature extractor, a diffusion process is performed using a neural network in the diffusion model (such as a neural network formed by n UNets) to obtain reconstructed features based on positive samples and reconstructed features based on negative samples. A subtraction function (f) is then applied to the reconstructed features based on positive and negative samples to obtain the reconstructed features of the third-person face image.
[0171] The texture fidelity preserver can extract global facial features from the reconstructed features of the third face image. Then, based on the facial features and global facial features output by the efficient ID feature extractor, it performs feature correction or modulation on the decoded features obtained by decoding the reconstructed features through the decoder of the diffusion model to obtain the modulated decoded features. Finally, it generates a fourth face image corresponding to the third face image based on the modulated decoded features.
[0172] According to embodiments of this disclosure, more detailed features for face restoration are obtained by referencing high-quality face images of the same user. Furthermore, the fidelity of the user's identity in the restored image is ensured by extracting efficient ID features and applying them to feature reconstruction using a diffusion model. Texture correction of the features reconstructed by the diffusion model using extracted texture features ensures the fidelity of the texture in the restored image. Modulation of the texture-corrected features using global information extracted from the reconstructed features guarantees the integrity and semantic consistency of the output features, further avoiding artifact generation. Constructed negative samples are introduced during the diffusion process of the diffusion model to guide the diffusion process and prevent artifact generation.
[0173] Figure 11 This is a flowchart of a method performed by an electronic device according to another exemplary embodiment of the present disclosure.
[0174] Reference Figure 11 In step S111, an artifact image corresponding to the user's third face image is obtained. The third face image can be the face image to be recovered.
[0175] As an example, based on a third-party face image and a pre-established standard face database, the region containing artifacts in the third-party face image can be determined, and then an artifact generation operation can be performed on that region to obtain an artifact image.
[0176] For example, facial feature points in a third-party face image can be aligned with facial feature points in a standard face database. Then, the semantic information of the face regions in the third-party face image is compared with the semantic information of the corresponding face regions in the standard face database. Face regions with inconsistent semantic information are identified as regions containing artifacts.
[0177] When performing artifact generation on a region containing artifacts, at least one singular operator corresponding to that region can be selected from a pre-established library of singular operator operations. The selected singular operator is then used to perform the artifact generation operation on that region. For example, it can be performed as described above. Figure 7 The described operations are used to obtain artifact images.
[0178] In step S112, the artifact image is degraded to obtain a negative sample related to the third face image.
[0179] As an example, the artifact image can be downsampled to obtain a downsampled artifact image, and then upsampled to obtain a degraded artifact image. The degraded artifact image can then be used as a negative sample. For example, it can be done according to the above reference... Figure 6 The described operation yields negative samples.
[0180] In step S113, based on the user's first face image, positive samples related to the third face image are obtained. The first face image can be a clear, high-quality face image belonging to the same user as the third face image. For example, the first face image can be used as a positive sample.
[0181] In step S114, image restoration is performed on the third face image based on negative and positive samples to obtain a fourth face image corresponding to the third face image.
[0182] As an example, a diffusion model can be used to perform a diffusion process on a third face image based on negative and positive samples to obtain a first reconstruction feature corresponding to the negative sample and a second reconstruction feature corresponding to the positive sample. Based on the first and second reconstruction features, the reconstruction features of the third face image are obtained, and a fourth face image corresponding to the third face image is obtained based on the reconstruction features.
[0183] For example, refer to Figure 10 The diffusion model can perform a diffusion process based on positive and negative samples through a neural network to obtain the first and second reconstruction features. The reconstruction features of the third face image are obtained by processing the first and second reconstruction features (such as cascading). The fourth face image corresponding to the third face image is obtained by the decoder of the diffusion model based on the reconstruction features.
[0184] According to embodiments of this disclosure, by performing image restoration based on negative samples used for image restoration, artifacts are avoided during the image restoration process, thereby obtaining clearer face images.
[0185] Figure 12 This is a flowchart of a method performed by an electronic device according to yet another exemplary embodiment of the present disclosure.
[0186] Reference Figure 12 In step S121, based on the user's third face image, the reconstructed features of the third face image are obtained. The third face image can be the face image to be restored.
[0187] In step S122, the decoding features of the third face image are obtained based on the reconstructed features of the third face image.
[0188] For example, refer to Figure 10 The diffusion model can be used to perform a diffusion process on a third face image to obtain the reconstructed features of the third face image. Then, the decoder of the diffusion model can be used to decode the reconstructed features to obtain the decoded features.
[0189] In step S123, texture correction is performed on the decoded features based on the user's personalized facial features to obtain texture-corrected decoded features.
[0190] Personalized facial features can be obtained through the following operations: obtaining a second facial image based on a user's first facial image, wherein the second facial image represents a facial image that does not include personalized facial features; and obtaining personalized facial features based on the first facial image, the second facial image, and the third facial image.
[0191] The second face image can be generated based on a pre-established library of face components for each face region in the first face image.
[0192] As an example, based on texture features in different pixel regions of personalized facial features, an adaptive convolutional kernel corresponding to each pixel region is generated. This generated adaptive convolutional kernel is then used to perform texture correction on the decoded features, resulting in texture-corrected decoded features. For example, the above reference can be followed... Figure 8 The described operations are used to obtain the decoded features after texture correction.
[0193] In step S124, a fourth face image corresponding to the third face image is obtained based on the decoded features after texture correction.
[0194] As an example, the user's global facial features can be obtained based on the reconstructed features of the third face image. The texture-corrected decoding features are then modulated based on the global facial features to obtain the modulated decoding features. A fourth face image is then generated based on the modulated decoding features.
[0195] Optionally, the global facial features may include a first global facial feature and a second global facial feature. In this case, the texture-corrected decoded features can be multiplied with the first global facial feature to obtain a first feature, and the first feature can be added with the second global facial feature to obtain a modulated decoded feature. The modulated decoded feature is then used to generate the fourth image. For example, it can be done according to the above reference. Figure 8 The described operations are used to obtain the modulated decoded features.
[0196] According to embodiments of this disclosure, texture correction is performed on the decoded features of the face image to be restored based on the user's personalized facial features, thereby ensuring the texture fidelity of the restored face image.
[0197] The method disclosed herein can be applied to scenarios such as zoom shooting and editing images in a photo album.
[0198] For example, when a user is far from a target object and uses zoom to photograph it, the face in the captured image may be blurry. In this case, the electronic device can recommend several clear face images that most closely resemble the target object's identity based on the captured image. The user can select one of the recommended clear face images that matches the target object's identity as a reference face image. The user can adjust the selected reference face image (such as adjusting the reference ratio of the reference face image) or use the default settings. Then, the method of this disclosure (e.g., referring to...) can be used. Figure 10 The described method performs face restoration on captured images to obtain a clear image of the target object.
[0199] For example, a user can select a face image to be restored from the electronic device's photo album, such as a low-quality face image containing multiple faces. In this case, the electronic device can first automatically detect all faces in the selected image, and then the user can select the target face to be restored. Next, the electronic device can recommend several clear face images that are closest to the identity of the target face to the user. The user can select a face image from the recommended clear face images that matches the identity of the target face as a reference face image. The user can adjust the selected reference face image (such as adjusting the reference ratio of the reference face image) or use the default settings. Then, the method of this disclosure (e.g., referring to...) can be used. Figure 10 The described method restores the desired face image to obtain an image with improved target face quality. After restoring the target face in the desired face image, the current operation can be terminated when the restored image meets the user's requirements. If the user wishes to continue restoring other faces in the image, another target face can be selected from previously detected faces, and image restoration can be performed using the above method.
[0200] The above example scenarios are merely illustrative, and this disclosure is not limited thereto.
[0201] The methods performed by an electronic device according to exemplary embodiments of the present disclosure have been described above.
[0202] The electronic device according to embodiments of the present disclosure will now be briefly described. Figure 13 This is a block diagram illustrating an electronic device according to exemplary embodiments of the present disclosure. (Refer to...) Figure 13 The electronic device 1100 may include a memory 1101 and a processor 1102, wherein the processor 1102 is coupled to the memory 1101 and configured to perform any of the methods described above.
[0203] This disclosure also provides an electronic device including at least one processor, and optionally, at least one transceiver coupled to the at least one processor and / or at least one memory, wherein the at least one processor is configured to perform the steps of the method provided in any optional embodiment of this disclosure.
[0204] Figure 14 The diagram shows a schematic representation of an electronic device to which embodiments of this disclosure apply. For example... Figure 14 As shown, Figure 14 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, each of the processor 4001, memory 4003, and transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this disclosure. Optionally, the electronic device may be a first network node, a second network node, or a third network node.
[0205] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with this disclosure. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0206] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 14The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0207] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.
[0208] The memory 4003 is used to store computer programs or executable instructions that execute the embodiments of this disclosure, and is controlled by the processor 4001 to execute them. The processor 4001 is used to execute the computer programs or executable instructions stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0209] This disclosure provides a computer-readable storage medium storing a computer program or instructions that, when executed by at least one processor, can perform or implement the steps and corresponding content of the aforementioned method embodiments.
[0210] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0211] The terms “first,” “second,” “third,” “fourth,” “1,” “2,” etc. (if present) in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in a sequence other than that shown in the figures or text.
[0212] It should be understood that although arrows indicate various operation steps in the flowcharts of the embodiments of this disclosure, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of the embodiments of this disclosure, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured as required, and the embodiments of this disclosure do not limit this.
[0213] The above text and accompanying drawings are provided as examples only to help the reader understand this disclosure. They are not intended and should not be construed as limiting the scope of this disclosure in any way. Although certain embodiments and examples have been provided, it will be apparent to those skilled in the art, based on the content disclosed herein, that changes can be made to the illustrated embodiments and examples, and other similar implementations based on the technical concept of this disclosure can be adopted without departing from the scope of this disclosure, and these modifications and modifications are also within the protection scope of the embodiments of this disclosure.
Claims
1. A method performed by an electronic device, comprising: Based on the user's first facial image, a second facial image of the user is obtained, wherein the second facial image represents a facial image that does not include the user's personalized facial features; Based on the first face image, the second face image, and the user's third face image, the user's personalized facial features are obtained; Based on the personalized facial features, the third face image is restored to obtain a fourth face image corresponding to the third face image.
2. The method according to claim 1, wherein, The second face image is generated based on a pre-established face component library for each face region in the first face image.
3. The method according to claim 1, wherein, Based on the first face image, the second face image, and the user's third face image, the user's personalized facial features are obtained, including: Based on the first face image, the second face image, and the third face image, the user's first facial features are obtained; The user's personalized facial features are obtained based on the first facial features and the second facial features extracted from the third facial image.
4. The method according to claim 3, wherein, The first facial feature includes at least one of the three-dimensional geometric features and texture features of the face; The second facial feature includes at least one of the following: brightness feature, color feature, expression feature, and posture feature of the face.
5. The method according to claim 1, wherein, Based on the user's first facial image, obtain the user's second facial image, including: The first face image is transformed into a first geometric structure feature corresponding to a set direction; For each part of the face corresponding to the first geometric structure feature, the features of each part are determined based on a pre-established face component library; The second face image is obtained based on the features of each part.
6. The method according to claim 5, wherein, For each part of the face corresponding to the first geometric structure feature, features of each part are determined based on a pre-established face component library, including: For each part, n cluster centers for each part are determined from the face component library, where n is a positive integer greater than or equal to 1; Based on each part and n cluster centers of each part, features of each part are obtained using a first attention network.
7. The method according to claim 6, wherein, Based on each part and n cluster centers of each part, features of each part are obtained using a first attention network, including: For each of the parts, the features of the part are used as the first query vector of the first attention network; Extract the first key vector and the first value vector of the first attention network from each of the n cluster centers in the part; Based on the first query vector and the first key vector and first value vector corresponding to the n cluster centers respectively, the first attention network is used to obtain n intermediate features; The features of the part are obtained by fusing the n intermediate features.
8. The method according to claim 3, wherein, Based on the first face image, the second face image, and the third face image, the user's first facial features are obtained, including: The first face image is transformed into a first geometric structure feature corresponding to a set direction; Based on the first geometric structure features and the texture features in the first face image, a face image with texture features is obtained; Based on the first face image and the second face image, the enhanced features of the user are obtained; The first facial feature is obtained based on the enhanced features and the texture and three-dimensional geometric features in the third face image.
9. The method according to claim 5 or 8, wherein, The set direction is the frontal direction of the face.
10. The method according to claim 8, wherein, Based on the enhanced features and the texture and three-dimensional geometric features in the third face image, the first facial features are obtained, including: Based on the pose features in the third face image and the pose features in the first face image, the three-dimensional geometric structure features in the enhancement features or the three-dimensional geometric structure features in the first face image are transformed into a second geometric structure feature corresponding to the pose in the third face image. Based on the texture features in the enhanced features and the second geometric structure features, perform texture mapping to obtain the mapped texture features; The first facial feature is obtained based on the second geometric structure feature, the mapped texture feature, and the three-dimensional geometric structure feature and texture feature in the third face image.
11. The method according to claim 10, wherein, Based on the second geometric structure features, the mapped texture features, and the three-dimensional geometric structure features and texture features in the third face image, the first facial features are obtained, including: Based on the three-dimensional geometric structure features and texture features in the third face image, the second query vector of the second attention network is obtained; Based on the second geometric structure features and the mapped texture features, the second key vector and the second value vector of the second attention network are obtained respectively. The first facial features are obtained using the second attention network based on the second query vector, the second key vector, and the second value vector.
12. The method according to any one of claims 3-11, wherein, Based on the first facial feature and the second facial feature extracted from the third facial image, the user's personalized facial features are obtained, including: Determine the weights applied to the first facial feature; Based on the weights and the first facial feature, a weighted first facial feature is obtained; The user's personalized facial features are obtained based on the weighted first facial features and the second facial features.
13. The method according to claim 12, wherein, Determining the weights applied to the first facial feature includes: The weight is determined based on at least one of the geometric structure consistency information between the third face image and the first face image and the weight control information input by the user.
14. The method according to claim 13, wherein, The weight is determined based on at least one of the geometric structural consistency information between the third face image and the first face image and the weight control information input by the user, including: Based on the pose features in the third face image and the pose features in the first face image, the three-dimensional geometric structure features in the enhancement features or the three-dimensional geometric structure features in the first face image are transformed into a second geometric structure feature corresponding to the pose in the third face image. Based on the second geometric structure features and the three-dimensional geometric structure features in the third face image, the geometric structure consistency information is obtained.
15. The method according to claim 14, wherein, Based on the second geometric structure features and the three-dimensional geometric structure features in the third face image, the geometric structure consistency information is obtained, including: Normalization processing is performed on the second geometric structure feature and the three-dimensional geometric structure feature in the third face image, respectively. Based on the normalized features, a fully connected network is used to obtain the geometric structure consistency information.
16. The method according to claim 1, wherein, Based on the personalized facial features, image restoration is performed on the third face image to obtain a fourth face image corresponding to the third face image, including: Obtain the artifact image corresponding to the third face image; The artifact image is degraded to obtain a degraded artifact image, which is used as a negative sample for the diffusion model for image restoration. Based on the negative sample, the first face image, and the personalized facial features, the diffusion model is used to perform a diffusion process on the third face image to obtain a first reconstructed feature corresponding to the negative sample and a second reconstructed feature corresponding to the first face image; Based on the first reconstruction feature and the second reconstruction feature, the reconstruction features of the third face image are obtained; A fourth face image corresponding to the third face image is obtained based on the reconstructed features.
17. A method performed by an electronic device, comprising: Obtain the artifact image corresponding to the user's third-party face image; The artifact image is degraded to obtain a negative sample related to the third face image; Based on the user's first face image, obtain positive samples related to the third face image; Based on the negative samples and the positive samples, image restoration is performed on the third face image to obtain a fourth face image corresponding to the third face image.
18. A method performed by an electronic device, comprising: Based on the user's third-party face image, the reconstructed features of the third-party face image are obtained; Based on the reconstructed features, the decoding features of the third face image are obtained; Based on the user's personalized facial features, the decoded features are texture-corrected to obtain texture-corrected decoded features; A fourth face image corresponding to the third face image is obtained based on the texture-corrected decoding features.
19. An electronic device comprising: At least one processor; as well as At least one memory that stores computer-executable instructions. Wherein, when the computer-executable instructions are executed by the at least one processor, they cause the at least one processor to perform the method as described in any one of claims 1 to 18.
20. A computer-readable storage medium for storing instructions, wherein, When the instruction is executed by at least one processor, it causes the at least one processor to perform the method as described in any one of claims 1 to 18.