A method for transferring visual texture and emotion between face images
The method addresses the limitations of generative AI models by using sparse representation and optimization to transfer facial textures and emotions efficiently, producing high-quality images on low-cost hardware without datasets, overcoming the challenges of high computational demands and biases.
Patent Information
- Application Number
- PCT/TR2024/051948
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-08-28
AI Technical Summary
Existing generative AI models, particularly GANs, require large datasets, high-cost hardware, and long training times, leading to inconsistent and noisy outputs, especially in transferring facial textures and emotions, and are prone to biases and high computational demands.
A method utilizing sparse representation and optimization techniques, including Discrete Cosine Transform (DCT) and Gradient Descent, processes two images to transfer emotional states and facial textures without a dataset, using a global transformation function and image pyramids for efficient, low-cost image generation.
Generates noise-free, consistent, and predictable human face images with accurate facial textures and emotions, reducing computational and financial costs while maintaining high output quality.
Smart Images

Figure TR2024051948_28082025_PF_FP_ABST
Abstract
Description
[0001] A METHOD FOR TRANSFERRING VISUAL TEXTURE AND EMOTION BETWEEN FACE IMAGES
[0002] Technical Field
[0003] This invention pertains to a method for transferring emotional states and facial textures between two human face photographs using optimization techniques, without requiring any dataset. It is applicable in various fields, including forensic investigations, the entertainment and social media sectors, pre-procedural studies in medical aesthetics, synthetic dataset generation for artificial intelligence software development, and identity anonymization in images.
[0004] Prior Art
[0005] Today, with the increasing number of generative artificial intelligence models and advancements in hardware capabilities (enabled by GPUs that allow parallel computing), software developments in synthetic image generation have significantly progressed. However, the complexity of generative Al models, the need for high-cost hardware, and the requirement for large datasets with numerous examples are among the disadvantages of these models. In addition to these challenges, the visual outputs produced by generative Al models may suffer from issues such as a lack of realism, inconsistencies in details, and noise-related distortions. Furthermore, the potential bias in datasets can directly impact the outputs of generative Al models. Controlling these disadvantages is not always possible; thus, the final generated images may contain inconsistencies and unrealistic details. Specifically, in generative Al models designed for creating human faces, unreal regions in the images can be easily noticeable. Moreover, there are software solutions, particularly used in the film industry, that create a deformable 3D model of a human face and map it back onto the face. However, in this approach, progress is made using 3D modeling and computer graphics techniques.
[0006] Previous emotion transfer methods developed using Generative Adversarial Networks (GANs) have shown a high rate of inconsistencies in their outputs, even though some examples appear flawless. The development of these methods required large datasets containing extensive data and high-cost hardware. Training generative Al models can take days or even months. Additionally, the outputs of GAN-based emotion transfer methods are not fully successful in transferring texture elements such as wrinkles on human skin. In contrast, our invention successfully transfers both emotional states and facial textures (such as wrinkles and dimples) between images, achieving a more accurate and reliable result.
[0007] Modem generative artificial intelligence models, particularly advanced techniques like Generative Adversarial Networks (GANs), require significant computational power. This necessitates the use of powerful hardware capable of parallel computation, such as GPUs (Graphics Processing Units), which are often very costly. The success of generative Al models heavily depends on the size and diversity of the datasets used. However, collecting, processing, and cleaning large and diverse datasets is a time-consuming and expensive process. Additionally, biases in datasets can introduce prejudices into the model's outputs, posing serious concerns, particularly when considering issues of social justice and ethics. A lack of demographic or cultural diversity in the training datasets can result in the model failing to reflect such diversity in its outputs. Generative Al models, especially those using GANs, often face issues such as a lack of realism and inconsistencies in details in the generated visuals. While the images generated by these models may resemble real-world objects and human faces, closer examination often reveals that the details lack authenticity. For example, a generative Al model working on human faces might struggle to accurately represent facial features such as eyes, nose, and mouth or fail to maintain proper proportional relationships between these features. These inconsistencies limit the practical applicability of the model's outputs in real-world scenarios.
[0008] In the current state of the art, studies aimed at similar purposes, such as image generation, exhibit a variety of approaches, with methods based on deep learning techniques standing out in this field. Artificial intelligence models such as Generative Adversarial Networks (GANs) and Autoencoders have achieved notable successes in this domain. GANs operate on the principle of competition between two opposing networks (a generator and a discriminator), while Autoencoders aim to compress input data and then reconstruct outputs as close as possible to the original input. Both systems have driven significant advancements in areas such as visual content generation, image editing, and restoration. However, despite their achievements, these technologies come with significant limitations and challenges. Firstly, models like GANs and Autoencoders typically require millions of example data points for training. This necessitates processes such as collecting, processing, and labeling large and diverse datasets. Preparing such datasets is not only time-consuming but also costly, posing a significant barrier for researchers and initiatives with limited resources.
[0009] Secondly, training these types of models requires significant computational power, typically provided by high-performance Graphics Processing Units (GPUs). However, the high cost of such hardware poses a serious challenge, especially for those with limited budgets. The long training times and hardware requirements are among the factors that limit the practicality of these techniques. Thirdly, the outputs produced by systems like GANs and Autoencoders can sometimes contain uncontrollable distortions and noise. In particular, images generated by GANs may exhibit unexpected deformations or unwanted noise in certain regions, highlighting the limitations of the model's generative capacity. These issues reduce the usability of the model's outputs, especially in applications requiring high levels of detail. These limitations present significant challenges for researchers and technology developers aiming to advance image generation. Therefore, it is of great importance to explore new methods and models that can be trained with less data, at lower costs, and in shorter timeframes, while maintaining high output quality.
[0010] The invention in this application has made a significant contribution to the field of generative models and image processing optimization by creating new images through the transfer of emotions and textures in human face images. The methods adopted in this invention involve transforming images into a sparse representation using a global transformation function, the Discrete Cosine Transform (DCT). This sparse representation is then processed at different resolution pyramid levels of the image using a specialized optimization algorithm, Gradient Descent, to generate the output image after a certain number of iterations. Additionally, the number of iterations is automatically adjusted based on image quality. Remarkably, these processes can be carried out using only two images, without requiring any dataset, highlighting a unique and distinctive feature of this invention that warrants protection.
[0011] The human face images generated by the invention provide noise-free, consistent, and predictable results while operating efficiently on low-cost hardware. The methodology, which combines sparse representation and optimization without requiring any dataset, sets this approach apart by offering advantages in cost, time efficiency, and stable outputs. Additionally, one of the innovations introduced by this invention is the application of sparse representation and optimization processes exclusively on 2D images. This offers a novel perspective in the fields of image processing and generative model optimization, broadening the range of practical applications for the invention.
[0012] In the current state of the art, there is no explanation regarding the technical features present in the invention or the technical effects it provides. Existing applications do not include a method for transferring visual texture and emotion between face images that possesses the technical characteristics described above.
[0013] Detailed Description of the Invention
[0014] The method for transferring visual texture and emotion between face images, implemented to achieve the purpose of this invention, is illustrated in the attached figure, which depicts;
[0015] Figure 1. A view of the flow diagram for the method of transferring visual texture and emotion between face images.
[0016] The steps of the flow diagram in Figure 1 are numbered, and the corresponding descriptions of these numbers are provided below.
[0017] 100. The method for transferring visual texture and emotion between face images
[0018] This invention relates to a method (100) for transferring emotional state and facial texture between two human face photographs using optimization techniques, without requiring any dataset. It can be utilized in forensic investigations, the entertainment and social media sectors, pre-procedural studies in medical aesthetics, synthetic dataset generation for artificial intelligence software, and image identity anonymization. Its distinguishing feature is , including processing steps consisting of four main sections: pre-processing and face detection (101), geometric transformation (102), sparse representation (103), and optimization (104), primarily utilizing two different input images, where one serves as the image whose texture and emotional state are to be altered, and the other serves as the reference image, the foundation of the method (100) for transferring emotional state and facial texture being the process of transferring the texture and emotional state of the reference image to the image whose attributes are to be altered, in the pre-processing and face detection step (101), the color histogram of the input image is transferred to the reference image (histogram equalization) to ensure color consistency in the resulting image to be generated (101.1), subsequently, detecting human faces in the input and reference images and determining the coordinates of facial landmarks using a pre-trained machine learning model (101.2), storing these coordinates for use in the geometric transformation (102) step (101.3), in the geometric transformation step (102), geometrically transforming the facial landmark points of the input image to align with the landmark points of the reference image (102.1), using the coordinates of the facial landmark points to geometrically manipulate the input image with Delaunay Triangulation and Voronoi Diagram methods, making it resemble the reference image (102.2), the sparse representation step, where the images are sparsely represented to enable the transfer of texture details on the human face, and the sparse representation of the images is performed using a selected universal dictionary component (103), in the sparse representation step (103), selecting a universal dictionary based on the Discrete Cosine Transform (DCT), Performing a transformation from the RGB color space to the HSI color space, Dividing the image into equal patches, Representing the divided patches sparsely, Merging the said patches,, Transforming the image back to the RGB color space, and Extracting facial textures using the difference operator
[0019] The optimization step (104), where the final output is obtained by filtering the reference texture image based on facial landmarks from the previous steps, and optimizing the geometrically transformed input and reference images using the Gradient Descent optimization algorithm through a number of iterations automatically determined, In the optimization step (104), creating two-level image pyramids to make the optimization process more effective and faster, and repeating the optimization at each level of the pyramid, o Filtering the texture image, o Designing the objective function, o Operating on optimization image pyramids to reduce computational load and process details at different resolutions during the optimization process, o Performing gradient descent optimization.
[0020] The invention, a method (100) for transferring emotional state and facial texture, has made a significant contribution to the fields of generative models and image processing optimization by creating new images through the transfer of emotions and textures in human face images. In this method (100), the adopted approach involves transforming images into a sparse representation using a global transformation function, the Discrete Cosine Transform (DCT). This sparse representation is then processed at different resolution pyramid levels using a specialized optimization algorithm, Gradient Descent, to generate the output image after a certain number of iterations, with the iteration count being automatically adjusted based on image quality. Remarkably, these processes can be executed using only two images without requiring any dataset, highlighting a unique and distinctive feature of the invention that deserves protection.
[0021] The method (100) for transferring emotional state and facial texture in the invention produces human face images that are noise-free, consistent, and predictable, while also being capable of running efficiently on low-cost hardware. The methodology, which combines sparse representation and optimization without the need for any dataset, sets this approach apart by offering significant advantages in terms of cost, time efficiency, and stable outputs. Additionally, one of the innovations introduced by the method (100) is the application of sparse representation and optimization exclusively on 2D images. This provides a novel perspective in the fields of image processing and generative model optimization, significantly broadening the range of practical applications for the invention.
[0022] In one application of the invention, the method (100) for transferring emotional state and facial texture comprises processing steps organized into four main sections: pre-processing and face detection (101), geometric transformation (102), sparse representation (103), and optimization (104). The method (100) fundamentally utilizes two different input images, where one serves as the image whose texture and emotional state are to be altered, and the other serves as the reference image. The core of the method (100) involves transferring the texture and emotional state of the reference image to the image whose attributes are to be modified.
[0023] In one application of the method (100) for transferring emotional state and facial texture, a series of detailed steps are carried out during the sparse representation step (103). First, the image is transformed from the RGB (Red, Green, Blue) color space to the HSI (Hue, Saturation, Intensity) color space. This transformation preserves color information while reducing the image from three channels to a single channel, the I (Intensity) channel. This step is crucial as it allows the sparse representation process to be performed solely on the I channel while ensuring the retention of color information.
[0024] The image transformed into the HSI color space is then divided into overlapping 8x8 pixel squares on the I channel. This division process breaks the image into more manageable and easily processed patches. These small patches are represented as sparse vectors using the DCT (Discrete Cosine Transform) dictionary. The distinctive feature of these vectors is that five of their elements are non-zero, while the remaining elements are zero. The selection of which five elements will be non-zero, along with their values, is determined by running the Orthogonal Matching Pursuit (OMP) algorithm on the DCT dictionary.
[0025] The sparsely represented patches are then merged to complete the I channel of the image. At this stage, the sparsely represented I channel is combined with the unprocessed H and S channels and transformed back into the RGB color space, resulting in a colored image. This color conversion ensures that the image processed through sparse representation retains its original colors.
[0026] Finally, when the sparsely represented images are mathematically subtracted from the original images, facial texture features (such as wrinkles, dimples, beards, and eyebrows) become more pronounced. This process, known in the literature as Cartoon-Texture Separation, enables the detailed extraction of facial textures. By highlighting the textural features of the face as a cartoon-like image, this method achieves a significant step in emotion and texture transfer. This process represents an innovative approach of the invention, capable of being performed between just two images without the need for any dataset. In one application of the method (100) for transferring emotional state and facial texture, the optimization step (104) includes several noteworthy actions. In the first step, the texture image obtained from the reference image is filtered before being transferred to the input image. During the filtering process, regions on the face that do not require transfer (such as hair and ears) are masked using black via a masking technique. This masking process is performed automatically with the aid of facial landmarks.
[0027] The objective function to be minimized during the optimization process of the invention includes three distinct loss functions: Mean Squared Error Loss, Total Variation Loss, and Color Distribution Loss. The sum of these three functions is referred to as the Combined Loss, which is optimized using the gradient descent algorithm. Notably, the Color Distribution Loss stands out as one of the innovative contributions of the invention to the literature.
[0028] The explanations for the symbols used in the formulas are as follows: 'A' symbolizes the input image being processed. 'B' denotes the reference image used for comparison. 'Ahw' represents the value of a specific pixel at the (h, w) coordinates of the input image. 'M' refers to the texture image filtered to include only the texture features intended for transfer. 'H' and 'W' correspond to the height and width of the image, respectively. The Mean Squared Error Loss is a metric that measures the mean of the squared differences between the pixels of two images. As shown in Formula 1, this calculation is weighted by the filtered texture image, ensuring that the operation is conducted only on the relevant texture regions. The difference between the two images is summed across the height (H) and width (W) dimensions and weighted using the M matrix.
[0029] Formula 2: Total Variation Loss, as defined in Formula 2, measures how free a given image is from signal noise. This is calculated by summing the variation between pixels in the image, helping to preserve smoothness in the image.
[0030] Formula 3:
[0031] HistA'- Represents the color distribution histogram of image A.
[0032] HistB'. Represents the color distribution histogram of image B . b: Represents the parameter that determines which pixels the distribution function will be based on.
[0033] Color Distribution Loss, as explained in Formula 3, measures the difference between the color histograms of two images. This calculation is expressed as the sum of the differences in the distributions of both images within a specific color range.
[0034] Formula 4:
[0035] Finally, the Combined Loss, defined in Formula 4, is the weighted sum of the three loss functions described above. In this sum, the values of a, P, and y are coefficients that determine the impact of each loss function on the combined loss, and these values are determined experimentally. The objective function operates on a two-level image pyramid created using Gaussian Image Pyramids to reduce computational load and optimize details at different resolutions. The gradient descent optimization algorithm is used to minimize this complex objective function.
[0036] In the method (100) for transferring emotional state and facial texture, which is the subject of this invention, visual inspection of the algorithm outputs reveals that the results are consistent.
[0037] Furthermore, there are no elements in the images that would create an unrealistic impression.
Claims
CLAIMS1. This invention relates to a method (100) for transferring emotional state and facial texture between two human face photographs using optimization techniques, without requiring any dataset. It is applicable in forensic investigations, the entertainment and social media sectors, pre -procedural studies in medical aesthetics, synthetic dataset generation for artificial intelligence software, and image identity anonymization, characterized by; including processing steps organized into four main sections: pre-processing and face detection (101), geometric transformation (102), sparse representation (103), and optimization (104), primarily utilizing two different input images, where one serves as the image whose texture and emotional state are to be altered, and the other serves as the reference image, the foundation of the method (100) for transferring emotional state and facial texture being the process of transferring the texture and emotional state of the reference image to the image whose attributes are to be modified, comprising these steps.
2. The method (100) for transferring emotional state and facial texture as mentioned in Claim 1, characterized by including a pre-processing and face detection step (101), o transferring the color histogram of the input image to the reference image (histogram equalization) to ensure color consistency in the resulting image o subsequently detecting human faces in both the input and reference images and identifying the coordinates of facial landmarks using a pre-trained machine learning model o storing these coordinates for use in the geometric transformation step (102) comprising these steps.
3. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by including a geometric transformation step (102), o geometrically transforming the facial landmark points of the input image to align with the landmark points of the reference image o geometrically manipulating the input image using Delaunay Triangulation and Voronoi Diagram methods based on the coordinates of the facial landmark points to resemble the reference imagecomprising these steps.
4. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by including a sparse representation step (103), where the images are sparsely represented to enable the transfer of texture details on the human face, and the sparse representation is performed using a selected universal dictionary component.
5. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by including, in the sparse representation step (103) o Selecting a universal dictionary based on Discrete Cosine Transform (DCT), o Converting the image from the RGB color space to the HSI color space, o Dividing the image into equal patches, o Representing the divided patches sparsely, o Merging the said patches, o Converting the image back to the RGB color space, and o Extracting facial textures using the difference operator, comprising these steps.
6. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by including an optimization step (104) where the final output is obtained by filtering the reference texture image, obtained in the previous steps, based on facial landmarks, and optimizing the geometrically transformed input and reference images using the Gradient Descent optimization algorithm through a number of iterations automatically determined.
7. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by; creating two-level image pyramids in the optimization step (104) to make the optimization process more effective and faster, and repeating the optimization at each level of the pyramid, o Filtering the texture image, o Designing the objective function, o Working on optimization image pyramids to reduce computational load and process details at different resolutions during the optimization process, o Performing gradient descent optimization, comprising these steps.
8. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by transforming images into a sparse representation using a global transformation function, the Discrete Cosine Transform(DCT), processing this sparse representation at different resolution pyramid levels of the image using a specialized optimization algorithm, Gradient Descent, to generate the output image after a certain number of iterations, and automatically adjusting the number of iterations based on image quality.
9. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by applying sparse representation and optimization processes exclusively on 2D images.
10. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by primarily utilizing two different input images, where one serves as the image whose texture and emotional state are to be altered, and the other serves as the reference image.
11. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by including, in the sparse representation step (103), first converting the image from the RGB (Red, Green, Blue) color space to the HSI (Hue, Saturation, Intensity) color space to enable sparse representation to be performed solely on the I channel while preserving color information, and subsequently dividing the image transformed into the HSI color space into overlapping 8x8 pixel squares on the I channel.
12. The method (100) for transferring emotional state and facial texture as mentioned in Claim-11, characterized by including the step of representing the obtained small patches as sparse vectors using a DCT (Discrete Cosine Transform) dictionary.
13. The method (100) for transferring emotional state and facial texture as mentioned in Claim- 12, characterized by the vectors having the property that five of their elements are non-zero while the remaining elements are zero, with the specific non-zero elements and their values being determined by running the Orthogonal Matching Pursuit (OMP) algorithm on the DCT dictionary.
14. The method (100) for transferring emotional state and facial texture as mentioned in Claim- 13, characterized by merging the sparsely represented patches to complete the I channel of the image, and transforming the sparsely represented I channel, along with the unprocessed H and S channels, back to the RGB color space to obtain a colored image.
15. The method (100) for transferring emotional state and facial texture as mentioned in Claim- 14, characterized by including the step of mathematically subtracting the sparsely represented images from the original images to enhance the visibility of facial texture features such as wrinkles, dimples, beards, and eyebrows.
16. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by the objective function to be minimized during the optimization process including three distinct loss functions: Mean Squared Error Loss, Total Variation Loss, and Color Distribution Loss, the sum of these three functions being referred to as the Combined Loss, and being optimized using the gradient descent algorithm.
17. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by the Mean Squared Error Loss, a metric that measures the mean of the squared differences between the pixels of two imagesIs calculated using the formula, where,'A' represents the input image being processed'B', denotes the reference image for comparison'Ahw' refers to the value of a specific pixel at the (h, w) coordinates in the input image.'M', represents the texture image that has been filtered to include only the desired texture features for transfer'H' ve 'W', are values representing the height and width of the image, respectively.
18. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by Total Variation Loss, which measures how free an image is from signal noise and helps preserve smoothness in the image,is calculated using the formula.
19. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by Color Distribution Loss, which measures the difference between the color histograms of two images and is expressed as the sum of the differences in the distributions of both images within a specific color rangeis calculated using the formula, where;HistA'- Represents the color distribution histogram of image A.HistB'. Represents the color distribution histogram of image B. b: Represents the parameter that specifies which pixels the distribution function is based on.
20. The method (100) for transferring emotional state and facial texture as mentioned in any of the preceding claims, characterized by the Combined Loss, which is the weighted sum of the Mean Squared Error Loss, Total Variation Loss, and Color Distribution Loss, operating on a two-level image pyramid determined using Gaussian Image Pyramids to reduce computational load and optimize details at different resolutions;Is calculated using the formula, where a, P, and y are coefficients that determine the impact of the respective loss functions on the Combined Loss, and these values are determined experimentally.
Citation Information
Patent Citations
Method and system for facial expression transfer
US20160004905A1
Training method for expression transfer model, expression transfer method and apparatus
US20220245961A1
Expression transfer method, model training method, and device
WO2023142886A1