Image emotion synthesis method based on diffusion model

By adopting a diffusion model-based method in image emotion synthesis, combining single-step inference, emotion injection and multi-emotional cross-entropy loss function, the problem of lack of flexible frameworks in the prior art to deal with emotion generation and editing tasks is solved, and high emotional accuracy and high quality image emotion synthesis is achieved.

CN120070668AActive Publication Date: 2025-05-30NANKAI UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510143788.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The prior art lacks a flexible framework that can handle both the tasks of image emotion generation and editing at the same time with high emotional accuracy and high quality.

Method used

Using a diffusion model-based image emotion synthesis method, the timing and optimization process of emotion labels are controlled by using single-step inference and emotion injection in hidden spaces, combining pre-trained emotion classifiers and multi-emotional cross-entropy loss function.

Benefits of technology

It realizes efficient completion of emotional image synthesis in emotional generation and editing tasks, improves emotional accuracy and image quality, and avoids the problem of image content offset or poor emotional effect caused by premature or late emotional guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070668A_ABST
    Figure CN120070668A_ABST
Patent Text Reader

Abstract

The invention provides an image emotion synthesis method based on a diffusion model, and belongs to the technical field of image emotion synthesis. The image emotion synthesis method based on the diffusion model comprises the following steps: S1, embedding an image to be generated in a hidden space; s2, performing emotion injection on the input prompt description; s3, for the clean image embedding in the S1, optimizing the emotion label in the S2 by minimizing emotion loss through reverse gradient optimization; s4, controlling the emotion injection opportunity according to the image-text similarity; and S5, optimizing the emotion label by using a multi-emotion cross entropy loss function. The image emotion synthesis framework optimizes emotion tags by minimizing cross entropy emotion loss, and introduces the tags at the best opportunity to guide image emotion synthesis. The image emotion synthesis framework is superior to an existing emotion image generation method in the aspects of emotion generation task accuracy and image quality, and meanwhile, the image emotion synthesis framework is also excellent in emotion editing task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image emotion synthesis. Specifically, it relates to an image emotion synthesis method based on a diffusion model. Background Art

[0002] Currently, image emotion synthesis is an emerging technology aimed at generating or editing image content according to text prompts to evoke specific emotions in users. Existing methods usually regard emotion generation and editing as independent tasks. Due to the subjectivity and abstractness of emotions, these methods usually associate the target emotion with specific attributes (such as color and style), and use large visual language models or fine-tune the model on emotion datasets to modify these attributes. The core idea of these methods usually strongly couples emotions with certain fixed elements, thus limiting the performance of emotion editing and generation tasks in terms of semantic diversity and other aspects.

[0003] Early emotion image editing methods mainly transmitted emotions by editing low-level features such as color and texture. For example, some methods selected theme candidates by establishing the relationship between color and emotion words, while others combined an emotion classifier for color constraint and used a coloring network to enhance emotion expression. In addition, higher-level style information is also important for emotion expression. Some studies have proposed multi-emotion semantic spaces to achieve multi-emotion transfer; some studies have also extracted emotion clues from text to guide color conversion.

[0004] Generative models (such as diffusion models) have been successfully applied to complex tasks such as image and video generation, and have been used for emotion generation for the first time. Researchers have realized the mapping between emotion and semantic space through a network, enabling abstract emotions to be transformed into specific image attributes. In addition, emotion editing methods based on large models have also attracted wide attention. For example, some methods use GPT-4 to extract an emotion factor tree to determine the most relevant emotion attributes, and then modify them through a specific model; The above existing technical solutions have the following defects: Image emotion synthesis aims to generate media content consistent with specific emotions, and its methods are mainly divided into two major tasks: image generation and image editing. Currently, there is still a lack of a flexible framework that can simultaneously process emotion generation and editing tasks with high emotion accuracy and high quality. Summary of the Invention

[0005] To make up for the above deficiencies, this application provides an image emotion synthesis method based on a diffusion model, aiming to improve the problem of lacking a flexible framework that can simultaneously process with high emotion accuracy and high quality.

[0006] The embodiments of this application provide an image emotion synthesis method based on a diffusion model, including the following steps: S1. For the image embedding to be generated in the hidden space, use single-step inference at each time step, approximate the inversion of the noise map, and obtain the estimated clean image embedding; S2. Inject emotion into the input prompt description, append additional emotion labels to the original text, and input it into the diffusion model during the emotion guidance process. Update the weight of the new emotion label at each inference step; S3. For the clean image embedding in S1, perform emotion classification using a pre-trained emotion classifier, use cross-entropy loss to calculate the gap with the target emotion, and optimize the emotion label in S2 by minimizing the emotion loss through backpropagation of the gradient; S4. Control the timing of emotion injection according to the text-image similarity. For the clean image embedding described in S1, obtain the text-image semantic similarity score at each step. When the semantic similarity between the synthesized image and the target reaches or exceeds the set threshold, introduce the emotion label to ensure the timing of emotion expression and the diversity of semantics; S5. For the emotion loss, use a multi-emotion cross-entropy loss function to optimize the emotion label.

[0007] In a preferred embodiment of the present invention, for the single-step inference of the noise map in S1, the deterministic sampling characteristic of the DDIM model is utilized, which allows for the rapid generation of high-quality images from a small number of sampling steps; without adding additional random noise, the original image's latent representation is restored by predicting the noise and gradually reducing the noise.

[0008] In a preferred embodiment of the present invention, for the emotion label appended after text embedding in S2, the emotion label is connected to the original text prompt using the same text encoder with additional text.

[0009] Do not directly backpropagate the predicted noise to avoid conflicts between the gradients of the emotion classifier and the gradients generated by the text prompt, more stably integrate emotion information, and maintain the consistency of the text content.

[0010] In a preferred embodiment of the present invention, for the emotion image generated at each time step in S3, after single-step inference, it is input into the emotion classifier to calculate its emotion label, and the similarity cross-entropy loss with the target emotion is calculated, and the gradient is optimized in the reverse direction to optimize the weight of the added emotion label.

[0011] In a preferred embodiment of the present invention, in the process of image generation in S4, cosine similarity is used to calculate the text and image features, and at the same time, a temperature factor is used to scale the result to obtain the text-image semantic similarity.

[0012] Better control is exerted over it. Because introducing emotional guidance prematurely will cause deviation in the image content, and the complexity and uncontrollability of emotions will conflict between the latent space and the image guidance, resulting in the loss of the content of the text prompt; introducing emotions too late will make it difficult for the image to achieve the expected emotional effect in the later stage of image generation.

[0013] In a preferred embodiment of the present invention, CLIP is used to obtain the graphic and text semantic similarity score, and the score is compared with a set threshold.

[0014] In a preferred embodiment of the present invention, the image and text encoders of CLIP are utilized and At each time step Measure the scaled CLIP score between the generated content and the text prompt (temperature factor ): Wherein, represents the temperature factor, represents the text prompt, represents from obtained denoised embedding.

[0015] In a preferred embodiment of the present invention, in S5, due to the diversity and complexity of emotions, the target emotion will be affected by the native emotion under the text description and the similar emotions in the emotion wheel. The new emotion loss adopts contrastive learning and respectively uses cross-entropy loss to suppress it.

[0016] In a preferred embodiment of the present invention, in order to make the synthesis process approach the target emotion, the following target emotion loss is proposed: Wherein, is a loss function based on the target emotion, is the inherent emotion loss, is the similar emotion loss, and are hyperparameters.

[0017] In a preferred embodiment of the present invention, based on the classifier-based emotion synthesis framework, the tasks of emotion image generation and editing are simultaneously completed in terms of the accuracy of emotion arousal and the image quality.

[0018] Beneficial effects: 1. The tasks of emotion image generation and emotion editing are efficiently completed through the three aspects of the method of emotion guidance, the timing of emotion guidance, and the accurate injection of emotions.

[0019] 2. Specifically, the synthesis is guided by the prior knowledge of the pre-trained sentiment classifier, and the predefined sentiment labels are optimized using gradient backpropagation to ensure synthesis stability.

[0020] 3. Both premature or too late sentiment guidance will damage the image content or fail to add the target sentiment; therefore, semantic similarity is used as a supervision signal to determine the best timing for applying sentiment guidance.

[0021] 4. A multi-sentiment cross-entropy loss is proposed to reduce the interference of inherent and similar sentiments, thereby improving the accuracy of sentiment synthesis; the image sentiment synthesis framework optimizes the sentiment labels by minimizing the cross-entropy sentiment loss and introduces these labels at the best timing to guide image sentiment synthesis. The image sentiment synthesis framework is superior to the current sentiment image generation methods in terms of the accuracy of the sentiment generation task and the image quality, and also performs excellently in the sentiment editing task. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is the overall structure of the sentiment synthesis framework proposed by the present invention; Figure 2 It is the comparison diagram of the semantic similarity control effect proposed by the present invention; Figure 3 It is the comparison diagram of the effect of the sentiment suppression method proposed by the present invention; Figure 4 It is the actual effect diagram of the sentiment image generated by the present invention; Figure 5 It is the actual effect diagram of the sentiment image editing by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] In the present invention, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may include the direct contact between the first and second features, or may include the indirect contact between the first and second features through other features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes the first feature being directly above and obliquely above the second feature, or simply indicating that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "below" and "beneath" the second feature includes the first feature being directly below and obliquely below the second feature, or simply indicating that the horizontal height of the first feature is lower than that of the second feature.

[0025] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.

[0026] Please refer to Figures 1-5 , the present invention provides an image emotion synthesis method based on a diffusion model, as Figure 1 shown, including the following steps: S1. For the image embedding to be generated in the latent space, use single-step inference at each time step, approximate the inversion of the noise map, and obtain the estimated clean image embedding; S2. Inject emotion into the input prompt description, append additional emotion labels to the original text, input it into the diffusion model during the emotion guidance process, and update the weight of the new emotion label at each inference step; S3. For the clean image embedding in S1, perform emotion classification using a pre-trained emotion classifier, calculate the gap with the target emotion using cross-entropy loss, and optimize the emotion label in S2 by minimizing the emotion loss through backpropagation of gradients; S4. Control the timing of emotion injection according to the text-image similarity. For the clean image embedding described in S1, obtain the text-image semantic similarity score at each step. When the semantic similarity between the synthesized image and the target reaches or exceeds the set threshold, introduce the emotion label to ensure the timing of emotion expression and ensure semantic diversity; S5. For the emotion loss, use a multi-emotion cross-entropy loss function to optimize the emotion label.

[0027] It allows for the rapid generation of high-quality images from a small number of sampling steps; without adding additional random noise, the latent representation of the original image is recovered by predicting the noise and gradually reducing the noise.

[0028] For the generation of an emotion image, a given initial random noise image is simultaneously input with the corresponding text prompt P. The text prompt P is input into the text encoder to obtain the initial text embedding Before emotion guidance, the text embedding and the initial random noise are input into the diffusion model to predict the next noise, where z 0 represents the initial latent embedding obtained by the given image I 0 through the VAE encoder, represents the gradually decreasing noise intensity: The image embedding at the next noise t-1 can be calculated from the predicted noise, Among them, the network parameter θ minimizes the mean squared error (MSE) between the predicted noise and the true noise, where represents sampling samples z from the standard normal distribution 0 , t: After the model training is completed, the initial embedding distribution is recovered from the noise samples z in an iterative manner. z T is sampled from the distribution z T is sampled from the distribution z T ~q(z T ). To introduce conditional information c during the generation process, a classifier-free guidance diffusion method is used, and a pre-trained classifier p(c∣x t ) is used to guide the noise ∈ θ in the inference process, and the guidance strength of the condition c is controlled by a scalar ω greater than 0 To support more complex conditions, the emotion guidance uses the method of replacing the explicit classifier with an implicit classifier: during the training process, no conditions are used, but the same network is used to calculate and This method is called "Classifier-Free Guidance (CFG)", and is also controlled by the scalar ω: Image synthesis emotion guidance: An additional pre-trained emotion classifier is used as guidance to utilize the learned emotion prior. The original text prompt P passes through the text encoder and is then concatenated with the emotion label S to obtain the emotion text embedding Here represents the concatenation operation of the embeddings. Then, the denoising embedding is calculated under the condition of c emo To make the emotion classifier pre-trained on clean images adapt well to the noisy data, the DDIM single-step inference method is adopted to obtain the estimated clean embedding from each time step t Specifically as follows: Among them Then, the emotion label S is optimized by the gradient of backpropagation, and the emotion loss between the target emotion y target and the predicted emotion is minimized during the denoising process As follows: The method involves using a pre-trained sentiment classifier to change the weights of sentiment cue words in the text input during each denoising step and performing inference on the current noisy image. The model can learn how to synthesize sentiment in the presence of noise, thereby improving performance and accuracy on noisy data.

[0029] Design optimization of additional sentiment labels Provides two main advantages: (1) Directly utilizes the prior knowledge of state-of-the-art sentiment classifiers, eliminating the need to create a large-scale sentiment manipulation dataset. Given the subjectivity of sentiment and the huge effort required to create such a dataset.

[0030] (2) Taking and together as the framework input allows to consider both semantic and sentiment information in the cues simultaneously when estimating noise. This prevents gradient inconsistency, avoids ignoring text information, and ensures stable synthesis.

[0031] Finally, since the sentiment labels only affect the input layer of the model, no further complex operations on the intermediate embeddings are required. The framework can be easily applied to image generation and editing tasks by simply adjusting the input (P, S) and noise sampling easily.

[0032] Applications in generation: Image sentiment generation can be divided into with and without text descriptions. Both cases start from random noise. Applications in editing: For sentiment editing tasks, the image is input with a text prompt according to its content , and the initial noise is obtained by performing DDIM inversion on the initial image .

[0033] The cosine similarity is used to calculate the text and image features, and the result is scaled using a temperature factor to obtain the text-image semantic similarity; Semantic similarity supervision: Introducing target sentiment guidance at the optimal stage can suppress the inherent sentiment expression while maintaining semantic accuracy, ensuring the expression of the desired sentiment. To achieve this without increasing the complexity of denoising, the image and text encoders of CLIP are utilized and to measure the scaled CLIP score between the generated content and the text prompt at each time step : (temperature factor ): where is obtained from The denoising embedding. During the reverse process, once exceeds a preset threshold , it indicates that the image content has been initially formed, and then emotional guidance is optimized and introduced : is an empty placeholder. This mechanism enables the present invention to introduce the target emotion at an appropriate stage, while ensuring the conveyance of the target emotion, maintaining the diversity and accuracy of semantics, as Figure 2 shows the comparison effect of whether semantic supervision is performed. Conducting emotional guidance at the appropriate time can ensure the quality and accuracy of the image.

[0034] Due to the diversity and complexity of emotions, the target emotion will be affected by the native emotion under the text description and the similar emotions in the emotion wheel. The new emotion loss adopts contrastive learning and uses cross-entropy loss respectively to suppress it; Inherent and similar emotion suppression: During the generation process, the target emotion of guided synthesis should be emphasized. There are two types of adverse emotions that will affect the expression of the target emotion: a) Inherent emotion: Derived from the source image or text prompt. To suppress the inherent emotion, the present invention applies a semantic similarity supervision mechanism to obtain the embedding at time steps before exceeding the threshold . Then use an emotion classifier to predict the inherent emotion from , and construct a loss term: where is the cross-entropy loss, and is the inherent emotion loss.

[0035] b) Similar emotions: Similar emotions are related to the target emotion. Different from other classification tasks, in the emotion classification task, the categories have relative similarity rather than completely independent and mutually exclusive features. They are adjacent in Mikel's emotion wheel mental model. Similar emotions may cause the guided emotion to shift towards them during the synthesis process. The similar emotions related to the target emotion can be obtained by querying the emotion wheel, and cross-entropy loss can be used to suppress them: is the similar emotion loss.

[0036] In summary, to suppress these two undesirable emotions and force the synthesis process to move closer to the target emotion, the following target emotion loss is proposed: Among them, , is a loss function based on the target emotion, while and are hyperparameters used to control the suppression intensity. This loss term can suppress the inherent emotion in the synthesized image. See similar emotions in Figure 3 .

[0037] The target emotion loss combines the suppression of the inherent emotion and the similar emotion, ensuring that the emotion expression in the synthesis process can more accurately conform to the target emotion. In this way, the model can more effectively guide the emotion expression when denoising and synthesizing images, reduce the influence of undesirable emotions, and thus improve the emotion accuracy and quality of the synthesized images.

[0038] In summary, the classifier-based emotion synthesis framework designed by the present invention can simultaneously complete the emotion image generation and editing tasks, and is superior to the existing methods in terms of the accuracy of emotion arousal and the image quality.

[0039] Table 1 Table 2 Table 1 shows the comparison of the present invention with the existing sota tasks in six evaluation metrics in the emotion generation task. The "↑" symbol in the table indicates that the higher the value of this metric, the better the effect, and "↓" indicates the lower the better.

[0040] Table 2 shows the comparison of the present invention with the existing sota tasks in four evaluation metrics in the emotion editing task. The "↑" symbol in the table indicates that the higher the value of this metric, the better the effect.

[0041] For the generation task, evaluation metrics including FID, LPIPS, Sem-C, Sem-D, and emotion accuracy (ACC1 and ACC2) are followed. FID and LPIPS measure the realism and diversity of images respectively, while Sem-C and Sem-D evaluate the semantic clarity and content richness of images. Two classifiers are used to measure emotion accuracy. For the editing task, in addition to the LPIPS, ACC1, and ACC2 metrics used in the generation task, the CLIP image score is also adopted to measure the similarity between the input image and the edited image. The table shows the comparison of the present invention with other sota methods in emotion image generation, image emotion generation under specific prompts, and emotion image editing tasks, including traditional methods and diffusion model-based methods. The results show that the present invention outperforms all published results in emotion image synthesis on multiple tasks.

[0042] The present invention is the first framework to simultaneously complete emotion generation and emotion editing. Figure 4 and Figure 5 respectively show the effects of the present invention in emotion image generation and editing for different emotions. Figure 4 In it, the upper name of each group of images represents the target emotion and the corresponding text prompt description. Figure 5 In, the left side of the input image represents the target emotion. It can be seen that by comparison with different methods, the present invention has better generation and editing effects than other image generation and editing methods, with obvious improvements in image quality, generation accuracy, and diversity of image content.

[0043] Efficiently complete the emotion image generation and emotion editing tasks through three aspects: emotion-guided method, timing of emotion guidance, and accurate injection of emotion.

[0044] Specifically, the synthesis is guided by the prior knowledge of the pre-trained emotion classifier, and the predefined emotion labels are optimized using gradient backpropagation to ensure synthesis stability.

[0045] Premature or too late emotion guidance will damage the image content or fail to add the target emotion; therefore, semantic similarity is used as a supervision signal to determine the best timing for applying emotion guidance.

[0046] A multi-emotion cross-entropy loss is proposed to reduce the interference of inherent emotions and similar emotions, thereby improving the accuracy of emotion synthesis; the image emotion synthesis framework optimizes emotion labels by minimizing the cross-entropy emotion loss and introduces these labels at the best timing to guide image emotion synthesis. The image emotion synthesis framework is superior to current emotion image generation methods in terms of accuracy and image quality in the emotion generation task, and also performs excellently in the emotion editing task.

[0047] A new framework is designed to address the shortcomings of existing emotional image synthesis tasks, such as the inability to accurately evoke specific emotions, lack of semantic diversity, and the separation of emotion generation and editing tasks. A simple and effective image emotion synthesis framework is proposed, which can significantly change the image content and is not dependent on specialized datasets; it does not limit semantic diversity; and it solves the generation and editing tasks through three key aspects: (1) how to generate emotions; (2) when to guide emotions; and (3) which emotions to guide the synthesis process. In addition to image and text prompt conditions, emotional labels are introduced to ensure stable emotion synthesis. The timing of emotion introduction is controlled by using image-text semantic similarity in the denoising process, and the emotional loss function is constructed by suppressing inherent and similar emotions in the synthesis process for optimization, thereby achieving high emotional accuracy. This method outperforms existing state-of-the-art methods in terms of emotional accuracy, semantic diversity, and image quality for generation and editing tasks, and does not require specialized datasets or large visual models.

[0048] The above description is only an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.

Claims

1. A method for image emotion synthesis based on a diffusion model, characterized in that: The steps include: S1. For the image embedding to be generated in the latent space, use single-step inference at each time step to invert the noise map and obtain the estimated clean image embedding. S2, inject emotion into the input prompt description, append additional emotion tags to the original text, input it into the diffusion model in the emotion guidance process, and update the new emotion tag weight at each reasoning step; S3, for the clean image embedding in S1, use the pre-trained sentiment classifier to perform sentiment classification, use cross entropy loss, calculate the gap with the target sentiment, and optimize the sentiment label in S2 by minimizing the sentiment loss through reverse gradient optimization; S4, controlling the timing of emotion injection according to the image-text similarity. For the clean image embedding described in S1, obtaining the image-text semantic similarity score at each step. When the semantic similarity between the synthesized image and the target reaches or exceeds the set threshold, the emotion label is introduced to ensure the time of emotion expression and the diversity of semantics. S5. For sentiment loss, the multi-sentiment cross entropy loss function is used to optimize the sentiment labels.

2. The image emotion synthesis method based on diffusion model according to claim 1, characterized in that: The single-step inference of the noise map in S1 exploits the deterministic sampling properties of the DDIM model, allowing high-quality images to be quickly generated from a small number of sampling steps; without adding additional random noise, the potential representation of the original image is recovered by predicting the noise and gradually reducing it.

3. The image emotion synthesis method based on diffusion model according to claim 1, characterized in that: In S2, the sentiment label is appended after the text embedding, and the sentiment label is concatenated with the original text prompt using the additional text through the same text encoder.

4. The image emotion synthesis method based on diffusion model according to claim 1, characterized in that: In S3, for the emotional image generated at each time step, after single-step inference, it is input into the sentiment classifier to calculate its sentiment label, and the similarity cross entropy loss is calculated with the target sentiment, the gradient is optimized in reverse, and the weight of the added sentiment label is optimized.

5. The image emotion synthesis method based on diffusion model according to claim 1, characterized in that: In the process of image generation, S4 uses cosine similarity to calculate text and image features, and uses a temperature factor to scale the results to obtain the semantic similarity between the image and the text.

6. The image emotion synthesis method based on diffusion model according to claim 1, characterized in that: Use CLIP to obtain the semantic similarity score of the image and text, and compare the score with the set threshold.

7. The image emotion synthesis method based on diffusion model according to claim 6 is characterized in that: Image and text encoder using CLIP and At each time step Measuring the scaled CLIP score between generated content and text prompts : in, represents the temperature factor, Indicates a text prompt. Indicates from The resulting denoised embedding.

8. The image emotion synthesis method based on diffusion model according to claim 1, characterized in that: In S5, due to the diversity and complexity of emotions, the target emotion will be affected by the original emotion under the text description and the similar emotion in the emotion wheel. The new emotion loss adopts contrastive learning and cross entropy loss to suppress it.

9. The image emotion synthesis method based on diffusion model according to claim 8, characterized in that: In order to make the synthesis process closer to the target emotion, the following target emotion loss is proposed: in, is a loss function based on the target sentiment, It's an inherent emotional loss. It’s a similar emotional loss. and is a hyperparameter.

10. The image emotion synthesis method based on diffusion model according to claim 1, characterized in that: A classifier-based emotion synthesis framework that simultaneously completes the emotion image generation and editing tasks in terms of emotion arousal accuracy and image quality.

Citation Information

Patent Citations

  • Image filter generation method based on sentiment analysis

    CN116910294A

  • Text-guided image processing method based on diffusion model

    CN116977489A

  • Facial action prediction method and device, medium and computer program product

    CN118314614A

  • Aspect-level emotion triple extraction method based on diffusion model

    CN119168044A

  • Method of training sentiment preference recognition model for comment information, recognition method, and device thereof

    US20240078384A1