Emotional image content generation method and system based on text editing and correcting
By constructing a data set with text and emotional value annotations and designing an emotion embedding network model, embeding emotion values into text features, the problem of difficulty in controlling content and inability to capture subtle emotional changes in image generation in the prior art is solved, and high-quality and controllable emotional image generation is achieved.
Patent Information
- Application Number
- CN202411951166.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-27
AI Technical Summary
In the prior art, image generation based on emotion tags is difficult to control image content, and image generation cannot be performed from continuous emotional dimensions, making it difficult to capture and express subtle emotional changes.
By constructing a data set with text and affective value annotations, design an emotion embedding network model, embed emotion values into text features, and trained using specific loss functions to generate images consistent with the emotion values entered by the user.
It achieves higher accuracy and controllability in emotional image generation, can generate high-quality images with strong emotional expression and continuity, and supports users' personalized customization of content.
Smart Images

Figure CN120030164A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image content generation, and in particular relates to a method and system for generating emotional image content based on text editing. Background Art
[0002] Emotions are fundamental to the human experience and play a key role in shaping how people perceive and interact with the world. Research shows that emotions influence memory and comprehension, which are essential for effective communication. As a result, content creators are increasingly aware of the importance of incorporating emotions to increase audience engagement and understanding.
[0003] For the technical route of using other conditions to control the generated image content instead of using emotions as control conditions, it is mainly divided into conditional control from the visual level and conditional control from the text control level. Conditional control from the visual level refers to inputting conditional information into the visual content generation structure of the image generation model. For example, IP-adapter integrates the features of the reference image through cross-attention to generate an image consistent with the attributes or identity of the reference image. Conditional control from the text control level refers to inputting conditional features into the text control structure of the text-to-image generation model. For example, by learning the specific concept of a placeholder, images of a specific style or field can be generated, but these methods may have difficulty capturing and expressing more subtle emotional changes.
[0004] Currently, research on image generation using emotion as a condition is still limited. Some early studies have explored emotion image generation techniques in specific fields, such as facial expressions or landscapes. EmoGen generates emotional images by mapping emotion features to the semantic features of the text-generated image model, that is, mapping an emotion concept to certain specific semantics. EmoGen has pioneered the generation of emotional image content. However, it has two key shortcomings:
[0005] (1) Images are generated based on emotional labels rather than textual cues, making the generated content difficult to control;
[0006] (2) Although the discrete emotion labels used by EmoGen are easy to understand, psychologists have not yet reached a consensus on emotion categories. The limited range of discrete emotion labels makes it difficult to capture subtle emotional changes. In contrast, continuous emotion models represent emotions through Cartesian coordinates and can more accurately capture the complexity and nuances of emotions. Therefore, it is very necessary to develop a general text-to-image model that is not limited to discrete emotion labels.
[0007] In addition, existing technologies also face challenges in capturing complex and subtle emotional changes because they mainly rely on specific discrete emotion categories, which limits the depth and breadth of emotional expression of generated images. Existing technologies are also limited in the collection and annotation of datasets. In order to train emotional image generation models, a large number of datasets with emotion annotations are required, but existing datasets often lack sufficient emotional dimension coverage and depth, which also limits the model's ability to learn emotional features. Summary of the invention
[0008] The technical problem to be solved by the present invention is to provide a method and system for generating emotional image content based on text editing, which solves the shortcomings of the image generation technology based on emotional tags in the prior art, namely, the inability to control the image content and the inability to generate images from a continuous emotional dimension.
[0009] The present invention adopts the following technical solutions to solve the above technical problems:
[0010] A method for generating emotional image content based on text editing.
[0011] First, we collect and construct a dataset with text and sentiment value annotations, which includes several images and corresponding sentiment value annotations, as well as text annotations with and without sentiment texts.
[0012] Secondly, a sentiment embedding network model is constructed to embed sentiment values into text features;
[0013] Then, the labeled data in the dataset and the pre-designed loss function are used to train the sentiment embedding network model to obtain the optimal model parameters;
[0014] Finally, the text prompts and sentiment values are input into the trained sentiment embedding network model, and the output of the sentiment embedding network model is passed through the image generation model to obtain the output sentiment image.
[0015] The sentiment embedding network model includes two inputs and one output, wherein one input is a sentiment value, the other input is a text without sentiment or a text feature without sentiment, and the output is a corresponding text with sentiment or a text feature embedded with sentiment.
[0016] The emotion embedding network model is trained by optimizing a loss function, and the loss function is used to constrain the output distribution of the emotion embedding network during the training process so that it fits the target distribution.
[0017] The sentiment embedding network model first encodes the input sentiment value into sentiment features, and then embeds the sentiment features into text features, thereby embedding the sentiment value into the sentiment-embedded text features.
[0018] The loss function includes a scaling residual learning strategy or a sentiment value density weighting strategy, wherein the scaling residual learning strategy linearly enlarges or reduces the residual between the emotion-free text feature and the emotion-containing text feature as the scaled target emotion text feature; the emotion value density weighting strategy weights the loss function by estimating the emotion value distribution density.
[0019] The emotional image generation model based on text editing is connected to the image generation model at the back end of the emotional embedding network model, and the emotional text or emotional text features output by the emotional embedding network model are converted into image form and output.
[0020] According to the output of the image generation model or the output of the emotion embedding network, the emotion embedding network model is cyclically optimized and trained to obtain the optimal network parameters.
[0021] The emotional image content generation system based on text editing includes a user input module, an emotional image generation module, and an output display module, wherein the emotional image generation module is the emotional image generation model based on text editing, and the user input module uses an input device to input text prompts and emotional values; after receiving the input information, the emotional image generation module embeds the emotional value into the encoded emotionless text features, and generates an image based on the emotion-embedded text features; the output display module displays the generated image to the user through a display device.
[0022] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, call all or part of the steps of the method.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] 1. A novel emotion embedding mapping network is proposed, which can embed continuous or discrete emotions into text features, so that it can fuse specific emotions according to the expected input prompts. This method not only improves the accuracy and controllability of emotional image generation, but also generates high-quality images consistent with the emotional values input by the user, with strong emotional expression and continuity.
[0025] 2. A new loss function is also developed to help enhance the emotional resonance of generated images. This loss function modifies the target features to learn more obvious emotional changes and considers the distribution of emotional values to offset the unevenness of the data. By constructing an emotional image dataset for training the proposed emotional embedding network, this technical scheme provides a new solution for the field of emotional image generation.
[0026] 3. It can realize the user's control over the content, so as to customize the personalized emotional image content. Under the condition of controlling both emotion and content, the generation of high-quality image content can be realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a flow chart of the method for generating emotional image content based on text editing according to the present invention.
[0028] Figure 2 This is a schematic diagram of the emotion injection block structure of the present invention.
[0029] Figure 3 This is a graph showing the significance test results of the emotional image content generation method based on text editing of the present invention.
[0030] Figure 4 It is a schematic diagram of the effectiveness results of different scaling factors of the present invention. DETAILED DESCRIPTION
[0031] The structure and working process of the present invention will be further described below in conjunction with the accompanying drawings.
[0032] This application proposes EmotiCrafter, which is a text-to-emotional image generation method and system based on sentiment values, which can automatically generate sentiment images based on text prompts and sentiment values. The method includes an sentiment image automatic generation pipeline, a deep neural network design that embeds sentiment values into text features, a loss function and training strategy, and a dataset collection method. The sentiment image automatic generation pipeline takes the text prompt and sentiment value input by the user as input and generates sentiment images consistent with the input sentiment value.
[0033] The method for generating emotional image content based on text editing is as follows:
[0034] First, we collect and construct a dataset with text and sentiment value annotations, which includes several images and corresponding sentiment value annotations, as well as text annotations with and without sentiment texts.
[0035] Secondly, a sentiment embedding network model is constructed to embed sentiment values into text features;
[0036] Then, the image generation model is trained using the labeled data in the dataset and the pre-designed loss function until the optimal model parameters are obtained;
[0037] Finally, the text prompts and sentiment values are input into the trained image generation model to obtain the output image.
[0038] Furthermore, we generate emotional images from two inputs : A text hint describing the desired image content and specify sentiment values The emotional value can be a discrete emotion, such as Ekman's six basic emotions, or a continuous emotion, such as valence-arousal. The text prompt is used to describe the picture content of the emotional image.
[0039] Then, the text prompt or the text prompt encoded based on the text encoder is embedded into the emotion through the emotion embedding model, which can be output in the following ways:
[0040] (1) The model directly outputs emotional text. The output of the emotional text is directly used as the input of the text-to-image generation model G (such as StableDiffusion, SDXL, etc.) to generate emotional images.
[0041] (2) The model directly outputs the sentiment-embedded text features, which are used as the input of the image generation model G to generate sentiment images ( Figure 1 (b.1)).
[0042] (3) The model first outputs sentiment-embedded text features, and then decodes the features into sentiment-embedded text. The sentiment-embedded text features can be used as the input of the image generation model G, and the sentiment-embedded text can also be used as the input of the image generation model G.
[0043] In order to train the model, a loss function must be designed to constrain the output distribution of the model. In order to enhance the emotional expression of the generated images, several strategies can be used ( Figure 1 (b.2)) By considering the distribution of sentiment values of the training data and the residual between the text features of the enhanced sentiment embedding and the text features without sentiment, the network training is optimized to ensure that the generated images convey the desired sentiment and content. To support the sentiment embedding network, an image dataset ( Figure 1 (a)). This dataset pairs emotional images with non-emotional and emotional texts and contains corresponding emotion value labels.
[0044] Dataset collection process
[0045] Before building the EmotiCrafter system, you first need to collect and build a dataset with emotion annotations. This dataset contains a large number of images and corresponding emotion values, which are labeled by manual annotators based on emotion models in psychology. The collection of the dataset involves obtaining images from multiple sources, such as public emotional image databases, authorized image libraries, etc., to ensure the diversity of images and the accuracy of emotion annotations. In order to ensure the quality and diversity of the dataset, the collection process requires strict quality control and review steps. At the same time, all images need to be annotated with and without emotional text. The annotation process can be done manually or using a large language model. Thus, the acquired dataset is recorded as , Contains pairs , respectively, are sentiment value, text without sentiment, and text with sentiment.
[0046] Model structure design
[0047] Emotion Embedding Network for the EmotiCrafter System , which is responsible for converting the input sentiment value Embedded into emotion-free text features middle, , The text encoder is the text encoder, while the image generation model generates images based on the sentiment-embedded text features. When designing the model structure, it is necessary to consider how to effectively integrate sentiment information and ensure that the generated images can accurately reflect these sentiments.
[0048] Sentiment values are first encoded into sentiment features.
[0049] For discrete emotion values, a coding value can be assigned to each discrete emotion. For example, for Ekman's six basic emotions, each emotion uses a separate feature parameter express.
[0050] For continuous emotions, you can use MLP, linear layer and other structures to encode them as features. For example, for valence-arousal values, they can be encoded as The emotional features are uniformly recorded as .
[0051] There are several feasible examples of integrating sentiment features into text features, but methods other than the following examples can support this process:
[0052] (1) Method of fusing multimodal information based on cross-attention mechanism ( Figure 2 ). Based on the transformation block of the traditional transformer network, a cross attention mechanism is added to fuse the sentiment features with the text features.
[0053] (2) Feature fusion based on addition operation. Add the sentiment feature directly to the non-sentiment text feature. , as the input of the neural network.
[0054] (3) Feature fusion based on splicing. Splice the sentiment features and the non-sentimental text features. ] and then input into the neural network.
[0055] For the structural design of the backbone network, you can choose a transformer, such as the design structure of GPT, LLAMA, etc. For example, a method based on cross-attention mechanism to embed continuous emotional features The network structure design method is as follows: First, the input non-emotional text features Projection into the transformer’s feature space:
[0056]
[0057] in, represents the initial hidden state, (·) is the projection layer, is the position embedding. The core of the Emotion Injection Transformer (EIT) consists of It consists of a series of emotion injection blocks (EIBs), each of which outputs a hidden state:
[0058]
[0059] in It is indivual The output of a design such as Figure 2 As shown, each EIB enhances the traditional transformer block through a criss-cross attention mechanism:
[0060]
[0061]
[0062]
[0063]
[0064] in, is an intermediate hidden variable; is the layer normalization used for normalization; represents self-attention, which is used to capture contextual dependencies; It is used to inject cross attention of valence features and arousal features. is a feed-forward network that introduces nonlinear transformations to adapt to the complexity of the sentiment embedding process. The output of the last emotion injection block pass (a linear layer and a layer normalization layer) projected back to the text feature space of SDXL as the text features of the sentiment embedding output by the model:
[0065]
[0066] In order to optimize the results of the model, the final output, that is, the final sentiment embedded text features, can be obtained through residual connection. This can be obtained by adding this model output to the original sentiment-free text features:
[0067]
[0068] Model training process
[0069] The model training process involves training the EmotiCrafter system using the collected dataset. During the training process, the model learns how to generate emotionally consistent images based on textual prompts and sentiment values. In order to optimize the output distribution of the model, the following loss functions can be selected but are not limited to: :
[0070] (1)
[0071] (2)
[0072] (3)
[0073] in, For expectations, is the number of characteristic elements, and is the optimal parameter obtained manually or experimentally. The target is emotional text features.
[0074] At the same time, this solution uses two strategies to enable the model to achieve the expected training effect, including: scaling residual learning and sentiment value density weighted learning. These two training strategies will be integrated into the loss function middle.
[0075] Scaling residual learning,To better capture the prominent emotional changes in the generated images,,the target residual is scaled:
[0076]
[0077] in, is the residual feature, It has emotional text features. is a scaling factor or learnable factor, which can be set to 1.5, for example.
[0078] Sentiment density weighting: In order to reduce the impact of imbalanced distribution of training samples, the loss is weighted. For example, according to the density of sentiment values of training samples The inverse ratio of is weighted. At this time, a loss function is
[0079]
[0080] It can be estimated using the definition of the density distribution function, or using the kernel density estimation (KDE) with a Gaussian kernel, or any other method that can estimate the density distribution function.
[0081] For example, the method of using Gaussian kernel to estimate density distribution function is as follows:
[0082]
[0083] in The bandwidth is 2D Gaussian kernel of; is the number of training samples; It is The sentiment value of training samples. Bandwidth It can be chosen using Silverman's rule of thumb to provide the best smoothing of the density estimate.
[0084] The training strategy is to obtain the optimal model parameters :
[0085]
[0086] A specific embodiment based on continuous emotion, valence-arousal values, such as Figures 1 to 4 As shown,
[0087] In order to further illustrate the benefits of this scheme compared with other image generation methods that can control image content, quantitative experiments based on valence-arousal value examples and user experiments are conducted to illustrate the benefits of this scheme, as follows:
[0088] 1. Quantitative Experiments: Four baselines are constructed based on the existing technologies, which are constructed through two strategies:
[0089] (1) Injecting emotional features directly into the image generation module in SDXL, such as UNet;
[0090] (2) Use sentiment features to change the input text prompt to control the content of the generated image (similar to the proposed method). All baselines include:
[0091] Cross Attention: Inject emotional features into SDXL’s UNet through a cross attention mechanism based on IP-Adapter;
[0092] Time Embedding: directly add sentiment features to the time embedding of UNet;
[0093] Textual Inversion: Use textual inversion technology to embed emotional features into prompt templates with predefined emotional placeholders;
[0094] GPT-4+SDXL (GPT-SD): Use GPT-4 to rewrite the input text according to the valence-arousal value as the SDXL prompt word for image generation.
[0095] Our approach is compared with several baselines using the following metrics:
[0096] (1) V / A-Error evaluates the absolute error between the predicted valence / arousal value of the generated image and the input valence / arousal value.
[0097] (2) CLIPScore evaluates the similarity between input text and generated images.
[0098] (3) CLIP-IQA uses the pre-trained CLIP model to evaluate image quality.
[0099] The method of this scheme achieves the best (lowest) V / A-Error on average. Cross Attention and TimeEmbedding obtain the best (highest) results on CLIP-IQA and CLIPScore respectively. The comparison results are shown in Table 1.
[0100] Table 1
[0101]
[0102] 2. User experiment: This example conducted two user studies on 20 college students to evaluate the effectiveness of the method and compare it with GPT-4+SDXL (baseline). The comparison results are shown in Table 2.
[0103] Table 2
[0104]
[0105] Experiment I. The first experiment tested whether the emotions expressed by the generated images were consistent with human perception. To this end, two sets of images were prepared to test the perception of arousal value and valence value. Each set consisted of 20 groups of images, 10 of which were generated by the method of this scheme and the other 10 were generated by the baseline. Each set contained 5 images with different arousal or valence values, and the order was randomly shuffled. Participants were asked to reorder the 5 images according to the perceived valence or arousal value. Kendall's To assess the consistency between the order given by the participants and the true order, indicates an exact match. During the experiment, participants were also asked to estimate and label the valence and arousal values of each image using perceived emotion. The absolute error between the result and the true value was calculated.
[0106] Experiment II. The second experiment aims to test whether the generated images can reveal controversial changes in emotions (i.e., valence-arousal values). To this end, 20 sets of images were generated (10 sets based on the proposed method and the other 10 sets based on the baseline). Each set contains 25 images generated according to the gradual change of valence-arousal values from -3 to +3. Participants were asked to rate each set of images (using a 5-point Likert scale) on two aspects: (1) the consistency between the change of valence-arousal value and the change of image content; (2) the smoothness of the change of image content.
[0107] The Shapiro-Wilk test was performed to assess normality and the Wilcoxon signed rank test was used to assess the significance of all results. Figure 3 ) show that our approach outperforms the baseline in all aspects and is statistically significant in valence / arousal ranking, consistency, and smoothness.
[0108] Based on the above results, the following conclusions can be drawn:
[0109] EmotiCrafter has demonstrated significant technical advantages in the field of emotional image generation. Compared with existing technologies, EmotiCrafter can more accurately capture and express continuous emotional changes and generate images consistent with the emotional values input by users. At the same time, EmotiCrafter can enable users to control the content, thereby customizing personalized emotional image content. Under the condition of controlling both emotions and content, EmotiCrafter can achieve the generation of high-quality image content.
[0110] The actual use process of the EmotiCrafter system is as follows:
[0111] Input: Users input text prompts and sentiment values through input devices (such as keyboards and mice). These inputs are submitted to the model through the system interface.
[0112] Sentiment Embedding: After the system receives the input, the sentiment embedding network starts working to embed the sentiment value into the text features and prepare the sentiment features for image generation.
[0113] Image Generation: Text prompts embedded with sentiment features are fed into an image generation model, which generates images based on these features.
[0114] Output: The generated image is presented to the user via a display device (such as a monitor), and the user can download and evaluate the image.
[0115] In actual implementation, this process may require the use of corresponding hardware facilities. For example, the input process may require an input device (such as a keyboard, mouse), while displaying the generated image may require a display device (such as a monitor). In order to run large models, a computer with sufficient computing power (such as a computer equipped with a GPU) may also be required. These hardware facilities ensure the efficient operation of the system and a good user experience.
[0116] The emotional image content generation system based on text editing includes a user input module, an emotional image generation module, and an output display module, wherein the emotional image generation module is the emotional image generation model based on text editing, and the user input module uses an input device to input text prompts and emotional values; after receiving the input information, the emotional image generation module embeds the emotional value into the encoded emotionless text features, and generates an image based on the text features embedded with emotions; the output display module displays the generated image to the user through a display device.
[0117] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, call all or part of the steps of the method.
[0118] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0119] It should be understood that the present solution is not limited to the above-mentioned specific implementation methods, and the devices and structures not described in detail should be understood to be implemented in a common manner in the art; any technician familiar with the art can use the above-disclosed methods and technical contents to make many possible changes and modifications to the technical solution of the present solution without departing from the scope of the technical solution of the present solution, or modify it into an equivalent embodiment with equivalent changes, which does not affect the essential content of the present solution. Therefore, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present solution without departing from the content of the technical solution of the present solution still falls within the scope of protection of the technical solution of the present solution.
Claims
1. A method for generating emotional image content based on text editing, characterized in that: First, we collect and construct a dataset with text and sentiment value annotations, which includes several images and corresponding sentiment value annotations, as well as text annotations with and without sentiment texts. Secondly, a sentiment embedding network model is constructed to embed sentiment values into text features; Then, the labeled data in the dataset and the pre-designed loss function are used to train the sentiment embedding network model to obtain the optimal model parameters; Finally, the text prompts and sentiment values are input into the trained sentiment embedding network model, and the output of the sentiment embedding network model is passed through the image generation model to obtain the output sentiment image.
2. The method for generating emotional image content based on text editing according to claim 1, characterized in that: The sentiment embedding network model includes two inputs and one output, wherein one input is a sentiment value, the other input is a text without sentiment or a text feature without sentiment, and the output is a corresponding text with sentiment or a text feature embedded with sentiment.
3. The method for generating emotional image content based on text editing according to claim 2, characterized in that: The emotion embedding network model is trained by optimizing a loss function, and the loss function is used to constrain the output distribution of the emotion embedding network during the training process so that it fits the target distribution.
4. The method for generating emotional image content based on text editing according to claim 3 is characterized in that: The sentiment embedding network model first encodes the input sentiment value into sentiment features, and then embeds the sentiment features into text features, thereby embedding the sentiment value into the sentiment-embedded text features.
5. The method for generating emotional image content based on text editing according to claim 3 is characterized in that: The loss function includes a scaling residual learning strategy or a sentiment value density weighting strategy, wherein the scaling residual learning strategy linearly enlarges or reduces the residual between the emotion-free text feature and the emotion-containing text feature as the scaled target emotion text feature; the emotion value density weighting strategy weights the loss function by estimating the emotion value distribution density.
6. The emotional image generation model based on text editing is characterized by: The image generation model is connected to the back end of the emotion embedding network model described in any one of claims 1 to 5, and the emotion text or emotion text features output by the emotion embedding network model are converted into image form and output.
7. The emotional image generation model based on text editing according to claim 6 is characterized by: According to the output of the image generation model or the output of the emotion embedding network, the emotion embedding network model is cyclically optimized and trained to obtain the optimal network parameters.
8. The emotional image content generation system based on text editing is characterized by: It includes a user input module, an emotion image generation module, and an output display module, wherein the emotion image generation module is the emotion image generation model based on text editing as described in claim 6, and the user input module uses an input device to input text prompts and emotion values; after receiving the input information, the emotion image generation module embeds the emotion value into the encoded emotion-free text features, and generates an image based on the emotion-embedded text features; the output display module displays the generated image to the user through a display device.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, all or part of the steps of the method described in any one of claims 1 to 5 are called.
Citation Information
Patent Citations
Text generation method and device based on multiple modes and model training method and device based on multiple modes
CN114298121A
Image filter generation method based on sentiment analysis
CN116910294A
Geospatial point data sampling method driven by emotional feature consistency
CN117312468A
Image joint text sentiment analysis method based on modal fusion graph convolutional network
CN117540023A
Transforming Audio Content into Images
US20200126584A1