Short text generation image model training method, system, short text to image generation method, electronic device and storage medium

By extracting and enhancing the features of short text, using denoising and object detection technology, an adaptive loss function is constructed, and the short text generation image model is optimized, which solves the problem of insufficient semantic understanding of short text to image generation in the prior art, and achieves a better match between generated images and daily common sense and the satisfaction of user expectations.

CN119181102BActive Publication Date: 2025-05-06EAST CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411667050.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-05-06
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

The existing short text-to-image generation methods have limited ability to understand strong semantic information from short texts, which leads to the generated images not matching daily common sense and difficult to meet user expectations.

Method used

By extracting the subject characteristics and common sense characteristics of short text, weight allocation and feature standardization are performed to obtain short text features enhanced by common sense. Then, using the verification noise data and denoising rules, the denoising characteristics are obtained, and input them into the preset object detection model to obtain the subject characteristics. Finally, the denoising loss function and the subject generation loss function are constructed to form an adaptive loss function, and the short text generation image model is optimized.

Benefits of technology

The model's ability to understand short text semantics is improved, the generated images are more in line with daily common sense, the relevance and quality of the images are enhanced, and the user's expectations can be better met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181102B_ABST
    Figure CN119181102B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of artificial intelligence, and discloses a training method, system, a method for generating an image from a short text, an electronic device, and a storage medium. The training method for generating an image from a short text includes: obtaining a short text training sample and extracting subject features from it, and obtaining common sense features through a large language model and filtering rules; performing weight distribution and feature standardization to obtain common sense-enhanced short text features; obtaining denoising features by verifying noise data and denoising rules; inputting a preset target detection model to obtain subject features, and constructing a denoising loss function and a subject generation loss function; constructing an adaptive loss function based on these two loss functions to optimize the short text generation image model, and finally obtaining a target model. This method can solve the problem that the existing short text to image generation method has limited ability to understand strong semantic information from short texts, resulting in the generated image not being consistent with common sense.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and computer vision technology, and in particular to a short text-to-image model training method and system, a short text-to-image generation method, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of intelligent computing and deep learning, AI models for text-generated images have made remarkable progress. In recent years, the demand for short text-generated images has also been growing, such as the generation of remote sensing images, the generation of multimedia content, and the synthesis of e-commerce product images. Many people, especially non-professionals and some ordinary users who usually lack professional design backgrounds, often rely on short text descriptions to express their visual needs, and their descriptions are often concise and incomplete, using vague concepts such as "two little white rabbits" or "there is a car in the city" to convey their needs.

[0003] In this case, short text to image generation can not only help users transform these initial ideas into concrete visual content, but also significantly improve creative efficiency, especially in scenarios where needs need to be expressed in simple text. Therefore, solving the task of generating images from short text is a necessary condition for image processing and intelligent computing.

[0004] However, the limitations of short text expressions have brought certain challenges to the task of generating short texts from images. Although AI models such as Stable Diffusion have made progress in image generation tasks, they have limited ability to understand strong semantic information from short texts compared to large language models (LLM) and humans. At the same time, it is generally believed that short texts often lack sufficient semantic information to describe complex image requirements in detail. For example, a user may provide a short description such as "cake on the table" without specifying key details such as whether the table has more than two legs, the color of the cake, or the overall atmosphere, resulting in the generated image having uncomfortable places that conflict with daily common sense and the user's wishes. For another example, the generated bear legs are strange, the facial expressions are abnormal, the arrangement of the aircraft is unreasonable, and the number is wrong. This not only affects the practical applicability of the image, but also may undermine the user's trust in automatic generation technology.

[0005] Therefore, existing short text-to-image generation methods have limited ability to understand strong semantic information from short texts, resulting in the generated images being inconsistent with daily common sense and difficult to meet user expectations. Summary of the invention

[0006] Based on this, the present application proposes a short text to image generation model training method, system, short text to image generation method, electronic device and storage medium, aiming to solve the problem that the existing short text to image generation method has limited ability to understand strong semantic information from short text, resulting in the generated images being inconsistent with daily common sense and difficult to meet user expectations.

[0007] The first aspect of the present application provides a short text generation image model training method, the method comprising:

[0008] Get short text training samples;

[0009] Extracting short text main body features according to the short text training samples, and obtaining multiple effective common sense features from the short text training samples according to the large language model and filtering rules;

[0010] Performing weight assignment and feature standardization on the main features of the short text and the multiple effective common sense features to obtain common sense enhanced short text features;

[0011] Obtaining denoising features according to the verification noise data and the denoising rule, wherein the denoising rule is constructed by the common sense-enhanced short text feature and the plurality of effective common sense features with increased weights;

[0012] Inputting the denoising features into a preset target detection model to obtain subject features;

[0013] According to the denoising rule, construct a denoising loss function;

[0014] Constructing a subject generation loss function according to the common sense enhanced short text feature, the multiple effective common sense features and the subject feature;

[0015] Constructing an adaptive loss function according to the denoising loss function and the subject generation loss function;

[0016] The short text generation image model is optimized according to the adaptive loss function to obtain a target short text generation image model.

[0017] As an optional implementation of the first aspect, the step of obtaining a plurality of effective common sense features from the short text training sample according to the large language model and the filtering rules includes:

[0018] Inputting the short text training sample into a first large language model to generate a plurality of first common sense features;

[0019] Inputting the short text training sample and the plurality of first common sense features into a second language model to obtain a plurality of screened second common sense features;

[0020] The plurality of second common sense features are scored using filtering rules formulated by an artificial feedback mechanism to obtain a plurality of valid common sense features.

[0021] As an optional implementation of the first aspect, after the step of obtaining a plurality of effective common sense features from the short text training sample according to the large language model and the filtering rules, the step further includes:

[0022] Inputting the third language model according to the main features of the short text and the multiple effective common sense features to obtain multiple effective common sense features oriented to the main features of the short text;

[0023] The multiple valid common sense features oriented to the main features of the short text are input into a text encoder to obtain text information, and the text information is used to generate the multiple valid common sense features and the common sense enhanced short text features when performing a common sense enhancement operation.

[0024] As an optional implementation of the first aspect, the step of performing weight assignment and feature standardization on the short text main features and the multiple valid common sense features to obtain common sense enhanced short text features includes:

[0025] Using a cross attention mechanism, capturing the long-distance dependency between the main features of the short text and the multiple effective common sense features and dynamically assigning weights; and using a residual connection mechanism, retaining the main features of the short text and obtaining multiple initial common sense enhanced features;

[0026] The multiple initial common sense enhanced features are passed through a projection network for feature normalization to obtain the common sense enhanced short text features, wherein the projection network includes a linear layer and a layer normalization layer.

[0027] As an optional implementation manner of the first aspect, the step of extracting short text main body features according to the short text training sample includes:

[0028] The short text main body features are extracted from the short text training samples using a pre-trained Chinese named entity recognition model.

[0029] As an optional implementation manner of the first aspect, the denoising rule includes:

[0030] Adding dynamic weights to the multiple effective common sense features, wherein the dynamic weights increase dynamically as the time step increases, to obtain the multiple effective common sense features with increased weights;

[0031] A denoising condition function is constructed based on the common sense-enhanced short text feature and the plurality of effective common sense features with increased weights, and the denoising condition function is used to restore the denoising feature from the verification noise data.

[0032] The second aspect of the present application provides a short text generation image model training system, the system comprising:

[0033] A sample acquisition module is used to obtain short text training samples;

[0034] A common sense extraction module, comprising a large model extraction unit and a small model extraction unit, wherein the small model extraction unit is used to extract short text main features according to the short text training sample, and the large model extraction unit is used to obtain multiple effective common sense features from the short text training sample according to the large language model and filtering rules;

[0035] The model learning module includes a common sense enhancement unit and an image denoising unit, wherein the common sense enhancement unit is used to perform weight assignment and feature standardization on the short text main features and the multiple effective common sense features to obtain common sense enhanced short text features, and the image denoising unit is used to obtain denoising features according to verification noise data and denoising rules, wherein the denoising rules are constructed by the common sense enhanced short text features and the multiple effective common sense features with increased weights, and the denoising features are input into a preset target detection model to obtain main features;

[0036] A model training module is used to construct an adaptive loss function based on a denoising loss function and a subject generation loss function, and optimize the short text generation image model according to the adaptive loss function to obtain a target short text generation image model; the model training module includes a first loss function unit and a second loss function unit, the first loss function unit is used to construct a denoising loss function according to the denoising rule, and the second loss function unit is used to construct a subject generation loss function according to the common sense enhanced short text features, the multiple effective common sense features and the subject features.

[0037] A third aspect of the present application provides a method for generating a short text into an image, the method comprising:

[0038] Get the target short text;

[0039] Inputting the target short text into a target short text generation image model, wherein the target short text generation image model is obtained by training using the short text generation image model training method as described above;

[0040] An image model is generated based on the target short text to obtain a target image corresponding to the target short text.

[0041] A fourth aspect of the present application provides a system for generating a short text into an image, the system comprising:

[0042] A short text acquisition module is used to acquire target short text;

[0043] A text input module, used for inputting the target short text into a target short text generation image model, wherein the target short text generation image model is obtained by training the short text generation image model training method as described above;

[0044] An image generation module is used to generate an image model based on the target short text to obtain a target image corresponding to the target short text.

[0045] The fifth aspect of the present application provides an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement the short text generation image model training method as described above or the short text to image generation method as described above.

[0046] The sixth aspect of the present application provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the short text generation image model training method as described above or the short text to image generation method as described above.

[0047] Compared with the prior art, the beneficial effects of the present invention include: 1. By extracting the main features of short texts, the model can understand the core content and theme of the text; the common sense features extracted by large language models and filtering rules can provide the model with background knowledge and contextual information to enhance its understanding ability. 2. Weight allocation and standardization of short text main features and effective common sense features help balance the impact of different features on model training, so that the model can better combine text content and common sense information when generating images; 3. By verifying noise data and denoising rules, the model can identify and remove noise in training data and improve data quality; the construction of denoising rules relies on common sense-enhanced short text features and effective common sense features to ensure that the denoising process can effectively retain useful information; inputting denoising features into the preset target detection model can extract more accurate main features, which will directly affect the quality and relevance of the generated image; 4. The construction of denoising loss function ensures that the model can effectively reduce the impact of noise on the generation results during training; the main generation loss function ensures that the generated image can accurately reflect the main features of the short text; combining the denoising loss function and the main generation loss function to form an adaptive loss function, so that the model can dynamically adjust the loss weight during training and optimize the generation effect; the short text generation image model is optimized through the adaptive loss function, and the target short text generation image model finally obtained will have a stronger generation ability and can generate images that are highly relevant to the input short text and of high quality. Therefore, this solution can solve the problem that the existing short text to image generation method has limited ability to understand strong semantic information from short text, resulting in the generated images being inconsistent with daily common sense and difficult to meet user expectations.

[0048] Additional aspects and advantages of the present application will be given in part in the following description, and in part will become apparent from the following description, or will be understood through the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 An exemplary system architecture of an embodiment of the present application;

[0050] Figure 2 A flowchart of a short text generation image model training method proposed in the first embodiment of the present application;

[0051] Figure 3 A framework diagram of a short text generation image model proposed in the first embodiment of the present application;

[0052] Figure 4 This is a framework diagram of the common sense extraction module in the first embodiment of the present application;

[0053] Figure 5 This is a framework diagram corresponding to the filtering rules in the first embodiment of the present application;

[0054] Figure 6 A schematic diagram of the structure of a short text generation image model training system proposed in the second embodiment of the present application;

[0055] Figure 7 A flowchart of a method for generating a short text to an image proposed in the third embodiment of the present application;

[0056] Figure 8 This is a schematic diagram of the structure of a short text to image generation system proposed in the fourth embodiment of the present application.

[0057] The following specific implementation methods will further illustrate the present application in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0058] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0059] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.

[0060] In order to facilitate the understanding of the present application, the system architecture on which the present application is based is first described. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown. Figure 1 As shown in FIG. 1 , the system architecture may include a server 101, a network 102, a terminal device 103, a terminal device 104, and a terminal device 105. The network 102 is used to provide a medium for a communication link between the terminal device 103, the terminal device 104, or the terminal device 105 and the server 101. The network 102 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.

[0061] The server 101 may be a server that provides various services, such as a background management server that provides support for devices operated by users using the terminal device 103, the terminal device 104, or the terminal device 105. The background management server may analyze and process the received request and other data, and feed back the processing results to the terminal device 103, the terminal device 104, or the terminal device 105.

[0062] Terminal device 103, terminal device 104 and terminal device 105 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a wearable smart device, a virtual reality device, an augmented reality device, etc., but are not limited thereto.

[0063] It should be understood that Figure 1 The number of terminal devices 103, terminal devices 104, terminal devices 105, networks 102 and servers 101 are merely illustrative. Server 101 may be a physical server, a server cluster consisting of multiple servers, or a cloud server. Depending on actual needs, any number of terminal devices, networks and servers may be provided.

[0064] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.

[0065] First, the terms involved in the present invention are explained.

[0066] Short text training samples: refers to short text data used to train natural language processing models. These samples usually contain a small number of words or sentences and are used to train the model to recognize specific patterns, keywords or phrases in the text. Short text training samples can be social media posts, news articles, product reviews, or any other type of short text data.

[0067] Main features of short texts: refers to the main characteristics and attributes of short text content. When analyzing and understanding short texts, we extract and summarize the main features of short texts. These features can summarize the core information of the text and help us understand the main content and intention of the text more quickly. These main features usually include the theme, emotional tendency, key information points, language style, rhetoric, etc. of the text.

[0068] Large Language Model (LLM): is a natural language processing model based on deep learning, designed to understand and generate human language. LLM is trained using massive text data and can perform a wide range of tasks, including text generation, text classification, text summarization, machine translation, sentiment analysis, etc. These models usually contain billions or even tens of billions of parameters and are large in scale, so they are also called large language models. The characteristic of LLM is that it can learn the grammar and semantics of natural language to generate human-readable text. They are usually based on neural network models, such as Transformer, and use self-attention mechanisms to capture contextual information in the text. This makes LLM perform well in processing various natural language tasks, especially when dealing with long texts and complex language phenomena.

[0069] Effective Common Sense Features: When analyzing and processing short text data, in addition to the entity features of the text itself (such as vocabulary, syntax, semantics, etc.), some basic common sense information that is generally recognized and accepted by people can also be used to enhance the effect of text analysis. For example: the basic attributes and relationships of common things, habits and norms in daily life, common concepts and knowledge in social and cultural backgrounds, general laws of human behavior and psychology, basic concepts such as time, space, and quantity. These common sense features are combined and refined to form common sense knowledge features that are widely accepted and applied in daily life and can effectively guide people's behavior or decision-making.

[0070] Common sense enhanced short text features: Combines entity features and common sense features in the text. Entity features refer to specific nouns, verbs, etc. that appear in short texts, while common sense features are a summary and interpretation of the text background and context based on general knowledge and experience. Specifically, common sense enhanced short text features can be achieved by introducing external knowledge bases, common sense reasoning, and other methods. External knowledge bases can provide background knowledge and context information related to short texts, while common sense reasoning can infer and supplement based on clues and known common sense in short texts.

[0071] Validation noise data: refers to input data with random noise or uncertainty that is intentionally introduced during model training and validation. This data is usually similar to real data, but contains a certain degree of random perturbations or errors to simulate the uncertainty in the real environment. By inputting the validation noise data into the alignment loss function together with the original input data, the model can be prompted to learn how to maintain a correct understanding and representation of the data in the presence of noise. In this way, the model can make more accurate predictions when faced with noise and interference in real applications.

[0072] Object Detection Model: It is a computer vision technology that can automatically identify and locate multiple objects in an image or video. This model is usually composed of a neural network, especially a convolutional neural network (CNN), which can learn features in an image and determine the location and category of objects in the image based on these features. In the object detection task, the model needs to output the bounding box of each detected object and the corresponding category label. Common object detection models include R-CNN, Fast R-CNN, Faster R-CNN, YOLO (YouOnly Look Once) and SSD (Single Shot MultiBox Detector).

[0073] Subject Features: refers to the characteristics exhibited by the main objects or key elements in the image. These characteristics can be visual information such as the shape, texture, color, size, position, etc. of the object. For example, when we want to extract the main features of the image from a short text, we can use the object detection model to complete this task. First, we need to convert the short text content into image data, and then use the object detection model to recognize and process these image data. The object detection model analyzes each area of ​​the image and tries to identify the objects in it. Through this process, we can extract the main features in the image, such as the shape, size, position, etc. of the object.

[0074] Example 1

[0075] See also Figure 2 , is a flowchart of a short text generation image model training method proposed in the first embodiment of the present application. The proposed method includes the following steps S01 to S09, please refer to Figure 3 , shown is a framework diagram of the short text generation image model proposed in an embodiment of the present application.

[0076] S01: Obtain short text training samples.

[0077] S02: extracting short text main features according to the short text training samples, and obtaining a plurality of effective common sense features from the short text training samples according to a large language model and filtering rules.

[0078] In this embodiment, in order to automatically extract more common sense knowledge with deep semantic information from short texts, the present application proposes a large language model (such as LLM) and a small model (a lightweight model for a specific task) to extract multiple effective common sense features and short text subject features respectively. Since the large language model has a strong language understanding ability and a wide range of common sense knowledge base, it can extract rich semantics and implicit common sense information from short texts; inspired by the chain of thought and people in the loop, the present application uses filtering rules based on manual feedback to screen out extensive and effective common sense. The small model focuses on the entity recognition task in the short text, and combines its results with the effective common sense knowledge provided by the large language model to provide common sense knowledge embedding for the short text subject. Although LLM can also do the work of subject extraction, the cost of LLM is often much higher than that of the basic small model during use, so in order to save time and resources, we assign this task to a small model. Therefore, we send the short text to the above-mentioned large language model and small model to obtain effective common sense knowledge for the subject, and finally embed the effective common sense knowledge for the subject through the text encoder.

[0079] It should be noted that in this step, the processing of the large language model and the small model can be represented by the framework diagram of the common sense extraction module, as shown in Figure 4 As shown, it is a framework diagram of the common sense extraction module in this embodiment.

[0080] In the large language model part: input the short text training sample into the first large language model to generate multiple first common sense features; input the short text training sample and multiple first common sense features into the second large language model to obtain multiple screened second common sense features; use the filtering rules formulated by the artificial feedback mechanism to assign scores to the multiple second common sense features to obtain multiple effective common sense features.

[0081] Specifically, the short text training sample is first sent to the first LLM (we take the Tongyi Qianwen 2.0 model as an example), where the first prompt contains: the definition of common sense, an instruction to generate less constrained broad common sense knowledge, and an example that concretizes the instruction. In this way, rich common sense features about the short text training sample can be obtained. However, the generated common sense features are too broad, and not every common sense feature is effective when generating images. For example, when the short text is "two red apples", the common sense "apples are usually round" is effective when drawing images because it is a specific description of the subject in the text, while another common sense "red is the primary color" is useless when generating images because it is purely theoretical knowledge and is powerless to construct scenes. At the same time, although the LLM is very capable, due to the randomness of the LLM, the common sense generated at one time is not completely related to the short text. In this case, filtering is required to obtain truly usable common sense knowledge.

[0082] Furthermore, the short text training samples and the common sense features generated for the first time are sent to the second LLM. The second prompt includes: instructions for filtering negative common sense knowledge (i.e. unreasonable common sense), filtering rules, and several examples of specific elaboration of the rules. For the instructions, we designed them to give scores and reasons for each common sense feature according to the rules that define each score range. For the content of the rules, we first wrote them manually, but the results were not optimistic. Knowledge such as "red is one of the three primary colors" still got the same score as positive knowledge. To solve this problem, inspired by people in the loop and multi-agent collaboration, we created an artificial feedback mechanism such as Figure 5 As shown, it is a framework diagram corresponding to the filtering rules in this embodiment, and the filtering rules are adaptively generated in limited iterations through human feedback.

[0083] The process of one iteration is as follows: We first use the first LLM to filter the common sense, and then pass the result to humans to manually select negative common sense knowledge that cannot help image generation. After that, we pass the selected negative common sense to the second LLM (using GPT4 at this time) to specifically summarize the universality of these common sense and infer the reasons why the selected common sense is negative. Finally, we use these reasons to adjust and enrich the rules and start the next iteration until the number of manually selected common sense is small enough.

[0084] After several iterations, we got the final filtering rules. Table 1 shows some of the rules and the corresponding filtering results during the entire iteration process. The final filtering rules are shown in the nth iteration. It is a list of 4 json-schema type elements, each of which represents a score range, as well as its description, criteria, and specific examples. After the detailed description and examples, it is clear to give a score to each common sense and give the corresponding reasons. From Table 1, we can see that as the iterations go deeper, the filtering rules become more and more detailed, and the remaining common sense gradually eliminates those abstract knowledge and theoretical knowledge that are not descriptive or relevant to the short text. Finally, we use the final filtering rules in the second prompt of the common sense filtering stage to successfully extract effective common sense features.

[0085] Table 1 Some iteration results of filtering rules

[0086]

[0087] In the small model part: using a pre-trained Chinese named entity recognition model, the short text main features are extracted from the short text training samples.

[0088] Specifically, in terms of entity recognition models, classic Chinese named entity recognition models (NER) usually extract entities with specific names, such as "Francisco" and "cupcake". However, in short texts, the subjects are often general simple words, such as "city", "cake", etc., which are not suitable for NER models to play a direct role. Therefore, we chose a pre-trained Chinese named entity recognition model (BERT-based NER model). The model has been trained on a token classification dataset and has achieved an accuracy of more than 96% in extracting subjects from short texts and constructing entity sets. For example, when the given short text is "two apples in the forest", the entity set contains "apple" and "forest".

[0089] Finally, in the text encoder part: the main features of the short text and multiple effective common sense features are input into the third language model to obtain multiple effective common sense features for the main features of the short text;

[0090] A plurality of effective common sense features oriented to the main features of the short text are input into a text encoder to obtain text information, and the text information is used to generate a plurality of effective common sense features and common sense enhanced short text features when performing a common sense enhancement operation.

[0091] Specifically, we send the short text subject features and effective common sense features to the third LLM, map each effective common sense feature to a corresponding subject, and realize the generation of subject-oriented common sense features. This kind of common sense feature pays more attention to the subject, and the subject is often the key point of the short text. Therefore, multiple effective common sense features for the short text subject features can better help generate images aligned with the short text. Finally, multiple effective common sense features for the short text subject features are feature encoded through the text encoder to obtain the text information required for the next step.

[0092] S03: performing weight assignment and feature standardization on the main features of the short text and the multiple effective common sense features to obtain common sense enhanced short text features.

[0093] S04: obtaining denoising features according to the verified noise data and denoising rules, wherein the denoising rules are constructed by the common sense-enhanced short text features and the plurality of effective common sense features with increased weights.

[0094] It should be noted that in this step, Figure 2 As shown, the common sense knowledge enhanced short text to image generation link includes two steps: the common sense enhancement process (corresponding to S03 above) and the denoising process of the diffusion model (corresponding to S04 above).

[0095] In the common sense enhancement process: a cross-attention mechanism is used to capture the long-distance dependencies between the main features of the short text and multiple effective common sense features and dynamically assign weights; and a residual connection mechanism is used to retain the main features of the short text and obtain multiple initial common sense enhanced features; multiple initial common sense enhanced features are passed through a projection network for feature normalization to obtain common sense enhanced short text features, and the projection network includes a linear layer and a layer normalization layer.

[0096] Specifically, since the obtained text information is just the concatenation of two types of text embeddings, without mining the close relationship and deep semantic meaning between them, in this state, common sense features cannot help the image generation model understand the deep meaning in the short text. At the same time, a previous study found that in the denoising step of image generation, the model cannot effectively learn the given information accurately and comprehensively when interacting with information from different modules. In this case, we developed a common sense enhancement process to establish a dynamic association between the main features of short texts and common sense features. Inspired by previous work, we first use the cross-attention mechanism to use common sense feature embedding to guide the generation of the main features of short texts, and keep the main features of short texts consistent with the common sense features by capturing the long-distance dependencies between different input sources and dynamically assigning different weights to different input blocks. At the same time, in order to prevent the gradient from vanishing, make it easier for information to flow in the model and maintain the key information of short texts, we use the residual connection mechanism to enhance it with common sense feature embedding while retaining the original features of short texts to obtain the initial common sense enhanced features. The process can be expressed as:

[0097] ,

[0098] ,

[0099] in, represents the initial common sense enhanced features, represents the cross attention mechanism, Q represents the main features of the short text , K and V represent common sense features .

[0100] Next, we will A projection network is included for feature normalization, which consists of several linear layers and layer normalization operations. This improves the expressiveness of the model by introducing more learning parameters, helps enrich the model's feature representation capabilities, and enables the model to handle complex inputs more flexibly. At the same time, due to and From different sources, projection can ensure the effective transfer of information between different dimensions, thereby improving training efficiency and effectively decomposing global text embedding Short text features as final common sense enhancement.

[0101] After completing the common sense enhancement process, we enter the denoising process of the diffusion model. In this process, image generation includes a multi-step denoising iterative process according to the denoising rules, gradually restoring the verification noise data (which is random noise) to a clear image. The denoising rules include: adding dynamic weights to the multiple valid common sense features, the dynamic weights dynamically increase with the increase of the time step, and obtaining the multiple valid common sense features with increased weights; constructing a denoising condition function based on the common sense enhanced short text features and the multiple valid common sense features with increased weights, and the denoising condition function is used to restore the denoising features from the verification noise data.

[0102] Specifically, the denoising mechanism formed by this progressive denoising rule allows the model to make corrections and optimizations at each time step, avoiding inaccuracies or distortions that may occur when generating the entire image at once. For the early time steps, the denoising mechanism allows the model to generate a rough outline of the image. As the time steps increase, the denoising mechanism allows the model to gradually add details to the image. In this case, Figure 2 As shown in the figure, we incorporate common sense enhanced short text features into the denoising process as part of the denoising conditional function y, and explicitly incorporates the common sense features embedding as part of the conditional y.

[0103] In addition, due to is only auxiliary information, and contains the key picture information that controls the generated image, so we Add dynamic weights. In the early time steps, Compare More necessary because it controls the basic outline of the image, however, It becomes increasingly important in further time steps, as commonsense knowledge constrains the details of the generated images, aligning them with human consensus and preventing uncomfortable regions. Therefore, we first assign A smaller weight , and increases dynamically with the increase of time step t. Then, the condition y in the denoising process can be defined as:

[0104] .

[0105] S05: Input the denoising features into a preset target detection model to obtain subject features.

[0106] S06: Construct a denoising loss function according to the denoising rule.

[0107] It should be noted that Stable diffusion is a generative method process that uses the concept of diffusion to synthesize high-quality images from text prompts, using the denoising diffusion probability model (DDPM) framework to achieve efficient and scalable image generation with enhanced control (such as text) during the generation process. Stable diffusion consists of two processes: the forward process and the reverse denoising process.

[0108] like Figure 2 As shown, the forward process starts from the original image (with validation noise data ), the image is gradually corrupted by adding Gaussian noise at each time step until the image becomes pure noise. The process can be described as:

[0109] ,

[0110] in, represents the denoising feature at time step t, represents the variance of the noise distribution, I represents noise, and N represents Gaussian distribution.

[0111] In the reverse denoising process learning, the model gradually learns to remove noise from the noisy image, ideally reversing the forward process. The reverse process is performed by the neural network model Parameterized, the network predicts the noise corresponding to each time step based on the condition y generated above, allowing us to iteratively denoise the image. The inverse process can be defined as:

[0112] ,

[0113] in, and Represent the model parameters The mean and variance of the predicted noise distribution, y represents the conditions that control the denoising process, and in the short text generation image task, it represents the text embedding. The ultimate goal of the diffusion model is to optimize the model parameters , which reduces the difference between the predicted noise and the added noise, can be defined as:

[0114] ,

[0115] in, represents the average loss across the entire training data distribution, and represent the noise distribution of the added noise and the noise predicted by the model, respectively.

[0116] S07: Constructing a subject generation loss function according to the common sense enhanced short text features, the multiple effective common sense features and the subject features.

[0117] It should be noted that the above denoising diffusion probability model is based on the reconstruction loss of the difference between the measured predicted noise and the added noise reflects the denoising ability of the model, but it does not explicitly focus on the generation quality of the subject in short texts. However, due to the limited contextual information of short texts, it is easy to ignore the subject words when generating images. In this case, we create a subject generation loss function, which calculates and the main features in the generated image The cosine similarity between . Figure 2 As shown, we use a lightweight and well-performing preset target detection model to extract the subject in the generated image, and the filtering threshold is less than 0.5 detection results, subject features The image features of the last decoded hidden state of the target area in the detection result. The main body generates the loss It can be expressed as:

[0118] ,

[0119] ,

[0120] in, is a sigmoid activation function that maps the loss to the range of 0 to 1. represents cosine similarity, A represents , B represents .

[0121] It is uncertain which loss function is more important, so we create an adaptive loss function L that dynamically balances different loss terms according to the training stage, thereby adopting different optimization strategies. As a result, the model can more flexibly adapt to complex tasks and data distributions, while improving the stability and efficiency of the training process.

[0122] S08: Constructing an adaptive loss function according to the denoising loss function and the subject generation loss function.

[0123] It should be noted that considering that it is uncertain which loss function is more important, we create an adaptive loss function L, which dynamically balances different loss terms according to the training stage, thereby adopting different optimization strategies. Therefore, the model can adapt to complex tasks and data distributions more flexibly, while improving the stability and efficiency of the training process. The final adaptive loss function can be expressed as:

[0124] ,

[0125] in, and Represents training parameters.

[0126] S09: Optimizing the short text generation image model according to the adaptive loss function to obtain a target short text generation image model.

[0127] In summary, the model training method provided by the present application first obtains short text training samples to provide a data basis for subsequent model training; extracts short text main features and obtains multiple effective common sense features from short text training samples. These features can help the model better understand the semantics and background knowledge of the text; weights are assigned and features are standardized for the short text main features and common sense features to obtain common sense enhanced short text features. This enables common sense features to play a more important role in the model and improve the generation effect of the model; denoising features are obtained based on the verification noise data and denoising rules. The denoising rules are constructed by common sense enhanced short text features and weighted common sense features, which helps to improve the robustness and denoising ability of the model to noise data; the denoising features are input into the preset target detection model to obtain the main features. The target detection model can help the model better capture the main information of the image and provide guidance for generating images; a denoising loss function is constructed according to the denoising rules, and a main generation loss function is constructed according to the common sense enhanced short text features, common sense features and main features. These two loss functions can respectively optimize the denoising ability and main generation ability of the model to improve the performance of the model; an adaptive loss function is constructed based on the denoising loss function and the main generation loss function. The adaptive loss function can be adjusted according to different tasks and data characteristics to improve the generalization ability and adaptability of the model; finally, the short text generation image model is optimized according to the adaptive loss function to obtain the target short text generation image model. Through continuous optimization and adjustment, the model can learn how to generate high-quality images based on short texts, thereby achieving the task of generating short texts from images.

[0128] Example 2

[0129] See also Figure 6 , which is a schematic diagram of the structure of a short text generation image model training system proposed in the second embodiment of the present application, and the system includes:

[0130] The sample acquisition module 10 is used to acquire short text training samples;

[0131] The common sense extraction module 20 includes a small model extraction unit 21 and a large model extraction unit 22, wherein the small model extraction unit 21 is used to extract the main features of the short text according to the short text training sample, and the large model extraction unit 22 is used to obtain multiple effective common sense features from the short text training sample according to the large language model and filtering rules;

[0132] The model learning module 30 includes a common sense enhancement unit 31 and an image denoising unit 32. The common sense enhancement unit 31 is used to perform weight assignment and feature standardization on the short text main features and the multiple effective common sense features to obtain common sense enhanced short text features. The image denoising unit 32 is used to obtain denoising features according to the verification noise data and denoising rules. The denoising rules are constructed by the common sense enhanced short text features and the multiple effective common sense features with increased weights. The denoising features are input into a preset target detection model to obtain main features.

[0133] The model training module 40 is used to construct an adaptive loss function according to the denoising loss function and the subject generation loss function, and optimize the short text generation image model according to the adaptive loss function to obtain a target short text generation image model; the model training module 40 includes a first loss function unit 41 and a second loss function unit 42, the first loss function unit 41 is used to construct a denoising loss function according to the denoising rule, and the second loss function unit 42 is used to construct a subject generation loss function according to the common sense enhanced short text features, the multiple effective common sense features and the subject features.

[0134] Example 3

[0135] See also Figure 7 , which is a flow chart of a method for generating a short text to an image according to a third embodiment of the present application, the method comprising:

[0136] S101: Obtain target short text;

[0137] S102: inputting a target short text into a target short text to generate an image model;

[0138] The target short text generation image model is obtained by training the short text generation image model training method in the first embodiment;

[0139] S103: Generate an image model based on the target short text to obtain a target image corresponding to the target short text.

[0140] Example 4

[0141] See also Figure 8 , which is a schematic diagram of the structure of a short text to image generation system proposed in the fourth embodiment of the present application, the system comprises:

[0142] A short text acquisition module 100 is used to acquire a target short text;

[0143] A text input module 200, used for inputting the target short text into a target short text generation image model, wherein the target short text generation image model is obtained by training the short text generation image model training method as described above;

[0144] The image generation module 300 is used to generate an image model based on the target short text to obtain a target image corresponding to the target short text.

[0145] The present application also proposes an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement the short text generation image model training method as described above or the short text to image generation method as described above.

[0146] The present application also proposes a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the short text generation image model training method as described above or the short text to image generation method as described above.

[0147] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0148] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0149] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A short text generation image model training method, characterized in that: The method comprises: Get short text training samples; Extracting short text main body features according to the short text training samples, and obtaining multiple effective common sense features from the short text training samples according to the large language model and filtering rules; Performing weight assignment and feature standardization on the main features of the short text and the multiple effective common sense features to obtain common sense enhanced short text features; Obtaining denoising features according to the verification noise data and the denoising rule, wherein the denoising rule is constructed by the common sense-enhanced short text feature and the plurality of effective common sense features with increased weights; Inputting the denoising features into a preset target detection model to obtain subject features; According to the denoising rule, construct a denoising loss function; Constructing a subject generation loss function according to the common sense enhanced short text feature, the multiple effective common sense features and the subject feature; Constructing an adaptive loss function according to the denoising loss function and the subject generation loss function; The short text generation image model is optimized according to the adaptive loss function to obtain a target short text generation image model.

2. The short text generation image model training method according to claim 1 is characterized in that: The step of obtaining a plurality of effective common sense features from the short text training sample according to the large language model and filtering rules comprises: Inputting the short text training sample into a first large language model to generate a plurality of first common sense features; Inputting the short text training sample and the plurality of first common sense features into a second language model to obtain a plurality of screened second common sense features; The plurality of second common sense features are scored using filtering rules formulated by an artificial feedback mechanism to obtain a plurality of valid common sense features.

3. The short text generation image model training method according to claim 1 is characterized in that: After the step of obtaining a plurality of effective common sense features from the short text training sample according to the large language model and the filtering rules, the step further includes: Inputting the third language model according to the main features of the short text and the multiple effective common sense features to obtain multiple effective common sense features oriented to the main features of the short text; The multiple valid common sense features oriented to the main features of the short text are input into a text encoder to obtain text information, and the text information is used to generate the multiple valid common sense features and the common sense enhanced short text features when performing a common sense enhancement operation.

4. The short text generation image model training method according to claim 1 is characterized in that: The steps of weighting and standardizing the main features of the short text and the multiple effective common sense features to obtain the common sense enhanced short text features include: Using a cross attention mechanism, capturing the long-distance dependency between the main features of the short text and the multiple effective common sense features and dynamically assigning weights; and using a residual connection mechanism, retaining the main features of the short text and obtaining multiple initial common sense enhanced features; The multiple initial common sense enhanced features are passed through a projection network for feature normalization to obtain the common sense enhanced short text features, wherein the projection network includes a linear layer and a layer normalization layer.

5. The short text generation image model training method according to claim 1 is characterized in that: The step of extracting short text main body features according to the short text training sample comprises: The short text main body features are extracted from the short text training samples using a pre-trained Chinese named entity recognition model.

6. The short text generation image model training method according to claim 1 is characterized in that: The denoising rules include: Adding dynamic weights to the multiple effective common sense features, wherein the dynamic weights increase dynamically as the time step increases, to obtain the multiple effective common sense features with increased weights; A denoising condition function is constructed based on the common sense-enhanced short text feature and the plurality of effective common sense features with increased weights, and the denoising condition function is used to restore the denoising feature from the verification noise data.

7. A short text generation image model training system, characterized in that: The system comprises: A sample acquisition module is used to obtain short text training samples; A common sense extraction module, comprising a large model extraction unit and a small model extraction unit, wherein the small model extraction unit is used to extract short text main features according to the short text training sample, and the large model extraction unit is used to obtain multiple effective common sense features from the short text training sample according to the large language model and filtering rules; The model learning module includes a common sense enhancement unit and an image denoising unit, wherein the common sense enhancement unit is used to perform weight assignment and feature standardization on the short text main features and the multiple effective common sense features to obtain common sense enhanced short text features, and the image denoising unit is used to obtain denoising features according to verification noise data and denoising rules, wherein the denoising rules are constructed by the common sense enhanced short text features and the multiple effective common sense features with increased weights, and the denoising features are input into a preset target detection model to obtain main features; A model training module is used to construct an adaptive loss function based on a denoising loss function and a subject generation loss function, and optimize the short text generation image model according to the adaptive loss function to obtain a target short text generation image model; the model training module includes a first loss function unit and a second loss function unit, the first loss function unit is used to construct a denoising loss function according to the denoising rule, and the second loss function unit is used to construct a subject generation loss function according to the common sense enhanced short text features, the multiple effective common sense features and the subject features.

8. A method for generating short text into an image, characterized in that: The method comprises: Get the target short text; Inputting the target short text into a target short text generation image model, wherein the target short text generation image model is obtained by training the short text generation image model training method according to any one of claims 1 to 6; An image model is generated based on the target short text to obtain a target image corresponding to the target short text.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the executable instructions to implement the short text generation image model training method as described in any one of claims 1 to 6 or the short text to image generation method as described in claim 8.

10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the short text generation image model training method as described in any one of claims 1 to 6 or the short text to image generation method as described in claim 8.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and storage medium

    CN117173497A

  • Training method of text generation graph model and text-based image generation method

    CN118115613A