Image-text data set construction method, system and device, storage medium and product
By constructing a knowledge base for prospective feature and background feature knowledge base, high-quality text-based graphic model prompt text is solved, and the problem of insufficient text quality in the existing technology is not high, and high-quality graphic data sets are generated.
Patent Information
- Application Number
- CN202510526455.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
When generating images, existing text-generated graphics models are highly dependent on the quality of the input text. If the text lacks accuracy, logic or detailed description, images that deviate from expectations may be generated, resulting in low quality of the graphics and text data set.
By extracting the entities, entity attributes and inter-entities in the original text corresponding to the original image, a knowledge base for the foreground feature is constructed; at the same time, a background description sample is extracted to construct a background feature knowledge base; then a target prompt text is generated based on these two knowledge bases, which is used to generate the corresponding target image and form a picture-text pair.
The quality of text prompts for text of the literary graphic model is improved, ensuring that the generated graphic data set has higher quality and accuracy.
Smart Images

Figure CN120069093A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of text processing, and particularly to a method, system, device, storage medium and product for constructing a text-image dataset. Background Art
[0002] A text-to-image model is a generative artificial intelligence model that can automatically generate an image semantically associated with the text description provided by the user. Such models typically adopt diffusion models, generative adversarial networks or autoregressive architectures, and are trained with a large amount of text-image pairs of data to learn the mapping relationship between text and visual features. For example, the MidJourney model is good at generating images with a strong artistic style and is suitable for illustration and creative design; the diffusion model, relying on the latent space diffusion technology, supports users to customize training and generate high-resolution images; while the DALL-E 3 model significantly improves the ability to restore details of complex texts by integrating CLIP semantic alignment and diffusion models. However, despite significant technological progress, the core performance of text-to-image models still highly depends on the quality of the input text. That is, if the prompt lacks accuracy, logic or detailed description, the text-to-image model may generate images that deviate from expectations due to semantic ambiguity or insufficient information, and thus it is difficult to construct a text-image dataset with high quality.
[0003] Therefore, how to construct a text-to-image model prompt text with high text quality so as to construct a text-image dataset with high quality is a technical problem yet to be solved by those skilled in the art. Summary of the Invention
[0004] The main purpose of this application is to provide a method, system, device, storage medium and product for constructing a text-image dataset, aiming to solve the technical problem of how to construct a text-to-image model prompt text with high text quality so as to construct a text-image dataset with high quality.
[0005] To achieve the above object, this application proposes a method for constructing a text-image dataset, the method comprising: extracting entities, entity attributes and relationships between entities based on the original text corresponding to each original image, and constructing a foreground feature knowledge base based on the entities, the entity attributes and the relationships between entities; extracting background description samples from each of the original texts, and constructing a background feature knowledge base based on the background description samples; constructing text based on the foreground feature knowledge base and the background feature knowledge base to obtain a target prompt text; generating a corresponding target image by using the target prompt text, and forming a text-image pair based on the target image and the corresponding target prompt text for constructing the text-image dataset.
[0006] In one embodiment, before the step of extracting entities, entity attributes, and relationships between entities based on the original text corresponding to each original image, the method for constructing the graphic-text dataset further includes: Segment and classify each original image to obtain respective visual features; Based on a preset semantic description knowledge base and each of the visual features, generate the original text corresponding to each original image.
[0007] In one embodiment, the method further includes: Perform data cleaning and spelling correction on the original text corresponding to each original image to obtain preprocessed text; The step of extracting entities, entity attributes, and relationships between entities based on the original text corresponding to each original image includes: Extract entities, entity attributes, and relationships between entities based on each of the preprocessed texts.
[0008] In one embodiment, the step of extracting entities, entity attributes, and relationships between entities based on the original text corresponding to each original image includes: Perform semantic splitting on the original text corresponding to each original image to obtain respective words corresponding to each original text; Identify entities in each of the words through a large language model, and based on the large language model and each of the original texts, obtain relationships between entities and entity attributes corresponding to each entity, where the entity attributes include: functional attribute, shape attribute, size attribute, color attribute, pattern attribute, texture attribute, material attribute, situation attribute, opacity attribute, direction attribute, action attribute, text attribute.
[0009] In one embodiment, after the step of obtaining relationships between entities and entity attributes corresponding to each entity based on the large language model and each of the original texts, the method further includes: Classify each of the entities to obtain entity categories corresponding to each entity; Store each of the entities, the relationships between entities, the entity attributes corresponding to each entity, and the entity categories corresponding to each entity.
[0010] In one embodiment, the step of constructing a background feature knowledge base based on the background description samples includes: Extract each background attribute in the background description samples, where the background attributes include: scene attribute, background subject attribute, picture style attribute, composition attribute, lighting attribute, color palette attribute, texture attribute, depth of field attribute, theme attribute, mood attribute, counting attribute, perspective attribute; Construct a background feature knowledge base according to each of the background attributes.
[0011] In one embodiment, the step of constructing text based on the foreground feature knowledge base and the background feature knowledge base to obtain a target prompt text includes: When it is detected that there is no original prompt text, determine a first random feature in the foreground feature knowledge base and a second random feature in the background feature knowledge base, where the original prompt text is the prompt text received in advance; Construct text based on the first random feature and the second random feature through a large language model to obtain a target prompt text.
[0012] In one embodiment, the step of constructing text based on the foreground feature knowledge base and the background feature knowledge base to obtain a target prompt text includes: When it is detected that there is the original prompt text, screen for a first target feature corresponding to the original prompt text in the foreground feature knowledge base through a large language model, and construct a foreground description text based on the original prompt text and the first target feature through the large language model; Screen for a second target feature corresponding to the original prompt text in the background feature knowledge base through the large language model, and construct a background description text based on the original prompt text and the second target feature through the large language model; Obtain a target prompt text according to the foreground description text and the background description text.
[0013] In one embodiment, the step of obtaining a target prompt text according to the foreground description text and the background description text includes: Splice the foreground description text and the background description text through the large language model to obtain an intermediate prompt text; Determine the words to be replaced in the intermediate prompt text according to a preset word replacement rule, generate target words based on the words to be replaced through the large language model, and replace the words to be replaced with the target words to obtain a target prompt text, where the word replacement rule at least includes: adjective replacement, positional relationship replacement, and quantitative relationship replacement.
[0014] In one embodiment, the step of generating a corresponding target image using the target prompt text and forming a text-image pair based on the target image and the corresponding target prompt text includes: Input the target prompt text into a preset text-to-image model to obtain a target image; Input the target image and the target prompt text into a contrastive language-image pre-training model to obtain the similarity between the target image and the target prompt text. When it is detected that the similarity is greater than or equal to a preset similarity threshold, the target image and the target prompt text are constructed into a text-image pair, and the text-image pair is stored.
[0015] In one embodiment, after the step of constructing the target image and the target prompt text into a text-image pair, the method further includes: Among various preset image types, determine the target image type corresponding to the target image, and divide the text-image pair into the text-image pair dataset corresponding to the target image type.
[0016] In addition, to achieve the above object, the present application further provides a system for constructing a text-image dataset, the system for constructing a text-image dataset includes: A foreground feature knowledge base construction module, configured to extract entities, entity attributes, and relationships between entities based on the original text corresponding to each original image, and construct a foreground feature knowledge base based on the entities, the entity attributes, and the relationships between entities; A background feature knowledge base construction module, configured to extract background description samples from each of the original texts, and construct a background feature knowledge base based on the background description samples; A text construction module, configured to construct a text based on the foreground feature knowledge base and the background feature knowledge base to obtain a target prompt text; A text-image pair construction module, configured to generate a corresponding target image by using the target prompt text, and form a text-image pair based on the target image and the corresponding target prompt text for constructing the text-image dataset.
[0017] In addition, to achieve the above object, the present application further provides an electronic device, the electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program is configured to implement the steps of the method for constructing a text-image dataset as described above.
[0018] In addition, to achieve the above object, the present application further provides a storage medium, the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the method for constructing a text-image dataset as described above are implemented.
[0019] In addition, to achieve the above object, the present application further provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the method for constructing a text-image dataset as described above are implemented.
[0020] In the embodiments of the present application, based on the original texts corresponding to each original image, entities, entity attributes, and relationships between entities are extracted, and a foreground feature knowledge base is constructed based on the entities, entity attributes, and relationships between entities. It is possible to extract the entities, entity attributes, and relationships between entities in the image from the publicly available images, and use the extracted information as foreground features to construct a foreground feature knowledge base; then, background description samples in each original text are extracted, and a background feature knowledge base is constructed based on the background description samples. It is possible to extract the background features in the image from the publicly available images and construct a background feature knowledge base based on the background features; then, texts are constructed based on the foreground feature knowledge base and the background feature knowledge base to obtain target prompt texts. The known foreground features and background features can be combined to obtain semantically rich and accurate target prompt texts; then, corresponding target images are generated using the target prompt texts, and a text-image pair is formed based on the target image and the corresponding target prompt text for constructing a text-image dataset. A text-image dataset with high quality can be obtained based on the accurate target prompt texts.
[0021] In the present application, information such as detailed descriptions of objects, overall environments, styles, etc. in publicly available images can be extracted, and a foreground feature knowledge base and a background feature knowledge base are established based on the extracted information. Then, through a large language model, rich prompt texts can be generated based on the foreground feature knowledge base and the background feature knowledge base, so that it is possible to construct a text-to-image model prompt text with high text quality. Furthermore, a target image with high quality can be obtained based on the text-to-image model prompt text with high quality, in order to construct a text-image dataset with high quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0023] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 It is a schematic flowchart of the first embodiment of the method for constructing a text-image dataset of the present application; Figure 2 It is a schematic flowchart of the construction process of the target prompt text in an embodiment of the method for constructing a text-image dataset of the present application; Figure 3 It is a schematic diagram of the construction process of the text-image pair in an embodiment of the method for constructing a text-image dataset of the present application; Figure 4Schematic diagram of the module structure of the construction system of the text and image dataset according to the embodiments of the present application; Figure 5 Schematic diagram of the device structure of the hardware operating environment involved in the construction method of the text and image dataset according to the embodiments of the present application.
[0025] The implementation, functional features and advantages of the present application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Detailed implementation manners
[0026] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0027] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the drawings in the specification and specific implementation manners.
[0028] It can be understood that the text-to-image model is a generative artificial intelligence model that can automatically generate an image semantically associated with the text description provided by the user. Such models usually adopt diffusion models, generative adversarial networks or autoregressive architectures, and are trained with a large amount of text-image pair data to learn the mapping relationship between text and visual features. For example, the MidJourney model is good at generating images with strong artistic styles and is suitable for illustration and creative design; the diffusion model, relying on the latent space diffusion technology, supports users to customize training and generate high-resolution images; while the DALL-E 3 model significantly improves the ability to restore details of complex texts by integrating CLIP semantic alignment and diffusion models. However, despite significant technological progress, the core performance of the text-to-image model still highly depends on the quality of the input text. That is, if the prompt lacks precision, logic or detailed description, the text-to-image model may generate images that deviate from expectations due to semantic ambiguity or insufficient information.
[0029] Therefore, how to construct high-quality text for the text-to-image model prompt is a technical problem that those skilled in the art still need to solve.
[0030] To solve the above problems, in this application, entities, entity attributes, and relationships between entities are extracted based on the original text corresponding to each original image, and a foreground feature knowledge base is constructed based on the entities, entity attributes, and relationships between entities. Entities, entity attributes, and relationships between entities in the image can be extracted based on publicly available images, and the extracted information is used as foreground features to construct a foreground feature knowledge base. Then, background description samples in each original text are extracted, and a background feature knowledge base is constructed based on the background description samples. Background features in the image can be extracted based on publicly available images, and a background feature knowledge base is constructed based on the background features. Then, text is constructed based on the foreground feature knowledge base and the background feature knowledge base to obtain a target prompt text. Known foreground features and background features can be combined to obtain a semantically rich and accurate target prompt text. Then, a corresponding target image is generated using the target prompt text, and an image-text pair is formed based on the target image and the corresponding target prompt text for constructing an image-text data set. A high-quality image-text data set can be obtained based on the accurate target prompt text.
[0031] In this application, information such as detailed descriptions of objects, overall environments, and styles in publicly available images can be extracted, and a foreground feature knowledge base and a background feature knowledge base are established based on the extracted information. Then, a large language model can generate rich prompt texts based on the foreground feature knowledge base and the background feature knowledge base, so that a high-quality text-to-image model prompt text can be constructed. Furthermore, a high-quality target image can be obtained based on the high-quality text-to-image model prompt text, so as to construct a high-quality image-text data set.
[0032] It should be noted that the execution subject of the method for constructing the image-text data set in this application can be an electronic device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a server, etc. In the following embodiments, the execution subject is omitted for elaboration.
[0033] Based on this, an embodiment of this application provides a method for constructing an image-text data set, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the method for constructing the image-text data set in this application.
[0034] In this embodiment, the method for constructing the image-text data set includes steps S10 to S40: Step S10, based on the original text corresponding to each original image, extract entities, entity attributes, and relationships between entities, and construct a foreground feature knowledge base based on the entities, the entity attributes, and the relationships between entities; It should be noted that the original image refers to the already public image, and the original text refers to the text used to describe the content of the original image. There is a one-to-one correspondence between the original text and the original image. For example, the original text can be the title, label, and descriptive text corresponding to the original image. An entity refers to an object with independent meaning or recognizable in the image, and the entity attribute is the information description of the entity. The relationship between entities refers to the relationship between each entity. For example, in a feasible implementation, the relationship between entities includes at least a spatial position relationship and an interaction relationship. Among them, the spatial position relationship is the position information of multiple entities in the spatial dimension presented in the original image, and the interaction relationship refers to the information in the action dimension presented by multiple entities in the original image. For example, the interaction relationship can be "little boy -> wearing -> shirt, woman -> holding -> flag".
[0035] It can be understood that the foreground object is usually the core object of concern in the image. Therefore, accurately grasping the foreground object features is crucial for generating image details that meet the requirements. Different objects vary greatly in attributes such as category, shape, and color. Separately constructing them can facilitate targeted storage and invocation of these foreground features. For example, if an image containing flowers needs to be generated, it is necessary to accurately know the foreground features such as the type of flower, the shape of the petals, and the color gradient in order to make the generated flowers lifelike. Therefore, in this embodiment, the original image and the original text corresponding to the original image can be directly obtained from the open source dataset at the same time. Then, the original text corresponding to each original image can be analyzed to extract the entities, entity attributes, and the spatial position relationship and interaction relationship between each entity from the original text. Then, the entities, entity attributes, and the spatial position relationship and interaction relationship between each entity can be summarized to obtain a foreground feature knowledge base, so as to generate a prompt text for guiding the text-to-image model to generate foreground objects according to the features in the foreground feature knowledge base.
[0036] Step S20, extract the background description samples from each of the original texts, and construct a background feature knowledge base based on the background description samples; It can be understood that the background description sample refers to the text in the original text that describes the background information of the image. In an image, the background information determines the overall atmosphere and style of the image. For example, indoor scenes, outdoor scenes, the brightness and darkness of light and shadow, and the warmth and coldness of tones will generate images with different atmospheres and styles. Therefore, constructing a background knowledge base separately can better control the image style from a macroscopic perspective and meet diverse scene requirements. For example, to generate a vintage-style indoor home image, it is necessary to guide the image generation based on information such as the vintage color combination and typical indoor layout in the background knowledge base. Therefore, in this embodiment, the background description samples can also be extracted from the original text corresponding to each original image through a large language model, and then the background features represented by the background description samples can be extracted, and the various background features can be summarized to obtain a background feature knowledge base for subsequent calls.
[0037] Step S30, construct text based on the foreground feature knowledge base and the background feature knowledge base to obtain a target prompt text; In this embodiment, when it is necessary to generate a specific image through a text-to-image model, rich prompt texts can be constructed based on the features in the foreground feature knowledge base and the features in the background feature knowledge base, and the constructed prompt texts are used as the target prompt texts to generate a specific image based on the target prompt texts.
[0038] In the embodiment of the present application, by extracting entities, entity attributes, and relationships between entities based on the original text corresponding to each original image, and constructing a foreground feature knowledge base based on the entities, entity attributes, and relationships between entities, the entities, entity attributes, and relationships between entities in the image can be extracted from the publicly available images, and the extracted information is used as foreground features to construct a foreground feature knowledge base; then the background description samples in each original text are extracted, and a background feature knowledge base is constructed based on the background description samples. The background features in the image can be extracted from the publicly available images, and a background feature knowledge base is constructed based on the background features; then text is constructed based on the foreground feature knowledge base and the background feature knowledge base to obtain a target prompt text, and the known foreground features and background features can be combined to obtain a target prompt text with rich and accurate semantics.
[0039] In this embodiment, information such as the detailed description, overall environment, and style of the objects in the original image can be extracted, and a foreground feature knowledge base and a background feature knowledge base are established based on the extracted information. Then, through the large language model, rich prompt texts can be generated based on the foreground feature knowledge base and the background feature knowledge base, so that a text-to-image model prompt text with high text quality can be constructed.
[0040] Step S40: Generate a corresponding target image using the target prompt text, and form an image-text pair based on the target image and the corresponding target prompt text for constructing the image-text dataset.
[0041] It can be understood that the target image refers to the image generated by the text-to-image model based on the target prompt text.
[0042] In this embodiment, after obtaining the target prompt text, an image that meets the requirements of the target prompt text can be generated based on the text-to-image model. Then, the target prompt text and its corresponding target image can be combined into an image-text pair, and further, an image-text dataset can be constructed through multiple image-text pairs.
[0043] In this embodiment, the present application can extract information such as the detailed description of the objects, the overall environment, and the style in the original image, and establish a foreground feature knowledge base and a background feature knowledge base based on the extracted information. Then, the large language model can generate rich prompt texts based on the foreground feature knowledge base and the background feature knowledge base, so as to construct high-quality text-to-image model prompt texts. Furthermore, high-quality target images can be obtained based on the high-quality text-to-image model prompt texts, so as to construct a high-quality image-text dataset.
[0044] Further, based on the first embodiment of the construction method of the image-text dataset of the present application, a second embodiment of the construction method of the image-text dataset of the present application is proposed.
[0045] In this embodiment, before the above step S10, the method further includes: Step S100: Perform image segmentation and classification labeling on each original image to obtain respective visual features; It should be noted that image segmentation refers to dividing an image into multiple objects, and classification labeling refers to classifying the divided objects according to existing categories and labeling the corresponding categories on the objects. Visual features refer to the respective objects after category labeling.
[0046] Step S200: Generate respective original texts corresponding to each of the original images based on a preset semantic description knowledge base and each of the visual features.
[0047] It should be noted that the semantic description knowledge base refers to a template for semantic description. In this embodiment, the visual features can be filled into the semantic description knowledge base, and the template filled with visual features can be sorted out and filled by the large language model, so as to obtain the respective original texts corresponding to each of the original images.
[0048] It is understandable that the method for obtaining the original text in the first embodiment is to obtain it directly from a public data source. However, in the public data source, the original image is usually not equipped with the original text. Therefore, the method of directly obtaining the original text from the public data source in the first embodiment will result in fewer samples, resulting in insufficient content in the obtained foreground feature knowledge base and background feature knowledge base.
[0049] In this embodiment, it is only necessary to obtain the original image. Specifically, the original image can be processed so that the original text corresponding to the original image can be automatically generated. Thus, even in a scenario where the sample of the original text is small, the present application can automatically generate the original text, thereby facilitating the enrichment of the background feature knowledge base and the foreground feature knowledge base, and further facilitating the construction of rich text prompts.
[0050] Furthermore, based on the first embodiment of the method for constructing a graphic and text dataset of the present application, a third embodiment of the method for constructing a graphic and text dataset of the present application is proposed.
[0051] In this embodiment, the method further includes: Step S300, performing data cleaning and spelling correction on the original text corresponding to each original image to obtain a pre-processed text; It is understandable that the original text obtained through the open source database usually contains noise information, such as garbled characters, special symbols, irrelevant punctuation, etc. Therefore, it is necessary to remove the noise information in the original text through data cleaning. Splicing error correction refers to correcting spelling errors in the original text to improve the accuracy of the original text. Preprocessed text refers to the original text obtained after data cleaning and spelling error correction steps.
[0052] Based on this, the above step S10 includes: Step S101, extracting entities, entity attributes and relationships between entities based on the preprocessed texts.
[0053] It can be understood that in this embodiment, more accurate entities, entity attributes, and spatial position relationships can be obtained based on the preprocessed text with higher accuracy, thereby improving the text quality of the target prompt text.
[0054] Furthermore, based on the first embodiment of the method for constructing a graphic and text dataset of the present application, a fourth embodiment of the method for constructing a graphic and text dataset of the present application is proposed.
[0055] In this embodiment, the above step S10 includes: Step S102, semantically splitting the original texts corresponding to the original images to obtain the words corresponding to the original texts; Step S103: Identify the entities in each of the said words through a large language model, and based on the large language model and each of the said original texts, obtain the relationships between entities and the entity attributes corresponding to each of the said entities respectively. The entity attributes include: function attribute, shape attribute, size attribute, color attribute, pattern attribute, texture attribute, material attribute, situation attribute, opacity attribute, direction attribute, action attribute, text attribute.
[0056] Exemplarily, the original text can be tokenized through a large language model, so that the continuous text sequence in the original text can be split into individual words or phrases according to semantics. For example, for the original text "A red round balloon is floating in the air", after tokenization, words such as "a", "red", "round", "balloon", "in the air", "floating" are obtained for subsequent analysis and feature extraction.
[0057] Then, entity recognition can be performed on the tokenized text through a large language model to determine the nouns or noun phrases describing objects therein as entities. For example, "balloon" in the above original text is an entity.
[0058] Then, the entity attributes in the original text can be extracted through a large language model based on the definitions of each entity attribute.
[0059] In a feasible implementation manner, the definitions of each entity attribute are as follows: Function attribute: Through semantic analysis and combined with context, understand the role of the entity in the image in the original text. For example, for the text "There is a table lamp on the table for lighting", identify that the function of the table lamp is "lighting"; if the text description is "The worker is holding a hammer and hammering a nail", then the function of the hammer is "hammering".
[0060] Shape attribute: Extract shape description words from the original text to judge the shape of the entity. For example, the entity can be circular, square, triangular in shape. For example, if the original text is "There is a rectangular painting hanging on the wall", it can be determined that the shape of the painting is "rectangular".
[0061] Size attribute: Search for words related to size comparison in the original text, such as "big", "small", "huge", "tiny", etc., and combined with the size reference of common objects, judge the size situation of the object relative to other objects or general cognition. For example, if the original text is "There is a big tree by a small pond", it can be known that the tree is "big" and the pond is "small"; if the original text is "Holding a tiny chip in the hand", then it can be reflected that the chip is "tiny".
[0062] Color Attribute: Extract color terms from the original text, including descriptions of solid colors such as "red", "blue", "green", etc., and variegated colors such as "colorful", "flowery", etc., to determine the color characteristics of the entity. When the original text is "She is wearing a pink dress", the color of the dress is "pink"; when the original text is "There are many variegated butterflies flying in the flowerbed", the color of the butterflies is "variegated".
[0063] Pattern Attribute: Analyze the vocabulary in the original text regarding the surface style of the entity, and judge whether the entity is solid, has patterns (such as stripes, plaids, floral patterns, etc.), or geometric patterns (such as circular pattern arrangements, triangular splicings, etc.). When the original text is "This wooden table has delicate carvings on its surface", the pattern attribute of the table is "carvings"; when the original text is "The floor is paved with black and white square tiles", the pattern attribute of the tiles is "black and white square geometric pattern".
[0064] Texture Attribute: Based on the descriptive vocabulary in the original text, judge whether the surface of the object is smooth (such as the surface of glass), rough (such as sandpaper), uneven (such as a cobblestone road surface), shiny (such as a polished metal surface), or dull (such as old fabric). When the original text is "A metal ball that feels very smooth", the texture attribute of the metal ball is "smooth"; when the original text is "The walls of this ancient castle appear uneven", the texture attribute of the castle walls is "uneven".
[0065] Material Attribute: Identify the material nouns mentioned in the original text to determine the material of the entity. Common materials include wood, metal, glass, plastic, ceramic, etc. When the original text is "A vase made of ceramic is placed on the table", the material attribute of the vase is "ceramic"; when the original text is "The bench in the park is made of wood", the material attribute of the bench is "wood".
[0066] Condition Attribute: Search for words describing the state of the entity such as new, good, damaged, worn, etc., to judge the situation of the object. For example, when the original text is "That old car is parked by the roadside", the condition attribute of the car is "old"; when the original text is "The newly opened book emits the fragrance of ink", the condition attribute of the book is "new".
[0067] Opacity Attribute: Judge from the original text whether the object is transparent (such as a glass water cup), semi - transparent (such as frosted glass, certain thin - gauze materials), or opaque (such as a metal block, a wooden board); for example, when the original text is "You can see the scenery outside through the transparent glass window", the opacity attribute of the glass window is "transparent"; when the original text is "Use an opaque curtain to block the stage", the opacity attribute of the curtain is "opaque".
[0068] Direction attribute: Analyze the descriptions of directions in the original text regarding the placement of physical objects, the postures of people, etc., and determine whether the object is upright (such as a standing flagpole), horizontal (such as a lying flat wooden board), inverted (such as a hanging bat), or inclined (such as a ladder leaning against the wall). For example, when the original text is "That upright telegraph pole towers into the clouds", the direction attribute of the telegraph pole is "upright"; when the original text is "He leaned the picture frame against the wall", the direction attribute of the picture frame is "inclined".
[0069] Action attribute: For people or movable objects, extract the words describing their actions and postures in the original text, such as "running", "jumping", "sitting", "waving", etc.; if the original text is "The athlete is running vigorously on the field", then the action attribute of the athlete is "running".
[0070] Text attribute: Search for information in the original text about whether there are words on the surface of the object, as well as the writing position, font, number of lines, paragraphs, etc. of the words; for example, when the original text is "The packaging box is printed with a prominent brand name, the font is bold regular script, and it is arranged in a single line", it can be determined that the text attribute of the packaging box is: there are words, and the position is on the "surface", the font is "bold regular script", and the number of lines is "single line".
[0071] Then, the spatial position relationship and / or interaction relationship between different entities in the original text can be analyzed. Specifically, semantic dependency analysis technology can be used to find the interactions and / or relative spatial relationships between entities; for example, when the original text is "The book is placed on the table", the spatial position relationship between the book and the table is "placed on"; when the original text is "The car is driving on the road", the spatial position relationship between the car and the road is "driving on".
[0072] Thus, the present application can obtain each component of the foreground feature knowledge base based on the original text corresponding to the image.
[0073] In a feasible implementation manner, after the above step S103, the method further includes: Step S104, classifying each of the entities to obtain the entity category corresponding to each of the entities; Step S105, storing each of the entities, the relationships between the entities, the entity attributes corresponding to each of the entities, and the entity categories corresponding to each of the entities.
[0074] Exemplarily, after obtaining the entities, the identified entities can be classified according to a predefined category system, where the above entity categories are the results of classification. Specifically, the category system can cover common object categories, such as: people, animals, plants, transportation vehicles, equipment and tools, etc. For example, "balloon" can be classified into the category of "equipment and tools". Then, the entities, the relationships between entities, the entity attributes of each entity, and the entity categories corresponding to each entity can be stored.
[0075] In a feasible implementation manner, a knowledge graph can be used to store entities, relationships between entities, entity attributes of each entity, and entity categories corresponding to each entity. Among them, the entity categories can exist in the form of labels in the knowledge graph. Thus, when the user needs to generate an image with specific categories and specific entity relationships, information such as target entities and target entity relationships can be filtered out from the knowledge graph based on the content input by the user, and then an image can be quickly generated based on the filtered information, so as to achieve the purpose of improving the image generation efficiency.
[0076] In a feasible implementation manner, the above step S20 includes: Step S201, extract each background attribute in the background description sample, where the background attributes include: scene attribute, background subject attribute, picture style attribute, composition attribute, lighting attribute, color palette attribute, texture attribute, depth of field attribute, theme attribute, emotion attribute, counting attribute, perspective attribute; Step S202, construct a background feature knowledge base according to each of the background attributes.
[0077] Exemplarily, statements and paragraphs describing the background can be screened out from the original text to obtain a background description sample. Then, the background description sample can be normalized, the text format can be unified, and typos can be corrected to improve the quality of the background description sample.
[0078] Then, based on the definitions of each background attribute, each background attribute can be extracted from the background description sample through a large language model.
[0079] Specifically, the definitions of each background attribute are: Scene attribute: Judge the scene type of the image from the background description sample; the scene attribute can be: indoor (mentioning indoor elements such as rooms, furniture, walls, etc.), outdoor (appearing natural elements such as sky, grassland, mountains, etc.), landscape (focusing on describing landscape scenery such as mountains, rivers, lakes, etc.), city street (including urban facilities such as high-rise buildings, roads, street lights, etc.); if the background description sample is "Sunshine shines on the quiet town street, there are ancient buildings and street lights by the street", it can be known that the scene attribute corresponding to this background description sample is "city street".
[0080] Background Subject Attribute: Identify the description of the background subject elements in the background description sample. For example, if the background description sample is "Deep in the dark forest, the trees are tall and dense", then the background subject attribute of this background description sample is "forest"; if the background description sample is "Against the backdrop of blue sky and white clouds, there is an endless prairie below", then the background subject attribute is "blue sky and white clouds, prairie".
[0081] Picture Style Attribute: Judge the picture style based on the keywords in the background description sample; the picture style attribute can be: realistic (detailed description, close to reality, without exaggeration or distortion, for example, a photo truly records the street scene), cartoon (appearance of cartoon characters, bright colors, simple lines, such as "The cartoon castle in the animation is colorful"), oil painting (mention of oil painting brushstrokes, thick color texture, such as "This painting has a strong oil painting style, with rich and layered colors"), retro (including retro elements, such as old furniture, old-time atmosphere, for example, "There is an old phonograph in the room, full of retro flavor"), sci-fi (involving future technology elements, such as spaceships, laser weapons, for example, "In the vast universe, interstellar battleships shuttle through it, with a strong sense of sci-fi"), etc.
[0082] Composition Attribute: Analyze the description of the arrangement of elements and the distribution of focus in the image in the background description sample. For example, the composition attribute can be: symmetrical (such as "The palace building is symmetrical left and right, solemn and majestic"), balanced (the elements are evenly and harmoniously distributed, "The figures and scenery in the picture are properly matched, visually very balanced"), or asymmetrical ("Modern art works, with elements randomly combined, presenting an asymmetrical beauty").
[0083] Illumination Attribute: Search for the description of the light source in the background description sample; the illumination attribute can be natural light source (such as sunlight, moonlight, "The morning sunlight shines through the leaves on the ground"), artificial light source (light, candlelight, "There is a dim light in the room"), and information such as the direction and intensity of the light source, such as "Strong light hits the model's face from the side, highlighting the three-dimensional sense".
[0084] Palette Attribute: Extract the main colors and the color matching relationship mentioned in the background description sample; for example, the palette attribute can be "The picture is mainly in warm colors, and the orange-red sunset and the golden wheat fields set off each other"; the palette attribute can also be "The cyan-blue sea water in cold tones beats against the black reefs".
[0085] Texture Attribute: Judge the texture of the overall picture or key elements according to the description in the background description sample; the texture attribute can be: smooth (such as "The smooth marble floor reflects light"), rough ("The surface of the ancient city wall is rough, full of a sense of historical vicissitudes"), shiny ("The metal-textured sculpture shines brightly in the sun"), or dull ("The dim basement, with the walls in dull colors").
[0086] Depth of field attributes: Analyze the description of the image focus and clarity range in the background description samples; depth of field attributes can be: the entire image is clear (such as the description of a panoramic landscape photo "the beautiful scenery in front of you is unobstructed, from the flowers in the near distance to the mountains in the distance are all clearly visible"), partial clarity ("close-up of the character, blurred background, highlighting the protagonist's expression") and the atmosphere created by the depth of field effect, such as "shallow depth of field makes the subject stand out, and the blurred background creates a dreamy feeling."
[0087] Theme attribute: summarizes the core theme of the image described by the background description sample, as well as the relationship between other elements and the theme; for example, if the background description sample is "with maternal love as the theme, the mother is holding her child tenderly, with the cradle and toys beside her dotted around", then the theme attribute is "maternal love", and other elements serve as a foil; if the background description sample is "environmental protection theme poster, with withered trees contrasting with new green buds", then the theme attribute is "environmental protection".
[0088] Emotional attributes: Understand the overall emotion conveyed by the image from the background description samples; emotional attributes can be: cheerful ("The children are laughing and playing in the playground, and the picture is full of cheerful atmosphere"), tranquil ("The quiet lake, only the breeze blows across the water, rippling"), mysterious ("The ancient castle is looming in the mist, exuding a mysterious atmosphere"), tense ("The battlefield is filled with smoke, the soldiers are nervous, and the situation is on the verge of breaking out"), etc.
[0089] Count attribute: Count the approximate number of main and secondary objects mentioned in the background description sample. If the background description sample is "there are three people in the picture, and several birds flying across the sky", the count attribute can be the number of people and birds, which is used as a reference for subsequent composition.
[0090] Perspective attribute: The camera shooting perspective is judged based on the background description sample keywords. If the background description sample is "from a horizontal perspective, showing the daily life of people on the street", it is at eye level; the background description sample is "from a high altitude, the prosperity of the city is in sight", it is a bird's-eye view; the background description sample is "looking up at the tall monument, it looks more majestic and spectacular", which is a low angle; the background description sample is "lying on the grass, shooting flowers from a worm's-eye perspective, presenting a different microscopic world", which is a worm's-eye perspective, etc.
[0091] Then, the background attributes in each background description sample may be summarized to obtain a background feature knowledge base, so as to obtain a rich prompt text based on the rich background feature knowledge base.
[0092] In a feasible implementation manner, the above step S30 includes: Step S301, when it is detected that there is no original prompt text, a first random feature is determined in the foreground feature knowledge base, and a second random feature is determined in the background feature knowledge base, wherein the original prompt text is a previously received prompt text; It is understandable that the original prompt text refers to the keywords or key sentences given by the user to generate a specific field or a specific scene. The first random feature refers to a feature randomly determined in the foreground feature knowledge base. The second random feature refers to a feature randomly determined in the background feature knowledge base. Specifically, the first random feature includes at least a randomly determined entity, a randomly determined entity attribute in the entity, and a randomly determined spatial position relationship or interaction relationship. The second random feature includes a randomly determined background attribute.
[0093] Step S302: constructing text based on the first random feature and the second random feature through a large language model to obtain a target prompt text.
[0094] It is understandable that when it is detected that the user has not given any key words or phrases, entities, entity attributes, relationships between entities, and background attributes can be randomly determined from the foreground feature knowledge base and the background feature knowledge base, and then the target prompt text can be generated based on the randomly determined features through the large language model.
[0095] In this embodiment, the application randomly selects entities and background information from two knowledge bases, and uses the rich text generation imagination of the large language model to generate rich and diverse target prompt texts that may transcend realism. Therefore, in this embodiment, the large language model can be driven to generate target prompt texts covering various scenarios and different styles by randomly determining features, thereby meeting the different creative needs of users.
[0096] Please refer to Figure 2 , Figure 2 This is a schematic diagram of the target prompt text construction process of an embodiment of the method for constructing a graphic and text data set of the present application. In a feasible implementation, the above step S30 also includes: Step S303, when the original prompt text is detected, the first target feature corresponding to the original prompt text is screened in the foreground feature knowledge base by using the large language model, and a foreground description text is constructed by using the large language model based on the original prompt text and the first target feature; It can be understood that the first target feature refers to the information related to the original prompt text in the foreground feature knowledge base. For example, when the original prompt text is a certain category, such as "person", the entity, entity attributes, and entity relationship of the person category can be used as the first target feature. Then, the large language model can construct a text describing the foreground of the image based on the original prompt text and the first target feature.
[0097] Step S304, filter the second target features corresponding to the original prompt text in the background feature knowledge base through the large language model, and construct a background description text based on the original prompt text and the second target features through the large language model; Similarly, the second target feature refers to the information related to the original prompt text in the background feature knowledge base. Specifically, the second target feature includes background attributes related to the original prompt text. Then, a background description text can still be constructed based on the original prompt text and the second target features through the large language model.
[0098] Step S305, obtain the target prompt text according to the foreground description text and the background description text.
[0099] In this embodiment, the foreground description text and the background description text can be spliced through the large language model, and the spliced text can be used as the target prompt text.
[0100] It can be understood that Figure 2 the steps of semantic enhancement of the original prompt text based on the foreground feature knowledge base and the background feature knowledge base refer to the steps of constructing the foreground description text, constructing the background description text, and splicing to obtain the target prompt text.
[0101] In addition, in a feasible implementation manner, the specific steps of the above step S303 may be: When it is detected that there is an original prompt text, if it is detected that the original prompt text does not match all categories in the foreground feature knowledge base, it is determined whether the original prompt text is an entity; in the case where it is determined that the original prompt text is an entity, the information related to the entity represented by the original prompt text in the foreground feature knowledge base can be used as the first target feature; in the case where it is determined that the original prompt text is not an entity, each entity in the foreground feature knowledge base can be matched with the original prompt text, and the information related to the entity with a higher matching degree can be used as the first target feature. Then, based on the language combination ability of the large language model, a semantically smooth foreground description text can be generated for the first target feature and the original prompt text through the large language model.
[0102] The specific steps of the above step S304 may be: For each known background attribute in the background feature knowledge base, through the natural language understanding ability of the large language model, the content with a higher matching degree with the original prompt text in each background attribute can be determined, and the content with a higher matching degree with the original prompt text in the background feature knowledge base can be used as the second target feature. Furthermore, based on the language combination ability of the large language model, the original prompt text and the second target feature can be integrated to obtain the background description text.
[0103] The specific steps of the above step S305 may be as follows: By combining the foreground description text and the background description text through a large language model, a target prompt text can be obtained.
[0104] Exemplarily, when the original prompt text is: Generate an image with a food theme, the large language model can search for entities of the food category in the foreground feature knowledge base. If there are entities of the food category in the foreground feature knowledge base, such as "cake", then extract the entity attributes of the entities in the food category. For example, the entity attributes can be: the shape of the cake is "round and multi-layered", the color is "pink cream with colorful candy decorations", and the materials are "flour, cream, sugar, etc.". Thus, the first target features are: "cake", "round and multi-layered", "pink cream with colorful candy decorations", "flour, cream, sugar, etc.". Furthermore, the foreground description text generated by the large language model based on "Generate an image with a food theme" and the above first target features can be: "A delicate round multi-layered cake with colorful candies dotted on the pink cream, emitting an attractive aroma". At the same time, the large language model can determine the background attributes in the background feature knowledge base that have a relatively high degree of match with the original prompt text of "Generate an image with a food theme". For example, this original prompt text has a relatively high degree of match with the "bright kitchen background" and "dining table scene" in the scene attributes, a relatively high degree of match with the "realistic and warm" in the picture style attributes, and a relatively high degree of match with the "warm tone-based, highlighting the color of food" in the color palette attributes. Then the second target features are: "bright kitchen background", "dining table scene", "realistic and warm", "warm tone-based, highlighting the color of food". The background description text generated by the large language model can be: "In a bright and warm kitchen, fresh fruits are placed on a wooden dining table, and the sun shines through the window on the food". Then, based on the language organization ability of the large language model, the foreground description text and the background description text can be combined to obtain the target prompt text. For example, the target prompt text can be: "In a bright and warm kitchen, a delicate round multi-layered cake is placed on a wooden dining table, with colorful candies dotted on the pink cream, emitting an attractive aroma, and there are also fresh fruits beside it, and the sun shines through the window on the food".
[0105] In this embodiment, the present application can generate high-quality prompt texts whether the user gives a prompt statement or not, thereby improving the scene adaptation ability of the present application.
[0106] In a feasible implementation manner, the above step S305 includes: Step S3051: Concatenate the foreground description text and the background description text through the large language model to obtain an intermediate prompt text; Step S3052: Determine the words to be replaced in the intermediate prompt text according to the preset word replacement rules, generate target words based on the words to be replaced through the large language model, and replace the words to be replaced with the target words to obtain a target prompt text. The word replacement rules at least include: adjective replacement, positional relationship replacement, and quantitative relationship replacement.
[0107] It can be understood that, in addition to directly using the text concatenated by the large language model as the target prompt text, the text concatenated by the large language model can also be used as the intermediate prompt text. Thus, adjectives, numerals, and positional relationship words in the intermediate prompt text can be replaced, and multiple target prompt texts can be obtained based on a single intermediate prompt text, and then multiple images can be obtained.
[0108] In this embodiment, the present application can expand the content of the prompt text by means of word replacement, so as to not only improve the quality of the target prompt text, but also ensure the diversity of the target prompt text.
[0109] Furthermore, based on each of the above embodiments of the construction method of the present application's text-image dataset, a fifth embodiment of the construction method of the present application's text-image dataset is proposed.
[0110] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the text-image pair construction process of an embodiment of the construction method of the present application's text-image dataset. In Figure 3 , the above step S40 includes: Step S401: Input the target prompt text into a preset text-to-image model to obtain a target image; Step S402: Input the target image and the target prompt text into a contrastive language-image pre-training model to obtain the similarity between the target image and the target prompt text; Step S403: When it is detected that the similarity is greater than or equal to a preset similarity threshold, construct the target image and the target prompt text into a text-image pair and store the text-image pair.
[0111] Exemplarily, the image encoder in a contrastive language-image pre-training model (such as CLIP) extracts features from the target image and converts the image into a high-dimensional image feature vector. Specifically, through structures such as convolutional neural networks, the pixel information of the image can be abstracted layer by layer to capture features such as the color, shape, texture, and object layout of the image; at the same time, the text encoder encodes the input target prompt text and converts the text into a text feature vector, understanding information such as the semantics, lexical relationships, and themes of the text through word embedding, multi-head attention mechanisms, etc. Then, the contrastive language-image pre-training model calculates the similarity between the image feature vector and the text feature vector, and filters out the low-similarity text-image data pairs, only retaining the text-image pairs with relatively high similarity to ensure the quality of the data stored in the database.
[0112] In this embodiment, after obtaining the target prompt text, an image can also be generated based on the target prompt text, and the similarity between the generated image and the target prompt text can be calculated to obtain high-quality text-image pairs. Thus, the present application can not only generate high-quality prompt texts, but also generate high-quality text-image pairs, laying a foundation for training an efficient image generation model.
[0113] In a feasible implementation manner, after the above step S403, the method further includes: Step S404, among the preset various image types, determine the target image type corresponding to the target image, and divide the text-image pair into the text-image pair dataset corresponding to the target image type.
[0114] Exemplarily, after obtaining the text-image pair, the target image can also be classified, so that the text-image pair can be divided into the text-image pair dataset corresponding to the type of the target image for subsequent use.
[0115] In another feasible implementation manner, metadata tags can also be added to the image. The metadata tags include information such as the category of the image (such as scenery, people, animals, etc.), style (such as realistic, cartoon, abstract, etc.), etc. Then, the sorted and labeled text-image pairs can be stored in the text-image synthesis dataset, so as to be more convenient for querying, screening, and classifying in the subsequent data usage and management processes, thereby improving the usability and management efficiency of the dataset.
[0116] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation on the construction method of the text-image dataset of the present application. Based on this technical concept, more forms of simple transformations are within the protection scope of the present application.
[0117] The present application also provides a construction system for a text-image dataset. Please refer to Figure 4 The construction system for the text-image dataset includes: The foreground feature knowledge base construction module 10 is configured to extract entities, entity attributes, and relationships between entities based on the original text corresponding to each original image, and construct a foreground feature knowledge base based on the entities, the entity attributes, and the relationships between entities; The background feature knowledge base construction module 20 is configured to extract background description samples from each of the original texts, and construct a background feature knowledge base based on the background description samples; The text construction module 30 is configured to construct a text based on the foreground feature knowledge base and the background feature knowledge base to obtain a target prompt text; The image-text pair construction module 40 is configured to generate a corresponding target image by using the target prompt text, and form an image-text pair based on the target image and the corresponding target prompt text for constructing the image-text data set.
[0118] In one embodiment, the construction system of the image-text data set further includes: The image processing module is configured to perform image segmentation and classification labeling on each original image to obtain respective visual features; The original text generation module is configured to generate the original text corresponding to each original image based on a preset semantic description knowledge base and each of the visual features.
[0119] In one embodiment, the construction system of the image-text data set further includes: The text preprocessing module is configured to perform data cleaning and spelling correction on the original text corresponding to each original image to obtain a preprocessed text; Based on this, the above-mentioned foreground feature knowledge base construction module 10 is further configured to: Extract entities, entity attributes, and relationships between entities based on each of the preprocessed texts.
[0120] In one embodiment, the above-mentioned foreground feature knowledge base construction module 10 is further configured to: Perform semantic splitting on the original text corresponding to each original image to obtain respective words corresponding to each of the original texts; Identify entities in each of the words through a large language model, and obtain relationships between entities and entity attributes corresponding to each of the entities based on the large language model and each of the original texts, where the entity attributes include: function attribute, shape attribute, size attribute, color attribute, pattern attribute, texture attribute, material attribute, situation attribute, opacity attribute, direction attribute, action attribute, text attribute.
[0121] In one embodiment, the construction system of the image-text data set further includes: An entity classification module, configured to classify each of the entities to obtain the entity category corresponding to each of the entities; A data storage module, configured to store each of the entities, the relationships between the entities, the entity attributes corresponding to each of the entities, and the entity categories corresponding to each of the entities.
[0122] In one embodiment, the above-mentioned background feature knowledge base construction module 20 is further configured to: Extract each background attribute in the background description sample, where the background attributes include: scene attribute, background subject attribute, picture style attribute, composition attribute, lighting attribute, color palette attribute, texture attribute, depth of field attribute, theme attribute, emotion attribute, counting attribute, perspective attribute; Construct a background feature knowledge base according to each of the background attributes.
[0123] In one embodiment, the text construction module 30 is configured to: When it is detected that there is no original prompt text, determine a first random feature in the foreground feature knowledge base and a second random feature in the background feature knowledge base, where the original prompt text is the prompt text received in advance; Construct text through a large language model based on the first random feature and the second random feature to obtain a target prompt text.
[0124] In one embodiment, the text construction module 30 is configured to: When it is detected that there is the original prompt text, screen a first target feature corresponding to the original prompt text in the foreground feature knowledge base through a large language model, and construct a foreground description text through the large language model based on the original prompt text and the first target feature; Screen a second target feature corresponding to the original prompt text in the background feature knowledge base through the large language model, and construct a background description text through the large language model based on the original prompt text and the second target feature; Obtain a target prompt text according to the foreground description text and the background description text.
[0125] In one embodiment, the text construction module 30 is configured to: Concatenate the foreground description text and the background description text through the large language model to obtain an intermediate prompt text; Determine the words to be replaced in the intermediate prompt text according to a preset word replacement rule, generate target words through the large language model based on the words to be replaced, and replace the words to be replaced with the target words to obtain a target prompt text, where the word replacement rule at least includes: adjective replacement, positional relationship replacement, and quantitative relationship replacement.
[0126] In one embodiment, the graphic-text pair construction module 40 is configured to: Input the target prompt text into a preset text-to-image model to obtain a target image; Input the target image and the target prompt text into a contrastive language-image pre-training model to obtain the similarity between the target image and the target prompt text; When it is detected that the similarity is greater than or equal to a preset similarity threshold, construct the target image and the target prompt text into a graphic-text pair, and store the graphic-text pair.
[0127] In one embodiment, the construction system of the graphic-text data set further includes: A graphic-text pair classification module, configured to determine the target image type corresponding to the target image among various preset image types, and divide the graphic-text pair into the graphic-text data set corresponding to the target image type.
[0128] The construction system of the graphic-text data set provided by this application adopts the construction method of the graphic-text data set in the above embodiment, and can solve the technical problem of how to construct a text-to-image model prompt text with higher text quality so as to construct a graphic-text data set with higher quality. Compared with the prior art, the beneficial effects of the construction system of the graphic-text data set provided by this application are the same as those of the construction method of the graphic-text data set provided in the above embodiment, and other technical features in the construction system of the graphic-text data set are the same as the features disclosed in the above embodiment method, and will not be elaborated here.
[0129] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the construction method of the graphic-text data set in the first embodiment above.
[0130] Next, refer to Figure 5 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of this application. Figure 5 The electronic device shown is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of this application.
[0131] As Figure 5As shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or had alternatively.
[0132] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0133] The electronic device provided by the present application, adopting the method for constructing the graphic and text data set in the above embodiments, can solve the technical problem of how to construct a text generation image model prompt text with relatively high text quality so as to construct a graphic and text data set with relatively high quality. Compared with the prior art, the beneficial effects of the electronic device provided by the present application are the same as those of the method for constructing the graphic and text data set provided by the above embodiments, and other technical features in the electronic device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0134] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0135] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0136] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the method for constructing a graphic and text data set in the above embodiments.
[0137] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM: Random Access Memory), read-only memory (ROM: Read Only Memory), erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0138] The above computer-readable storage medium can be included in an electronic device; or it can exist separately without being assembled into the electronic device.
[0139] The above computer-readable storage medium stores one or more programs which, when executed by an electronic device, cause the electronic device to: extract entities, entity attributes, and relationships between entities based on the original text corresponding to each original image, and construct a foreground feature knowledge base based on the entities, the entity attributes, and the relationships between entities; extract background description samples from each of the original texts, and construct a background feature knowledge base based on the background description samples; and construct text based on the foreground feature knowledge base and the background feature knowledge base to obtain a target prompt text.
[0140] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0142] The modules described in the embodiments of this application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0143] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-described method for constructing a graphic and text data set, and can solve the technical problem of how to construct a text-to-image model prompt text with relatively high text quality so as to construct a graphic and text data set with relatively high quality. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the method for constructing a graphic and text data set provided in the above embodiments, and will not be elaborated here.
[0144] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method for constructing a graphic and text data set as described above are implemented.
[0145] The computer program product provided by this application can solve the technical problem of how to construct a text-to-image model prompt text with relatively high text quality so as to construct a graphic and text data set with relatively high quality. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as those of the method for constructing a graphic and text data set provided in the above embodiments, and will not be elaborated here.
[0146] The above are only some embodiments of this application, and thus do not limit the patent scope of this application. Any equivalent structural transformation made under the technical concept of this application by using the content of the specification and drawings of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.
Claims
1. A method for constructing a graphic and text data set, characterized in that: include: Based on the original texts corresponding to the original images, entities, entity attributes and relationships between entities are extracted, and a foreground feature knowledge base is constructed based on the entities, the entity attributes and the relationships between entities; Extracting background description samples from each of the original texts, and building a background feature knowledge base based on the background description samples; Constructing text based on the foreground feature knowledge base and the background feature knowledge base to obtain target prompt text; The target prompt text is used to generate a corresponding target image, and a picture-text pair is formed based on the target image and the corresponding target prompt text to construct the picture-text data set.
2. The method for constructing a graphic and text data set according to claim 1, characterized in that: Before the step of extracting entities, entity attributes and relationships between entities based on the original texts corresponding to the original images, the method for constructing the image-text dataset further includes: Perform image segmentation and classification labeling on each original image to obtain various visual features; Based on a preset semantic description knowledge base and each of the visual features, the original text corresponding to each of the original images is generated.
3. The method for constructing a graphic and text data set according to claim 1, characterized in that: The method further comprises: Perform data cleaning and spelling correction on the original text corresponding to each original image to obtain a preprocessed text; The step of extracting entities, entity attributes and relationships between entities based on the original texts corresponding to the original images comprises: Based on the preprocessed texts, entities, entity attributes and relationships between entities are extracted.
4. The method for constructing a graphic and text data set according to claim 1, characterized in that: The step of extracting entities, entity attributes and relationships between entities based on the original texts corresponding to the original images comprises: Semantically splitting the original texts corresponding to the original images to obtain the words corresponding to the original texts; Entities in each of the words are identified through a large language model, and based on the large language model and the original texts, the relationship between entities and the entity attributes corresponding to each of the entities are obtained, and the entity attributes include: function attributes, shape attributes, size attributes, color attributes, pattern attributes, texture attributes, material attributes, situation attributes, opacity attributes, direction attributes, action attributes, and text attributes.
5. The method for constructing a graphic and text data set according to claim 4, characterized in that: After the step of obtaining the relationship between entities and the entity attributes corresponding to each of the entities based on the large language model and the original texts, the method further includes: Classifying each of the entities to obtain entity categories corresponding to each of the entities; The entities, the relationships between the entities, the entity attributes corresponding to the entities, and the entity categories corresponding to the entities are stored.
6. The method for constructing a graphic and text data set according to claim 1, characterized in that: The step of constructing a background feature knowledge base based on the background description sample comprises: Extracting various background attributes from the background description sample, the background attributes including: scene attributes, background subject attributes, picture style attributes, composition attributes, lighting attributes, palette attributes, texture attributes, depth of field attributes, theme attributes, emotion attributes, count attributes, and viewing angle attributes; A background feature knowledge base is constructed according to each of the background attributes.
7. The method for constructing a graphic and text data set according to claim 1, characterized in that: The step of constructing text based on the foreground feature knowledge base and the background feature knowledge base to obtain the target prompt text includes: When it is detected that there is no original prompt text, determining a first random feature in the foreground feature knowledge base and determining a second random feature in the background feature knowledge base, the original prompt text being a previously received prompt text; A text is constructed based on the first random feature and the second random feature through a large language model to obtain a target prompt text.
8. The method for constructing a graphic and text data set according to claim 7, characterized in that: The step of constructing text based on the foreground feature knowledge base and the background feature knowledge base to obtain the target prompt text includes: When the original prompt text is detected, a first target feature corresponding to the original prompt text is selected in the foreground feature knowledge base by using a large language model, and a foreground description text is constructed by using the large language model based on the original prompt text and the first target feature; Using the large language model to select a second target feature corresponding to the original prompt text in the background feature knowledge base, and using the large language model to construct a background description text based on the original prompt text and the second target feature; A target prompt text is obtained according to the foreground description text and the background description text.
9. The method for constructing a graphic and text data set according to claim 8, characterized in that: The step of obtaining the target prompt text according to the foreground description text and the background description text comprises: The foreground description text and the background description text are spliced together by the large language model to obtain an intermediate prompt text; According to preset word replacement rules, the words to be replaced in the intermediate prompt text are determined, and target words are generated based on the words to be replaced through the large language model, and the words to be replaced are replaced based on the target words to obtain a target prompt text, wherein the word replacement rules at least include: adjective replacement, position relationship replacement and quantity relationship replacement.
10. The method for constructing a graphic and text data set according to claim 1, characterized in that: The step of generating a corresponding target image using the target prompt text, and forming an image-text pair based on the target image and the corresponding target prompt text, comprises: Inputting the target prompt text into a preset text image model to obtain a target image; Inputting the target image and the target prompt text into a comparative language image pre-training model to obtain the similarity between the target image and the target prompt text; When it is detected that the similarity is greater than or equal to a preset similarity threshold, the target image and the target prompt text are constructed as an image-text pair, and the image-text pair is stored.
11. The method for constructing a graphic and text data set according to claim 10, characterized in that: After the step of constructing the target image and the target prompt text into an image-text pair, the method further includes: In each preset image type, a target image type corresponding to the target image is determined, and the image-text pairs are divided into an image-text pair data set corresponding to the target image type.
12. A system for constructing a graphic and text data set, characterized in that: The construction system of the graphic and text dataset includes: A foreground feature knowledge base construction module is used to extract entities, entity attributes and inter-entity relationships based on the original texts corresponding to the original images, and to construct a foreground feature knowledge base based on the entities, the entity attributes and the inter-entity relationships; A background feature knowledge base construction module, used for extracting background description samples from each of the original texts, and constructing a background feature knowledge base based on the background description samples; A text construction module, used to construct text based on the foreground feature knowledge base and the background feature knowledge base to obtain a target prompt text; The image-text pair construction module is used to generate a corresponding target image using the target prompt text, and form an image-text pair based on the target image and the corresponding target prompt text, so as to construct the image-text data set.
13. An electronic device, characterized in that: The electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for constructing a graphic and text data set as described in any one of claims 1 to 11.
14. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the method for constructing a graphic and text data set according to any one of claims 1 to 11 are implemented.
15. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the method for constructing a graphic and text data set as claimed in any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
A task-oriented text generation image network model
CN111858954A
Image association system based on AIGC cue word
CN117633285A
Picture generation model construction method and device, equipment and readable medium
CN117788637A
Image processing method and device and computer readable storage medium
CN117975497A
Image generation method and device, equipment and storage medium
CN118941717A