Integration method based on unified expression and fusion of design information
By filtering and fusing the image text pairs of the generation model, dynamically adjusting the expression of the propt, the problem of difficulty in the existing technology of prompt optimization and unmet personalized needs is solved, and an efficient and personalized image generation process is achieved.
Patent Information
- Application Number
- CN202411734069.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-29
AI Technical Summary
The prior art is difficult to effectively optimize the propt of the generation model, which makes it difficult for users to obtain ideal image generation results, and the existing methods do not consider user personalized needs and subsequent optimization of the propt.
By obtaining the original image text pairs of the dataset and the text input by the user, filtering image text pairs with high matching degrees, using the large language model and the literary and genomic graph model to generate a new prompt, and filtering images through semantic similarity to achieve iterative optimization.
It improves the quality and user experience of generated images, supports the rapid iteration and innovation of creative content, and meets users' personalized needs.
Smart Images

Figure CN119963671A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large models and human-computer interaction, and in particular to an integration method based on unified expression and fusion of design information. Background Art
[0002] With the continuous development of generative large model technology, top generative models such as Stable Diffusion and DALL·E have been able to generate highly relevant and high-quality images based on text prompts, greatly lowering the threshold for image creation and gradually becoming a popular and efficient creation tool.
[0003] Building on the remarkable achievements of these generative models, researchers and developers have further explored a human-computer interaction technique called “prompts”. This technique allows users to formulate natural language prompts describing desired image features (such as themes and styles) and adjust the model’s hyperparameters during the creation process to help users achieve ideal generation results. However, for novice users, writing effective prompts that can accurately optimize the model’s generation output is a challenging task. This is because the complexity and ambiguity of natural language make it difficult to fully express the user’s creative ideas. In addition, prompts based on different model hyperparameters may produce very different images. It is difficult to evaluate the quality of prompts under limited hyperparameter experiment conditions. When receiving image results that do not meet expectations, users may be confused about whether and how to adjust the prompts or model hyperparameters.
[0004] Existing prompt optimization methods mainly focus on generating prompts quickly and automatically, a process that is usually end-to-end and completed in one go. However, this method does not take into account the personalized needs of users, and is not convenient for subsequent optimization and improvement of prompts. Summary of the invention
[0005] The purpose of the present invention is to provide an integrated method based on unified expression and innovative fusion of design information, comprising the following steps:
[0006] 1) Obtain the original image-text pair dataset D = [D 1 ,D 2 ,D 3 ,…,D n ] and the text P entered by the user when performing text generation. Among them, the original image text pair D i Include description text i And the image generated by the text I i , i is the sequence number of the image-text pair, and n is the number of original image-text pairs.
[0007] 2) Based on the generated image I iWith description text P i The matching degree of the original image-text pair dataset D is filtered, and a subset of the original image-text pair dataset D is constructed. Among them, n 0 is a positive integer less than n.
[0008] 3) Based on the description text P in the subset D′ i ′ and the semantic relevance of the input text P, and select the image-text pair dataset with the highest relevance to the input text P from the subset D′ Among them, n 1 is less than n 0 A positive integer.
[0009] 4) Build a large language model and text graph model.
[0010] 5) The image text with the highest correlation is matched to the description text P in the dataset D″ i ″ and the input text P are sent to the large language model to generate the fused text P′.
[0011] 6) Send the fused text P′ into the text-generated graph model to generate an image dataset GI = [GI 1 ,GI 2 ,…,GI k ], k is the number of generated images.
[0012] 7) According to the semantic similarity between the generated image dataset GI and the text P′, the generated image dataset GI is divided into qualified image sets GI q and the unqualified image set GI nq .
[0013] 8) The user selects a qualified image set GI q Select a relatively satisfactory image I like , and generate a new text P next .
[0014] 9) Let the new text P next The text P as input returns to step 3) until the user obtains a satisfactory image.
[0015] Furthermore, the element D i It also includes the hyperparameter H for text generation images i ′.
[0016] The hyperparameter H i ' includes image size, random seed, and number of iterations.
[0017] Further, the generated image I i With description text P iThe matching degree of the original image-text pair dataset D is filtered, and a subset of the original image-text pair dataset D is constructed. The steps are as follows:
[0018] 2.1) Calculate the random sample size as follows:
[0019]
[0020] Where n s is the sample size randomly drawn, Z is the standard normal distribution score with a% confidence level, a is a positive integer. P is the expected proportion. E is the error limit.
[0021] 2.2) Based on the randomly selected sample size, randomly sample n from the original image-text pair dataset D s Sample data as sample set
[0022] 2.3) Construct a multimodal large model CLIP and DSG evaluation framework.
[0023] 2.4) The sample set D s Medium Element Description text and generate images Input the multimodal large model CLIP to obtain the corresponding semantic vector and visual vector
[0024] The multimodal large model CLIP outputs the calculated Clip score
[0025] 2.5) The sample set D s Medium Element Description text Enter the DSG assessment framework and calculate the DSG score
[0026] 2.6) Plot data based on Clip score and Dsg score and determine Clip threshold T clip and Dsg threshold T dsg .
[0027] 2.7) According to Clip threshold T clip and Dsg threshold T dsg , filter the original image-text pair dataset D, and construct a subset D′ of the original image-text pair dataset D.
[0028] Furthermore, the calculation formula of the Dsg score is as follows:
[0029]
[0030] Where i is the sequence number of the image-text pair. i is the Dsg score of the i-th image-text pair. j is the description text The number of questions generated after entering the DSG assessment framework. To describe the text The jth question generated after entering the DSG evaluation framework. score i To generate images A function to determine whether the condition is met.
[0031] The calculation formula of the Clip score is as follows:
[0032]
[0033] In the formula, Clip i Indicates the image I generated from the i-th image-text pair i and description text P i Clip score of . ‖·‖ is the Euclidean norm.
[0034] Further, a subset D′ of the original image-text pair dataset D is as follows:
[0035]
[0036] In the formula, i is the sequence number of the image-text pair, and reatin(i) is the filtering function. i is the Clip score of the i-th image-text pair. DSG i is the Dsg score of the i-th image-text pair. clip is the Clip threshold. dsg is the Dsg threshold.
[0037] Furthermore, the element D in the subset D′ i ′ also includes the semantic vector SV i ′.
[0038] Further, the description text P in the subset D′ i ′ and the semantic relevance of the input text P, and select the image-text pair dataset with the highest relevance to the input text P from the subset D′ The steps are as follows:
[0039] 3.1) Perform named entity recognition on the input text P and obtain the entity array of the input text P Among them, j 0 is the number of entity words in the input text P.
[0040] 3.2) Filter the subset D′ to obtain the data set D′ temp , as shown below:
[0041] D′ temp ={(I i ′,P i ′,SV i ′)∈D′|contains(P i ′,E)=True} (7)
[0042]
[0043] In the formula, contains(P i ′,E) is a Boolean function. I′ i , P i ′、SV i ′ are the generated pictures, description texts, and semantic vectors in subset D′ respectively. e is the entity word.
[0044] 3.3) Build a multimodal large model CLIP.
[0045] 3.4) Input the input text P into the multimodal large model CLIP to obtain the semantic encoding PV.
[0046] 3.5) Calculate the data set D′ temp The semantic vector SV′ in temp The semantic similarity with the semantic code PV is obtained by obtaining the semantic similarity score array S = [S 1 ,S 2 ,S 3 ,…,S m ], where m is the data set D′ temp size.
[0047] The calculation formula of the semantic similarity is as follows:
[0048]
[0049] In the formula, S(PV,SV′ temp ) is the semantic similarity. PV·SV′ temp is the semantic code PV and semantic vector SV′ temp is the dot product function of . ‖·‖ is the Euclidean norm.
[0050] 3.6) Sort the semantic similarity score array S in descending order and take the first n 1 The data corresponding to the elements constitute the image-text pair dataset D″.
[0051] Furthermore, the fused text P′ is as follows:
[0052] P′=fuse(Prompts) (10)
[0053]
[0054] In the formula, Prompts is the description text P i ″ and the set of input text P. fuse is the fusion function.
[0055] Furthermore, the qualified image set GI q and the unqualified image set GI nq As shown below:
[0056] GI q ={i∈ GI|qualify(i)=1} (12)
[0057] GI nq ={i∈ GI|qualify(i)=0} (13)
[0058]
[0059] Where i is the number of the image-text pair and qualify(i) is the screening function. are the Clip score set and Dsg score set between the generated image and the text P′ in the generated image dataset GI. clip is the Clip threshold. dsg is the Dsg threshold.
[0060] Further, the user selects from a qualified image set GI q Select a relatively satisfactory image I like , and generate a new text P next The steps are as follows:
[0061] 8.1) Build a large visual model.
[0062] 8.2) Select the image I like Input the visual model to obtain image I like The visual encoding is as follows:
[0063] IV like =encode(I like ) (15)
[0064] Where, IV like For image I like The visual encoding of , encode is the simulated image encoding function.
[0065] 8.3) According to visual coding IV likefeatures, generate new text P next , as shown below:
[0066] P next =GeneratePrompt(IV like ) (16)
[0067] Where GeneratePrompt is the conversion function.
[0068] The technical effect of the present invention is unquestionable. The present invention provides an interactive, iterative and efficient method to optimize the prompt of user input, thereby quickly improving the quality of generated images.
[0069] Based on the original text-image big model, the present invention realizes the screening and preprocessing of the original image-text pair data set by introducing CLIP and DSG scoring, thus forming a new high-quality data set. Furthermore, the method extracts the most relevant image-text pairs from the new data set according to semantic similarity based on the user input prompt, and sends them together with the user prompt into the predefined big model for fusion to generate a new optimized prompt. This optimized prompt is used to guide the image generation model, batch generate images, and further screen images based on semantic similarity. The user can select a satisfactory image and generate the next prompt based on it, thereby achieving efficient and high-precision control of the image generation process and supporting rapid iteration, fusion and innovation of creative content.
[0070] The invention has a reliable design and broad prospects, and has great application prospects in the fields of human-computer interaction and innovative design.
[0071] This paper proposes a research on the integration of unified expression of design information and innovative methods. The research adopts an iterative interactive prompt optimization method, which can dynamically adjust the expression of design information to better meet the needs of different scenarios. This method not only improves the accuracy and adaptability of information expression, but also stimulates design innovation, thereby achieving efficient integration of unified expression of design information and innovative methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 This is the overall framework diagram of the integrated approach based on unified expression and innovative integration of design information. DETAILED DESCRIPTION
[0073] The present invention is further described below in conjunction with the embodiments, but it should not be understood that the above subject matter of the present invention is limited to the following embodiments. Without departing from the above technical ideas of the present invention, various substitutions and changes are made according to the common technical knowledge and customary means in the art, which should all be included in the protection scope of the present invention.
[0074] Embodiment 1:
[0075] See also Figure 1 , an integrated method based on unified expression and innovative integration of design information, including the following steps:
[0076] 1) Obtain the original image-text pair dataset D = [D 1 ,D 2 ,D 3 ,…,D n ] and the text P entered by the user when performing text generation. Among them, the original image text pair D i Include description text i And the image generated by the text I i , i is the sequence number of the image-text pair, and n is the number of original image-text pairs.
[0077] 2) Based on the generated image I i With description text P i The matching degree of the original image-text pair dataset D is filtered, and a subset of the original image-text pair dataset D is constructed. Among them, n 0 is a positive integer less than n.
[0078] 3) Based on the description text P in the subset D′ i ′ and the semantic relevance of the input text P, and select the image-text pair dataset with the highest relevance to the input text P from the subset D′ Among them, n 1 is less than n 0 A positive integer.
[0079] 4) Build a large language model and text graph model.
[0080] 5) The image text with the highest correlation is matched to the description text P in the dataset D″ i ″ and the input text P are sent to the large language model to generate the fused text P′.
[0081] 6) Send the fused text P′ into the text-generated graph model to generate an image dataset GI = [GI 1 ,GI 2 ,…,GI k ], k is the number of generated images.
[0082] 7) According to the semantic similarity between the generated image dataset GI and the text P′, the generated image dataset GI is divided into qualified image sets GI q and the unqualified image set GI nq .
[0083] 8) The user selects a qualified image set GI q Select a relatively satisfactory image I like , and generate a new text P next .
[0084] 9) Let the new text P next The text P as input returns to step 3) until the user obtains a satisfactory image.
[0085] Embodiment 2:
[0086] Based on the integrated method of unified expression and innovative integration of design information, the main technical content is shown in Example 1. Further, the element D i It also includes the hyperparameter H for text generation images i ′.
[0087] The hyperparameter H i ' includes image size, random seed, and number of iterations.
[0088] Embodiment 3:
[0089] The integrated method based on unified expression and innovative integration of design information, the main technical content of which is shown in any one of Embodiments 1 to 2, further, the method based on generating image I i With description text P i The matching degree of the original image-text pair dataset D is filtered, and a subset of the original image-text pair dataset D is constructed. The steps are as follows:
[0090] 2.1) Calculate the random sample size as follows:
[0091]
[0092] Where n s is the sample size randomly drawn, Z is the standard normal distribution score with a% confidence level, a is a positive integer. P is the expected proportion. E is the error limit.
[0093] 2.2) Based on the randomly selected sample size, randomly sample n from the original image-text pair dataset D s Sample data as sample set
[0094] 2.3) Construct a multimodal large model CLIP and DSG evaluation framework.
[0095] 2.4) The sample set D s Medium Element Description text and generate images Input the multimodal large model CLIP to obtain the corresponding semantic vector and visual vector
[0096] The multimodal large model CLIP outputs the calculated Clip score
[0097] 2.5) The sample set D s Medium Element Description text Enter the DSG assessment framework and calculate the DSG score
[0098] 2.6) Plot data based on Clip score and Dsg score and determine Clip threshold T clip and Dsg threshold T dsg .
[0099] 2.7) According to Clip threshold T clip and Dsg threshold T dsg , filter the original image-text pair dataset D, and construct a subset D′ of the original image-text pair dataset D.
[0100] Embodiment 4:
[0101] An integrated method based on unified expression and innovative integration of design information, the main technical content of which is shown in any one of Embodiments 1 to 3. Further, the calculation formula of the Dsg score is as follows:
[0102]
[0103] Where i is the sequence number of the image-text pair. i is the Dsg score of the i-th image-text pair. j is the description text The number of questions generated after entering the DSG assessment framework. To describe the text The jth question generated after entering the DSG evaluation framework. score i To generate images A function to determine whether the condition is met.
[0104] The calculation formula of the Clip score is as follows:
[0105]
[0106] In the formula, Clip i Indicates the image I generated from the i-th image-text pair i and description text P i Clip score of . ‖·‖ is the Euclidean norm.
[0107] Embodiment 5:
[0108] The integrated method based on unified expression and innovative fusion of design information, the main technical content of which is shown in any one of Embodiments 1 to 4, further, the subset D′ of the original image text pair dataset D is as follows:
[0109]
[0110]
[0111] In the formula, i is the sequence number of the image-text pair, and reatin(i) is the filtering function. i is the Clip score of the i-th image-text pair. DSG i is the Dsg score of the i-th image-text pair. clip is the Clip threshold. dsg is the Dsg threshold.
[0112] Embodiment 6:
[0113] An integrated method based on unified expression and innovative integration of design information, the main technical content of which is shown in any one of Embodiments 1 to 5. Further, the element D in the subset D′ i ′ also includes the semantic vector SV i ′.
[0114] Embodiment 7:
[0115] The integrated method based on unified expression and innovative integration of design information, the main technical content of which is shown in any one of Embodiments 1 to 6, further, the description text P in the subset D′ i ′ and the semantic relevance of the input text P, and select the image-text pair dataset with the highest relevance to the input text P from the subset D′ The steps are as follows:
[0116] 3.1) Perform named entity recognition on the input text P and obtain the entity array of the input text P Among them, j 0 is the number of entity words in the input text P.
[0117] 3.2) Filter the subset D′ to obtain the data set D′ temp , as shown below:
[0118] D′ temp ={(I i ′,P i ′,SV i ′)∈D′|contains(P i ′,E)=True} (7)
[0119]
[0120] In the formula, contains(P i ′,E) is a Boolean function. i ′、P i ′、SV i ′ are the generated pictures, description texts, and semantic vectors in subset D′ respectively. e is the entity word.
[0121] 3.3) Build a multimodal large model CLIP.
[0122] 3.4) Input the input text P into the multimodal large model CLIP to obtain the semantic encoding PV.
[0123] 3.5) Calculate the data set D′ temp The semantic vector SV in t ' emp The semantic similarity with the semantic code PV is obtained by obtaining the semantic similarity score array S = [S 1 ,S 2 ,S 3 ,…,S m ], where m is the data set D′ temp size.
[0124] The calculation formula of the semantic similarity is as follows:
[0125]
[0126] In the formula, S(PV,SV t ' emp ) is the semantic similarity. PV·SV t ' emp is the semantic code PV and semantic vector SV t ' emp is the dot product function of . ‖·‖ is the Euclidean norm.
[0127] 3.6) Sort the semantic similarity score array S in descending order and take the first n 1 The data corresponding to the elements constitute the image-text pair dataset D″.
[0128] Embodiment 8:
[0129] The integrated method based on unified expression and innovative fusion of design information, the main technical content of which is shown in any one of Embodiments 1 to 7, further, the fused text P′ is as follows:
[0130] P′=fuse(Prompts) (10)
[0131]
[0132] In the formula, Prompts is the description text P i ″′ and the set of input text P. fuse is the fusion function.
[0133] Embodiment 9:
[0134] Based on the integrated method of unified expression and innovative fusion of design information, the main technical content is shown in any one of embodiments 1 to 8. Further, the qualified image set GI q and the unqualified image set GI nq As shown below:
[0135] GI q ={i∈ GI|qualify(i)=1} (12)
[0136] GI nq ={i∈ GI|qualify(i)=0} (13)
[0137]
[0138] Where i is the number of the image-text pair and qualify(i) is the screening function. are the Clip score set and Dsg score set between the generated image and the text P′ in the generated image dataset GI. clip is the Clip threshold. dsg is the Dsg threshold.
[0139] Embodiment 10:
[0140] Based on the integrated method of unified expression and innovative fusion of design information, the main technical content is shown in any one of embodiments 1 to 9. Further, the user selects a qualified image set GI q Select a relatively satisfactory image I like , and generate a new text P next The steps are as follows:
[0141] 8.1) Build a large visual model.
[0142] 8.2) Select the image I like Input the visual model to obtain image I like The visual encoding is as follows:
[0143] IV like =encode(I like ) (15)
[0144] Where, IV like For image I likeThe visual encoding of , encode is the simulated image encoding function.
[0145] 8.3) According to visual coding IV like features, generate new text P next , as shown below:
[0146] P next =GeneratePrompt(IV like ) (16)
[0147] Where GeneratePrompt is the conversion function.
[0148] Embodiment 11:
[0149] See also Figure 1 , an integrated method based on unified expression and innovative integration of design information, including the following steps:
[0150] 1) Obtain the original image-text pair dataset D = [D 1 ,D 2 ,D 3 ,…,D n ] and the text P entered by the user when performing text generation, for example, "generate a car". i Include description text i And the image generated by the text I i , i is the sequence number of the image-text pair, and n is the number of original image-text pairs.
[0151] The original image-text pair dataset is Diffusion DB, which is the first large-scale text-to-image prompt dataset. The dataset has two versions. This method uses the 2M version, which contains 2 million data items. Each data item consists of three parts: prompt, hyperparameters, and generated images.
[0152] 2) Based on the generated image I i With description text P i The matching degree of the original image-text pair dataset D is filtered, and a subset of the original image-text pair dataset D is constructed. Among them, n 0 is a positive integer less than n.
[0153] 3) Based on the description text P in the subset D′ i ′ and the semantic relevance of the input text P, and select the image-text pair dataset with the highest relevance to the input text P from the subset D′ Among them, n 1 The value is 5.
[0154] 4) Build a large language model and text graph model.
[0155] 5) The image text with the highest correlation is matched to the description text P in the dataset D″ i ″ and the input text P are sent to the large language model to generate the fused text P′.
[0156] 6) Send the fused text P′ into the text-generated graph model to generate an image dataset GI = [GI 1 ,GI 2 ,…,GI k ], k is the number of generated images.
[0157] 7) According to the semantic similarity between the generated image dataset GI and the text P′, the generated image dataset GI is divided into qualified image sets GI q and the unqualified image set GI nq .
[0158] 8) The user selects a qualified image set GI q Select a relatively satisfactory image I like , and generate a new text P next .
[0159] The relatively satisfactory image I like It is to judge which picture among a bunch of pictures is the best and most pleasing to the eye from a subjective perspective, but it is certain that the results generated in the first few times are often flawed, so the word "relatively satisfactory" is used to describe this image.
[0160] 9) Let the new text P next The text P as input returns to step 3) until the user obtains a satisfactory image.
[0161] The satisfactory image is a subjective standard. If the user feels satisfied, no further iterative modification is required.
[0162] Embodiment 12:
[0163] Based on the integrated method of unified expression and innovative integration of design information, the main technical content is shown in Example 11. Further, the element D i It also includes the hyperparameter H for text generation images i ′.
[0164] The hyperparameter H i ' includes image size, random seed, and number of iterations.
[0165] Hyperparameters refer to each piece of data in the dataset D. In addition to the image-text pair, it also contains the corresponding input hyperparameters for generating images, such as size, temperature, random seed, etc.
[0166] Embodiment 13:
[0167] The integrated method based on unified expression and innovative integration of design information, the main technical content of which is shown in any one of Embodiments 11 to 12, further, the method based on generating image I i With description text P i The matching degree of the original image-text pair dataset D is filtered, and a subset of the original image-text pair dataset D is constructed. The steps are as follows:
[0168] 2.1) Calculate the random sample size as follows:
[0169]
[0170] Where n s is the sample size randomly drawn, with a value of 10,000, Z is the standard normal distribution score at a 95% confidence level, and 1.96 is the Z value at a 95% confidence level. P is the expected proportion, usually 0.5. E is the error limit, with a value of 0.01.
[0171] 2.2) Based on the randomly selected sample size, randomly sample n from the original image-text pair dataset D s Sample data as sample set
[0172] 2.3) Construct a multimodal large model CLIP and DSG evaluation framework.
[0173] For multimodal large model CLIP:
[0174] Assume we have m pairs of image-text pairs, and the image encoding is represented as I m , the text encoding is represented by T m , then the similarity matrix S between image-text pairs (Clip score) is:
[0175]
[0176] The element s of S ij represents the semantic similarity between the i-th image and the j-th text.
[0177] For the evaluation framework DSG:
[0178] DSG (Davidsonian Scene Graph) is a structured approach for evaluating the performance of text-to-image generation models. The DSG framework works in the following ways:
[0179] a) Structured representation: DSG decomposes text descriptions into structured atomic propositions that cover entities, attributes, relations, and global features.
[0180] b) Dependency graph: DSG uses directed acyclic graph (DAG) to represent the dependency relationships between these atomic propositions, ensuring the logical consistency of the problem.
[0181] c) Question Generation: DSG generates specific natural language questions from atomic propositions, which can be used by Visual Question Answering (VQA) models to evaluate generated images.
[0182] d) Evaluation: DSG evaluates the performance of image generation models by comparing the consistency between the answers given by the VQA model based on generated images and the expected answers based on text prompts.
[0183] The core of the DSG framework is that it combines natural language processing and computer vision techniques, especially leveraging pre-trained large language models (LLMs) and visual question answering models (VQA) to provide a fine-grained and interpretable way to evaluate the accuracy and reliability of text-to-image generation models.
[0184] 2.4) The sample set D s Medium Element Description text and generate images Input the multimodal large model CLIP to obtain the corresponding semantic vector and visual vector
[0185] The multimodal large model CLIP outputs the calculated Clip score
[0186] 2.5) The sample set D s Medium Element Description text Enter the DSG assessment framework and calculate the DSG score
[0187] 2.6) Plot data based on Clip score and Dsg score and determine Clip threshold T clip and Dsg threshold T dsg .
[0188] 2.7) According to Clip threshold T clip and Dsg threshold T dsg , filter the original image-text pair dataset D, and construct a subset D′ of the original image-text pair dataset D.
[0189] Embodiment 14:
[0190] An integrated method based on unified expression and innovative integration of design information, the main technical content of which is shown in any one of Examples 11 to 13. Further, the calculation formula of the Dsg score is as follows:
[0191]
[0192] Where i is the sequence number of the image-text pair. i is the Dsg score of the i-th image-text pair. j is the description text The number of questions generated after entering the DSG assessment framework. To describe the text The jth question generated after entering the DSG evaluation framework, such as is there a bicycle in the picture, is the bicycle red, is the bicycle leaning against the wall. i To generate images A function to determine whether the condition is met.
[0193] The calculation formula of the Clip score is as follows:
[0194]
[0195] In the formula, Clip i Indicates the image I generated from the i-th image-text pair i and description text P i Clip score of . ‖·‖ is the Euclidean norm.
[0196] Embodiment 15:
[0197] The integrated method based on unified expression and innovative fusion of design information, the main technical content of which is shown in any one of Embodiments 11 to 14, further, the subset D′ of the original image text pair dataset D is as follows:
[0198]
[0199] In the formula, i is the sequence number of the image-text pair, and reatin(i) is the filtering function. i is the Clip score of the i-th image-text pair. DSG i is the Dsg score of the i-th image-text pair. clip is the Clip threshold. dsg is the Dsg threshold.
[0200] Embodiment 16:
[0201] An integrated method based on unified expression and innovative integration of design information, the main technical content of which is shown in any one of Embodiments 11 to 15. Further, the element D in the subset D′ i′ also includes the semantic vector SV i ′.
[0202] Embodiment 17:
[0203] The integrated method based on unified expression and innovative integration of design information, the main technical content of which is shown in any one of Embodiments 11 to 16, further, the description text P in the subset D′ i ′ and the semantic relevance of the input text P, and select the image-text pair dataset with the highest relevance to the input text P from the subset D′ The steps are as follows:
[0204] 3.1) Perform named entity recognition on the input text P and obtain the entity array of the input text P Among them, j 0 is the number of entity words in the input text P.
[0205] 3.2) Filter the subset D′ to obtain the data set D′ temp , as shown below:
[0206] D′ temp ={(I i ′,P i ′,SV i ′)∈D′|contains(P i ′,E)=True} (7)
[0207]
[0208] In the formula, contains(P i ′,E) is a Boolean function. i ′、P i ′、SV i ′ are the generated pictures, description texts, and semantic vectors in subset D′ respectively. e is the entity word.
[0209] 3.3) Build a multimodal large model CLIP.
[0210] 3.4) Input the input text P into the multimodal large model CLIP to obtain the semantic encoding PV.
[0211] 3.5) Calculate the data set D′ temp The semantic vector SV in t ' emp The semantic similarity with the semantic code PV is obtained by obtaining the semantic similarity score array S = [S 1 ,S 2 ,S 3 ,…,S m ], where m is the data set D′ tempsize.
[0212] The calculation formula of the semantic similarity is as follows:
[0213]
[0214] In the formula, S(PV,SV t ' emp ) is the semantic similarity. PV·SV t ' emp is the semantic code PV and semantic vector SV t ' emp is the dot product function of . ‖·‖ is the Euclidean norm.
[0215] 3.6) Sort the semantic similarity score array S in descending order and take the first n 1 The data corresponding to the elements constitute the image-text pair dataset D "" .
[0216] Embodiment 18:
[0217] The integrated method based on unified expression and innovative fusion of design information, the main technical content of which is shown in any one of Embodiments 11 to 17, further, the fused text P′ is as follows:
[0218] P′=fuse(Prompts) (10)
[0219]
[0220] In the formula, Prompts is the description text P i "" and the set of input text P. fuse is the fusion function.
[0221] Embodiment 19:
[0222] An integrated method based on unified expression and innovative fusion of design information, the main technical content of which is shown in any one of Embodiments 11 to 18. Further, the qualified image set GI q and the unqualified image set GI nq As shown below:
[0223] GI q ={i∈ GI|qualify(i)=1} (12)
[0224] GI nq ={i∈ GI|qualify(i)=0} (13)
[0225]
[0226] Where i is the number of the image-text pair and qualify(i) is the screening function. are the Clip score set and Dsg score set between the generated image and the text P′ in the generated image dataset GI. clip is the Clip threshold. dsg is the Dsg threshold.
[0227] Embodiment 20:
[0228] Based on the integrated method of unified expression and innovative fusion of design information, the main technical content is shown in any one of embodiments 11 to 19. Further, the user selects a qualified image set GI q Select a relatively satisfactory image I like , and generate a new text P next The steps are as follows:
[0229] 8.1) Build a large visual model.
[0230] The visual macro model implements image coding using CLIP.
[0231] 8.2) Select the image I like Input the visual model and obtain image I like The visual encoding is as follows:
[0232] IV like =encode(I like ) (15)
[0233] Where, IV like For image I like The visual encoding is a 512-dimensional vector, and encode is a simulated image encoding function.
[0234] 8.3) According to visual coding IV like features, construct a text P in the vector space with high semantic relevance that can accurately describe the visual content of the image next , as shown below:
[0235] P next =GeneratePrompt(IV like ) (16)
[0236] Where GeneratePrompt is the conversion function.
[0237] Embodiment 21:
[0238] See also Figure 1 , based on the integrated method of unified expression and innovative integration of design information, the main technical contents include:
[0239] In the research on the integration of the expression of multimodal design information and innovative methods, achieving efficient information transfer and innovative generation is a key issue. To this end, a research on the integration of unified expression of design information and innovative methods is proposed. This research adopts the method of iterative interactive Prompt optimization, which can dynamically adjust the expression of design information to better meet the needs of different scenarios. Through this method, not only the accuracy and adaptability of information expression are improved, but also design innovation is stimulated, thus realizing the efficient integration of the unified expression of design information and innovative methods.
[0240] The integrated method based on the unified expression and innovation integration of design information includes the following steps:
[0241] 1) Obtain the dataset D = [D 1 , D 2 , D 3 , …, D n of original image-text pairs. n is the number of original image-text pairs.
[0242] 2) Screen the original dataset based on the matching degree between the image and the text. The specific operation is to evaluate the semantic relevance or consistency between each pair of images and the corresponding text, and then select the pairs that meet the criteria, that is, those pairs where the image content is closely related to the descriptive text. Construct a subset D′ = [D′ 1 , D′ 2 , D′ 3 , …, D′ i of the dataset D of original image-text pairs; i < n; the element D′ i consists of I i ′ (generated picture), P i ′ (prompt), H i ′ (hyperparameter), and SV i ′ (semantic vector).
[0243] 3) Obtain the original prompt P input by the user. P is the text input by the user when generating an image from text. For example, "Generate a car".
[0244] 4) Based on the semantic relevance between the prompt P′ in D ’ and the prompt P, screen and extract a set of image-text pairs D″ = [D″ 1 , D″ 2 , D″ 3 , …, D″ 5 from D′ that is most relevant to P.
[0245] 5) The prompt text in D″ and the user input promptP are sent together into the pre-fine-tuned large language model Qwen:32b to generate a new prompt P′.
[0246] 6) Send P′ to the Wensheng graph model to batch generate images GI = [GI 1 ,GI 2 ,GI,…,GI k ]; k is the number of generated images.
[0247] 7) According to the semantic similarity between the generated image GI and P′, GI is divided into GI q ,GI nq GI q Represents a qualified generated image, GI nq Indicates an unqualified generated image. Obviously, GI = GI q ∪GI nq ;
[0248] 8) The user selects a relatively satisfactory image I like PromptP used to generate the next step next .
[0249] 9) Repeat steps 4) to 8) until the user obtains a suitable prompt to generate a satisfactory image.
[0250] The original image-text pair dataset is Diffusion DB, which is the first large-scale text-to-image prompt dataset. This dataset has two versions. This method uses the 2M version, which has a total of 2 million data. Each data consists of three parts: prompt, hyperparameters, and generated images.
[0251] Construct a subset D′=[D′ 1 ,D′ 2 ,D′ 3 ,…,D′ I The steps of ] include:
[0252] 1) According to the basic sample size calculation formula In a sample with a total size of 2 million, in order to achieve a 95% confidence level and maintain a 1% error range, and in the absence of prior information about the proportion, the calculation results show that at least 9604 samples need to be randomly selected in order to reliably reflect the data characteristics of the population. Where:
[0253] n is the desired sample size
[0254] Z is the standard normal distribution score for the desired confidence level (e.g., 1.96 is the Z value for a 95% confidence level)
[0255] P is the expected proportion (if there is no prior information, 0.5 is often used as this gives the largest sample size)
[0256] E is the acceptable error margin (e.g., 0.01 or 1%)
[0257] 2) Based on the calculation results of 1), use the random sampling method to randomly sample 10,000 data from the original data set D as the sample set in Represents the selected i-th sample, (Generate image), (prompt), (Hyperparameter) composition.
[0258] 3) For sample set D s Traverse and semantically encode the image-text pairs of the samples in turn. The encoding method is to convert I s and P s Send it to the pre-trained multimodal large model CLIP (Contrastive Language-Image Pre-training) to obtain the corresponding semantic vector SV = [SV 1 ,SV 2 ,SV 3 ,…,SV 10000 ] and visual vector IV = [IV 1 ,IV 2 ,IV 3 ,…,IV 10000 ].
[0259] 4) For sample set D s Traverse and calculate the Clip score and Dsg score of the sample image-text pairs in turn to evaluate their semantic similarity. For the Clip score, just replace the paired SV i and IV i It can be calculated by sending it into the CLIP large model. For the Dsg score, first we need to Feed it into the pre-trained DSG evaluation framework to get a series of questions j is the number of generated questions. For example, if For "a red bicycle leaning against the wall", then the generated series of questions Q i It could be as follows:
[0260] Is there a bicycle in the picture?
[0261] Is the bicycle red?
[0262] Is this bicycle leaning against a wall?
[0263] The Dsg score calculation formula for the i-th sample is as follows:
[0264]
[0265] The calculated score is Clip = [Clip 1 ,Clip 2 ,Clip 3 ,…,Clip 10000 ] and DSG = [DSG 1 ,DSG 2 ,DSG 3 ,…,DSG 10000 ].
[0266] 5) Plot the Clip data and DSG data. After comprehensively considering the number and quality of the screened images, the screening thresholds of the two indicators are determined as T clip ,T dsg .
[0267] 6) According to the established screening threshold, the original data set D is traversed, and samples are retained only if both indicators are greater than the specified threshold, otherwise they will be discarded. The remaining samples constitute D′. The formula is as follows:
[0268]
[0269] Where i is any sample in the original dataset D, Clip i and DSG i is the Clip score and Dsg score of the sample. Then the filtered data set D′ can be expressed as a set of samples that meet all retention conditions:
[0270]
[0271] Note that at this time, there is an additional attribute SV in D′, that is, D′ consists of I′ (generated image), P′ (prompt), H′ (hyperparameter), and SV′ (semantic vector).
[0272] The method of screening and extracting a set of image-text pairs D″=[D″] that are most relevant to P from D′ 1 ,D″ 2 ,D″ 3 ,…,D″ 5 The steps of ] include:
[0273] 1) Perform named entity recognition on P and obtain the entity array E of P = [E 1 ,E2 ,E 3 ,…,E j ], j is the number of entity words in P.
[0274] 2) Filter D′ and retain only those samples whose prompt contains any element in E, and obtain the data set D′ temp The screening process is as follows:
[0275] Define a Boolean function:
[0276]
[0277] Where P is the prompt of the sample in D′.
[0278] Construct a new dataset D′ temp , contains all samples that meet the conditions:
[0279] D′ temp ={(I,P,SV)∈D′|contains(P,E)=True}
[0280] 3) Send P to CLIP to obtain its semantic code PV and traverse D′ temp , calculate the semantic similarity between PV and each sample SV in turn, and get the semantic similarity score array S = [S 1 ,S 2 ,S 3 ,…,S m ], m is the data set D′ temp Size. The formula for calculating semantic similarity is:
[0281]
[0282] Where:
[0283] PV·SV is the dot product of vectors PV and SV
[0284] ‖PV‖ and ‖SV‖ are the Euclidean norms of vectors PV and SV (i.e., the lengths of the vectors)
[0285] 4) Sort S in descending order and take the samples corresponding to the first five elements as D″=[D″ 1 ,D″ 2 ,D″ 3 ,…,D″ 5 ].
[0286] The process of fusing the prompt text in D″ and the user input promptP to generate a new prompt P′ is as follows:
[0287] 1) Combine all prompt texts in D″ with prompt P entered by the user. Define a set Prompts to contain these texts:
[0288] Prompts = [P 1 ,P 2 ,P 3 ,..,P 5 ,P]
[0289] 2) Input the set of prompts into the pre-fine-tuned large language model for fusion to obtain P′. This process can be expressed by a fusion function fuse:
[0290] P′=fuse(Prompts)#(5)
[0291] In the formula, fuse represents the encoding fusion process of prompts in the large model.
[0292] According to the semantic similarity between the generated image GI and the generated text P′, GI is divided into GI q ,GI nq Here are the steps:
[0293] 1) Traverse GI, calculate the Clip score and Dsg score between each generated image and P′, and get the set Clip generated and DSG generated .
[0294] 2) Traverse the collection Clip generated and Dsg generated , each image must simultaneously meet the preset Clip and Dsg score thresholds T clip and T dsg In this way, we can clearly distinguish the images that meet the quality standards (qualified image set GI q ) and images that fail to meet the standards (GI nq ). The formula is as follows:
[0295] Define a filter function:
[0296]
[0297] Define two generated image sets:
[0298] GI q ={i∈GI|qualify(i)=1}
[0299] GI nq ={i∈GI|qualify(i)=0}
[0300] The image I selected by the user like Generate the next prompt P next The steps are as follows:
[0301] 1) Input the image I selected by the user like into a pre-defined visual large model to obtain the visual encoding of the image:
[0302] IV like = encode(I like )#(7)
[0303] In the formula, the encode function is used to simulate the image encoding process.
[0304] 2) The large model uses its powerful natural language generation ability to construct a prompt P with a high semantic relevance and capable of accurately describing the visual content of the image in the vector space according to the features of IV like : next
[0305] P next = GeneratePrompt(I like )#(8)
[0306] In the formula, GeneratePrompt is a function representing the processing process of the large model, which converts the input image into a descriptive text highly similar to the image.
[0307] Example 22:
[0308] See Figure 1 , the integrated method based on the unified expression and innovative integration of design information, the main technical contents include:
[0309] 1) Obtain the original image-text pair dataset D = [D 1 , D 2 , D 3 , …, D n ; n is the number of original image-text pairs;
[0310] 2) Screen the original dataset based on the matching degree of the image-text pairs to construct a subset D' = [D' 1 , D' 2 , D' 3 , …, D' i ; i < n; the element D' i contains an image-text pair and hyperparameters, as well as the semantic encoding corresponding to the text;
[0311] 3) Obtain the original prompt P input by the user;
[0312] 4) Based on the semantic relevance to promptP, select and extract from D′ a set of image-text pairs D″=[D″ 1 ,D″ 2 ,D″ 3 ,…,D″ 5 ];
[0313] 5) Send the prompt text in D″ and the user input promptP together into the pre-fine-tuned large language model to generate a new prompt P′;
[0314] 6) Send P′ to the Wensheng graph model to batch generate images GI = [GI 1 ,GI 2 ,GI,…,GI k ]; k is the number of generated images;
[0315] 7) According to the semantic similarity between the generated image GI and the generated text P′, GI is divided into GI q ,GI nq GI q Represents a qualified generated image, GI nq Indicates an unqualified generated image. Obviously, GI = GI q ∪GI nq ;
[0316] 8) The user selects a relatively satisfactory image I like PromptP used to generate the next step next ;
[0317] 9) Repeat steps 4) to 8) until the user obtains a suitable prompt to generate a satisfactory image;
[0318] The original image-text pair dataset is Diffusion DB, which is the first large-scale text-to-image prompt dataset. The dataset has two versions. This method uses the 2M version, which contains 2 million data items. Each data item consists of three parts: prompt, hyperparameters, and generated images.
[0319] Construct a subset D′=[D′ 1 ,D′ 2 ,D′ 3 ,…,D′ I The steps of ] include:
[0320] 1) According to the basic sample size calculation formula In a sample with a total size of 2 million, in order to achieve a 95% confidence level and maintain a 1% error range, and in the absence of prior information about the proportion, the calculation results show that at least 9604 samples need to be randomly selected in order to reliably reflect the data characteristics of the population. Where:
[0321] n is the desired sample size
[0322] Z is the standard normal distribution score for the desired confidence level (e.g., 1.96 is the Z value for a 95% confidence level)
[0323] P is the expected proportion (if there is no prior information, 0.5 is often used as this gives the largest sample size)
[0324] E is the acceptable error margin (e.g., 0.01 or 1%)
[0325] 2) Based on the calculation results of 1), use the random sampling method to randomly sample 10,000 data from the original data set D as the sample set in Represents the selected i-th sample, (Generate image), (prompt), (Hyperparameter) composition;
[0326] 3) For sample set D s Traverse and semantically encode the image-text pairs of the samples in turn. The encoding method is to convert I s and P s Send it to the pre-trained multimodal large model CLIP (Contrastive Language-Image Pre-training) to get the corresponding semantic vector SV = [SV 1 ,SV 2 ,SV 3 ,…,SV 10000 ] and visual vector IV = [IV 1 ,IV 2 ,IV 3 ,…,IV 10000 ];
[0327] 4) For sample set D s Traverse and calculate the Clip score and Dsg score of the sample image-text pairs in turn to evaluate their semantic similarity. For the Clip score, just replace the paired SV i and IV i It can be calculated by sending it into the CLIP large model. For the Dsg score, first we need to Feed it into the pre-trained DSG evaluation framework to get a series of questions j is the number of generated questions. For example, if For "a red bicycle leaning against the wall", then the generated series of questions Q i It could be as follows:
[0328] Is there a bicycle in the picture?
[0329] Is the bicycle red?
[0330] Is this bicycle leaning against a wall?
[0331] The Dsg score calculation formula for the i-th sample is as follows:
[0332]
[0333] The calculated score is Clip = [Clip 1 ,Clip 2 ,Clip 3 ,…,Clip 10000 ] and DSG = [DSG 1 ,DSG 2 ,DSG 3 ,…,DSG 10000 ];
[0334] 5) Plot the Clip data and DSG data. After comprehensively considering the number and quality of the screened images, the screening thresholds of the two indicators are determined as T clip ,T dsg ;
[0335] 6) According to the established screening threshold, the original data set D is traversed, and samples are retained only if both indicators are greater than the specified threshold, otherwise they will be discarded. The remaining samples constitute D′. The formula is as follows:
[0336]
[0337] Where i is any sample in the original dataset D, Clip i and DSG i is the Clip score and Dsg score of the sample. Then the filtered data set D′ can be expressed as a set of samples that meet all retention conditions:
[0338]
[0339] Note that at this time, D′ has an additional attribute SV, that is, D′ consists of I′ (generated image), P′ (prompt), H′ (hyperparameter), and SV′ (semantic vector);
[0340] Filter and extract a set of image-text pairs D″=[D″] that are most relevant to P from D′ 1 ,D″ 2 ,D″ 3 ,…,D″ 5 The steps of ] include:
[0341] 1) Perform named entity recognition on P and obtain the entity array E of P = [E 1 ,E 2 ,E 3 ,…,E j ], j is the number of entity words in P;
[0342] 2) Filter D′ and retain only those samples whose prompt contains any element in E, and obtain the data set D′ temp The screening process is as follows:
[0343] Define a Boolean function:
[0344]
[0345] Where P is the prompt of the sample in D′.
[0346] Construct a new dataset D′ temp , contains all samples that meet the conditions:
[0347] D′ temp ={(I,P,SV)∈D′|contains(P,E)=True}
[0348] 3) Send P to CLIP to obtain its semantic code PV and traverse D′ temp , calculate the semantic similarity between PV and each sample SV in turn, and get the semantic similarity score array S = [S 1 ,S 2 ,S 3 ,…,S m ], m is the data set D′ temp Size. The formula for calculating semantic similarity is:
[0349]
[0350] Where:
[0351] PV·SV is the dot product of vectors PV and SV
[0352] ‖PV‖ and ‖SV‖ are the Euclidean norms of vectors PV and SV (i.e., the lengths of the vectors)
[0353] 4) Sort S in descending order and take the samples corresponding to the first five elements as D″=[D″ 1 ,D″2 ,D″ 3 ,…,D″ 5 ];
[0354] The process of fusing the prompt text in D″ and the user input promptP to generate a new prompt P′ is as follows:
[0355] 1) Combine all prompt texts in D″ with prompt P entered by the user. Define a set Prompts to contain these texts:
[0356] Prompts = [P 1 ,P 2 ,P 3 ,..,P 5 ,P]
[0357] 2) Input the set of prompts into the pre-fine-tuned large language model for fusion to obtain P′. This process can be expressed by a fusion function fuse:
[0358] P′=fuse(Prompts)#(5)
[0359] In the formula, fuse represents the encoding fusion process of prompts in the large model;
[0360] According to the semantic similarity between the generated image GI and the generated text P′, GI is divided into GI q ,GI nq Here are the steps:
[0361] 1) Traverse GI, calculate the Clip score and Dsg score between each generated image and P′, and get the set Clip generated and DSG generated ;
[0362] 2) Traverse the collection Clip generated and Dsg generated , each image must simultaneously meet the preset Clip and Dsg score thresholds T clip and T dsg In this way, we can clearly distinguish the images that meet the quality standards (qualified image set GI q ) and images that fail to meet the standards (GI nq ). The formula is as follows:
[0363] Define a filter function:
[0364]
[0365] Define two generated image sets:
[0366] GI q ={i∈GI|qualify(i)=1}
[0367] GI nq ={i∈GI|qualify(i)=0}
[0368] Based on the image selected by the user like Generate the next step prompt P next The steps are as follows:
[0369] 1) The image I selected by the user like Input into the predefined visual model to obtain the visual encoding of the image:
[0370] IV like =encode(I like )#(7)
[0371] The encode function is used to simulate the image encoding process.
[0372] 2) The large model uses its powerful natural language generation capabilities to like features, construct a prompt P in the vector space with high semantic relevance that can accurately describe the visual content of the image bext :
[0373] P next =GeneratePrompt(IV like )#(8)
[0374] Where GeneratePrompt is a function that represents the processing of the large model.
Claims
1. An integrated method based on unified expression and fusion of design information, characterized in that: The following steps are involved: 1) Get the original image-text pair dataset D = [D1, D2, D3, ..., D n ] and the text P entered by the user when performing text generation; among them, the original image text pair D i Include description text i And the image generated by the text I i , i is the sequence number of the image-text pair, n is the number of original image-text pairs; 2) Based on the generated image I i With description text P i The matching degree of the original image-text pair dataset D is filtered, and a subset of the original image-text pair dataset D is constructed. Wherein, n0 is a positive integer less than n. 3) Based on the description text P in the subset D′ i ′ and the semantic relevance of the input text P, and select the image-text pair dataset with the highest relevance to the input text P from the subset D′ Wherein, n1 is a positive integer less than n0. 4) Build a large language model and text graph model; 5) The image text with the highest correlation is matched to the description text P in the dataset D″ i ″ and the input text P are sent to the large language model to generate the fused text P′; 6) The fused text P′ is fed into the text-generated graph model to generate an image dataset GI = [GI1, GI2, …, GI k ], k is the number of generated images; 7) According to the semantic similarity between the generated image dataset GI and the text P′, the generated image dataset GI is divided into qualified image sets GI q and the unqualified image set GI nq ; 8) The user selects a qualified image set GI q Select a relatively satisfactory image I like , and generate a new text P next ; 9) Let the new text P next The text P as input returns to step 3) until the user obtains a satisfactory image.
2. The integrated method based on unified expression and innovative integration of design information according to claim 1 is characterized in that: The element D i It also includes the hyperparameter H for text generation images i ′; The hyperparameter H i ' includes image size, random seed, and number of iterations.
3. The integrated method based on unified expression and innovative integration of design information according to claim 1 is characterized in that: The generated image I i With description text P i The matching degree of the original image-text pair dataset D is filtered, and a subset of the original image-text pair dataset D is constructed. The steps are as follows: 2.1) Calculate the random sample size as follows: Where n s is the sample size randomly drawn, Z is the standard normal distribution score with a% confidence level, a is a positive integer; P is the expected proportion; E is the error limit; 2.2) Based on the randomly selected sample size, randomly sample n from the original image-text pair dataset D s Sample data as sample set 2.3) Construct a multimodal large model CLIP and DSG evaluation framework; 2.4) The sample set D s Medium Element Description text and generate images Input the multimodal large model CLIP to obtain the corresponding semantic vector and visual vector The multimodal large model CLIP outputs the calculated Clip score 2.5) The sample set D s Medium Element Description text Enter the DSG assessment framework and calculate the DSG score 2.6) Plot data based on Clip score and Dsg score and determine Clip threshold T clip and Dsg threshold T dsg ; 2.7) According to Clip threshold T clip and Dsg threshold T dsg , filter the original image-text pair dataset D, and construct a subset D′ of the original image-text pair dataset D.
4. The integrated method based on unified expression and innovative integration of design information according to claim 3 is characterized in that: The calculation formula of the Dsg score is as follows: Where i is the sequence number of the image-text pair; DSG i is the Dsg score of the i-th image-text pair; j is the description text The number of questions generated after entering the DSG assessment framework; To describe the text The jth question generated after entering the DSG evaluation framework; score i To generate images Whether the condition is met or not; The calculation formula of the Clip score is as follows: In the formula, Clip i Indicates the image I generated from the i-th image-text pair i and description text P i Clip score; ‖·‖ is the Euclidean norm.
5. The integrated method based on unified expression and innovative integration of design information according to claim 3 is characterized in that: The subset D′ of the original image-text pair dataset D is as follows: Where i is the image-text pair number, reatin(i) is the filtering function; Clip i is the Clip score of the i-th image-text pair; DSG i is the Dsg score of the i-th image-text pair; T clip is the Clip threshold; T dsg is the Dsg threshold.
6. The integrated method based on unified expression and innovative integration of design information according to claim 3 is characterized in that: The element D' in the subset D' i It also includes semantic vector SV i ′.
7. The integrated method based on unified expression and innovative integration of design information according to claim 1 is characterized in that: The description text P in the subset D′ i ′ and the semantic relevance of the input text P, and select the image-text pair dataset with the highest relevance to the input text P from the subset D′ The steps are as follows: 3.1) Perform named entity recognition on the input text P and obtain the entity array of the input text P Among them, j0 is the number of entity words in the input text P; 3.2) Filter the subset D′ to obtain the data set D′ temp , as shown below: D′ temp ={(I i ′,P i ′,SV i ′)∈D′|contains(P i ′,E)=True} (7) In the formula, contains(P i ′,E) is a Boolean function; I i ′、P i ′、SV i ′ are the generated pictures, description texts, and semantic vectors in subset D′ respectively; e is the entity word; 3.3) Construct a multimodal large model CLIP; 3.4) Input the input text P into the multimodal large model CLIP to obtain the semantic encoding PV; 3.5) Calculate the data set D′ temp The semantic vector SV′ in temp The semantic similarity with the semantic encoding PV is obtained to obtain the semantic similarity score array S = [S1, S2, S3, ..., S m ], where m is the data set D′ temp size; The calculation formula of the semantic similarity is as follows: In the formula, S(PV,SV′ temp ) is the semantic similarity; PV·SV′ temp is the semantic code PV and semantic vector SV′ temp The dot product function of ; ‖·‖ is the Euclidean norm; 3.6) Sort the semantic similarity score array S in descending order, and take the data corresponding to the first n1 elements to form the image-text pair dataset D″.
8. The integrated method based on unified expression and innovative integration of design information according to claim 1 is characterized in that: The fused text P′ is as follows: P′=fuse(Prompts) (10) In the formula, Prompts is the description text P i ″ and the set of input text P; fuse is the fusion function.
9. The integrated method based on unified expression and innovative integration of design information according to claim 1, characterized in that: The qualified image set GI q and the unqualified image set GI nq As shown below: GI q ={i∈GI|qualify(i)=1} (12) GI nq ={i∈GI|qualify(i)=0} (13) Where i is the sequence number of the image-text pair; qualify(i) is the filtering function; are the Clip score set and Dsg score set between the generated image and the text P′ in the generated image dataset GI; T clip is the Clip threshold; T dsg is the Dsg threshold.
10. The integrated method based on unified expression and innovative integration of design information according to claim 1, characterized in that: The user selects from a qualified image set GI q Select a relatively satisfactory image I like , and generate a new text P next The steps are as follows: 8.1) Build a large visual model; 8.2) Select the image I like Input the visual model to obtain image I like The visual encoding is as follows: IV like =encode(I like ) (15) Where, IV like For image I like Visual encoding, encode is a simulated image encoding function; 8.3) According to visual coding IV like features, generate new text P next , as shown below: P next =GeneratePrompt(IV like ) (16) Where GeneratePrompt is the conversion function.
Citation Information
Patent Citations
Multimedia event extraction method based on multi-modal low-dimensional feature representation space
CN118762261A
Unsupervised cross-modal hash retrieval method based on CLIP and attention fusion mechanism
CN118861327A
Image-text data enhancement method, text-to-graph model training method and image generation method
CN118968214A