Integrated method based on unified expression and fusion of design information
By constructing large language models and text-to-image models, and combining the CLIP and DSG evaluation frameworks, the prompts of generative large models are filtered and optimized, solving the problem that novice users often encounter unexpected results in generative large models. This enables efficient and high-precision image generation and rapid iteration of creative content.
Patent Information
- Application Number
- CN202411734069.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-11-29
AI Technical Summary
In image creation, existing generative large models make it difficult for novice users to write accurate prompts, and existing methods fail to effectively consider users' personalized needs, resulting in unexpected generated results and a lack of subsequent optimization methods.
By constructing a large language model and a text-to-image model, the dataset of image-text pairs with the highest relevance to the user's input text is selected, the fused text is generated, and the image is generated. The CLIP and DSG evaluation frameworks are used for image selection and optimization. Users can iteratively select satisfactory images to generate new prompts, thus achieving efficient optimization of the prompts.
It improves the quality of generated images and enhances user control over the generation process, supporting rapid iteration and innovation of creative content, and achieving efficient and high-precision image generation.
Smart Images

Figure CN119963671B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of large models and human-computer interaction, specifically to an integrated method based on the unified expression and innovative fusion of design information. Background Technology
[0002] With the continuous development of generative large model technology, top generative models such as Stable Diffusion and DALL·E are now able to generate highly relevant and high-quality images based on text prompts, which has greatly lowered the threshold for image creation and has gradually become a popular and efficient creative tool.
[0003] Building on the significant achievements of these generative models, researchers and developers have further explored a human-computer interaction technique called "prompts." This technique allows users to create natural language prompts describing desired image features (such as theme and style) and adjust the model's hyperparameters during the creative process, thereby helping them achieve their desired generated results. However, for novice users, writing effective prompts that accurately optimize the model's output is a challenging task. This is because the complexity and ambiguity of natural language make it difficult to fully express the user's creative ideas. Furthermore, prompts based on different model hyperparameters can produce drastically different images. Under limited hyperparameter testing conditions, it is difficult to evaluate the quality of the prompts. When receiving image results that do not meet expectations, users may feel confused about whether and how to adjust the prompts or model hyperparameters.
[0004] Existing prompt optimization methods primarily focus on generating prompts quickly and automatically, a process typically completed end-to-end in a single step. However, this approach does not consider users' personalized needs and is not conducive to subsequent optimization and improvement of the prompts. Summary of the Invention
[0005] The purpose of this invention is to provide an integrated method based on the unified expression and innovative fusion of design information, comprising the following steps:
[0006] 1) Obtain the original image-text pair dataset D = [D1, D2, D3, ..., D n The original image text is the text P input by the user when creating the image. The original image text is the text pair D. i Including the descriptive text P i And images generated from text I i , where i is the image-text pair index and n is the original number of image-text pairs.
[0007] 2) Based on generated image I i With description text P iBased on the degree of matching, filter the original image-text pair dataset D, and construct a subset of the original image-text pair dataset D. Where n0 is a positive integer less than n.
[0008] 3) Based on the description text P′ in subset D′ i The semantic relevance of the image-text pairs to the input text P is used to select the image-text pairs from the subset D′ that have the highest relevance to the input text P. Where n1 is a positive integer less than n0.
[0009] 4) Construct a large language model and a text-to-graph model.
[0010] 5) P represents the descriptive text P in dataset D″, which has the highest relevance between the image and text pairs. i The input text P and the input text P are fed into a large language model to generate the fused text P′.
[0011] 6) Feed the fused text P′ into the text-to-image model to generate an image dataset GI = [GI1, GI2, ..., GI...]. k ], k is the number of images generated.
[0012] 7) Based on the semantic similarity between the generated image dataset GI and the text P′, the generated image dataset GI is divided into qualified image sets GI. q and unqualified image sets GI nq .
[0013] 8) Users from qualified image sets GI q Select a relatively satisfactory image I like And generate a new text P next .
[0014] 9) Let the new text P next The text P, which is the input, returns to step 3) until the user obtains a satisfactory image.
[0015] Furthermore, the original image text pair D i It also includes the hyperparameter H for generating images from text. i ′.
[0016] The hyperparameter H i This includes image size, random seed, and number of iterations.
[0017] Furthermore, the method based on generated image I i With description text P i Based on the degree of matching, filter the original image-text pair dataset D, and construct a subset of the original image-text pair dataset D. The steps are as follows:
[0018] 2.1) Calculate the random sample size, as shown below:
[0019]
[0020] where n s is the random sample size, Z is the standard normal distribution score of a% confidence level, a is a positive integer. P is the expected proportion. E is the error limit.
[0021] 2.2) Based on the random sample size, randomly sample n s sample data from the original image text pair dataset D as the sample set
[0022] 2.3) Construct a multi-modal large model CLIP and DSG evaluation framework.
[0023] 2.4) Input the description text s and the generated picture of the element in the sample set D into the multi-modal large model CLIP to obtain the corresponding semantic vector and visual vector
[0024] The multi-modal large model CLIP outputs the calculated Clip score
[0025] 2.5) Input the description text s of the element in the sample set D into the DSG evaluation framework to calculate the Dsg score
[0026] 2.6) Based on the Clip score and the Dsg score, draw a data plot, and determine the Clip threshold T clip and the Dsg threshold T dsg .
[0027] 2.7) According to the Clip threshold T clip and the Dsg threshold T dsg , filter the original image text pair dataset D to construct the subset D' of the original image text pair dataset D.
[0028] Further, the calculation formula of the Dsg score is as follows:
[0029]
[0030] where i is the image text pair number. DSG i is the Dsg score of the i-th image text pair. j is the description text Number of questions generated after inputting the DSG evaluation framework. To describe the text The jth question generated after inputting the DSG evaluation framework. score i To generate pictures Judgment function whether the condition is met.
[0031] The calculation formula of the Clip score is as follows:
[0032]
[0033] In the formula, Clip i represents the Clip score of the ith image-text pair, which generates the picture I i and the description text P i . ‖·‖ is the Euclidean norm.
[0034] Further, the subset D' of the original image-text pair dataset D is as follows:
[0035]
[0036] In the formula, i is the image-text pair number, and reatin(i) is the screening function. Clip i is the Clip score of the ith image-text pair. DSG i is the Dsg score of the ith image-text pair. T clip is the Clip threshold value. T dsg is the Dsg threshold value.
[0037] Further, the element D' in the subset D' includes a semantic vector SV i . i
[0038] Further, the step of screening the image-text pair dataset from the subset D' based on the semantic relevance of the description text P' in the subset D' to the input text P is as follows: i
[0039] 3.1) Perform named entity recognition on the input text P to obtain the entity array of the input text P where j0 is the number of entity words in the input text P.
[0040] 3.2) Screen the subset D' to obtain the dataset D' temp as follows:
[0041] D' temp = {(I' i , P'i ,SV′ i )∈D′|contains(P i ′,E)=True} (7)
[0042]
[0043] In the formula, contains(P′) i E) is a Boolean function. I′ i 、P′ i ,SV′ i These represent the generated image, descriptive text, and semantic vector in subset D′, respectively. e represents an entity word.
[0044] 3.3) Construct a multimodal large model CLIP.
[0045] 3.4) Input the input text P into the multimodal large model CLIP to obtain the semantic code PV.
[0046] 3.5) Calculate dataset D′ temp semantic vector SV′ in temp The semantic similarity scores are obtained by comparing the semantic similarity scores of the semantically encoded PV with those of the PV, resulting in a semantic similarity score array S = [S1, S2, S3, ..., S...]. m ], where m is the dataset D′ temp Size.
[0047] The formula for calculating semantic similarity is as follows:
[0048]
[0049] In the formula, S(PV,SV′) temp ) represents semantic similarity. PV·SV′ temp Semantic encoding PV and semantic vector SV′ temp The dot product function. ‖·‖ is the Euclidean norm.
[0050] 3.6) Sort the semantic similarity score array S in descending order, and take the data corresponding to the first n1 elements to form the image-text pair dataset D″.
[0051] Furthermore, the fused text P′ is shown below:
[0052] P′=fuse(Prompts) (10)
[0053]
[0054] In the formula, Prompts is the descriptive text P″ i The set of input text P. `fuse` is the fusion function.
[0055] Further, the qualified image set GI q and the unqualified image set GI nq As follows:
[0056] GI q = {i e GI | qualify(i) = 1} (12)
[0057] GI nq = {i e GI | qualify(i) = 0} (13)
[0058]
[0059] In the formula, i is the image text pair serial number. qualify(i) is the screening function. are the Clip score set and the Dsg score set between the generated image and the text P' in the generated image data set GI, respectively. clip is the Clip threshold value. dsg is the Dsg threshold value.
[0060] Further, the user selects a relatively satisfactory image I q from the qualified image set GI like , and generates a new text P next The steps are as follows:
[0061] 8.1) Construct a visual large model.
[0062] 8.2) Input the selected image I like into the visual large model to obtain the visual encoding of the image I like , as follows:
[0063] IV like = encode(I like ) (15)
[0064] In the formula, IV like is the visual encoding of the image I like , and encode is the analog image encoding function.
[0065] 8.3) According to the characteristics of the visual encoding IV like , a new text P next is generated, as follows:
[0066] P next = GeneratePrompt(IV like ) (16)
[0067] In the formula, GeneratePrompt is the conversion function.
[0068] The technical effect of the present application is self-evident. The present application provides an interactive, iterative and efficient method to optimize the user input prompt, thereby rapidly improving the quality of generated images.
[0069] The present application introduces CLIP and DSG scores on the basis of the original image-text pair large model, realizes the screening and preprocessing of the original image-text pair dataset, and forms a new high-quality dataset. Further, the method extracts the most relevant image-text pairs from the new dataset according to semantic similarity based on the user input prompt, and inputs them into the pre-defined large model together with the user prompt to generate a new optimized prompt. This optimized prompt is used to guide the image generation model to generate images in batches, and further screen the images according to semantic similarity. The user can select the satisfactory images to generate the next prompt, thereby realizing efficient and high-precision control of the image generation process, supporting rapid iteration and integration of creative content and innovation.
[0070] The present application is reliable in design and has broad prospects, and has great application prospects in the field of human-computer interaction and innovative design.
[0071] The present application proposes an integrated research based on unified expression of design information and innovation method integration, which adopts an iterative interactive prompt optimization method, can dynamically adjust the expression mode of design information to better adapt to the needs of different scenarios. Through this method, not only the accuracy and adaptability of information expression are improved, but also design innovation is stimulated, thereby realizing the unified expression of design information and the efficient integration of innovation methods. BRIEF DESCRIPTION OF DRAWINGS
[0072] Figure 1 FIG. 1 is a schematic diagram of the overall framework of the integrated method based on unified expression of design information and innovation integration. DETAILED DESCRIPTION
[0073] The present application will be further described below in conjunction with examples, but should not be understood as limiting the above-mentioned subject matter of the present application to the following examples. Various substitutions and modifications can be made according to ordinary technical knowledge and conventional means in the art without departing from the above-mentioned technical idea of the present application, and all should be included in the protection scope of the present application.
[0074] Example 1:
[0075] Referring to Figure 1 , the integrated method based on unified expression of design information and innovation integration includes the following steps:
[0076] 1) Obtain the original image-text pair dataset D = [D1, D2, D3, …, Dn], where n is the number of image-text pairs in the dataset. nand the text P input by the user when generating the picture. Wherein, the original image-text pair D i comprises the description text P i and the picture I generated by the text i , i is the serial number of the image-text pair, and n is the number of original image-text pairs.
[0077] 2) Based on the matching degree of the generated picture I i and the description text P i , the original image-text pair dataset D is screened to construct a subset D′ of the original image-text pair dataset D Wherein, n0 is a positive integer less than n.
[0078] 3) Based on the semantic relevance of the description text P i ′ in the subset D′ and the input text P, the image-text pair dataset D″ with the highest relevance to the input text P is screened from the subset D′ Wherein, n1 is a positive integer less than n0.
[0079] 4) Construct a large language model and a picture generation model.
[0080] 5) The description text P i ″ in the image-text pair dataset D″ with the highest relevance and the input text P are input into the large language model to generate the fused text P′.
[0081] 6) The fused text P′ is input into the picture generation model to generate the image dataset GI = [GI1, GI2, …, GI k ], and k is the number of generated images.
[0082] 7) According to the semantic similarity of the generated image dataset GI and the text P′, the generated image dataset GI is divided into the qualified image set GI q and the unqualified image set GI nq .
[0083] 8) The user selects the relatively satisfactory image I q from the qualified image set GI like , and generates a new text P next .
[0084] 9) Let the new text P next be the input text P and return to step 3) until the user obtains a satisfactory image.
[0085] Embodiment 2:
[0086] Based on the integrated method of unified expression and innovation fusion of design information, the main technical content is shown in embodiment 1, further, the original image-text pair D iAlso included is a hyperparameter H for generating pictures from text i ′.
[0087] The hyperparameter H i ′ includes image size, random seed, number of iteration steps.
[0088] Embodiment 3:
[0089] The integrated method based on the unified expression of design information and innovation fusion, the main technical content is any one of embodiments 1 to 2, further, the integrated method based on the generated picture I i The matching degree of the description text P i , filter the original image text pair dataset D, and construct a subset of the original image text pair dataset D The steps are as follows:
[0090] 2.1) Calculate the sample size of random sampling, as follows:
[0091]
[0092] In the formula, n s is the sample size of random sampling, Z is the standard normal distribution score of a% confidence level, a is a positive integer. P is the expected proportion. E is the error limit.
[0093] 2.2) Based on the sample size of random sampling, randomly sample n s sample data from the original image text pair dataset D as a sample set
[0094] 2.3) Construct a multi-modal large model CLIP and DSG evaluation framework.
[0095] 2.4) Input the description text s and the generated picture of the elements in the sample set D into the multi-modal large model CLIP to obtain the corresponding semantic vector and visual vector
[0096] The multi-modal large model CLIP outputs the calculated Clip score
[0097] 2.5) Input the description text s of the elements in the sample set D into the DSG evaluation framework to calculate the Dsg score
[0098] 2.6) Data plotting based on Clip score and Dsg score, and determining Clip threshold T clip and Dsg threshold T dsg .
[0099] 2.7) Screening original image-text pair dataset D according to Clip threshold T clip and Dsg threshold T dsg , to build subset D' of original image-text pair dataset D.
[0100] Embodiment 4:
[0101] The integrated method based on unified expression of design information and innovation fusion, the main technical content of which is any one of embodiments 1 to 3, further, the calculation formula of the Dsg score is as follows:
[0102]
[0103] In the formula, i is the image-text pair number. DSG i is the Dsg score of the i-th image-text pair. j is the description text number generated after inputting into the DSG evaluation framework. is the j-th question generated after inputting the description text into the DSG evaluation framework. score i is the judgment function whether the generated picture meets the condition.
[0104] The calculation formula of the Clip score is as follows:
[0105]
[0106] In the formula, Clip i represents the Clip score of the i-th image-text pair generated picture I i and description text P i . ‖·‖ is the Euclidean norm.
[0107] Embodiment 5:
[0108] The integrated method based on unified expression of design information and innovation fusion, the main technical content of which is any one of embodiments 1 to 4, further, the subset D' of the original image-text pair dataset D is as follows:
[0109]
[0110] In the formula, i is the image-text pair number, and reatin(i) is the screening function. Clip i is the Clip score of the i-th image-text pair. DSG iDsg score of the ith image-text pair. T clip Clip threshold. T dsg Dsg threshold.
[0111] Embodiment 6:
[0112] The integrated method based on the unified expression of design information and the fusion of innovation, the main technical content of any one of embodiments 1 to 5, further, the elements D i ′ in the subset D′ also include semantic vectors SV i ′.
[0113] Embodiment 7:
[0114] The integrated method based on the unified expression of design information and the fusion of innovation, the main technical content of any one of embodiments 1 to 6, further, the semantic relevance of the description text P i ′ in the subset D′ and the input text P, the image-text pair data set with the highest relevance to the input text P is screened out from the subset D′ The steps are as follows:
[0115] 3.1) Perform named entity recognition on the input text P to obtain an entity array of the input text P Where j0 is the number of entity words in the input text P.
[0116] 3.2) Screen the subset D′ to obtain the data set D′ temp As follows:
[0117] D′ temp ={(I′ i ,P′ i ,SV′ i )∈D′|contains(P i ′,E)=True} (7)
[0118]
[0119] In the formula, contains(P i ′,E) is a Boolean function. I i ′, P i ′, SV i ′ are the generated picture, description text and semantic vector in the subset D′, respectively. e is an entity word.
[0120] 3.3) Construct a multi-modal large model CLIP.
[0121] 3.4) Input the input text P into the multi-modal large model CLIP to obtain the semantic encoding PV.
[0122] 3.5) Calculate dataset D′ temp semantic vector SV′ in temp The semantic similarity scores are obtained by comparing the semantic similarity scores of the semantically encoded PV with those of the PV, resulting in a semantic similarity score array S = [S1, S2, S3, ..., S...]. m ], where m is the dataset D′ temp Size.
[0123] The formula for calculating semantic similarity is as follows:
[0124]
[0125] In the formula, S(PV,SV′) temp ) represents semantic similarity. PV·SV′ temp Semantic encoding PV and semantic vector SV′ temp The dot product function. ‖·‖ is the Euclidean norm.
[0126] 3.6) Sort the semantic similarity score array S in descending order, and take the data corresponding to the first n1 elements to form the image-text pair dataset D″.
[0127] Example 8:
[0128] The integration method based on unified expression and innovative fusion of design information, the main technical contents of which are described in any one of Examples 1 to 7, and further, the fused text P′ is as follows:
[0129] P′=fuse(Prompts) (10)
[0130]
[0131] In the formula, Prompts is the descriptive text P. i The set of "" and the input text P. `fuse` is the fusion function.
[0132] Example 9:
[0133] Based on the integrated method of unified expression and innovative fusion of design information, the main technical contents are described in any one of Examples 1 to 8. Furthermore, the qualified image set GI... q and unqualified image sets GI nq As shown below:
[0134] GI q ={i∈GI|qualify(i)=1} (12)
[0135] GI nq ={i∈GI|qualify(i)=0} (13)
[0136]
[0137] where i is the image-text pair index. qualify(i) is a filter function. are the set of Clip scores and the set of Dsg scores between the generated images and the texts P' in the generated image dataset GI, respectively. clip is the Clip threshold. dsg is the Dsg threshold.
[0138] Embodiment 10:
[0139] The integrated method based on the unified expression of design information and the fusion of innovation, the main technical content is any one of embodiments 1 to 9, further, the user selects a relatively satisfactory image I q from the qualified image set GI like , and generates a new text P next The steps are as follows:
[0140] 8.1) Construct a visual large model.
[0141] 8.2) Input the selected image I like into the visual large model to obtain the visual encoding of the image I like , as follows:
[0142] IV like = encode(I like ) (15)
[0143] where IV like is the visual encoding of the image I like , and encode is the analog image encoding function.
[0144] 8.3) According to the characteristics of the visual encoding IV like , a new text P next is generated, as follows:
[0145] P next = GeneratePrompt(IV like ) (16)
[0146] where GeneratePrompt is the conversion function.
[0147] Embodiment 11:
[0148] Referring to Figure 1 , the integrated method based on the unified expression of design information and the fusion of innovation includes the following steps:
[0149] 1) Obtain the original image-text pair dataset D = [D1, D2, D3, …, D nand the text P input by the user when generating the image, for example, “generate a car”. Among them, the original image text pair D i includes the description text P i and the picture I generated by the text i , i is the image text pair serial number, and n is the number of original image text pairs.
[0150] The original image text pair dataset is Diffusion DB, which is the first large-scale text-to-image prompt dataset; this dataset has two versions, and the method uses the 2M version, which has 2 million data, and each data is composed of prompt, hyperparameter, and generated picture.
[0151] 2) Based on the matching degree of the generated picture I i and the description text P i , the original image text pair dataset D is filtered to construct a subset D′ of the original image text pair dataset D , wherein n0 is a positive integer less than n.
[0152] 3) Based on the semantic relevance of the description text P′ in the subset D′ i and the input text P, the image text pair dataset D″ with the highest relevance to the input text P is selected from the subset D′ , wherein n1 is 5.
[0153] 4) Construct a large language model and a text-to-image model.
[0154] 5) The description text P i ″ in the image text pair dataset D″ with the highest relevance and the input text P are input into the large language model to generate the fused text P′.
[0155] 6) The fused text P′ is input into the text-to-image model to generate the image dataset GI = [GI1, GI2, …, GI k ], and k is the number of generated images.
[0156] 7) According to the semantic similarity of the generated image dataset GI and the text P′, the generated image dataset GI is divided into the qualified image set GI q and the unqualified image set GI nq .
[0157] 8) The user selects a relatively satisfactory image I q from the qualified image set GI like , and generates a new text P mext .
[0158] The relatively satisfactory image I likeIt is a subjective perspective to determine which picture is the best in a pile of pictures, the most eye-catching, but it is certain that the first few generated results are flawed, so the image is described as relatively satisfactory.
[0159] 9) Let new text P next The text P as input returns to step 3) until the user gets a satisfactory image.
[0160] The satisfactory image is a subjective standard, and the user is satisfied, and no further iteration is needed to be satisfied.
[0161] Example 12:
[0162] The integrated method based on the unified expression of design information and innovation fusion, the main technical content is seen in any one of embodiments 11 to 12, further, the original image text pair D i Also includes the hyperparameters H i ′ for text generation pictures.
[0163] The hyperparameters H i ′ include image size, random seed, iteration step.
[0164] Hyperparameters refer to each piece of data in the data set D, in addition to the image text pair, also contains the input hyperparameters when generating the corresponding image, such as size, temperature, random seed, etc.
[0165] Example 13:
[0166] The integrated method based on the unified expression of design information and innovation fusion, the main technical content is seen in any one of embodiments 11 to 12, further, the original image text pair D i The matching degree of the description text P i , filter the original image text pair data set D, and construct a subset of the original image text pair data set D The steps are as follows:
[0167] 2.1) Calculate the sample size of random sampling, as follows:
[0168]
[0169] In the formula, n s is the sample size of random sampling, taking the value of 10,000, Z is the standard normal distribution score of 95% confidence level, 1.96 is the Z value under the 95% confidence level. P is the expected proportion, usually using 0.5. E is the error limit, taking the value of 0.01.
[0170] 2.2) Based on the sample size of random sampling, randomly sample n s sample data from the original image text pair data set D as a sample set
[0171] 2.3) Constructing the multi-modal large model CLIP and DSG evaluation framework.
[0172] Where for the multi-modal large model CLIP:
[0173] Suppose we have m pairs of image-text pairs, the image encoding representation is I m , and the text encoding representation is T m , then the similarity (Clip score) matrix S between image-text pairs is:
[0174]
[0175] Where the element s ik of S represents the semantic similarity between the i-th image and the j-th text.
[0176] Where for the evaluation framework DSG:
[0177] DSG (Davidsonian Scene Graph) is a structured approach to evaluating the performance of text-to-image generation models. The DSG framework works by:
[0178] a) Structured representation: DSG decomposes the textual description into structured atomic propositions that cover entities, attributes, relationships, and global features.
[0179] b) Dependency graph: DSG uses a directed acyclic graph (DAG) to represent the dependencies between these atomic propositions, ensuring the logical consistency of the problem.
[0180] c) Question generation: DSG generates specific natural language questions from the atomic propositions, which can be used by visual question answering (VQA) models to evaluate the generated images.
[0181] d) Evaluation: DSG evaluates the performance of the image generation model by comparing the consistency of the answers given by the VQA model based on the generated images with the expected answers based on the textual prompts.
[0182] The core of the DSG framework lies in its combination of natural language processing and computer vision techniques, particularly the use of pre-trained large language models (LLMs) and visual question answering models (VQA) to provide a fine-grained and interpretable way to evaluate the accuracy and reliability of text-to-image generation models.
[0183] 2.4) The description text s and generated image of the elements in the sample set D Input the multi-modal large model CLIP to obtain the corresponding semantic vector and visual vector
[0184] The multi-modal large model CLIP outputs the calculated Clip score
[0185] 2.5) The sample set D s The description text of the element is input into the DSG evaluation framework to calculate the Dsg score
[0186] 2.6) Based on the Clip score and the Dsg score, data mapping is performed, and the Clip threshold T clip and the Dsg threshold T dsg are determined.
[0187] 2.7) According to the Clip threshold T clip and the Dsg threshold T dsg , the original image-text pair dataset D is screened to construct a subset D' of the original image-text pair dataset D.
[0188] Embodiment 14:
[0189] The integrated method based on the unified expression of design information and the integration of innovation, the main technical content is any one of embodiments 11 to 13, further, the calculation formula of the Dsg score is as follows:
[0190]
[0191] In the formula, i is the serial number of the image-text pair. DSG i is the Dsg score of the i-th image-text pair. j is the number of questions generated after inputting the description text into the DSG evaluation framework. is the j-th question generated after inputting the description text into the DSG evaluation framework, such as whether there is a bicycle in the picture, whether the bicycle is red, and whether the bicycle is leaning against the wall. score i is a judgment function for whether the generated picture meets the conditions.
[0192] The calculation formula of the Clip score is as follows:
[0193]
[0194] In the formula, Clip i represents the i-th image-text pair generated picture I i and description text Pi Clip score of the i-th image-text pair.‖·‖ is the Euclidean norm.
[0195] Embodiment 15:
[0196] The integrated method based on unified expression of design information and innovation fusion, the main technical content of which is seen in any one of embodiments 11 to 14, further, the subset D' of the data set D based on the original image-text pairs is as follows:
[0197]
[0198] wherein i is the serial number of the image-text pair, reatin(i) is the screening function. Clip i is the Clip score of the i-th image-text pair. DSG i is the Dsg score of the i-th image-text pair. T clip is the Clip threshold value. T dsg is the Dsg threshold value.
[0199] Embodiment 16:
[0200] The integrated method based on unified expression of design information and innovation fusion, the main technical content of which is seen in any one of embodiments 11 to 15, further, the element D' in the subset D' based on the description text P i further comprises a semantic vector SV i '.
[0201] Embodiment 17:
[0202] The integrated method based on unified expression of design information and innovation fusion, the main technical content of which is seen in any one of embodiments 11 to 16, further, the semantic relevance of the description text P i ' in the subset D' based on the input text P, the image-text pair data set with the highest relevance to the input text P is screened out from the subset D' based on the semantic relevance of the description text P The steps are as follows:
[0203] 3.1) Perform named entity recognition on the input text P to obtain an entity array of the input text P wherein j0 is the number of entity words in the input text P.
[0204] 3.2) Screen the subset D ′ to obtain the data set D' temp as follows:
[0205] D' temp = {(I' i , P' i , SV' i ) ∈ D' | contains (P' i,E)=True} (7)
[0206]
[0207] In the formula, contains(P′) i E) is a Boolean function. I′ i 、P′ i ,SV′ i These represent the generated image, descriptive text, and semantic vector in subset D′, respectively. e represents an entity word.
[0208] 3.3) Construct a multimodal large model CLIP.
[0209] 3.4) Input the input text P into the multimodal large model CLIP to obtain the semantic code PV.
[0210] 3.5) Calculate dataset D′ temp semantic vector SV′ in temp The semantic similarity scores are obtained by comparing the semantic similarity scores of the semantically encoded PV with those of the PV, resulting in a semantic similarity score array S = [S1, S2, S3, ..., S...]. m ], where m is the dataset D′ temp Size.
[0211] The formula for calculating semantic similarity is as follows:
[0212]
[0213] In the formula, S(PV,SV′) temp ) represents semantic similarity. PV·SV′ temp Semantic encoding PV and semantic vector SV′ temp The dot product function. ‖·‖ is the Euclidean norm.
[0214] 3.6) Sort the semantic similarity score array S in descending order, and take the data corresponding to the first n1 elements to form the image-text pair dataset D″.
[0215] Example 18:
[0216] The integration method based on unified expression and innovative fusion of design information, the main technical contents of which are described in any one of Examples 11 to 17, and further, the fused text P′ is as follows:
[0217] P′=fuse(Prompts) (10)
[0218]
[0219] In the formula, Prompts is the descriptive text P. iand a set of input texts P. fuse is a fusion function.
[0220] Embodiment 19
[0221] The integrated method for unified expression and innovation fusion based on design information, the main technical content of which is any one of embodiments 11 to 18, further, the qualified image set GI q and the unqualified image set GI nq As shown below:
[0222] GI q ={i∈GI|qualify(i)=1} (12)
[0223] GI nq ={i∈GI|qualify(i)=0} (13)
[0224]
[0225] In the formula, i is the image text pair serial number. qualify(i) is a screening function. respectively, the Clip score set and the Dsg score set between the generated image in the generated image data set GI and the text P'. T clip is the Clip threshold value. T dsg is the Dsg threshold value.
[0226] Embodiment 20
[0227] The integrated method for unified expression and innovation fusion based on design information, the main technical content of which is any one of embodiments 11 to 19, further, the user selects a relatively satisfactory image I q from the qualified image set GI like and generates a new text P next The steps are as follows:
[0228] 8.1) Construct a visual large model.
[0229] The visual large model is used to realize image coding by using CLIP.
[0230] 8.2) Input the selected image I like into the visual large model to obtain the visual coding of the image I like , as shown below:
[0231] IV like =encode(I like ) (15)
[0232] In the formula, IV like is the image I likeThe visual encoding is a 512-dimensional vector, which is encoded as an analog image encoding function.
[0233] 8.3) According to the visual encoding IV like features, construct a text P in the vector space that has a high semantic relevance and can accurately describe the visual content of the image next , as follows:
[0234] P next = GeneratePrompt(IV like ) (16)
[0235] In the formula, GeneratePrompt is a conversion function.
[0236] Example 21:
[0237] See Figure 1 , the integrated method based on the unified expression and innovative integration of design information, the main technical content includes:
[0238] In the research on the integrated integration of the expression and innovative methods of multi-modal design information, realizing the efficient transmission and innovative generation of information is a key issue. For this reason, a research on the integrated integration of the unified expression of design information and innovative methods is proposed. This research adopts the method of iterative interactive Prompt optimization, which can dynamically adjust the expression mode of design information to better meet the needs of different scenarios. Through this method, not only the accuracy and adaptability of information expression are improved, but also design innovation can be stimulated, thus realizing the efficient integration of the unified expression of design information and innovative methods.
[0242] 3) Obtain the original prompt P from the user input, where P is the text entered by the user when generating the image. For example, "Generate a car".
[0243] 4) Based on the semantic relevance between promptP′ and promptP in D′, filter and extract the set of image-text pairs D″=[D″1,D″2,D″3,…,D″5] that are most relevant to P from D′.
[0244] 5) The prompt text in D″ and the user input promptP are fed into the pre-fine-tuned large language model Qwen:32b and fused to generate a brand new prompt P′.
[0245] 6) Input P′ into the Wensheng image model to generate images GI = [GI1, GI2, GI, ..., GI] in batches. k ]; k is the number of images generated.
[0246] 7) Based on the semantic similarity between the generated image GI and P′, GI is divided into GI 2 and GI 3. q ,GI nq Among them, GI q A qualified generated image, GI nq This indicates a generated image that is not up to standard. Clearly, GI = GI q ∪GI nq ;
[0247] 8) The user selects a relatively satisfactory image I like Used to generate the next step promptP next .
[0248] 9) Repeat steps 4) through 8) until the user receives the appropriate prompt to generate a satisfactory image.
[0249] The original image-text pair dataset is Diffusion DB, the first large-scale text-to-image cue dataset. This dataset has two versions; this method uses the 2M version, containing 2 million data entries. Each entry consists of three parts: prompt, hyperparameters, and the generated image.
[0250] Construct a subset D′ = [D′1, D′2, D′3, ..., D′] of the original image-text pair dataset D. I The steps include:
[0251] 1) Based on the basic sample size calculation formula In a sample of 2 million in total, to achieve a confidence level of 95% and maintain an error range of 1%, and in the absence of prior information about the proportion, the calculation shows that at least 9604 samples need to be randomly selected in order to reliably reflect the data characteristics of the population. In the formula:
[0252] n is the required sample size
[0253] Z is the standard normal distribution score of the required confidence level (for example, 1.96 is the Z value at a 95% confidence level)
[0254] P is the expected proportion (if there is no prior information, 0.5 is usually used, as this will give the largest sample size)
[0255] E is the acceptable error limit (for example, 0.01 or 1%)
[0256] 2) Based on the calculation result of 1), use a random sampling method to randomly sample 10,000 data from the original data set D as a sample set where represents the i-th selected sample, which is composed of (Generate pictures), (prompt), (hyperparameters).
[0257] 3) Traverse the sample set D s , and sequentially encode the image-text pairs of the samples semantically. The encoding method is to input I s and P s into the pre-trained multi-modal large model CLIP (Contrastive Language-Image Pre-training) to obtain the corresponding semantic vector SV = [SV1, SV2, SV3, …, SV 10000 ] and visual vector IV = [IV1, IV2, IV3, …, IV 10000 ].
[0258] 4) Traverse the sample set D s , and sequentially calculate the Clip score and Dsg score for the image-text pairs of the samples to evaluate their semantic similarity. For the Clip score, it can be calculated by simply inputting the paired SV i and IV i into the CLIP large model. For the Dsg score, first input into the pre-trained DSG evaluation framework to obtain a series of questions j is the number of generated questions. For example, if is "a red bicycle leaning against the wall", then some of the generated questions Qi It could be as follows:
[0259] Is there a bicycle in the picture?
[0260] Is the bicycle red?
[0261] Is this bicycle leaning against the wall?
[0262] The formula for calculating the Dsg score of the i-th sample is as follows:
[0263]
[0264] The calculated score is Clip = [Clip1, Clip2, Clip3, ..., Clip] 10000 ] and DSG = [DSG1, DSG2, DSG3, ..., DSG 10000 ].
[0265] 5) Plot the Clip data and DSG data. After comprehensively considering the quantity and quality of the images selected, determine the selection threshold for both indicators as T. clip ,T dsg .
[0266] 6) Based on the established screening thresholds, iterate through the original dataset D. Only samples with both metrics exceeding the specified thresholds are retained; otherwise, they are discarded. The remaining samples form D′. The formula is as follows:
[0267]
[0268] In the formula, i is any sample in the original dataset D, and Clip i and DSG i Let be the Clip score and Dsg score of this sample. Then the filtered dataset D′ can be represented as the set of all samples that meet the retention criteria:
[0269]
[0270] Note that D′ now has an additional attribute SV, meaning that D′ is composed of I′ (generated image), P′ (prompt), H′ (hyperparameters), and SV′ (semantic vector).
[0271] The step of filtering and extracting the set of image-text pairs D″=[D″1,D″2,D″3,…,D″5] most relevant to P from D′ includes:
[0272] 1) Perform named entity recognition on P to obtain the entity array E = [E1, E2, E3, ..., E j ], where j is the number of entity words in P.
[0273] 2) Filter D′, retaining only samples whose prompts contain any element from E, thus obtaining dataset D′. temp The screening process is as follows:
[0274] Define a Boolean function:
[0275]
[0276] In the formula, P is the prompt of the sample in D′.
[0277] Construct a new dataset D′ temp It contains all samples that meet the conditions:
[0278] D′ temp ={(I,P,SV)∈D′|contains(P,E)=True}
[0279] 3) Input P into CLIP to obtain its semantic encoding PV, and iterate through D′. temp The semantic similarity between PV and each sample SV is calculated sequentially to obtain a semantic similarity score array S = [S1, S2, S3, ..., S...]. m ], m is the dataset D′ temp Size. The formula for calculating semantic similarity is:
[0280]
[0281] In the formula:
[0282] PV·SV is the dot product of vectors PV and SV.
[0283] ‖PV‖ and ‖SV‖ are the Euclidean norms (i.e., the lengths of the vectors PV and SV).
[0284] 4) Sort S in descending order and take the samples corresponding to the first five elements as D″=[D″1,D″2,D″3,…,D″5].
[0285] The process of merging the prompt text in D″ and the user input promptP to generate a completely new prompt P′ is as follows:
[0286] 1) Integrate all prompt text in D″ with the user-input prompt P. Define a set Prompts to contain this text:
[0287] Prompts = [P1, P2, P3, ..., P5, P]
[0288] 2) input Prompts into the pre-finetuned large language model to fuse, and get P'. This process can be expressed by a fusion function fuse:
[0289] P' = fuse(Prompts) #(5)
[0290] In the formula, fuse represents the encoding fusion process of the large model for Prompts.
[0291] According to the semantic similarity between the generated image GI and the generated text P', the GI is divided into GI q , GI nq The steps are as follows:
[0292] 1) Traverse GI, calculate the Clip score and Dsg score between each generated image and P', and get the set Clip generated and DSG generated .
[0293] 2) Traverse the set Clip generated and Dsg generated , each image must meet the preset Clip and Dsg score thresholds T clip and T dsg to be identified as qualified. Through this method, we can clearly distinguish the images that meet the quality standards (qualified image set GI q ) and the images that fail to meet the standards (GI nq ). The formula is as follows:
[0294] Define a screening function:
[0295]
[0296] Define two generated image sets:
[0297] GI q = {i e GI | qualify(i) = 1}
[0298] GI nq = {i e GI | qualify(i) = 0}
[0299] According to the user-selected image I like , the next prompt P next is generated as follows:
[0300] 1) input the user-selected image I like into the pre-defined visual large model to obtain the visual encoding of the image:
[0301] IV like= encode(I like )#(7)
[0302] Where the encode function is used to simulate the image encoding process.
[0303] 2) The large model utilizes its powerful natural language generation ability to construct a prompt P with a high semantic relevance and capable of precisely describing the visual content of the image in the vector space based on the features of IV like : next P
[0304] = GeneratePrompt(I next )#(8) like
[0305] Where GeneratePrompt is a function representing the processing of the large model, which converts the input image into descriptive text highly similar to the image.
[0306] Example 22:
[0307] Figure 1 Refer to , for the integrated method based on the unified expression and innovative integration of design information, the main technical contents include:
[0308] 1) Obtain the original image-text pair dataset D = [D1, D2, D3,..., D n ; n is the number of original image-text pairs;
[0309] 2) Screen the original dataset based on the matching degree of image-text pairs to construct a subset D' = [D'1, D'2, D'3,..., D' i of the original image-text pair dataset D; i < n; the element D' i contains an image-text pair and hyperparameters, as well as the semantic encoding corresponding to the text;
[0310] 3) Obtain the original prompt P input by the user;
[0311] 4) Screen and extract a set of image-text pairs D'' = [D''1, D''2, D''3,..., D''5] with the highest relevance to P from D' based on the semantic relevance to prompt P;
[0312] 5) Feed the prompt text in D'' and the user input prompt P into a pre-fine-tuned large language model to fuse and generate a new prompt P';
[0313] 6) Feed P' into the text-to-image model to batch generate images GI = [GI1, GI2, GI,..., GI k ; k is the number of generated images;
[0314] 7) Based on the semantic similarity between the generated image GI and the generated text P′, GI is divided into GI 2 and GI 3. q ,GI nq Among them, GI q A qualified generated image, GI nq This indicates a generated image that is not up to standard. Clearly, GI = GI q ∪GI nq ;
[0315] 8) The user selects a relatively satisfactory image I like Used to generate the next step promptP next ;
[0316] 9) Repeat steps 4) through 8) until the user receives a suitable prompt to generate a satisfactory image;
[0317] The original image-text pair dataset is Diffusion DB, which is the first large-scale text-to-image cue dataset. There are two versions of this dataset. This method uses the 2M version, which contains 2 million data entries. Each data entry consists of three parts: prompt, hyperparameters, and generated image.
[0318] Construct a subset D′ = [D′1, D′2, D′3, ..., D′] of the original image-text pair dataset D. I The steps include:
[0319] 1) Based on the basic sample size calculation formula In a sample of 2 million, to achieve a 95% confidence level and maintain a 1% error margin, and in the absence of prior information about proportions, calculations show that at least 9604 samples need to be randomly drawn to reliably reflect the data characteristics of the population. Where:
[0320] n is the required sample size.
[0321] Z is the standard normal distribution score for the required confidence level (e.g., 1.96 is the Z-value at a 95% confidence level).
[0322] P is the expected proportion (if there is no prior information, 0.5 is usually used because this gives the maximum sample size).
[0323] E is the acceptable error limit (e.g., 0.01 or 1%).
[0324] 2) Based on the calculation results in 1), 10,000 data points are randomly sampled from the original dataset D using a random sampling method as the sample set. in denotes the i-th selected sample, where (Generate pictures), (prompt), (hyperparameters) are composed;
[0325] 3) Traverse the sample set D s , and sequentially perform semantic encoding on the image-text pairs of the samples. The encoding method is to input I s and P s into the pre-trained multi-modal large model CLIP (Contrastive Language-Image Pre-training) to obtain the corresponding semantic vector SV = [SV1, SV2, SV3, …, SV 10000 ] and visual vector IV = [IV1, IV2, IV3, …, IV 10000 ];
[0326] 4) Traverse the sample set D s , and sequentially calculate the Clip score and Dsg score for the image-text pairs of the samples to evaluate their semantic similarity. For the Clip score, it can be calculated by simply inputting the paired SV i and IV i into the CLIP large model. For the Dsg score, first input into the pre-trained DSG evaluation framework to obtain a series of questions j is the number of generated questions. For example, if is “a red bicycle leans against the wall”, then some of the generated question series Q i may be as follows:
[0327] Is there a bicycle in the picture?
[0328] Is the bicycle red?
[0329] Is the bicycle leaning against the wall?
[0330] The Dsg score of the i-th sample is calculated as follows:
[0331]
[0332] The calculated scores are Clip = [Clip1, Clip2, Clip3, …, Clip 10000 ] and DSG = [DSG1, DSG2, DSG3, …, DSG 10000 ];
[0333] 5) Plot the Clip data and DSG data, and after comprehensive consideration of the number and quality of the selected pictures, determine the screening threshold of the two indicators as T clip,T dsg ;
[0334] 6) Based on the established screening thresholds, iterate through the original dataset D. Only samples with both metrics exceeding the specified thresholds are retained; otherwise, they are discarded. The remaining samples form D′. The formula is as follows:
[0335]
[0336] In the formula, i is any sample in the original dataset D, and Clip i and DSG i Let be the Clip score and Dsg score of this sample. Then the filtered dataset D′ can be represented as the set of all samples that meet the retention criteria:
[0337]
[0338] Note that D′ now has an additional attribute SV, meaning that D′ is composed of I′ (generated image), P′ (prompt), H′ (hyperparameters), and SV′ (semantic vector).
[0339] The steps to filter and extract the set of image-text pairs D″=[D″1,D″2,D″3,…,D″5] that are most relevant to P from D′ include:
[0340] 1) Perform named entity recognition on P to obtain the entity array E = [E1, E2, E3, ..., E j ], where j is the number of entity words in P;
[0341] 2) Filter D′, retaining only samples whose prompts contain any element from E, thus obtaining dataset D′. temp The screening process is as follows:
[0342] Define a Boolean function:
[0343]
[0344] In the formula, P is the prompt of the sample in D′.
[0345] Construct a new dataset D′ temp It contains all samples that meet the conditions:
[0346] D′ temp ={(I,P,SV)∈D′|contains(P,E)=True}
[0347] 3) Input P into CLIP to obtain its semantic encoding PV, and iterate through D′. temp, the semantic similarity between PV and each sample SV is calculated in turn, and a semantic similarity score array S = [S1, S2, S3, …, S m ] is obtained, where m is the size of the data set D' temp . The formula for calculating semantic similarity is:
[0348]
[0349] In the formula:
[0350] PV·SV is the dot product of vectors PV and SV
[0351] ‖PV‖ and ‖SV‖ are the Euclidean norms (i.e., lengths) of vectors PV and SV
[0352] 4) Sort S in descending order, and take the top five elements corresponding to the samples as D″ = [D″1, D″2, D″3, …, D″5];
[0353] The prompt text in D″ and the user input prompt P are fused to generate a new prompt P′ as follows:
[0354] 1) Integrate all prompt texts in D″ and the user input prompt P. Define a set Prompts to include these texts:
[0355] Prompts = [P1, P2, P3,.., P5, P]
[0356] 2) Input the set Prompts together into the pre-tuned large language model for fusion to obtain P′. This process can be expressed by a fusion function fuse:
[0357] P′ = fuse(Prompts) #(5)
[0358] In the formula, fuse represents the encoding and fusion process of the large model for Prompts;
[0359] According to the semantic similarity between the generated image GI and the generated text P′, GI is divided into GI q ,GI nq The steps are as follows:
[0360] 1) Traverse GI, calculate the Clip score and Dsg score between each generated image and P′, and obtain sets Clip generated and DSG generated ;
[0361] 2) Traverse sets Clip generated and Dsg generated, each image must satisfy the preset Clip and Dsg score thresholds T simultaneously clip and T dsg are considered qualified. In this way, we can clearly distinguish between images that meet the quality standards (the qualified image set GI q ) and those that fail to meet the standards (GI nq ). The formula is as follows:
[0362] Define a screening function:
[0363]
[0364] Define two image sets:
[0365] GI q = {i e GI | qualify(i) = 1}
[0366] GI nq = {i e GI | qualify(i) = 0}
[0367] According to the user's selected image I like , the next step is to generate the prompt P next as follows:
[0368] 1) Input the user's selected image I like into the pre-defined visual large model to obtain the visual encoding of the image:
[0369] IV like = encode(I like ) #(7)
[0370] Where the encode function is used to simulate the image encoding process.
[0371] 2) The large model uses its powerful natural language generation capabilities to construct a prompt P next with high semantic relevance and accurate description of the visual content of the image in the vector space based on the characteristics of IV like :
[0372] P next = GeneratePrompt(IV like ) #(8)
[0373] Where GeneratePrompt is a function representing the processing of the large model.
Claims
1. An integrated method based on unified expression and innovative fusion of design information, characterized in that, Includes the following steps: 1) Obtain the original image-text pair dataset D = [D1, D2, D3, ..., D n The text P input by the user when creating the image; where the original image text is paired with D. i Including the descriptive text P i And images generated from text I i i is the image-text pair index, and n is the original number of image-text pairs; 2) Based on generated image I i With description text P i Based on the degree of matching, filter the original image-text pair dataset D, and construct a subset of the original image-text pair dataset D. Where n0 is a positive integer less than n; 3) Based on the description text P in subset D′ i The semantic relevance of the image-text pairs to the input text P is used to select the image-text pairs from the subset D' that have the highest relevance to the input text P. Where n1 is a positive integer less than n0; 4) Construct a large language model and a text-to-graph model; 5) P represents the descriptive text P in dataset D″, which has the highest relevance between the image and text pairs. i " and the input text P are fed into the large language model to generate the fused text P′; 6) Feed the fused text P′ into the text-to-image model to generate an image dataset GI = [GI1, GI2, ..., GI...]. k ], k is the number of images generated; 7) Based on the semantic similarity between the generated image dataset GI and the text P′, the generated image dataset GI is divided into qualified image sets GI. q and unqualified image sets GI nq ; 8) Users from qualified image sets GI q Select a relatively satisfactory image I like And generate a new text P next ; 9) Let the new text P next The text P, which is the input, returns to step 3) until the user obtains a satisfactory image.
2. The integration method based on unified expression and innovative fusion of design information according to claim 1, characterized in that, The original image text pair D i It also includes the hyperparameter H for generating images from text. i ′; The hyperparameter H i This includes image size, random seed, and number of iterations.
3. The integration method based on unified expression and innovative fusion of design information according to claim 1, characterized in that, The image-based generation I i With description text P i Based on the degree of matching, filter the original image-text pair dataset D, and construct a subset of the original image-text pair dataset D. The steps are as follows: 2.1) Calculate the randomly selected sample size as follows: In the formula, n s Z is the sample size randomly selected; Z is the standard normal distribution score at the a% confidence level, where a is a positive integer; P is the expected proportion; and E is the error limit. 2.2) Based on the randomly selected sample size, n samples are randomly drawn from the original image-text pair dataset D. s Each sample data is used as a sample set 2.3) Construct the CLIP and DSG evaluation frameworks for multimodal large models; 2.4) The sample set D s medium elements Description text and generating images Inputting a large multimodal model CLIP yields the corresponding semantic vectors. and visual vectors The Clip score is calculated from the CLIP output of the multimodal large model. 2.5) The sample set D s medium elements Description text Input the DSG assessment framework to calculate the DSG score. 2.6) Plot the data based on the Clip score and Dsg score, and determine the Clip threshold T. clip and Dsg threshold T dsg ; 2.7) Based on the Clip threshold T clip and Dsg threshold T dsg Filter the original image-text pair dataset D to construct a subset D′ of the original image-text pair dataset D.
4. The integration method based on unified expression and innovative fusion of design information according to claim 3, characterized in that, The formula for calculating the Dsg score is as follows: In the formula, i is the image-text pair number; DSG i Let be the Dsg score of the i-th image-text pair; j is the descriptive text. The number of questions generated after inputting the DSG evaluation framework; To describe the text The j-th question generated after inputting into the DSG evaluation framework; score i To generate an image A function to determine whether the condition is met; The formula for calculating the Clip score is as follows: In the formula, Clip i Indicates the generation of image I from the i-th image-text pair. i and description text P i The Clip fraction; ||·| is the Euclidean norm.
5. The integration method based on unified expression and innovative fusion of design information according to claim 3, characterized in that, The subset D′ of the original image-text pair dataset D is shown below: In the formula, i is the image-text pair index, and reatin(i) is the filtering function; Clip i The Clip score for the i-th image-text pair; DSG i T is the Dsg score of the i-th image-text pair; clip The Clip threshold; T dsg The threshold value for Dsg is [value].
6. The integration method based on unified expression and innovative fusion of design information according to claim 3, characterized in that, The element D′ in the subset D′ i It also includes semantic vectors (SVs). i ′.
7. The integration method based on unified expression and innovative fusion of design information according to claim 1, characterized in that, The descriptive text P based on subset D′ i The semantic relevance of the image-text pairs to the input text P is used to select the image-text pairs from the subset D' that have the highest relevance to the input text P. The steps are as follows: 3.1) Perform named entity recognition on the input text P to obtain the entity array of the input text P. Where j0 is the number of entity words in the input text P; 3.2) Filter the subset D′ to obtain the dataset D′. temp As shown below: D′ temp ={(I i ′,P i ′,SV i ′)∈D′|contains(P i ′,E) =True} (7) In the formula, contains(P) i I′, E) is a Boolean function; i P i ′、SV′ i These represent the generated image, descriptive text, and semantic vector in subset D′, respectively; e represents the entity word. 3.3) Construct a multimodal large model CLIP; 3.4) Input the input text P into the multimodal large model CLIP to obtain the semantic code PV; 3.5) Calculate dataset D′ temp semantic vector SV′ in temp The semantic similarity scores are obtained by comparing the semantic similarity scores of the semantic encoding PV with those of the semantic similarity scores of S = [S1, S2, S3, ..., S...]. m ], where m is the dataset D′ temp Size; The formula for calculating semantic similarity is as follows: In the formula, S(PV,SV′) temp ) represents semantic similarity; PV·SV′ temp Semantic encoding PV and semantic vector SV′ temp The dot product function; ||·| is the Euclidean norm; 3.6) Sort the semantic similarity score array S in descending order, and take the data corresponding to the first n1 elements to form the image-text pair dataset D″.
8. The integration method based on unified expression and innovative fusion of design information according to claim 1, characterized in that, The fused text P′ is shown below: P′=fuse(Prompts) (10) In the formula, Prompts is the description text P. i " and the set of input text P; fuse is the fusion function.
9. The integration method based on unified expression and innovative fusion of design information according to claim 1, characterized in that, The qualified image set GI q and unqualified image sets GI nq As shown below: GIVE q ={i∈ GI|qualify(i)=1} (12) GIVE nq ={i∈ GI|qualify(i)=0} (13) In the formula, i is the image-text pair index; qualify(i) is the filtering function; These represent the Clip score set and Dsg score set, respectively, between the generated image and text P′ in the generated image dataset GI; T clip The Clip threshold; T dsg The threshold value for Dsg is [value].
10. The integration method based on unified expression and innovative fusion of design information according to claim 1, characterized in that, The user is from the qualified image set GI q Select a relatively satisfactory image I lik And generate a new text P next The steps are as follows: 8.1) Construct a large visual model; 8.2) Select the image I like Input a large visual model and obtain image I like The visual encoding is as follows: IV like =encode(I like ) (15) In the formula, IV like For image I like The visual encoding, where encode is a simulated image encoding function; 8.3) Based on visual encoding IV like Features that generate new text P next As shown below: P next =GeneratePrompt(IV like ) (16) In the formula, GeneratePrompt is the conversion function.
Citation Information
Patent Citations
Unsupervised cross-modal hash retrieval method based on CLIP and attention fusion mechanism
CN118861327A
Image-text data enhancement method, text-to-graph model training method and image generation method
CN118968214A