Knitted product image generation method and device based on text adjustment and visual feedback

By adaptively adjusting the user input text and optimizing the image generation model based on the generative flow network GFlowNet and visual feedback, the problems of texture detail and text consistency in knitted product image generation in the existing technology are solved, and high-quality knitted product image generation is achieved.

CN120219553BActive Publication Date: 2025-09-09HUAQIAO UNIVERSITY +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510695629.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-09
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing knitted product image generation technologies have significant deficiencies in texture details, input adaptability, and text-image consistency, making it difficult to generate high-quality knitted product images.

Method used

A language model based on the generative flow network GFlowNet is used to adaptively adjust the user input text, and the image generation model Stable diffusion is optimized through visual feedback. Visual question answering and attention matrix optimization technology are used to ensure the alignment of image and text.

Benefits of technology

The generation quality and user experience of knitted product images have been significantly improved, ensuring the consistency between the generated images and texts, and improving the realism and detail expression of the images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219553B_ABST
    Figure CN120219553B_ABST
Patent Text Reader

Abstract

A method and device for generating knitted product images based on text adjustment and visual feedback, relating to the field of computer vision, includes: constructing a language model to adjust input text; inputting knitting-related text received from a user into a trained language model to obtain text that is adaptively adjusted to the user input text; inputting the adaptively adjusted text into a trained text-image model to generate a corresponding image; formatting the adaptively adjusted text to determine whether the generated image matches the adaptively adjusted text; inputting the formatted text and the generated image into a large visual language model for visual question answering to obtain a score; outputting a knitted product image if the score meets the expectation; and regenerating the image by minimizing the distance between the text self-attention matrix and the cross-attention similarity matrix to optimize potential noise variables if the score does not meet the expectation. This invention significantly improves the quality of knitted product image generation and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a method and device for generating knitted product images based on text adjustment and visual feedback. Background Art

[0002] Knitted products play a crucial role in modern apparel design and production. Their complex textures and diverse material properties make image generation a challenging task. Traditional methods for generating knitted product images rely primarily on manual design or fixed template-based rendering techniques. These methods have significant limitations when dealing with complex texture details and dynamic variations. For example, existing technologies often fail to accurately capture subtle variations in knitted textures, resulting in a lack of realism and detail in the generated images.

[0003] Furthermore, image generation for knitted products requires adapting to diverse input conditions, such as product descriptions, material parameters, and user requirements. However, user-entered text often has limitations and struggles to meet diverse needs, resulting in a poor user experience. Furthermore, existing methods often fail to generate images that match user-entered text describing knitwear.

[0004] In summary, the existing knitted product image generation technology has significant shortcomings in texture details, input adaptability and text-image consistency. Summary of the Invention

[0005] The purpose of the present invention is to provide a knitted product image generation method and device based on text adjustment and visual feedback, which significantly improves the generation quality of knitted product images and user experience by dynamically adjusting the input text and optimizing the image generation model.

[0006] The present invention adopts the following technical solutions:

[0007] On the one hand, a method for generating a knitted product image based on text adjustment and visual feedback is characterized by comprising:

[0008] S101, building a language model based on the generative flow network GFlowNet to adjust the input text;

[0009] S102, inputting the knitting-related text received from the user into the trained language model to obtain text after adaptively adjusting the knitting-related text input from the user;

[0010] S103, inputting the adaptively adjusted text into the trained text-graph model Stable diffusion to generate a corresponding image;

[0011] S104, formatting the adaptively adjusted text into text for determining whether the generated image matches the adaptively adjusted text, inputting the formatted text and the generated image into a large visual language model for visual question answering to obtain a score;

[0012] S105: If the score meets expectations, the knitted product image is output; if the score does not meet expectations, the potential noise variable is optimized by minimizing the distance between the text self-attention matrix and the cross-attention similarity matrix, and the image is generated again by optimizing the stable diffusion.

[0013] Preferably, the training process of the language model for adjusting the input text based on the generative flow network GFlowNet is as follows:

[0014] Use a pre-trained language model as the initial model ; Where x represents the initial knitted text, y represents the improved text y generated based on the initial text x, represents the pre-trained language model;

[0015] Initialize the forward strategy of the generated flow network GFlowNet and stream function ;

[0016] The improved text y is input into the Stable diffsion model of the text-graph model to generate an image; through the reward function To calculate the reward, the reward function is defined as follows:

[0017] ;

[0018] in, Indicates expected value; Represents the diffusion model from text to image based on text x The generated image; Represents the diffusion model from text to image according to text y The generated image; and Representing a text-to-image diffusion model; indicates aesthetic rewards; Relevance reward.

[0019] The reward function Decomposed into a step-by-step reward signal, as follows:

[0020] ;

[0021] in, Represents the text segment generated to step t The reward value; Represents the text segment generated to the tth step; Indicates that the pre-trained language model generates text fragments The conditional probability of represents the conditional probability of the pre-trained language model generating the improved text y; represents a hyperparameter; Indicates the total number of steps to generate text y; exp() represents the exponential function;

[0022] Based on the reward signal, we extend the forward-backward balance objective to optimize GFlowNet, and the loss function is as follows:

[0023] ; ;

[0024] in, represents the overall loss function, which is used to measure the performance of the initial model in generating text y; Represents the local loss function at each step, which is used to evaluate the performance when generating text to step t+1; represents the value of the stream function at step t; represents the value of the stream function at step t+1; represents the conditional probability of the forward strategy at step t+1; represents the value of the reward function at step t; Represents the value of the reward function at step t+1; Represents the parameters of the language model; Represents the text generated up to step t+1; Represents the word generated in step t; Represents the word generated in step t+1;

[0025] Minimize the above loss function and update the forward strategy of the generated flow network GFlowNet and stream function , and update the parameters of the language model at the same time.

[0026] Preferably, update the forward strategy of the generated flow network GFlowNet and stream function , and after updating the parameters of the language model, it also includes:

[0027] A flow reactivation mechanism is introduced, which periodically resets the last layer of the GFlowNet flow function every M steps; ultimately, the language model regulated by the generative flow network GFlowNet learns to generate text proportional to the reward.

[0028] Preferably, the adaptively adjusted text is input into the trained text-graph model Stable diffusion to generate the corresponding image, as follows:

[0029] ;

[0030] Among them, prompt represents the adaptively adjusted text; Sd represents the trained text-graph model Stablediffusion; I represents the generated image.

[0031] Preferably, the adaptively adjusted text is formatted into text for determining whether the generated image matches the adaptively adjusted text, and the formatted text and the generated image are input into a large visual language model for visual question answering to obtain a score, specifically including:

[0032] For the adaptively adjusted text, a language model is used to generate formatted text to determine whether the generated image matches the adaptively adjusted text;

[0033] The generated image and formatted text are fed into a large visual language model to calculate the visual question answering score as follows:

[0034] ;

[0035] Among them, VQA represents the score of visual question answering, which is an indicator used to measure the degree of alignment between the generated image and text; I represents the generated image; text represents the formatted text; P represents the probability of the visual question answering output "Yes", and "Yes" represents the answer output by the visual question answering when the image matches the content of the text.

[0036] Preferably, if the score meets expectations, the knitted product image is output; if the score does not meet expectations, the potential noise variable is optimized by minimizing the distance between the text self-attention matrix and the cross-attention similarity matrix, and the image is generated again by optimizing the stable diffusion, specifically including:

[0037] Set a score threshold. If the score exceeds the threshold, the knitted product image is output. If the score does not exceed the threshold, the input text is first converted into dense vectors through a text editor and processed through a multi-head self-attention layer.

[0038] Self-attention matrix The calculation formula for the elements in is as follows:

[0039] ;

[0040] in, represents the self-attention weight of the i-th token to the j-th token in the l-th layer and the h-th head of the text encoder; exp() represents the exponential function; represents the attention weight between the i-th token and the j-th token; represents the attention weight between the i-th token and the k-th token; represents the key vector of the i-th token in the l-th layer of the text encoder; represents the key vector of the jth token in the lth layer; represents the pre-trained weights in the lth layer and hth head of the text encoder;

[0041] The self-attention matrices of all layers and heads are averaged, and the attention weights for special tags are removed, and then renormalized as follows:

[0042] ;

[0043] ;

[0044] in, represents the averaged self-attention matrix; Indicates the number of layers of the text encoder; Indicates the number of heads in each layer of the text encoder; Represents the renormalized self-attention matrix Elements of Represents the averaged self-attention matrix Elements of Indicates that the elements from the 2nd to the 3rd in the nth row are The sum of the elements is normalized, m is a column index; Indicates the length of the text sequence, that is, the number of tokens the text is segmented into, ;

[0045] In the diffusion model, the text and image data are connected through multi-head cross attention; the a-th query vector in the cross attention layer l is defined as ;in, , N c represents the length of the query vector in the cross-attention layer, D c is the hidden dimension of each head, H c Indicates the number of heads in the cross-attention layer; Cross-attention map The element at the hth head is defined as:

[0046] ;

[0047] in, represents the a-th query vector in the l-th layer of the cross-attention layer; represents the key vector of the i-th token of the cross-attention layer l; represents the pre-trained weight matrix in the crisscross attention layer l and head h; express and Unnormalized attention weights between ; express and The normalized attention weights between them, l represents the index of the cross attention layer, h represents the head index in the l-th layer, a represents the a-th query vector in the query sequence, i represents the i-th token in the text, and j represents the j-th token in the text; represents a real matrix;

[0048] The cross attention maps between the head and all layers are averaged as follows:

[0049] ;

[0050] Among them, A represents the average attention matrix, L M Indicates the number of cross attention layers when the query sequence length is 256. represents the cross attention map between the lth layer and the hth head;

[0051] Define a similarity matrix whose values ​​represent the similarity between the cross-attention graph pairs; the calculation formula of the elements in the similarity matrix S is as follows:

[0052] ;

[0053] in, represents the elements in the similarity matrix S; represents the cosine similarity, which is used to measure the similarity between the key vectors of the i-th token and the j-th token; A represents the similarity between the key vectors of the i-th and k-th tags; ai A represents the attention weight of the a-th query vector to the key vector of the i-th tag in the cross attention matrix; aj N represents the attention weight of the a-th query vector to the j-th token key vector in the cross attention matrix; c represents the length of the query vector in the cross-attention layer;

[0054] Optimize the potential noise by minimizing the distance between the similarity matrix S of the cross attention map and the text self-attention matrix T , which is achieved by minimizing the following loss function:

[0055] ;

[0056] in, represents the loss function; represents the weight coefficient; The elements of the text self-attention matrix T represent the self-attention weights between the i-th and j-th tokens; represents the index; Represents the elements in the cross-attention similarity matrix;

[0057] Optimizing latent noise via gradient descent ,as follows:

[0058] ;

[0059] in, represents the learning rate; represents the potential noise variable; Represents the loss function right The gradient of , used to update the potential noise variable; represents the latent noise variable after being updated by gradient descent;

[0060] Through the above process, the grammatical relationship of text embedding is effectively transferred to the cross attention, so that the semantics of the regenerated image and text remain aligned.

[0061] Preferably, after optimizing the Stable diffusion and generating the image again, the process further includes: returning to S104 to obtain the score again.

[0062] On the other hand, a knitted product image generation device based on text adjustment and visual feedback includes:

[0063] The text fine-tuning language model construction module is used to build a language model based on the generative flow network GFlowNet to adjust the input text;

[0064] a text adaptive adjustment module, configured to input the knitting-related text received from the user into the trained language model to obtain text after adaptively adjusting the knitting-related text input from the user;

[0065] The image generation module is used to input the adaptively adjusted text into the trained text-to-graph model Stablediffusion to generate the corresponding image;

[0066] A score acquisition module is used to format the adaptively adjusted text into text to determine whether the generated image matches the adaptively adjusted text, and input the formatted text and the generated image into a large visual language model for visual question answering to obtain a score;

[0067] The latent noise variable optimization module is used to determine if the score meets expectations and output the knitted product image. If the score does not meet expectations, the latent noise variable is optimized by minimizing the distance between the text self-attention matrix and the cross-attention similarity matrix, and the stable diffusion is optimized to generate the image again.

[0068] Compared with the prior art, the present invention has the following beneficial effects:

[0069] (1) The present invention fine-tunes the language model based on GFlowNet, enabling the language model to adaptively adjust the text input by the user, so that the adjusted text can accurately and multi-facetedly describe the characteristics of knitted products, providing a strong guarantee for the generation of images by the generative model;

[0070] (2) The present invention optimizes the image generation model through visual feedback from a large visual language model, so that the generated image and the input text semantics are consistent, ensuring the alignment of the image and text, and generating high-quality knitted images that are consistent with the text. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 A flowchart of a method for generating a knitted product image based on text adjustment and visual feedback provided by an embodiment of the present invention;

[0072] Figure 2 A detailed flowchart of a method for generating a knitted product image based on text adjustment and visual feedback provided by an embodiment of the present invention;

[0073] Figure 3 A structural block diagram of a knitted product image generation device based on text adjustment and visual feedback provided by an embodiment of the present invention;

[0074] Figure 4 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0075] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0076] See also Figure 1 and Figure 2 As shown, this embodiment provides a method for generating a knitted product image based on text adjustment and visual feedback, including the following steps.

[0077] A method for generating a knitted product image based on text adjustment and visual feedback, characterized by comprising:

[0078] S101, constructing a language model based on the generative flow network GFlowNet to adjust the input text.

[0079] It should be noted that the language model here refers to a language model that only has language processing capabilities and is mainly used for text generation tasks.

[0080] Use a pre-trained language model as the initial model ; where x represents the initial knitted text, and y represents the improved text y generated based on the initial text x (the improved text y here refers to the text obtained in each step of the training process). Represents a pre-trained language model. This language model has been trained on large-scale text data and has the ability to understand and generate natural language. Based on this, GFlowNet is used to fine-tune the language model and redefine the problem of text adaptation as probabilistic inference, that is, using probabilistic methods to predict how to adjust the text. Using a pre-trained language model As the initial strategy. Initialize the forward strategy of GFlowNet and stream function The initial text x is input, and the pre-trained language model generates the improved text y. The improved text y is input into the text-to-graph model Stable diffsion to generate an image.

[0081] It's important to note that the Stable diffsion model used here is only used for image generation: text input and image output. This can be considered an optimization of text generation. Therefore, by optimizing the reward function as part of the loss function, the language model parameters are adjusted to achieve higher rewards for the generated text, enabling the language model to learn how to adjust the initial text.

[0082] Through the reward function Calculate the reward, where the reward function is defined as follows:

[0083] ;

[0084] in, Indicates the expected value, Represents the diffusion model from text to image based on the initial text x The generated image, Represents the diffusion model from text to image according to text y The generated image, and represents the text-to-image diffusion model, Indicates aesthetic rewards, Relevance reward.

[0085] In order to increase the diversity of text and alleviate mode collapse, the reward function Decomposed into a step-by-step reward signal, the formula is as follows:

[0086] ;

[0087] in, Represents the text segment generated to step t The reward value, Indicates that the pre-trained language model generates text fragments The conditional probability of represents the conditional probability of the pre-trained language model generating the complete text y, represents the reward function, represents the hyperparameter, Indicates the total number of steps to generate text, Represents the text segment generated up to the tth step, and exp() represents the exponential function.

[0088] The above process enables the introduction of a step-by-step reward signal obtained by decomposing the reward function at each step to extend the use of the forward-backward balance objective to optimize GFlowNet. The loss function formula is as follows:

[0089] ; ;

[0090] in, represents the overall loss function, which is used to measure the performance of the initial model in generating text y; Represents the local loss function at each step, which is used to evaluate the performance when generating text to step t+1; represents the value of the stream function at step t; represents the value of the stream function at step t+1; represents the conditional probability of the forward strategy at step t+1; represents the value of the reward function at step t; Represents the value of the reward function at step t+1; Represents the parameters of the language model; Represents the text generated up to step t+1; Represents the word generated in step t; Represents the word generated in step t+1.

[0091] By minimizing the above loss function, the forward strategy of the generated flow network GFlowNet is updated and stream function , while simultaneously updating the language model's parameters. To prevent neurons from gradually deactivating, a flow reactivation mechanism is introduced. This reactivation mechanism periodically resets the last layer of the GFlowNet flow function every M steps. Ultimately, the language model fine-tuned based on the GFlowNet learns to generate text proportional to the reward.

[0092] S102: inputting the knitting-related text received from the user into the trained language model to obtain text after adaptively adjusting the knitting-related text input from the user.

[0093] S103 , inputting the adaptively adjusted text into the trained text-to-graph model Stable diffusion to generate a corresponding image.

[0094] Specifically, the user inputs knitting-related text into the language model fine-tuned by GFlowNet, which then adaptively adjusts the text. The adjusted text is then fed into the pre-trained Stable diffusion model to generate the relevant image. The formula is as follows:

[0095] ;

[0096] ;

[0097] Among them, x represents the initial text entered by the user, LM represents the language model fine-tuned by GFlowNet, represents the text after adaptive adjustment, Sd represents the pre-trained text-graph model Stable diffusion, and I represents the generated image.

[0098] S104: formatting the adaptively adjusted text into text for determining whether the generated image matches the adaptively adjusted text, inputting the formatted text and the generated image into a large visual language model for visual question answering to obtain a score.

[0099] Based on the adjusted text, a language model is used to generate formatted text: "Does this image and {text} match? Answer yes or no." The generated image and formatted text are input into a large visual language model to calculate the visual question answering score:

[0100] ;

[0101] VQA represents the visual question answering score, a metric used to measure the alignment between the generated image and text. I represents the generated image, text represents the formatted text, and P represents the probability, which is the probability that the visual question answering system will output "Yes." "Yes" indicates the answer the visual question answering system outputs when the image matches the text content.

[0102] It should be noted that the large-scale visual language model of this embodiment can be LVLM, which combines large-scale visual and language processing capabilities and can understand and generate multimodal data such as images and texts.

[0103] S105: If the score meets expectations, the knitted product image is output; if the score does not meet expectations, the potential noise variable is optimized by minimizing the distance between the text self-attention matrix and the cross-attention similarity matrix, the stable diffusion is optimized to generate the image again, and the process returns to S104.

[0104] Set a score threshold. If the score exceeds the score threshold, the knitted product image is output. If the score does not exceed the score threshold, the input text is first converted into a dense vector through a text editor and processed through a multi-head self-attention layer. Self-attention matrix The calculation formula is as follows:

[0105] Self-attention matrix The calculation formula for the elements in is as follows:

[0106] ;

[0107] in, represents the self-attention weight of the i-th token to the j-th token in the l-th layer and the h-th head of the text encoder; exp() represents the exponential function; represents the attention weight between the i-th token and the j-th token; represents the attention weight between the i-th token and the k-th token; represents the key vector of the i-th token in the l-th layer of the text encoder; represents the key vector of the jth token in the lth layer; represents the pre-trained weights in the l-th layer and h-th head of the text encoder; l represents the l-th layer of the text encoder, i represents the i-th token, j represents the j-th token, h represents the h-th head, and k represents the index for calculating the normalization factor.

[0108] In order to obtain the final text self-attention matrix , average the self-attention matrices of all layers and heads, remove the attention weights for special tags, and then renormalize them as follows:

[0109] ;

[0110] ;

[0111] in, represents the averaged self-attention matrix; Indicates the number of layers of the text encoder; Indicates the number of heads in each layer of the text encoder; Represents the renormalized self-attention matrix Elements of Represents the averaged self-attention matrix Elements of Indicates that the elements from the 2nd to the 3rd in the nth row are The sum of the elements is normalized, m is a column index; Indicates the length of the text sequence, that is, the number of tokens the text is segmented into, .

[0112] In the diffusion model, the text and image data are connected through multi-head cross attention; the a-th query vector in the cross attention layer l is defined as ;in, , N c represents the query length of a given cross-attention layer, D c is the hidden dimension of each head, H c Indicates the number of heads in the cross attention layer;

[0113] In the diffusion model, the text and image data are connected through multi-head cross attention; the a-th query vector in the cross attention layer l is defined as ;in, , N c represents the length of the query vector in the cross-attention layer, D c is the hidden dimension of each head, H c Indicates the number of heads in the cross-attention layer; Cross-attention map The element at the hth head is defined as:

[0114] ;

[0115] in, represents the a-th query vector in the l-th layer of the cross-attention layer; represents the key vector of the i-th token of the cross-attention layer l; represents the pre-trained weight matrix in the crisscross attention layer l and head h; express and Unnormalized attention weights between ; express and The normalized attention weights between them, l represents the index of the cross attention layer, h represents the head index in the l-th layer, a represents the a-th query vector in the query sequence, i represents the i-th token in the text, and j represents the j-th token in the text; represents a real matrix, Indicates the length of the text sequence, that is, the number of tokens the text is segmented into.

[0116] It should be noted that the multi-head attention mechanism is used when calculating the text self-attention matrix, so there are the concepts of the l-th layer and the h-th head; the multi-head attention mechanism is also used when calculating the cross-attention of text and images, so there are also the concepts of the l-th layer and the h-th head, and the two correspond to each other.

[0117] The cross attention maps between the head and all layers are averaged as follows:

[0118] ;

[0119] Among them, A represents the average attention matrix, L M Indicates the number of cross attention layers when the query sequence length is 256. represents the cross-attention map of the l-th layer and the h-th head, l represents the index of the cross-attention layer, and h represents the index of the head in the l-th layer.

[0120] Define a similarity matrix whose values ​​represent the similarity between the cross-attention graph pairs. The calculation formula of the elements in the similarity matrix S is as follows:

[0121] ;

[0122] Among them, among them, represents the elements in the similarity matrix S; represents the cosine similarity, which is used to measure the similarity between the key vectors of the i-th token and the j-th token; A represents the similarity between the key vectors of the i-th and k-th tags; ai A represents the attention weight of the a-th query vector to the key vector of the i-th tag in the cross attention matrix; aj N represents the attention weight of the a-th query vector to the j-th token key vector in the cross attention matrix; c Represents the length of the query vector in the crisscross attention layer.

[0123] Optimize the potential noise by minimizing the distance between the similarity matrix S of the cross attention map and the text self-attention matrix T , which is achieved by minimizing the following loss function:

[0124] ;

[0125] in, represents the loss function; represents the weight coefficient; The elements of the text self-attention matrix T represent the self-attention weights between the i-th and j-th tokens; represents the index; Represents the elements in the cross-attention similarity matrix.

[0126] Optimizing latent noise via gradient descent ,as follows:

[0127] ;

[0128] in, represents the learning rate; represents the potential noise variable; Represents the loss function right The gradient of , used to update the potential noise variable; represents the latent noise variable after being updated by gradient descent.

[0129] Through the above process, the grammatical relationship of text embedding is effectively transferred to the cross attention, so that the generated image and text semantics remain aligned.

[0130] In summary, see Figure 2 As shown, this embodiment first receives the user input of the knitted text "a sitting white knitted puppy". After fine-tuning the language model, the adaptively adjusted text is "a sitting cute white knitted puppy doll, with clear details, 4k". The adaptively adjusted text is input into the pre-trained Stable diffusion to generate an image, and the adjusted prompt is formatted as "Does this image match {a sitting cute white knitted puppy doll, with clear details, 4k.}? Answer yes or no." The formatted prompt and the generated image are input into the large visual language model LVLM for VQA scoring. If the score is greater than 0.75, the image is output. If the score is less than 0.75, the latent noise variable of Stable dffusion is optimized and the image is regenerated until the score is greater than 0.75. By dynamically adjusting the input text and optimizing the image generation model, the generation quality of knitted product images and user experience are significantly improved.

[0131] See also Figure 3 As shown, the present invention also discloses a knitted product image generation device based on text adjustment and visual feedback, comprising:

[0132] A text fine-tuning language model construction module 301 is used to construct a language model for adjusting input text based on a generative flow network GFlowNet;

[0133] The text adaptive adjustment module 302 is configured to input the knitting-related text received from the user into the trained language model to obtain text after adaptively adjusting the knitting-related text input from the user;

[0134] An image generation module 303 is configured to input the adaptively adjusted text into the trained text-image model Stablediffusion to generate a corresponding image;

[0135] A score acquisition module 304 is configured to format the adaptively adjusted text into text for determining whether the generated image matches the adaptively adjusted text, and input the formatted text and the generated image into a large-scale visual language model for visual question answering to obtain a score;

[0136] The latent noise variable optimization module 305 is used to determine if the score meets expectations and output the knitted product image; if the score does not meet expectations, the latent noise variable is optimized by minimizing the distance between the text self-attention matrix and the cross-attention similarity matrix, and the stable diffusion is optimized to generate the image again.

[0137] The specific implementation of each module of the device for generating a knitted product image based on text adjustment and visual feedback is the same as the method for generating a knitted product image based on text adjustment and visual feedback, and will not be repeated in this embodiment.

[0138] Figure 4 FIG. 1 is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. Figure 4 As shown, the electronic device of this embodiment includes: a processor 401 and a memory 402; wherein the memory 402 is used to store computer-executable instructions; and the processor 401 is used to execute the computer-executable instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant description of the above method embodiment.

[0139] Optionally, the memory 402 may be independent or integrated with the processor 401 .

[0140] When the memory 402 is independently provided, the electronic device further includes a bus 403 for connecting the memory 402 and the processor 401 .

[0141] An embodiment of the present invention further provides a computer storage medium, in which computer execution instructions are stored. When the processor 401 executes the computer execution instructions, the above method is implemented.

[0142] An embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 401, the above method is implemented.

[0143] In the embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not implemented. In addition, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or module, which may be electrical, mechanical or other forms.

[0144] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of these modules may be selected to implement the solution of this embodiment based on actual needs.

[0145] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each module may exist physically separately, or two or more modules may be integrated into a single unit. The units formed by the above modules may be implemented in the form of hardware or hardware plus software functional units.

[0146] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or processor 401 to perform some steps of the methods of various embodiments of the present application.

[0147] It should be understood that the processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), or application-specific integrated circuits (ASIC). A general-purpose processor may be a microprocessor, or the processor 401 may be any conventional processor 401. The steps of the method disclosed in the present invention may be directly implemented as being executed by the hardware processor 401, or may be implemented by a combination of hardware and software modules in the processor 401.

[0148] The memory 402 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.

[0149] Bus 403 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Bus 403 can be classified as an address bus, a data bus, a control bus, etc. For ease of illustration, the bus 403 in the drawings of this application is not limited to a single bus 403 or a single type of bus 403.

[0150] The storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0151] An exemplary storage medium is coupled to the processor 401, so that the processor 401 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor 401. The processor 401 and the storage medium can be located in an application-specific integrated circuit (ASIC). Of course, the processor 401 and the storage medium can also exist as discrete components in an electronic device or a host control device.

[0152] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating knitted product images based on text adjustment and visual feedback, characterized in that: include: S101, building a language model based on the generative flow network GFlowNet to adjust the input text; S102, inputting the knitting-related text received from the user into the trained language model to obtain text after adaptively adjusting the knitting-related text input from the user; S103, inputting the adaptively adjusted text into the trained text-to-graph model Stablediffusion to generate a corresponding image; S104, formatting the adaptively adjusted text into text for determining whether the generated image matches the adaptively adjusted text, inputting the formatted text and the generated image into a large visual language model for visual question answering to obtain a score; S105, if the score meets the expectation, output the knitted product image; if the score does not meet the expectation, optimize the potential noise variable by minimizing the distance between the text self-attention matrix and the cross-attention similarity matrix, and optimize the stable diffusion to generate the image again; The training process of the language model based on the generative flow network GFlowNet to adjust the input text is as follows: Use the pre-trained language model as the initial model P θ (y|x); where x represents the initial knitted text, y represents the improved text y generated based on the initial text x, P θ represents the pre-trained language model; Initialize the forward strategy P of the generated flow network GFlowNet F and stream function F θ ; Input the improved text y into the text-graph model Stablediffsion to generate an image; calculate the reward using the reward function r(x,y); Decompose the reward function r(x,y) into a step-by-step reward signal as follows: Among them, r(x,y 0:t ) represents the text segment y generated to step t 0:t The reward value of y 0:t represents the text segment generated up to step t; p ref (y 0:t |x) represents the pre-trained language model generating the text segment y 0:t The conditional probability of p ref (y|x) represents the conditional probability of the pre-trained language model generating the improved text y; β represents a hyperparameter; T represents the total number of steps to generate the text y; exp() represents an exponential function; Based on the reward signal, we extend the forward-backward balance objective to optimize GFlowNet, and the loss function is as follows: in, represents the overall loss function, which is used to measure the performance of the initial model in generating text y; Represents the local loss function at each step, which is used to evaluate the performance when generating text to step t+1; represents the value of the stream function at step t; represents the value of the stream function at step t+1; P F (y t+1 |x,y 0:t ; θ) represents the conditional probability of the forward strategy at step t+1; r(x,y 0:t ) represents the value of the reward function at step t; r(x,y 0:t+1 ) represents the value of the reward function at step t+1; θ represents the parameters of the language model; y 0:t+1 Indicates the text generated to step t+1; y t represents the word generated in step t; y t+1 Represents the word generated in step t+1; Minimize the above loss function and update the forward strategy P of the generated flow network GFlowNet F and stream function F θ , and update the parameters of the language model at the same time.

2. The knitted product image generation method based on text adjustment and visual feedback according to claim 1, characterized in that: The reward function r(x,y) is defined as follows: in, represents the expected value; i x Represents the diffusion model p from text to image according to text x ψ (·|x) generated image; i y Represents the diffusion model p from text to image according to text y ψ (·|y) generated image; p ψ (·|x) and p ψ (·|y) represents the diffusion model from text to image; r aes (i x ,i y ) represents aesthetic reward; r rel (x,i y ) represents the relevance reward.

3. The method for generating knitted product images based on text adjustment and visual feedback according to claim 2, characterized in that: Update the forward strategy P of the generated flow network GFlowNet F and stream function F θ , and after updating the parameters of the language model, it also includes: A flow reactivation mechanism is introduced, which periodically resets the last layer of the GFlowNet flow function every M steps; ultimately, the language model regulated by the generative flow network GFlowNet learns to generate text proportional to the reward.

4. The method for generating knitted product images based on text adjustment and visual feedback according to claim 1, characterized in that: The adaptively adjusted text is input into the trained text-to-graph model Stablediffusion to generate the corresponding image, as follows: I = Sd(prompt); Among them, prompt represents the text after adaptive adjustment; Sd represents the trained text-graph model Stablediffusion; I represents the generated image.

5. The method for generating knitted product images based on text adjustment and visual feedback according to claim 1, characterized in that: The adaptively adjusted text is formatted to determine whether the generated image matches the adaptively adjusted text. The formatted text and the generated image are input into a large visual language model for visual question answering to obtain a score, which includes: For the adaptively adjusted text, a language model is used to generate formatted text to determine whether the generated image matches the adaptively adjusted text; The generated image and formatted text are fed into a large visual language model to calculate the visual question answering score as follows: VQA(I,y)=P("Yes"∣I,text); Among them, VQA represents the score of visual question answering, which is an indicator used to measure the degree of alignment between the generated image and text; I represents the generated image; text represents the formatted text; P represents the probability of the visual question answering output "Yes", and "Yes" represents the answer output by the visual question answering when the image matches the content of the text.

6. The method for generating knitted product images based on text adjustment and visual feedback according to claim 1, characterized in that: If the score meets expectations, the knitted product image is output; if the score does not meet expectations, the potential noise variables are optimized by minimizing the distance between the text self-attention matrix and the cross-attention similarity matrix, and the image is generated again by optimizing Stable Diffusion. Specifically, the following steps are performed: Set a score threshold. If the score exceeds the threshold, the knitted product image is output. If the score does not exceed the threshold, the input text is first converted into dense vectors through a text editor and processed through a multi-head self-attention layer. Self-attention matrix The calculation formula for the elements in is as follows: in, represents the self-attention weight of the i-th token to the j-th token in the l-th layer and the h-th head of the text encoder; exp() represents the exponential function; ω ij represents the attention weight between the i-th token and the j-th token; ω ik represents the attention weight between the i-th token and the k-th token; represents the key vector of the i-th token in the l-th layer of the text encoder; represents the key vector of the jth token in the lth layer; represents the pre-trained weights in the lth layer and hth head of the text encoder; The self-attention matrices of all layers and heads are averaged, and the attention weights for special tags are removed, and then renormalized as follows: Among them, T′ represents the averaged self-attention matrix; L e Indicates the number of layers of the text encoder; H e T represents the number of heads in each layer of the text encoder; ij Represents the elements of the renormalized self-attention matrix Τ; Τ′ ij Represents the elements of the averaged self-attention matrix T′; Indicates normalizing the sum of the elements from the 2nd to the sth in the nth row, where m is a column index; s represents the length of the text sequence, that is, the number of tokens into which the text is segmented, n∈[1,s]; In the diffusion model, the text and image data are connected through multi-head cross attention; the a-th query vector in the cross attention layer l is defined as Where a=1,…,N c , N c represents the length of the query vector in the cross-attention layer, D c is the hidden dimension of each head, H c Indicates the number of heads in the cross attention layer; Cross-Attention Map The element at the hth head is defined as: in, represents the a-th query vector in the l-th layer of the cross-attention layer; represents the key vector of the i-th token of the cross-attention layer l; represents the pre-trained weight matrix in the crisscross attention layer l and head h; Ω ai express and Unnormalized attention weights between ; express and The normalized attention weights between them, l represents the index of the cross attention layer, h represents the head index in the l-th layer, a represents the a-th query vector in the query sequence, i represents the i-th token in the text, and j represents the j-th token in the text; represents a real matrix; The cross attention maps between the head and all layers are averaged as follows: Among them, A represents the average attention matrix, L M Indicates the number of cross attention layers when the query sequence length is 256. represents the cross attention map between the lth layer and the hth head; Define a similarity matrix whose values ​​represent the similarity between the cross-attention graph pairs; the calculation formula of the elements in the similarity matrix S is as follows: Among them, S ij Represents the elements in the similarity matrix S; C ij represents the cosine similarity, which is used to measure the similarity between the key vectors of the i-th token and the j-th token; C ik A represents the similarity between the key vectors of the i-th and k-th tags; ai A represents the attention weight of the a-th query vector to the key vector of the i-th tag in the cross attention matrix; aj N represents the attention weight of the a-th query vector to the j-th token key vector in the cross attention matrix; c represents the length of the query vector in the cross-attention layer; Optimize the potential noise z by minimizing the distance between the similarity matrix S of the cross attention map and the text self-attention matrix T t , which is achieved by minimizing the following loss function: in, represents the loss function; ρ i represents the weight coefficient; T ij Represents the element of the text self-attention matrix T, which represents the self-attention weight between the i-th and j-th tokens; γ represents the exponent; S ij (z t ) represents the elements in the cross-attention similarity matrix; Optimize the latent noise z by gradient descent t ,as follows: Where α represents the learning rate; represents the potential noise variable; Represents the loss function z t The gradient of z is used to update the potential noise variable; t ′ represents the latent noise variable after being updated by gradient descent; Through the above process, the grammatical relationship of text embedding is effectively transferred to the cross attention, so that the semantics of the regenerated image and text remain aligned.

7. The method for generating knitted product images based on text adjustment and visual feedback according to claim 1, characterized in that: After optimizing Stablediffusion to generate the image again, the process further includes: returning to S104 to obtain the score again.

8. A knitted product image generation device based on text adjustment and visual feedback, characterized in that: The method for generating a knitted product image according to any one of claims 1 to 7 comprises: The text fine-tuning language model construction module is used to build a language model based on the generative flow network GFlowNet to adjust the input text; a text adaptive adjustment module, configured to input the knitting-related text received from the user into the trained language model to obtain text after adaptively adjusting the knitting-related text input from the user; The image generation module is used to input the adaptively adjusted text into the trained text-to-graph model Stablediffusion to generate the corresponding image; A score acquisition module is used to format the adaptively adjusted text into text to determine whether the generated image matches the adaptively adjusted text, and input the formatted text and the generated image into a large visual language model for visual question answering to obtain a score; The latent noise variable optimization module is used to determine if the score meets expectations and output the knitted product image; if the score does not meet expectations, the latent noise variable is optimized by minimizing the distance between the text self-attention matrix and the cross-attention similarity matrix, and the image is generated again by optimizing the Stable Diffusion.

Citation Information

Patent Citations

  • Cross-modal image-text retrieval method fusing semantic similarity embedding and metric learning

    CN114817596A

  • Text image generation method based on attention modulation and text redescription

    CN119991855A