Image generation method and device driven by image retrieval
Through the image retrieval-driven method, the problem of lack of effective image retrieval and real-time update resources in image generation is solved, and more efficient and controllable image generation is achieved, which improves the diversity, accuracy and consistency of generated images.
Patent Information
- Application Number
- CN202510111014.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art lacks an effective image retrieval mechanism in image generation, which leads to the generation of images that may lack accurate context information, affecting image quality and diversity, and the image resources cannot be updated in real time, resulting in the generated content being outdated or lack of innovation.
Through the image retrieval-driven method, the text needs input by the user are obtained, the text needs are understood and expanded, the keywords are extracted for image retrieval, the images are filtered and re-screened, and the image common description is integrated to generate image prompt words, and finally the image that meets the user's needs is generated.
It improves the diversity and accuracy of image generation, reduces hallucination problems, ensures that the generated images are consistent with the text description, can update image resources in real time, and the generated image content is more innovative and time-consuming.
Smart Images

Figure CN120179846A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to an image generation method and device driven by image retrieval. Background Art
[0002] Existing retrieval-augmented generation (RAG) technologies are mainly used for recall and augmented generation based on text knowledge bases, and are widely used in tasks such as question-answering systems and text generation. This technology improves the accuracy and richness of the model by introducing external document information into the generation process, and is mainly optimized for text data. However, for visual generation tasks such as text-to-image, the application of existing RAG technologies is still relatively scarce, and the following problems are faced:
[0003] 1. Technical barriers between image retrieval and generation: Traditional RAG technologies mainly rely on text recall, and retrieve relevant information from a large number of text databases to enhance the context understanding in generation tasks. However, in the text-to-image task, how to effectively retrieve image content and combine it with relevant models to enhance the generation effect is still an unsolved problem. Existing technologies often lack a direct text+image recall mechanism, which may lead to a lack of accurate context information during image generation, thus affecting the quality of the generated images and even possibly producing results inconsistent with the text description.
[0004] 2. Difficulty in cross-modal combination of images and text: For text-to-image tasks, existing technologies usually rely on simple text conditions and the generation ability of the model itself to create images, but do not fully utilize external image information as a reference, thus limiting the diversity and accuracy of the generated images. The lack of an image retrieval mechanism causes the generation model to be unable to obtain relevant visual information from diverse image data, which may in turn result in the generated images lacking sufficient variability or showing low accuracy when dealing with complex scenes.
[0005] 3. Lack of real-time updated image resources: Many image generation models rely on historical datasets, which are fixed when constructed and cannot quickly adapt to emerging scenarios or requirements. For example, image generation models in certain specific fields may be unable to generate accurate up-to-date content due to the timeliness of the datasets. The limitation of such static resources makes the image generation system unable to effectively reflect current dynamic changes, resulting in the content of the generated images being possibly outdated or lacking innovation. Summary of the Invention
[0006] The present disclosure provides an image generation method and device driven by image retrieval.
[0007] According to a first aspect of the present disclosure, an image generation method driven by image retrieval is provided.
[0008] The method includes:
[0009] Obtaining the text requirements of the image to be generated input by the user;
[0010] Understanding and expanding the text requirements to obtain expanded requirements; and extracting keywords from the text requirements, and driving image retrieval according to the extracted keywords to obtain a primary image set;
[0011] According to the expanded requirements, performing primary screening and secondary screening on the primary image set until the number of images selected is one, and taking the finally selected image as the target image;
[0012] Fusing the expanded requirements and the common descriptions of the images obtained in the secondary screening stage to obtain an image prompt;
[0013] Generating an image that meets the user's requirements according to the image prompt and the target image.
[0014] For the aspects and any possible implementation manners as described above, a further implementation manner is provided.
[0015] Performing primary screening on the primary image set according to the expanded requirements, including:
[0016] Grouping and label assignment are performed on all the images in the primary image set to obtain multiple groups of labeled images; the labels are used to identify the group numbers and in-group encodings to which the respective images belong;
[0017] Inputting the expanded requirements and the multiple groups of images into a vision large model, and outputting the labels of the images that match the expanded requirements as the primary screening labels;
[0018] Determining the corresponding images from the primary image set according to the primary screening labels as the initial image set in the secondary screening stage;
[0019] The number of primary screening times is one.
[0020] For the aspects and any possible implementation manners as described above, a further implementation manner is provided.
[0021] Performing secondary screening on the primary image set according to the expanded requirements, including:
[0022] Grouping and label assignment are performed on the initial image set obtained in the primary screening stage to obtain at least one group of labeled images; the labels are used to identify the group numbers and in-group encodings to which the respective images belong;
[0023] Inputting the expanded requirements and the at least one group of images into a vision large model, and outputting the labels of the images that match the expanded requirements as the secondary screening labels;
[0024] Determine the corresponding image from the corresponding image set according to the re-screening label, and use it as the image input to the visual large model in the next round of the re-screening stage;
[0025] The number of re-screening times is at least once.
[0026] For the above aspects and any possible implementation manners, a further implementation manner is provided.
[0027] The commonality description is the common feature of the images obtained from at least one round of re-screening; the common features include scenes, emotions, and object categories;
[0028] When inputting the expansion requirement and the at least one set of images into the visual large model and outputting the label of the image matching the expansion requirement as the re-screening label, it further includes:
[0029] Output the commonality description of the at least one set of images.
[0030] For the above aspects and any possible implementation manners, a further implementation manner is provided. The image set corresponding to the re-screening label is: the image set for which the current grouping and label assignment are performed in the current re-screening round.
[0031] For the above aspects and any possible implementation manners, a further implementation manner is provided.
[0032] Fuse the expansion requirement and the commonality description of the images obtained in the re-screening stage to obtain a picture prompt, including:
[0033] Use the language large model to fuse the common features of the images obtained from at least one round of re-screening with the expansion requirement to obtain a picture prompt.
[0034] For the above aspects and any possible implementation manners, a further implementation manner is provided. The generation of an image that meets the user's needs according to the picture prompt and the target image includes:
[0035] Input the picture prompt and the target image into the text-to-image model, and output the first image that meets the user's needs;
[0036] Through the visual large model, verify the requirement matching of the first image according to the picture prompt to obtain a matching verification result;
[0037] Determine the final image that meets the user's needs according to the matching verification result.
[0038] For the above aspects and any possible implementation manners, a further implementation manner is provided.
[0039] The matching verification result includes passing the matching verification and failing the matching verification;
[0040] Determining the image that ultimately meets the user's needs based on the matching verification result includes:
[0041] If the matching verification is passed, the image corresponding to the passed matching verification is used as the image that ultimately meets the user's needs;
[0042] If the matching verification fails, obtain the improvement suggestions output by the vision large model, update the picture prompt according to the improvement suggestions, and regenerate the image according to the updated picture prompt.
[0043] According to a second aspect of the present disclosure, an image generation device driven by image retrieval is provided.
[0044] The device includes:
[0045] A data acquisition module for acquiring the text requirements of the picture to be generated input by the user;
[0046] An image generation module for understanding and expanding the text requirements to obtain expanded requirements; and extracting keywords from the text requirements, and driving image retrieval according to the extracted keywords to obtain a primary image set;
[0047] The image generation module is further configured to perform a primary screening and a secondary screening on the primary image set according to the expanded requirements until the number of the screened images is one, and use the finally screened image as the target image;
[0048] The image generation module is further configured to fuse the common descriptions of the expanded requirements and the images obtained in the secondary screening stage to obtain a picture prompt;
[0049] The image generation module is further configured to generate an image that meets the user's needs according to the picture prompt and the target image.
[0050] According to a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: a memory and a processor, and a computer program is stored on the memory, and when the processor executes the program, the method as described above is implemented.
[0051] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method as described in the first aspect of the present disclosure is implemented.
[0052] The image retrieval-driven image generation method and apparatus provided by the embodiments of the present disclosure utilize a large model to analyze user requirements, locate visual resources, and optimize the generation effect of the text-to-image model through examples, achieving more efficient and controllable image retrieval and screening. The entire technical framework effectively solves the hallucination problem existing in the prior art in image generation, improves the diversity and accuracy of the generated images, and significantly enhances the accuracy and consistency of the generation results, especially in complex scenarios.
[0053] It should be understood that the content described in the summary of the invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings
[0054] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. In the drawings, the same or similar reference numerals represent the same or similar elements, where:
[0055] Figure 1 shows a flowchart of an image retrieval-driven image generation method according to an embodiment of the present disclosure;
[0056] FIG. 2 shows a flowchart of the work process in the preliminary screening stage according to an embodiment of the present disclosure, where (a) represents a flowchart of one round of the work process in the preliminary screening stage; (b) represents a flowchart of the overall work process in the preliminary screening stage;
[0057] Figure 3 shows a flowchart of the work process in the re-screening stage according to an embodiment of the present disclosure;
[0058] FIG. 4 shows a flowchart of image generation after information fusion according to an embodiment of the present disclosure, where (a) represents a general flowchart of text commonality and extended text image generation in a single-round re-screening stage; (b) represents a general flowchart of text commonality and extended text image generation in a multi-round re-screening stage;
[0059] Figure 5 shows a flowchart of the verification and improvement link of the picture according to an embodiment of the present disclosure;
[0060] Figure 6 shows a block diagram of an image retrieval-driven image generation apparatus according to an embodiment of the present disclosure;
[0061] Figure 7 shows a schematic block diagram of an exemplary electronic device capable of implementing the embodiments of the present disclosure. Detailed Description of the Embodiments
[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0063] In addition, the term "and / or" in this document merely describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after.
[0064] The present disclosure reduces the hallucination problem by combining image retrieval and RAG technology, thereby improving the quality of the finally generated images. Among them, hallucination is a common problem existing in all text-to-image models, specifically manifested as inaccuracies or misleadingness in the image content. The most common hallucination problem is a logical problem: the spatial relationship, logic, or physical attributes of the objects in the image do not conform to common sense.
[0065] Figure 1 The flowchart of an image retrieval-driven image generation method 100 according to an embodiment of the present disclosure is shown. The method 100 includes:
[0066] Step 110, obtaining the text requirement of the picture to be generated input by the user.
[0067] In some embodiments, obtaining the text requirement of the picture to be generated input by the user, for example, the input:
[0068] {
[0069] ‘text’: Please help me generate a picture with the content: Visiting the Thatched Cottage Three Times
[0070] }
[0071] Step 120, understanding and expanding the text requirement to obtain an expanded requirement; and extracting keywords from the text requirement, and driving image retrieval according to the extracted keywords to obtain a primary image set.
[0072] In some embodiments, extracting keywords from the text requirement input by the user, and outputting:
[0073] {
[0074] ‘keywords’: Visiting the Thatched Cottage Three Times
[0075] }
[0076] In some embodiments, the text requirements input by the user are understood and expanded. Input:
[0077] {
[0078] ‘text’: Please help me generate a picture with the content: Visiting the Thatched Cottage Three Times
[0079] }
[0080] Output:
[0081] {
[0082] ‘expanded’: Deep in a dense bamboo forest, Zhu's thatched cottage stands quietly at the foot of the mountain. Liu, Guan, and Zhang stand in front of the thatched cottage, with respectful and determined expressions. Liu is about to pay his third visit, hoping to persuade Zhu to come out and help. The sky is gloomy, and the gentle breeze stirs the bamboo leaves, adding a bit of solemn atmosphere. The thatched cottage is simple and tidy, surrounded by green plants and trees, creating a peaceful atmosphere for a hermit.
[0083] }
[0084] In some embodiments, the most relevant pictures among all the pictures corresponding to the keywords are obtained through search, for example, 20 pictures are used as preselected pictures. Then, the pictures that cannot be downloaded are filtered out and deleted, and all the pictures that can be obtained are used as candidate pictures.
[0085] By expanding and polishing the user input, the accuracy and multi-dimensional understanding ability of the input description are significantly improved, providing clearer guidance for subsequent generation.
[0086] Step 130, according to the expanded requirements, conduct primary screening and secondary screening on the primary image set until only one image is selected, and the finally selected image is used as the target image.
[0087] In some embodiments, according to the expanded requirements, the primary screening of the primary image set includes: grouping all the images in the primary image set and assigning labels, obtaining multiple groups of images with labels; the labels are used to identify the group numbers and the in-group codes of each image; inputting the expanded requirements and the multiple groups of images into a vision large model, and outputting the labels of the images that match the expanded requirements as the primary screening labels; determining the corresponding images from the primary image set according to the primary screening labels as the initial image set in the secondary screening stage; the number of primary screening times is one.
[0088] In some embodiments, for example, the candidate images are divided into multiple rounds, with at least 1 and at most 5 images in each round. When the number of candidate images is greater than or equal to 5, ensure that there are 5 images in that round; if there are candidate images currently but the number of candidate images is less than 5, the number of images in that round can be less than 5. As shown in Figure 2, all images will be assigned a label (A to E) and saved in the form of key-value pairs (for example: {"A": image 1, "B": image 2}). The visual large model will analyze the translated extended text and compare it with all the images in that round one by one, select the image that best meets the user's needs, and return its label (for example, "A"), find the corresponding image through the key-value pair for this label, and temporarily save the image until the re-screening stage. It is equivalent to randomly grouping the images manually or by machine, with each group having, for example, 5 images, and the label is used to determine which image in which group.
[0089] Among them, an example code of the key-value pair is as follows:
[0090] {
[0091] ‘A’: image 1,
[0092] ‘B’: image 2,
[0093] ‘C’: image 3,
[0094] ‘D’: image 4,
[0095] ‘E’: image 5,
[0096] }
[0097] In some embodiments, according to the expansion requirements, the initial image set is re-screened, including: grouping and label assignment for the initial image set obtained in the initial screening stage to obtain at least one group of images with labels; the label is used to identify the group number and the in-group coding to which each image belongs; inputting the expansion requirements and the at least one group of images into the visual large model, and outputting the label of the image that matches the expansion requirements as the re-screening label; determining the corresponding image from the corresponding image set according to the re-screening label as the image input into the visual large model in the next round of the re-screening stage; the number of re-screening times is at least once.
[0098] In some embodiments, when inputting the expansion requirements and the at least one group of images into the visual large model and outputting the label of the image that matches the expansion requirements as the re-screening label, it further includes: outputting a common description of the at least one group of images.
[0099] In some embodiments, the image set corresponding to the re-screening label is: in the current re-screening round, the image set for the current grouping and label assignment.
[0100] In some embodiments, the secondary screening is to select at least one more picture. The visual large model will divide the pictures initially selected into several rounds, with a maximum of five pictures and at least one picture in each round. Then, the model will select one picture that best meets the requirements in each round, compare these selected pictures with each other, and finally select the one that best meets the requirements to prepare for generating the final picture. In each round, the visual large model uses its picture description ability to generate similarities of these pictures based on the input text and the pictures in the current round, form a common description and save it for reference in subsequent picture generation. As Figure 3 shown, if there are more than 5 pictures in the secondary screening stage, they are divided into several rounds, with a maximum of 5 pictures in each round. For example, if there are still 20 pictures in the secondary screening stage, they are divided into 4 rounds. First, select one best picture in each round, and then continue to select from these selected pictures until only one picture remains. If there are less than 5 pictures in the secondary screening stage, a direct secondary screening is sufficient (for example, in the current example, there are 20 pictures in the primary screening stage, so there are at most 4 pictures in the secondary screening stage, and thus a single secondary screening is enough). It should be noted that the number of pictures obtained by searching keywords can be dynamically adjusted. The specific number is not limited to 20, for example, it can be set to 10 to 100, or determined by system configuration, user requirements, algorithm optimization results, etc. All picture number settings that can achieve the same retrieval effect fall within the scope of protection of this patent.
[0101] Among them, the sample code for secondary screening is as follows:
[0102] Prompt:
[0103]
[0104] Input:
[0105] {
[0106] ‘text’:’Deep within a dense bamboo forest, Moumou Liang's thatched cottage quietly sits at the foot of a mountain. Liu Mou, Guan Mou, and Zhang Mou stand before the cottage, their expressions respectful and resolute. Liu Mou is preparing for his third visit, hoping to persuade Moumou Liang to come out of seclusion and join him. The sky is overcast, and a gentle breeze stirs the bamboo leaves, adding to the solemn atmosphere. The cottage is simple yet tidy, surrounded by lush greenery, creating a tranquil and reclusive ambiance.’
[0107] ‘A’:’Picture 1’,
[0108] ‘B’:’Picture 2’,
[0109] ‘C’:’Picture 3’,
[0110] ‘D’:’Picture 4’,
[0111] ‘E’:’Picture 5’,
[0112] }
[0113] Output:
[0114] {
[0115] ‘option’:’A’
[0116] 'summarize': 'These images collectively portray the classic scene from the Three Kingdoms period, where Liu Mou visits the recluse Moumou Liang three times. The characters typically include Liu Mou, Guan Mou, Zhang Mou, and Moumou Liang, with the background often showing a simple thatched cottage or a mountainous setting, reflecting an atmosphere of respect, humility, and the earnest desire to recruit talent. The overall style is likely traditional, with subtle color tones and a focus on details and the expressions of the characters, highlighting the reverence and sincerity in the three visits.'
[0117] }
[0118] In some embodiments, the vision large model will extract the key information in the text and the common features between the images in the current round based on the translated and extended text, such as the scene, emotion, or object category. Through precise retrieval and multi-round screening of visual resources, it provides reliable and up-to-date input materials for the text-to-image task, while reducing the hallucination problem and improving the relevance of the generated results. Moreover, by extracting the common descriptions of multiple images, the information is fused into a more diverse and expressive input description, significantly broadening the application scope of text-to-image.
[0119] Step 140, fuse the expansion requirement and the common description of the images obtained in the re-screening stage to obtain an image prompt.
[0120] In some embodiments, the common description is the common features of the images obtained through at least one round of re-screening; the common features include the scene, emotion, and object category.
[0121] In some embodiments, the common descriptions of the image obtained in the expansion requirement and the re-screening stage are fused to obtain a picture prompt, including: fusing the common features of the images obtained by at least one round of re-screening with the expansion requirement through a language model to obtain a picture prompt.
[0122] In some embodiments, the language model fuses these information into a summary statement with clear logic and smooth grammar according to the user's requirements and the common picture descriptions output in each round of the re-screening stage, which can not only accurately reflect the core content of the text and pictures, but also meet the user's requirements. As shown in Figure 4.
[0123] For example, the prompts for forming a summary statement with clear logic and smooth grammar are as follows:
[0124] Input:
[0125] {
[0126] ‘text’:’Deep within a dense bamboo forest,Moumou Liang's thatchedcottage quietly sits at the foot of a mountain.Liu Mou,Guan Mou,and Zhang Moustand before the cottage,their expressions respectful and resolute.Liu Mou ispreparing for his third visit,hoping to persuade Moumou Liang to come out ofseclusion and join him.The sky is overcast,and a gentle breeze stirs thebamboo leaves,adding to the solemn atmosphere.The cottage is simple yet tidy,surrounded by lush greenery,creating a tranquil and reclusive ambiance.’,
[0127] ‘description1’:’From the pictures,There are three people in a cottage,seeming that they are waiting for someone.The pictures are in a solemn atmosphere...’,
[0128] ‘description2’:’......’, ......
[0130] }
[0131] Output:
[0132] {
[0133] “text”:’In a serene and solemn bamboo forest setting,Moumou Liang's modest cottage sits quietly at the mountain's foot.Liu Mou,accompanied by Guan Mou and Zhang Mou,stands respectfully before the cottage,preparing for a third visit to persuade Moumou Liang to leave seclusion.The tranquil environment,marked by overcast skies and the gentle rustle of bamboo leaves,adds to the atmosphere of anticipation.’
[0134] }
[0135] Step 150, generate an image that meets the user's requirements according to the picture prompt and the target image.
[0136] In some embodiments, generating an image that meets the user's requirements according to the picture prompt and the target image includes: inputting the picture prompt and the target image into a text-to-image model to output a first image that meets the user's requirements; through a vision large model, according to the picture prompt, perform requirement matching verification on the first image to obtain a matching verification result; determine the final image that meets the user's requirements according to the matching verification result.
[0137] In some embodiments, the matching verification result includes passing the matching verification and failing the matching verification. Determining the image that finally meets the user's needs according to the matching verification result includes: if the matching verification is passed, using the image corresponding to the passed matching verification as the image that finally meets the user's needs; if the matching verification fails, obtaining the improvement suggestions output by the vision large model, updating the picture prompt according to the improvement suggestions, and regenerating the image again according to the updated picture prompt.
[0138] In some embodiments, a vision large model is used to check whether the text information accepted by the text-to-image model and the pictures it outputs are consistent. Specifically, the vision large model obtains the text information accepted by the text-to-image model and the pictures it outputs, and analyzes their consistency in semantic and visual features. If they are consistent, it returns True; if they are inconsistent, it returns False and also returns where improvement is needed. When the vision large model returns False, the next step is to improve the picture. Each time, the text-to-image model will, according to the improvement suggestions of the large model and in combination with the text information (including a logically clear and grammatically correct summary statement with common descriptions), use the previous round of pictures as the input set to regenerate a new picture. Then, the vision large model will check this new picture to see if it meets the requirements. If it meets the requirements, it returns True; if it does not meet the requirements, it returns False and indicates where improvement is needed. This process continues until the picture meets all the requirements (i.e., the vision large model finally returns True). As Figure 5 shown. The above uses the method of having the model return True / False for determination. It should be noted that the judgment method of this application can be a hard judgment to return True / False or a soft judgment. For example, the large model scores and judges whether the text and the picture meet the standards according to a set threshold.
[0139] Among them, the verification example code is as follows:
[0140] 1) For example, the current prompt for forming a logically clear and grammatically correct summary statement is as follows:
[0141] Input: {
[0142] ‘text’:’Deep within a dense bamboo forest, Moumou Liang's thatched cottage quietly sits at the foot of a mountain. Liu Mou, Guan Mou, and Zhang Mou stand before the cottage, their expressions respectful and resolute. Liu Mou is preparing for his third visit, hoping to persuade Moumou Liang to come out of seclusion and join him. The sky is overcast, and a gentle breeze stirs the bamboo leaves, adding to the solemn atmosphere. The cottage is simple yet tidy, surrounded by lush greenery, creating a tranquil and reclusive ambiance.’
[0143] ‘image’:’xxx’(base64 format)
[0144] }
[0145] Output: {
[0146] ‘similar’: True / False
[0147] ‘reason’:’The picture completely matches the text / The text requires stating that there are three people, but only two people are shown in the picture’
[0148] }
[0149] Then, at this time, the text-to-image model will, according to the improvement suggestions of the large model, combine the text information (including a logically clear and grammatically correct summary statement with common descriptions), and use the previous round of pictures as the input set to regenerate a new picture. The sample code for generating pictures is as follows:
[0150] Input:
[0151] 1. Text: ‘Deep within a dense bamboo forest, Moumou Liang's thatched cottage quietly sits at the foot of a mountain. Liu Mou, Guan Mou, and Zhang Mou stand before the cottage, their expressions respectful and resolute. Liu Mou is preparing for his third visit, hoping to persuade Moumou Liang to come out of seclusion and join him. The sky is overcast, and a gentle breeze stirs the bamboo leaves, adding to the solemn atmosphere. The cottage is simple yet tidy, surrounded by lush greenery, creating a tranquil and reclusive ambiance.’ Improvement: ‘The text specifies three people, but the image only depicts two.’
[0152] 2. The previous round of pictures
[0153] Output:
[0154] New pictures
[0155] Then, the vision large model will check this new picture to see if it meets the requirements. Until the generated picture meets the requirements, it is considered that the finally generated picture meets the user's needs. Based on this, through the verification and improvement mechanism of the generated picture, a closed-loop optimization process is formed to ensure that the output result is accurate and of high quality, fully meeting the user's needs.
[0156] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present disclosure.
[0157] The above is an introduction to the method embodiments. The following further illustrates the solution of the present disclosure through device embodiments.
[0158] Figure 6 Fig. shows a block diagram of an image retrieval-driven image generation device 600 according to an embodiment of the present disclosure. As Figure 6 shown, the device 600 includes:
[0159] A data acquisition module 610, configured to acquire the text requirement of the picture to be generated input by the user;
[0160] An image generation module 620, configured to understand and expand the text requirement to obtain an expanded requirement; and extract keywords from the text requirement, and drive image retrieval according to the extracted keywords to obtain a primary image set;
[0161] The image generation module 620 is further configured to perform a primary screening and a secondary screening on the primary image set according to the expanded requirement until only one image is selected, and use the finally selected image as the target image;
[0162] The image generation module 620 is further configured to fuse the common descriptions of the expanded requirement and the images obtained in the secondary screening stage to obtain a picture prompt;
[0163] The image generation module 620 is further configured to generate an image that meets the user's requirements according to the picture prompt and the target image.
[0164] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0165] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0166] Figure 7 Fig. shows a schematic block diagram of an electronic device 700 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0167] The electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in the ROM 702 or a computer program loaded from the storage unit 708 into the RAM 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The I / O interface 705 is also connected to the bus 704.
[0168] Multiple components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0169] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of method 100 described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute method 100 in any other appropriate manner (e.g., by means of firmware).
[0170] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0171] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0172] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include electrical connections based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0173] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0174] The systems and techniques described here can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described here), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0175] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0176] It should be understood that the various forms of the processes shown above can be reordered, added to, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0177] The above - described specific embodiments do not limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An image generation method driven by image retrieval, characterized in that: include: Get the text requirements of the image to be generated entered by the user; Understanding and expanding the text requirements to obtain expanded requirements; and, extracting keywords from the text requirements, and driving image retrieval based on the extracted keywords to obtain a preliminary image set; According to the expansion requirement, the preliminary selected image set is preliminarily screened and rescreened until only one image is screened out, and the finally screened out image is used as the target image; The expansion requirement and the common description of the image obtained in the re-screening stage are integrated to obtain the picture prompt words; An image that meets user needs is generated according to the picture prompt word and the target image.
2. The method according to claim 1, characterized in that According to the expansion requirements, the preliminary image set is screened, including: Grouping and labeling all images in the preliminary image set to obtain multiple groups of images with labels; the labels are used to identify the group number and the intra-group code to which each image belongs; Inputting the expansion requirement and the multiple groups of images into a visual macro model, and outputting labels of images matching the expansion requirement as preliminary screening labels; Determining corresponding images from the preliminary selection image set according to the preliminary screening labels as the initial image set for the rescreening stage; The number of initial screenings is once.
3. The method according to claim 2, characterized in that According to the expansion requirements, the preliminary image set is re-screened, including: The initial image set obtained in the initial screening stage is grouped and labeled to obtain at least one group of images with a label; the label is used to identify the group number and the intra-group code to which each image belongs; Inputting the expansion requirement and the at least one group of images into a visual macro model, and outputting labels of images matching the expansion requirement as re-screening labels; Determine a corresponding image from the corresponding image set according to the rescreening label as an image for inputting the visual macro model in the next round of the rescreening stage; The number of rescreening is at least once.
4. The method according to claim 3, characterized in that: The common description is the common features of the images obtained from at least one round of re-screening; the common features include scenes, emotions, and object categories; When the expansion requirement and the at least one group of images are input into the visual macro model and the labels of the images matching the expansion requirement are output as the re-screening labels, the method further includes: A commonality description of the at least one group of images is output.
5. The method according to claim 3, characterized in that: The image set corresponding to the rescreening label is: the image set that is currently grouped and labeled in the current rescreening round.
6. The method according to claim 4, characterized in that The common description of the image obtained in the expansion requirement and the re-screening stage is integrated to obtain the picture prompt words, including: The common features of the images obtained from at least one round of re-screening are integrated with the expansion requirements through a large language model to obtain picture prompt words.
7. The method according to claim 1, characterized in that The step of generating an image that meets the user's needs according to the picture prompt word and the target image includes: Input the picture prompt word and the target image into a text-based image model, and output a first image that meets the user's needs; Through the visual macro model, according to the picture prompt word, the first image is verified for demand matching to obtain a matching verification result; An image that ultimately meets the user's needs is determined based on the matching verification result.
8. The method according to claim 7, characterized in that The matching verification result includes passing the matching verification and failing the matching verification; Determining the image that ultimately meets the user's needs according to the matching verification result includes: If the matching verification is passed, the image corresponding to the matching verification will be used as the final image that meets the user's needs; If the matching verification fails, improvement suggestions output by the large visual model are obtained, the image prompt words are updated according to the improvement suggestions, and the image is regenerated according to the updated image prompt words.
9. An image generation device driven by image retrieval, characterized in that: include: The data acquisition module is used to obtain the text requirements of the image to be generated input by the user; An image generation module is used to understand and expand the text requirements to obtain expanded requirements; and to extract keywords from the text requirements and drive image retrieval based on the extracted keywords to obtain a preliminary image set; The image generation module is further used to perform preliminary screening and rescreening on the preliminary selected image set according to the expansion requirement, until only one image is screened out, and then use the finally screened out image as the target image; The image generation module is also used to merge the expansion requirements and the common description of the image obtained in the re-screening stage to obtain the picture prompt words; The image generation module is also used to generate an image that meets the user's needs based on the picture prompt word and the target image.
10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 8.