Image generation method, system and equipment based on large model and medium

By generating and filtering image prompt word collections through big models, the problem of mismatch between the image generated by the literary image model and the text is solved, and high-quality image generation is achieved.

CN120339428APending Publication Date: 2025-07-18GUANGZHOU YUNCONG INFORMATION TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510398285.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There is a big difference between the images generated by the existing literary and biographical graphics model and the input text, which leads to poor quality of the generated images, making it difficult to formulate clear rules to determine the quality of the text, and it is difficult for users to operate.

Method used

Use the big model to generate candidate prompt words, and generate images for each prompt word, filter candidate prompt words that meet preset requirements, build an image prompt word set, and improve the quality of prompt words through vector encoding and mapping relationship tables.

Benefits of technology

It reduces the difficulty of user operation, improves the quality and aesthetics of generated images, and solves the problem of mismatch in graphics and text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339428A_ABST
    Figure CN120339428A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image generation, particularly provides an image generation method, system and device based on a large model, and a medium, and aims to solve the problem of image-text mismatching of an image generated by a text generation model according to an input text. In order to achieve the purpose, the image cue word generation method based on the large model comprises the steps that a plurality of candidate cue words are generated according to preset information by means of the large model, a group of images corresponding to each candidate cue word are generated, the candidate cue words meeting preset requirements are screened according to the group of images corresponding to each candidate cue word, and the candidate cue words are generated according to the screened candidate cue words. And obtaining an image cue word set according to the candidate cue words meeting the preset requirements. By using the method to obtain the image cue word set, the operation difficulty of a user is greatly reduced, and the quality of the cue words input into the text generation graph model is effectively improved, so that the problem of image-text mismatching when the text generation graph model generates the image according to the input text is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image generation technology, and in particular to a large model-based image generation method, system, device and medium. Background Art

[0002] In image generation application scenarios, such as the automatic generation of marketing materials, it is necessary to combine the text input by the user, that is, the prompt words of the text-generated graph model, to generate the corresponding image. However, the images generated by the existing text-generated graph models are very uncontrollable. The generated images may be very different from the input text, and even have distortion, deformation, and incomplete image elements that greatly affect the appearance. This method of generating images based on text has high requirements for the input text. The quality of the generated image is very dependent on the input text, and it is difficult to formulate clear rules to determine the quality of the text.

[0003] Accordingly, the art needs a new large-model-based image prompt word generation solution to solve the above problems. Summary of the invention

[0004] In order to overcome the above-mentioned defects, the present application is proposed to solve or at least partially solve the problem of mismatch between image and text when a text-to-image model generates an image based on input text.

[0005] In a first aspect, a method for generating image prompt words based on a large model is provided, the method comprising: using a large model to generate a number of candidate prompt words according to preset information; generating a corresponding set of images for each candidate prompt word; screening candidate prompt words that meet preset requirements according to a set of images corresponding to each candidate prompt word; and obtaining an image prompt word set according to the candidate prompt words that meet the preset requirements.

[0006] In a technical solution of the above-mentioned method for generating image prompt words based on a large model, the method further includes: performing vector encoding on each image prompt word in the image prompt word set to obtain a corresponding prompt word vector.

[0007] In a technical solution of the above-mentioned method for generating image prompt words based on a large model, the method further includes: constructing a mapping relationship table according to the corresponding relationship between the image prompt words and the prompt word vectors.

[0008] In a technical solution of the above-mentioned large-model-based image prompt word generation method, the preset information includes product features, and the preset requirements include that the image quality meets preset standards and that the matching degree between the image and the product features meets preset standards.

[0009] In a second aspect, there is provided an image generation method based on the large model-based image prompt generation method in the above first aspect or any corresponding technical solution thereof, including: obtaining text information input by a user and an image prompt set; selecting the image prompt most matching the text information from the image prompt set; and generating an image using a text-to-image model according to the most matching image prompt.

[0010] In a technical solution of the above image generation method, the selecting the image prompt most matching the text information from the image prompt set includes: converting the text information into a text vector; obtaining a mapping relation table; selecting, using vector retrieval, the prompt vector most similar to the text vector from the mapping relation table; and determining the image prompt most matching the text information according to the most similar prompt vector and the mapping relation table.

[0011] In a third aspect, there is provided a large model-based image prompt generation system, the system including: a first generation module for generating a number of candidate prompts using a large model according to preset information; a second generation module for generating a corresponding set of images for each candidate prompt; a screening module for screening candidate prompts meeting preset requirements according to a set of images corresponding to each candidate prompt; and a obtaining module for obtaining an image prompt set according to the candidate prompts meeting the preset requirements.

[0012] In a fourth aspect, there is provided an image generation system based on the large model-based image prompt generation method in the above first aspect or any corresponding technical solution thereof, including: an obtaining module for obtaining text information input by a user and an image prompt set; a selecting module for selecting the image prompt most matching the text information from the image prompt set; and a generating module for generating an image using a text-to-image model according to the most matching image prompt.

[0013] In a fifth aspect, there is provided an intelligent device, the intelligent device including at least one processor; and a memory communicatively connected to the at least one processor; wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the large model-based image prompt generation method in the above first aspect or any corresponding technical solution thereof and the image generation method in the above second aspect or its corresponding technical solution are implemented.

[0014] In a sixth aspect, a computer-readable storage medium is provided, which stores multiple pieces of program code. The program code is adapted to be loaded and run by a processor to execute the method for generating image prompt words based on a large model in the first aspect or any corresponding technical solution thereof, and the method for generating images in the second aspect or the corresponding technical solution thereof.

[0015] One or more of the above technical solutions of the present application have at least one or more of the following beneficial effects:

[0016] In implementing the technical solution provided by the present application, a large model is used to generate a number of candidate prompt words according to preset information, and a corresponding set of images is generated for each candidate prompt word. Then, the candidate prompt words that meet the preset requirements are screened according to the set of images corresponding to each candidate prompt word, and finally an image prompt word set is obtained. Since the large model has strong learning ability and applicability, a rich set of candidate prompt words can be obtained by using the large model. By generating images for the candidate prompt words to screen the candidate prompt words, the problem of difficult to formulate clear rules to judge the text quality is solved. Using the above method to obtain the image prompt word set greatly reduces the operation difficulty of users, effectively improves the quality of the prompt words input into the text-to-image model, and thus solves the problem of text-image mismatch when the text-to-image model generates images according to the input text. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Referring to the accompanying drawings, the disclosure of the present application will become more understandable. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the protection scope of the present application. Among them:

[0018] Figure 1 is a schematic diagram of the main step flow of the method for generating image prompt words based on a large model according to an embodiment of the present application;

[0019] Figure 2 is a schematic diagram of the main step flow of the method for generating images according to an embodiment of the present application;

[0020] Figure 3 is a schematic diagram of the main structural block diagram of the system for generating image prompt words based on a large model according to an embodiment of the present application;

[0021] Figure 4 is a schematic diagram of the main structural block diagram of the system for generating images according to an embodiment of the present application;

[0022] Figure 5 is a schematic diagram of the overall step flow of the method for generating image prompt words based on a large model according to an embodiment of the present application;

[0023] Figure 6It is a schematic diagram of the overall step flow of an image generation method according to an embodiment of the present application;

[0024] Figure 7 It is a schematic diagram of the main structure of an intelligent device according to an embodiment of the present application.

[0025] Reference numerals:

[0026] 11: Memory; 12: Processor. Detailed implementation manners

[0027] The following describes some implementation manners of the present application with reference to the accompanying drawings. Those skilled in the art should understand that these implementation manners are only used to explain the technical principle of the present application and are not intended to limit the protection scope of the present application.

[0028] In the description of the present application, terms such as "first", "second", etc. are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and can also be the communication between two elements. It can be a wireless connection or a wired connection.

[0029] In addition, a "module" and a "processor" can include hardware, software, or a combination of both. A module can include a hardware circuit, various appropriate sensors, communication ports, a memory, and can also include a software part, such as program code, or a combination of software and hardware. A processor can be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor can be implemented in software, in hardware, or in a combination of both. A computer-readable storage medium includes any suitable medium that can store program code, such as a magnetic disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, and so on.

[0030] In addition, if the meaning of "and / or" appears in this application, it includes three parallel scenarios. Taking "A and / or B" as an example, it includes Scenario A, or Scenario B, or the scenario where both A and B are satisfied simultaneously. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the premise that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or unachievable, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application. The term "at least one of A or B" or "at least one of A and B" has a meaning similar to "A and / or B" and can include only A, only B, or both A and B. The singular terms "a" and "this" can also include the plural form.

[0031] This application attaches great importance to the security of users' personal information and has taken security protection measures that meet industry standards and are reasonable and feasible to protect users' information and prevent personal information from being accessed, publicly disclosed, used, modified, damaged, or lost without authorization.

[0032] Here, some terms related to this application will be explained first.

[0033] Large Language Model (LLM): Also known as a large language model / big model, it is an artificial intelligence model designed to understand and generate human language. They are trained on a large amount of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and so on.

[0034] Prompt: The text or instruction provided to the model as input to guide it to generate a specific output. It is the text passage provided by the user when interacting with the model and is used to describe the information, answer, text, etc. that the user wants to obtain from the model. The purpose of the Prompt is to guide the model to generate the required response in order to better control the generated output.

[0035] Embedding: The Embedding technology is a method of mapping high-dimensional data into a low-dimensional space, usually used to convert discrete and non-continuous data into continuous vector representations for easy processing by computers. The Embedding technology usually captures the semantic information of the data. In Natural Language Processing (NLP), similar words or phrases will be closer in the embedding space, while different words or phrases will be farther apart from each other. This helps the model understand the meaning and semantic relationships of the language.

[0036] Vector retrieval: Vector retrieval is to retrieve the K vectors similar to the query vector in a given vector dataset according to a certain metric method. The essence of vector retrieval is to calculate the distance between vectors, and the magnitude of the vector distance directly reflects the similarity between two vectors.

[0037] Vector distance metric: used to measure the distance between two vectors. The smaller the distance, the higher the similarity between the vectors; conversely, the lower the similarity.

[0038] Text-to-image model: A model that generates corresponding images based on the input text content. Typical text-to-image models include, for example, DALL·E, Diffusion, Imagen, etc.

[0039] In the application scenario of automatically generating marketing materials (such as posters), it is usually necessary to generate corresponding images in combination with the text description input by the user. However, there may be a large difference between the images generated by existing text-to-image models and the text input by the user, and problems such as distortion and mutilation may even occur. This phenomenon results in a very high failure rate in the generation of marketing materials, leading to a poor user experience. In addition, this interaction method has high requirements for the user input. The user needs to imagine reasonable picture elements to represent the marketing content, but in many cases, it is very difficult for the user to do so.

[0040] To address the above problems, the present application provides a method for generating image prompts based on a large model. Refer to the attached Figure 1 , Figure 1 is a schematic diagram of the main steps of a method for generating image prompts based on a large model according to an embodiment of the present application. As Figure 1 shown, the method mainly includes the following steps S2 to S8:

[0041] Step S2, using the large model to generate a number of candidate prompts according to the preset information.

[0042] In this embodiment, taking the application scenario of poster generation as an example, the scenario of generating other types of image prompts is similar. The preset information can be information related to the product to be promoted, such as product name, product features, product uses, etc. Give full play to the associative ability of the large language model to generate multiple candidate prompts (prompts) for a class of products.

[0043] In one implementation, the large language model includes, but is not limited to, models such as GPT, BERT, Llama, and their derivative models.

[0044] Step S4, generating a corresponding set of images for each candidate prompt.

[0045] In this embodiment, a corresponding set of images is generated for each candidate prompt obtained in step S2.

[0046] In one embodiment, assume the product is sunglasses, and the product feature is to protect the eyes. Use a large model to generate multiple candidate prompt words based on the above product features, such as: anti-ultraviolet, reducing glare, enhancing color contrast, sunshading, etc. Then input each of the above candidate prompt words into the text-to-image model to obtain a corresponding set of images.

[0047] Step S6: Screen the candidate prompt words that meet the preset requirements according to the set of images corresponding to each candidate prompt word.

[0048] In this embodiment, screen the candidate prompt words that meet the preset requirements (reach the preset standard) according to the set of images corresponding to the candidate prompt words.

[0049] In one embodiment, the preset requirements can be set according to the product requirements, such as the quality of the generated images, whether the images accurately express the product features, etc.

[0050] Step S8: Obtain an image prompt word set according to the candidate prompt words that meet the preset requirements.

[0051] In this embodiment, construct an image prompt word set according to all the candidate prompt words that meet the preset requirements.

[0052] Based on the method from step S2 to step S8 above, use a large model to generate several candidate prompt words according to the preset information, and generate a corresponding set of images for each candidate prompt word. Then screen the candidate prompt words that meet the preset requirements according to the set of images corresponding to each candidate prompt word, and finally obtain an image prompt word set. Since the large model has strong learning ability and applicability, rich candidate prompt words can be obtained by using the large model. By generating images for the candidate prompt words to screen the candidate prompt words, the problem of difficult to formulate clear rules to judge the text quality is solved. Using the above method to obtain the image prompt word set greatly reduces the operation difficulty of users, effectively improves the quality of the prompt words input into the text-to-image model, and thus solves the problem of mismatch between the text and the image generated by the text-to-image model according to the input text.

[0053] In an alternative embodiment, after the above step S8, the following step S10 can be further included:

[0054] Step S10: Perform vector encoding on each image prompt word in the image prompt word set to obtain the corresponding prompt word vector.

[0055] In this embodiment, use the Embedding model to encode each image prompt word in the image prompt word set into a prompt word vector.

[0056] In one embodiment, the Embedding model can map the original data (such as image prompts) from a high-dimensional space to a low-dimensional space, thereby reducing the complexity of the data and the demand for computing resources. In NLP, similar words or phrases are closer in the Embedding vector space, and different words or phrases are farther apart from each other.

[0057] In one embodiment, the Embedding model includes, but is not limited to, the VSM vector space model, Locality-Sensitive Hashing, LSA / LDA topic models, word2vec / doc2vec models, Bert / ELMo / GPT and their variants, AVG / DNN / RNN / CNN / AE models, and so on.

[0058] In an alternative embodiment, after the above step S10, the following step S12 may further be included:

[0059] Step S12: Construct a mapping relation table according to the correspondence between the image prompts and the prompt vectors.

[0060] In this embodiment, the relationship between the image prompts and the prompt vectors is stored to facilitate subsequent lookup of the image prompts corresponding to the prompt vectors.

[0061] The method for generating image prompts based on a large model provided in this embodiment greatly reduces the requirements for the image description text input by the user, effectively improves the quality of the prompts input into the text-to-image model, reduces the probability of distortion and mutilation of the generated images, and achieves the effect of improving the quality and aesthetics of the generated images.

[0062] On the other hand, the present application also provides an image generation method based on the method for generating image prompts based on a large model in the above first aspect or any corresponding technical solution thereof. Refer to the attached Figure 2 , Figure 2 is a schematic diagram of the main step flow of the method for generating image prompts based on a large model according to an embodiment of the present application. As Figure 2 shown, the method mainly includes the following steps S14 to step S18:

[0063] Step S14: Obtain the text information input by the user and the set of image prompts.

[0064] In this embodiment, when the user's request to generate a poster occurs, the text information input by the user and the set of image prompts obtained by the method for generating image prompts based on a large model in the above first aspect or any corresponding technical solution thereof are obtained.

[0065] Step S16: Select the image prompt that best matches the text information from the set of image prompts.

[0066] In this embodiment, methods such as vector retrieval and vector distance measurement are used to select the image prompt word that best matches the text information from the set of image prompt words.

[0067] Step S18, use the text-to-image model to generate an image according to the most matching image prompt word.

[0068] In this embodiment, according to the image prompt word that best matches the text information input by the user obtained in step S16, use the text-to-image model to generate the corresponding (poster) image.

[0069] In an alternative embodiment, the above step S18 may further include the following steps S182 to S188:

[0070] Step S182, convert the text information into a text vector.

[0071] In this embodiment, use the Embedding model to calculate the text vector corresponding to the text information (such as the text description of the poster image) input by the user.

[0072] Step S184, obtain the mapping relation table.

[0073] In this embodiment, obtain the mapping relation table established according to the above-mentioned method for generating image prompt words based on the large model.

[0074] Step S186, use vector retrieval to select the prompt word vector that is most similar to the text vector from the mapping relation table.

[0075] In this embodiment, determine the prompt word vector that is most similar to the text vector by calculating the vector distance between the text vector and each prompt word vector in the mapping relation table.

[0076] In one embodiment, the measurement methods of vector distance include but are not limited to Euclidean distance, standardized Euclidean distance, cosine distance, Hamming distance, Manhattan distance, Gerrard distance, Tanimoto distance, etc.

[0077] Step S188, determine the image prompt word that best matches the text information according to the most similar prompt word vector and the mapping relation table.

[0078] In this embodiment, according to the prompt word vector that is most similar to the text vector obtained in step S186, look up its corresponding image prompt word in the mapping relation table.

[0079] In this embodiment, the user only needs to generally describe the main content of the picture, and can recall and match the appropriate prompt (image prompt word) through the text, achieving the effect of improving the quality of the generated image while reducing the operation difficulty of the user.

[0080] It should be noted that although the above embodiments describe the various steps in a specific order, those skilled in the art can understand that in order to achieve the effects of the present application, it is not necessary for different steps to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders, and these adjusted solutions are equivalent technical solutions to the technical solutions described in the present application, and thus will also fall within the protection scope of the present application.

[0081] On the other hand, the present application also provides an image prompt word generation system based on a large model, as Figure 3 shown, including: a first generation module 2, configured to use the large model to generate a number of candidate prompt words according to preset information; a second generation module 4, configured to generate a corresponding set of images for each candidate prompt word; a screening module 6, configured to screen the candidate prompt words that meet the preset requirements according to a set of images corresponding to each candidate prompt word; and a obtaining module 8, configured to obtain an image prompt word set according to the candidate prompt words that meet the preset requirements.

[0082] The above image prompt word generation system based on a large model is used to execute Figure 1 the embodiments of the image prompt word generation method based on a large model shown. The technical principles, the technical problems solved, and the technical effects produced by both are similar. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process and relevant descriptions of the system can refer to the content described in the embodiments of the method, and will not be elaborated here.

[0083] On the other hand, the present application also provides an image generation system based on the image prompt word generation method in the first aspect above or any corresponding technical solution thereof, as Figure 4 shown, including: an obtaining module 14, configured to obtain the text information input by the user and the image prompt word set; a selecting module 16, configured to select the image prompt word that best matches the text information from the image prompt word set; and a generating module 18, configured to generate an image according to the best-matching image prompt word by using a text-to-image model.

[0084] The above image generation system is used to execute Figure 2 the embodiments of the image generation method shown. The technical principles, the technical problems solved, and the technical effects produced by both are similar. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process and relevant descriptions of the system can refer to the content described in the embodiments of the method, and will not be elaborated here.

[0085] In an example of an application scenario of the present application, a method for automatically generating a poster picture is provided, including Figure 5 the image prompt word generation method shown and Figure 6The image generation method shown in FIG. Figure 5 It is an offline phase and can be executed in advance. Figure 6 It is the online stage of automatic poster image generation, that is, the steps required to generate a poster.

[0086] Specifically, Figure 5 The offline phase shown mainly includes the following steps:

[0087] Step S101, batch-generate texts describing a number of image elements through a large language model as a candidate set for built-in prompts.

[0088] The goal of this step is to generate several prompts that describe the content of the poster image. This requires the use of a large language model (LLM). The input of the model is information related to the product to be promoted, and multiple prompts can be generated for the same type of product.

[0089] Step S102, generating a set of pictures for each prompt, and filtering the prompt set according to the generated pictures, and forming a set A through the filtered prompts.

[0090] The goal of this step is to filter the prompt candidate set in step S101 and select prompts that can generate high-quality poster images that meet the characteristics of the corresponding product. Specifically, each prompt in the candidate set can be input into the Wensheng graph model to generate a set of poster images, and then determine whether the quality of the set of images meets the standards and whether it accurately expresses the characteristics of the product. Finally, the selected prompts will form a set A, i.e., an image prompt word set.

[0091] Step S103: Use the Embedding model to generate a vector code for each prompt in set A to form vector set B.

[0092] The goal of this step is to vectorize the prompt in set A of step S102 to obtain vector set B. This facilitates the subsequent recall of relevance with the text description entered by the user. This process requires the use of an Embedding model, which encodes the input text into vectors. The vectors corresponding to texts with similar semantics have higher similarity. In addition, the mapping relationship between the vector and the prompt needs to be stored to facilitate the subsequent search for the prompt corresponding to the vector.

[0093] Furthermore, Figure 6 The online phase shown mainly includes the following steps:

[0094] Step S201, performing vector encoding on the user's input text through the Embedding model to generate a vector m.

[0095] The function of this step is to calculate the Embedding vector m corresponding to the text description of the poster image input by the user when the user's request to generate a poster occurs. The Embedding model used is the same as that described in step S103.

[0096] Step S202: Calculate the vector distance metric between vector m and the vector set B, select the vector d with the closest distance therefrom, and query the corresponding prompt description p of this vector.

[0097] The goal of this step is to recall the vector most similar to vector m in step S201 from the vector set B described in step S103, and find the most similar prompt p according to the mapping relationship between the prompt and the Embedding vector.

[0098] Step S203: Input p into the text-to-image model to generate the final image.

[0099] This step inputs the prompt p recalled in step S202 into the text-to-image model to generate the final poster image. The text-to-image model includes, but is not limited to, machine-related derivative models such as StableDiffusion, DALL·E, Diffusion, and Imagen.

[0100] This embodiment provides a method / system for recalling built-in text-to-image prompts in combination with text search technology, achieving the purpose of making the text-to-image effect controllable, and solving problems such as text-image mismatch, image distortion, and mutilation that exist in image generation using only text-to-image models. By combining text-to-image and text recall algorithms, the quality and relevance of the generated images can be effectively improved. This embodiment innovatively introduces text search technology into the poster generation process, solving problems such as text-image mismatch, image distortion, and mutilation that exist in image generation using only text-to-image models, and achieving the effect of improving the quality and aesthetics of the generated images.

[0101] Those skilled in the art can understand that all or part of the processes in the method of the above-mentioned embodiment of the present application can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium, etc., that can carry the computer program code.

[0102] Another aspect of the present application also provides a computer-readable storage medium.

[0103] In an embodiment of a computer-readable storage medium according to the present application, the computer-readable storage medium may be configured to store a program for executing the large model-based image prompt word generation method in the above-mentioned first aspect or any corresponding technical solution thereof and the image generation method in the above-mentioned second aspect or its corresponding technical solution. This program can be loaded and run by a processor to implement the above-mentioned large model-based image prompt word generation method and image generation method. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The computer-readable storage medium may be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiments of the present application is a non-transitory computer-readable storage medium.

[0104] Another aspect of the present application also provides an intelligent device.

[0105] In an embodiment of an intelligent device according to the present application, the intelligent device may include at least one processor; and a memory communicatively connected to the at least one processor; wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the method described in any of the above embodiments is implemented. Refer to the attached Figure 7 , Figure 7 It is exemplarily shown in the figure that the memory 11 and the processor 12 are communicatively connected through a bus.

[0106] In some embodiments of the present application, the intelligent device may further include at least one sensor for sensing information. The sensor is communicatively connected to any type of processor mentioned in the present application. Optionally, the intelligent device described in the present application may be, but is not limited to, a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, a vehicle-mounted device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc. The embodiments of the present application do not make any limitations thereto.

[0107] So far, the technical solution of the present application has been described in conjunction with an embodiment shown in the accompanying drawings. However, those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. Without departing from the principle of the present application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present application.

Claims

1. An image prompt generation method based on a large model, characterized in that, The method includes: Using a large model to generate a number of candidate prompt words according to preset information; Generating a corresponding set of images for each candidate prompt word; Screening the candidate prompt words that meet the preset requirements according to the set of images corresponding to each candidate prompt word; Obtaining an image prompt word set according to the candidate prompt words that meet the preset requirements.

2. The method according to claim 1, wherein The method further includes: Performing vector encoding on each image prompt word in the image prompt word set to obtain a corresponding prompt word vector.

3. The method according to claim 2, wherein The method further includes: Constructing a mapping relation table according to the corresponding relationship between the image prompt word and the prompt word vector.

4. The method according to claim 1, wherein The preset information includes product features, and the preset requirements include that the image quality meets the preset standard and the matching degree between the image and the product features meets the preset standard.

5. An image generation method based on the large model-based image prompt word generation method according to any one of claims 1 to 4, characterized in that, The method includes: Obtaining the text information input by the user and the image prompt word set; Selecting the image prompt word that best matches the text information from the image prompt word set; Using a text-to-image model to generate an image according to the most matching image prompt word.

6. The method according to claim 5, wherein The selecting the image prompt word that best matches the text information from the image prompt word set includes: Converting the text information into a text vector; Obtaining the mapping relation table; Using vector retrieval to select the prompt word vector most similar to the text vector from the mapping relation table; Determining the image prompt word that best matches the text information according to the most similar prompt word vector and the mapping relation table.

7. An image prompt generation system based on a large model, characterized in that, The system includes: A first generation module for using a large model to generate a number of candidate prompt words according to preset information; A second generation module for generating a corresponding set of images for each candidate prompt word; A screening module for screening the candidate prompt words that meet the preset requirements according to the set of images corresponding to each candidate prompt word; An obtaining module for obtaining an image prompt word set according to the candidate prompt words that meet the preset requirements.

8. An image generation system based on the large model-based image prompt generation method according to any one of claims 1 to 4, characterized in that, The system includes: An obtaining module for obtaining the text information input by the user and the image prompt word set; A selecting module for selecting the image prompt word that best matches the text information from the image prompt word set; A generation module for using a text-to-image model to generate an image according to the most matching image prompt word.

9. An intelligent device, characterized in that, Includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, it implements the large model-based image prompt word generation method according to any one of claims 1 to 4 and the image generation method according to claim 5 or 6.

10. A computer-readable storage medium storing multiple program codes, characterized in that, The program code is adapted to be loaded and run by a processor to execute the large model-based image prompt word generation method according to any one of claims 1 to 4 and the image generation method according to claim 5 or 6.

Citation Information

Cited By

  • Multi-type image intelligent generation method, system and equipment based on large model

    CN122368260A