Business image automatic generation method and device based on large language model, storage medium and electronic device
Patent Information
- Application Number
- CN202410690980.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-12-02
AI Technical Summary
然而,文字生成图像模型的提示词大多需要人工填写,且提示词通常需要描述的很具体,甚至需要用到垂直领域(比如装修设计效果图,或者菜谱设计效果图)的专业术语,门槛较高
[0021]本申请还提供一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如上述任一种所述基于大语言模型的业务图像自动生成方法的步骤。
Smart Images

Figure CN121053232A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home technology, and in particular to a method, apparatus, storage medium and electronic device for automatic generation of business images based on a large language model. Background Technology
[0002] With the development of text-to-image models, text-to-image models are widely used in creative design, education, entertainment, medical and other fields.
[0003] In related technologies, users only need to input appropriate prompts, and the text-to-image model can generate a matching image based on the prompts. However, the prompts for text-to-image models mostly need to be filled in manually, and the prompts usually need to be very specific, even using professional terminology from vertical fields (such as interior design renderings or recipe design renderings), which has a high barrier to entry.
[0004] Therefore, there is an urgent need for a method to automatically generate business images, simplify the process for users to use text-to-image models, and reduce the difficulty of using text-to-image models. Summary of the Invention
[0005] The purpose of this application is to provide a method, apparatus, storage medium, and electronic device for automatically generating business images based on a large language model, which simplifies the process of users using text-to-image models and reduces the difficulty of using text-to-image models.
[0006] This application provides a method for automatically generating business images based on a large language model, including: The system acquires first business information based on the user's first input and a guidance word template corresponding to the target business, and concatenates and merges the first business information with the guidance word template to obtain the target guidance word. The guidance word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model. The system inputs the target guidance word into the natural language model to guide the natural language model to generate a first prompt word. The system inputs the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information.
[0007] Optionally, obtaining the first business information based on the user's first input and the guide word template corresponding to the target business includes: displaying the target page; the target page includes: multiple business information options related to the target business; each business information candidate includes multiple candidate options; in response to the user's first input on a candidate option of at least one of the multiple business information options, the first business information is obtained; wherein the first business information is generated based on the selection result of the candidate option for the at least one business information option.
[0008] Optionally, generating the corresponding target guidance word based on the first business information includes: obtaining the guidance word template corresponding to the target business, and concatenating and fusing the first business information with the guidance word template to obtain the target guidance word; wherein, the guidance word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model.
[0009] Optionally, the step of inputting the target guide word into the natural language model and guiding the natural language model to generate the first prompt word includes: calling the application programming interface provided by the natural language model and inputting the target guide word into the natural language model to obtain the first prompt word generated by the natural language model based on the target guide word.
[0010] Optionally, the step of inputting the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information includes: calling the application programming interface provided by the text-to-image model and inputting the first prompt word into the text-to-image model to obtain the first image generated by the text-to-image model based on the first prompt word.
[0011] Optionally, after inputting the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information, the method further includes: in response to a second input from the user for a candidate of at least one of the plurality of business information options, obtaining second business information; wherein the second business information is generated based on the result of a secondary selection of the candidate of the at least one business information option.
[0012] Optionally, after obtaining the second business information in response to a user's second input regarding at least one of the candidate business information options, the method further includes: comparing the difference between the second business information and the first business information, generating difference information, and generating an adjustment dialogue text based on the difference information; calling the application programming interface provided by the natural language model and inputting the adjustment dialogue text into the natural language model to obtain a second prompt word obtained by the natural language model after adjusting the first prompt word based on the adjustment dialogue text; and calling the application programming interface provided by the text-to-image model and inputting the second prompt word into the text-to-image model to obtain a second image generated by the text-to-image model based on the second prompt word.
[0013] This application also provides an automatic business image generation device based on a large language model, comprising: The module includes an acquisition module for acquiring first business information based on the user's first input and a guidance word template corresponding to the target business; a generation module for concatenating and fusing the first business information with the guidance word template to obtain the target guidance word; the guidance word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model; a calling module for inputting the target guidance word into the natural language model to guide the natural language model to generate a first prompt word; the calling module is also used to input the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information.
[0014] Optionally, the device further includes: a display module; the display module is used to display the target page; the target page includes: multiple business information options related to the target business; each business information candidate includes multiple candidate options; the acquisition module is specifically used to obtain the first business information in response to a first input from a user regarding a candidate option of at least one of the multiple business information options; wherein the first business information is generated based on the selection result of the candidate option for the at least one business information option.
[0015] Optionally, the calling module is specifically used to call the application programming interface provided by the natural language model, and input the target guide word into the natural language model to obtain the first prompt word generated by the natural language model based on the target guide word.
[0016] Optionally, the calling module is specifically used to call the application programming interface provided by the text-to-image model, and input the first prompt word into the text-to-image model to obtain the first image generated by the text-to-image model based on the first prompt word.
[0017] Optionally, the acquisition module is specifically configured to obtain second business information in response to a second input from the user regarding a candidate option of at least one of the plurality of business information options; wherein the second business information is generated based on the result of a secondary selection of the candidate option of the at least one business information option.
[0018] Optionally, the generation module is further configured to compare the difference between the second business information and the first business information, generate difference information, and generate adjusted dialogue text based on the difference information; the calling module is further configured to call the application programming interface provided by the natural language model, and input the adjusted dialogue text into the natural language model to obtain a second prompt word obtained by the natural language model after adjusting the first prompt word based on the adjusted dialogue text; the calling module is further configured to call the application programming interface provided by the text-to-image model, and input the second prompt word into the text-to-image model to obtain a second image generated by the text-to-image model based on the second prompt word.
[0019] This application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute, through the computer program, steps of implementing the business image automatic generation method based on any of the above-described large language models.
[0020] This application also provides a computer-readable storage medium comprising a stored program, wherein the program, when executed, implements the steps of any of the above-described methods for automatically generating business images based on a large language model.
[0021] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described methods for automatically generating business images based on a large language model.
[0022] The method, apparatus, storage medium, and electronic device for automatically generating business images based on a large language model provided in this application firstly acquire first business information obtained based on a user's first input and a guidance word template corresponding to the target business, and then concatenate and fuse the first business information and the guidance word template to obtain the target guidance word. The guidance word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model. Next, the target guidance word is input into the natural language model to guide the natural language model to generate a first prompt word. Finally, the first prompt word is input into the text-to-image model to obtain a first image related to the business content indicated by the first business information. This simplifies the user's process of using the text-to-image model and reduces the difficulty of using it. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0024] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a schematic diagram of the hardware environment for an automatic business image generation method based on a large language model according to an embodiment of this application; Figure 2 This is one of the flowcharts illustrating the automatic generation method for business images based on a large language model provided in this application; Figure 3 This is the second flowchart of the business image automatic generation method based on a large language model provided in this application; Figure 4 This is a schematic diagram of the structure of the business image automatic generation device based on a large language model provided in this application; Figure 5 This is a schematic diagram of the electronic device provided in this application. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] According to one aspect of the embodiments of this application, a method for automatically generating business images based on a large language model is provided. This method is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned method for automatically generating business images based on a large language model can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0029] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0030] In related technologies, text-to-image generation models are mainly divided into two categories: models based on Generative Adversarial Networks (GANs) and models based on diffusion models. A GAN model consists of a generator and a discriminator. The generator attempts to produce realistic images, while the discriminator attempts to distinguish between real and generated images. The advantage of GAN models is their ability to generate high-resolution and diverse images, but they also suffer from problems such as training instability and mode collapse. Diffusion models are models that generate data samples from random noise. They generate images by progressively removing noise and recovering the signal. The advantages of diffusion models are stable training, no need for adversarial loss, and ease of control, but they also suffer from slow generation speed and unclear image details.
[0031] To address the technical challenge of manually writing prompts for text-to-image models in related technologies, this application provides a method for automatically generating business images based on a large language model. This method can provide different pages for different businesses, allowing users to select various business information on the page. Subsequently, the system automatically generates guiding words and prompts based on the user's selections, and obtains the final image. This greatly simplifies the process of using text-to-image models and reduces the difficulty of using them.
[0032] The following description, in conjunction with the accompanying drawings, details the method for automatically generating business images based on a large language model provided in this application, through specific embodiments and application scenarios.
[0033] like Figure 2 As shown in the embodiment of this application, an automatic business image generation method based on a large language model is provided. This method may include the following steps 201 to 203: Step 201: Obtain the first business information based on the user's first input and the guide word template corresponding to the target business, and concatenate and merge the first business information and the guide word template to obtain the target guide word.
[0034] Wherein, the first input is: the input used by the user to select business content on the target page corresponding to the target business; the target business is any one of multiple businesses; one business corresponds to one page. The guiding word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model.
[0035] For example, in this embodiment of the application, the target page corresponding to the target business is used to provide users with multiple business information options, facilitating the user's selection of business information on the target page. Then, the system generates first business information based on the user's selection, and generates corresponding target guidance words based on the first business information. These target guidance words are used to guide the natural language model to generate corresponding prompt words.
[0036] It should be noted that large language models are models that utilize artificial intelligence technology to learn and generate human language from large amounts of text data. They can perform various tasks, such as text summarization, translation, and sentiment analysis. Large language models are characterized by their massive scale, containing billions or more parameters, and their ability to capture complex patterns and representations in language data. The natural language model in the embodiments of this application can be any of the following large language models: GPT-3, GPT-4, LaMDA, Imagen, Parti, and PaLM.
[0037] Specifically, step 201 above may also include the following steps 201a1 and 201a2: Step 201a1: Display the target page.
[0038] The target page includes multiple business information options related to the target business. Each business information candidate includes multiple candidate options.
[0039] For example, the target page mentioned above contains multiple business information options, and each business information option corresponds to multiple candidate options. For instance, taking interior design as the target business, the target page corresponding to this target business may include business information options such as decoration style and decoration color scheme. Furthermore, each business information option on the target page also contains multiple candidate options for the user to choose from.
[0040] Step 201a2: In response to the user's first input regarding a candidate option for at least one of the plurality of service information options, the first service information is obtained.
[0041] The first business information is generated based on the selection result of the candidate options for the at least one business information option.
[0042] For example, after a user makes a selection based on their actual needs on the target page, the system can generate corresponding first business information based on the user's selection. This first business information is used to generate guiding keywords.
[0043] For example, in this embodiment of the application, each service has a corresponding guide word template. After obtaining the first service information, the corresponding target guide word can be generated based on the guide word template corresponding to the target service.
[0044] For example, taking the generation of interior design renderings using a natural language model as an example, the following guiding text can be obtained based on the first business information mentioned above and the guiding text template: "You are an expert in using AI text to generate images and are proficient in writing AI program prompts. Here is a prompt for generating an image. You can analyze it and break it down into a reasonable prompt structure, which includes some replaceable parts and some fixed parts. Below, I will provide you with a prompt for interior design. Please break it down into a reasonable prompt template, including various parameters that may be involved in interior design, and then use this template to generate a new prompt. Please output this prompt in English."
[0045] Step 202: Input the target prompt word into the natural language model to guide the natural language model to generate the first prompt word.
[0046] For example, in this application embodiment, a natural language model can be used via remote access.
[0047] Specifically, step 202 above may also include step 202a: Step 202a: Call the application programming interface provided by the natural language model and input the target guide word into the natural language model to obtain the first prompt word generated by the natural language model based on the target guide word.
[0048] Step 203: Input the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information.
[0049] For example, after obtaining the first prompt word, the first prompt word can be input into the text-to-image model to obtain the image desired by the user.
[0050] It should be noted that the text generation image model used in the embodiments of this application can be any of the following: Imagen, Parti, and Stable Diffusion, etc.
[0051] Specifically, step 203 above may also include the following step 203a: Step 203a: Call the application programming interface provided by the text-to-image model and input the first prompt word into the text-to-image model to obtain the first image generated by the text-to-image model based on the first prompt word.
[0052] For example, after inputting the first prompt word into the text generation image model, an image generated by the text generation image model based on the first prompt word can be obtained.
[0053] Optionally, in this embodiment, if the user is not satisfied with the image generated by the text-to-image model, the prompt words can be adjusted using the multi-turn dialogue capability of the natural language model.
[0054] For example, after step 203 above, the automatic generation method for business images based on a large language model provided in this application embodiment may further include the following step 204: Step 204: In response to the user's second input regarding a candidate for at least one of the plurality of service information options, obtain the second service information.
[0055] The second business information is generated based on the result of a secondary selection of the candidate options for the at least one business information option.
[0056] For example, if a user is not satisfied with the generated image, the candidate options in each business information option on the target page can be adjusted to obtain the adjusted second business information.
[0057] For example, after step 204 above, the automatic generation method for business images based on a large language model provided in this application embodiment may further include the following steps 205 to 207: Step 205: Compare the differences between the second business information and the first business information, generate difference information, and generate adjusted dialogue text based on the difference information.
[0058] For example, the above-mentioned adjusted dialogue text is used to guide the natural language model to adjust the first prompt word generated earlier.
[0059] In one possible implementation, the prompt words can include fixed and adjustable parts. When the natural language model adjusts the prompt words generated during the previous dialogue, it can adjust only the adjustable parts of the prompt words. That is, the adjusted dialogue text generated based on the difference information is used to adjust the adjustable parts of the first prompt word.
[0060] Step 206: Call the application programming interface provided by the natural language model and input the adjusted dialogue text into the natural language model to obtain the second prompt word obtained by the natural language model after adjusting the first prompt word based on the adjusted dialogue text.
[0061] Step 207: Call the application programming interface provided by the text-to-image model and input the second prompt word into the text-to-image model to obtain the second image generated by the text-to-image model based on the second prompt word.
[0062] For example, when a user needs to adjust an image generated by a text-to-image model, the system will not generate new prompts, but will directly generate adjustment dialogue text. Based on this adjustment dialogue text, the system will use the multi-turn dialogue capability of the natural language model to adjust the first prompt and obtain the second prompt.
[0063] For example, after obtaining the second prompt word, the text-to-image model is input based on the second prompt word, so that the text-to-image model generates a new image.
[0064] It should be noted that users can repeat steps 204 to 207 above until the image generated by the text-to-image model meets the user's requirements.
[0065] For example, such as Figure 3 As shown, the system first retrieves business information from the target page based on user input. Then, it guides a natural language model to generate prompts, which are then input into a text-to-image model. Finally, the text-to-image model generates an image based on the prompts. If the user is not satisfied with the image, the above steps can be repeated until a satisfactory image is obtained. It should be noted that because the natural language model has multi-turn dialogue capabilities, apart from the initial input of prompts, subsequent processes only require inputting adjustment information text to control the natural language model to adjust the previously output prompts.
[0066] The method for automatically generating business images based on a large language model provided in this application first obtains first business information based on a user's first input and a guidance word template corresponding to the target business. The first business information and the guidance word template are then concatenated and fused to obtain the target guidance word. The guidance word template includes at least one of the following: identity-qualifying text describing the identity of the natural language model; image model description text describing the function of the text-to-image model; and image style description text describing the style of the image generated by the text-to-image model. Next, the target guidance word is input into the natural language model to guide it in generating a first prompt word. Finally, the first prompt word is input into the text-to-image model to obtain a first image related to the business content indicated by the first business information. This simplifies the user's process of using the text-to-image model and reduces its difficulty of use.
[0067] It should be noted that the execution entity of the service image automatic generation method based on a large language model provided in this application embodiment can be a service image automatic generation device based on a large language model, or a control module in the service image automatic generation device based on a large language model for executing the service image automatic generation method based on a large language model. This application embodiment uses the execution of the service image automatic generation method based on a large language model by the service image automatic generation device as an example to illustrate the service image automatic generation device based on a large language model provided in this application embodiment.
[0068] It should be noted that, in the embodiments of this application, the accompanying drawings of the various methods described above illustrate the automatic generation method of business images based on a large language model, all of which are exemplified by referring to one of the accompanying drawings in the embodiments of this application. In specific implementation, the automatic generation method of business images based on a large language model shown in the accompanying drawings of the various methods described above can also be implemented in conjunction with any other accompanying drawings that can be combined as illustrated in the above embodiments, which will not be elaborated here.
[0069] The following describes the automatic business image generation device based on a large language model provided in this application. The description below can be referred to in conjunction with the automatic business image generation method based on a large language model described above.
[0070] Figure 4 A schematic diagram of the structure of an automatic business image generation device based on a large language model provided in an embodiment of this application is shown below. Figure 4 As shown, it specifically includes: The acquisition module 401 is used to acquire first business information obtained based on the user's first input and a guidance word template corresponding to the target business; the generation module 402 is used to concatenate and fuse the first business information with the guidance word template to obtain the target guidance word; the guidance word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model; the invocation module 403 is used to input the target guidance word into the natural language model to guide the natural language model to generate a first prompt word; the invocation module 403 is also used to input the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information.
[0071] Optionally, the device further includes: a display module; the display module is used to display the target page; the target page includes: multiple business information options related to the target business; each business information candidate includes multiple candidate options; the acquisition module 401 is specifically used to obtain the first business information in response to a user's first input on a candidate option of at least one of the multiple business information options; wherein the first business information is generated based on the selection result of the candidate option of the at least one business information option.
[0072] Optionally, the calling module 403 is specifically used to call the application programming interface provided by the natural language model and input the target guide word into the natural language model to obtain the first prompt word generated by the natural language model based on the target guide word.
[0073] Optionally, the calling module 403 is specifically used to call the application programming interface provided by the text-to-image model and input the first prompt word into the text-to-image model to obtain the first image generated by the text-to-image model based on the first prompt word.
[0074] Optionally, the acquisition module 401 is specifically used to obtain second business information in response to a second input from the user regarding a candidate option of at least one of the plurality of business information options; wherein the second business information is generated based on the result of a secondary selection of the candidate option of the at least one business information option.
[0075] Optionally, the generation module 402 is further configured to compare the difference between the second business information and the first business information, generate difference information, and generate adjusted dialogue text based on the difference information; the calling module 403 is further configured to call the application programming interface provided by the natural language model, and input the adjusted dialogue text into the natural language model to obtain a second prompt word obtained by the natural language model after adjusting the first prompt word based on the adjusted dialogue text; the calling module 403 is further configured to call the application programming interface provided by the text-to-image model, and input the second prompt word into the text-to-image model to obtain a second image generated by the text-to-image model based on the second prompt word.
[0076] The business image automatic generation device based on a large language model provided in this application first obtains first business information based on a user's first input and a guidance word template corresponding to the target business, and then concatenates and fuses the first business information with the guidance word template to obtain the target guidance word. The guidance word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model. Then, the target guidance word is input into the natural language model to guide the natural language model to generate a first prompt word. Finally, the first prompt word is input into the text-to-image model to obtain a first image related to the business content indicated by the first business information. This simplifies the user's process of using the text-to-image model and reduces the difficulty of using it.
[0077] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a business image automatic generation method based on a large language model. The method includes: acquiring first business information obtained based on a user's first input and a guide word template corresponding to the target business, and concatenating and fusing the first business information with the guide word template to obtain the target guide word; the guide word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model; inputting the target guide word into the natural language model to guide the natural language model to generate a first prompt word; inputting the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information.
[0078] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0079] On the other hand, this application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute the business image automatic generation method based on a large language model provided by the above methods. The method includes: acquiring first business information obtained based on a user's first input and a guide word template corresponding to the target business, and concatenating and fusing the first business information with the guide word template to obtain the target guide word; the guide word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model; inputting the target guide word into the natural language model to guide the natural language model to generate a first prompt word; inputting the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information.
[0080] In another aspect, this application also provides a computer-readable storage medium, which includes a stored program, wherein the program executes the business image automatic generation method based on a large language model provided by the above methods when running. The method includes: acquiring first business information obtained based on a user's first input and a guide word template corresponding to the target business, and concatenating and fusing the first business information with the guide word template to obtain the target guide word; the guide word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model; inputting the target guide word into the natural language model to guide the natural language model to generate a first prompt word; inputting the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information.
[0081] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0082] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for automatically generating business images based on a large language model, characterized in that, include: Obtain first business information based on the user's first input and a guide word template corresponding to the target business, and then concatenate and fuse the first business information with the guide word template to obtain the target guide word; The guide word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model. The target prompt word is input into the natural language model, which then generates the first prompt word. The first prompt word is input into the text-to-image model to obtain a first image related to the business content indicated by the first business information.
2. The method for automatically generating business images based on a large language model according to claim 1, characterized in that, The step of obtaining the first business information based on the user's first input includes: The target page is displayed; the target page includes: multiple business information options related to the target business; each business information candidate includes multiple candidate options; In response to a user’s first input regarding a candidate for at least one of the plurality of service information options, the first service information is obtained; The first business information is generated based on the selection result of the candidate options for the at least one business information option.
3. The method for automatically generating business images based on a large language model according to claim 1, characterized in that, The step of inputting the target guidance word into the natural language model and guiding the natural language model to generate the first prompt word includes: The application programming interface provided by the natural language model is invoked, and the target prompt word is input into the natural language model to obtain the first prompt word generated by the natural language model based on the target prompt word.
4. The method for automatically generating business images based on a large language model according to claim 1, characterized in that, The step of inputting the first prompt word into a text-to-image model to obtain a first image related to the business content indicated by the first business information includes: The application programming interface provided by the text-to-image model is invoked, and the first prompt word is input into the text-to-image model to obtain the first image generated by the text-to-image model based on the first prompt word.
5. The method for automatically generating business images based on a large language model according to claim 2, characterized in that, After inputting the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information, the method further includes: In response to a second input from the user regarding a candidate for at least one of the plurality of service information options, second service information is obtained; The second business information is generated based on the result of a secondary selection of the candidate options for the at least one business information option.
6. The method for automatically generating business images based on a large language model according to claim 5, characterized in that, After obtaining the second business information in response to a second input from a user regarding a candidate for at least one of the plurality of business information options, the method further includes: Compare the differences between the second business information and the first business information, generate difference information, and generate adjusted dialogue text based on the difference information; The application programming interface provided by the natural language model is called, and the adjusted dialogue text is input into the natural language model to obtain a second prompt word obtained by the natural language model after adjusting the first prompt word based on the adjusted dialogue text; The application programming interface provided by the text-to-image model is invoked, and the second prompt word is input into the text-to-image model to obtain the second image generated by the text-to-image model based on the second prompt word.
7. A business image automatic generation device based on a large language model, characterized in that, The device includes: The acquisition module is used to acquire first business information based on the user's first input and the guide word template corresponding to the target business; The generation module is used to concatenate and fuse the first business information with the guide word template to obtain the target guide word; the guide word template includes at least one of the following: identity-qualifying text for describing the identity of the natural language model, image model description text for describing the function of the text-to-image model, and image style description text for describing the style of the image generated by the text-to-image model. The calling module is used to input the target prompt word into the natural language model and guide the natural language model to generate the first prompt word; The calling module is further configured to input the first prompt word into the text-to-image model to obtain a first image related to the business content indicated by the first business information.
8. The apparatus according to claim 7, characterized in that, The device further includes: a display module; The display module is used to display the target page; the target page includes: multiple business information options related to the target business; each business information candidate includes multiple candidate options; The acquisition module is specifically used to obtain the first business information in response to the user's first input on the candidate option for at least one of the plurality of business information options; The first business information is generated based on the selection result of the candidate options for the at least one business information option.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the business image automatic generation method based on a large language model as described in any one of claims 1 to 6.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the business image automatic generation method based on a large language model according to any one of claims 1 to 6 through the computer program.