Image generation method based on generative pre-training model GPT and electronic device
By introducing the generative pre-trained model GPT and literary graphics model in the image generation method, and combining text construction and comparison results, the problem of low image generation accuracy in the prior art is solved, and higher image generation accuracy is achieved.
Patent Information
- Application Number
- CN202311762778.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
The existing machine learning-based literary graphics model is prone to errors when processing user spoken text descriptions, resulting in low accuracy of generated images.
The image generation method based on the generative pre-trained model GPT is adopted. The initial input text is predicted through the GPT model, the target text is generated, and the target text is image-drawn using the text-generated graph model. Then, by performing text construction processing on the first image, verification text is obtained, and the first image is adjusted according to the comparison result to improve the accuracy of the image.
Through this method, the accuracy of image generation can be significantly improved, so that the generated image is more in line with the user's intentions and needs.
Smart Images

Figure CN120182402A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to an image generation method and an electronic device based on the generative pre-trained model GPT. Background Art
[0002] With the development of society, the professional drawing requirements in various industries are more creative and practical, and there is a high demand for the final drawings in the field, including details, colors, professionalism, innovation, etc. Therefore, it is an extremely high challenge for professional painters and it is difficult to complete the corresponding creation in a short time.
[0003] Currently, the text-to-image model based on machine learning is mostly used to achieve intelligent drawing. For example, users convert their ideas into text descriptions and then use the text-to-image model to generate the images they want according to the text descriptions.
[0004] However, the user's colloquial text descriptions are relatively casual, which may cause errors in the text-to-image model, resulting in low accuracy of the drawn images. Therefore, there is an urgent need for a technical solution that can improve the accuracy of drawing. Summary of the Invention
[0005] In view of this, this application provides an image generation method and an electronic device based on the generative pre-trained model GPT to improve the accuracy of drawing, as follows:
[0006] An image generation method based on the generative pre-trained model GPT, the method includes:
[0007] Obtain an initial input text;
[0008] Use the GPT model to perform text prediction processing according to the initial input text to obtain a target text;
[0009] Use the text-to-image model to perform image drawing processing according to the target text to generate a first image;
[0010] Use the text-to-image model to perform text construction processing on the first image to obtain a verification text;
[0011] Obtain a comparison result at least according to the verification text, and the comparison result at least characterizes whether the first image matches the initial input text;
[0012] Adjust the first image according to the comparison result to obtain a second image.
[0013] In the above method, preferably, the comparison result is the comparison result between the verification text and the target text;
[0014] Among them, according to the comparison result, adjusting the first image to obtain a second image includes:
[0015] According to the comparison result, performing model optimization processing on the text-to-image generation model;
[0016] Using the optimized text-to-image generation model to perform image drawing processing on the target text to obtain a second image.
[0017] In the above method, preferably, the comparison result is the comparison result between the verification text and the initial input text;
[0018] Among them, according to the comparison result, adjusting the first image to obtain a second image includes:
[0019] According to the comparison result, performing model optimization processing on the GPT model;
[0020] Using the optimized GPT model to perform text prediction processing on the initial input text to obtain a new text;
[0021] Using the text-to-image generation model to perform image drawing processing according to the new text to obtain a second image.
[0022] In the above method, preferably, the comparison result includes: the comparison result between the verification text and the initial input text, and the comparison result between the verification text and the target text;
[0023] Among them, according to the comparison result, adjusting the first image to obtain a second image includes:
[0024] According to the comparison result between the verification text and the initial input text, performing model optimization processing on the GPT model;
[0025] According to the comparison result between the verification text and the target text, performing model optimization processing on the text-to-image generation model;
[0026] Using the optimized GPT model to perform text prediction processing on the initial input text to obtain a new text;
[0027] Using the optimized text-to-image generation model to perform image drawing processing on the new text to obtain a second image.
[0028] In the above method, preferably, the comparison result is the comparison result between the verification text and the target text;
[0029] Among them, according to the comparison result, adjusting the first image to obtain a second image includes:
[0030] According to the comparison result, perform text adjustment on the target text to obtain a modified text;
[0031] Use the text-to-image model to perform image drawing processing on the modified text to obtain a second image.
[0032] In the above method, preferably, the comparison result is the comparison result between the verification text and the initial input text;
[0033] Among them, according to the comparison result, adjusting the first image to obtain a second image includes:
[0034] According to the comparison result, perform text adjustment on the initial input text to obtain a new input text;
[0035] Use the GPT model to perform text prediction processing according to the new input text to obtain a new text;
[0036] Use the text-to-image model to perform image drawing processing according to the new text to generate a second image.
[0037] In the above method, preferably, it further includes:
[0038] Obtain a first modified input text;
[0039] Use the GPT model to perform text prediction processing according to the first modified input text and the initial input text to obtain an integrated text;
[0040] Use the text-to-image model to perform image drawing processing on the integrated text to obtain a third image.
[0041] In the above method, preferably, it further includes:
[0042] Obtain a second modified input text;
[0043] Use the second modified input text to perform text adjustment processing on the target text to obtain a target modified text;
[0044] Use the text-to-image model to perform image drawing processing on the target modified text to obtain a fourth image.
[0045] A computer-readable storage medium, the computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of the above when running.
[0046] An electronic device, including a memory and a processor, the memory stores a computer program, and the processor is configured to execute the method described in any one of the above through the computer program.
[0047] As can be seen from the above technical solutions, in an image generation method and an electronic device based on the generative pre-trained model GPT disclosed in this application, after obtaining the initial input text, the GPT model is first used to perform text prediction processing on the initial input text. After obtaining the target text, the text-to-image model is used to perform image drawing processing on the target text to obtain the first image. Here, the text-to-image model is trained with sample texts and sample images. After that, in order to improve the image accuracy, the text-to-image model is used to perform text construction processing on the first image and obtain a comparison result based on the obtained verification text to represent whether the first image matches the initial input text. Finally, the first image is adjusted according to the comparison result to obtain a second image that more meets the user's drawing requirements. It can be seen that in this application, the generative ability of the GPT model is used to process the initial input text, so that the target text input to the text-to-image model can more accurately express the characteristics of the drawn image. After that, the first image is further adjusted based on the comparison result of whether the first image matches the initial input text, so as to obtain a more accurate second image, thereby improving the accuracy of drawing. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0049] In order to more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1a It is a schematic diagram of the hardware environment composed of the terminal device and the server applicable to this application;
[0051] Figure 1b It is a flowchart of an image generation method based on the generative pre-trained model GPT provided in Embodiment 1 of this application;
[0052] Figure 2 It is an example diagram of the text input interface in the embodiment of this application;
[0053] Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 and Figure 7 They are respectively flowcharts for adjusting the first image in an image generation method based on the generative pre-trained model GPT provided in Embodiment 1 of this application;
[0054] Figure 8 It is a partial flowchart of an image generation method based on the generative pre-trained model GPT provided in the first embodiment of the present application;
[0055] Figure 9 It is an example diagram of a text modification interface in the embodiment of the present application;
[0056] Figure 10 It is another partial flowchart of an image generation method based on the generative pre-trained model GPT provided in the second embodiment of the present application;
[0057] Figure 11 It is a schematic structural diagram of an image generation device based on the generative pre-trained model GPT provided in the second embodiment of the present application;
[0058] Figure 12 And Figure 13 They are respectively another schematic structural diagram of an image generation device based on the generative pre-trained model GPT provided in the second embodiment of the present application;
[0059] Figure 14 It is a schematic structural diagram of an electronic device provided in the third embodiment of the present application;
[0060] Figure 15 It is a schematic diagram of the overall architecture of local text-to-image micro-adjustment in the form of GPT Q&A in the embodiment of the present application;
[0061] Figure 16 It is an example diagram of realizing drawing and fine-tuning in the embodiment of the present application;
[0062] Figure 17 It is a flowchart of realizing drawing and fine-tuning in the embodiment of the present application;
[0063] Figure 18 It is an example diagram of a chat session between a user and a GPT robot in the embodiment of the present application;
[0064] Figure 19 It is an example diagram of an image generated by the GPT robot in combination with a text-to-image model and the fine-tuned image in the embodiment of the present application;
[0065] Figure 20 It is a schematic diagram of the overall architecture of local text-to-image micro-adjustment in the form of GPT Q&A provided by the present application;
[0066] Figure 21 It is another example diagram of realizing drawing and fine-tuning in the embodiment of the present application. Detailed implementation manners
[0067] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0068] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0069] This application proposes an image generation method and an electronic device based on a generative pre-trained model GPT, which can be applied to smart home, smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the image generation method of the generative pre-trained model GPT can be applied to the hardware environment composed of terminal devices and servers, such as Figure 1a As shown in . The server is connected to the terminal device through the network and can be used to provide services such as image drawing application services for the terminal or the client installed on the terminal. Users can log in to the client to use the image drawing service on the server. In addition, a database is set on the server or independently of the server to provide data storage services for the server. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server.
[0070] The above network may include but is not limited to at least one of the following: wired network, wireless network. The above wired network may include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network, and the above wireless network may include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device may not be limited to a PC, a mobile phone, a tablet computer, etc.
[0071] refer to Figure 1bAs shown in the figure, it is a flowchart of an implementation of an image generation method based on the generative pre-trained model GPT provided in the first embodiment of the present application. This method can be applied to electronic devices capable of data processing, such as computers or servers. The technical solution in this embodiment is mainly used to improve the accuracy of images generated based on text.
[0072] Specifically, the method in this embodiment may include the following steps:
[0073] Step 101: Obtain the initial input text.
[0074] Among them, the initial input text contains at least one initial input statement.
[0075] Specifically, in this embodiment, a text input interface may be provided for the user, as Figure 2 shown in the figure, so that the user can input corresponding description statements in the text input interface. These description statements characterize the content and characteristics of the image that the user wants to draw. For example, "draw a picture of a primary school student doing homework", "the primary school student should wear a school uniform", and so on.
[0076] Based on this, in this embodiment, the input operation of the user on the text input interface can be received, and then the input content in the input operation is parsed to obtain at least one initial input statement, which forms the initial input text.
[0077] Step 102: Use the GPT model to perform text prediction processing on the initial input text to obtain the target text.
[0078] Among them, the target text contains at least one target statement.
[0079] It should be noted that the GPT model is a Generative Pre-Trained Transformer model. In this embodiment, the text of the current drawing scene can be used to pre-train the GPT model in advance, so that the GPT model can output a target text that better conforms to the current drawing scene for the initial input text, so as to improve the accuracy of the target text.
[0080] Specifically, the GPT model is configured with a constraint condition prompt. The GPT model obtains the initial input text based on the constraint condition and performs text prediction processing on the initial input text to obtain the target text.
[0081] Among them, the constraint conditions can be generated based on the basic theories of drawing, such as multi-professional factors like hue, light, transmission power, image ratio, etc. Through the constraint conditions, the GPT model outputs an element request to prompt the user to input corresponding content in the text input interface. For example, if the GPT model needs to generate a picture, the GPT model determines whether there are elements missing for drawing, such as format or style, according to the constraint conditions, and then outputs an element request to prompt the user to supplement the corresponding element content in the text input interface. Based on this, the GPT model performs text prediction processing based on the element content input by the user, that is, the initial input text, and thus obtains the target text.
[0082] In addition, the GPT model is trained using training samples, and these training samples are related to the constraint conditions. For example, in the training samples, according to the content described in the constraint conditions, questions are asked about the missing key factors:
[0083] The formats of two samples are as follows:
[0084] 1. {
[0085] Content: "Draw an image containing a hummingbird";
[0086] Subject: Hummingbird;
[0087] Need to supplement: ["background color", "image ratio", "style", etc...];
[0088] }
[0089] 2. {
[0090] Content: "Draw an image containing a hummingbird in Nordic style";
[0091] Subject: Hummingbird;
[0092] Style: Nordic;
[0093] Need to supplement: ["background color", "image ratio", etc...];
[0094] }
[0095] Based on this, the GPT model can output an element request "Need to supplement: ["background color", "image ratio", "style", etc...]" for the user input content "Draw an image containing a hummingbird" according to the constraint conditions. After the user inputs the corresponding content, the GPT model can generate the target text based on this content. Among them, the training samples can be organized and written in advance according to the knowledge of drawing experts.
[0096] It should be noted that the accuracy of the target text refers to the degree of matching between the target text and the initial input text. The higher the degree of matching between the target text and the initial input text, that is, the more accurate the target text.
[0097] Step 103: Use the text-to-image model to perform image drawing processing according to the target text to generate the first image.
[0098] Among them, the text-to-image model is trained according to the first training sample. The input samples in the first training sample include the first sample text, the first sample text contains at least one sample statement, and the output samples in the first training sample contain the first sample images.
[0099] Specifically, the text-to-image model can be a model based on machine learning, such as a neural network model. After being trained with the first training sample, the text-to-image model can output the corresponding first image for the input target text.
[0100] It should be noted that the first training sample in this embodiment can be a training sample for the current drawing scenario, so that the text-to-image model can output a first image that more conforms to the current drawing scenario for the target text, so as to improve the accuracy of the first image.
[0101] Step 104: Use the text-to-image model to perform text construction processing on the first image to obtain the verification text.
[0102] Among them, the verification text contains at least one verification statement.
[0103] It should be noted that the text-to-image model is trained according to the second training sample. The input samples in the second training sample contain the second sample images, and the output samples in the second training sample contain the second sample text, and the second sample text contains at least one sample statement.
[0104] Specifically, the image recognition model can be a model based on machine learning, such as a neural network model. After being trained with the second training sample, the image recognition model can perform image interpretation on the input first image, such as interpreting the content and characteristics in the first image, so as to output the corresponding verification text.
[0105] Step 105: Obtain a comparison result based on at least the verification text. The comparison result at least characterizes whether the first image matches the initial input text.
[0106] Among them, the comparison result may include: the comparison result between the verification text and the initial input text; or, the comparison result may include: the comparison result between the verification text and the target text; or, the comparison result may include: the comparison result between the verification text and the initial input text, and the comparison result between the verification text and the target text. Thus, the comparison structure can characterize whether the first image can present image content matching the initial input text.
[0107] It can be seen that in this embodiment, the first image drawn by the text-to-image model is interpreted, and then the interpreted verification text is compared with the initial input text, that is, the text representing the image that the user wants to draw, so as to obtain a comparison result indicating whether the drawn first image matches the image that the user wants to draw.
[0108] Step 106: Adjust the first image according to the comparison result to obtain a second image.
[0109] Among them, the comparison result characterizes whether the first image matches the initial input text. Based on this, the content and characteristics in the first image are adjusted according to the comparison result, such as adding content in the image, modifying the attributes of the content in the image, etc., to obtain a second image, so that the second image can match the image that the user wants to draw to meet the user's drawing requirements.
[0110] As can be seen from the above solution, in an image generation method based on the generative pre-trained model GPT provided in the first embodiment of the present application, after obtaining the initial input text, the GPT model is first used to perform text prediction processing on the initial input text. Thus, after obtaining the target text, the text-to-image model is used to perform image drawing processing on the target text to obtain a first image. Here, the text-to-image model is trained with sample texts and sample images. After that, in order to improve the image accuracy, the text-to-image model is used to perform text construction processing on the first image and a comparison result is obtained according to the obtained verification text to characterize whether the first image matches the initial input text. Finally, the first image is adjusted according to the comparison result to obtain a second image that more meets the user's drawing requirements. It can be seen that in this embodiment, the generative ability of the GPT model is used to process the initial input text, so that the target text input to the text-to-image model can more accurately express the characteristics of the drawn image. Then, further adjustment is made to the first image based on the comparison result of whether the first image matches the initial input text, so as to obtain a more accurate second image, thereby improving the accuracy of drawing.
[0111] Based on the above implementation solution, in one implementation, the comparison result is the comparison result between the verification text and the target text;
[0112] Based on this, when adjusting the first image according to the comparison result in step 106, it can be achieved in the following ways, as Figure 3 shown below:
[0113] Step 301: Optimize the text-to-image generation model according to the comparison result.
[0114] Specifically, in this embodiment, the optimization method of the model parameters during the training process of the text-to-image generation model can be referred to, and the corresponding model parameters in the text-to-image generation model can be adjusted up or down according to the comparison result, so that the text-to-image generation model can learn the matching relationship between the first image and the initial input text, thereby enabling the text-to-image generation model to output a new image different from the first image for the target text, and the new image will be more matched with the initial input text.
[0115] Step 302: Use the optimized text-to-image generation model to perform image drawing processing on the target text to obtain a second image.
[0116] It can be seen that in this embodiment, according to the matching relationship between the first image drawn by the text-to-image generation model last time and the initial input text, the text-to-image generation model is optimized, and then the optimized text-to-image generation model is used to reprocess the target text, so that the text-to-image generation model outputs a new image, that is, the second image, which is more matched with the initial input text, so as to improve the accuracy of image drawing.
[0117] In another implementation, the comparison result is the comparison result between the verification text and the initial input text;
[0118] Based on this, when adjusting the first image according to the comparison result in step 106, it can be achieved in the following ways, as Figure 4 shown below:
[0119] Step 401: Optimize the GPT model according to the comparison result.
[0120] Specifically, in this embodiment, the optimization method of the model parameters during the training process of the GPT model can be referred to, and the corresponding model parameters in the GPT model can be adjusted up or down according to the comparison result, so that the GPT model can learn the matching relationship between the first image and the initial input text, thereby enabling the GPT model to output a new text different from the target text for the initial input text, the new text will be more matched with the initial input text, and it can also enable the text-to-image generation model to output a new image different from the first image, and the new image will be more matched with the initial input text.
[0121] Step 402: Use the optimized GPT model to perform text prediction processing on the initial input text to obtain a new text.
[0122] Step 403: Use the text-to-image model to perform image drawing processing on the new text to obtain a second image.
[0123] It can be seen that in this embodiment, the GPT model is optimized according to the matching relationship between the first image drawn by the text-to-image model last time and the initial input text, and then the optimized GPT model is used to reprocess the initial input text. In this way, the text-to-image model outputs a new image, i.e., the second image, that is more matched with the initial input text based on the new text output by the GPT model, so as to improve the accuracy of image drawing.
[0124] In another implementation, the comparison result includes: the comparison result between the verification text and the initial input text, and the comparison result between the verification text and the target text;
[0125] Based on this, in step 106, when adjusting the first image according to the comparison result, it can be achieved in the following way, as Figure 5 shown in:
[0126] Step 501: Perform model optimization processing on the GPT model according to the comparison result between the verification text and the initial input text.
[0127] Among them, the optimization processing method of the GPT model in step 501 can refer to the Figure 4 optimization processing method of the GPT model in.
[0128] Step 502: Perform model optimization processing on the text-to-image model according to the comparison result between the verification text and the target text.
[0129] Among them, the optimization processing method of the text-to-image model in step 502 can refer to the Figure 3 optimization processing method of the text-to-image model in.
[0130] Step 503: Use the optimized GPT model to perform text prediction processing on the initial input text to obtain new text.
[0131] Step 504: Use the optimized text-to-image model to perform image drawing processing on the new text to obtain a second image.
[0132] It can be seen that in this embodiment, both the GPT model and the text-to-image model are optimized according to the matching relationship between the first image drawn by the text-to-image model last time and the initial input text, and then the optimized GPT model is used to reprocess the initial input text. In this way, the optimized text-to-image model outputs a new image, i.e., the second image, that is more matched with the initial input text based on the new text output by the GPT model, so as to improve the accuracy of image drawing.
[0133] In another implementation, the comparison result includes the comparison result between the verification text and the target text.
[0134] Based on this, when adjusting the first image according to the comparison result in step 106, it can be achieved in the following manner, as Figure 6 shown below:
[0135] Step 601: Adjust the target text according to the comparison result to obtain a modified text.
[0136] Specifically, in this embodiment, according to the comparison result, the sentences in the target text related to the image content and characteristics can be adjusted. For example, change "red line" in the target text to "pink line", thereby obtaining the modified text, so that the modified text is more matched with the initial input text, and can enable the text-to-image model to output a new image different from the first image, and the new image will be more matched with the initial input text.
[0137] Step 602: Use the text-to-image model to perform image drawing processing on the modified text to obtain a second image.
[0138] It can be seen that in this embodiment, according to the matching relationship between the first image drawn by the text-to-image model last time and the initial input text, the target text is optimized, and then the text-to-image model is used to process the modified text obtained after optimization again, so that the text-to-image model outputs a new image, that is, the second image, which is more matched with the initial input text, to improve the accuracy of image drawing.
[0139] In another implementation, the comparison result includes the comparison result between the verification text and the initial input text.
[0140] Based on this, when adjusting the first image according to the comparison result in step 106, it can be achieved in the following manner, as Figure 7 shown below:
[0141] Step 701: Adjust the initial input text according to the comparison result to obtain a new input text.
[0142] Specifically, in this embodiment, according to the comparison result, the sentences in the initial input text related to the image content and characteristics can be adjusted. For example, change "line" in the initial input text to "pink line", thereby obtaining the new input text, so that the new input text better meets the user's drawing requirements. Thus, the GPT model can output a new text different from the target text for the new input text, and the new text will better meet the user's drawing requirements, and can enable the text-to-image model to output a new image different from the first image, and the new image will better meet the user's drawing requirements.
[0143] Step 702: Use the GPT model to perform text prediction processing on the new input text to obtain new text.
[0144] Step 703: Use the text-to-image model to perform image drawing processing on the new text to generate a second image.
[0145] It can be seen that in this embodiment, according to the matching relationship between the first image drawn by the text-to-image model last time and the initial input text, the initial input text is optimized, and then the GPT model is used to process the new input text, so that the text-to-image model outputs a new image that better meets the user's drawing requirements, namely the second image, to improve the accuracy of image drawing.
[0146] In one implementation, after step 103, the method in this embodiment may further include the following steps, as Figure 8 shown in
[0147] Step 107: Obtain the first modified input text.
[0148] Among them, the first modified input text contains at least one modified input statement.
[0149] Specifically, in this embodiment, a text modification interface can be provided for the user, as Figure 9 shown in, so that the user can input corresponding description statements in the text modification interface. These description statements characterize the content and characteristics of the image that the user wants to modify the first image, for example, "change the primary school student's homework to math homework", etc.
[0150] Based on this, in this embodiment, the input operation of the user on the text modification interface can be received, and then the input content in the input operation is parsed to obtain at least one modified input statement, which forms the first modified input text.
[0151] Step 108: Use the GPT model to perform text prediction processing on the first modified input text and the initial input text to obtain an integrated text.
[0152] Specifically, in this embodiment, the modified input text and the initial input text can be merged, and then the GPT model is used to process the merged text to obtain an integrated text different from the target text. For example, on the basis of the initial input text, according to the first modified input text, a modification operation is performed on the initial input text. The modification operation can include any one or more of adding description statements, deleting description statements, and modifying description statements. After that, the GPT model is used to process the modified text to obtain an integrated text. The obtained integrated text can characterize the text content of the initial input text after being adjusted by the modified input text.
[0153] Step 109: Use a text-to-image model to perform image drawing processing on the integrated text to obtain a third image.
[0154] It can be seen that in this embodiment, a text modification service can be provided for the user. The user inputs the corresponding modified input text according to the drawing intention. In this way, a new image that meets the user's needs, namely the third image, can be drawn through the GPT model and the text-to-image model, so as to improve the accuracy of image drawing.
[0155] In one implementation, after step 103, the method in this embodiment may further include the following steps, as Figure 10 shown in
[0156] Step 110: Obtain a second modified input text.
[0157] Wherein, the second modified input text includes at least one modified input statement.
[0158] Specifically, in this embodiment, a text modification interface can be provided for the user, as Figure 9 shown in, so that the user can input corresponding description statements in the text modification interface. These description statements characterize the content and characteristics of the image that the user wants to modify the first image. For example, "change the primary school student's homework to math homework", etc.
[0159] Based on this, in this embodiment, the input operation of the user on the text modification interface can be received, and then the input content in the input operation can be parsed to obtain at least one modified input statement, which constitutes the second modified input text.
[0160] Step 111: Use the second modified input text to perform text adjustment processing on the target text to obtain a target modified text.
[0161] Specifically, in this embodiment, the second modified input text and the initial input text can be merged. For example, on the basis of the target text, according to the second modified input text, a modification operation is performed on the target text. The modification operation can include any one or any combination of adding description statements, deleting description statements, and modifying description statements, thereby obtaining the target modified text. The obtained target modified text can characterize the text content of the target text after being adjusted by the second modified input text.
[0162] Step 112: Use the text-to-image model to process the target modified text to obtain a fourth image.
[0163] It can be seen that in this embodiment, a text modification service can be provided for users. The users input corresponding modified input text according to their drawing intentions. After modifying the target text with the modified input text, a new image that meets the user's needs, namely the fourth image, is drawn using a text-to-image model, so as to improve the accuracy of image drawing.
[0164] Reference Figure 11 , which is a schematic structural diagram of an image generation device based on the generative pre-trained model GPT provided in the second embodiment of the present application. This device can be configured in an electronic device capable of data processing, such as a computer or a server, etc. The technical solution in this embodiment is mainly used to improve the accuracy of images generated based on text.
[0165] Specifically, the device in this embodiment may include the following units:
[0166] An initial acquisition unit 1101, configured to acquire an initial input text;
[0167] A text processing unit 1102, configured to perform text prediction processing on the initial input text using a GPT model to obtain a target text;
[0168] An image drawing unit 1103, configured to perform image drawing processing on the target text using a text-to-image model to generate a first image;
[0169] A text construction unit 1104, configured to perform text construction processing on the first image using a text-to-image model to obtain a verification text;
[0170] A text comparison unit 1105, configured to obtain a comparison result at least according to the verification text, where the comparison result at least characterizes whether the first image matches the initial input text;
[0171] An image adjustment unit 1106, configured to adjust the first image according to the comparison result to obtain a second image.
[0172] As can be seen from the above solution, in an image generation device based on the generative pre-trained model GPT provided in the second embodiment of the present application, after obtaining the initial input text, the GPT model is first used to perform text prediction processing on the initial input text. After obtaining the target text, the text-to-image model is used to perform image drawing processing on the target text to obtain the first image. Here, the text-to-image model is trained with sample texts and sample images. After that, in order to improve the image accuracy, the text-to-image model is used to perform text construction processing on the first image and obtain a comparison result according to the obtained verification text to represent whether the first image matches the initial input text. Finally, the first image is adjusted according to the comparison result to obtain a second image that more meets the user's drawing requirements. It can be seen that in this embodiment, the generative ability of the GPT model is used to process the initial input text, so that the target text input to the text-to-image model can more accurately express the characteristics of the drawn image. After that, the first image is further adjusted based on the comparison result of whether the first image matches the initial input text, so as to obtain a more accurate second image, thereby improving the accuracy of drawing.
[0173] In one implementation, the comparison result is the comparison result between the verification text and the target text;
[0174] Among them, the image adjustment unit 1106 is specifically configured to: perform model optimization processing on the text-to-image model according to the comparison result; use the optimized text-to-image model to perform image drawing processing on the target text to obtain a second image.
[0175] In one implementation, the comparison result is the comparison result between the verification text and the initial input text;
[0176] Among them, the image adjustment unit 1106 is specifically configured to: perform model optimization processing on the GPT model according to the comparison result; use the optimized GPT model to perform text prediction processing on the initial input text to obtain a new text; use the text-to-image model to perform image drawing processing according to the new text to obtain a second image.
[0177] In one implementation, the comparison result includes: the comparison result between the verification text and the initial input text, and the comparison result between the verification text and the target text;
[0178] Among them, the image adjustment unit 1106 is specifically configured to: perform model optimization processing on the GPT model according to the comparison result between the verification text and the initial input text; perform model optimization processing on the text-to-image model according to the comparison result between the verification text and the target text; use the optimized GPT model to perform text prediction processing on the initial input text to obtain a new text; use the optimized text-to-image model to perform image drawing processing on the new text to obtain a second image.
[0179] In one implementation, the comparison result is the comparison result between the verification text and the target text;
[0180] Among them, the image adjustment unit 1106 is specifically configured to: perform text adjustment on the target text according to the comparison result to obtain a modified text; use the text-to-image model to perform image drawing processing on the modified text to obtain a second image.
[0181] In one implementation, the comparison result is the comparison result between the verification text and the initial input text;
[0182] Among them, the image adjustment unit 1106 is specifically configured to: perform text adjustment on the initial input text according to the comparison result to obtain a new input text; use the GPT model to perform text prediction processing according to the new input text to obtain a new text; use the text-to-image model to perform image drawing processing according to the new text to generate a second image.
[0183] In one implementation, the device in this embodiment may further include the following units, as Figure 12 shown in
[0184] The first modification unit 1107 is configured to obtain a first modified input text; use the GPT model to perform text prediction processing according to the first modified input text and the initial input text to obtain an integrated text; use the text-to-image model to perform image drawing processing on the integrated text to obtain a third image.
[0185] In one implementation, the device in this embodiment may further include the following units, as Figure 13 shown in
[0186] The second modification unit 1108 is configured to obtain a second modified input text; use the second modified input text to perform text adjustment processing on the target text to obtain a target modified text; use the text-to-image model to perform image drawing processing on the target modified text to obtain a fourth image.
[0187] It should be noted that the specific implementation of each unit in this embodiment can refer to the corresponding content in the previous text and will not be elaborated here.
[0188] In addition, the embodiments of the present application also claim to protect a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it executes the image generation method based on the generative pre-trained model GPT as described in any of the previous embodiments.
[0189] Reference Figure 14 , which is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present application. The electronic device can be an electronic device capable of data processing, such as a computer or a server, etc. The technical solution in this embodiment is mainly used to improve the accuracy of images generated based on text.
[0190] Specifically, the electronic device in this embodiment may include the following structure:
[0191] A memory 1401, configured to store a computer program and the data generated when the computer program runs;
[0192] A processor 1402, configured to implement, through the computer program: obtaining an initial input text; using a GPT model to perform text prediction processing on the initial input text to obtain a target text; using an image generation model based on text to perform image drawing processing on the target text to generate a first image; using the image generation model based on text to perform text construction processing on the first image to obtain a verification text; obtaining a comparison result at least based on the verification text, where the comparison result at least characterizes whether the first image matches the initial input text; and adjusting the first image according to the comparison result to obtain a second image.
[0193] As can be seen from the above technical solution, in an electronic device provided in the third embodiment of the present application, after obtaining the initial input text, the GPT model is first used to perform text prediction processing on the initial input text. After obtaining the target text, the text-to-image model is used to perform image drawing processing on the target text to obtain the first image. Here, the text-to-image model is trained with sample texts and sample images. After that, in order to improve the image accuracy, the text-to-image model is used to perform text construction processing on the first image and obtain a comparison result based on the obtained verification text to represent whether the first image matches the initial input text. Finally, the first image is adjusted according to the comparison result to obtain a second image that better meets the user's drawing requirements. It can be seen that in this embodiment, the generation ability of the GPT model is used to process the initial input text, so that the target text input to the text-to-image model can more accurately express the characteristics of the drawn image. Then, based on the comparison result of whether the first image matches the initial input text, the first image is further adjusted to obtain a more accurate second image, thereby improving the accuracy of drawing.
[0194] Taking the scenario of character drawing as an example, a specific implementation solution of the present application in actual application will be described in detail below:
[0195] As Figure 15 shown in
[0196] Specifically, the user inputs the original Q&A statement, the GPT service provides Q&A services for the user, the GPT service outputs the final text, and the text-to-image model provides text-to-image services to generate pictures that meet the user's needs.
[0197] Regarding the overall GPT text-to-image Q&A architecture, the GPT service includes the GPT model. The GPT model is trained and fine-tuned with data in the field. At the same time, through the prompt, the GPT service can output according to the rules and methods defined by the user. Just use the text-to-image model to generate images for the output text. If you are not satisfied with the output image, you can gradually generate the text that meets the requirements through the Q&A service.
[0198] Among them, in GPT, "prompt" refers to the initial message input by the user when initiating a conversation, which serves as the starting point for interacting with the model. The prompt input by the user can be a question, a sentence, a paragraph, or a complete conversation history. The model will generate a reply based on this prompt and then continuously interact with the user to generate a continuous conversation.
[0199] The GPT model is a chatbot based on the GPT series of models that can generate meaningful responses according to the prompts input by users. In the application of chatbots, the design and selection of prompts are crucial for the generation quality and interaction experience. Good prompts can guide the model to generate responses more accurately and specifically, making the interaction more natural and smooth. Some best practices include using concise language, avoiding overly vague or abstract descriptions, and providing sufficient context information to help the model understand the user's intentions and needs.
[0200] In addition, in this embodiment, historical professional data can be organized to fine-tune the text-to-image model according to the text-to-image model and professional knowledge in the field of drawing, so that it can understand the user's description of the image more quickly and accurately and can accurately give professional text-to-image text.
[0201] In summary, communicate with the Q&A service according to the image to be generated, describe the details, and finally the GPT model gives the text of the generated image, and the image is generated through the text-to-image model.
[0202] As Figure 16 shown, first, the user makes a preliminary input of the basic text of the required image "Query1: The external is a white refrigerator, the background is the sea, there are red apples inside, and the position is in the lower right corner"; then, the GPT service provides the Q&A service, and the Q&A service repairs according to the content through the service dialogue between the GPT services, asks relevant questions such as parameters, etc., and the user supplements certain points according to needs. After generating the intermediate text Answer2, it enters the text-to-image service for image generation.
[0203] In addition, the present application can also be fine-tuned according to the generated effect. If the user is not satisfied with a certain part, the modified text is input, and then the GPT service will organize and modify the text. After multiple conversations and modifications, until the generated image meets the requirements. For example, the GPT service obtains the Q&A data of image parameters, the Q&A data of object attributes, and various prompts, such as Q&A forms and rules, logic, etc., and then performs model fine-tuning and prompt adjustment based on this.
[0204] Specifically as Figure 17 shown:
[0205] First, the user tells the GPT service what elements and artistic conceptions are in the image to be generated according to their own needs;
[0206] Then, the GPT service conducts professional communication with the user according to the prompts pre-entered by the user or the fine-tuned GPT model, so as to increase or modify the corresponding parameters or other elements in the image;
[0207] After that, the text generated through multiple rounds of conversations is used in the tattoo pattern model to generate multiple images, and the user makes a selection according to their needs;
[0208] Finally, if there are parts in the image that the user is not satisfied with or need to be fine-tuned locally, the user can inform the GPT-based robot which parts need to be modified. Based on the robot's understanding, the robot conducts a question-and-answer session according to the prompt, modifies the text of the original text-to-image generation, and sends the modified text into the text-to-image model to generate a new image. This cycle continues until the user is satisfied.
[0209] For example:
[0210] As Figure 18 shown, the user has a chat session with the GPT robot, and the GPT robot outputs corresponding text according to the content of the multiple rounds of conversations input by the user;
[0211] As Figure 19 shown, the text generated by the GPT robot is sent to the text-to-image model, and the text-to-image model outputs an image based on the text; additionally, if the user is not satisfied with the output image, they can interact with the GPT robot again to complete local fine-tuning.
[0212] It can be seen that this application provides a text-to-image question-and-answer service architecture system containing GPT, as well as a service process. This makes drawing more intelligent and convenient, enhances the user's drawing experience, and greatly improves the drawing efficiency. Overall, it realizes a technical architecture from imagination to actual images, truly enabling users who do not understand drawing to also depict the images they need to meet the needs of actual life and production.
[0213] Specifically, in this application, the fine-tuning and training of GPT-based image-related knowledge Q&A pairs; the formulation and design of GPT-based prompt rules. For example, the content format to be returned by the GPT service and the expression form of important content to be returned, etc. Moreover, for the regeneration of the text for generating images, in response to the deficiencies of the returned images, the GPT service is deeply processed and generated through the question-and-answer service, and through the above processing, the text for generating images is continuously iteratively generated.
[0214] In summary, in the present application, by using image-related data within the field to fine-tune and train the GPT model, the GPT service can communicate with users more professionally, obtain feedback from users on relevant drawing key points, generate text based on the user's expression, and this text can quickly generate relevant images through the text-to-image service. Moreover, the purpose of quickly describing and quickly generating images is achieved, making the drawing not only fast, but also imaginative and innovative, eliminating the previous drawing process and achieving the effect of "what is described is what is obtained". And, to ensure obtaining ideal images, the image text regeneration technology is adopted to fine-tune the local part according to the generated image, continuously adjust and optimize the text of the generated image until the image meets the requirements. In addition, the present application also overcomes the troubles brought by cross-professions, enabling users who do not understand drawing to only automatically adjust the output text according to their descriptions and continuous questions and answers, and finally obtain ideal images.
[0215] Further, as Figure 20 shown in, the present application also provides an overall architecture for local fine-tuning of text-to-image in the form of GPT Q&A. The following steps are mainly described according to the architecture:
[0216] 1. The user and the GPT robot conduct descriptive Q&A according to the original image through the Q&A service;
[0217] 2. The GPT cluster calls the GPT service to process the Q&A;
[0218] 3. The GPT cluster generates image description text through the Q&A server;
[0219] 4. Enter the text-to-image service: Provide the image description text to the text-to-image model;
[0220] 5. Generate pictures;
[0221] 6. After the image is displayed, the user selects the image;
[0222] 7. After the image is displayed, the picture enters the text-to-image service again to prepare for generating image text;
[0223] 8. Use the image recognition model corresponding to the text-to-image model to recognize the image text;
[0224] 9. Automatically compare the image text with the image description text;
[0225] 10. Start a new conversation according to the comparison result, and use the GPT robot to regenerate the image description text.
[0226] It can be seen that Figure 20The overall architecture of the GPT text-to-image Q&A is described. The GPT service includes the GPT model, which is trained and fine-tuned with data in the field. At the same time, through prompts, the GPT service can output according to the rules and methods defined by the user. The text output can be used to generate images using the text-to-image model. Additionally, users can select images with relatively high recognition generated by text-to-image, and use the text-to-image model to generate text descriptions of the images. These descriptions are compared with the original text descriptions generated by GPT, and GPT can generate new text descriptions based on these comparisons. Users can communicate with the GPT service as needed to add, modify, or other content.
[0227] In summary, in this application, the generated images are communicated with the Q&A service to depict details. GPT finally gives the text of the generated image. Images are generated through the text-to-image model. Text descriptions are generated in reverse based on the images generated by text-to-image and compared automatically with the original text generated by GPT. The differences are informed to the user, and the user can add, modify, and improve the text content in a dialogue form until the image meets the requirements.
[0228] For example:
[0229] Such as Figure 21 As shown, first, the user makes a preliminary input of the basic text of the required image through the Q&A service, such as: "The external is a white refrigerator, the background is the sea, there are red apples inside, and the position is in the lower right corner." The GPT service will repair according to the content and ask relevant questions such as parameters. The user supplements some points as needed to generate intermediate text, such as "answer2: A minimalist still life...". Then, it enters the text-to-image model to generate images through the text-to-image service and provides them to the user for image selection. Additionally, fine-tuning can be performed according to the generated effect. If the user is not satisfied with a certain part, the modified text is input, and the text description of the image is generated through the corresponding image recognition model of the text-to-image model. This description is compared with the intermediate text. Based on this, the GPT service will organize and modify the text. After multiple dialogues and modifications, the generated image meets the requirements.
[0230] In summary, this application describes the text-to-image Q&A service architecture system including GPT and the service process. It makes drawing more intelligent and convenient, improves the user's drawing experience, and greatly improves the drawing efficiency. Overall, it realizes the technical architecture from imagination to actual images, and truly enables users who do not understand painting to depict the required images to meet the needs of actual life and production.
[0231] It can be seen that the technical solution provided by this application is based on the fine-tuning and training of GPT's image-related knowledge Q&A pairs, and on the formulation and design of GPT's prompt rules. For example, the content format returned by the GPT service and the expression form of important content to be returned. Moreover, for the regeneration of the text of the generated image, in view of the deficiencies of the returned image, the GPT service is deeply processed and generated through the Q&A service. After the above processing, the text of the generated image is continuously iterated. In addition, an image description text comparison mechanism and a similarity and difference point comparison algorithm are adopted in this application.
[0232] In summary, in this application, the GPT model is fine-tuned and trained using image-related data in the field, enabling the GPT service to communicate with users more professionally, obtaining feedback from users on relevant drawing key points, and generating text based on the user's expression. This text can quickly generate relevant images through the text-to-image service. The specific advantages are as follows:
[0233] First of all, this application achieves the goal of quickly describing and generating images, making the drawing not only fast, but also imaginative and innovative, eliminating the previous drawing process and achieving the effect of "what you say is what you get".
[0234] Moreover, to ensure obtaining an ideal image, this application adopts the image text regeneration technology to fine-tune the local area according to the generated image, continuously adjusting and optimizing the text of the generated image until the image meets the requirements.
[0235] In addition, this application overcomes the troubles brought by cross-professions, enabling users who do not understand drawing to only need to automatically adjust the output text according to their descriptions and continuous Q&A, and finally an ideal image can be obtained.
[0236] Finally, an automated text reverse comparison mechanism is adopted in this application to continuously iterate the pictures to make the pictures more in line with the user's requirements.
[0237] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. An image generation method based on the generative pre-trained model GPT, characterized in that, The method includes: Obtaining an initial input text; Using a GPT model to perform text prediction processing based on the initial input text to obtain a target text; Using an image generation model to perform image drawing processing based on the target text to generate a first image; Using the image generation model to perform text construction processing on the first image to obtain a verification text; Obtaining a comparison result at least based on the verification text, where the comparison result at least characterizes whether the first image matches the initial input text; Adjusting the first image according to the comparison result to obtain a second image.
2. The method according to claim 1, characterized in that, The comparison result is the comparison result between the verification text and the target text; Wherein, adjusting the first image according to the comparison result to obtain a second image includes: Performing model optimization processing on the image generation model according to the comparison result; Using the optimized image generation model to perform image drawing processing on the target text to obtain a second image.
3. The method according to claim 1, characterized in that, The comparison result is the comparison result between the verification text and the initial input text; Wherein, adjusting the first image according to the comparison result to obtain a second image includes: Performing model optimization processing on the GPT model according to the comparison result; Using the optimized GPT model to perform text prediction processing on the initial input text to obtain a new text; Using the image generation model to perform image drawing processing based on the new text to obtain a second image.
4. The method according to claim 1, characterized in that, The comparison result includes: the comparison result between the verification text and the initial input text, and the comparison result between the verification text and the target text; Wherein, adjusting the first image according to the comparison result to obtain a second image includes: Performing model optimization processing on the GPT model according to the comparison result between the verification text and the initial input text; Performing model optimization processing on the image generation model according to the comparison result between the verification text and the target text; Using the optimized GPT model to perform text prediction processing on the initial input text to obtain a new text; Using the optimized image generation model to perform image drawing processing on the new text to obtain a second image.
5. The method according to claim 1, characterized in that, The comparison result is the comparison result between the verification text and the target text; Wherein, adjusting the first image according to the comparison result to obtain a second image includes: Performing text adjustment on the target text according to the comparison result to obtain a modified text; Using the image generation model to perform image drawing processing on the modified text to obtain a second image.
6. The method according to claim 1, characterized in that, The comparison result is the comparison result between the verification text and the initial input text; Wherein, adjusting the first image according to the comparison result to obtain a second image includes: Performing text adjustment on the initial input text according to the comparison result to obtain a new input text; Using the GPT model to perform text prediction processing based on the new input text to obtain a new text; Performing image drawing processing according to the new text using the text-to-image model to generate a second image.
7. The method according to claims 1, 2, 3, 4, 5, and 6, characterized in that, Further comprising: Obtaining a first modified input text; Performing text prediction processing using the GPT model according to the first modified input text and the initial input text to obtain an integrated text; Performing image drawing processing on the integrated text using the text-to-image model to obtain a third image.
8. The method according to claims 1, 2, 3, 4, 5, and 6, characterized in that, Further comprising: Obtaining a second modified input text; Performing text adjustment processing on the target text using the second modified input text to obtain a target modified text; Performing image drawing processing on the target modified text using the text-to-image model to obtain a fourth image.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when running, executes the method according to any one of claims 1 to 8.
10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 8 through the computer program.