Document generation method, apparatus, system, device, and storage medium

By constructing the information to be rewritten and using a large language model to generate target prompt words, the problems of high user input requirements and high costs in existing text-to-image technology are solved, and high-quality image generation and personalized interaction are achieved.

CN119810259BActive Publication Date: 2025-10-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411877476.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-10-10
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Existing text-based image technology requires high user input information and relies on traditional natural language processing technology, resulting in high maintenance and expansion costs, low flexibility, and difficulty in generating high-quality images.

Method used

By obtaining the target conversation text and historical conversation situations to construct the information to be rewritten, using a large language model to generate target prompt words, and combining the text-to-graph model to generate images, the dependence on the natural language processing module is reduced.

Benefits of technology

It improves the quality and flexibility of image generation, reduces maintenance and expansion costs, and enhances personalized interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810259B_ABST
    Figure CN119810259B_ABST
Patent Text Reader

Abstract

The present disclosure provides a text-to-image method, device, system, equipment and storage medium, and relates to the technical field of artificial intelligence, in particular to the technical field of computer vision, deep learning, large model and the like, and is applied to the scene of AIGC content generation based on artificial intelligence. The specific implementation scheme is as follows: obtaining target dialogue text for generating an image of a current round of dialogue of a target object; constructing to-be-rewritten information based on the target dialogue text and historical dialogue conditions; inputting the to-be-rewritten information into a large language model to obtain a target prompt word generated by the large language model according to the requirements of the target dialogue text; the target prompt word is obtained based on the text rewriting capability of the large language model; and generating an image of the current round of dialogue based on the target prompt word.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as computer vision, deep learning, and large models, and is applied to scenarios such as AIGC (Artificial Intelligence Generated Content) content generation based on artificial intelligence. Background Art

[0002] With the continuous development of artificial intelligence (AI), the technology of generating images from text has received unprecedented exposure and attention. A series of innovative tools and products have emerged in this field, not only significantly improving the ability to automatically generate high-quality images from text descriptions, but also opening up new possibilities for numerous application scenarios such as artistic creation, design conception, and virtual reality. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, system, device, and storage medium for creating a text map.

[0004] According to one aspect of the present disclosure, a method for generating a cultural image is provided, comprising:

[0005] Obtain the target conversation text for generating images in the current round of conversation of the target object;

[0006] Construct the information to be rewritten based on the target dialogue text and historical dialogue situations;

[0007] Input the information to be rewritten into the large language model to obtain target prompt words generated by the large language model according to the requirements of the target dialogue text; the target prompt words are obtained based on the text rewriting ability of the large language model;

[0008] Generate an image of the current round of dialogue based on the target prompt word.

[0009] According to another aspect of the present disclosure, there is provided a cultural image device, comprising:

[0010] An acquisition module, configured to acquire target conversation text of a current conversation round of a target object for generating an image;

[0011] A construction module is used to construct the information to be rewritten based on the target dialogue text and historical dialogue situations;

[0012] The first generation module is configured to input the information to be rewritten into the large language model to obtain target prompt words generated by the large language model according to the requirements of the target dialogue text; the target prompt words are obtained based on the text rewriting capability of the large language model;

[0013] The second generation module is used to generate an image of the current round of dialogue based on the target prompt word.

[0014] According to another aspect of the present disclosure, a cultural graph system is provided, comprising:

[0015] The prompt word optimization agent is used to obtain the target dialogue text for generating images in the current round of dialogue with the target object; construct the information to be rewritten based on the target dialogue text and historical dialogue situations; input the information to be rewritten into the large language model to obtain the target prompt words generated by the large language model according to the requirements of the target dialogue text; send the target prompt words to the text-to-image tool; the target prompt words are obtained based on the text rewriting capability of the large language model;

[0016] The text image tool is used to generate images based on target prompt words and the text image model in the text image tool.

[0017] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0018] at least one processor; and

[0019] a memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0021] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0022] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0023] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0025] Figure 1 is a flow chart of a method for producing a text graph according to an embodiment of the present disclosure;

[0026] Figure 2 This is a schematic diagram of constructing information to be rewritten in different rounds of dialogue according to an embodiment of the present disclosure;

[0027] Figure 3 1 is a flow chart of training a large language model according to an embodiment of the present disclosure;

[0028] Figure 4 1 is a schematic diagram of the overall framework of the Wensheng diagram method provided according to an embodiment of the present disclosure;

[0029] Figure 5 1 is a schematic diagram of the architecture of a cultural graph system provided according to an embodiment of the present disclosure;

[0030] Figure 6 is a processing flow chart of a Wensheng graph system provided according to an embodiment of the present disclosure;

[0031] Figure 7 is a schematic structural diagram of a Wensheng graph device provided according to an embodiment of the present disclosure;

[0032] Figure 8 4 is a block diagram of an electronic device for implementing the Wensheng diagram method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0033] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0034] The terms "first," "second," and the like in this disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. Furthermore, the terms "including," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or elements. A method, system, product, or apparatus is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.

[0035] With the emergence of large models and the birth of multimodal models, text-based image technology has gained significant exposure, and numerous text-based image products have emerged. The key feature of these text-based image products is that they generate one or more images from a single text input. Within the text-based image product category, conversational interactive image generation has also emerged. This involves chatting with the system, where users continuously organize their language to accurately express their needs. The text-based image product then understands the user's needs and ultimately produces the desired image.

[0036] Therefore, the related art text-based image products place extremely high demands on the user's input information, that is, the user is required to have good language expression ability.

[0037] Moreover, in related technologies, traditional natural language processing technology is used to extract, identify, and classify user input information in order to understand the intention of the user information. For example, the text input by the user is "draw a girl". In related technologies, traditional natural language processing technology is needed to extract the subject "girl" from it, and identify and classify it as "character" and "single subject". Then, based on the identified content, according to the fixed format of describing the character, the character's face, clothing and other modifications are added to obtain prompt words for use in the text-based image model reasoning. These supplementary methods use the dictionary method, which requires a large number of professionals to collect and organize, and has a fixed pattern and low flexibility.

[0038] Furthermore, traditional natural language processing solutions employed in related technologies require the adjustment of multiple functional modules, such as extraction, recognition, and classification, as needed. This hinders their application to other raw image scenarios. For example, when drawing people, animals, objects, or posters, the functional effects of each module must be tailored to the specific application scenario, requiring individual adjustments and optimizations for each scenario.

[0039] In summary, among the related technologies, the prompt words required for literary images still rely on traditional natural language processing technology, which requires high costs in maintenance, iteration, and expansion.

[0040] In view of this, the present disclosure provides a method for generating a text graph. Figure 1 As shown in FIG, it is a flow chart of the method, which mainly includes the following contents:

[0041] S101, obtaining a target dialogue text for generating an image in a current round of dialogue of a target object.

[0042] The target conversation text refers to the various descriptions provided by the target participant in the current conversation. These descriptions are used to generate images. The target participant can express their image requirements through text input or voice input. If the target participant expresses their image requirements through text input, the target conversation text is the text they entered. If the target participant expresses their image requirements through voice input, the voice is converted into text to serve as the target conversation text.

[0043] Regardless of the method used above, the target dialogue text can be treated as the target subject's current input query and processed accordingly. The target dialogue text may include descriptions of the desired image's theme, compositional elements, visual effects, color matching, and stylistic preferences.

[0044] S102, construct to-be-revised information based on the target dialogue text and the historical dialogue situation.

[0045] In the embodiments of the present disclosure, the historical dialogue situation can be divided into two cases: the current round dialogue is the first round dialogue, and the current round dialogue is the non-first round dialogue.

[0046] For the non-first round dialogue case, that is, the multi-round dialogue case, it refers to multiple rounds of interaction with the target object to generate the expected image. In the multi-round dialogue case, each round of dialogue as the current round of dialogue is revised by the large language model to obtain the corresponding prompt word.

[0047] Through the historical dialogue situation, the historical dialogue text of the target object and the historical prompt word generated by the large language model for the historical dialogue text are mainly sorted out. Among them, the historical dialogue text is the image generation requirement expressed by the target object before the current round of dialogue for the image to be generated. The historical dialogue situation mainly focuses on the historical prompt word, and the historical dialogue text is auxiliary information for understanding the context to help the large language model understand and sort out the historical prompt word to fully understand the requirements for generating the image.

[0048] Based on the target dialogue text of the current round and the historical dialogue situation, to-be-revised information is constructed, which integrates the specific requirements of the target object for the image to be generated and the preferences expressed before for the image to be generated, and is a more comprehensive understanding of the requirements of the target object for the image to be generated.

[0049] S103, input the to-be-revised information into the large language model to obtain a target prompt word generated by the large language model according to the requirements of the target dialogue text; the target prompt word is obtained based on the text revision capability of the large language model.

[0050] That is, the constructed to-be-revised information is input into the large language model, which generates a target prompt word according to the requirements of the dialogue text of the current round of dialogue by analyzing the to-be-revised information based on the text revision capability of the large language model. The target prompt word can more accurately express the requirements for generating the image after being revised by the large language model. For example, the requirements for generating the image are more specific, detailed, and more professional.

[0051] S104, generate an image of the current round of dialogue based on the target prompt word.

[0052] After receiving the target prompt word, a pre-trained text-to-image model is used to generate an image. The generated image should reflect the key content and core elements in the target dialogue text, ensuring that the visual output of the generated image is highly related to the theme expressed in the target dialogue text.

[0053] In the embodiments of the present disclosure, the latest requirements of the target object in the to-be-rewritten information are fused, that is, the requirements proposed by the target dialogue text and the historical requirements of the historical dialogue situation on the generated image. The to-be-rewritten information is input into the large language model, and the target prompt word can be generated by using the language understanding ability, language organization ability and knowledge of the large language model itself. By means of the text rewriting ability of the large language model, some colloquial, ambiguous and non-standard expressions in the original target dialogue text can be converted into prompt words that meet the requirements of the image generation technology and are clear and explicit in expression, so that high-quality prompt words can be generated to finally improve the quality of the generated image. In addition, relying on the language expression ability of the large language model, it is not necessary to perform extraction, identification and classification operations as in the traditional natural language processing scheme, and it is also not necessary to collect and arrange various template expressions by using a dictionary method. Therefore, the scheme provided in the embodiments of the present disclosure has low maintenance cost, is convenient for iterative optimization and has better scalability.

[0054] In the embodiments of the present disclosure, the to-be-rewritten information is constructed based on the target dialogue text and the historical dialogue situation, which can be implemented as follows: in the case that the current round of dialogue is the first round of dialogue, as shown in S201, the target dialogue text and the target character are spliced according to a preset format to obtain the to-be-rewritten information; wherein the value of the target character is empty, which is used to represent that the historical dialogue content is empty. Figure 2

[0055] That is, in the case of the first round of dialogue, the target dialogue text and the target character are spliced into structured to-be-rewritten information according to the preset format.

[0056] In implementation, the preset format can be set as the format of <target dialogue text, target character>. In the case of the first round of dialogue, the format meets the following requirements:

[0057] 1) The target field is used to express the target dialogue text, and the target dialogue text is the value of the target field. The dialogue start symbol and the dialogue end symbol are used to express the start and end of the target dialogue text, so as to facilitate the large language model to identify and verify the integrity of the target dialogue text.

[0058] 2) The target character is used to represent the historical dialogue record, and the target character is empty, which means that the current round of dialogue is the first round of dialogue and there is no historical dialogue record.

[0059] In the embodiments of the present disclosure, the target dialogue text and the target character are spliced according to the preset format, which can ensure that the obtained to-be-rewritten information has regularity and logicality, so as to facilitate the large language model to understand and sort out the requirements of the current round of dialogue on the generation of the image and the specific content of the historical situation, reduce the understanding error of the large model, improve the accuracy of the target prompt word, and thus improve the quality of the generated image.

[0060] ​In the case that the current round of dialogue is the latest round of dialogue in the multi-round dialogue, that is, the multi-round dialogue successor, the to-be-rewritten information is constructed based on the dialogue text and the historical dialogue situation, which can be realized through the following steps, as shown in the following: Figure 2

[0061] S202, obtaining historical dialogue text provided by a target object in a historical dialogue record, and historical prompt words generated by a large language model according to requirements of the historical dialogue text.

[0062] S203, splicing the target dialogue text and the historical dialogue record according to a preset format to obtain to-be-rewritten information; wherein in the to-be-rewritten information, the historical dialogue text and the historical prompt words of the historical dialogue text are assembled into information tuples according to an association relationship, and the information tuples are values of target characters.

[0063] Based on the content described in the foregoing, the preset format can be set as the format of <target dialogue text, target character>. In the case that the current round of dialogue is the latest round of dialogue in the multi-round dialogue, and the current round of dialogue is the second round of dialogue, the format meets the following requirements:

[0064] 1), the target field is used to express the target dialogue text of the current round of dialogue, and the target dialogue text is the value of the target field. The dialogue start symbol and the dialogue end symbol are used to express the start and end of the target dialogue text, so as to facilitate the large language model to identify and verify the integrity of the target dialogue text.

[0065] 2), the target character is used to represent the historical dialogue record, and the target character contains two key fields, namely a first subfield and a second subfield.

[0066] 3), the first subfield is used to represent the historical dialogue text of the first round of dialogue, and the historical dialogue text is the value of the first subfield. Moreover, the historical dialogue text uses the first start symbol and the first end symbol to represent the start and end of the historical dialogue text, which is used to identify and verify the integrity of the historical dialogue text. Similarly, the second subfield represents the historical prompt words of the first round of dialogue, and the historical prompt words are the value of the second subfield. Moreover, the historical prompt words use the second start symbol and the second end symbol to represent the start and end of the historical prompt words, which is used to identify and verify the integrity of the historical prompt words.

[0067] 4), the target start symbol and the target end symbol of the historical dialogue record are used to represent the start and end positions of the value of the target character, so as to facilitate the identification and verification of the integrity of the value of the target character.

[0068] ​In the disclosed embodiment, in a multi-round conversation scenario, the conversation text and historical conversation records are spliced ​​together according to a preset format. This allows the collection of image generation requirements expressed in the historical conversation records and then combines them with the latest round of conversation text to generate the information to be rewritten. Within the information to be rewritten, the information tuples are formatted to facilitate the large model's understanding of drawing requirements and preferences, thereby improving the quality of the generated target prompts.

[0069] Accordingly, during the creation process, the target user may refer to the generated image and change their image requirements. This may require multiple rounds of dialogue to iterate on the drawing requirements and ultimately achieve the desired image. If the historical dialogue record includes multiple rounds of historical dialogue, there will be multiple sets of information tuples in the historical dialogue record. The historical dialogue text and corresponding historical prompt words of each historical dialogue round are constructed into corresponding information tuples.

[0070] Taking a multi-round historical conversation as an example, the format of the prompt information to be rewritten meets the following requirements:

[0071] 1) Use the target field to represent the target conversation text for the current round of conversation. The target conversation text is the value of this target field. Use conversation start and end symbols to indicate the start and end of the target conversation text, making it easier for the text graph model to identify and verify the integrity of the target conversation text.

[0072] 2) Use target characters to represent historical conversation records. The target characters contain two key fields, namely the first subfield and the second subfield.

[0073] 3) In the case of multiple historical conversations, each historical conversation includes a corresponding pair of first and second subfields. The first and second subfields of the same historical conversation are represented by the third start and end symbols. This allows the text graph model to accurately understand the historical conversation text and historical prompt words of the same historical conversation.

[0074] 4) In each historical conversation, the corresponding first subfield is used to represent the historical conversation text of that conversation, and the historical conversation text is the value of the first subfield. Furthermore, the historical conversation text uses a first start symbol and a first end symbol to indicate the beginning and end of the historical conversation text, which is used to identify and verify the integrity of the historical conversation text. Similarly, the second subfield represents the historical prompt word of that historical conversation, and the historical prompt word is the value of the second subfield. Furthermore, the historical prompt word uses a second start symbol and a second end symbol to indicate the beginning and end of the historical prompt word, which is used to identify and verify the integrity of the historical prompt word.

[0075] 5) Use the target start symbol and target end symbol recorded in the historical conversation to indicate the start and end positions of the target character value, so as to facilitate identification and verification of the integrity of the target character value.

[0076] In the disclosed embodiment, the historical conversation text of each round of historical conversation and the corresponding historical prompt words are constructed as corresponding information tuples, which can provide a structured storage method for conversation information. By arranging these information tuples by round, the evolution process of the target object's raw image requirements can be clearly expressed. This enables the large language model to accurately capture the target object's raw image requirements and preferences, clarify the rewriting direction and purpose of the generated historical prompt words, and improve the quality of the rewritten target prompt words.

[0077] In the disclosed embodiment, the target object can use a piece of text to generate one or more images. Each image has its own characteristics and can also have its own corresponding core content. Therefore, the prompt words used to generate each image are different and independent. Based on this, when the first target round of dialogue in the historical dialogue record generates multiple images, the historical prompt words of the first target round of dialogue include sub-prompt words corresponding to the multiple images.

[0078] That is, the multiple sub-prompt words of the multiple images and the historical dialogue text of the first target round dialogue are constructed as the information tuple of the first target round dialogue.

[0079] Taking any round of historical dialogue as an example, assuming that multiple images are generated in any round of historical dialogue, in addition to meeting the above requirements, the expression of the historical prompt words in any round of historical dialogue must meet the following requirements:

[0080] 1) The historical prompt word in the second subfield includes multiple sub-prompt words, and the multiple sub-prompt words are separated by separators to facilitate identification of the sub-prompt words of each image.

[0081] 2) Ability to express the correspondence between each sub-prompt word and its corresponding image. During implementation, the sub-prompt words can be sorted in the same order as the generated multiple images, thereby establishing the correspondence between the sub-prompt words and images through their sorting positions. Furthermore, special descriptors can be added to the second subfield to express the correspondence between the sub-prompt words and images. For example, a first special symbol and a second special symbol can be used to express that these two special symbols contain a sub-prompt word and a corresponding image identifier. It is understood that the ability to identify the correspondence between the sub-prompt words and images is sufficient, and this is not limited in the present embodiment.

[0082] Thus, a set of information tuples is constructed from the multiple sub-prompt words of the multiple images and the historical dialogue text of the first target round of dialogue.

[0083] In the disclosed embodiment, by constructing multiple sub-prompt words of multiple images and the historical dialogue text of the first target round of dialogue into the information tuple of the first target round of dialogue, the large language model can understand the description of each image in each round of dialogue, so that the large language model can clarify the target object's image requirements through the structured information tuple, clarify the rewriting purpose and rewriting direction, and improve the quality of the generated target prompt words.

[0084] In some embodiments, the target user may upload a reference image and wish to generate an image similar in style or content based on the reference image. In this case, in the current round of conversation, if the target user has provided a reference image, a text description of the reference image can be generated. This text description can then be used as the content of the historical prompt words corresponding to the previous round of conversation, to construct the information to be rewritten.

[0085] That is, if the target subject provides a reference image, a corresponding text description can be generated for this reference image by combining image recognition technology and natural language generation technology. The generated text description can be determined based on the target conversation text of the current round of conversation. For example, if the target subject specifies that a similar subject should be generated based on a certain subject in the reference image, the subject can be identified using image recognition technology and a corresponding description can be generated. If the target subject specifies that the image should be generated based on the overall layout style of the reference image, the layout style of the reference image can be classified using image recognition technology to obtain the corresponding layout style category.

[0086] The text description of the reference image can serve as an important basis for generating the image. In the embodiment of the present disclosure, the text description of the reference image is used as the content of the historical prompt words corresponding to the previous round of dialogue of the current round, so that the large language model can sort out the target object's image generation requirements in the form of multiple rounds of dialogue. The reference image can be processed in a compatible dialogue format, so that the large model can understand and sort out the target dialogue text using one mode.

[0087] In the disclosed embodiment, the reference image uploaded by the target object is converted into a textual description of the reference image, making it easier for the text graph model to understand the content of the reference image. The textual description of the reference image is used as the content of the historical prompt words corresponding to the previous round of dialogue before the current round to construct the information to be rewritten, so that the information contained in the reference image can be integrated with the previous dialogue information. Using an information tuple representation allows the large language model to sort out and understand the rewriting task, thereby ensuring that the generated target prompt words meet expectations and improving the quality of the generated image.

[0088] In a multi-round conversation, as the number of conversation turns increases, the time and resource consumption costs for processing and analyzing historical conversation data will increase accordingly. Therefore, in the embodiment of the present disclosure, when the number of conversation turns in the historical conversation record is greater than m, the historical conversation text of m historical conversation turns is selected from the historical conversation record; where m is a positive integer greater than 1.

[0089] During implementation, the most recent m rounds of historical conversations can be retained. This ensures that the generated image is more focused on the target object's current drawing needs, thus better meeting the target object's expectations. Furthermore, selecting m rounds of historical conversations from a large number of historical conversations effectively reduces the memory resources required to store and process these large amounts of historical conversations, thereby improving image generation efficiency.

[0090] Based on the above, the historical conversation refers to the series of exchanges with the target object related to image generation before the current round of conversation. In other words, the historical conversation includes all the historical conversation text and the corresponding target prompt words before the current round of conversation.

[0091] Among them, when the current round of dialogue is the first round of dialogue, the target prompt word is the prompt word obtained by the large language model rewriting the dialogue text of the first round of dialogue.

[0092] For example, the target audience's first round of dialogue text might be, "Please draw a 2D anime-style picture: cyberpunk style, urban setting, a man with a robotic hand holding a combat knife, and a chip implanted in his spine." Using the large language model to rewrite the dialogue text, the resulting prompt might be, "The picture depicts a 2D anime-style scene. A man is in a cyberpunk-style city. One of his hands has been transformed into a robotic hand, tightly gripping a combat knife, looking majestic and powerful. A chip is clearly visible implanted in his spine, exuding a mysterious fusion of high technology and humanity." This demonstrates that after rewriting using the large language model, the prompt requirements are more specific and the style is more prominent.

[0093] In the embodiment of the present disclosure, when the current round of dialogue is the first round of dialogue, the target prompt word is the prompt word obtained by the large language model rewriting the dialogue text. Based on the rewriting ability of the large language model, the dialogue text of the target object can be converted into standardized and professional instructions that are more in line with the image generation technical requirements of the text-based graph model, which helps the text-based graph model to better understand the target object's image requirements.

[0094] When the current round of dialogue is the latest round of dialogue in multiple rounds of dialogue, the target prompt word is the prompt word obtained by rewriting the designated prompt word by the large language model; wherein the designated prompt word is the historical prompt word of the second target round of dialogue indicated by the dialogue text.

[0095] Taking the second round of dialogue as an example, assuming the current round is the latest in a multi-round conversation, if the target participant in the current round enters the text "The knife should be longer, like the one in the comic," the designated prompt is the prompt obtained by rewriting the first round of target participant input using the large language model, i.e., "The screen shows a scene in the style of anime. A man is in a cyberpunk-themed city. One of his hands has been transformed into a robotic hand, tightly gripping a combat knife, looking majestic and powerful. A chip is clearly visible implanted in his spine, exuding a mysterious fusion of high technology and humanity."

[0096] The prompt word rewritten by the large language model can be: "The screen shows a scene in the style of anime. In a cyberpunk urban setting, a man stands. One of his hands has been transformed into a majestic mechanical hand, holding a longer than usual combat knife, the blade shining with a cold light. A chip is implanted in his spine. The combination of high technology and humanity presents a mysterious sense of the future."

[0097] That is, when generating an image, if the target text of the current conversation doesn't explicitly specify which prompt word to rewrite, the default prompt word is the previous conversation's historical prompt word. Similarly, when generating an image, if the target text of the current conversation explicitly specifies a historical prompt word, that historical prompt word is the designated prompt word.

[0098] When generating multiple images, if the target dialogue text of the current round does not explicitly specify which prompt word to rewrite and does not specify which image's prompt word to rewrite, the default prompt word is the sub-prompt word of a randomly selected image from the previous round. If the target dialogue text of the current round explicitly specifies the prompt word for a specific image in that round, the sub-prompt word of the image in that specified round is used as the designated prompt word.

[0099] In the disclosed embodiment, in multiple rounds of conversations, the prompt words obtained by rewriting the specified prompt words through the large language model can be integrated with new raw image requirements based on historical prompt words. The rewritten prompt words can provide more comprehensive and accurate instructions for the raw image model.

[0100] It should be noted that when multiple images are generated in a historical round of conversation, the user can specify the historical prompt word corresponding to at least one of the multiple images in the target conversation text of the current round of conversation to rewrite the prompt word. If no specification is made, the large language model will default to randomly selecting the historical prompt word corresponding to one of the multiple images and rewriting it to obtain the prompt word. Alternatively, the large language model can intelligently identify and select the historical prompt word corresponding to the image with the highest visual quality to rewrite it to obtain the prompt word.

[0101] In some embodiments, for specific objects that the text graph model has never learned or that the text graph model cannot accurately process, the text graph model may encounter generation difficulties, resulting in a large gap between the generated image and the expectations of the target object. To this end, during implementation, professionals can provide professional preset descriptions of this type of specific object. The preset description may include a description of the appearance, material, characteristics, etc. of the specific object. The preset descriptions of these specific objects can be stored in a preset set. When the target dialogue text includes a specific object in the preset set, the preset description of the specific object is obtained; then, the information to be rewritten is constructed based on the preset description of the specific object, the dialogue text, and the historical dialogue situation.

[0102] During implementation, the preset description may be a manually defined detailed description of the specific object, which covers information such as the core features, typical forms, unique attributes, common associated elements, and general presentation methods of the specific object in different situations.

[0103] In the disclosed embodiment, the information to be rewritten is constructed based on a preset description of a specific object, the target conversation text, and historical conversation situations. The preset description of the specific object can assist the text-based graph model to better understand and generate an image that is more in line with the expectations of the target object, thereby reducing image deviations caused by the model's unfamiliarity with the specific object.

[0104] With the continuous development of artificial intelligence, in some scenarios, although related technologies can provide responses at the same time as the generated images, these responses are difficult to personalize and cannot effectively improve the user experience and satisfaction of the target audience.

[0105] In light of this, in the disclosed embodiments, a large language model can be fine-tuned to enable it to rewrite generated prompt words. Furthermore, this large language model can generate personalized responses based on this fine-tuning. Specifically, this can be implemented by obtaining the response text generated by the large language model for the conversation text based on the information to be rewritten; and feeding the response text and the generated image back to the target audience.

[0106] Specifically, the large language model generates a response tailored to the target conversation text. The generated image is then fed back to the target audience, providing them with a more intuitive and rich feedback experience. For example, if the target audience requests an image, such as "Generate an ink painting," the large language model will generate corresponding responses based on this request, such as "I've created an ink painting in the style of your request. I hope you like it." or "Based on your inspiration and requirements, I've carefully created a work of art imbued with the charm of ink painting. I hope it will win your favor."

[0107] In the disclosed embodiment, a large language model is used to generate targeted reply text for the target conversation text based on the information to be rewritten, and a more personalized reply can be fed back according to the target object's current round of raw image needs, thereby increasing the degree of personalization of the interaction between the target object and the raw image model and improving the target object's user experience.

[0108] In the embodiment of the present disclosure, in order to enable the large language model for generating images to generate high-quality images based on the rewritten prompt words, the large language model can be obtained by training a base model with a certain basic framework, and the training method can be as follows: Figure 3 Shown, including:

[0109] S301, supervised training is performed on the base model based on the first training corpus set to obtain an intermediate model, wherein the sample prompt words of the training first training corpus set are aligned to generate images according to the format requirements of the training corpus of the text-graph model.

[0110] The first training set is a dataset containing a large number of labeled input-output pairs, used to guide the model in learning a specific task. In the rewriting task, the first training set can be composed of query samples and their corresponding excellent prompt word samples.

[0111] The base model is a pre-trained large language model that has not been optimized for the specific task of rewriting prompt words, but has a certain basic framework.

[0112] The base model is trained in a supervised manner based on the first training corpus to obtain an intermediate model. Specifically, the base model is adjusted and optimized by learning the mapping relationship between the input and expected output in the first training corpus. During implementation, the model can be trained using SFT (Supervised Fine-Tunning), with the final optimized base model serving as the intermediate model.

[0113] In order to align the training corpus of the text graph model used for the text graph, the sample prompt words of the first training corpus set are trained to align with the format requirements of the training corpus of the text graph model.

[0114] For example, the prompt word samples can be organized in the format and order of the training corpus of the pre-text graph model. In this way, the optimized intermediate model can align the expression and understanding of the pre-text graph model, ensuring that the generated target prompt word adapts to the pre-text graph model to retain the original functional performance of the pre-text graph model.

[0115] S302, using a direct preference optimization training method, the intermediate model is optimized based on the second training corpus set to obtain a large language model.

[0116] The second training corpus set is different from the first training corpus set in that it may focus more on refining the model's capabilities in certain specific aspects, for example, it can be different output examples corresponding to multiple groups of similar input, and with preference evaluation information of these output examples.

[0117] Using the preference relationship information embodied in the second training corpus set, the parameters of the intermediate model are adjusted through the DPO (Direct Preference Optimization) training method. The performance and expression of the intermediate model are further improved, and finally a large language model optimized is obtained.

[0118] In the embodiments of the present disclosure, the base model is supervised trained based on the first training corpus set, which can make the obtained intermediate model more suitable for the needs of specific scenarios. By aligning the pre-text graph model, not only the text processing capability of the intermediate model can be improved, but also the original image generation processing capability of the pre-text graph model can be retained. Further, the intermediate model is directly preference optimized based on the second training corpus set, which can further optimize the model, so that the large language model obtained can avoid the short board of the model, and the generated content can better meet the needs of the target object.

[0119] It should be noted that in the case where the obtained large language model supports text rewriting and obtaining target prompt words, and also supports generating personalized replies, in the SFT stage, the base model can be trained through rewriting tasks and reply tasks. After obtaining the respective losses of the two tasks, the total loss is determined based on the respective losses, and then the parameters of the base model are optimized based on the total loss to obtain the intermediate model.

[0120] In the DPO stage, only the rewriting task of the intermediate model is trained, and the reply task can not be trained.

[0121] In some embodiments, after aligning the base model and optimizing the prompt word, the content in the target prompt word includes at least one of the following and satisfies the following described order:

[0122] Description of the picture style, description of the picture subject and limitations, description of the picture details, description of the picture background modification, special effects, composition, color tone, clarity description, and quality description.

[0123] If the target prompt includes multiple types of content, the relative order of the multiple types of content meets the above-mentioned description order requirements. For example, if the target prompt includes: composition, picture style description, picture subject and limitation description, picture background modification description, special effects, then the description order of the various types of information follows the above-mentioned description order requirements, and the final description order is: picture style description, picture subject and limitation description, picture background modification description, special effects, composition.

[0124] The disclosed embodiments can adapt to the expression specifications of the text-based graph model by standardizing the content of target prompt words and the order of different contents, so as to generate high-quality target prompt words and improve the quality of the generated image.

[0125] In summary, the structure of the text graph method provided by the embodiment of the present disclosure is as follows: Figure 4 As shown, in S401, a sample of information to be rewritten can be constructed based on the format of <target dialogue text, target characters>, and sample prompt words and sample replies can be constructed based on the format of <prompt word example, reply text example>, thereby obtaining a training corpus. After constructing the training corpus, in S401, the base model is fine-tuned using SFT and DPO training methods to obtain a large language model. The large language model can provide an interface to the outside world for providing inference and prediction services. The main function of this service is to realize text rewriting, obtain high-quality target prompt words and personalized replies, and provide them to the downstream text-based graph model for use. Figure 4 In S403, the conversation text entered by the target user is obtained. In S404, the conversation text and historical conversation records are processed according to the format of <target conversation text, target characters> to obtain the information to be rewritten. In S405, the information to be rewritten is input into the large language model to obtain the target prompt word and response information, which are then provided to the downstream text graph model. After the downstream text graph model generates an image based on the target prompt word, it uses the response information generated by the large language model to feed back to the target user.

[0126] In addition, during the online application process, in S406, prompt word examples and response examples can be collected and added to the training corpus to iteratively optimize the large language model. For example, for a new scenario, it is only necessary to build the corresponding training corpus to optimize the large language model. This includes the following core content:

[0127] (1) Information organization

[0128] That is, after the target subject enters the target conversation text used to generate the image, the target conversation text and the historical conversations are organized into a tuple of <target conversation text, target character>. If the current round is the first round of conversation, the target character is used to indicate that the historical conversation content is empty; if the current round is not the first round of conversation, the target character is used to indicate that the historical conversation content is at least one information tuple formed by the historical conversation text and the historical prompt words in the historical conversation text according to the association relationship.

[0129] (2) Construction of training corpus

[0130] That is, based on the tuple of <target dialogue text, target character> organized and constructed by information organization, a high-quality rewriting model is used to rewrite and reply, and obtain prompt word examples obtained by rewriting the dialogue text and reply text examples generated by the target dialogue text.

[0131] Then, the obtained prompt word examples and response text examples are manually optimized and corrected to obtain the correct results, which are used as training corpus.

[0132] Of course, you can also collect actual conversation texts and corresponding replies and prompt words to construct training corpus.

[0133] (3) Data rewriting and acquisition

[0134] Specifically, we set rewriting rules for the dialogue text used to generate images. These rules follow the following format and order to align with the base model: description of image style, description of the main image and its limitations, description of image details, description of image background modifications, special effects, composition, color tone, clarity description, and quality description. This ensures better alignment with the training corpus and the image description, resulting in high-quality images.

[0135] The description of the image style refers to the artistic style of the image, presenting the representative artistic characteristics of the image, such as two-dimensional, illustration, Chinese style, ink painting, cartoon, anime, sketch, stick figure, landscape painting, etc. Realistic descriptions such as photograph, realism, and reality are not artistic styles.

[0136] The subject of the image and the limited description refer to the subject presented in the image, which can include people, animals, plants, buildings, scenery, etc., as well as limited descriptions, such as an 18-year-old girl, blue sky, and white clouds.

[0137] Detailed description refers to the detailed description of the subject, further refining and enriching the subject. For example, an 18-year-old girl has a delicate face and wears comfortable clothes with beautiful patterns.

[0138] Background modification refers to adding background details to enhance the overall aesthetics of the image, preventing it from appearing simple or empty. For example, a green meadow with flowers and trees could be a background. Alternatively, a blank, transparent, or solid-color background can enhance the image's aesthetics.

[0139] Special effects are used to enhance and enrich the visuals, making them more refined and enhancing their impact and appeal. These include terms like "impressionist effects," "light," and "focus," such as "impasto," "perfect lighting," "texture," "ray tracing," "Tyndall effect," "blur," "sharp focus," "light and shadow," "reflection," "CG (Computer Graphics) rendering," "3D (Three Dimensions) rendering," "multiple exposure," and "double exposure."

[0140] Composition refers to the positioning of the subject, specifically how the subject is positioned, organized appropriately to create a coherent and complete picture. There are terms like lens, perspective, and depth of field, such as long shot, close-up, center shot, close-up, wide angle, frontal, and bird's-eye view.

[0141] Tone refers to the relative lightness and darkness of the picture, depicting the color of the subject and the background color, such as: white, gold, color, rich color, colorful, rich color, colorful.

[0142] Clarity description refers to the clarity of the overall image quality, such as: high definition, ultra high definition, and 8k high definition.

[0143] Quality descriptions refer to the texture and subjective feeling of the entire picture, such as: masterpiece, highest quality, master work, award-winning, immersive, simple, and dilapidated.

[0144] It is understandable that the supervised training of the base model is to align the training corpus of the text graph model with the rewritten format and order requirements.

[0145] Therefore, during the inference phase, the rewriting capability of the large language model is also to align the format and order of the text graph model.

[0146] When rewriting, the earlier the rewrite order, the more important it is. The content included in the rewrite order depends on the target conversation text. For less important content, if the target conversation text doesn't mention it, it can be excluded from the target prompt. If the target conversation text requires image quality, the rewritten target prompt will include a description of image quality. This applies to both the training and inference stages.

[0147] (4) Risk Control

[0148] That is, the risk control model is used to judge the intention of the dialogue text input by the target object for generating an image. If the judgment result shows that the image cannot be drawn, involves risks, or violates the law, the corresponding reply text is returned, that is, the image cannot be generated, and the reason for the inability to generate the image can be returned.

[0149] (5) Training module

[0150] Based on the constructed training corpus, the base model is trained using the SFT method to meet new painting requirements and align the training corpus of the text-based image model. For specific scenarios and long-tail problems, retraining can be performed using the DPO method to avoid model shortcomings and obtain a large language model.

[0151] After training a large language model, it can be provided to the outside world.

[0152] (6) Inference and prediction

[0153] The trained large language model is deployed to provide online services. When the target object inputs a dialogue text request for generating an image, the information organization module organizes the target dialogue text and historical dialogues used to generate the image into a tuple of <target dialogue text, target characters> based on the target object input. This is then input into the large language model. This allows us to obtain prompt words obtained by rewriting the target dialogue text and reply text generated by the dialogue text. The target prompt words can be input into the text-to-graph model to obtain the generated image.

[0154] (7) Effect correction module

[0155] The prompt words obtained by rewriting the dialogue text and the reply text generated by the dialogue text can be collected, and the effects can be corrected again manually to check whether there are any problems with the rewritten and reply content. The problematic marks will be manually corrected again and put into the training corpus for the next iterative training.

[0156] In summary, the Wenshengtu method provided by the embodiments of this disclosure can improve the rendering quality and user experience of Wenshengtu products, with the advantages of rapid iteration and easy expansion. It can quickly support the productization of Wenshengtu models at low cost and generate revenue.

[0157] Based on the same technical concept, the present disclosure also provides a cultural graph system, such as Figure 5 The following is a schematic diagram of the system architecture, including:

[0158] The prompt word optimization agent 501 is used to obtain the target dialogue text for generating an image in the current dialogue of the target subject; construct information to be rewritten based on the target dialogue text and historical dialogue situations; input the information to be rewritten into the large language model to obtain target prompt words generated by the large language model according to the requirements of the target dialogue text; send the target prompt words to the text-to-image tool; the target prompt words are obtained based on the text rewriting capability of the large language model;

[0159] The text graph tool 502 is used to generate an image based on the target prompt word and the text graph model in the text graph tool.

[0160] The prompt word optimization agent in the embodiment of the present disclosure constructs the information to be rewritten with reference to the large language model mentioned above, and performs a rewriting operation on it to obtain the target prompt word, which will not be repeated here.

[0161] It should be noted that the text-to-graph tool in the embodiment of the present disclosure can send the target conversation text input by the target object to the prompt word optimization agent 501 for rewriting operation to obtain the target prompt word adapted to the text-to-graph model.

[0162] In order to generate personalized reply content for the target object, the prompt word optimization agent 501 in the embodiment of the present disclosure can also use a large language model to generate a reply text for the target object text based on the method described above.

[0163] like Figure 6 As shown, in S601 , the prompt word optimization agent 501 generates the reply text while generating the target prompt word, and sends the target prompt word and the reply text together to the text map tool 502 .

[0164] In S602 , the text graph tool 502 calls the text graph model according to the target prompt word to generate an image, outputs the image to the target object 503 , and optimizes the reply text provided by the agent 501 in conjunction with the output prompt word.

[0165] In S603 , the target object 503 outputs the generated image and the reply text at the same time.

[0166] For example, if the target text of the current conversation is "I want the knife to be longer, like the ones in Japanese anime pictures," the generated response text will be "According to your request, a long Japanese anime-style knife has been generated." If the target text of the next conversation is "Change the previous one to a Chinese style with an ink painting effect," the response text will be "The image has been modified to have a Chinese style with an ink painting effect."

[0167] As a result, the reply text received by the target object is targeted at the content of the current round of dialogue, is closely related to the current round of dialogue, and can improve the user experience.

[0168] Based on the same technical concept, the embodiment of the present disclosure also proposes a text-to-image device 700, as shown in Figure 7 The device 700 comprises:

[0169] An acquisition module 701 is configured to acquire target dialogue text for generating an image of a current round of dialogue of a target object;

[0170] A construction module 702 is configured to construct to-be-revised information based on the target dialogue text and a historical dialogue situation;

[0171] A first generation module 703 is configured to input the to-be-revised information into a large language model to obtain a target prompt word generated by the large language model according to a requirement of the target dialogue text; the target prompt word is obtained based on text revision capability of the large language model;

[0172] A second generation module 704 is configured to generate an image of the current round of dialogue based on the target prompt word.

[0173] In some embodiments, the construction module is configured to:

[0174] A first splicing unit is configured to splice the target dialogue text and a target character according to a preset format to obtain the to-be-revised information, in a case where the current round of dialogue is a first round of dialogue;

[0175] The value of the target character is empty, and is used to indicate that the historical dialogue content is empty.

[0176] In some embodiments, the construction module comprises:

[0177] An extraction unit is configured to acquire historical dialogue text provided by the target object in the historical dialogue record and historical prompt words generated by the large language model according to a requirement of the historical dialogue text, in a case where the current round of dialogue is a latest round of dialogue in multiple rounds of dialogue;

[0178] A second splicing unit is configured to splice the target dialogue text and the historical dialogue record according to a preset format to obtain the to-be-revised information;

[0179] In the to-be-revised information, the historical dialogue text and the historical prompt words of the historical dialogue text are assembled into information tuples according to an association relationship, and the information tuples are values of the target character.

[0180] In some embodiments, in a case where the historical dialogue record comprises multiple rounds of historical dialogue, the historical dialogue text and the corresponding historical prompt words of each round of historical dialogue are constructed into corresponding information tuples.

[0181] In some embodiments, in a case where a first target round of dialogue in the historical dialogue record generates multiple images, the historical prompt words of the first target round of dialogue comprise sub-prompt words corresponding to the multiple images respectively.

[0182] In some embodiments, the constructing module comprises:

[0183] The image generation unit is configured to generate a text description of the reference image in a case where the target object is provided with the reference image in the current round of dialogue.

[0184] The determining unit is configured to take the text description of the reference image as content of a historical prompt word corresponding to a previous round of dialogue of the current round of dialogue, for constructing the information to be rewritten.

[0185] In some embodiments, the extracting unit is specifically configured to:

[0186] In a case where the number of dialogue rounds in the historical dialogue record is greater than m, the historical dialogue text of m rounds of historical dialogue is selected from the historical dialogue record.

[0187] wherein m is a positive integer greater than 1.

[0188] In some embodiments, in a case where the current round of dialogue is the first round of dialogue, the target prompt word is a prompt word obtained by performing a rewriting operation on the target dialogue text by the large language model.

[0189] In some embodiments, in a case where the current round of dialogue is the latest round of dialogue in the multi-round dialogue, the target prompt word is a prompt word obtained by performing a rewriting operation on the specified prompt word by the large language model.

[0190] wherein the specified prompt word is a historical prompt word of a second target round of dialogue indicated by the target dialogue text.

[0191] In some embodiments, the constructing module is further configured to:

[0192] In a case where the target dialogue text includes a specific object in the preset set, obtaining a preset description of the specific object.

[0193] Constructing the information to be rewritten based on the preset description of the specific object, the target dialogue text, and the historical dialogue.

[0194] In some embodiments, further comprising a feedback module configured to:

[0195] Obtaining a reply text generated by the large language model for the target dialogue text based on the information to be rewritten;

[0196] Feeding back the reply text to the target object together with the generated image.

[0197] In some embodiments, further comprising a training module configured to train the large language model based on the following manner:

[0198] The base model is supervisedly trained based on a first training corpus set to obtain an intermediate model; wherein the sample prompt words of the first training corpus set are aligned to generate the format requirements of the training corpus of the text-graph model of the image, wherein, when the target prompt words include multiple contents of the following, the relative order of the multiple contents meets the following description order requirements;

[0199] A direct preference optimization training method is used to optimize the intermediate model based on the second training corpus to obtain a large language model.

[0200] In some embodiments, the content of the target prompt word includes at least one of the following, and satisfies the following description order:

[0201] Description of the picture style, description of the picture subject and limitations, description of the picture details, description of the picture background modification, special effects, composition, color tone, clarity description, and quality description.

[0202] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0203] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0204] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0205] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0206] like Figure 8As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0207] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0208] The computing unit 801 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the Vincent diagram method. For example, in some embodiments, the Vincent diagram method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the Vincent diagram method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the Vincent diagram method by any other suitable means (e.g., via firmware).

[0209] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0210] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.

[0211] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0212] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0213] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0214] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0215] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0216] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for generating a cultural image, comprising: Obtain the target conversation text for generating images in the current round of conversation of the target object; Constructing information to be rewritten based on the target dialogue text and historical dialogue situations; Inputting the information to be rewritten into a large language model to obtain target prompt words generated by the large language model according to the requirements of the target dialogue text; The target prompt word is obtained based on the text rewriting capability of the large language model; Based on the target prompt word, generating an image of the current round of dialogue; In a case where the current round of dialogue is the latest round of dialogue in multiple rounds of dialogue, the target prompt word is a prompt word obtained by rewriting a designated prompt word by the large language model; Wherein, the designated prompt word is a historical prompt word of the second target round of dialogue indicated by the target dialogue text; The historical prompt words of the second target round of dialogue are historical prompt words corresponding to the previous round of historical dialogue text of the target dialogue text, or historical prompt words corresponding to the historical dialogue text referred to by the target dialogue text.

2. The method according to claim 1, wherein The constructing the information to be rewritten based on the target dialogue text and the historical dialogue situation includes: In a case where the current round of dialogue is the first round of dialogue, the target dialogue text and the target characters are concatenated according to a preset format to obtain the information to be rewritten; The value of the target character is empty, which is used to indicate that the historical conversation content is empty.

3. The method according to claim 1, wherein The constructing the information to be rewritten based on the target dialogue text and the historical dialogue situation includes: When the current round of dialogue is the latest round of dialogue in multiple rounds of dialogue, obtaining a historical dialogue text provided by the target object in the historical dialogue record, and a historical prompt word generated by the large language model according to the requirements of the historical dialogue text; splicing the target conversation text and the historical conversation record according to a preset format to obtain the information to be rewritten; Among them, in the information to be rewritten, the historical conversation text and the historical prompt words of the historical conversation text are organized into an information tuple according to an association relationship, and the information tuple is the value of the target character.

4. The method according to claim 3, wherein: In the case where the historical conversation record includes multiple rounds of historical conversations, the historical conversation text of each round of historical conversation and the corresponding historical prompt words are constructed into corresponding information tuples.

5. The method according to claim 3, wherein In a case where a plurality of images are generated in the first target-round dialogue in the historical dialogue record, the historical prompt words of the first target-round dialogue include sub-prompt words corresponding to the plurality of images respectively.

6. The method according to claim 2 or 3, further comprising: In the current round of dialogue, if the target object provides a reference image, generating a text description of the reference image; The text description of the reference image is used as the content of the historical prompt words corresponding to the previous round of dialogue of the current round of dialogue, so as to construct the information to be rewritten.

7. The method according to claim 3, wherein: The step of obtaining the historical conversation text provided by the target object in the historical conversation record includes: If the number of conversation rounds in the historical conversation record is greater than m, selecting historical conversation texts of m rounds of historical conversation from the historical conversation record; Wherein, m is a positive integer greater than 1.

8. The method according to claim 1, wherein In a case where the current round of dialogue is the first round of dialogue, the target prompt word is a prompt word obtained by rewriting the target dialogue text by the large language model.

9. The method according to claim 1, wherein The constructing the information to be rewritten based on the target dialogue text and the historical dialogue situation includes: When the target dialogue text includes a specific object in a preset set, obtaining a preset description of the specific object; The information to be rewritten is constructed based on a preset description of the specific object, the target dialogue text, and the historical dialogue situation.

10. The method according to claim 1, further comprising: Obtaining a reply text generated by the large language model for the target dialogue text based on the information to be rewritten; The reply text and the generated image are fed back to the target object.

11. The method according to claim 1, wherein The large language model is trained based on the following method: The base model is supervisedly trained based on the first training corpus set to obtain an intermediate model; wherein the sample prompt words of the first training corpus set are aligned to generate the format requirements of the training corpus of the text-graph model of the image; The large language model is obtained by optimizing the intermediate model based on the second training corpus using a direct preference optimization training method.

12. The method according to claim 1, wherein The target prompt word includes at least one of the following, and satisfies the following description order: Description of the picture style, description of the picture subject and limitations, description of the picture details, description of the picture background modification, special effects, composition, color tone, clarity description, and quality description.

13. A Wensheng diagram device, comprising: An acquisition module, configured to acquire target conversation text of a current conversation round of a target object for generating an image; A construction module, configured to construct information to be rewritten based on the target conversation text and historical conversation situations; A first generating module is configured to input the information to be rewritten into a large language model to obtain target prompt words generated by the large language model according to the requirements of the target dialogue text; The target prompt word is obtained based on the text rewriting capability of the large language model; A second generation module is used to generate an image of the current round of dialogue based on the target prompt word; In a case where the current round of dialogue is the latest round of dialogue in multiple rounds of dialogue, the target prompt word is a prompt word obtained by rewriting a designated prompt word by the large language model; Wherein, the designated prompt word is a historical prompt word of the second target round of dialogue indicated by the target dialogue text; The historical prompt words of the second target round of dialogue are historical prompt words corresponding to the previous round of historical dialogue text of the target dialogue text, or historical prompt words corresponding to the historical dialogue text referred to by the target dialogue text.

14. The device according to claim 13, wherein The building blocks are used to: A first splicing unit is configured to, when the current round of dialogue is the first round of dialogue, splice the target dialogue text and target characters according to a preset format to obtain the information to be rewritten; The value of the target character is empty, which is used to indicate that the historical conversation content is empty.

15. The device according to claim 13, wherein The building blocks include: an extraction unit, configured to obtain, when the current round of dialogue is the latest round of dialogue in a multi-round dialogue, a historical dialogue text provided by the target party in the historical dialogue record, and historical prompt words generated by the large language model according to requirements of the historical dialogue text; A second splicing unit is used to splice the target conversation text and the historical conversation record according to a preset format to obtain the information to be rewritten; Among them, in the information to be rewritten, the historical conversation text and the historical prompt words of the historical conversation text are organized into an information tuple according to an association relationship, and the information tuple is the value of the target character.

16. The device according to claim 15, wherein In the case where the historical conversation record includes multiple rounds of historical conversations, the historical conversation text of each round of historical conversation and the corresponding historical prompt words are constructed into corresponding information tuples.

17. The device according to claim 15, wherein In a case where a plurality of images are generated in the first target-round dialogue in the historical dialogue record, the historical prompt words of the first target-round dialogue include sub-prompt words corresponding to the plurality of images respectively.

18. The device according to claim 14 or 15, wherein the building block comprises: A picture-to-text unit, configured to generate a text description of a reference image when the target object provides a reference image in the current round of dialogue; The determining unit is configured to use the text description of the reference image as the content of the historical prompt words corresponding to the previous round of dialogue before the current round of dialogue, so as to construct the information to be rewritten.

19. The device according to claim 15, wherein The extraction unit is specifically used for: If the number of conversation rounds in the historical conversation record is greater than m, selecting historical conversation texts of m rounds of historical conversation from the historical conversation record; Wherein, m is a positive integer greater than 1.

20. The apparatus according to claim 13, wherein In a case where the current round of dialogue is the first round of dialogue, the target prompt word is a prompt word obtained by rewriting the target dialogue text by the large language model.

21. The apparatus according to claim 13, wherein The building block is further used to: When the target dialogue text includes a specific object in a preset set, obtaining a preset description of the specific object; The information to be rewritten is constructed based on a preset description of the specific object, the target dialogue text, and the historical dialogue situation.

22. The apparatus according to claim 13, further comprising a feedback module for: Obtaining a reply text generated by the large language model for the target dialogue text based on the information to be rewritten; The reply text and the generated image are fed back to the target object.

23. The apparatus according to claim 13, further comprising a training module, configured to train the large language model based on the following method: Perform supervised training on the base model based on the first training corpus to obtain an intermediate model; in, The format requirements of the training corpus for generating the text-graph model of the image by aligning the sample prompt words of the first training corpus set are met, wherein, when the target prompt words include multiple contents of the following, the relative order of the multiple contents meets the requirements of the following description order; The large language model is obtained by optimizing the intermediate model based on the second training corpus using a direct preference optimization training method.

24. The apparatus according to claim 13, wherein The target prompt word includes at least one of the following, and satisfies the following description order: Description of the picture style, description of the picture subject and limitations, description of the picture details, description of the picture background modification, special effects, composition, color tone, clarity description, and quality description.

25. A cultural graph system comprising: The prompt word optimization agent is used to obtain the target dialogue text for generating images in the current dialogue round of the target object; Constructing information to be rewritten based on the target conversation text and historical conversation situations; inputting the information to be rewritten into a large language model to obtain target prompt words generated by the large language model according to the requirements of the target conversation text; sending the target prompt words to a text-generating diagram tool; the target prompt words are obtained based on the text rewriting capability of the large language model; A text image tool, configured to generate an image based on the target prompt word and a text image model in the text image tool; In a case where the current round of dialogue is the latest round of dialogue in multiple rounds of dialogue, the target prompt word is a prompt word obtained by rewriting a designated prompt word by the large language model; Wherein, the designated prompt word is a historical prompt word of the second target round of dialogue indicated by the target dialogue text; The historical prompt words of the second target round of dialogue are historical prompt words corresponding to the previous round of historical dialogue text of the target dialogue text, or historical prompt words corresponding to the historical dialogue text referred to by the target dialogue text.

26. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

27. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.

28. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and storage medium

    CN116797684A

  • Image generation method and device, electronic equipment and storage medium

    CN116843795A