Image display method, device and equipment, computer readable medium and program product

By performing graph generation intent preprocessing and information augmentation on the target dialogue, and using a pre-trained language model to generate images, the problems of insufficient training samples and insufficient customization in existing technologies are solved, and efficient and accurate image generation is achieved.

CN121414865APending Publication Date: 2026-01-27JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511500915.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing technologies require a large number of training samples for text-to-image tasks, resulting in inaccurate image generation and an inability to customize the images to the dialogue content, leading to low accuracy of the generated images.

Method used

By preprocessing the target dialogue with graph generation intent, extracting graph generation element information and performing information enhancement, and using the pre-trained first language model and text-generated graph model to generate images, an image generation framework based on target dialogue is constructed.

Benefits of technology

Without requiring a large number of training samples, it accurately generates images corresponding to the target dialogue, achieving customized image generation and ensuring semantic consistency between the generated images and the dialogue content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121414865A_ABST
    Figure CN121414865A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image display method, device and equipment, a computer readable medium and a program product. A specific embodiment of the method comprises the following steps: in response to acquired graph generation information for a target dialogue, performing graph generation intention preprocessing on the target dialogue to obtain a preprocessed dialogue; generating prompt information for extracting graph generation element information corresponding to the preprocessed dialogue and performing information enhancement on the graph generation element information; inputting the prompt information into a pre-trained first large language model to obtain graph generation element enhancement information corresponding to the preprocessed dialogue; generating an image corresponding to the preprocessed dialogue by using a pre-trained text graph model according to the graph generation element enhancement information; and displaying the image in a dialogue page corresponding to the target dialogue. The implementation mode is related to artificial intelligence, the image under the target conversation can be accurately and efficiently generated, and the image requirement of the interaction object is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of artificial intelligence technology, specifically to image display methods, apparatus, devices, computer-readable media, and program products. Background Technology

[0002] Currently, with the continuous development of artificial intelligence, text-to-image (TPE) tasks are becoming increasingly widespread. TPE tasks involve generating corresponding images based on descriptive text. For image generation during dialogue, existing methods often require a large number of training samples to train the graph generation model, enabling it to generate images from dialogue scenarios, which incurs significant data costs. Furthermore, existing methods often cannot customize image generation based on dialogue content, resulting in low accuracy of the generated images. Summary of the Invention

[0003] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0004] Some embodiments of this disclosure provide image display methods, apparatuses, devices, computer-readable media, and program products to address the technical problems mentioned in the background section above.

[0005] In a first aspect, some embodiments of this disclosure provide an image display method, comprising: in response to obtaining graph generation information for a target dialogue, performing graph generation intent preprocessing on the target dialogue to obtain a preprocessed dialogue; generating information on graph generation elements corresponding to the preprocessed dialogue and information enhancement information on the graph generation elements; inputting the prompt information into a pre-trained first language model to obtain graph generation element enhancement information corresponding to the preprocessed dialogue; generating an image corresponding to the preprocessed dialogue using a pre-trained text-to-graph model based on the graph generation element enhancement information; and displaying the image on a dialogue page corresponding to the target dialogue.

[0006] Optionally, the above-mentioned graph generation intent preprocessing of the target dialogue to obtain a preprocessed dialogue includes: extracting at least one graph generation intent corresponding to the target dialogue; determining the intent proposal time corresponding to each of the at least one graph generation intent; converting the target dialogue into dialogue text in a target format; and adding the graph generation intent whose intent proposal time satisfies the target time condition as the actual graph generation intent to the dialogue text to obtain the preprocessed dialogue.

[0007] Optionally, the first language model further generates: the graph generation element information; and before performing graph generation intent preprocessing on the target dialogue in response to obtaining graph generation information for the target dialogue to obtain a preprocessed dialogue, the method further includes: querying the dialogue graph database to see if there is dialogue graph data whose matching information with the graph generation element information satisfies the target matching condition; in response to determining that it exists, generating an image corresponding to the preprocessed dialogue based on the image in the dialogue graph data; and performing graph generation intent preprocessing on the target dialogue in response to obtaining graph generation information for the target dialogue to obtain a preprocessed dialogue includes: in response to receiving the graph generation information and not finding dialogue graph data that satisfies the target matching condition, performing graph generation intent preprocessing on the target dialogue to obtain a preprocessed dialogue.

[0008] Optionally, after inputting the above prompt information into the pre-trained first language model to obtain the graph generation element enhancement information corresponding to the preprocessed dialogue, the method further includes: using the above first language model to convert the graph generation element enhancement information into enhancement information in the target language to obtain converted enhancement information; and determining the converted enhancement information as graph generation element enhancement information.

[0009] Optionally, before performing graph generation intent preprocessing on the target dialogue in response to obtaining graph generation information for the target dialogue to obtain a preprocessed dialogue, the method further includes: using a pre-trained second language model to determine whether the target dialogue has a graph generation requirement; in response to determining that it exists, using the second language model to determine the graph generation type corresponding to the target dialogue; and in response to determining that the graph generation type is an open graph generation type, generating graph generation information corresponding to the target dialogue.

[0010] Optionally, the above prompt information may also include: if the extracted graph generation element information is empty, the key dialogue information corresponding to the target dialogue is used as a sub-prompt information of the graph generation element information, and at least one limitation prompt information regarding image generation quality limitations.

[0011] Optionally, after generating the image corresponding to the preprocessed dialogue using the pre-trained text-based graph model based on the graph generation element enhancement information, the method further includes: using a pre-trained third language model to determine first matching information between the semantic content of the first image corresponding to the image and the graph generation element information, and to determine second matching information between the semantic content of the first image and the dialogue content corresponding to the target dialogue; determining whether the image is an image to be displayed based on the first and second matching information; in response to determining that it is not an image to be displayed, generating a regenerated image using the first and the text-based graph model based on the first and second matching information; using the third language model to determine third matching information between the semantic content of the second image corresponding to the regenerated image and the graph generation element information, and to determine fourth matching information between the semantic content of the second image and the dialogue content corresponding to the target dialogue; determining whether the regenerated image is an image to be displayed based on the third and fourth matching information; and in response to determining that it is an image to be displayed, identifying the regenerated image as an image.

[0012] Secondly, some embodiments of this disclosure provide an image display apparatus, including: a preprocessing unit configured to preprocess the target dialogue for graph generation intent in response to obtaining graph generation information for a target dialogue, thereby obtaining a preprocessed dialogue; a first generation unit configured to generate prompt information that extracts graph generation element information corresponding to the preprocessed dialogue and enhances the graph generation element information; an input unit configured to input the prompt information into a pre-trained first language model to obtain graph generation element enhancement information corresponding to the preprocessed dialogue; a second generation unit configured to generate an image corresponding to the preprocessed dialogue based on the graph generation element enhancement information and using a pre-trained text-to-graph model; and a display unit configured to display the image on a dialogue page corresponding to the target dialogue.

[0013] Optionally, the preprocessing unit can be configured to: extract at least one graph generation intent corresponding to the target dialogue; determine the intent proposal time corresponding to each of the at least one graph generation intents; convert the target dialogue into dialogue text in the target format; and add the graph generation intents whose intent proposal times satisfy the target time conditions as actual graph generation intents to the dialogue text to obtain the preprocessed dialogue.

[0014] Optionally, the first language model further generates: the graph generation element information; and the apparatus further includes: querying the dialogue graph database to see if there is dialogue graph data whose matching information with the graph generation element information satisfies the target matching condition; in response to determining that there is, generating an image corresponding to the preprocessed dialogue based on the image in the dialogue graph data; and the above-mentioned graph generation intent preprocessing of the target dialogue in response to obtaining graph generation information for the target dialogue to obtain a preprocessed dialogue includes: in response to receiving the graph generation information and not finding dialogue graph data that satisfies the target matching condition, performing graph generation intent preprocessing of the target dialogue to obtain a preprocessed dialogue.

[0015] Optionally, the apparatus further includes: using the first language model described above to convert the graph generation element enhancement information into enhancement information in the target language to obtain converted enhancement information; and determining the converted enhancement information as graph generation element enhancement information.

[0016] Optionally, the apparatus further includes: using a pre-trained second language model to determine whether the target dialogue has a graph generation requirement; in response to determining that it exists, using the second language model to determine the graph generation type corresponding to the target dialogue; and in response to determining that the graph generation type is an open graph generation type, generating graph generation information corresponding to the target dialogue.

[0017] Optionally, the apparatus further includes: using a pre-trained third language model to determine first matching information between the semantic content of the first image corresponding to the image and the graph generation element information, and to determine second matching information between the semantic content of the first image and the dialogue content corresponding to the target dialogue; determining whether the image is an image to be displayed based on the first matching information and the second matching information; in response to determining that it is not an image to be displayed, generating a regenerated image using the first language model and the text-generated graph model based on the first matching information and the second matching information; using the third language model to determine third matching information between the semantic content of the second image corresponding to the regenerated image and the graph generation element information, and to determine fourth matching information between the semantic content of the second image and the dialogue content corresponding to the target dialogue; determining whether the regenerated image is an image to be displayed based on the third matching information and the fourth matching information; and in response to determining that it is an image to be displayed, identifying the regenerated image as an image.

[0018] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.

[0019] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0020] Fifthly, some embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0021] The above embodiments of this disclosure have the following beneficial effects: the image display method of some embodiments of this disclosure can accurately and efficiently generate images under the target dialogue to meet the image needs of the interactive object. Specifically, the reason why the images under the target dialogue are not accurate enough is that: training graph generation models with insufficient training samples to enable them to generate images from dialogue scenarios often results in inaccurate image generation. In addition, customized image generation for dialogue content is often not possible, leading to low accuracy of the generated images. Based on this, the image display method of some embodiments of this disclosure firstly, in response to obtaining graph generation information for the target dialogue, performs graph generation intent preprocessing on the target dialogue to obtain a preprocessed dialogue. Here, by performing graph generation intent preprocessing, the subsequent text-generated graph model can clearly understand what kind of image to generate under the target dialogue (i.e., what kind of intent the image should be). Then, it generates and extracts graph generation element information corresponding to the preprocessed dialogue and provides information enhancement prompts for the graph generation element information. Here, by constructing prompts, the first language model can be used to extract each graph generation element in the preprocessed dialogue and enhance the semantics of the graph generation elements. Next, the aforementioned prompt information is input into the pre-trained first language model. This model can accurately generate graph generation element enhancement information corresponding to the pre-processed dialogue without requiring a large number of training samples. This allows the subsequent text-to-graph model to clearly understand the rich elements of the generated image. Furthermore, based on the graph generation element enhancement information, the pre-trained text-to-graph model can accurately generate the image corresponding to the pre-processed dialogue. Finally, the image is displayed on the dialogue page corresponding to the target dialogue. In summary, by using intent preprocessing, the first language model for extracting and semantically enhancing graph generation element information, and the application of the text-to-graph model, an image generation framework based on the target dialogue can be constructed without requiring a large number of training samples. This framework enables customized image generation based on the semantic content of the target dialogue and the actual image generation intent, ensuring semantic consistency between the generated image and the dialogue content corresponding to the target dialogue. Attached Figure Description

[0022] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0023] Figure 1 This is a schematic diagram of an application scenario of an image display method according to some embodiments of the present disclosure; Figure 2 This is a flowchart of some embodiments of the image display method according to the present disclosure; Figure 3 This is a flowchart of some other embodiments of the image display method according to the present disclosure; Figure 4 These are schematic diagrams illustrating the structure of some embodiments of the image display device according to the present disclosure; Figure 5 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0025] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0026] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0027] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0028] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0029] Before performing any of the operations involving the collection, storage, or use of user personal information (such as target conversations) disclosed in this disclosure, the relevant organizations or individuals shall fulfill their obligations, including conducting personal information security impact assessments, informing personal information subjects, and obtaining prior authorization and consent from personal information subjects.

[0030] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0031] Figure 1 This is a schematic diagram of an application scenario of an image display method according to some embodiments of the present disclosure.

[0032] exist Figure 1 In this application scenario, firstly, in response to obtaining graph generation information for the target dialogue 102, the electronic device 101 can perform graph generation intent preprocessing on the target dialogue 102 to obtain a preprocessed dialogue 103. In this application scenario, the target dialogue 102 can be "[{'user': 'Then if he answers incorrectly, you ask him,' 'agent': 'Yes, I've sent it to you, did you see it? Is it the breakfast picture you wanted?'}, {'user': 'I saw it, send me another picture of spring, then flowers, a stream, and a forest,', 'agent': ''}]". The preprocessed dialogue 103 can be "User outputs: Then if he answers incorrectly, you ask him, Agent outputs: Yes, I've sent it to you, did you see it? Is it the breakfast picture you wanted? User's latest output:: I saw it, send me another picture of spring, then flowers, a stream, and a forest." Then, electronic device 101 can generate graph generation element information corresponding to the preprocessed dialogue 103 and prompt information 104 that enhances the graph generation element information. Next, electronic device 101 can input the prompt information 104 into the pre-trained first large language model 105 to obtain graph generation element enhancement information 106 corresponding to the preprocessed dialogue. In this application scenario, graph generation element enhancement information 106 could be "a spring scene photo with flowers, a stream, and a forest. The flowers are in full bloom, colorful, the stream flows gently, and the trees are tall, providing shade. The atmosphere is tranquil and refreshing, capturing the essence of spring." Furthermore, electronic device 101 can generate an image 108 corresponding to the preprocessed dialogue 103 based on the graph generation element enhancement information 106 and using the pre-trained text-based graph model 107. Finally, electronic device 101 can display the image 108 on the dialogue page corresponding to the target dialogue 102.

[0033] It should be noted that the aforementioned electronic device 101 can be either hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the electronic device is software, it can be installed in the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here.

[0034] It should be understood that Figure 1 The number of electronic devices shown is merely illustrative. Any number of electronic devices can be used depending on the implementation requirements.

[0035] Continue to refer to Figure 2 The diagram illustrates a flow 200 of some embodiments of an image display method according to the present disclosure. This image display method includes the following steps: Step 201: In response to obtaining graph generation information for the target dialogue, perform graph generation intent preprocessing on the target dialogue to obtain a preprocessed dialogue.

[0036] In some embodiments, in response to obtaining graph generation information for the target dialogue, the execution entity of the above image display method (e.g., Figure 1 The electronic device 101 shown can perform graph generation intent preprocessing on the target dialogue to obtain a preprocessed dialogue. The target dialogue can be a dialogue between a user and a question-answering robot. In practice, the target dialogue can be a multi-turn dialogue used to subsequently generate a corresponding descriptive graph. This multi-turn dialogue can be the content of a multi-turn dialogue between the user and the question-answering robot. Graph generation information can be a graph generation request generated based on the dialogue content corresponding to the target dialogue. Graph generation information can be request information in a predetermined format. Graph generation intent preprocessing can be intent processing on at least one graph generation intent corresponding to the target dialogue. The target dialogue corresponds to at least one graph generation intent. For example, for the target dialogue "Generate an image of a bird first, then generate an image of a dog," the corresponding at least one graph generation intent can include: "Intent to generate a bird" and "Intent to generate a dog." The graph generation intent can be the intent content for generating the graph. The preprocessed dialogue can be the dialogue content after graph generation intent preprocessing. In practice, at least one graph generation intent corresponding to the target dialogue can be extracted using a pre-trained intent recognition model. For example, the intent recognition model can be a Transformer model.

[0037] As an example, firstly, the aforementioned execution entity can input the target dialogue into the intent recognition model to obtain at least one graph generation intent. Then, it generates at least one intent key-value pair corresponding to the at least one graph generation intent. Finally, it combines the at least one intent key-value pair with the target dialogue to obtain a preprocessed dialogue.

[0038] In some optional implementations of certain embodiments, the aforementioned execution entity may perform graph generation intent preprocessing on the aforementioned target dialogue to obtain a preprocessed dialogue, including the following steps: The first step is to extract at least one graph corresponding to the target dialogue to generate intent.

[0039] As an example, the aforementioned execution entity can utilize the first major language model to extract at least one graph generation intent corresponding to the target dialogue.

[0040] The second step is to determine the intent proposal time corresponding to each of the at least one graph generation intent mentioned above. The intent proposal time can be the time when the target dialogue proposes generating an image based on the graph generation intent.

[0041] As an example, firstly, the aforementioned execution entity can utilize the first major language model to locate the dialogue text position corresponding to each graph generation intent in the target dialogue, thereby obtaining at least one dialogue text position. Then, the dialogue time of the text content corresponding to each dialogue text position is used as the intent presentation time, thereby obtaining at least one intent presentation time.

[0042] The third step involves adding the graph generation intent whose corresponding intent release time meets the target time condition from at least one of the above graph generation intents to the corresponding dialogue text of the target dialogue, thus obtaining the preprocessed dialogue. The target time condition is that the release time of the graph generation intent is the latest among the release times of at least one intent. The dialogue text can be in text format. The actual graph generation intent corresponds to the actual image to be generated subsequently.

[0043] For example, the target dialogue text might be: [{'user': 'Then if he answers incorrectly, just ask him,' 'agent': 'Okay, I've already sent it to you, did you see it? Is it the breakfast picture you wanted?'}, {'user': 'I saw it, send me another picture of spring, with flowers, a stream, and a forest,' 'agent': ''}]. The corresponding preprocessed dialogue would be: User outputs: Then if he answers incorrectly, just ask him. Agent outputs: Okay, I've already sent it to you, did you see it? Is it the breakfast picture you wanted? User's latest output: I saw it, send me another picture of spring, with flowers, a stream, and a forest.

[0044] In some optional implementations of certain embodiments, before step 201, the steps further include: The first step is to use a pre-trained second large language model to determine whether the target dialogue has a graph generation requirement. This second large language model can be a large language model that performs various processing on the target dialogue. In practice, these processing methods can include: determining the graph generation requirement and determining the graph generation type. The graph generation requirement can be that the dialogue text contains a need to generate a graph.

[0045] As an example, firstly, the aforementioned execution entity can generate a requirement confirmation message to confirm whether the target dialogue has a graph generation requirement. Then, the requirement confirmation message is input into the second major language model to obtain the graph generation requirement.

[0046] The second step, in response to the confirmed existence, utilizes the second major language model mentioned above to determine the graph generation type corresponding to the target dialogue. The graph generation type can be the type of graph generated. In practice, the graph generation type is one of the following: basic selfie generation type, contextualized selfie generation type, or open-ended graph generation type. A basic selfie generation type can be a type where the user only needs a virtual avatar selfie generated by AI, without specific scene or style requirements, and the system randomly returns results from a preset image library. A contextualized selfie generation type can be a type where the user provides a text description and reference images, requiring the AI ​​to combine the two to generate a selfie image that conforms to a specific scene. An open-ended graph generation type can be a type where the user generates a completely new image based solely on a text description, without involving real people or existing image references, and requires an open-ended generation process.

[0047] As an example, firstly, type generation prompts are generated to determine the graph generation type corresponding to the target dialogue. Then, the type generation prompts are input into the second language model to obtain the graph generation type.

[0048] The third step is to generate graph generation information corresponding to the target dialogue in response to determining that the above graph generation type is an open graph generation type.

[0049] As an example, in response to determining that the above graph generation type is an open graph generation type, a graph generation request for open graph generation for the target dialogue is generated as graph generation information.

[0050] Step 202: Generate graph generation element information corresponding to the above preprocessed dialogue and prompt information for information enhancement of the above graph generation element information.

[0051] In some embodiments, the execution entity can generate graph generation element information corresponding to the preprocessed dialogue and prompts for enhancing the graph generation element information. The graph generation element information can be the information of various elements in the generated graph. For example, the graph generation element information may include: descriptive object, image style, background information, and action information. Enhancing the graph generation element information can involve enhancing the semantic content corresponding to each element in the graph generation element information. The prompts can be prompt words.

[0052] As an example, first, a prompt word template is obtained. Then, the preprocessed dialogue is added to the prompt word template to obtain the graph generation feature information corresponding to the preprocessed dialogue and the prompt information that enhances the graph generation feature information.

[0053] In some optional implementations of certain embodiments, the above-mentioned prompting information further includes: sub-prompting information such as the key dialogue information corresponding to the target dialogue as the key dialogue information when the extracted graph generation element information is empty, and restrictive prompting information regarding whether the image generation quality meets the target quality condition. Here, the key dialogue information can be key text content in the dialogue text corresponding to the target dialogue. Image generation quality can be the quality of the image content of the generated image. The target quality condition can be a condition characterizing the image quality of the generated image as high quality. In practice, image quality can be defined by various indicators. For example, it can be characterized by image sharpness and image subject integrity. The corresponding target quality conditions include: a sub-condition that the corresponding image sharpness is higher than the target sharpness and a sub-condition that the image subject integrity is subject integrity.

[0054] Step 203: Input the above prompt information into the pre-trained first language model to obtain the graph generation element enhancement information corresponding to the above pre-processed dialogue.

[0055] In some embodiments, the executing entity can input the aforementioned prompt information into a pre-trained first large language model to obtain graph generation element enhancement information corresponding to the preprocessed dialogue. The first large language model can be a large language model that supports graph generation element information extraction and semantic content enhancement of the graph generation element information. In practice, the first large language model can be a commercially available large language model. The first large language model can be a pre-fine-tuned large language model. The graph generation element enhancement information can be information after semantic content enhancement of the graph generation elements corresponding to the preprocessed dialogue.

[0056] In some optional implementations of certain embodiments, after step 203, the steps further include: The first step is to use the first major language model described above to convert the enhanced information of the generated graph elements into enhanced information in the target language, thus obtaining the transformed enhanced information. The target language can be the data language supported by the subsequent text-generated graph model for data input. For example, the target language could be English.

[0057] As an example, firstly, conversion prompts are generated to translate the augmented information of the generated graph features into the target language. Then, the conversion prompts are input into the first large language model to obtain the converted augmented information.

[0058] The second step is to identify the above-mentioned transformation enhancement information as graph generation feature enhancement information.

[0059] Step 204: Based on the above graph generation feature enhancement information, use the pre-trained text-to-graph model to generate the image corresponding to the preprocessed dialogue.

[0060] In some embodiments, the aforementioned execution entity can generate images corresponding to preprocessed dialogues using a pre-trained text-to-graph model based on the graph generation feature enhancement information. The text-to-graph model can be a model that generates corresponding images based on descriptive text. In practice, the text-to-graph model can be a diffusion model or a multimodal large model.

[0061] As an example, firstly, graph generation prompts corresponding to the graph generation feature enhancement information are generated. Then, the graph generation prompts are input into the text-to-graph model to obtain the image.

[0062] In some optional implementations of certain embodiments, after step 204, the steps further include: The first step involves using a pre-trained third language model to determine the first matching information between the semantic content of the first image corresponding to the aforementioned image and the graph generation element information, and to determine the second matching information between the semantic content of the first image and the dialogue content corresponding to the target dialogue. The third language model can be a large language model that supports determining the matching status between semantic content. The third language model can be the same as the first or second language model, or it can be a different large language model. The first matching information represents the degree of matching between the semantic content of the image and each graph generation element in the graph generation element information. The first matching information can be in numerical form; a larger value indicates a higher degree of matching. The second matching information guarantees the degree of matching between the semantic content of the first image and the dialogue content. The second matching information can also be in numerical form; a larger value indicates a higher degree of matching.

[0063] The second step is to determine whether the image is the one to be displayed, based on the first matching information and the second matching information mentioned above.

[0064] As an example, firstly, the aforementioned executing entity can perform a weighted summation of the first matching information and the second matching information to obtain a weighted sum value. Then, in response to determining that the weighted sum value is greater than a target value, the aforementioned image is determined to be the image to be displayed. In response to determining that the weighted sum value is less than or equal to the aforementioned target value, the aforementioned image is determined not to be the image to be displayed.

[0065] The third step is to generate a regenerated image based on the first matching information and the second matching information, using the first large language model and the text-generated image model.

[0066] As an example, firstly, image adjustment prompts are generated to enhance the image content based on first and second matching information. Then, these prompts are input into a first large language model to obtain the adjusted image generation feature enhancement information. Next, based on this enhanced information, a regenerated image is generated using a text-to-image model.

[0067] Fourth, using the third language model described above, determine the third matching information between the semantic content of the second image corresponding to the regenerated image and the graph generation element information, and determine the fourth matching information between the semantic content of the second image and the dialogue content corresponding to the target dialogue. For an explanation of the third matching information, please refer to the explanation of the first matching information. For an explanation of the fourth matching information, please refer to the explanation of the second matching information.

[0068] Fifth, based on the third and fourth matching information mentioned above, determine whether the regenerated image is the image to be displayed. The specific implementation method will not be elaborated further; please refer to the implementation method for determining whether an image is the image to be displayed.

[0069] Step 6: In response to determining that it is an image to be displayed, the above-generated image is identified as an image.

[0070] Step 205: Display the above image on the dialogue page corresponding to the target dialogue.

[0071] In some embodiments, the aforementioned executing entity may display the aforementioned image on the dialogue page corresponding to the aforementioned target dialogue. The dialogue page may be the page where the user and the question-answering robot conduct their conversation in the target dialogue.

[0072] The above embodiments of this disclosure have the following beneficial effects: the image display method of some embodiments of this disclosure can accurately and efficiently generate images under the target dialogue to meet the image needs of the interactive object. Specifically, the reason why the images under the target dialogue are not accurate enough is that: training graph generation models with insufficient training samples to enable them to generate images from dialogue scenarios often results in inaccurate image generation. In addition, customized image generation for dialogue content is often not possible, leading to low accuracy of the generated images. Based on this, the image display method of some embodiments of this disclosure firstly, in response to obtaining graph generation information for the target dialogue, performs graph generation intent preprocessing on the target dialogue to obtain a preprocessed dialogue. Here, by performing graph generation intent preprocessing, the subsequent text-generated graph model can clearly understand what kind of image to generate under the target dialogue (i.e., what kind of intent the image should be). Then, it generates and extracts graph generation element information corresponding to the preprocessed dialogue and provides information enhancement prompts for the graph generation element information. Here, by constructing prompts, the first language model can be used to extract each graph generation element in the preprocessed dialogue and enhance the semantics of the graph generation elements. Next, the aforementioned prompt information is input into the pre-trained first language model. This model can accurately generate graph generation element enhancement information corresponding to the pre-processed dialogue without requiring a large number of training samples. This allows the subsequent text-to-graph model to clearly understand the rich elements of the generated image. Furthermore, based on the graph generation element enhancement information, the pre-trained text-to-graph model can accurately generate the image corresponding to the pre-processed dialogue. Finally, the image is displayed on the dialogue page corresponding to the target dialogue. In summary, by using intent preprocessing, the first language model for extracting and semantically enhancing graph generation element information, and the application of the text-to-graph model, an image generation framework based on the target dialogue can be constructed without requiring a large number of training samples. This framework enables customized image generation based on the semantic content of the target dialogue and the actual image generation intent, ensuring semantic consistency between the generated image and the dialogue content corresponding to the target dialogue.

[0073] Further reference Figure 3 The diagram illustrates a flow 300 of another embodiment of the image display method according to the present disclosure. This image display method includes the following steps: Step 301: Query the dialogue graph database to see if there is dialogue graph data whose matching information with the above graph generation feature information meets the target matching conditions.

[0074] In some embodiments, the execution entity (e.g. Figure 1The electronic device 101 shown can query the dialogue graph database to see if dialogue graph data exists that matches the aforementioned graph generation element information to meet the target matching condition. The dialogue graph database can store a database of each dialogue and its corresponding generated image. That is, the dialogue graph data can include: dialogue information, generated image, graph generation element information, and image description information. The target matching condition can be the dialogue graph data with the highest matching information. The matching information can be the content similarity between graph generation elements. The first major language model also generates the aforementioned graph generation element information.

[0075] Step 302: In response to the determination of existence, generate an image corresponding to the preprocessed dialogue based on the image in the dialogue graph data.

[0076] In some embodiments, in response to determining that dialogue graph data that satisfies the target matching conditions exists, the execution entity can generate an image corresponding to the preprocessed dialogue based on the image in the dialogue graph data.

[0077] As an example, the aforementioned execution entity can directly identify the image in the dialogue graph data as the image corresponding to the preprocessed dialogue.

[0078] As another example, firstly, the aforementioned execution entity can combine the image semantic feature information corresponding to the image and the feature semantic feature information corresponding to the graph generation feature enhancement information to obtain combined semantic feature information. Then, based on the combined semantic feature information, an image is generated using a text-to-graph model.

[0079] Step 303: In response to receiving the above graph generation information and not finding any dialogue graph data that meets the above target matching conditions, perform graph generation intent preprocessing on the above target dialogue to obtain a preprocessed dialogue.

[0080] In some embodiments, in response to receiving the above graph generation information and not finding dialogue graph data that meets the above target matching conditions, the target dialogue is preprocessed with graph generation intent to obtain a preprocessed dialogue.

[0081] Step 304: Generate graph generation element information corresponding to the above preprocessed dialogue and prompt information for information enhancement of the above graph generation element information.

[0082] Step 305: Input the above prompt information into the pre-trained first language model to obtain the graph generation element enhancement information corresponding to the above pre-processed dialogue.

[0083] Step 306: Based on the above graph generation feature enhancement information, use the pre-trained text-to-graph model to generate the image corresponding to the preprocessed dialogue.

[0084] Step 307: Display the above image on the dialogue page corresponding to the target dialogue.

[0085] In some embodiments, the specific implementation of steps 304-307 and the resulting technical effects can be found in [reference needed]. Figure 2 Steps 202-205 in the corresponding embodiments will not be repeated here.

[0086] from Figure 3 It can be seen from this that, with Figure 2 Compared to the description of some corresponding embodiments, Figure 3 In some corresponding embodiments, the image display method process 300 can be based on a dialogue graph database, which can quickly and accurately generate images corresponding to the target object.

[0087] Further reference Figure 4 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an image display device, which are similar to... Figure 2 Corresponding to the method embodiments shown, this image display device can be specifically applied to various electronic devices.

[0088] like Figure 4 As shown, an image display device 400 includes: a preprocessing unit 401, a first generation unit 402, an input unit 403, a second generation unit 404, and a display unit 405. The preprocessing unit 401 is configured to, in response to obtaining graph generation information for a target dialogue, perform graph generation intent preprocessing on the target dialogue to obtain a preprocessed dialogue; the first generation unit 402 is configured to generate graph generation element information corresponding to the preprocessed dialogue and prompt information for enhancing the graph generation element information; the input unit 403 is configured to input the prompt information into a pre-trained first large language model to obtain graph generation element enhancement information corresponding to the preprocessed dialogue; the second generation unit 404 is configured to generate an image corresponding to the preprocessed dialogue based on the graph generation element enhancement information and using a pre-trained text-to-graph model; and the display unit 405 is configured to display the image on a dialogue page corresponding to the target dialogue.

[0089] In some optional implementations of some embodiments, the preprocessing unit 401 may be further configured to: extract at least one graph generation intent corresponding to the target dialogue; determine the intent proposal time corresponding to each of the at least one graph generation intents; and add the graph generation intent whose corresponding intent proposal time satisfies the target time condition as the actual graph generation intent to the dialogue text corresponding to the target dialogue to obtain the preprocessed dialogue.

[0090] In some optional implementations of certain embodiments, the first language model further generates: the graph generation element information; and the apparatus 400 further includes: a query unit and a third generation unit (not shown in the figure). The query unit can be configured to: query the dialogue graph database to see if dialogue graph data exists that matches the graph generation element information with the target matching conditions. The third generation unit can be configured to: in response to receiving the graph generation information and not finding dialogue graph data that matches the target matching conditions, perform graph generation intent preprocessing on the target dialogue to obtain a preprocessed dialogue.

[0091] In some optional implementations of certain embodiments, the apparatus 400 further includes a conversion unit and a first determining unit (not shown in the figure). The conversion unit can be configured to: use the first language model to convert the graph generation feature enhancement information into enhancement information in the target language, thereby obtaining converted enhancement information. The first determining unit can be configured to: determine the converted enhancement information as graph generation feature enhancement information.

[0092] In some optional implementations of certain embodiments, the apparatus 400 further includes a second determining unit, a third determining unit, and a fourth generating unit (not shown in the figure). The second determining unit may be configured to: use a pre-trained second language model to determine whether the target dialogue requires graph generation. The third determining unit may be configured to: in response to determining that the target dialogue requires graph generation, use the second language model to determine the graph generation type corresponding to the target dialogue. The fourth generating unit may be configured to: in response to determining that the graph generation type is an open graph generation type, generate graph generation information corresponding to the target dialogue.

[0093] In some optional implementations of certain embodiments, the apparatus 400 further includes: a fourth determining unit, a fifth determining unit, a fifth generating unit, a sixth determining unit, a seventh determining unit, and an eighth determining unit (not shown in the figure). The fourth determining unit can be configured to: use a pre-trained third language model to determine first matching information between the semantic content of the first image corresponding to the image and the graph generation element information, and to determine second matching information between the semantic content of the first image and the dialogue content corresponding to the target dialogue. The fifth determining unit can be configured to: determine whether the image is an image to be displayed based on the first matching information and the second matching information. The fifth generating unit can be configured to: in response to determining that it is not an image to be displayed, generate a regenerated image based on the first matching information and the second matching information, using the first language model and the text-to-graph model. The sixth determining unit can be configured to: use the third language model to determine third matching information between the semantic content of the second image corresponding to the regenerated image and the graph generation element information, and to determine fourth matching information between the semantic content of the second image and the dialogue content corresponding to the target dialogue. The seventh determining unit can be configured to: determine whether the regenerated image is an image to be displayed based on the third matching information and the fourth matching information. The eighth determining unit can be configured to: determine the regenerated image as an image in response to determining that it is an image to be displayed.

[0094] It is understandable that the units described in the image display device 400 are related to the reference. Figure 2 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the image display device 400 and the units contained therein, and will not be repeated here.

[0095] The following is for reference. Figure 5 It illustrates electronic devices suitable for implementing some embodiments of this disclosure (e.g., Figure 1 A schematic diagram of the structure of electronic device 101)500. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0096] like Figure 5As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory 502 or a program loaded from a storage device 508 into a random access memory 503. The random access memory 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, the read-only memory 502, and the random access memory 503 are interconnected via a bus 504. An input / output interface 505 is also connected to the bus 504.

[0097] Typically, the following devices can be connected to the input / output interface 505: input devices 506 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 508 including, for example, magnetic tape, hard disk, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 5 Each box shown can represent a device or multiple devices as needed.

[0098] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a read-only memory 502. When the computer program is executed by the processing device 501, it performs the functions defined above in the methods of some embodiments of this disclosure.

[0099] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0100] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0101] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: in response to obtaining graph generation information for a target dialogue, perform graph generation intent preprocessing on the target dialogue to obtain a preprocessed dialogue; generate prompts that extract graph generation element information corresponding to the preprocessed dialogue and enhance the graph generation element information; input the prompts into a pre-trained first language model to obtain graph generation element enhancement information corresponding to the preprocessed dialogue; generate an image corresponding to the preprocessed dialogue using a pre-trained text-to-graph model based on the graph generation element enhancement information; and display the image on the dialogue page corresponding to the target dialogue.

[0102] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0104] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a preprocessing unit, a first generation unit, an input unit, a second generation unit, and a display unit. The names of these units do not necessarily limit the specific unit; for example, the second generation unit may also be described as "a unit that generates an image corresponding to a preprocessed dialogue based on the aforementioned graph generation feature enhancement information and using a pre-trained text-based graph model."

[0105] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0106] Some embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the image display methods described above.

[0107] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. An image display method, comprising: In response to obtaining graph generation information for the target dialogue, the target dialogue is preprocessed with graph generation intent to obtain a preprocessed dialogue; Generate graph generation element information corresponding to the preprocessed dialogue and prompt information for information enhancement of the graph generation element information; The prompt information is input into the pre-trained first language model to obtain the graph generation element enhancement information corresponding to the preprocessed dialogue; Based on the graph generation feature enhancement information, a pre-trained text-generated graph model is used to generate an image corresponding to the pre-processed dialogue. The image will be displayed on the dialogue page corresponding to the target dialogue.

2. The method according to claim 1, wherein, The step of performing graph generation intent preprocessing on the target dialogue to obtain a preprocessed dialogue includes: Extract at least one graph corresponding to the target dialogue to generate an intent; Determine the intent proposal time corresponding to each of the at least one graph generation intents; The graph generation intent whose corresponding intent proposal time meets the target time condition is taken as the actual graph generation intent and added to the dialogue text corresponding to the target dialogue to obtain the preprocessed dialogue.

3. The method according to claim 1, wherein, The first large language model also generates: graph generation element information; as well as Before performing graph generation intent preprocessing on the target dialogue in response to obtaining graph generation information for the target dialogue to obtain a preprocessed dialogue, the method further includes: Query the dialogue graph database to see if there is dialogue graph data whose matching information with the graph generation element information meets the target matching conditions; In response to the determination of existence, an image corresponding to the preprocessed dialogue is generated based on the image in the dialogue graph data; and In response to obtaining graph generation information for the target dialogue, the target dialogue is subjected to graph generation intent preprocessing to obtain a preprocessed dialogue, including: In response to receiving the graph generation information and not finding any dialogue graph data that meets the target matching conditions, the target dialogue is preprocessed with graph generation intent to obtain a preprocessed dialogue.

4. The method according to claim 1, wherein, After inputting the prompt information into the pre-trained first large language model to obtain the graph generation feature enhancement information corresponding to the preprocessed dialogue, the method further includes: Using the first large language model, the graph generation element enhancement information is converted into enhancement information in the target language to obtain the converted enhancement information; The transformation enhancement information is identified as graph generation feature enhancement information.

5. The method according to claim 1, wherein, Before performing graph generation intent preprocessing on the target dialogue in response to obtaining graph generation information for the target dialogue to obtain a preprocessed dialogue, the method further includes: Using a pre-trained second-largest language model, determine whether the target dialogue has a graph generation requirement; In response to the determination of existence, the graph generation type corresponding to the target dialogue is determined using the second major language model; In response to determining that the graph generation type is an open graph generation type, graph generation information corresponding to the target dialogue is generated.

6. The method according to claim 1, wherein, The prompt information also includes: when the extracted graph generation element information is empty, using the key dialogue information corresponding to the target dialogue as a sub-prompt information of the graph generation element information; and a restriction prompt information regarding whether the image generation quality meets the target quality conditions.

7. The method according to claim 1, wherein, After generating feature enhancement information based on the graph and using a pre-trained text-based graph model to generate an image corresponding to the preprocessed dialogue, the method further includes: Using a pre-trained third language model, a first matching information is determined between the semantic content of the first image corresponding to the image and the graph generation element information, and a second matching information is determined between the semantic content of the first image and the dialogue content corresponding to the target dialogue. Based on the first matching information and the second matching information, determine whether the image is an image to be displayed; In response to determining that it is not an image to be displayed, a regenerated image is generated based on the first matching information and the second matching information, using the first large language model and the text-generated image model; Using the third language model, a third matching information is determined between the semantic content of the second image corresponding to the regenerated image and the graph generation element information, and a fourth matching information is determined between the semantic content of the second image and the dialogue content corresponding to the target dialogue. Based on the third matching information and the fourth matching information, determine whether the regenerated image is the image to be displayed; In response to determining that it is an image to be displayed, the regenerated image is determined as an image.

8. An image display device, comprising: The preprocessing unit is configured to perform graph generation intent preprocessing on the target dialogue in response to obtaining graph generation information for the target dialogue, thereby obtaining a preprocessed dialogue. The first generation unit is configured to generate graph generation element information corresponding to the preprocessed dialogue and prompt information for enhancing the graph generation element information; The input unit is configured to input the prompt information into a pre-trained first large language model to obtain graph generation element enhancement information corresponding to the preprocessed dialogue; The second generation unit is configured to generate feature enhancement information based on the graph and generate an image corresponding to the pre-processed dialogue using a pre-trained text-to-graph model. The display unit is configured to display the image on the dialogue page corresponding to the target dialogue.

9. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

10. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.

11. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.