Image generation method and apparatus, storage medium and program product
By utilizing the prompts from previous dialogue rounds and the input from the current dialogue round in the image generation method, more consistent images are generated, solving the problem of inconsistent images in multi-turn dialogue scenarios and improving the user experience.
Patent Information
- Application Number
- PCT/CN2024/108820
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-02-05
AI Technical Summary
In multi-turn dialogue scenarios, the lack of consistency between the images generated by the dialogue between the user and the intelligent agent leads to a poor user experience.
By obtaining the prompt information of the first image generated in the previous dialogue rounds and combining it with the input content of the current dialogue round, a second prompt information is determined to generate a more consistent image.
It improves the consistency of image generation in multi-turn dialogue scenarios and enhances the user experience.
Smart Images

Figure CN2024108820_05022026_PF_FP_ABST
Abstract
Description
Image generation method and device, storage medium, and program product TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to an image generation method and device, a storage medium and a program product. BACKGROUND
[0002] With the development of artificial intelligence technology, related applications have penetrated into various aspects of our life, such as intelligent question answering, AI (Artificial Intelligence) image generation, etc. AI image generation is a method of generating images by using artificial intelligence technology.
[0003] In the related art, a user can have multi-round dialogues with an intelligent agent in a dialogue interface, and the intelligent agent returns an image based on input content of the user in each dialogue round.
[0004] SUMMARY
[0005] This summary is provided to introduce a selection of concepts, which will be described in greater detail below in the detailed description section. This summary does not necessarily describe all of the important features or functions of the claimed technology, and can not be used to define the scope of the claimed technology.
[0006] According to a first aspect of some embodiments of the present disclosure, an image generation method is provided, comprising:
[0007] In response to receiving input content from a user in a current dialogue round, first prompt information for a first image is obtained, wherein the first image is generated in a historical dialogue round;
[0008] According to the input content and the first prompt information, second prompt information corresponding to the current dialogue round is determined;
[0009] According to the second prompt information, a second image corresponding to the current dialogue round is generated.
[0010] According to a second aspect of some embodiments of the present disclosure, an image generation device is provided, comprising:
[0011] The obtaining module is configured to, in response to receiving input content from a user in a current dialogue round, obtain first prompt information for a first image, wherein the first image is generated in a historical dialogue round;
[0012] The determining module is configured to determine, according to the input content and the first prompt information, second prompt information corresponding to the current dialogue round;
[0013] The generating module is configured to generate a second image corresponding to the current dialogue turn according to the second prompt information.
[0014] According to a third aspect of some embodiments of the present disclosure, there is provided an image generation apparatus, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the image generation method of any of the embodiments described in the present disclosure based on instructions stored in the memory.
[0015] According to a fourth aspect of some embodiments of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, performs the image generation method of any of the embodiments described in the present disclosure.
[0016] According to a fifth aspect of some embodiments of the present disclosure, there is provided a computer program product which, when executed on a computer, causes the computer to implement the image generation method of any of the embodiments.
[0017] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of the exemplary embodiments of the present disclosure with reference to the following drawings. BRIEF DESCRIPTION OF DRAWINGS
[0018] The preferred embodiments of the present disclosure will be described below with reference to the accompanying drawings. The accompanying drawings are used to provide further understanding of the present disclosure, and together with the specific description below, form a part of the description of the present disclosure, and are used to explain the present disclosure. It should be understood that the accompanying drawings described below only relate to some embodiments of the present disclosure, and do not constitute a limitation on the present disclosure. In the drawings:
[0019] FIG. 1 is a flowchart illustrating an image generation method according to some embodiments of the present disclosure;
[0020] FIG. 2 is a flowchart illustrating an image generation method according to some other embodiments of the present disclosure;
[0021] FIG. 3A is a flowchart illustrating an image generation method according to yet some other embodiments of the present disclosure;
[0022] FIG. 3B is a flowchart illustrating an image generation method according to yet some other embodiments of the present disclosure;
[0023] FIG. 4A is a flowchart illustrating an image generation method according to yet some other embodiments of the present disclosure;
[0024] FIG. 4B is a flowchart illustrating an image generation method according to yet some other embodiments of the present disclosure;
[0025] FIG. 5A is a flowchart illustrating an image generation method according to yet some other embodiments of the present disclosure;
[0026] FIG. 5B is a flow diagram illustrating an image generation method according to yet some embodiments of the present disclosure;
[0027] FIG. 5C is a flow diagram illustrating an image generation method according to yet some embodiments of the present disclosure;
[0028] FIG. 6 is a block diagram illustrating an image generation apparatus according to some embodiments of the present disclosure;
[0029] FIG. 7 is a block diagram illustrating an image generation apparatus according to some other embodiments of the present disclosure;
[0030] FIG. 8 illustrates a block diagram of an electronic device according to some embodiments of the present disclosure.
[0031] It should be understood that the dimensions of the various portions shown in the attached drawings are shown for purposes of convenience and do not necessarily correspond to the actual proportions. Identical or similar reference numerals are used to indicate identical or similar components throughout the attached drawings. Thus, once a component is defined in one drawing, it can not be further discussed in subsequent drawings. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The description of the embodiments below is actually only illustrative, and should not be construed as any limitation on the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms, and should not be construed as being limited to the embodiments described herein.
[0033] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions, and numerical values set forth in these embodiments should be interpreted as merely exemplary, and not limiting to the scope of the present disclosure.
[0034] The term "comprise" and variations thereof used in the present disclosure means an open term that includes at least the recited elements / features, but does not exclude other elements / features, i.e., "including but not limited to". In addition, the term "include" and variations thereof used in the present disclosure means an open term that includes at least the recited elements / features, but does not exclude other elements / features, i.e., "including but not limited to". Therefore, include and comprise are synonymous. The term "based on" means "based at least in part on".
[0035] Reference throughout this specification to "one embodiment", "an embodiment", or "embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places in the specification are not necessarily all referring to the same embodiment, although it can. Furthermore, the term "a couple" or "a few" means "one or more" or "two or more".
[0036] It should be noted that the terms "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units. Unless otherwise specified, the terms "first", "second", etc. are not intended to imply a given order or any other manner of given order in time, space, ranking or any other manner.
[0037] It should be noted that the modification of "one" or "multiple" mentioned in the present disclosure is illustrative rather than limiting, and those skilled in the art should understand that unless otherwise specified in the context, it should be understood as "one or more".
[0038] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0039] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes can not be described in some embodiments. In addition, in one or more embodiments, specific features, structures or characteristics can be combined by any suitable means from the present disclosure that will be clear to those skilled in the art.
[0040] In the related art, there can be a context association between multiple rounds of dialogues between a user and an intelligent agent, and the multiple rounds of dialogues are independent of each other when generating images, resulting in poor consistency between images generated by multiple dialogue rounds and poor user experience.
[0041] The present disclosure provides a technical solution that can improve the consistency of image generation in a multi-round dialogue scenario and improve user experience.
[0042] FIG. 1 is a flow diagram illustrating an image generation method according to some embodiments of the present disclosure.
[0043] As shown in FIG. 1, the image generation method comprises: in response to receiving input content from a user in a current dialogue turn, obtaining first prompt information for a first image, wherein the first image is generated in a historical dialogue turn; in step S20, determining second prompt information corresponding to the current dialogue turn according to the input content and the first prompt information; and in step S30, generating a second image corresponding to the current dialogue turn according to the second prompt information. For example, the image generation method is executed by an image generation apparatus.
[0044] A dialogue turn refers to a process in which a user sends input content and receives an image corresponding to the input content. Prompt information is text that guides the generation of an image that the user expects. The historical dialogue turn occurs before the current dialogue turn.
[0045] Taking the generation of an image including a small duck as an example, in the historical dialogue turn, the image generation apparatus receives the input content "draw a small duck" of the user in the historical dialogue turn, and generates a first image including a yellow small duck standing on the river bank for the user. In the current dialogue turn, the image generation apparatus receives the input content "wear sunglasses" of the user in the current dialogue turn, and obtains the first prompt information "a yellow small duck stands on the river bank" for the first image. The image generation apparatus determines the second prompt information "a yellow small duck wearing sunglasses stands on the river bank" according to the first prompt information "a yellow small duck stands on the river bank" and the input content "wear sunglasses" of the user in the current dialogue turn. The image generation apparatus generates a second image including a yellow small duck wearing sunglasses standing on the river bank according to the second prompt information "a yellow small duck wearing sunglasses stands on the river bank".
[0046] In the above embodiment, in the process of generating an image for the user in the current dialogue turn, the second prompt information is generated on the basis of the input content of the user, in combination with the first prompt information of the first image generated in the historical dialogue turn, to guide the generation of the second image, so that the image generated in the current dialogue turn integrates the context intention of the user, thereby improving the consistency of image generation in the multi-turn dialogue scenario and improving the user experience.
[0047] The image generation method in other embodiments of the present disclosure will be described in detail below with reference to FIGS. 2-5C.
[0048] FIG. 2 is a flow diagram illustrating an image generation method according to some other embodiments of the present disclosure. FIG. 2 differs from FIG. 1 in that steps S21, S23 and S25 in FIG. 2 are an implementation of step S21 in FIG. 1. Only the differences between FIG. 2 and FIG. 1 will be described below, and the same parts will not be described again.
[0049] As shown in FIG. 2, in step S21, image features of the first image are extracted to obtain first image features.
[0050] In step S23, text description information of the first image is generated according to the first image features.
[0051] In step S25, the second prompt information is determined according to the input content, the first prompt information and the text description information of the first image.
[0052] For example, the image features of the first image can be extracted by using a feature extraction model to obtain the first image features. The feature extraction model can be a convolutional neural network or a recurrent neural network, and the present disclosure does not limit the same.
[0053] Taking the first image features as feature vectors for example, the text generator can be used to map the feature vectors of the first image to a text sequence, so as to obtain the text description information of the first image. The text generator can be implemented by using a recurrent neural network, a long short-term memory network or attention.
[0054] Taking the generation of the image including the duck as an example, the image generation device extracts features of the first image including a yellow duck standing on the river bank to obtain first image features, and further generates the text description information of the first image according to the first image features, that is, “There is a small duck with yellow and brown feathers standing on the grassland on the river bank, and the river water is clear.” Further, the image generation device determines the second prompt information “There is a small duck with yellow and brown feathers standing on the grassland on the river bank, and the river water is clear” according to the input content “Wear sunglasses” in the current dialogue turn, the first prompt information “A small yellow duck stands on the river bank” and the text description information of the first image “There is a small duck with yellow and brown feathers standing on the grassland on the river bank, and the river water is clear.” The image generation device generates the second image including a small duck with yellow and brown feathers standing on the grassland on the river bank wearing sunglasses according to the second prompt information “A small duck with yellow and brown feathers stands on the grassland on the river bank wearing sunglasses”.
[0055] In this embodiment, in the process of determining the second prompt information for generating the second image, on the basis of the input content, not only the first prompt information of the first image generated in the historical dialogue turn is integrated, but also the text description information of the first image generated based on the image features of the first image is integrated, so that more context intention information is provided for the generation of the second image in the current dialogue turn, thereby further improving the consistency of image generation in the multi-turn dialogue scene and further improving the user experience.
[0056] FIG. 3A is a flowchart illustrating an image generation method according to yet some embodiments of the present disclosure. FIG. 3A differs from FIG. 1 in that steps S22, S24 and S26 of FIG. 3A are an implementation of step S20 in FIG. 1. Hereinafter, only the differences between FIG. 3A and FIG. 1 will be described, and the same will not be described again.
[0057] As shown in FIG. 3A, taking an example that the input content includes content of multiple modalities, in step S22, the content of multiple modalities is fused to obtain multi-modal fusion information. In step S24, according to the multi-modal fusion information, content in a text format corresponding to the content of multiple modalities is generated. In step S26, according to the content in a text format corresponding to the content of multiple modalities and the first prompt information, the second prompt information is determined.
[0058] The content of multiple modalities includes but is not limited to images, audios, texts, etc., which are not limited by the present disclosure. For example, the user can input at least one of images, audios and texts in the current dialogue turn to express the user's intention more richly.
[0059] Taking an example that an image including a duckling is generated, in the historical dialogue turn, the image generation apparatus receives that the input content of the user in the current dialogue turn includes input text "wear sunglasses" and input image including a duckling. The image generation apparatus fuses the input text "wear sunglasses" in the text modality and the input image in the image modality to obtain multi-modal fusion information, and further generates content in a text format corresponding to the input content based on the multi-modal fusion information.
[0060] Taking an example that the input image including a duckling has a duckling with yellow and brown feathers, the content in a text format corresponding to the input content can be "a duckling with yellow and brown feathers wearing sunglasses". The image generation apparatus can further determine the second prompt information "a duckling with yellow and brown feathers wearing sunglasses standing by the river" according to the content in a text format corresponding to the input content "a duckling with yellow and brown feathers wearing sunglasses" and the first prompt information "a duckling with yellow standing by the river".
[0061] The image generation apparatus generates a second image including a duckling with yellow and brown feathers wearing sunglasses standing by the river according to the second prompt information "a duckling with yellow and brown feathers wearing sunglasses standing by the river".
[0062] In this embodiment, in response to the input content of the user including content of multiple modalities, through multi-modal fusion, the user's intention can be better learned, the understanding ability of the user's intention can be improved, and thus the accuracy of generating the second prompt information can be improved, thereby further improving the consistency of image generation in the multi-round dialogue scene and further improving the user experience.
[0063] The multi-modal fusion is, for example, feature-level fusion based on multiple modalities, that is, semantic features are extracted from the content of different modalities, and then the semantic features corresponding to the information of different modalities are fused to obtain fused features as multi-modal fusion information. Extracting semantic features may, for example, include representing the content of each modality as a feature vector, and fusing semantic features includes merging the feature vectors of the content of multiple modalities in the feature space.
[0064] The multi-modal fusion may, for example, also be model-level fusion based on multiple modalities, that is, the content of each modality is processed by an independent model, and then the outputs of the models corresponding to the content of multiple modalities are combined to obtain multi-modal fusion information.
[0065] The multi-modal fusion may, for example, also be other types of fusion, which are not limited by the present disclosure.
[0066] FIG. 3B is a flow diagram illustrating an image generation method according to yet some embodiments of the present disclosure. FIG. 3B differs from FIG. 1 in that steps S22' and S24' of FIG. 3B are an implementation of step S20 in FIG. 1. Only the differences between FIG. 3B and FIG. 1 will be described below, and the same parts will not be described again.
[0067] As shown in FIG. 3B, taking the input content including content of a non-text format as an example, in step S22', the content of a non-text format in the input content is converted into content of a text format.
[0068] In step S24', the second prompt information is determined according to the text format content obtained through the conversion and the first prompt information.
[0069] In the case where the input content also includes input text, the second prompt information is determined according to the text format content obtained through the conversion, the input text included in the input content, and the first prompt information
[0070] For example, the input content can include the input text "wear sunglasses" in the foregoing embodiment and an input image including a small duck, and the image generation apparatus can convert the input image in an image format into content in a text format. Taking an example of the input image including a small duck with yellow and brown feathers, the content in the text format corresponding to the input image can be "a small duck with yellow and brown feathers". The image generation apparatus can determine the second prompt information "a small duck with yellow and brown feathers wearing sunglasses standing by the river" according to the content in the text format corresponding to the input image "a small duck with yellow and brown feathers", the input text "wear sunglasses", and the first prompt information "a small duck with yellow standing by the river".
[0071] In this embodiment, the non-text format content can be directly and simply converted into content in a text format, and the efficiency of image generation is improved under the premise of ensuring consistency of image generation in a multi-round dialogue scenario.
[0072] In some embodiments, the non-text format includes an image format, and converting the non-text format content in the input content into content in a text format includes the following steps.
[0073] First, feature extraction is performed on the content in the image format in the input content to obtain second image features.
[0074] Then, according to the second image features, a text description information corresponding to the content in the image format in the target content is generated.
[0075] Taking an example of the input content including the input image including a small duck in the foregoing embodiment, the image generation apparatus performs feature extraction on the input image to obtain second image features, and according to the second image features, a text description information corresponding to the input image can be generated. Taking an example of the input image including a small duck with yellow and brown feathers, the text description information corresponding to the input image can be "a small duck with yellow and brown feathers wearing sunglasses".
[0076] In some embodiments, the non-text format includes content in an audio format, and converting the non-text format content in the target content into content in a text format can be achieved in the following manner.
[0077] Voice recognition is performed on the content in the audio format in the target content to obtain text description information corresponding to the content in the audio format.
[0078] The input content includes input audio, and the content of the input audio is "the duck stands on the grassland by the river". The input audio can be directly subjected to speech recognition to obtain the text description information corresponding to the input audio, i.e., "the duck stands on the grassland by the river".
[0079] FIG. 4A is a flowchart illustrating an image generation method according to yet some embodiments of the present disclosure. FIG. 4 is different from FIG. 1 in that FIG. 4A illustrates other steps S40-S60 of the image generation method in some embodiments of the present disclosure. Hereinafter, only the differences between FIG. 4A and FIG. 1 will be described, and the same parts will not be described again.
[0080] As shown in FIG. 4A, in step S40, the second prompt information is sent to the user.
[0081] In step S50, in response to receiving adjustment information from the user for the second prompt information, the second prompt information is adjusted according to the adjustment information for the second prompt information to obtain adjusted second prompt information.
[0082] In step S60, the adjusted second image is generated according to the adjusted second prompt information.
[0083] Taking the second prompt information "a yellow duck wearing sunglasses stands by the river" in the foregoing embodiments as an example, the image generation apparatus sends the second prompt information "a yellow duck wearing sunglasses stands by the river" to the user, the adjustment information for the second prompt information from the user is "change the yellow duck to a brown duck", and the image generation apparatus receives the adjustment information for the second prompt information and adjusts the second prompt information to obtain adjusted second prompt information "a brown duck wearing sunglasses stands by the river". Then, the adjusted second image is generated according to the adjusted second prompt information "a brown duck wearing sunglasses stands by the river". The duck in the adjusted second image is brown.
[0084] In this embodiment, by sending the second prompt information to the user, the user can provide adjustment information according to the user's intention, so that the second prompt information can be adjusted according to the user's adjustment information, and the flexibility and convenience of image generation are improved on the premise of ensuring the consistency of image generation.
[0085] FIG. 4B is a flowchart illustrating an image generation method according to yet some embodiments of the present disclosure. FIG. 4B is different from FIG. 1 in that FIG. 4B illustrates other steps S70-S90 of the image generation method in some embodiments of the present disclosure. Hereinafter, only the differences between FIG. 4B and FIG. 1 will be described, and the same parts will not be described again.
[0086] As shown in FIG. 4B, in step S70, adjustment information for one or more elements in the second image is received from the user.
[0087] In step S80, the second prompt information is adjusted according to the adjustment information for one or more elements in the second image and the second image, to obtain adjusted second prompt information.
[0088] In step S90, the one or more elements in the second image are adjusted according to the adjusted second prompt information, to obtain an adjusted second image.
[0089] The adjustment information for the second image can be adjustment information for one or more elements in the second image. Taking the second image of a yellow duck wearing sunglasses standing on the riverbank as an example, after receiving the second image, the user sends adjustment information “change the duck’s feathers in the image to brown” for the element of the duck to the image generation device. The image generation device adjusts the second prompt information “a yellow duck wearing sunglasses stands on the riverbank” according to the adjustment information “change the duck’s feathers in the image to brown” and the second image, to obtain adjusted second prompt information “a brown duck wearing sunglasses stands on the riverbank”, and further adjusts the duck’s feathers in the second image to brown according to the adjusted second prompt information, to obtain an adjusted second image.
[0090] The adjustment information for the second image can also be adjustment information for the image style of the second image. Taking the second image of a yellow duck wearing sunglasses standing on the riverbank as an example, after receiving the second image, the user sends adjustment information “change the duck to an anime-style duck” for the element of the duck to the image generation device. The image generation device adjusts the second prompt information “a yellow duck wearing sunglasses stands on the riverbank” according to the adjustment information “change the duck to an anime-style duck” and the second image, to obtain adjusted second prompt information “a yellow duck wearing sunglasses stands on the riverbank, the duck is in an anime style”, and further changes the duck in the second image to an anime style according to the adjusted second prompt information, to obtain an adjusted second image.
[0091] The adjustment information for the second image can also be adjustment information for other aspects of the second image, including but not limited to expanding the image, adjusting the background, etc. The present disclosure does not limit this.
[0092] In this embodiment, based on the adjustment information for the second image input by the user, the adjustment of the second prompt information can be implemented, so that based on the adjusted second prompt information, the user can adjust the second image, the flexibility and convenience of image generation can be improved under the premise of ensuring the consistency of image generation.
[0093] FIG. 5A is a flow diagram illustrating an image generation method according to yet some embodiments of the present disclosure. FIG. 5A differs from FIG. 1 in that step S31 of FIG. 5A is an implementation of step S30 in FIG. 1. In the following, only the differences between FIG. 5A and FIG. 1 will be described, and the same parts will not be described again.
[0094] As shown in FIG. 5A, in step S31, the second image is generated by using a graph-to-image model according to the second prompt information and the first image.
[0095] For example, the second prompt information and the first image can be input into the graph-to-image model to obtain the second image. The graph-to-image model can process the second prompt information and the first image to generate the second image.
[0096] In this embodiment, the second prompt information and the first image are processed together by the graph-to-image model, which can fuse more feature information of the first image, thereby further improving the consistency of image generation in the multi-round dialogue scenario.
[0097] FIG. 5B is a flow diagram illustrating an image generation method according to yet some embodiments of the present disclosure. FIG. 5B differs from FIG. 1 in that step S31' of FIG. 5B is an implementation of step S30 in FIG. 1. In the following, only the differences between FIG. 5B and FIG. 1 will be described, and the same parts will not be described again.
[0098] As shown in FIG. 5B, in step S31', the second image is generated by using a text-to-graph model according to the second prompt information.
[0099] In this embodiment, the second prompt information is processed by the text-to-graph model to generate the second image, which can improve the efficiency of image generation under the premise of ensuring the consistency of image generation in the multi-round dialogue scenario.
[0100] FIG. 5C is a flow diagram illustrating an image generation method according to yet some embodiments of the present disclosure. FIG. 5C differs from FIG. 1 in that step S11 of FIG. 5C is an implementation of step S10 in FIG. 1. In the following, only the differences between FIG. 5C and FIG. 1 will be described, and the same parts will not be described again.
[0101] As shown in FIG. 5C, in step S11, in response to receiving input content from a user in a current dialogue turn, the first prompt information for generating the first image to which the input content is directed is obtained, wherein the first image is generated in a historical dialogue turn.
[0102] In this embodiment, the first image is an image corresponding to the input content of the user in the current dialogue turn, so that the second image generated in the current dialogue turn is more targeted, improving the consistency and accuracy of image generation.
[0103] In some embodiments, the historical dialogue turn in any of the preceding embodiments includes at least one historical dialogue turn adjacent to the current dialogue turn. For example, the historical dialogue turn includes the previous dialogue turn of the current dialogue turn. The number of historical dialogue turns can be determined according to actual needs.
[0104] In some embodiments, the second prompt information in any of the preceding embodiments is obtained through natural language processing. For example, the input content and the first prompt information are subjected to natural language processing to obtain user intent information, and then the second prompt information is obtained by rewriting and / or expanding the input content and the first prompt information based on the user intent information using a model.
[0105] In some embodiments, the generation of the second prompt information and the generation of the second image in any of the preceding embodiments can be implemented by means of a machine learning model. The machine learning model may, for example, include a large language model (LLM), a natural language processing model, etc.
[0106] In some embodiments, the image generation method can be executed by an image generation device, which can or can not include an agent. In the case where the image generation device does not include an agent, the user can interact with the image generation device through an agent.
[0107] The above is the image generation method provided by some embodiments of the present disclosure. The content publishing device in some embodiments of the present disclosure will be described below in conjunction with FIG. 6.
[0108] FIG. 6 is a block diagram of an image generation device according to some embodiments of the present disclosure.
[0109] As shown in FIG. 6, the image generation device 6 includes an obtaining module 61, a determining module 62, and a generating module 63.
[0110] The obtaining module 61 is configured to, in response to receiving input content from a user in a current dialogue turn, obtain first prompt information for a first image, wherein the first image is generated in a historical dialogue turn; the determining module 62 is configured to determine second prompt information corresponding to the current dialogue turn according to the input content and the first prompt information; and the generating module 63 is configured to generate a second image corresponding to the current dialogue turn according to the second prompt information.
[0111] The image generation apparatus 6 can be configured to perform steps S10-S30 of FIG. 1. In some embodiments, the image generation apparatus 6 can also perform any of the steps shown in FIGS. 2-5C.
[0112] It should be noted that the above-mentioned modules are only logical modules according to the specific functions implemented by them, and are not intended to limit the specific implementation manners, for example, they can be implemented in software, hardware or a combination of software and hardware. In actual implementation, the above-mentioned modules can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, the above-mentioned modules are indicated by dashed lines in the drawings, indicating that these modules can not actually exist, and the operations / functions implemented by them can be implemented by the processing circuit itself.
[0113] FIG. 7 is a block diagram of an image generation apparatus according to some embodiments of the present disclosure.
[0114] As shown in FIG. 7, the image generation apparatus 7 includes a memory 71 and a processor 72 coupled to the memory 71, the processor 72 being configured to perform the image generation method of any of the preceding embodiments based on instructions stored in the memory 71.
[0115] The memory 71 is configured to store one or more computer-readable instructions. The memory 71 can include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), flash memory. The memory 71 may, for example, store an operating system, an application program, a boot loader, a database, and other programs, and can also store various application programs and various data.
[0116] The processor 72 is configured to run the computer-readable instructions to implement the image generation method of any of the preceding embodiments. The specific implementation of each step of the image generation method can be referred to the above-mentioned embodiments, and the repeated parts will not be described here.
[0117] The processor 72 and the memory 71 can communicate with each other directly or indirectly. For example, the processor 72 and the memory 71 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 72 and the memory 71 can also communicate with each other through a system bus, without limitation.
[0118] It should be noted that the components of the image generation apparatus 7 shown in FIG. 7 are exemplary and not limiting, and the image generation apparatus 7 can also have other components according to actual application needs. The processor 72 can communicate with other components of the image generation apparatus 7 to perform desired functions.
[0119] The image generation apparatus can be implemented in software, firmware, and / or hardware, and can be integrated in an electronic device in which a related application is installed.
[0120] The above is the image generation apparatus in some embodiments of the present disclosure.
[0121] FIG. 8 shows a block diagram of an electronic device according to some embodiments of the present disclosure.
[0122] The electronic device 8 shown in FIG. 8 can be a computer system having a dedicated hardware structure, and can perform corresponding functions when a related application is installed.
[0123] The electronic device includes, but is not limited to, a mobile terminal such as a smartphone, a notebook computer, a personal digital assistant (PDA), a tablet personal computer (Tablet PC), a PMP (portable multimedia player), a vehicle terminal (e.g., a car navigation terminal), a wearable device, and the like, and a stationary terminal such as a digital television, a desktop computer, and the like.
[0124] As shown in FIG. 8, a central processing unit (CPU) 81 performs various processes according to a program stored in a read-only memory (ROM) 82 or a program loaded from a storage portion 88 to a random access memory (RAM) 83. In the RAM 83, data required when the CPU 81 performs various processes and the like is stored as needed. The central processing unit is merely exemplary, and can be other types of processors such as various processors described above. The ROM 82, the RAM 83, and the storage portion 88 can be various forms of computer readable storage media. It should be noted that although the ROM 82, the RAM 83, and the storage portion 88 are shown separately in FIG. 8, one or more of them can be combined or located in the same or different memory or storage module.
[0125] The CPU 81, the ROM 82, and the RAM 83 are connected to each other via the bus 84. The input / output interface 85 is also connected to the bus 84.
[0126] The following components are connected to the input / output interface 85: an input portion 86, such as a touch panel, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, and the like; an output portion 87 including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage portion 88 including a hard disk, a magnetic tape, and the like; and a communication portion 89 including a network interface card, such as a LAN card, a modem, and the like. The communication portion 89 allows communication processing to be performed via a network, such as the Internet. It is easily understood that, although the respective devices or modules in the electronic device 8 are shown in FIG. 8 to communicate through the bus 84, they can also communicate through a network or other means, where the network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0127] A drive 810 is also connected to the input / output interface 85 as necessary. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, is attached to the drive 810 as necessary, so that a computer program read therefrom is installed in the storage portion 88 as necessary.
[0128] In the case where the above series of processes are realized by software, the program constituting the software can be installed from a network or a storage medium 811 or the like.
[0129] According to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when run on a computer, causes the computer to implement the image generation method described in any of the preceding embodiments. The computer program product includes a computer program carried on a computer-readable medium, which contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication portion 89, or installed from the storage portion 88, or installed from the ROM 82. When the computer program is executed by the CPU 81, the image generation method of the embodiment of the present disclosure is executed.
[0130] Note that, in the context of the present disclosure, the computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0131] The computer readable medium can be a computer readable storage medium or a computer readable signal medium, or any combination thereof.
[0132] The computer readable storage medium includes, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this disclosure, a computer readable storage medium can be any tangible medium that contains or stores a program used by an instruction execution system, apparatus, or device. The program stored by the computer readable storage medium is executed by a processor of the instruction execution system, apparatus, or device to implement the image generation method described in any of the preceding embodiments.
[0133] The computer readable signal medium can include a data signal traveling in a baseband or a carrier wave traveling in a baseband, in which the computer readable program code is contained. This data signal can take many forms, including, but not limited to, electro-magnetic, optical, or any suitable combination of the foregoing. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer readable medium can be transmitted using any suitable medium, including, but not limited to, wire, cable, RF, etc., or any suitable combination of the foregoing.
[0134] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled into the electronic device.
[0135] In some embodiments, a computer program product is also provided, which, when executed on a computer, causes the computer to implement the image generation method described in any of the preceding embodiments.
[0136] In some embodiments, a computer program is also provided, which includes instructions, which, when executed by a processor, cause the processor to perform the image generation method of any of the preceding embodiments. For example, the instructions can be embodied in a computer program code.
[0137] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0138] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0139] The functions described above can be implemented in at least part by one or more hardware logic components. For example, and without limitation, illustrative hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0140] While certain aspects of the present disclosure have been described with reference to one or more particular embodiments thereof, those skilled in the art will understand that many alternative embodiments can be made therefrom. In general, embodiments of the present disclosure are applicable to any suitable electronic device, system, or architecture. In addition, unless otherwise indicated, the functions performed by the various components described herein can be implemented using electronic components, software, firmware, or any suitable combination thereof. In addition, it will be understood that various hardware and software components, or modules, of the embodiments described herein can be combined, divided, recombined, or otherwise rearranged, unless otherwise indicated.
Claims
1. An image generation method, comprising: obtaining first prompt information for a first image in response to receiving input content from a user in a current dialogue turn, wherein the first image is generated in a historical dialogue turn; determining second prompt information corresponding to the current dialogue turn according to the input content and the first prompt information; generating a second image corresponding to the current dialogue turn according to the second prompt information.
2. The image generation method of claim 1, wherein, The determining of the second prompt information corresponding to the current dialogue turn according to the input content and the first prompt information comprises: extracting image features of the first image to obtain first image features; generating text description information of the first image according to the first image features; determining the second prompt information according to the input content, the first prompt information and the text description information of the first image.
3. The image generation method according to claim 1 or 2, wherein The input content comprises content of multiple modalities, and the determining of the second prompt information corresponding to the current dialogue turn according to the input content and the first prompt information comprises: performing multi-modal fusion on the content of multiple modalities to obtain multi-modal fusion information; generating content in a text format corresponding to the content of multiple modalities according to the multi-modal fusion information; determining the second prompt information according to the content in a text format corresponding to the content of multiple modalities and the first prompt information.
4. The image generation method of claim 1 or 2, further comprising: sending the second prompt information to the user; in response to receiving adjustment information from the user for the second prompt information, adjusting the second prompt information according to the adjustment information for the second prompt information to obtain adjusted second prompt information; generating an adjusted second image according to the adjusted second prompt information.
5. The image generation method according to claim 1 or 2, wherein The input content comprises non-text format content, and the determining of the second prompt information corresponding to the current dialogue turn according to the input content and the first prompt information comprises: converting the non-text format content in the input content into text format content; determining the second prompt information according to the text format content obtained through the conversion and the first prompt information.
6. The image generation method of claim 5, wherein, The non-text format comprises image format, and the conversion of the non-text format content in the input content into text format content comprises: performing feature extraction on the image format content in the input content to obtain second image features; generating text description information corresponding to the image format content in the target content according to the second image features.
7. The image generation method of claim 5, wherein, The non-text format comprises audio format content, and the conversion of the non-text format content in the target content into text format content comprises: performing speech recognition on the audio format content in the target content to obtain text description information corresponding to the audio format content.
8. The image generation method according to claim 1 or 2, wherein The generating of the second image corresponding to the current dialogue turn according to the second prompt information comprises: generating the second image according to the second prompt information and the first image by using a graph-to-graph model; or generating the second image according to the second prompt information by using a text-to-graph model. 9.The image generation method of claim 1 or 2, further comprising: receiving adjustment information from the user for the second image; adjusting the second prompt information according to the adjustment information for the second image and the second image, to obtain adjusted second prompt information; adjusting the second image according to the adjusted second prompt information, to obtain an adjusted second image.
10. The image generation method according to claim 1 or 2, wherein The obtaining the first prompt information for the first image comprises: obtaining the first prompt information for generating the first image to which the input content is directed.
11. The image generation method according to claim 1 or 2, wherein The historical dialogue turns comprise at least one historical dialogue turn adjacent to the current dialogue turn.
12. The image generation method of claim 1 or 2, wherein, The second prompt information is obtained through natural language processing. 13.An image generation apparatus, comprising: an obtaining module configured to, in response to receiving input content from a user in a current dialogue turn, obtain first prompt information for a first image, wherein the first image is generated in a historical dialogue turn; a determining module configured to determine, according to the input content and the first prompt information, second prompt information corresponding to the current dialogue turn; a generating module configured to generate, according to the second prompt information, a second image corresponding to the current dialogue turn. 14.An image generation apparatus, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute an image generation method according to any one of claims 1 to 12 based on instructions stored in the memory. 15.A computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the image generation method of any one of claims 1 to 12. 16.A computer program product, which, when executed on a computer, causes the computer to implement the image generation method of any one of claims 1 to 12.
Citation Information
Patent Citations
Image generation method and terminal device
CN110136216A
Image generation method, electronic equipment and computer readable storage medium
CN117689749A
Conversation method based on visual image, electronic equipment and storage medium
CN117827047A
Visual storyline generation from text story
US20200257763A1