Image batch generation method and device, electronic equipment, storage medium and product
Through semantic understanding of user image generation instructions, multiple different prompt information are generated and machine learning models are used to solve the problem of inefficient batch generation of images in the prior art, and efficient batch image generation is achieved.
Patent Information
- Application Number
- CN202510555583.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, users need to input different instructions multiple times to achieve batch generation of images, resulting in inefficient creation.
By obtaining user's image generation instructions, semantic understanding is carried out, the intention to batch generate images, and multiple different prompt information are generated. The machine learning model is used to generate images corresponding to each prompt information, and batch image generation based on a single instruction is realized.
It improves the efficiency of users to generate images in batches, simplifies user operations, and realizes the generation of multiple different images based on a single instruction.
Smart Images

Figure CN120374779A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a method, apparatus, electronic device, storage medium, and product for batch generating images. Background Art
[0002] With the development of artificial intelligence technologies, users can efficiently perform various types of creations using computer devices. For example, a user can issue an instruction to describe an image to be generated. After processing the user's instruction, a machine learning model will return an image created by the machine to the user. Summary of the Invention
[0003] According to some embodiments of the present disclosure, there is provided a method for batch generating images, including: obtaining an image generation instruction input by a user, where the image generation instruction is associated with text; determining, according to a semantic understanding result of the text, whether the image generation instruction includes an intention to batch generate images; in response to the image generation instruction including the intention to batch generate images, generating a plurality of different prompt messages based on the text; using a machine learning model to process each prompt message to generate an image corresponding to each prompt message; and displaying the plurality of generated images.
[0004] According to some other embodiments of the present disclosure, there is provided an apparatus for batch generating images, including: an obtaining module configured to obtain an image generation instruction input by a user, where the image generation instruction is associated with text; a determining module configured to determine, according to a semantic understanding result of the text, whether the image generation instruction includes an intention to batch generate images; a prompt generation module configured to, in response to the image generation instruction including the intention to batch generate images, generate a plurality of different prompt messages based on the text; an image generation module configured to use a machine learning model to process each prompt message to generate an image corresponding to each prompt message; and a display module configured to display the plurality of generated images.
[0005] According to some embodiments of the present disclosure, there is provided an electronic device, including: a memory; and a processor coupled to the memory, where the processor is configured to execute the method for batch generating images according to any one of the embodiments of the present disclosure based on instructions stored in the memory.
[0006] According to some embodiments of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, where the program, when executed by a processor, executes the method for batch generating images according to any one of the embodiments of the present disclosure.
[0007] According to some embodiments of the present disclosure, there is provided a computer program product, which, when running on a computer, causes the computer to implement the method for batch generating images according to any one of the embodiments of the present disclosure.
[0008] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Embodiments of the present disclosure will be described below with reference to the accompanying drawings. It should be understood that the drawings in the following description only relate to some embodiments of the present disclosure and do not constitute a limitation on the present disclosure. In the drawings:
[0010] Figure 1 A flowchart showing a method for batch generating images according to some embodiments of the present disclosure is shown.
[0011] Figure 2 A flowchart showing a method for generating prompt information according to some embodiments of the present disclosure is shown.
[0012] Figure 3 A comparison diagram showing the batch generation effect of the images of the present disclosure and the image generation effect in the related art is shown.
[0013] Figure 4 A flowchart showing a method for determining the number of text units according to some embodiments of the present disclosure is shown.
[0014] Figure 5 Another comparison diagram showing the batch generation effect of the images of the present disclosure and the image generation effect in the related art is shown.
[0015] Figure 6 A flowchart showing a method for batch editing images according to some embodiments of the present disclosure is shown.
[0016] Figure 7 A schematic diagram of an interaction interface according to some embodiments of the present disclosure is shown.
[0017] Figure 8 A flowchart showing a method for generating storyboard images according to some embodiments of the present disclosure is shown.
[0018] Figure 9 A schematic diagram of an interaction interface according to some other embodiments of the present disclosure is shown.
[0019] Figure 10 A schematic diagram of the structure of a device for batch generating images according to some embodiments of the present disclosure is shown.
[0020] Figure 11 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0021] Figure 12 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0023] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard. Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments should be construed as merely exemplary and do not limit the scope of the present disclosure.
[0024] The term "including" and its variants used in the present disclosure mean an open term that includes at least the subsequent elements / features but does not exclude other elements / features, that is, "including but not limited to". The term "based on" means "at least partially based on".
[0025] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependent relationships. Unless otherwise specified, the concepts such as "first" and "second" are not intended to imply that the objects so described must be in a given order in terms of time, space, ranking, or any other way.
[0026] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0027] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and do not limit the scope of these messages or information.
[0028] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0029] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. However, the present disclosure is not limited to these specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner that will be apparent to those of ordinary skill in the art from the present disclosure.
[0030] After research, it is found that image generation based on artificial intelligence technology often has the following application modes.
[0031] One mode is that after receiving an instruction sent by a user, a single image created based on the instruction is returned to the user.
[0032] Another mode is that after receiving an instruction sent by a user, multiple images created based on the instruction are returned to the user. However, through more in-depth analysis, each of these multiple images is generated based on the entire content of the instruction. For example, when the user instructs "generate five expressions", the model may return several generated images to the user, and each image contains 5 expressions. That is, although this mode can provide multiple images for the user, these multiple images are generated by making the model generate images multiple times with the same prompt information in consideration of the randomness of the model when generating images, to make up for the lack of generation quality of some images and provide the user with more choices, facilitating the user to select the most satisfactory one after comparison.
[0033] As the usage requirements and usage scenarios of artificial intelligence-generated images become more and more diverse, some users have a need for batch image generation. That is, after sending an instruction, the user expects to obtain images created based on different requirements in the instruction. However, as described above, if the user only sends one instruction, the entire instruction will be used as a generation instruction to generate one image or multiple images based on the same instruction. If the user needs to generate images in batches, different instructions need to be input multiple times. This will seriously reduce the user's creation efficiency.
[0034] To improve the efficiency of image generation, the present disclosure provides a method for batch image generation, which generates multiple different prompt information based on a single image generation instruction of a user to achieve batch image generation based on a single user instruction. The following refers to Figure 1 Describe the embodiments of the method for batch image generation.
[0035] Figure 1 shows a schematic flowchart of a method for batch image generation according to some embodiments of the present disclosure. As Figure 1 shown, the method for batch image generation in this embodiment includes steps S11 to S15.
[0036] In step S11, an image generation instruction input by the user is obtained, and the image generation instruction is associated with text.
[0037] The user can input an image generation instruction through the interaction interface of a web page or an application. This instruction can be text input by the user through an input box, voice input through a voice input control, or a shortcut instruction provided in the interface.
[0038] The text associated with the image generation instruction can be the image generation instruction itself. For example, if the user sends the text "Generate an image: sea in winter, wind in summer" to an agent for image generation, then "Generate an image: sea in winter, wind in summer" serves as both the image generation instruction and the text associated with the image generation instruction.
[0039] The text associated with the image generation instruction can also be text including the image generation instruction. For example, "Generate an image" can be the image generation instruction, and "Generate an image: sea in winter, wind in summer" can be the text associated with the image generation instruction.
[0040] The text associated with the image generation instruction can also be text in the same message as the image generation instruction. For example, "Generate an image" can be the image generation instruction, and "sea in winter, wind in summer" can be the text associated with the image generation instruction.
[0041] The text associated with the image generation instruction can also be text referenced by the image generation instruction. For example, if the user instructs the agent to create a piece of text 1, then the user can reference this text 1 and send a message "Generate an image based on this text" to the agent, and this message serves as the image generation instruction. Or, the user can directly send the message "Generate an image based on this text" as the image generation instruction, and according to the context of the conversation, it can be determined that text 1 is associated with this image generation instruction.
[0042] The text associated with the image generation instruction can be from the user or from an agent that interacts with the user. However, in either case, when the user has a need for batch creation, there is no need for the user to split the text associated with the image generation instruction by themselves.
[0043] In step S12, according to the semantic understanding result of the text, it is determined whether the image generation instruction includes the intention of batch generating images.
[0044] Since the text associated with the image generation instruction is freely input by the user in natural language or may be generated by the agent according to the user's previous text generation instruction, it is necessary to perform semantic understanding on the text to determine whether the user hopes to batch generate images.
[0045] In step S13, in response to the image generation instruction including the intention of batch generating images, based on the text, multiple different prompt messages are generated.
[0046] Each prompt message can be generated based on part or all of the content of the text. That is, in the process of generating each prompt message, the target content (i.e., part or all of the content) can be extracted from the text first. On this basis, the extracted target content can also be modified, such as adding some indication information or descriptive information related to the target content, etc.
[0047] Through the above processing, each prompt message can reflect the user's generation requirements for multiple different images carried by the image generation instruction and the text.
[0048] In step S14, a machine learning model is used to process each prompt message to generate an image corresponding to each prompt message.
[0049] The machine learning model can be a generative model. For example, a model that generates images based on text. The model can include large language models, foundation models, etc., to generate images matching the prompt messages by using the ability to understand natural language.
[0050] The machine learning model can be trained using training images and description information of the images. For example, the following process is iteratively executed until the convergence condition is reached: the description information of the images is input into the machine learning model, and the parameters of the machine learning model are adjusted according to the images generated by the machine learning model and the training images.
[0051] In step S15, the multiple generated images are displayed.
[0052] In addition to displaying the multiple generated images, the description information of the generated images can also be displayed. The description information can be the overall description information of the multiple generated images or the description information of each generated image separately.
[0053] In the above embodiment, by performing semantic understanding on the text associated with the image generation instruction, it is determined whether the instruction includes the intention of batch generating images. In the scenario of batch generating images, different prompt messages are generated based on the text to batch generate multiple different images based on a single user instruction. Thus, the efficiency of the user's batch image generation can be improved.
[0054] The text associated with the image generation instruction can indicate the intention of batch generating images in various ways. For example, the text includes an explicit indication of generating multiple images, or the text includes semantically independent description information.
[0055] The following describes these two scenarios exemplarily.
[0056] In some embodiments, determining whether an image generation instruction includes the intention of batch generating images based on the semantic understanding result of the text includes: performing semantic understanding on the text to determine whether the text includes an indication of the number of images to be generated; in response to the text including an indication of the number of images to be generated and the number of images to be generated being greater than 1, determining that the image generation instruction includes the intention of batch generating images.
[0057] The indication of the number of images to be generated can be an explicit quantity indication, such as "5 pictures", or an indication that does not have an explicit quantity but can be determined to be multiple, such as "multiple pictures", etc.
[0058] Some users, in the indication, do not separately describe the generation requirements for each of the multiple pictures, but only give an overall indication of generating multiple pictures. For example, "Generate 5 meme-style emojis". Although the user does not clearly specify what each emoji looks like, based on the semantic understanding of the emojis, it can be inferred that the expressions of the 5 emoji images to be generated should be different. Therefore, before generating each of the multiple images, the above text needs to be expanded to generate different prompt messages. The following describes an embodiment of a method for generating prompt messages by way of example. Figure 2 Exemplarily describe an embodiment of a method for generating prompt messages.
[0059] Figure 2 FIG. shows a schematic flowchart of a method for generating prompt messages according to some embodiments of the present disclosure. As Figure 2 shown, the generation method of this embodiment includes steps S131 to S132.
[0060] In step S131, according to the text, determine multiple groups of personalized features that match the text, and the number of groups of personalized features is consistent with the number of images to be generated indicated by the text.
[0061] Each group of personalized features may include one or more personalized features. The personalized features may be a more detailed description of the content in the text. For example, the content of the text is "5 emojis", among which, "emoji" can be further refined into "emoji expressing happiness", "emoji expressing sadness", "emoji expressing crying", "emoji expressing doubt", "emoji expressing anger", then the 5 groups of personalized features corresponding to the 5 images to be generated include, for example, "happiness", "sadness", "crying", "doubt", "anger".
[0062] Thus, the text can be reasonably expanded to obtain a detailed description of each of the multiple images to be generated, and the descriptions of each image are different.
[0063] In step S132, for each of the multiple groups of personalized features, prompt information is generated based on the text and the personalized features, and the number of pieces of prompt information is consistent with the number of images to be generated indicated by the text.
[0064] For example, the text and the personalized features can be concatenated or fused so that the generated prompt information includes the text and the personalized features. Of course, other content can also be included in the prompt information as needed, which will not be elaborated here.
[0065] Through the above processing, the information in the text can be expanded according to the intention of batch generating multiple images to generate different prompt information, and each piece of prompt information is used to generate one image. Thus, batch generation of images with different contents can be achieved, improving the operation efficiency of the user for batch generating images.
[0066] Figure 3 A schematic comparison diagram of the batch generation effect of the images of the present disclosure and the image generation effect in the related art is shown.
[0067] In Figure 3 Part (a) shows a schematic diagram of the image batch generation effect of the method according to some other embodiments of the present disclosure. In the interaction interface 31 between the user and the intelligent agent, the user sends an image generation instruction 311 "generate 4 expressions". After semantic understanding, this instruction has the intention of generating 4 images. Then, based on "expression", 4 different pieces of prompt information are generated respectively based on different personalized features, such as 4 pieces of prompt information including "expression of happiness", "expression of sadness", "expression of crying", and "expression of doubt". Based on these 4 pieces of prompt information, 4 images representing different emotions are generated respectively, as shown in 312 to 315. Thus, the above embodiments can perform multiple different semantic expansions based on the text associated with the image generation instruction (in this example, the image generation instruction itself), generate multiple different pieces of prompt information, and then generate images with different contents based on the multiple different pieces of prompt information.
[0068] In Figure 3 Part (b) shows a schematic diagram of the image generation effect of the method according to the related art. In the interaction interface 32 between the user and the intelligent agent, the user sends an image generation instruction 321 "generate 4 expressions". One piece of prompt information is generated based on the instruction 321, and this piece of prompt information is input into the model 2 times (the specific number of times is determined according to the number of images provided to the user by the application each time) to obtain 2 images returned by the model based on the same piece of prompt information, as shown in 322 and 323. In the images 322 and 323, each image includes 4 expressions at the same time.
[0069] Thus, through the embodiments of the present disclosure, it is possible to achieve batch generation of images with different contents, improving the operation efficiency of users for batch image generation.
[0070] In some embodiments, determining whether an image generation instruction includes an intention of batch image generation according to the semantic understanding result of the text includes: determining the number of text units in the text according to the semantic understanding result of the text; in response to the text including multiple text units, determining that the image generation instruction includes an intention of batch image generation, where the multiple text units are used to generate different images.
[0071] A text unit refers to a continuous segment of text in the text. For example, it can be multiple consecutive words, a sentence, or multiple consecutive sentences, etc. If the text includes descriptions of multiple images, then the descriptions of different images will have semantic differences. For example, if the text is "the sea in winter, shells, beach, cold tones, the wind in summer, green leaves, soda", then according to the semantics, "the sea in winter, shells, beach, cold tones" can be regarded as one text unit, and "the wind in summer, green leaves, soda" can be regarded as one text unit. The following refers to Figure 4 and exemplarily describes a method for determining the number of text units.
[0072] Figure 4 FIG. shows a flowchart of a method for determining the number of text units according to some embodiments of the present disclosure. As Figure 4 shown, the determination method of this embodiment includes steps S121 to S123.
[0073] In step S121, the text is segmented into multiple text elements.
[0074] The text elements can be divided according to characters, words, or punctuation marks. That is, each text element can be a character, a word, or a clause.
[0075] In step S122, based on the semantic relevance of the multiple text elements, the multiple text elements are grouped, and each group includes one or more consecutive text elements and corresponds to one text unit.
[0076] One text unit includes one or more consecutive text elements. Thus, the text elements in one text unit are semantically related.
[0077] For example, the text elements can be clustered according to semantic relevance, and it is specified that the text elements in the same class need to be consecutive elements in the text.
[0078] In step S123, according to the grouping result, the number of text units is determined.
[0079] For example, the number of groups is the number of text units.
[0080] In the above manner, when the user does not explicitly indicate to generate multiple images in the image generation instruction or text, the intention to batch generate images can be determined through semantic analysis. Thus, the user's instruction is simplified and the efficiency of batch generating images is improved.
[0081] Figure 5 Another comparison schematic diagram showing the batch generation effect of the images of the present disclosure and the image generation effect in the related art is shown.
[0082] In Figure 5 Part (a) shows a schematic diagram of the image batch generation effect of the method according to some embodiments of the present disclosure. In the interaction interface 51 between the user and the agent, the user sends an image generation instruction 511 "Generate images, sea in winter, shells, beach, cold tone, wind in summer, green leaves, soda". After semantic understanding, this instruction means to generate an image depicting winter and an image depicting summer. Then, a prompt message 1 is generated based on "sea in winter, shells, beach, cold tone", and another prompt message 2 is generated based on "wind in summer, green leaves, soda". Based on the prompt message 1, an image depicting winter is generated and displayed on the interface, as shown in Figure 512; based on the prompt message 2, an image depicting summer is generated and displayed on the interface, as shown in Figure 513. Thus, the above embodiments can segment the text associated with the image generation instruction (which is the image generation instruction itself in this example) into different parts according to semantics, and respectively generate different prompt messages based on the different segmented parts, and then generate images with different contents based on the different prompt messages.
[0083] In Figure 5 Part (b) shows a schematic diagram of the image generation effect of the method according to the related art. In the interaction interface 52 between the user and the agent, the user sends an image generation instruction 521 "sea in winter, shells, beach, cold tone, wind in summer, green leaves, soda". A prompt message is generated based on the instruction 521, and the prompt message is input into the model to obtain the image returned by the model based on the prompt message, as shown in Figure 522. Figure 522 includes elements representing both winter and summer at the same time.
[0084] Through the embodiments of the present disclosure, the intention of the user to batch generate images can be determined based on semantic analysis, and then images can be generated based on different parts of the text associated with the instruction. Thus, the efficiency of batch generating images is improved.
[0085] For the case where each text unit corresponds to the generation of an image, based on the text, multiple different prompt messages are generated, including: processing each of the multiple text units separately to obtain the prompt message corresponding to each text unit. For example, each text unit can be directly used as a prompt message. Alternatively, each text unit can be concatenated with a fixed instruction to obtain the prompt message corresponding to each text unit. Thus, the description information of the image in the prompt messages obtained in this way all comes from the text associated with the image generation instruction. It can be understood that by splitting the text, multiple key contents in the multiple prompt messages are obtained.
[0086] After returning multiple images generated in batch to the user, a batch editing function can be provided to the user. Figure 6 The flowchart showing the method for batch editing images according to some embodiments of the present disclosure is as follows. Figure 6 As shown, the method for batch editing images in this embodiment includes steps S61 to S63.
[0087] In step S61, in response to the generation of multiple images, a batch editing control is displayed.
[0088] That is, in the case where it is detected that multiple images have been generated, the batch editing function can be automatically recommended to the user.
[0089] In step S62, in response to the batch editing control being triggered, each of the multiple images is edited, and the editing method corresponds to the batch editing control.
[0090] The batch editing control can be associated with specific batch editing methods, such as batch image matting, batch changing the background of the image, batch modifying the expressions of the characters in the image, and so on.
[0091] The editing process for each image can be automatically completed by a machine. For example, after the batch editing control is triggered, each image can be processed according to the business logic corresponding to the editing method.
[0092] In step S63, the multiple edited images are displayed.
[0093] Through the above embodiments, it is possible to facilitate the user to perform unified processing on the generated multiple images again. Thus, in the case where the user has unified requirements for the batch-generated images, and when the generated images need to be adjusted, the user can efficiently adjust them through batch editing operations. In this way, the efficiency of batch generating images is further improved as a whole.
[0094] Two scenarios of batch editing are described below by way of example.
[0095] In some embodiments, after the first batch editing of the multiple images, a second batch editing method can continue to be recommended to the user.
[0096] Let the batch editing control be the first batch editing control. The editing method of the first batch editing control is to remove the background. In response to the batch editing control being triggered, the editing of each of the multiple images includes: in response to the batch editing control being triggered, removing the background of each of the multiple images; displaying the multiple images after the background is removed and the second batch editing control. The editing methods of the second batch editing control include at least one of adding a similar background, unifying the image style, and modifying the human features.
[0097] Removing the background can also be referred to as matting. Since only the foreground content in each generated image is retained after removing the background, other batch editing operations can continue to be recommended to the user to help the user further improve the image. For example, uniformly adding a new and unified background to these foreground contents, unifying the styles of these images, or batch-changing features such as the clothing, expressions, and hairstyles of the people in the foreground. Thus, according to the characteristics of the images edited in the first batch, a second editing method can continue to be recommended to the user.
[0098] In some embodiments, on the basis of batch generating images, videos can be further batch generated.
[0099] The editing method of the batch editing control is to generate videos. In response to the batch editing control being triggered, the editing of each of the multiple images includes: in response to the batch editing control being triggered, generating videos respectively according to each of the multiple images; displaying the multiple generated videos.
[0100] In the process of generating videos according to each image, a model based on image-to-video generation can be used. The user also sends further description information about the videos to be generated before, after, or at the same time as triggering the batch editing control. Thus, videos can be generated based on this description information and the pre-generated images, making the generated videos more in line with the user's needs.
[0101] In addition to the above batch editing method, users can also perform manual editing. For example, in a scenario where a user sends an image generation instruction through interaction with an agent, the agent usually returns the generated image through a dialogue interface. On this basis, some embodiments of the present disclosure can also display an editing interface for the generated image. In some embodiments, an interaction interface between the user and the agent is displayed, and the interaction interface includes a dialogue interface and a preview interface; in the dialogue interface, the multiple images sent by the agent to the user are displayed; in the preview interface, an editing interface for the first image among the multiple images and a preview of the multiple images are displayed; in response to the triggering of the preview of the second image among the multiple images, in the preview interface, the editing interface is switched from the first image to the second image.
[0102] Figure 7 The figure shows a schematic diagram of an interaction interface according to some embodiments of the present disclosure. Figure 7 The shown interaction interface 7 includes a dialogue interface 71 and a preview interface 72. The user can send messages to the agent through the dialogue interface 71 to generate images or conduct conversations, etc. For example, the user sends the message 711 "Generate character diagrams of 4 characters". After receiving this message, the agent generates 4 images 713 to 717 based on the intention of batch generating images in this message. Optionally, the agent sends the message 712 "Okay, 4 images have been generated for you, with one character in each image" to explain the generated images. In response to the generation of images 713 to 717, the image being edited 721 is displayed in the preview interface 72. For example, Figure 7 the image being edited shown in the figure is image 714. The preview interface 72 also includes previews 722 to 725 corresponding to images 713 to 717. For example, in response to the user triggering the preview 722 corresponding to image 713, image 713 and its edited effect are displayed in 721.
[0103] The user can input an editing instruction through the input control 717 in the dialogue interface 71. The input editing instruction can also be a batch editing instruction or an editing instruction for a certain image. In addition, the user can also trigger an editing control in the editing area to perform operations.
[0104] Thus, the user can efficiently edit the batch-generated images.
[0105] Some embodiments of the present disclosure can also efficiently assist users in creating storyboard images. Figure 8 The figure shows a schematic flowchart of a method for generating a storyboard image according to some embodiments of the present disclosure. As Figure 8 shown, the generation method of this embodiment includes steps S81 to S86.
[0106] In step S81, according to the storyboard creation instruction input by the user, a storyboard script is generated, and the storyboard script includes description information of multiple shots.
[0107] For example, the user can send a message to the intelligent agent, and this message describes that the user wants to create a storyboard script, as well as the user's description of the entire script, or the description of a specific shot. The intelligent agent can process the message sent by the user through a text generation model to generate a storyboard script and return it to the user.
[0108] The description information of each shot includes, for example, shot number, scene type, picture content, duration, and so on. As needed, some of the above description information can be replaced with other information, or other information can be further included based on the above description information.
[0109] In step S82, an instruction for generating storyboard images based on the storyboard script is obtained from the user input, and the text associated with the instruction for generating storyboard images is the storyboard script.
[0110] In step S83, according to the semantic understanding result of the storyboard script, it is determined that the image generation instruction includes the intention of generating images in batches.
[0111] Since the storyboard script is the description information of different shots and pictures, the scene corresponding to generating images in batches can be clearly determined.
[0112] In step S84, based on the storyboard script, multiple different prompt messages are generated.
[0113] The description information of each shot in the storyboard script is used to generate a prompt message. That is, the number of generated prompt messages is consistent with the number of shots involved in the storyboard script.
[0114] In some embodiments, based on the description information of each shot in the storyboard script, the shot features of each shot are determined, and the shot features include at least one of scene type, perspective, character description information, picture composition, and picture color; according to the description information and shot features of each shot, a prompt message corresponding to the shot is generated. That is, when generating the prompt message, a formatted template of the prompt message can be used, and the content required by the formatted template is extracted from the description information in the storyboard script. Thus, the generated prompt message can more fully describe the picture content of the sub-shot to be generated.
[0115] In step S85, a machine learning model is used to process each prompt message to generate a storyboard image corresponding to each prompt message.
[0116] In step S86, the multiple generated storyboard images are displayed.
[0117] Through the above embodiments, users can efficiently achieve continuous creation. That is, first, use a machine learning model to create a storyboard script, and then based on this script, instruct the machine learning model to create storyboard images. Thus, users can generate a storyboard script and batch generate storyboard images with at least two interactions, without the need for users to create the description information and storyboard images of each shot one by one, thereby improving the creation efficiency of storyboard images and greatly enhancing the creation efficiency of film and television animation works.
[0118] After obtaining the storyboard images, automated video creation can be further performed. In some embodiments, in response to a user's video generation instruction, a video is generated according to multiple images and the duration of each shot. For example, multiple images can be input into a video generation model to obtain the generated video. Or, two adjacent images can be used as the start and end frames of a video to generate a video segment, and after generating the video segments corresponding to each group of adjacent images among the multiple generated images, they are spliced into a complete video. Thus, users can automatically generate a storyboard script, storyboard images, and a video based on the storyboard script and storyboard images with at least three instructions. Compared with directly generating a video, this method enables users to understand detailed information in each process of video creation, so that in the case of dissatisfaction with the generated result, they can resend instructions or issue modification instructions to make the generated video of higher quality.
[0119] Figure 9 Shows a schematic diagram of an interaction interface according to other embodiments of the present disclosure. In Figure 9 In the shown interface 9, the user sends a message 91 to the agent: "Please create a storyboard script for a story about..., which can fully show the exciting highlights..., and please help me create a storyboard script presented in a table." The agent sends a message 92 to the user: "I will show the highlights with an intense battle scene and generate a storyboard script for you" and the generated storyboard script 93, which is presented in the form of a table plus text. Of course, it can also be presented in pure text form. The user further sends a message 94: "Generate storyboard images." Then, the agent returns several storyboard images 95 generated based on the storyboard script to the user. The user then sends a message 96 to the agent: "Generate a video." The agent returns a video 97 generated based on the storyboard images to the user. Thus, through simple interactions, users can efficiently generate a video according to the creation process of multimedia content.
[0120] The above introduces the method embodiments of the present disclosure. Next, the apparatus for implementing the methods of each embodiment will be further described.
[0121] Figure 10 Shows a schematic structural diagram of an apparatus for batch generating images according to some embodiments of the present disclosure. As Figure 10As shown, the batch image generation device 10 of this embodiment includes: an acquisition module 101 configured to acquire an image generation instruction input by a user, where the image generation instruction is associated with text; a determination module 102 configured to determine whether the image generation instruction includes an intention to generate images in batches based on the semantic understanding result of the text; a prompt generation module 103 configured to, in response to the image generation instruction including an intention to generate images in batches, generate a plurality of different prompt messages based on the text; an image generation module 104 configured to use a machine learning model to process each prompt message to generate an image corresponding to each prompt message; and a display module 105 configured to display the generated plurality of images.
[0122] In the above embodiment, by performing semantic understanding on the text associated with the image generation instruction, it is determined whether the instruction includes an intention to generate images in batches. In the scenario of generating images in batches, different prompt messages are generated based on the text, so as to generate a plurality of images with different contents in batches based on one instruction of the user. Thus, the efficiency of the user in generating images in batches can be improved.
[0123] In some embodiments, the determination module 102 is further configured to: perform semantic understanding on the text to determine whether the text includes an indication of the number of images to be generated; in response to the text including an indication of the number of images to be generated and the number of images to be generated being greater than 1, determine that the image generation instruction includes an intention to generate images in batches.
[0124] In some embodiments, the prompt generation module 103 is further configured to: determine multiple sets of personalized features that match the text according to the text, where the number of sets of personalized features is consistent with the number of images to be generated indicated by the text; for each set of the multiple sets of personalized features, generate a prompt message according to the text and the personalized features, where the number of prompt messages is consistent with the number of images to be generated indicated by the text.
[0125] In some embodiments, the determination module 102 is further configured to: determine the number of text units in the text based on the semantic understanding result of the text; in response to the text including multiple text units, determine that the image generation instruction includes an intention to generate images in batches, where the multiple text units are used to generate different images.
[0126] In some embodiments, the determination module 102 is further configured to: divide the text into multiple text elements; group the multiple text elements based on the semantic relevance of the multiple text elements, where each group includes one or more consecutive text elements and corresponds to one text unit; and determine the number of text units according to the grouping result.
[0127] In some embodiments, the prompt generation module 103 is further configured to process each of the multiple text units separately to obtain the prompt information corresponding to each text unit.
[0128] In some embodiments, the display module 105 is further configured to display a batch editing control in response to generating multiple images; the device 10 for batch generating images further includes an editing module 106, which is configured to edit each of the multiple images in response to the batch editing control being triggered, and the editing method corresponds to the batch editing control; the display module 105 is further configured to display the multiple edited images.
[0129] In some embodiments, the batch editing control is a first batch editing control, and the editing method of the first batch editing control is to remove the background. The editing module 106 is further configured to: in response to the batch editing control being triggered, remove the background of each of the multiple images; the display module 105 is further configured to: display the multiple images with the background removed and a second batch editing control, and the editing methods of the second batch editing control include at least one of adding a similar background, unifying the image style, and modifying the character features.
[0130] In some embodiments, the editing method of the batch editing control is to generate a video. The editing module 106 is further configured to: in response to the batch editing control being triggered, generate a video respectively according to each of the multiple images; the display module 105 is further configured to: display the multiple generated videos.
[0131] In some embodiments, the display module 105 is further configured to: display an interaction interface between the user and the agent, and the interaction interface includes a dialogue interface and a preview interface; in the dialogue interface, display the multiple images sent by the agent to the user; in the preview interface, display an editing interface of the first image among the multiple images and a preview of the multiple images; in response to the preview of the second image among the multiple images being triggered, in the preview interface, switch from the editing interface of the first image to the editing interface of the second image.
[0132] In some embodiments, the device 10 for batch generating images further includes a text generation module 107, which is configured to: generate a storyboard script according to the storyboard creation instruction input by the user, and the storyboard script includes description information of multiple shots; the acquisition module 101 is further configured to: acquire the instruction input by the user for generating storyboard images based on the storyboard script, and the text associated with the instruction for generating storyboard images is the storyboard script.
[0133] In some embodiments, the prompt generation module 103 is further configured to: based on the description information of each shot in the storyboard script, determine the shot characteristics of each shot, where the shot characteristics include at least one of shot type, perspective, character description information, picture composition, and picture color; generate prompt information corresponding to the shot according to the description information and shot characteristics of each shot.
[0134] In some embodiments, the image batch generation device 10 further includes a video generation module 108, configured to: in response to a user's video generation instruction, generate a video according to a plurality of images and the duration of each shot.
[0135] Figure 11 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0136] The memory 111 is used to store one or more computer-readable instructions. The memory 111 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. The memory 111 may store, for example, an operating system, application programs, a boot loader (BootLoader), a database, and other programs, and may also store various application programs and various data.
[0137] The processor 112 is used to run the computer-readable instructions to implement the method described in any of the foregoing embodiments. For the specific implementation of each step of the method, reference may be made to the above embodiments, and the repeated parts will not be elaborated here. Thus, in the scenario of batch generating images, different prompt information is generated based on the text, so as to batch generate multiple images with different contents based on a single instruction of the user. Thus, the efficiency of the user in batch generating images can be improved.
[0138] The processor 112 may be configured to execute the steps of the foregoing embodiments. The processor 112 may be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) may be of the X86 or ARM architecture, etc.
[0139] The processor 112 and the memory 111 can communicate with each other directly or indirectly. For example, the processor 112 and the memory 111 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 112 and the memory 111 can also communicate with each other through a system bus, and the present disclosure does not limit this.
[0140] It should be noted that Figure 11 The components of the electronic device 11 shown are merely exemplary and not restrictive. According to actual application needs, the electronic device 11 may also have other components. The processor 112 can control other components in the electronic device 11 to perform desired functions.
[0141] The electronic device 11 can be implemented in software, firmware, and / or hardware, and can be integrated into a device installed with relevant application programs.
[0142] Figure 12 A block diagram of an electronic device according to other embodiments of the present disclosure is shown.
[0143] Figure 12 The electronic device 12 shown can be a computer system with a dedicated hardware structure, and can perform corresponding functions when installed with relevant application programs.
[0144] The electronic device includes but is not limited to mobile terminals such as smart phones, laptop computers, personal digital assistants (PDAs), tablet personal computers (Tablet PCs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0145] As Figure 12 shown, the central processing unit (CPU) 121 executes various processes according to the programs stored in the read-only memory (ROM) 122 or the programs loaded from the storage section 128 into the random access memory (RAM) 123. In the RAM 123, data required when the CPU 121 executes various processes, etc. is stored as needed. The central processing unit is merely exemplary, and it can also be other types of processors, such as the various processors described above. The ROM 122, the RAM 123, and the storage section 128 can be various forms of computer-readable storage media. It should be noted that although Figure 12 the ROM 122, the RAM 123, and the storage section 128 are shown separately, one or more of them can be combined, or located in the same or different memories or storage modules.
[0146] The CPU 121, ROM 122, and RAM 123 are connected to each other via a bus 124. An input / output interface 125 is also connected to the bus 124.
[0147] The following components are connected to the input / output interface 125: an input section 126 such as a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output section 127 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage section 128 including a hard disk, a magnetic tape, etc.; and a communication section 129 including a network interface card such as a LAN card, a modem, etc. The communication section 129 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 12 some of the components shown in the electronic device 12 communicate via the bus 124, they can also communicate via a network or other means, where the network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0148] As needed, a drive 1210 is also connected to the input / output interface 125. A removable medium 1211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 1210 as needed, so that a computer program read therefrom is installed in the storage section 128 as needed.
[0149] In the case where the above-described series of processes are implemented by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 1211.
[0150] According to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when run on a computer, causes the computer to implement the method described in any one of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium and includes program code for performing the method shown in the flowchart. In such an embodiment, the computer instructions can be downloaded and installed from a network via the communication section 129, or installed from the storage section 128, or installed from the ROM 122. When the computer program is executed by the CPU 121, the method of the embodiment of the present disclosure is executed.
[0151] It should be noted that, in the context of the present disclosure, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0152] A computer-readable medium can be a computer-readable storage medium, a computer-readable signal medium, or any combination of the above two.
[0153] Computer-readable storage media include, but are not limited to: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. Computer instructions are stored on the computer-readable storage medium, and when executed by a processor, the instructions implement the method described in any of the foregoing embodiments.
[0154] Thus, in the scenario of batch generating images, different prompt messages are generated based on the text, so as to batch generate multiple images with different contents based on a single instruction from the user. Thus, the efficiency of the user in batch generating images can be improved.
[0155] A computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0156] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0157] In some embodiments, a computer program is also provided, including: instructions that, when executed by a processor, cause the processor to execute the method described in any of the foregoing embodiments. For example, the instructions can be embodied as computer program code.
[0158] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0160] The functions described above may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0161] According to some embodiments of the present disclosure, a method for batch generating images is provided, including: obtaining an image generation instruction input by a user, where the image generation instruction is associated with text; determining whether the image generation instruction includes an intention to batch generate images according to the semantic understanding result of the text; in response to the image generation instruction including the intention to batch generate images, generating a plurality of different prompt messages based on the text; using a machine learning model to process each prompt message to generate an image corresponding to each prompt message; and displaying the plurality of generated images.
[0162] In some embodiments, determining whether the image generation instruction includes an intention to batch generate images according to the semantic understanding result of the text includes: performing semantic understanding on the text to determine whether the text includes an indication of the number of images to be generated; in response to the text including an indication of the number of images to be generated and the number of images to be generated being greater than 1, determining that the image generation instruction includes an intention to batch generate images.
[0163] In some embodiments, in response to the image generation instruction including an intention to batch generate images, generating a plurality of different prompt messages based on the text includes: determining multiple sets of personalized features that match the text according to the text, where the number of sets of personalized features is consistent with the number of images to be generated indicated by the text; for each set of the multiple sets of personalized features, generating a prompt message according to the text and the personalized features, where the number of prompt messages is consistent with the number of images to be generated indicated by the text.
[0164] In some embodiments, determining whether the image generation instruction includes an intention to batch generate images according to the semantic understanding result of the text includes: determining the number of text units in the text according to the semantic understanding result of the text; in response to the text including multiple text units, determining that the image generation instruction includes an intention to batch generate images, where the multiple text units are used to generate different images.
[0165] In some embodiments, determining the number of text units in the text according to the semantic understanding result of the text includes: splitting the text into multiple text elements; grouping the multiple text elements based on the semantic relevance of the multiple text elements, where each group includes one or more consecutive text elements and corresponds to a text unit; and determining the number of text units according to the grouping result.
[0166] In some embodiments, in response to the image generation instruction including an intention to batch generate, generating a plurality of different prompt messages based on the text includes: processing each of the multiple text units separately to obtain a prompt message corresponding to each text unit.
[0167] In some embodiments, the batch generation method further includes: in response to generating multiple images, displaying a batch editing control; in response to the batch editing control being triggered, editing each of the multiple images, where the editing method corresponds to the batch editing control; and displaying the multiple edited images.
[0168] In some embodiments, the batch editing control is a first batch editing control, and the editing method of the first batch editing control is to remove the background. In response to the batch editing control being triggered, editing each of the multiple images includes: in response to the batch editing control being triggered, removing the background of each of the multiple images; and displaying the multiple images with the background removed and a second batch editing control, where the editing method of the second batch editing control includes at least one of adding a similar background, unifying the image style, and modifying the character features.
[0169] In some embodiments, the editing method of the batch editing control is to generate videos. In response to the batch editing control being triggered, editing each of the multiple images includes: in response to the batch editing control being triggered, generating videos respectively based on each of the multiple images; and displaying the multiple generated videos.
[0170] In some embodiments, displaying the multiple generated images includes: displaying an interaction interface between the user and the agent, where the interaction interface includes a dialogue interface and a preview interface; in the dialogue interface, displaying the multiple images sent by the agent to the user; in the preview interface, displaying an editing interface of the first image among the multiple images and a preview of the multiple images; and in response to a trigger on the preview of the second image among the multiple images, switching from the editing interface of the first image to the editing interface of the second image in the preview interface.
[0171] In some embodiments, the batch generation method further includes: generating a storyboard script according to the storyboarding creation instruction input by the user, where the storyboard script includes description information of multiple shots; and obtaining the image generation instruction input by the user includes: obtaining the instruction input by the user to generate storyboard images based on the storyboard script, and the text associated with the instruction to generate storyboard images is the storyboard script.
[0172] In some embodiments, generating multiple different prompt messages based on the text includes: determining the shot features of each shot based on the description information of each shot in the storyboard script, where the shot features include at least one of shot scale, perspective, character description information, picture composition, and picture color; and generating prompt messages corresponding to the shots according to the description information and shot features of each shot.
[0173] In some embodiments, the storyboard script further includes the duration of each shot, and the image batch generation method further includes: in response to the user's video generation instruction, generating a video according to the multiple images and the duration of each shot.
[0174] According to some embodiments of the present disclosure, there is provided an apparatus for batch generating images, including: an acquisition module configured to acquire an image generation instruction input by a user, the image generation instruction being associated with text; a determination module configured to determine whether the image generation instruction includes an intention to batch generate images according to the semantic understanding result of the text; a prompt generation module configured to generate a plurality of different prompt messages based on the text in response to the image generation instruction including the intention to batch generate images; an image generation module configured to process each prompt message by using a machine learning model to generate an image corresponding to each prompt message; and a display module configured to display the plurality of generated images.
[0175] According to some embodiments of the present disclosure, there is provided an electronic device, including: a memory; and a processor coupled to the memory, the processor being configured to execute the method for batch generating images according to any one of the embodiments of the present disclosure based on instructions stored in the memory.
[0176] According to some embodiments of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor, implements the method for batch generating images according to any one of the embodiments of the present disclosure.
[0177] According to some embodiments of the present disclosure, there is provided a computer program product, which when run on a computer, causes the computer to implement the method for batch generating images according to any one of the embodiments of the present disclosure.
[0178] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration purposes and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments can be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for batch generating images, comprising: Obtaining an image generation instruction input by a user, where the image generation instruction is associated with text; Determining whether the image generation instruction includes an intention to batch generate images according to the semantic understanding result of the text; In response to the image generation instruction including the intention to batch generate images, generating a plurality of different prompt messages based on the text; Using a machine learning model to process each prompt message to generate an image corresponding to each prompt message; Displaying the generated plurality of images.
2. The batch generation method according to claim 1, wherein The determining whether the image generation instruction includes an intention to batch generate images according to the semantic understanding result of the text includes: Performing semantic understanding on the text to determine whether the text includes an indication of the number of generated images; In response to the text including an indication of the number of generated images and the number of generated images being greater than 1, determining that the image generation instruction includes the intention to batch generate images.
3. The batch generation method according to claim 2, wherein The generating a plurality of different prompt messages based on the text in response to the image generation instruction including the intention to batch generate images includes: According to the text, determining multiple groups of personalized features that match the text, where the number of groups of personalized features is consistent with the number of generated images indicated by the text; For each of the multiple groups of personalized features, generating a prompt message according to the text and the personalized features, where the number of prompt messages is consistent with the number of generated images indicated by the text.
4. The batch generation method according to claim 1, wherein, The determining whether the image generation instruction includes an intention to batch generate images according to the semantic understanding result of the text includes: Determining the number of text units in the text according to the semantic understanding result of the text; In response to the text including the multiple text units, determining that the image generation instruction includes the intention to batch generate images, where the multiple text units are used to generate different images.
5. The batch generation method according to claim 4, wherein, The determining the number of text units in the text according to the semantic understanding result of the text includes: Segmenting the text into multiple text elements; Grouping the multiple text elements based on the semantic relevance of the multiple text elements, where each group includes one or more consecutive text elements and corresponds to one text unit; Determining the number of text units according to the grouping result.
6. The batch generation method according to claim 4 or 5, wherein, The generating a plurality of different prompt messages based on the text in response to the image generation instruction including the intention to batch generate includes: Processing each of the multiple text units respectively to obtain a prompt message corresponding to each text unit.
7. The batch generation method according to claim 1, further comprising: In response to generating the multiple images, displaying a batch editing control; In response to the batch editing control being triggered, editing each of the multiple images, where the editing method corresponds to the batch editing control; Displaying the edited multiple images.
8. The batch generation method according to claim 7, wherein, The batch editing control is the first batch editing control, and the editing method of the first batch editing control is to remove the background. Responding to the triggering of the batch editing control and editing each of the multiple images includes: Responding to the triggering of the batch editing control, removing the background of each of the multiple images; Displaying the multiple images after removing the background and a second batch editing control, and the editing method of the second batch editing control includes at least one of adding a similar background, unifying the image style, and modifying the character features.
9. The batch generation method according to claim 7 or 8, wherein The editing method of the batch editing control is to generate a video. Responding to the triggering of the batch editing control and editing each of the multiple images includes: Responding to the triggering of the batch editing control, generating a video respectively according to each of the multiple images; Displaying the generated multiple videos.
10. The batch generation method according to claim 1 or 7, wherein, The displaying the generated multiple images includes: Displaying an interaction interface between the user and the agent, and the interaction interface includes a dialogue interface and a preview interface; In the dialogue interface, displaying the multiple images sent by the agent to the user; In the preview interface, displaying an editing interface of the first image among the multiple images and a preview of the multiple images; Responding to the triggering of the preview of the second image among the multiple images, switching from the editing interface of the first image to the editing interface of the second image in the preview interface.
11. The batch generation method according to claim 1, wherein: The batch generation method further includes: generating a storyboard script according to a storyboard creation instruction input by the user, and the storyboard script includes description information of multiple shots; The obtaining the image generation instruction input by the user includes: obtaining the instruction input by the user to generate storyboard images based on the storyboard script, and the text associated with the instruction to generate storyboard images is the storyboard script.
12. The batch generation method according to claim 11, wherein, The generating multiple different prompt messages based on the text includes: Based on the description information of each shot in the storyboard script, determining the shot characteristics of each shot, and the shot characteristics include at least one of shot type, perspective, character description information, picture composition, and picture color; Generating a prompt message corresponding to the shot according to the description information and shot characteristics of each shot.
13. The batch generation method according to claim 11 or 12, wherein, The storyboard script further includes the duration of each shot, and the batch generation method further includes: Responding to the video generation instruction of the user, generating a video according to the multiple images and the duration of each shot.
14. An apparatus for batch generation of images, comprising: An obtaining module, configured to obtain an image generation instruction input by a user, and the image generation instruction is associated with text; A determining module, configured to determine whether the image generation instruction includes an intention to batch generate images according to a semantic understanding result of the text; A prompt generation module, configured to respond to the image generation instruction including an intention to batch generate images, and generate multiple different prompt messages based on the text; An image generation module, configured to use a machine learning model to process each prompt message to generate an image corresponding to each prompt message; A display module, configured to display a plurality of generated images.
15. An electronic device, comprising: A memory; And A processor coupled to the memory, the processor being configured to execute the method for batch generation of images according to any one of claims 1 to 13 based on instructions stored in the memory.
16. A computer-readable storage medium, having stored thereon a computer program, which when executed by a processor, implements the method for batch generation of images according to any one of claims 1 to 13.
17. A computer program product, which when run on a computer, causes the computer to implement the method for batch generation of images according to any one of claims 1 to 13.