Image generation method and device, medium, electronic equipment and program product
By generating multiple series of images from reference information, the problem of being unable to generate images with a unified style and a story in batches in existing technologies is solved, achieving the effect of realistically reproducing photo albums in real-world scenes and improving the user experience.
Patent Information
- Application Number
- CN202511725675.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-10
AI Technical Summary
Existing AI image generation technology cannot simulate the multiple scene transitions and the evolution of people's states in real photo collections, nor can it generate images with a consistent style and a storytelling quality in batches.
By acquiring reference information, multiple series of images related to the first object are generated, ensuring that the visual style of the images is similar to that of the reference information and revolves around the theme indicated by the reference information. Multiple series of images can be generated at once.
It achieves stylistic consistency and thematic coherence across multiple image series, enhancing the user's image creation experience and realistically reproducing the scenes in a photo album.
Smart Images

Figure CN121505064A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image technology, and more specifically, to an image generation method, apparatus, medium, electronic device, and program product. Background Technology
[0002] Among related technologies, AI (Artificial Intelligence) image generation only supports isolated generation of a single image at a time, and cannot simulate the multi-scene switching, character state evolution and storytelling common in real photo albums. Therefore, we continue to provide a solution that can generate images in batches. Summary of the Invention
[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, this disclosure provides an image generation method, including: Obtain reference information; A first image set is displayed, comprising multiple series of images related to a first object generated based on the reference information. The visual style of the multiple series of images is similar to the visual style corresponding to the reference information, and the image content of the multiple series of images revolves around the theme indicated by the reference information.
[0005] Secondly, this disclosure provides an image generation apparatus, comprising: The acquisition module is configured to acquire reference information. The display module is configured to display a first image set, which includes multiple series of images related to a first object generated based on the reference information. The visual style of the multiple series of images is similar to the visual style corresponding to the reference information, and the image content of the multiple series of images revolves around the theme indicated by the reference information.
[0006] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0007] Fourthly, this disclosure provides an electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0008] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0009] Based on the above technical solution, by acquiring reference information and displaying a first image set including multiple series of images related to the first object generated based on the reference information, it is possible to enable users to generate multiple series of images at once by inputting reference information. The visual style of the generated series of images is similar to that of the reference information, and the image content of the series of images revolves around the theme indicated by the reference information. This ensures that the generated first image set can realistically reproduce the scene of the user taking a photo album in the real world, so that the multiple series of images can belong to the same theme, maintain a similar style, and have empty shots and storytelling, greatly improving the user's image creation experience.
[0010] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a flowchart illustrating an image generation method according to some embodiments.
[0012] Figure 2 This is a schematic diagram illustrating a first image set according to some embodiments.
[0013] Figure 3 This is a schematic diagram of a second display interface shown according to some embodiments.
[0014] Figure 4 This is a schematic diagram of a first interface shown according to some embodiments.
[0015] Figure 5 This is a schematic diagram of a fourth interface shown according to some embodiments.
[0016] Figure 6 This is a schematic diagram of a second interface shown according to some embodiments.
[0017] Figure 7 This is a schematic diagram of a third interactive control shown according to some embodiments.
[0018] Figure 8 This is a schematic diagram of a third interface shown according to some embodiments.
[0019] Figure 9 This is a schematic diagram illustrating the production progress according to some embodiments.
[0020] Figure 10 This is a logical schematic diagram illustrating the generation of a first image set according to some embodiments.
[0021] Figure 11 This is a schematic diagram of a first image set shown according to some embodiments.
[0022] Figure 12 This is a schematic diagram of the structure of an image generation apparatus according to some embodiments.
[0023] Figure 13 This is a schematic diagram of the structure of an electronic device according to some embodiments. Detailed Implementation
[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0026] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0028] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0030] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0031] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0032] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0033] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0034] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0035] Figure 1 This is a flowchart illustrating an image generation method according to some embodiments. For example... Figure 1 As shown, this disclosure provides an image generation method, which can be executed by an image generation device, which can be implemented in software and / or hardware. Figure 1 As shown, the method may include the following steps.
[0036] In step 110, reference information is obtained.
[0037] Here, the reference information can be determined by the user. For example, the user can input specific content as reference information or select specific content as reference information as needed. The reference information is used to instruct the generation of an image set related to the first object based on the reference information.
[0038] The first object can refer to the user who inputs reference information, or a person, animal, still life, etc., specified by the user. For example, when user A needs to generate a first image set corresponding to user A, then user A is the first object. When user A needs to generate a first image set corresponding to user B, then user B is the first object. When user A needs to generate a first image set of animal C, then animal C is the first object.
[0039] In some embodiments, the reference information may include reference images and / or reference text. That is, reference images and / or reference text can be input as reference information for generating the image set. Here, the image set can be understood as the photo set that the user needs to generate.
[0040] Reference images provide visual features, appearance, and structural layout data as the basis for image generation. For example, when a user needs to generate a portrait album of people, they can provide one or more portrait photos of people as reference images. When a user needs to generate a pet portrait album, they can provide one or more pet photos as reference images. Reference text can refer to the text content entered by the user, which describes the image content or style corresponding to the image set to be generated. For example, the reference text could be, "The first image in the portrait album depicts an afternoon coffee time, and the second image shows someone wearing a red dress standing under a cherry blossom tree."
[0041] In other words, in this embodiment of the disclosure, users can generate multiple series of images by inputting a reference image, inputting reference text, or inputting both a reference image and reference text.
[0042] In step 120, a first image set is displayed, which includes a series of images related to the first object generated based on reference information.
[0043] Here, the first image set includes multiple series of images. The visual style of these multiple series of images is similar to the visual style of the reference information, and the image content of these multiple series of images revolves around the theme indicated by the reference information.
[0044] For example, the visual style corresponding to multiple series of images can refer to image style, such as retro style, Chinese trend style, forest-themed healing style, etc. The multiple series of images included in the first image set can be multiple images organized together according to a certain correlation (such as chronological order, scene changes, style unity, or story evolution) around the theme indicated by the reference information (such as subject or narrative logic). The first image set can be a photo album of a first object, which conveys a complete visual information or story through multiple series of images with a unified style.
[0045] Furthermore, the series of images can be a series of images related to the first object. The series of images being related to the first object means that all the images in the series include the first object or items related to the first object. For example, when user A needs to generate their own photo album, user A is the first object, and the series of images related to the first object are the series of images related to user A. As another example, the first object can be a specified scene, which can be a virtual scene or a real scene. By using the scene as the first object, subsequent developments in generating scene images can be extended based on the scene images input by the user.
[0046] It is worth noting that the series of multiple images may include at least one image that does not contain the first object. For example, the series of multiple images may include an establishing shot image. An establishing shot image refers to a scene without the first object, conveying visual information through nature or still life. Through establishing shots, the theme can be metaphorically deepened in the series of multiple images, and it also has narrative and lyrical functions.
[0047] In some embodiments, reference information can be input into the trained image generation model to obtain a first set of images output by the image generation model. The image generation model outputs a series of images based on the reference information.
[0048] It should be understood that, with reference information, the image generation model can understand the reference information and generate multiple series of images at once.
[0049] For example, multiple series of images included in the first image set can be displayed sequentially. For instance, multiple series of images can be displayed in the form of a carousel. A carousel displays multiple images or content cards sequentially within a fixed area, either automatically or manually.
[0050] In some embodiments, the first image set may further include a cover image obtained by stitching together multiple series of images. Exemplarily, the cover image may be obtained by stitching together multiple series of images using a preset stitching layout. The preset stitching layout may be a layout selected by the user according to their needs. Of course, the preset stitching layout may also be determined based on the number of multiple series of images and / or the aspect ratio of the series of images. For example, aspect ratios of 3:4, 1:1, and 4:3 may each correspond to a different stitching layout.
[0051] Accordingly, a cover image can be displayed, followed by a series of images displayed sequentially. It should be understood that the cover image and the series of images can be displayed in the form of a carousel.
[0052] Figure 2 This is a schematic diagram illustrating a first image set according to some embodiments. For example... Figure 2 As shown, in the first display interface 200, the first image set, including cover image 201, series image 1, series image 2, series image 3, and series image 4, can be displayed sequentially. It should be understood that the user can control the display of series image 1, series image 2, series image 3, and series image 4 through a swiping operation. Cover image 201 is obtained by stitching together series image 1, series image 2, series image 3, and series image 4 using a four-grid mosaic layout.
[0053] It's worth noting that users can choose to share and / or save the first set of images, depending on their needs. For example... Figure 2 As shown, the first display interface 200 includes a first interactive control 202. Users can use the sub-controls in the first interactive control 202 to choose to save the first image set locally and / or share the first image set to the application (APP). For example, users can use the first interactive control 202 to publish the generated first image set to the community.
[0054] Figure 3 This is a schematic diagram of a second display interface shown according to some embodiments. For example... Figure 3 As shown, when a user shares the first image set to the community, the first image set 301 can be displayed on the second display interface 300. It should be understood that the second display interface 300 can be the application's homepage interface; when a user generates the first image set in the application, the generated first image set can be published in the application. Displaying the first image set 301 can be done as a cover image; when a user clicks on the cover image, they can view the complete first image set 301.
[0055] Therefore, by acquiring reference information and displaying a first image set including multiple series of images related to the first object generated based on the reference information, it is possible to enable users to generate multiple series of images at once by inputting reference information. The visual style of the generated series of images is similar to that of the reference information, and the image content of the series of images revolves around the theme indicated by the reference information. This ensures that the generated first image set can realistically reproduce the scene of the user taking the photo album in the real world, so that the multiple series of images belong to the same theme, maintain a similar style, and have empty shots and storytelling, greatly improving the user's image creation experience.
[0056] In some feasible implementations, where the reference information includes reference images, the reference images may include images associated with a first object, images associated with a second object, or a second set of images associated with a second object.
[0057] Here, the first object and the second object can be different users, animals, or still life objects. For example, the first object can be user A, and the second object can be user B. The first object can be animal C, and the second object can be animal D.
[0058] The image related to the first object can refer to an image with the first object as its main subject. For example, when the first user takes or generates the first image of the first object, if the first user thinks the first image is good, the first user can choose to use the first image as reference information and generate a series of images based on the first image. The first image and the series of images form the first image set of the first object. It should be understood that the first object can be the first user; that is, the first user can use a photograph of themselves as reference information to generate a photo set belonging to the first user. Of course, the first object can also be other people, animals, or still lifes, which can be specified by the first user.
[0059] Images related to the second object can refer to images with the second object as the main subject. For example, when a first user views a second image related to the second object in the community, if the first user finds the composition or style of the second image appealing, the first user can select the second image as reference information and generate a series of multiple images related to the first object based on it. These multiple series of images form the first image set of the first object. It should be understood that the first object can be the first user, and the second object can be the second user. That is, the first user can use a photo of a second user as reference information to generate a photo set belonging to the first user. Of course, the first object can also be other people, animals, or still lifes, which can be specified by the first user.
[0060] The second image set related to the second object can refer to an image set with the second object as the main subject. It should be understood that the second image set related to the second image set can also be generated using the image generation method provided in this disclosure embodiment. For example, when a first user views a second image set related to the second object in a community, if the first user finds the composition or style of the second image set appealing, the first user can select the second image set as reference information and generate a series of multiple images related to the first object based on the second image set. These multiple series of images form the first image set of the first object. It should be understood that the first object can be the first user, and the second object can be the second user; that is, the first user can use a photo album published by the second user as reference information to generate a photo album belonging to the first user. Of course, the first object can also be other people, animals, or still lifes, which can be specified by the first user.
[0061] In other words, in this embodiment of the disclosure, user A can use user A's image as a reference image to generate user A's first image set, which can satisfy the user's need to expand a relatively good image into a first image set. User A can use user B's image as a reference image to generate user A's first image set, which can satisfy the user's need to quickly generate their own first image set based on images published by other users. User A can use user B's second image set as a reference image to generate user A's second image set, which can satisfy the user's need to quickly generate their own first image set based on second image sets published by other users.
[0062] In some feasible implementations, where the reference image includes an image related to the first object, the image related to the first object can be displayed in a first interface, and then, in response to a first triggering operation triggered in the first interface, the image related to the first object can be determined as reference information, the first triggering operation being used to instruct the generation of a first image set based on the image related to the first object.
[0063] Here, the first interface can be an image display interface used to show images generated related to the first object. For example, when user A generates an image of user A, the generated image of user A can be displayed on the first interface. The first triggering operation triggered on the first interface can be a click operation on a control in the first interface. For example, the first interface includes a control for triggering the generation of a first image set, and the user can trigger the first triggering operation by clicking the control. Upon detecting the first triggering operation, the images related to the first object displayed on the first interface are determined as reference information, and a first image set is generated based on the images related to the first object.
[0064] Figure 4 This is a schematic diagram of a first interface shown according to some embodiments. For example... Figure 4 As shown, in the first interface 400, a first image 401 related to the first object can be displayed. When the user thinks the generated first image 401 is good, the user can click the "photo album" control 402 to trigger the first trigger operation, using the first image 401 as reference information to generate a first image set based on the first image 401.
[0065] Of course, the first interface 400 can also include a download control 404 and a publish control 403. When the user clicks the download control 404, the first image 401 can be downloaded to the local machine. When the user clicks the publish control 403, the first image 401 can be published to the application's community.
[0066] In some embodiments, in response to a first triggering operation, a fourth interface may be displayed. The fourth interface includes an image associated with a first object and a text input area for obtaining reference text. Then, in response to a fourth triggering operation triggered on the fourth interface, the reference text entered in the text input area and the image associated with the first object are determined as reference information.
[0067] Upon detecting the first trigger operation, the user can jump from the first interface to the fourth interface. The fourth interface can be a creation interface for editing reference information. The fourth interface can display an image related to the first object and a text input area for obtaining reference text.
[0068] In the fourth interface, users can enter reference text in the text input area to further edit the reference information. It should be understood that the text input area can be a text input box or a text editor.
[0069] The fourth trigger operation can be an action to instruct the generation of the first image set, such as a click operation. When the fourth trigger operation is detected, the reference text entered in the text input area and the image related to the first object can be identified as reference information. In other words, through the text input area of the fourth interface, the user can further refine the reference information used to generate the first image set by entering reference text, so that the generated first image set can better meet the user's actual needs.
[0070] Figure 5 This is a schematic diagram illustrating a fourth interface according to some embodiments. For example... Figure 5 As shown, when the user clicks Figure 4 When the "Photo Album" control 402 is displayed, a fourth interface 500 can be switched to show the image 401, a text input area 501, and a photo album generation control 502. The user can enter reference text in the text input area 501. When the user clicks the photo album generation control 502, a fourth trigger operation is activated, using the reference text entered in the text input area 501 and the first image 401 as reference information.
[0071] It should be understood that this embodiment actually provides a method for generating a series of images related to a first object based on a user-provided reference image related to the first object. When the user is satisfied with the portrait image they created, they use the created portrait image as a reference image, thereby expanding and generating a first image set based on the user-created portrait image. Based on this, the user's need to expand and generate their own first image set based on the images they created can be met.
[0072] In some feasible implementations, where the reference image includes an image related to the second object, the image related to the second object can be displayed in the second interface, and then, in response to a second triggering operation triggered in the second interface, the image related to the second object can be identified as reference information. The second triggering operation is used to instruct the generation of a first image set based on the image related to the second object.
[0073] Here, the second interface can be an interface used to display images related to the second object. For example, when a user views images related to the second object posted by other users on the community's homepage, the user can click on these images to trigger the display of the second interface. The second triggering operation triggered on the second interface can be a click operation on a control within the second interface. For example, the second interface includes a control for triggering the generation of a first image set; the user can trigger the second triggering operation by clicking this control. Upon detecting the second triggering operation, the images related to the second object displayed on the second interface are determined as reference information to generate the first image set based on these images.
[0074] Figure 6 This is a schematic diagram of a second interface shown according to some embodiments. For example... Figure 6 As shown, in the second interface 600, a second image 601 related to the second object can be displayed. When the second image 601 is good, the user can click the "Make the same style" control 602 to trigger the second trigger operation, using the second image 601 as reference information to generate the first image set based on the second image 601.
[0075] Of course, the second interface 600 may also include a second interactive control 603. The second interactive control 603 may refer to an element used to receive interactive operations performed by the user on the second image 601. For example, the second interactive control 603 may be a like control for liking the second image 601, a forward control for sharing the second image 601, a comment control for commenting on the second image 601, and so on.
[0076] Of course, in other embodiments, upon detecting the second trigger operation, a third interactive control can be displayed on the second interface in response to the second trigger operation, and then, in response to a fifth trigger operation triggered in the third interactive control, an image related to the second object can be determined as reference information. The third interactive control can be used to allow the user to select an image generation mode, and this third interactive control can be a pop-up window displayed on the second interface.
[0077] Figure 7 This is a schematic diagram illustrating a third interactive control according to some embodiments. For example... Figure 7 As shown, when the user clicks Figure 6When the "Make the Same Style" control 602 is displayed, a third interactive control 604 can be shown in the second interface 600. The third interactive control 604 can include various image generation modes. For example, image generation modes can include AI portrait mode, AI photo album mode, and AI raw image mode. Specifically, AI portrait mode is used to generate a single portrait based on the second image 601, AI photo album mode is used to generate multiple portraits based on the second image 601, and AI raw image mode is used to generate different images based on the second image 601. The user can click the control 605 corresponding to the AI photo album mode to select it. Then, when the user clicks the "Make the Same Style with One Click" control 606, a fifth trigger operation is triggered, determining the second image 601 related to the second object as reference information.
[0078] It should be understood that this embodiment actually provides a method for generating a series of images related to a first object based on images posted by other users that are related to a second object. For example, when a user is satisfied with a photo posted by another user, the user can use the photo as a reference image to expand and generate their own first image set based on the photo posted by other users. Based on this, the user's need to expand and generate their own first image set based on their own images can be met. Therefore, the user's need to quickly generate their own first image set based on images posted by other users that are related to a second object can be met.
[0079] In some feasible implementations, when the reference image includes a second image set, the second image set can be displayed in a third interface, and then, in response to a third triggering operation triggered in the third interface, the second image set can be identified as reference information, the third triggering operation being used to instruct the generation of a first image set based on the second image set.
[0080] Here, the third interface can be an interface used to display a second image set related to the second object. For example, when a user views a second image set related to the second object posted by another user on the community's homepage, the user can click on the second image set posted by another user to trigger the display of the third interface. The third trigger operation triggered on the third interface can be a click operation on a control within the third interface. For example, the third interface includes a control for triggering the generation of a first image set, and the user can trigger the third trigger operation by clicking this control. Upon detecting the third trigger operation, the second image set related to the second object displayed on the third interface is determined as reference information, and a first image set is generated based on the second image set related to the second object.
[0081] Figure 8 This is a schematic diagram illustrating a third interface according to some embodiments. For example... Figure 8As shown, in the third interface 800, a second image set 801 related to the second object and a fourth interactive control 802 can be displayed. A fifth interactive control 803 can be displayed in the fourth interactive control 802. When the user clicks the fifth interactive control 803, a third trigger operation is triggered, using the second image set 801 as reference information to generate a first image set based on the second image set 801.
[0082] It should be understood that the second image set 801 can be displayed in the form of a carousel on the third interface 800. Users can view the series of images in the second image set 801 by swiping.
[0083] Of course, the third interface 800 may also include a fifth interactive control 804, which is the same as the second interactive control 603 in the above embodiment, and will not be described again here.
[0084] In other implementations, the production progress of the first image set can also be displayed on a third interface. Figure 9 This is a schematic diagram illustrating the production progress according to some embodiments. For example... Figure 9 As shown, the production progress 805 can be displayed in the form of a progress bar in the fourth interactive control 802. Of course, Figure 9 The method shown for displaying the production progress is merely an example; other methods can also be used. Displaying the production progress allows users to understand the real-time progress of the first image set, thus improving the user experience.
[0085] It should be understood that this embodiment actually provides a method for generating a series of images related to a first object based on a second image set published by other users that is related to a second object. For example, when a user is satisfied with a set of photos published by other users, the user can use the set of photos published by other users as reference images, thereby quickly generating a first image set belonging to the user based on the set of photos published by other users. Based on this, the user's need to quickly generate their own first image set based on a second image set published by other users that is related to a second object can be met.
[0086] In some feasible implementations, image description information for generating the first image set can be determined based on reference information, and then the first image set can be generated based on the image description information and the digital clone corresponding to the first object.
[0087] Here, after determining the reference information (such as reference images and / or references), image description information can be generated based on the understanding of the reference information. The image description information describes the image content corresponding to a series of images to be generated. Multiple series of images to be generated and their corresponding image description information can be determined through the reference information; each image description information describes the image content corresponding to its respective series of images.
[0088] For example, assuming the reference information represents "a girl wearing a white sweater playing in the snow," image description information corresponding to multiple images in a series can be generated based on this reference information. For instance, the image description information corresponding to the first image in the series could be "a girl wearing a white sweater having a snowball fight in the snow," the image description information corresponding to the second image could be "a girl wearing a white sweater jumping and spreading her arms in the snow," and the image description information corresponding to the third image could be "a girl wearing a white sweater walking in the heavy snow." For example, the generated image description could be: "The first image is a close-up, showing the cat sitting in the center of the frame, looking out at the sea from behind; the second is a medium shot, showing the cat standing sideways on the beach with a small fish in its mouth, its fur disheveled, looking at the camera; the third is a close-up, showing the cat's face on the left side of the frame, looking to the right. The overall tone is cool, with a Fujifilm effect, overexposed, resulting in a rough and cool-toned image. Details in the shadows are well preserved, and highlights exhibit natural bokeh. All images use soft, diffused light, without harsh shadows, creating an artistic atmosphere full of self-exploration."
[0089] For example, reference information can be input into the Visual Large Model (VLM) to obtain image description information output by the Visual Large Model. The Visual Large Model is used to associate multiple series of images to be generated based on the reference information to obtain image description information corresponding to them.
[0090] The visual big data model receives reference information input by the user and, based on this reference information, makes associations to obtain image description information corresponding to a series of images to be generated. In other words, the visual big data model, based on its understanding of the reference information, associates and obtains image description information corresponding to a series of images to be generated, so as to accurately describe the image content corresponding to each series of images to be generated.
[0091] Then, based on the image description information and the digital clone corresponding to the first object, a first image set related to the first object can be generated. The digital clone is used to define the image of the object described by the image description information.
[0092] A digital clone can refer to a virtual image of a primary object. It's a virtual representation of a primary object generated by learning and reproducing its visual characteristics based on an image provided by the user. For example, when the primary object is the user, the digital clone can be understood as the user's digital persona.
[0093] It should be understood that the object described in image description information is generally a broad concept or an object directly referenced in a reference image. For example, in the image description "a girl in a white sweater having a snowball fight in the snow," "girl" is the object described in the image description information. Through digital cloning, the object described in the image description information, "girl," can be defined as the first object corresponding to the digital cloning. For example, assuming the first object is user A, when generating the first image set for user A, multiple images of user A can be generated at once based on the digital cloning and the image description information.
[0094] It's worth noting that digital clones can be pre-built. For example, a user can pre-input multiple images of a first object, and based on these images, a digital clone corresponding to that first object can be generated. Then, when generating the first image set, the pre-built digital clone can be directly invoked.
[0095] For example, image description information and digital clones can be input into an image generation model to obtain a first set of images output by the image generation model. It should be understood that the image description information is actually equivalent to a prompt input into the image generation model, and the image generation model outputs a series of multiple images related to the first object based on the image description information and digital clones.
[0096] Figure 10 This is a logical schematic diagram illustrating the generation of a first image set, based on some embodiments. For example... Figure 10 As shown, reference information can be input into the visual large model, the visual large model outputs image description information based on the reference information, and then the image description information and digital clone are input into the image generation model, the image generation model outputs the first image set based on the image description information and digital clone.
[0097] Therefore, through the above implementation method, the reference information can be understood and image description information can be generated, thereby ensuring the effect of the series of images to be generated. Moreover, through digital cloning, the first image set corresponding to the first object can be accurately generated.
[0098] In some feasible implementations, when the reference information includes a reference image, the object skeleton corresponding to the object included in the reference object can be displayed. Then, in response to the adjustment operation of the target skeletal key points in the object skeleton, the adjusted object skeleton is obtained, and a first image set is generated based on the adjusted object skeleton, image description information, and digital clone.
[0099] Here, the object skeleton includes skeletal keypoints obtained by extracting keypoints from objects included in the reference image. Specifically, the object skeleton is obtained by extracting keypoints from objects in the reference image. For example, skeletal keypoints of objects in the reference image can be extracted, and then the object skeleton can be constructed based on these keypoints. It should be understood that the object skeleton can be the skeleton of a person or animal.
[0100] For example, following the above embodiments, when the object included in the reference image is a second object, the skeletal key points corresponding to the second object can be displayed so that the user can adjust the shooting posture corresponding to the second object.
[0101] After obtaining the object skeleton, it can be displayed. Users can then adjust the key points of the skeleton by targeting the target bones, thus obtaining the adjusted object skeleton.
[0102] It should be understood that the adjustment operation can be a drag-and-drop operation targeting key points of the target skeleton. Users can adjust the key points of the skeleton in the object by dragging the target key points, thereby adjusting the pose of the object corresponding to the skeleton. For example, by adjusting the key points of the skeleton, users can control the object skeleton to present a specific pose.
[0103] Next, the adjusted object skeleton, image description information, and digital clone can be input into the image generation model to obtain the first image set output by the image generation model. The adjusted object skeleton is used at least to adjust the object pose described by the image description information to the object pose corresponding to the object skeleton.
[0104] It is worth noting that the digital clone determines that the appearance features of objects in the generated series of images are the appearance features of the first object, while the adjusted object skeleton determines that the object pose of objects in the generated series of images is the object pose corresponding to the adjusted object skeleton.
[0105] Therefore, by adjusting the object skeleton, users can freely determine the pose of the object in the final series of images, allowing them to freely choose the desired shooting action.
[0106] Figure 11 This is a schematic diagram of a first image set shown according to some embodiments. For example... Figure 11As shown, a first image set 1102 can be generated from a reference image 1101. The first image set 1102 includes four series of images, which are stylistically similar to the reference image 1101. Furthermore, the reference image 1101 is an image related to a second object, and the four series of images in the first image set 1102 are images related to the first object. In other words, user A can generate multiple series of images belonging to user A based on a single photograph of user B.
[0107] In some feasible implementations, in step 120, auxiliary description information input by the user can be obtained and a first image set can be displayed, the first image set being generated based on reference information and auxiliary description information.
[0108] Here, the auxiliary descriptive information can be similar to the reference information, both of which are used to indicate the generation of a series of images. Exemplarily, the auxiliary descriptive information can be in the form of text or images, and is not specifically limited in this embodiment.
[0109] After selecting reference information, the user can further enrich the information provided to the image generation model by adding auxiliary descriptive information. Following the above implementation, the auxiliary descriptive information can be the reference text entered in the text input area as described in the above embodiments. That is, after selecting a reference image as reference information, the user can also input text information as auxiliary descriptive information to determine the reference image and auxiliary descriptive information as the final reference information.
[0110] In some feasible implementations, prompt information corresponding to each series of images can also be displayed, and then, in response to an adjustment operation on the target prompt information, a series of images generated based on the adjusted target prompt information can be displayed.
[0111] Here, the cue words can be generated based on reference information, and these cue words are used to instruct the generation of a series of images. It should be understood that the cue words can refer to information input to the image generation model, instructing the model to generate the required images.
[0112] Following the above embodiments, the prompt information can be a portion of the image description information described in the embodiments. For example, if the image description information corresponding to the first image in the series to be generated is "a girl wearing a white sweater is having a snowball fight in the snow," then the prompt information corresponding to the generated first image in the series is "a girl wearing a white sweater is having a snowball fight in the snow." If the image description information corresponding to the second image in the series to be generated is "a girl wearing a white sweater is jumping and spreading her arms in the snow," then the prompt information corresponding to the generated second image in the series is "a girl wearing a white sweater is jumping and spreading her arms in the snow." If the image description information corresponding to the third image in the series to be generated is "a girl wearing a white sweater is walking in the snow," then the prompt information corresponding to the generated third image in the series is "a girl wearing a white sweater is jumping and spreading her arms in the snow."
[0113] In this embodiment of the disclosure, when displaying the first image set, prompt information corresponding to each series of images can be displayed simultaneously. Alternatively, prompt information corresponding to each series of images can be displayed in response to a user-triggered operation instructing adjustment of the series of images, or prompt information corresponding to the series of images the user needs to adjust.
[0114] It should be understood that the specific display method of the prompt information is not limited in the embodiments disclosed herein. For example, the prompt information corresponding to a series of images can be displayed in the adjacent area of the series of images. Alternatively, the prompt information corresponding to each series of images can be displayed sequentially in a specific area (such as a sidebar). Yet another example is that the prompt information corresponding to each series of images or a series of images that need to be adjusted can be displayed on a single page.
[0115] Users can modify the target prompt information using adjustment options, instructing the user to regenerate the corresponding image series based on the adjusted prompts. These adjustments can be edit operations. For example, prompt information for the image series can be displayed in a text input box, which the user can then edit.
[0116] Accordingly, in response to adjustments made to the target cue information, a series of images generated based on the adjusted target cue information can be displayed.
[0117] Therefore, through the above implementation method, users can adjust the series of images generated based on reference information by modifying the prompt information of at least one series of images, so as to ensure the effect of the generated series of images and greatly improve the user experience.
[0118] In some feasible implementations, a storyline corresponding to multiple series of images can also be displayed, and then, in response to an adjustment operation on the storyline, multiple series of images generated based on the adjusted storyline can be displayed.
[0119] Here, the storyline is determined based on the theme indicated by the reference information, and it guides the content of multiple series of images to unfold around the storyline. In other words, before generating multiple series of images, a corresponding storyline can be generated based on the reference information, and then multiple series of images can be generated around that storyline.
[0120] A storyline can be a coherent narrative thread constructed based on a theme specified in user-provided reference information. Multiple series of images are then organized around this storyline, with the content of each image serving to tell the story, forming a logical and plot-driven visual narrative sequence. It should be understood that the storyline plays a unifying and explanatory role among the multiple images; it indicates that the series of images are not isolated but collectively tell the same story, with each image representing a fragment of that story.
[0121] For example, suppose the theme of the reference information is "a kitten finding its way home in the rain." Based on this reference information, a series of images are generated as follows: "First image: The kitten stands on a street corner in the rain, soaked to the bone; Second image: The kitten takes shelter in a cardboard box; Third image: The kitten encounters a little girl with an umbrella; Fourth image: The kitten is taken home by the little girl and curled up on a warm blanket." Correspondingly, the storyline of these four images could be: "A lost kitten experiences difficulties in the rain but is eventually rescued by a kind person and safely returns home."
[0122] When displaying the first image set, a storyline corresponding to multiple images in a series can be shown. Alternatively, the storyline can be displayed when a user-triggered instruction to adjust the series of images is detected. It should be noted that in this embodiment, the specific display location of the storyline is not limited. The storyline can be displayed within the interface containing the first image set, or it can be displayed in a separate interface.
[0123] The adjustment operation for the storyline can be an editing operation on the storyline. Users can adjust the storyline of multiple series of images through the adjustment operation to instruct the regeneration of multiple series of images.
[0124] For example, the storyline "A lost kitten endures hardships in the rain, but is eventually rescued by a kind person and returns home safely" can be modified by the user to "A lost kitten shivers in the rain, experiencing hiding, hunger, and loneliness, before being found and taken home by a kind little girl, regaining warmth and belonging."
[0125] Of course, adjusting the storyline can also be done by instructing the user to activate the digital assistant. In response to this instruction, the digital assistant will be displayed, allowing the user to adjust the storyline through dialogue. For example, the user could ask the digital assistant, "Please optimize the storyline of 'A lost kitten overcomes difficulties in the rain but is eventually rescued by a kind person and returns home safely.'" The digital assistant will output the optimized storyline. Once the user confirms the optimized storyline is usable, it can be designated as the adjusted storyline.
[0126] It is worth noting that, following the above implementation method, multiple series of images can be generated based on the image description information provided in the above embodiments, or they can be obtained based on the theme association indicated by the reference information, and are used to control the image content of the series of images.
[0127] After obtaining the adjusted storyline, multiple series of images can be regenerated based on the adjusted storyline. It should be understood that image description information corresponding to each series of images to be generated can be generated based on the adjusted storyline, and the corresponding series of images can be generated based on the image description information.
[0128] Therefore, in the above implementation, users can regenerate multiple series of images by modifying the storylines of the generated series of images, so that users can optimize the storytelling between the final obtained series of images and obtain a first image set that better meets their needs.
[0129] In some feasible implementations, the regenerated target series images can also be displayed in response to a generation operation that instructs the regeneration of a target series of images in a series of multiple images.
[0130] Here, when displaying the first image set, the user can select a target image series from multiple images using a selection operation. For example, if the user feels that the effect of a certain image series in the multiple image series is not good enough, the user can select that series of images as the target image series.
[0131] The regenerated target series images are generated based on reference information and the remaining series images from multiple series images. The remaining series images are the series images in the first image set excluding the target series images.
[0132] When displaying the first image set, if the user is not satisfied with some of the series of images in the multiple series of images included in the first image set, the user can use the above generation operation to determine the partial series of images as the target series of images, and then regenerate the target series of images.
[0133] The regenerated target series images are generated based on the adjustment information indicated by the generation operation and / or the remaining series images from multiple series images. Following the above implementation method, the adjustment information and / or the remaining series images can be input into the image generation model to generate a new target series images.
[0134] The generation operation used to instruct the regeneration of a target series of images from a multi-image series can be a fine-tuning operation for the target series of images from the multi-image series. This fine-tuning operation determines adjustment information to instruct adjustments to the target series of images, and then, based on the adjustment information and / or the remaining series of images, a new target series of images is regenerated.
[0135] It should be understood that by combining the remaining series of images to regenerate the target series of images, it can be ensured that the regenerated target series of images maintains stylistic and thematic consistency with the remaining series of images. For example... Figure 11 As shown, if a user is not satisfied with one or more target series images in the first image set 1102, a new target series of images can be generated.
[0136] Fine-tuning a target series of images within a multi-image series refers to inputting adjustment information to instruct adjustments to the target series. When a user selects the target series, a text input box is displayed, allowing the user to enter adjustment information. This adjustment information can be understood as more refined and personalized image descriptions.
[0137] It's worth noting that adjusting the target series of images through the generation operation can essentially be understood as refining the target series of images. After outputting multiple series of images, if the user is not satisfied with a particular image in the target series, it is possible to regenerate the target series of images.
[0138] Therefore, through the above implementation method, users can adjust the generated series of images to ensure the effect of the generated series of images, which greatly improves the user experience.
[0139] Figure 12 This is a schematic diagram of the structure of an image generation apparatus according to some embodiments. For example... Figure 12 As shown, this embodiment of the disclosure provides an image generation apparatus 1200, which includes: Module 1201 is configured to acquire reference information. The display module 1202 is configured to display a first image set, which includes multiple series of images related to a first object generated based on the reference information. The visual style of the multiple series of images is similar to the visual style corresponding to the reference information, and the image content of the multiple series of images revolves around the theme indicated by the reference information.
[0140] Optionally, the reference information includes reference images and / or reference text.
[0141] Optionally, when the reference information includes the reference image, the reference image includes an image related to the first object, an image related to the second object, or a second set of images related to the second object.
[0142] Optionally, the acquisition module 1201 is specifically configured as follows: If the reference image includes an image related to the first object, the image related to the first object is displayed in the first interface; In response to a first triggering operation triggered on the first interface, an image associated with the first object is determined as the reference information, the first triggering operation being used to instruct the generation of the first image set based on the image associated with the first object; or If the reference image includes an image related to the second object, the image related to the second object is displayed in the second interface; In response to a second triggering operation triggered on the second interface, an image associated with the second object is determined as the reference information, the second triggering operation being used to instruct the generation of the first image set based on the image associated with the second object; or If the reference image includes the second image set, the second image set is displayed in the third interface; In response to a third triggering operation triggered on the third interface, the second image set is determined as the reference information, the third triggering operation being used to instruct the generation of the first image set based on the second image set.
[0143] Optionally, the acquisition module 1201 is specifically configured as follows: In response to the first triggering operation, a fourth interface is displayed, the fourth interface including an image related to the first object and a text input area for obtaining reference text; In response to a fourth trigger operation triggered on the fourth interface, the reference text entered in the text input area and the image associated with the first object are determined as the reference information.
[0144] Optionally, the first image set further includes a cover image obtained by stitching together the multiple series of images; the display module 1202 is specifically configured as follows: The cover image is displayed, followed by a series of images displayed sequentially.
[0145] Optionally, the display module 1202 is further configured to: Obtain auxiliary description information input by the user; A first image set is displayed, which is generated based on the reference information and the auxiliary description information.
[0146] Optionally, the image generation apparatus 1200 further includes: The prompt word display unit is configured to display prompt word information corresponding to the series of images. The prompt word information is generated based on the reference information and is used to indicate the generation of the series of images. The first adjustment unit is configured to display a series of images generated based on the adjusted target cue information in response to an adjustment operation on the target cue information.
[0147] Optionally, the image generation apparatus 1200 further includes: The storyline display unit is configured to display the storyline corresponding to the multiple series of images; The second adjustment unit is configured to display a series of images generated based on the adjusted storyline in response to the adjustment operation for the storyline.
[0148] Optionally, the image generation apparatus 1200 further includes: The third adjustment unit is configured to display the regenerated target series images in response to a generation operation that instructs the regeneration of a target series image in the plurality of series images, the regenerated target series images being generated based on adjustment information indicated by the generation operation and / or the remaining series images in the plurality of series images.
[0149] Optionally, the image generation apparatus 1200 further includes: The determining module is configured to determine image description information for generating a first image set based on the reference information, wherein the image description information is used to describe the image content corresponding to a series of images to be generated. The generation module is configured to generate the first image set based on the image description information and the digital clone corresponding to the first object, wherein the digital clone is used to define the image of the object described by the image description information.
[0150] Optionally, the generation module is specifically configured as follows: When the reference information includes a reference image, the object skeleton corresponding to the object included in the reference object is displayed. The object skeleton includes skeletal key points obtained by extracting key points from the object in the reference image. In response to an adjustment operation on the target skeletal key points included in the object skeleton, an adjusted object skeleton is obtained; Based on the adjusted object skeleton, the image description information, and the digital clone, the first image set is generated. The adjusted object skeleton is used at least to adjust the object pose described by the image description information to the object pose corresponding to the object skeleton.
[0151] The functional logic executed by each functional module in the image generation apparatus 1200 has been described in detail in the section on methods, and will not be repeated here.
[0152] The following is for reference. Figure 13 The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 1300 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 13 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0153] like Figure 13 As shown, electronic device 1300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1301, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1302 or a program loaded from storage device 1308 into random access memory (RAM) 1303. The RAM 1303 also stores various programs and data required for the operation of electronic device 1300. The processing device 1301, ROM 1302, and RAM 1303 are interconnected via bus 1304. Input / output (I / O) interface 1305 is also connected to bus 1304.
[0154] Typically, the following devices can be connected to I / O interface 1305: input devices 1306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1309. Communication device 1309 allows electronic device 1300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 13 An electronic device 1300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0155] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1309, or installed from storage device 1308, or installed from ROM 1302. When the computer program is executed by processing device 1301, it performs the functions defined in the methods of embodiments of this disclosure.
[0156] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0157] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0158] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0159] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire reference information; and display a first image set, the first image set comprising multiple series of images related to a first object generated based on the reference information, the visual style corresponding to the multiple series of images being similar to the visual style corresponding to the reference information, and the image content of the multiple series of images revolving around the theme indicated by the reference information.
[0160] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0162] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the module itself.
[0163] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0164] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0165] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0166] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0167] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. An image generation method, characterized in that, include: Obtain reference information; A first image set is displayed, comprising multiple series of images related to a first object generated based on the reference information. The visual style of the multiple series of images is similar to the visual style corresponding to the reference information, and the image content of the multiple series of images revolves around the theme indicated by the reference information.
2. The method according to claim 1, characterized in that, The reference information includes reference images and / or reference text.
3. The method according to claim 2, characterized in that, When the reference information includes the reference image, the reference image includes an image related to the first object, an image related to the second object, or a second set of images related to the second object.
4. The method according to claim 1, characterized in that, The first image set also includes a cover image obtained by stitching together the multiple series of images; displaying the first image set includes: The cover image is displayed, followed by a series of images displayed sequentially.
5. The method according to any one of claims 1-4, characterized in that, The display of the first image set includes: Obtain auxiliary description information input by the user; A first image set is displayed, which is generated based on the reference information and the auxiliary description information.
6. The method according to any one of claims 1-4, characterized in that, The method further includes: Display prompt information corresponding to the series of images. The prompt information is generated based on the reference information and is used to indicate the generation of the series of images. In response to adjustments made to the target cue word information, a series of images generated based on the adjusted target cue word information are displayed.
7. The method according to any one of claims 1-4, characterized in that, The method further includes: Show the storyline corresponding to the multiple series of images; In response to the adjustment operation for the storyline, a series of images generated based on the adjusted storyline are displayed.
8. The method according to any one of claims 1-4, characterized in that, The method further includes: In response to a generation operation that instructs the regeneration of a target series of images from the plurality of series of images, the regenerated target series of images is displayed, the regenerated target series of images being generated based on adjustment information indicated by the generation operation and / or the remaining series of images from the plurality of series of images.
9. The method according to any one of claims 1-4, characterized in that, The first image set was obtained through the following steps: Based on the reference information, image description information for generating the first image set is determined, wherein the image description information is used to describe the image content corresponding to a series of images to be generated; Based on the image description information and the digital clone corresponding to the first object, the first image set is generated, and the digital clone is used to define the image of the object described by the image description information.
10. The method according to claim 9, characterized in that, The step of generating the first image set based on the image description information and the digital clone corresponding to the first object includes: When the reference information includes a reference image, the object skeleton corresponding to the object included in the reference object is displayed. The object skeleton includes skeletal key points obtained by extracting key points from the object included in the reference image. In response to an adjustment operation on the target skeletal key points in the object skeleton, an adjusted object skeleton is obtained; based on the adjusted object skeleton, the image description information, and the digital clone, a first image set is generated, wherein the adjusted object skeleton is at least used to adjust the object pose described by the image description information to the object pose corresponding to the object skeleton.
11. An image generation apparatus, characterized in that, include: The acquisition module is configured to acquire reference information; The display module is configured to display a first image set, which includes multiple series of images related to a first object generated based on the reference information. The visual style of the multiple series of images is similar to the visual style corresponding to the reference information, and the image content of the multiple series of images revolves around the theme indicated by the reference information.
12. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processing device, it implements the steps of the method according to any one of claims 1-10.
13. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-10.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-10.