DYNAMIC STORYBOARD CREATION WITH ONE CLICK USING TEXT INSTRUCTIONS

The system uses a speech and zero-shot image generation model to efficiently and accurately generate storyboards from text prompts, addressing computational inefficiencies and incoherent outputs in existing technologies, allowing for flexible image element modification.

DE102025131507A1Pending Publication Date: 2026-04-16ADOBE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Existing image generation models for storyboard creation are computationally intensive, time-consuming, and struggle to accurately generate high-resolution synthetic images for complex tasks, often resulting in visually incoherent output and requiring manual editing, with generalizability affected by training dataset quality and variety.

Method used

A system utilizing a speech generation model and a zero-shot image generation model to produce scene prompts and synthetic images from a text prompt, ensuring character identity preservation and narrative consistency, with the ability to modify image elements independently or collectively.

Benefits of technology

Efficient and accurate storyboard generation from a single text prompt, maintaining narrative coherence and enabling flexible modification of image elements, reducing the need for manual editing and improving system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method, a device, a non-transient computer-readable medium, and an image processing system comprise capturing a text prompt describing a story, generating a first scene prompt and a second scene prompt based on the text prompt, wherein the first scene prompt describes a first scene of the story and the second scene prompt describes a second scene of the story, and generating a first synthetic image and a second synthetic image based on the first scene prompt and the second scene prompt, wherein the first synthetic image represents the first scene and the second synthetic image represents the second scene.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] The following generally concerns image processing and, in particular, storyboard creation using a machine learning model. Image processing refers to the use of a computer to manipulate an image using an algorithm or processing network. In some cases, image processing software can be used for various image processing tasks, such as image restoration, image recognition, image composition, image editing, image generation, and storyboard creation. For example, storyboard creation involves using the machine learning model to generate a series of images arranged in a sequence to depict the events of a story described by the text prompt.

[0002] A storyboard is a visual representation of a narrative. For example, a storyboard comprises a series of images arranged in a sequence to depict the key scenes, actions, and events of a story. In some cases, each panel or image of the storyboard represents a specific scene from the story, and each panel may include a caption summarizing the actions depicted within that panel. In some cases, one or more elements of an image within the storyboard can be changed by modifying the text prompt. SUMMARY

[0003] Aspects of this revelation provide a method and system for storyboard creation. In one aspect, the system receives a text prompt describing a story and generates a storyboard based on the text prompt. In another aspect, the storyboard includes multiple synthetic images representing elements of the story, as well as multiple titles, each associated with the synthetic images. According to some aspects, the system includes a speech generation model and an image generation model. In one aspect, the speech generation model receives the text prompt and generates a series of scene prompts, each describing a scene of the story. The image generation model is configured to receive the series of scene prompts and generate a series of synthetic images, each representing the scenes and elements described by the series of scene prompts.In some embodiments, the image generation model is configured to receive an image prompt in order to generate the series of synthetic images. In some cases, the identity of a character depicted in the synthetic images is the same as the character depicted in the image prompt. In some aspects, a storyboard component is configured to receive the series of synthetic images and a set of captions, each corresponding to one of the synthetic images, in order to generate the storyboard.

[0004] A method, a device, a non-transitory computer-readable medium, and an image processing system comprise capturing a text prompt describing a story, generating—using a language generation model—a first scene prompt and a second scene prompt based on the text prompt, wherein the first scene prompt describes a first scene of the story and the second scene prompt describes a second scene of the story, and generating—using an image generation model—a first synthetic image and a second synthetic image based on the first scene prompt and the second scene prompt, wherein the first synthetic image represents the first scene and the second synthetic image represents the second scene.

[0005] A method, apparatus, non-transitory computer-readable medium and image processing system comprising capturing a text prompt and an image prompt; generating – using a speech generation model – a first scene prompt and a second scene prompt based on a text prompt; and generating – using an image generation model – a first synthetic image and a second synthetic image based on the image prompt and based on the first scene prompt and the second scene prompt.

[0006] An image processing device and system comprise a storage unit and a processing device coupled to the storage unit, the processing device being configured to perform operations that include capturing a text prompt and an image prompt, wherein the text prompt contains a story and an image prompt, respectively.The story describes and the image prompt represents an element of the story; the generation of a first scene prompt and a second scene prompt based on the text prompt, wherein the first scene prompt describes a first scene of the story and the second scene prompt describes a second scene of the story; and the generation of a first synthetic image and a second synthetic image based on the image prompt, the first scene prompt and the second scene prompts, wherein the first synthetic image represents the first scene including the element and the second synthetic image represents the second scene including the element. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 shows an example of an image processing system according to aspects of the present disclosure. Fig. Figure 2 shows an example of a method for conditional image generation according to aspects of the present disclosure. Fig. Figure 3 shows an example of storyboard creation according to aspects of the present revelation. Fig. Figure 4 shows an example of a procedure for generating a storyboard using a text prompt according to aspects of the present revelation. Fig. Figure 5 shows an example of an image processing device according to aspects of the present disclosure. Fig. Figure 6 shows an example of a storyboard generation system according to aspects of the present revelation. Fig. Figure 7 shows an example of an image generation model according to aspects of the present revelation. Fig. Figure 8 shows an example of a U-Net architecture according to aspects of the present disclosure. Fig. Figure 9 shows an example of a diffusion process according to aspects of the present revelation. Fig. Figure 10 shows an example of a procedure for modifying a scene prompt according to aspects of the present revelation. Fig. Figure 11 shows an example of a flowchart that presents an algorithm as a step-by-step procedure in an example implementation of executable operations for training a machine learning model according to aspects of the present disclosure. Fig. Figure 12 shows an example of a procedure for training a diffusion model according to aspects of the present disclosure. Fig. Figure 13 shows an example of a computing device according to aspects of the present disclosure. DETAILED DESCRIPTION

[0007] The following concerns image processing, and in particular storyboard creation using generative machine learning. Embodiments of the disclosure relate to a storyboard generation system that efficiently and accurately generates a storyboard from an input text prompt describing a story. In one aspect, the system includes a language generation model configured to generate a set of scene prompts describing scenes of the story. In another aspect, the system includes an image generation model configured to receive an image prompt showing a character and the set of scene prompts to generate a set of synthetic images representing the scenes and the character. By using the image prompt to guide the image generation model, the system ensures accurate image generation that corresponds to the textual description of the scene prompts.

[0008] In some embodiments, the system includes a speech generation model configured to produce a series of scene prompts, each describing a scene of the story based on the text prompt. In some embodiments, the system includes an image generation model configured to produce a series of synthetic images based on the series of scene prompts. In some embodiments, the series of synthetic images is generated based on an image prompt depicting a character. In some cases, the image prompt is generated, for example, based on an identity text prompt using the image generation model. In some cases, the image embedding of the image prompt is combined with each of the text embeddings of each individual scene prompt to guide the image generation process in producing the synthetic images.

[0009] In some aspects, the system includes a storyboard component configured to generate a storyboard with a set of panels. For example, a set of captions, each corresponding to a set of synthetic images, and the set of synthetic images themselves are input into the storyboard component. In other cases, the storyboard comprises a set of panels, each containing a synthetic image and a corresponding caption. In some cases, the panels are arranged sequentially based on the story.

[0010] A subfield of image processing involves storyboarding. For example, creating storyboards is an essential step in the production of animated graphics, films, and other digital media for storytelling. Traditional storytelling is performed by a human screenwriter who creates a story script, and a human artist then transforms the storyboard into story panels that represent the style of the image. In some cases, the storyboarding process involves several skilled artists in the fields of screenwriting and visual design. In some instances, storyboarding can extend over several weeks of work and iterations.

[0011] Some established storyboard creation methods involve the use of machine learning models. However, generating storyboards using established systems with deep learning architectures such as a Transformer or a Convolutional Neural Network (CNN) can be computationally intensive and time-consuming. These systems may struggle, for example, to generate high-resolution synthetic images for complex tasks like creating detailed scene compositions. In some cases, these systems may result in delayed user feedback and reduce overall system efficiency.

[0012] Some systems known from the state of the art cannot accurately understand the complex scene descriptions of a text prompt and generate the corresponding synthetic image during text-to-image conversion. For example, these systems may produce inaccurate or ambiguous images or pixels when receiving abstract text instructions. Consequently, these systems can lead to visually incoherent output. This, in turn, can result in additional and undesirable manual editing by the user.

[0013] In some cases, the performance of these systems depends heavily on the quality and variety of the training data. For example, if a model is trained on a limited or biased dataset, its output may be less diverse and representative. Similarly, if the model is trained on a specific genre, it may struggle to generate storyboards for different styles. Accordingly, the generalizability of these state-of-the-art systems can be affected by the training dataset.

[0014] Implementations of the revelation improve upon prior art image generation models by generating a storyboard more efficiently and accurately based on an input text prompt describing a story. This is achieved using a system comprising a speech generation model and an image generation model (for example, a zero-shot image generation model). In one aspect, the speech generation model is configured to produce a set of scene prompts describing one or more scenes, actions, and / or events of the story. This set of scene prompts is provided to the zero-shot image generation model to ensure diverse image generation while maintaining the narrative (or sequence of events) of the story.In one aspect, the image generation model takes an image prompt as input to ensure that the identity of the character described in the story is preserved and remains consistent across the frames of the storyboard.

[0015] In one aspect, the image generation model is a zero-shot image generation model. For example, the zero-shot image generation model generates a synthetic image based on one or more input prompts (for example, a text prompt or an image prompt) without specific training on the input prompt. In another aspect, the model can generate images from text prompts that the model has not previously seen explicitly. For example, if the model has never been trained on the phrase "a red panda playing guitar," it can still generate an accurate image based on its understanding of "red panda" and "guitar." In some cases, the model is pre-trained with diverse data in the shared latent space between text and image. Consequently, the model can generalize across different combinations of objects, actions, and styles, thereby generating new visual concepts based on the text description.

[0016] In some aspects, the image embedding of the image prompt and each of the text embeddings of the scene prompts are combined or chained together. This allows the image generation model to generate exactly one or more synthetic images that preserve the identity of the character depicted in the image prompt while simultaneously producing accurate image content that corresponds to the scenes described by the scene prompts. In some embodiments, a modified scene prompt can be obtained based on a user modification command. For example, the modified scene prompt can describe a change to an image element. By using the modified scene prompt, the image generation model can generate a modified synthetic image that represents the change to the image element.

[0017] An exemplary system of the invention concept in image processing is described in relation to the Fig. 1 and Fig. 13. An example application of the invention concept in image processing is described in relation to the Fig. 2-3 described. Details of the architecture of an image processing device are given in relation to the Fig. 5-8 described. An example of an image processing method is given in relation to the Fig. Sections 4 and 9-10 describe an exemplary training process. Fig. 11-12 provided.

[0018] Accordingly, the present disclosure provides a system and a method that improves upon prior art systems by efficiently and accurately generating a storyboard with a set of story panels from a single input text prompt describing a story. In some embodiments, the system receives an image prompt depicting a character to generate the synthetic images of the storyboard. By guiding the system's image generation model with the image prompt, the character depicted in the image prompt can be integrated into one or more synthetic images in the story panels. In some aspects, one or more image elements of the synthetic images in the story panels can be modified independently or collectively by manipulating one or more corresponding elements of the scene prompts generated by the speech generation model. Storyboard generation

[0019] As in the Fig. Figures 1-4 and 9-10 comprise a method, a device, a non-transitory computer-readable medium, and an image processing system for capturing a text prompt describing a story; generating—using a language generation model—a first scene prompt and a second scene prompt based on the text prompt, wherein the first scene prompt describes a first scene of the story and the second scene prompt describes a second scene of the story; and generating—using an image generation model—a first synthetic image and a second synthetic image based on the first scene prompt and the second scene prompt, wherein the first synthetic image represents the first scene and the second synthetic image represents the second scene.

[0020] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include capturing an image prompt depicting an element of the story, with the first synthetic image and the second synthetic image being generated based on the image prompt and representing the element. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include capturing an identity text prompt describing the element. Some examples further include generating the image prompt based on the identity text prompt. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include capturing a preliminary image of a person. Some examples further include cropping the preliminary image to obtain the image prompt, with the element comprising a face of the person.

[0021] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include capturing a style image that displays a style, wherein the first synthetic image and the second synthetic image are generated based on the style image and contain the style. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include capturing a first noise input and a second noise input. Some examples further include denoising the first noise input based on the first scene prompt and the second noise input based on the second scene prompt to obtain the first synthetic image and the second synthetic image.

[0022] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include generating a storyboard using the first synthetic image and the second synthetic image. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include generating—using the language generation model—a first caption and a second caption based on the first synthetic image and the second synthetic image, wherein the storyboard comprises a first panel containing the first synthetic image and the first caption, and a second panel containing the second synthetic image and the second caption.

[0023] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include receiving a modification command specifying the initial scene and a modified element. Some examples further include generating—using the speech generation model—a modified scene prompt based on the modification command. Some examples further include generating—using the image generation model—a modified synthetic image based on the initial scene and the modified element.

[0024] Fig. Figure 1 shows an example of an image processing system according to aspects of the present disclosure. The example shown includes user 100, user device 105, image processing device 110, cloud 115, database 120, and display panel 125. In some aspects, the user device 105 includes the display panel 125. The image processing device 110 is an example of the corresponding element that, with respect to Fig. 5 is described.

[0025] With reference to Fig. In one embodiment, the user 100 provides a text prompt to the image processing device 110 via the user device 105 and the cloud 115. In some cases, the text prompt may be a description of a story theme. For example, the text prompt reads "Blueberry's Interdimensional Adventure." In other embodiments, the user 100 provides an image prompt (for example, a real image or a synthetic image) to the image processing device 110 via the user device 105 and the cloud 115. For example, the image prompt shows a character (for example, Blueberry) to be generated in one or more story panels of the storyboard. In some cases, the image prompt is generated based on an identity text prompt using an image generation model.

[0026] In some embodiments, the image processing device 110 includes a speech generation model configured to generate a set of scene prompts based on the text prompt. In some embodiments, the image processing device 110 includes an image generation model configured to receive the set of scene prompts and the image prompt to generate a set of synthetic images. For example, one or more of the synthetic images represent the character from the image prompt in the scene described by the corresponding scene prompt. In some embodiments, the image processing device 110 includes a storyboard component configured to combine each of the synthetic images and each of the image titles corresponding to the synthetic images to generate the storyboard.The image processing device 110 displays the storyboard via the display panel 125 of the user device 105 to the user 100 via the cloud 115.

[0027] The user device 105 may be a personal computer, a laptop computer, a mainframe computer, a palmtop computer, a personal assistant, a mobile device, or any other suitable processing device. In some examples, the user device 105 includes software that contains an image processing application. In some examples, the image processing application on the user device 105 may include functions of the image processing device 110. In some cases, the user device 105 may include a user interface that performs functions of the image processing device 110.

[0028] A user interface can enable the user 100 to interact with the user device 105. In some embodiments, the user interface can include an audio device such as an external speaker system, an external display device such as a screen, or an input device (for example, a remotely controllable device connected to the user interface directly or via an I / O control module). In some cases, a user interface can be a graphical user interface (GUI). In some examples, a user interface can be represented in code, with the code being sent to the user device 105 and rendered locally by a browser. The process of using the image processing device 110 is further described with reference to Fig. 2 described.

[0029] The image processing device 110 is an example of, or comprises, aspects of the corresponding element which relates to Fig. 5 is described. According to some aspects, the image processing device 110 comprises a computer-implemented network that includes a machine learning model, a speech generation model, an image generation model, and a storyboard component. The image processing device 110 further comprises a processing unit, a storage unit, an I / O module, a user interface, and a training component. In some embodiments, the image processing device 110 further comprises a communication interface, user interface components, and a bus, as described in relation to Fig. 13. Additionally or alternatively, the image processing device 110 communicates with the user device 105 and the database 120 via the cloud 115. Further details on the operation of the image processing device 110 are described in relation to Fig. 2 described.

[0030] In some cases, the image processing device 110 is implemented on a server. A server provides one or more functions to one or more users over one or more of the various networks. In some cases, the server comprises a single microprocessor board containing a microprocessor responsible for controlling aspects of the server. In some cases, a server uses the microprocessor and protocols to exchange data with other devices / users over one or more of the networks using Hypertext Transfer Protocol (HTTP) and Simple Mail Transfer Protocol (SMTP), although other protocols such as File Transfer Protocol (FTP) and Simple Network Management Protocol (SNMP) may also be used. In some cases, a server is configured to send and receive Hypertext Markup Language (HTML)-formatted files.

[0031] Cloud 115 is a computer network configured to provide on-demand availability of computing resources such as data storage and processing power. In some examples, Cloud 115 provides resources without active management by the user (for example, User 100). The term "cloud" is sometimes used to describe data centers that are accessible to many users via the internet. Some large cloud networks have capabilities distributed across multiple locations of central servers. A server is referred to as an edge server if it has a direct or close connection to a user. In some cases, Cloud 115 is limited to a single organization. In other examples, Cloud 115 is available to many organizations. In one example, Cloud 115 comprises a multi-layered communications network that includes multiple edge routers and core routers.In some examples, Cloud 115 is based on a local arrangement of switches at a single physical location.

[0032] According to some aspects, database 120 stores training data. Database 120 is an organized collection of data. For example, database 120 stores data in a defined format called a schema. Database 120 can be structured as a single database, a distributed database, multiple distributed databases, or a disaster recovery backup database. In some cases, a database controller can manage data storage and processing in database 120. In some cases, a user (for example, user 100) interacts with the database controller. In other cases, the database controller can operate automatically without user interaction.

[0033] Fig. Figure 2 shows an example of a method 200 for conditional image generation according to aspects of the present disclosure. In some examples, these operations are performed by a system comprising a processor that executes a set of codes to control functional elements of a device. Additionally or alternatively, certain processes are performed using specialized hardware. Generally, these operations are performed according to the methods and procedures described in accordance with aspects of the present disclosure. In some cases, the operations described herein consist of various substeps or are performed in conjunction with other operations.

[0034] In step 205, the system provides a text prompt. In some cases, the operations of this step relate to a user or can be performed by a user, as in relation to Fig. 1. In some cases, the user provides the text prompt to the system's language generation model. In some aspects, the text prompt describes a high-level overview of the story. For example, the text prompt might read "Blueberry's Interdimensional Adventure." In some cases, the user can provide additional input, such as the number of panels to be generated in a storyboard. In some implementations, the system performs prompt engineering to generate an input text prompt. In some cases, the input text prompt might include the text prompt with the number of panels appended to it. For example, the input text prompt might read: "Create a storyboard of {story} with {n} panels," where {story} specifies the text prompt and {n} specifies the number of panels to be generated in the storyboard. In some implementations, the system can receive additional text input from the user describing one or more elements of the story. In some implementations, the user provides an image prompt showing a character from the story to preserve the identity of the character generated in the storyboard. Further details regarding identity preservation are discussed in relation to Fig. 6 described.

[0035] In step 210, the system generates a conditional guidance embedding. In some cases, the operations of this step relate to or can be performed by an image processing device, as in the case of Fig. 1 and Fig. 5 described. In some cases, the operations of this step relate to or can be performed by an image generation model, as described in relation to Fig. 5 and Fig. 6. In some embodiments, the speech generation model is configured to generate a set of scene prompts based on the input text prompt, the number of scene prompts corresponding to the number specified in the input text prompt. In some aspects, the image generation model includes a text encoder, an image encoder, a multimodal encoder, a prior model, or a combination thereof. In some embodiments, the image generation model includes a text encoder configured to encode the set of scene prompts to generate a set of scene prompt embeddings. In some embodiments, an image encoder or a prior model of the image generation model generates a set of image embeddings based on the set of scene prompt embeddings, the set of image embeddings being used to generate a set of synthetic images.

[0036] In some embodiments, the image generation model receives an image prompt representing a character. In some cases, the image prompt is generated based on an identity text prompt. In some cases, the image prompt is a real image. In some embodiments, the image encoder of the image generation model encodes the image prompt to generate an identity image embed. In one embodiment, the identity image embed is combined or chained with each of the set of image embeds of the set of scene prompts. Accordingly, the system can ensure the preservation of the character's identity in the synthetic images.

[0037] In step 215, the system initializes the noise input. In some cases, the operations of this step relate to, or can be performed by, an image processing device, as in relation to Fig. 1 and Fig. 5 described. In some cases, the operations of this step relate to or can be performed by an image generation model, as described in relation to Fig. Described in sections 5-7. In some cases, the noise input, which contains random noise, is initialized. The noise input can reside in a latent space. By initializing the image generation model with random noise, different variations of a synthetic image can be generated that contain the content described by the text conditioning (for example, the text prompt).

[0038] In step 220, the system generates media content. In some cases, the operations of this step relate to, or can be performed by, an image processing device, as in the case of Fig. 1 and Fig. 5 described. In some cases, the operations of this step relate to or can be performed by an image generation model, as described in relation to Fig. 5 and Fig. 6. In some cases, the media content includes a storyboard. In some cases, the storyboard includes a set of synthetic images and a set of captions. For example, the synthetic image represents the scene and elements described by the scene prompts. For example, captions describe each of the synthetic images at a high level. In some cases, a synthetic image includes image pixels generated by the image generation model.

[0039] Fig. Figure 3 shows an example of generating the storyboard 325 according to aspects of the present disclosure. The example shown includes a storyboard generation system 300, a text prompt 305, an image prompt 310, a style prompt 315, a machine learning model 320, and a storyboard 325. In some embodiments, the storyboard generation system 300 is implemented in a user interface.

[0040] With reference to Fig. 3. The storyboard generation system 300 receives the text prompt 305 and generates the storyboard 325. In some embodiments, the storyboard generation system 300 receives the text prompt 305 and the image prompt 310 as inputs to generate the storyboard 325. In some embodiments, the storyboard generation system 300 receives the text prompt 305, the image prompt 310, and the style prompt 315 as inputs to generate the storyboard 325. In some aspects, the storyboard 325 comprises a set of panels, each panel containing a synthetic image and an image title corresponding to the synthetic image.

[0041] In some embodiments, the machine learning model 320 receives the text prompt 305 as input. For example, the text prompt 305 describes a story such as "Blueberry's Interdimensional Adventure." In some aspects, the machine learning model 320 includes a speech generation model configured to generate a series of scene prompts based on the text prompt 305. For example, each scene prompt describes a scene, event, and / or action of the story. For example, the first scene prompt of the text prompt 305 might read: "Blueberry is inspired by a book that depicts the outside world. Blueberry is in a room and is ready to go outside. Blueberry wants to go out into the world for an adventure."“For example, a middle scene prompt (for example, the second scene prompt) of text prompt 305 might be: “Blueberry arrives in the human world and is overwhelmed by the things Blueberry has never seen before. Blueberry moves his eyes around to take in the magnificent scenes.” For example, a final scene prompt (for example, the third scene prompt) of the text prompt might be: “Blueberry returns to his world and tells his companions what he has seen. Blueberry plays guitar around the campfire and is surrounded by his companions.”

[0042] In some aspects, the Machine Learning Model 320 includes an image generation model configured to receive a series of scene prompts and generate a set of synthetic images corresponding to that series. In some cases, for example, the number of scene prompts and the number of synthetic images are equal. In some aspects, the synthetic image represents elements described by the scene prompts. In some embodiments, each synthetic image is assigned a caption (or image title) (for example, by a user). In some cases, the language generation model of the Machine Learning Model 320 generates a set of captions corresponding to the set of synthetic images.The machine learning model 320 combines each of the synthetic images and each of the captions to generate a set of panels (or storyboard panels) to create storyboard 325. Further details on the image generation model are provided in relation to [reference missing]. Fig. 5 and Fig. 6 described.

[0043] According to some embodiments, the machine learning model 320 receives the text prompt 305 and the image prompt 310 to generate the storyboard 325. For example, the image prompt 310 can be a real image or a synthetic image showing the character described in the story. In some embodiments, the image prompt 310 is generated using a text-to-image generation model or the image generation model of the machine learning model 320 based on an identity text prompt. For example, the identity text prompt describes the character in the story. In some cases, the image prompt 310 is a cropped image showing the character's face. The image generation model of the machine learning model 320 can extract identity information about the character based on the image prompt 310.For example, an image encoder of the image generation model generates an image embed based on image prompt 310, and the set of synthetic images is generated based on this image embed. Accordingly, the set of synthetic images represents the character from image prompt 310 in a scene described by the scene prompts, with each synthetic image having a consistent identity for the character. In some aspects, storyboard 325 includes the set of synthetic images and the corresponding set of captions.

[0044] In some embodiments, the machine learning model 320 receives the text prompt 305, the image prompt 310, and the style prompt 315 to generate the storyboard 325. In some cases, the style prompt 315 is an image displaying a specific style, such as a color style, a texture style, or a picture style. For example, the style prompt 315 might show a red bicycle against a black and white background. The image generation model of the machine learning model 320 can extract the style information based on the style prompt 315. For example, an image encoder of the image generation model generates an image embed based on the style prompt 315, and the set of synthetic images is generated based on this image embed. In some cases, the image embedding of style prompt 315 is combined (for example, chained) with the image embedding of image prompt 310, and the combined image embedding is entered into the image generation model.Accordingly, the set of synthetic images shows, for example, the color style from style prompt 315. For instance, the character depicted in the synthetic images is red, and the background scene is black and white. In some aspects, storyboard 325 includes the set of synthetic images and the corresponding set of captions.

[0045] Text prompt 605 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 3 and Fig. 7 is described. The language generation model 610 is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 5 is described. The image generation model 625 is an example of, or includes, aspects of the corresponding element, which relates to Fig. 5 is described.

[0046] Fig. Figure 4 shows an example of a Method 400 for generating a storyboard using a text prompt according to aspects of this disclosure. In some examples, these operations are performed by a system that includes a processor executing a set of codes to control functional elements of a device. Additionally or alternatively, certain processes are performed using specialized hardware. Generally, these operations are performed according to the methods and procedures described in accordance with aspects of this disclosure. In some cases, the operations described herein consist of various substeps or are performed in conjunction with other operations.

[0047] In step 405, the system captures a text prompt that describes a story. In some cases, the operations of this step relate to, or can be performed by, a language generation model, as in the case of... Fig. 5 and Fig. Section 6 describes this process. In some cases, the text prompt describes one or more image elements to be generated in a synthetic image. For example, an image element is an image component or feature that makes up the overall composition of an image, such as an object, entity, subject, shape, color, texture, pattern, background scene, visual attributes, and / or style. For example, the image element could be an animal like a cat or dog, a person, an object like a hat or table, a scene like a beach or mountaintop, or a combination thereof.

[0048] In some cases, for example, the story is a structured sequence of one or more events, actions, or scenes in chronological order. In some cases, an event refers to a key experience or a change in the state of the world or the characters. For example, an event might refer to a discovery, a conflict, a resolution, and so on. In some cases, an action refers to a decision or behavior by a character that drives the plot forward. For example, an action might consist of the character carrying out or doing something. In some cases, a scene refers to a specific setting or moment in which an event or action takes place, such as a location or an interaction between a character and the location, or between one character and another.In some cases, a character can be, for example, a fictional person, an animal, an object, an entity, or a real person.

[0049] In step 410, the system generates a first scene prompt and a second scene prompt based on the text prompt, where the first scene prompt describes the first scene of the story and the second scene prompt describes the second scene of the story. In some cases, the operations of this step relate to or can be performed by a language generation model, as in relation to Fig. 5 and Fig. 6. In some cases, a scene (for example, the first scene or the second scene) refers to an element of the story (for example, an event, an action, a scene, or a combination thereof). In some cases, the scene prompts may be a subtopic or a smaller unit within the story. In some cases, the scene represents a specific moment, event, or interaction that takes place at a particular time and place in the story. In some cases, one or more scenes are connected to the story. In some cases, the story encompasses the entire sequence of events (or scenes).

[0050] In step 415, the system generates a first synthetic image and a second synthetic image based on the first scene prompt and the second scene prompt, respectively, where the first synthetic image represents the first scene and the second synthetic image represents the second scene. In some cases, the operations of this step are related to, or can be performed by, an image generation model, as in the case of Fig. 5 and Fig. 6. In some cases, the system generates a multitude of synthetic images based on a multitude of scene prompts. For example, the synthetic image comprises image pixels generated by the image generation model.

[0051] In some cases, the system generates a multitude of text embeddings based on the multitude of scene prompts (for example, the first scene prompt and the second scene prompt). In some cases, a text embedding is a numerical vector that captures the semantic meaning of the text by encoding words, phrases, or sentences into a dense, continuous space. For example, the text embedding is encoded into a text embedding space, which is a low-dimensional vector space. The text embedding is generated by passing the text prompt through an encoder (for example, a text encoder or a multimodal encoder) that learns the relationships between words based on the context in large text corpora. In some cases, the text embedding represents textual features (for example, the semantic meaning, the relationship between words, or lexical features) of the text prompt.

[0052] In some cases, the system receives an image prompt and generates an image embed based on that prompt. For example, the image embed is a numerical (or vector) representation of an image in a high-dimensional vector space. For instance, the image embed captures the essential visual features or properties of an image, such as color, texture, shape, and spatial relationships.

[0053] In some cases, a text embedding space is a continuous, low-dimensional vector space in which each vector represents the semantic meaning of the text. Points in the text embedding space are organized so that texts with similar meanings are located close to each other, reflecting the relationships between different words, phrases, or sentences based on contextual usage.

[0054] In some cases, an image embedding space is a high-dimensional vector space in which each point corresponds to a visual representation of an image. In the image embedding space, the distance between points reflects the similarity of the visual features of the images. In some cases, similar images are located closer together based on the features encoded in the image embeddings.

[0055] In some cases, text embedding and image embedding are combined in a multimodal embedding space within the image generation model. For example, the multimodal embedding space (also called a shared embedding space) is a high-dimensional space in which different types of data (modalities), such as text, images, audio, or video, are represented uniformly. In the shared embedding space, data from different modalities are encoded into vectors that can be directly compared and related to each other, even though the data originates from different sources. For example, the text embedding of the text description "a cute cat" and the image embedding of the image of a cute cat would be mapped to adjacent points in the shared embedding space.In some cases, the shared embedding space includes a shared semantic space designed to capture shared semantic meanings across modalities, where text input can be mapped to an image or vice versa.

[0056] In some cases, a story element refers to a character in the story. For example, the character could be a real person, an object, an entity, or a fictional person. In some cases, a preliminary image refers to a real image depicting a real person. In some cases, a style image shows a style such as a color style, image style, texture style, etc. In some cases, a storyboard comprises a set of panels, with each panel containing a synthetic image representing an event, scene, and / or action from the story. In some cases, each panel includes a corresponding image title that summarizes the scene depicted in each synthetic image. System architecture

[0057] In the Fig. References 5-8 and 13 comprise an image processing device and system, a storage unit, and a processing device coupled to the storage unit, the processing device being configured to perform operations that include: capturing a text prompt and an image prompt, wherein the text prompt describes a story and the image prompt shows an element of the story; generating a first scene prompt and a second scene prompt based on the text prompt, wherein the first scene prompt describes a first scene of the story and the second scene prompt describes a second scene of the story; and generating a first synthetic image and a second synthetic image based on the image prompt, the first scene prompt, and the second scene prompt.where the first synthetic image represents the first scene including the element, and the second synthetic image represents the second scene including the element.

[0058] In some aspects, the language generation model includes a transformer model. In some aspects, the image generation model includes a diffusion model. In some aspects, the image generation model includes a text encoder, an image encoder, a multimodal encoder, a prior model, or a combination thereof.

[0059] Fig. Figure 5 shows an example of an image processing device 500 according to aspects of the present disclosure. The example shown comprises the image processing device 500, a processor unit 505, an I / O module 510, and a storage unit 515. In one aspect, the storage unit 515 comprises a speech generation model 520, an image generation model 525, and a storyboard component 530.

[0060] According to some embodiments of the present disclosure, the image processing device 500 comprises a computer-implemented artificial neural network (ANN). An ANN is a hardware or software component comprising a number of connected nodes (for example, artificial neurons) that loosely correspond to the neurons in a human brain. Each connection or edge transmits a signal from one node to another (similar to the physical synapses in a brain). When a node receives a signal, the node processes the signal and then transmits the processed signal to other connected nodes. In some cases, the signals between the nodes include real numbers, and the output of each node is computed by a function of the sum of the inputs.In some examples, nodes can determine the output using other mathematical algorithms (for example, selecting the maximum of the inputs as the output) or another suitable algorithm to activate the node. Each node and edge is assigned one or more node weights that determine how the signal is processed and transmitted. The image processing device 500 is an example of this and includes aspects of the corresponding element, which relates to... Fig. 1 is described.

[0061] The Processor Unit 505 is an intelligent hardware device (for example, a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, or a combination thereof). In some cases, the Processor Unit 505 is configured to operate a memory array using a memory controller. In other cases, a memory controller is integrated into the Processor Unit 505. In some cases, the Processor Unit 505 is configured to execute computer-readable instructions stored in memory to perform various functions.In some embodiments, the 505 processor unit includes specialized components for modem processing, baseband processing, digital signal processing, or transmission processing. The 505 processor unit is an example of, or includes, aspects of the processor that relate to... Fig. 13 is described.

[0062] The I / O Module 510 (for example, an input / output interface) can include an I / O controller. An I / O controller can manage input and output signals for a device. An I / O controller can also manage peripheral devices that are not integrated into a device. In some cases, an I / O controller can represent a physical connection or port to an external peripheral device. In some cases, an I / O controller can use an operating system such as iOS®, Android®, MS-DOS®, MS-Windows®, OS / 2®, UNIX®, Linux®, or another well-known operating system. In other cases, an I / O controller can represent or interact with a modem, keyboard, mouse, touchscreen, or similar device. In some cases, an I / O controller can be implemented as part of a processor.In some cases, a user can interact with a device via an I / O controller or via hardware components controlled by an I / O controller.

[0063] In some examples, the I / O module 510 includes a user interface. A user interface allows a user to interact with a device. In some embodiments, the user interface may include an audio device such as an external speaker system, an external display device such as a monitor, or an input device (for example, a remote control device connected to the user interface directly or via an I / O controller module). In some cases, a user interface may be a graphical user interface (GUI). In some examples, a communication interface operates at the boundary between communicating units and the channel and may also record and process communications. A communication interface is provided here to enable a processing system coupled with a transceiver (for example, a transmitter and / or a receiver).In some examples, the transceiver is configured to send and receive signals for a communication device via an antenna. The I / O module 510 is an example of this and includes aspects of the I / O interface related to [the specific application]. Fig. 13 is described.

[0064] Examples of the 515 memory unit include random-access memory (RAM), read-only memory (ROM), or a hard disk. Examples of the 515 memory unit include solid-state storage and a hard disk drive. In some examples, the 515 memory unit is used to store computer-readable, computer-executable software, including instructions that, when executed, cause a processor to perform various functions described herein.

[0065] In some cases, the 515 memory unit includes a Basic Input / Output System (BIOS) that controls fundamental hardware or software operations, such as interaction with peripheral components or devices. In some cases, a memory controller manages memory cells. For example, the memory controller might include a row decoder, a column decoder, or both. In some cases, memory cells within the 515 memory unit store information in the form of a logical state.

[0066] In one aspect, storage unit 515 comprises a machine learning model. In another aspect, the machine learning model comprises the speech generation model 520, the image generation model 525, and the storyboard component 530. Storage unit 515 is an example of, or comprises, aspects of the storage subsystem that relate to... Fig. 13 is described.

[0067] In some cases, the machine learning model is a computational algorithm, model, or system designed to recognize patterns, make predictions, or perform a specific task (for example, image processing) without being explicitly programmed. Depending on the aspect, the machine learning model is implemented as software stored in the 515 memory unit and executable by the 505 processor unit, as firmware, as one or more hardware circuits, or as a combination thereof.

[0068] According to some embodiments of the present disclosure, the machine learning model includes an ANN, which is a hardware or software component comprising a number of connected nodes (for example, artificial neurons) loosely corresponding to the neurons in a human brain. Each connection or edge transmits a signal from one node to another (similar to the physical synapses in a brain). When a node receives a signal, it processes the signal and then transmits the processed signal to other connected nodes. In some cases, the signals between the nodes include real numbers, and the output of each node is computed by a function of the sum of the inputs. In some examples, nodes can determine the output using other mathematical algorithms (for example, selecting the maximum of the inputs as the output) or another suitable algorithm for activating the node.Each node and edge is assigned one or more node weights that determine how the signal is processed and transmitted.

[0069] During the training process, one or more node weights are adjusted to increase the accuracy of the result (for example, by minimizing a loss function that, in a sense, represents the difference between the current result and the target result). The weight of an edge increases or decreases the strength of the signal transmitted between nodes. In some cases, nodes have a threshold below which a signal is not transmitted at all. In some examples, the nodes are grouped into layers. Different layers perform different transformations on their respective inputs. The initial layer is known as the input layer, and the final layer is known as the output layer. In some cases, signals pass through certain layers multiple times.

[0070] In some embodiments, the machine learning model includes a computer-implemented CNN. A CNN is a class of neural networks commonly used in computer vision or image classification systems. In some cases, a CNN can enable the processing of digital images with minimal preprocessing. A CNN can be characterized by the use of convolutional (or cross-correlative) layers. These layers apply a convolutional operation to the input before passing the result to the next layer. Each convolutional node can process data for a limited input field (for example, the receptive field). During a forward pass of the CNN, filters in each layer can be convolutional over the input volume, computing the dot product between the filter and the input.During the training process, the filters can be modified so that they are activated when the filters detect a specific feature within the input.

[0071] One aspect of a machine learning model is its use of machine learning parameters. Machine learning parameters, also known as model parameters or weights, are variables that determine the behavior and properties of the machine learning model. Machine learning parameters can be learned from training data or estimated and are used to make predictions or perform tasks based on learned patterns and relationships within the data.

[0072] Machine learning parameters are adjusted during a training process to minimize a loss function or maximize a performance metric. The goal of the training process is to find optimal values ​​for the parameters that enable the machine learning model to make accurate predictions or perform the given task well.

[0073] For example, during the training process, an algorithm adjusts the machine learning parameters to minimize errors or loss between predicted outputs and actual targets, using optimization techniques such as gradient descent, stochastic gradient descent, or other optimization algorithms. Once the machine learning parameters have been learned from the training data, they are used to make predictions on new, unseen data.

[0074] In some implementations, the machine learning model includes a computer-implemented recurrent neural network (RNN). An RNN is a class of ANNs in which the connections between nodes form a directed graph along an ordered (e.g., temporal) sequence. This allows an RNN to model temporally dynamic behavior, such as predicting which element should come next in a sequence. Thus, an RNN is suitable for tasks involving ordered sequences, such as text recognition (where words in a sentence are ordered). In some cases, an RNN comprises one or more finite impulse recurrent networks (characterized by nodes forming a directed acyclic graph), one or more infinite impulse recurrent networks (characterized by nodes forming a directed cyclic graph), or a combination thereof.

[0075] In some embodiments, the machine learning model includes a transformer (or a transformer model or transformer network), where the transformer is a type of neural network model used for natural language processing tasks. A transformer network transforms one sequence into another using an encoder and a decoder. The encoder and decoder comprise modules that can be stacked multiple times. The modules include multi-head attention and feed-forward layers. The inputs and outputs (target sentences) are initially embedded in an n-dimensional space. Positional encoding of the different words (for example, to assign a relative position to each word / part in a sequence, since the sequence depends on the order of the elements) is added to the embedded representation (n-dimensional vector) of each word.In some examples, a transformer network includes an attention mechanism where the attention considers an input sequence and, at each step, decides which other parts of the sequence are important. The attention mechanism includes a query, keys, and values, denoted by Q, K, and V, respectively. Q is a matrix containing the query (vector representation of a word in the sequence), K are the keys (vector representations of the words in the sequence), and V are the values, again the vector representations of the words in the sequence. For the encoder and decoder multi-head attention modules, V consists of the same word sequence as Q. However, for the attention module that considers the encoder and decoder sequences, V is different from the sequence represented by Q. In some cases, values ​​in V are multiplied by certain attention weights a and summed.

[0076] In machine learning, an attention mechanism (for example, implemented in one or more ANNs) is a method for assigning different levels of importance to different elements of an input sequence. Computing attention can involve three basic steps. First, a similarity between the query and key vectors obtained from the input is calculated to generate attention weights. Similarity functions that can be used for this process include the dot product, splice, detector, and the like. Next, a softmax function is used to normalize the attention weights. Finally, the attention weights are weighted along with their corresponding values. In the context of an attention network, the key and value are vectors or matrices used to represent the input data.The key is used to determine which parts of the input the attention mechanism should focus on, while the value is used to represent the data actually processed.

[0077] An attention mechanism is a key component in some KNN architectures, particularly KNNs used in natural language processing (NLP) and sequence-to-sequence tasks. This mechanism allows a KNN to focus on different parts of an input sequence when making predictions or generating outputs. Some sequence models (such as RNNs) process an input sequence sequentially and maintain an internal hidden state that captures information from previous steps. However, in some cases, this sequential processing leads to difficulties in capturing dependencies over long distances or in focusing on specific parts of the input sequence.

[0078] The attention mechanism addresses these difficulties by allowing a KNN to selectively focus on different parts of an input sequence and assign different levels of importance or attention to each part. The attention mechanism achieves this selective focusing by considering the relevance of each input element to the current state of the KNN.

[0079] The term "self-attention" refers to a machine learning model in which representations of the input interact to determine attentional weights for the input. Self-attention can be distinguished from other attentional models in that the attentional weights are at least partially determined by the input.

[0080] According to some aspects, the 520 speech generation model is implemented as software stored in the 515 memory unit and executed by the 505 processor unit, as firmware, as one or more hardware circuits, or as a combination thereof. According to some aspects, the 520 speech generation model captures a text prompt that describes a story. In some examples, the 520 speech generation model generates a first scene prompt and a second scene prompt based on the text prompt, with the first scene prompt describing a first scene of the story and the second scene prompt describing a second scene of the story.

[0081] In some examples, the Language Generation Model 520 generates a first caption and a second caption based on the first synthetic image and the second synthetic image, respectively, with the storyboard containing a first panel with the first synthetic image and the first caption, and a second panel with the second synthetic image and the second caption. In some examples, the Language Generation Model 520 receives a modification command specifying the first scene and a modified element. In some examples, the Language Generation Model 520 generates a modified scene prompt based on the modification command.

[0082] According to some aspects, the Language Generation Model 520 captures a text prompt and an image prompt, where the text prompt describes a story and the image prompt represents an element of the story. In some examples, the Language Generation Model 520 generates a first scene prompt and a second scene prompt based on the text prompt, where the first scene prompt describes a first scene of the story and the second scene prompt describes a second scene of the story. In some aspects, the Language Generation Model 520 includes a transformer model. The Language Generation Model 520 is an example of, or includes, aspects of the corresponding element that relates to Fig. 6 is described.

[0083] According to some aspects, the 525 image generation model is implemented as software stored in the 515 memory unit and executed by the 505 processor unit, as firmware, as one or more hardware circuits, or as a combination thereof. According to some aspects, the 525 image generation model generates a first synthetic image and a second synthetic image based on the first and second scene prompts, the first synthetic image representing the first scene and the second synthetic image representing the second scene. In some examples, the 525 image generation model captures an image prompt showing an element of the story, with the first and second synthetic images being generated based on the image prompt and representing that element.

[0084] In some examples, the 525 image generation model captures an identity text prompt that describes the element. In some examples, the 525 image generation model generates the image prompt based on the identity text prompt. In some examples, the 525 image generation model captures a preliminary image of a person. In some examples, the 525 image generation model crops the preliminary image to obtain the image prompt, where the element includes a face of the person. In some examples, the 525 image generation model captures a style image that displays a style, where the first synthetic image and the second synthetic image are generated based on the style image and include the style.

[0085] In some examples, the 525 image generation model captures a first noise input and a second noise input. In some examples, the 525 image generation model denoises the first noise input based on the first scene prompt and the second noise input based on the second scene prompt to obtain the first synthetic image and the second synthetic image. In some examples, the 525 image generation model generates a modified synthetic image based on the first scene and the modified element.

[0086] According to some aspects, the 525 image generation model generates a first synthetic image and a second synthetic image based on the image prompt, the first scene prompt, and the second scene prompt, where the first synthetic image represents the first scene including the element, and the second synthetic image represents the second scene including the element. In some aspects, the 525 image generation model includes a diffusion model.

[0087] According to some aspects, the image generation model 525 includes a text encoder, an image encoder, a multimodal encoder, a prior model, or a combination thereof. In some examples, the text encoder is a neural network that converts a text prompt (for example, words, sentences, etc.) into a text embed (for example, a numerical vector representation) that captures the semantic meaning of the text prompt. In some examples, the image encoder is a neural network that receives an input image or image prompt and generates an image embed that includes visual features of the image encoded in a low-dimensional vector space. In some cases, the image encoder includes a convolutional neural network (CNN). In some examples, the multimodal encoder receives input data from multiple modalities (for example, text, image, audio, video, etc.) and generates an embed with a unified representation.In some cases, the multimodal encoder can, for example, generate a text embedding based on input text and an image embedding based on input image. The multimodal encoder can combine the text embedding and the image embedding to generate a combined embedding in the same vector space (or embedding space).

[0088] In some cases, the Prior model transforms a text embedding into an image embedding. In some cases, the Prior model is a diffusion-based Prior model, which involves an iterative process that starts from a noisy state (for example, a noisy version of the text embedding) and gradually reduces the noise to obtain the final output (for example, the predicted image embedding). In other cases, the Prior model is a transformer-based Prior model that transforms the text embedding into an image embedding. For example, the transformer includes a variety of transformer layers that take the text embedding and model complex relationships and dependencies across the dimensions of the embedding, mapping the text embedding to the image embedding, which captures more visual information than the text embedding. The image generation model 525 is an example of this.includes aspects of the relevant element that relates to . Fig. 6 is described.

[0089] According to some aspects, the Storyboard Component 530 is implemented as software stored in the Memory Unit 515 and executed by the Processor Unit 505, as firmware, as one or more hardware circuits, or as a combination thereof. According to some aspects, the Storyboard Component 530 generates a storyboard using the first synthetic image and the second synthetic image. In some aspects, the storyboard includes a first panel with the first synthetic image and a first caption, and a second panel with the second synthetic image and a second caption. The Storyboard Component 530 is an example of, or includes, aspects of the corresponding element related to Fig. 6 is described.

[0090] According to some aspects, the image processing device 500 includes a training component. The training component is implemented as software stored in the memory unit 515 and executed by the processor unit 505, as firmware, as one or more hardware circuits, or as a combination thereof. According to some embodiments, the training component is implemented as software stored in a memory unit and executed by a processor in the processor unit of a separate computing device, as firmware in the separate computing device, as one or more hardware circuits of the separate computing device, or as a combination thereof. In some examples, the training component is part of a device other than the image processing device 500 and communicates with the image processing device 500. In other examples, the training component is part of the image processing device 500.According to some aspects, the training component trains the speech generation model 520 and the image generation model 525 together or separately.

[0091] Fig. Figure 6 shows an example of a storyboard generation system 675 according to aspects of the present revelation. The example shown includes the machine learning system 600, the text prompt 605, the speech generation model 610, the first scene prompt 615, the second scene prompt 620, the image generation model 625, the identity text prompt 630, the preliminary image generation model 635, the preliminary output image 640, the image prompt 645, the first synthetic image 650, the second synthetic image 655, the storyboard component 660, the first caption 665, the second caption 670, and the storyboard 675. In some aspects, the machine learning system 600 includes the speech generation model 610, the image generation model 625, and the storyboard component 660. In some embodiments, the machine learning system 600 also includes the preliminary image generation model. 635.

[0092] With reference to Fig. In 6, the machine learning system 600 receives the text prompt 605 and generates the storyboard 675. In some embodiments, the speech generation model 610 receives the text prompt 605 and generates the first scene prompt 615 and the second scene prompt 620. For example, the text prompt 605 describes a story such as "Blueberry's Interdimensional Adventure." In one aspect, the speech generation model 610 includes a transformer model. For example, the transformer architecture enables the speech generation model 610 to understand and generate natural language output. In some cases, the speech generation model 610 includes an input embedding layer configured to convert a text input (for example, the text prompt 605) into a text embedding, where the text embedding is a numerical representation of words that represent meaning and relationship.In some cases, the 610 language generation model includes positional encoding, which encodes positional information via the order of the tokens in the input text.

[0093] In some aspects, the 610 language generation model includes a multi-head self-attention layer that focuses on different parts of the input sequence when processing each token. In some cases, this attentional mechanism allows the model to learn different relationships between words in parallel. Each input token pays attention to every other token in the sequence, enabling the model to understand long-range dependencies and relationships. In some aspects, the 610 language generation model includes a feedforward network that forwards the processed tokens from the multi-head self-attention layer. In some embodiments, the 610 language generation model includes a normalization layer that scales the output to a target output.In some embodiments, the speech generation model 610 includes one or more stacked transformer layers, which contain one or more self-attention and feedforward layers. In some aspects, the speech generation model 610 includes a decoder configured to generate predicted tokens in the sequence based on the input text. In some cases, the decoder generates a predicted text embedding of output text. In some aspects, the speech generation model 610 includes an output layer configured to generate the output text based on the input text. For example, the output text is the first scene prompt 615 and the second scene prompt 620. In some embodiments, the speech generation model generates a variety of scene prompts based on the text prompt 605.

[0094] In some embodiments, the machine learning system 600 performs prompt engineering using text prompt 605 to obtain a modified text prompt. For example, the modified text prompt may include text prompt 605 with the number of panels prepended or appended. For example, the modified text prompt may read: "Create a storyboard of {story} with {n} panels," where {story} specifies text prompt 605 and {n} specifies the number of panels to be generated in storyboard 675. In some cases, the number of scene prompts is generated based on the number specified in the modified text prompt.

[0095] In some embodiments, the first scene prompt 615 and the second scene prompt 620 are provided to the image generation model 625 to generate the first synthetic image 650 and the second synthetic image 655, respectively. In some embodiments, the image generation model 625 further receives the image prompt 645 to generate the first synthetic image 650 and the second synthetic image 655. For example, the preliminary image generation model 635 receives the identity text prompt 630, which describes a character, and generates a preliminary output image 640 that represents the character. For example, the identity text prompt 630 might read: "An image of a cartoonish blue furball." In some embodiments, the preliminary image generation model 635 is a pretrained image generation model. In some embodiments, the preliminary image generation model 635 is the image generation model 625.In some cases, an image prompt 645 is generated based on the preliminary output image 640. For example, a facial area of ​​the figure depicted in the preliminary output image 640 is cropped to obtain the image prompt 645. The image prompt 645 is provided to the image generation model 625 to ensure that the figure depicted in the synthetic images (for example, the first synthetic image 650 and the second synthetic image 655) is consistent and that the figure's identity is preserved.

[0096] In some embodiments, the preliminary output image 640 is used as the image prompt 645. For example, if the identity text prompt 630 describes a fictional character (for example, a cartoonish blue furball), the preliminary output image 640, which represents the entire fictional character, is provided to the image generation model 625 to accommodate the fictional character's attributes, such as ears, clothing, shape, and size. In some embodiments, if the preliminary output image 640 represents a human character (for example, a model-generated image of a human or a real image of a human), the facial area of ​​the human character depicted in the preliminary output image 640 is cropped to obtain the image prompt 645. Accordingly, the identity of the human character can be preserved.

[0097] According to some embodiments, the image generation model 625 comprises a text encoder, an image encoder, a multimodal encoder, a prior model, or a combination thereof. In some embodiments, the image generation model 625 comprises a text encoder configured to encode the set of scene prompts (for example, the first scene prompt 615 and the second scene prompt 620) to generate a set of scene prompt embeddings. In some embodiments, the set of scene prompt embeddings is used to generate a set of synthetic images (the first synthetic image 650 and the second synthetic image 655). In some embodiments, an image encoder or a prior model of the image generation model 625 each generates a set of image embeddings based on the set of scene prompt embeddings, the set of image embeddings being used to generate a set of synthetic images.

[0098] In some embodiments, the image generation model 625 receives an image prompt 645 representing a figure. In some cases, the image prompt 645 is a real image. In some embodiments, the image encoder of the image generation model 625 encodes the image prompt 645 to generate an identity image embed. In some aspects, the identity image embed includes information about the identity of the figure represented in the image prompt 645. In one embodiment, the identity image embed is combined or chained with each of the set of image embeds of the set of scene prompts. In one aspect, the set of synthetic images is generated based on the chained image embeds.

[0099] According to some aspects, the first synthetic image 650, the second synthetic image 655, the first caption 665, and the second caption 670 are provided to a storyboard component 660 to generate the storyboard 675. In some embodiments, the first caption 665 and the second caption 670 are provided by a user. In some embodiments, the first caption 665 and the second caption 670 are generated based on the language generation model 610. In some cases, the storyboard component 660 combines the first synthetic image 650 and the first caption 665 to generate a first panel of the storyboard 675. In some cases, the storyboard component 660 combines the second synthetic image 655 and the second caption 670 to generate a second panel of the storyboard 675.In some aspects, the first and second panels are combined to create the storyboard. In other cases, the storyboard comprises a multitude of panels.

[0100] Text prompt 605 is an example of, or encompasses, aspects of the corresponding element, which relates to the Fig. 3 and Fig. 7 is described. The language generation model 610 is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 5 is described. The image generation model 625 is an example of, or includes, aspects of the corresponding element, which relates to Fig. 5 is described.

[0101] Image prompt 645 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 3 is described. Storyboard component 660 is an example of, or includes, aspects of the corresponding element, which relates to Fig. 5 is described. Storyboard 675 is an example of, or includes, aspects of the corresponding element, which relates to Fig. 3 is described.

[0102] Fig. Figure 7 shows an example of an image generation model according to aspects of the present disclosure. The example shown includes the diffusion model 700, the original image 705, the pixel space 710, the image encoder 715, the original image feature 720, the latent space 725, the forward diffusion process 730, the noisy feature 735, the reverse diffusion process 740, the denoised image feature 745, the image decoder 750, the output image 755, the text prompt 760, the text encoder 765, the guidance feature 770, and the guidance space 775.

[0103] Diffusion models are a class of generative neural networks that can be trained to generate new data with features similar to those in training data. In particular, diffusion models can be used to generate novel images. Diffusion models can be used for various image generation tasks, including image super-resolution, generating images with perceptual metrics, conditional generation (for example, generation based on text guidance, color guidance, style guidance, and image guidance), image inpainting, and image manipulation.

[0104] Types of diffusion models include Denoising Diffusion Probabilistic Models (DDPMs) and Denoising Diffusion Implicit Models (DDIMs). In DDPMs, the generative process involves inverting a stochastic Markov diffusion process. DDIMs, on the other hand, use a deterministic process, so the same input leads to the same output. Diffusion models can also be characterized by whether the noise is added directly to the image or to image features generated by an encoder (for example, latent diffusion).

[0105] Diffusion models work by iteratively adding noise to data during a forward process and then learning to reconstruct the data by removing the noise during a reverse process. For example, during training, the diffusion model 700 can take an original image 705 in pixel space 710 as input and apply an image encoder 715 to convert the original image 705 into an original image feature 720 in latent space 725. Then, a forward diffusion process 730 incrementally adds noise to the original image feature 720 to obtain the noisy feature 735 (also in latent space 725) at different noise levels.

[0106] Next, a reverse diffusion process 740 (for example, a U-Net KNN) incrementally removes the noise from the noisy feature 735 at the various noise levels to obtain the denoised image feature 745 in the latent space 725. In some examples, the denoised image feature 745 is compared to the original image feature 720 at each of the different noise levels, and parameters of the diffusion model's reverse diffusion process 740 are updated based on the comparison. Then, an image decoder 750 decodes the denoised image feature 745 to obtain an output image 755 in pixel space 710. In some cases, an output image 755 is generated for each of the different noise levels. The output image 755 can be compared to the original image 705 to train the reverse diffusion process 740. In some cases, output image 755 refers to the synthetic image (for example, as in relation to Fig. 6 described).

[0107] In some cases, the image encoder 715 and the image decoder 750 are pre-trained before the reverse diffusion process 740 is trained. In some examples, the image encoder 715 and the image decoder 750 are trained together, or the image encoder 715 and the image decoder 750 are fine-tuned together with the reverse diffusion process 740.

[0108] The reverse diffusion process 740 can also be controlled based on a text prompt 760 or another guidance prompt such as an image, layout, style, color, or segmentation map. The text prompt 760 can be encoded using a text encoder 765 (for example, a multimodal encoder) to obtain a guidance feature 770 in guidance space 775. The guidance feature 770 can be combined with the noisy feature 735 in one or more layers of the reverse diffusion process 740 to ensure that the output image 755 contains content described by the text prompt 760. For example, the guidance feature 770 can be combined with the noisy feature 735 using a cross-attention block within the reverse diffusion process 740.

[0109] Cross-attention, also known as multi-head attention, is an extension of the attention mechanism used in some KNNs, for example, in NLP tasks. In some cases, cross-attention simultaneously attends to multiple parts of an input sequence and captures interactions and dependencies between different elements. Cross-attention involves two input sequences: a query sequence and a key-value sequence. The query sequence represents the elements requiring attention, while the key-value sequence contains the elements to be observed. In some cases, to compute cross-attention, the cross-attention block transforms each element of the query sequence into a "query" representation (for example, using linear projection), while the elements of the key-value sequence are transformed into "key" and "value" representations.

[0110] The cross-attention block calculates attention scores by measuring the similarity between each query representation and the key representations, with higher similarity indicating that a key element receives more attention. An attention score indicates the importance or relevance of each key element to its corresponding query element.

[0111] The cross-attention block then normalizes the attention scores to obtain attention weights (for example, using a softmax function), with the attention weights determining how much information from each value element is included in the final attention-based representation. By simultaneously paying attention to different parts of the key-value sequence, the cross-attention block captures relationships and dependencies between the input sequences, enabling the machine learning model to understand the context and produce more accurate and contextually relevant outputs.

[0112] In some examples, diffusion models are based on a neural network architecture known as a U-Net. The U-Net takes input features with an initial resolution and an initial number of channels and processes the input features using an initial neural network layer (for example, a convolutional network layer) to generate intermediate features. The intermediate features are then downsampled using a downsampling layer, so that the downsampled features have a lower resolution than the initial resolution and a higher number of channels than the initial number of channels.

[0113] This process is repeated multiple times and then reversed. For example, downsampled features are upsampled to obtain upsampled features. These upsampled features can be combined with intermediate features that have the same resolution and number of channels via a skip connection. These inputs are then processed by a final neural network layer to generate output features. In some cases, the output features will have the same resolution and number of channels as the initial features.

[0114] In some cases, a U-Net accepts additional input features to generate conditionally generated outputs. For example, the additional input features might include a vector representation of an input prompt. These additional input features can be combined with intermediate features within the neural network in one or more layers. For instance, a cross-attention module can be used to combine the additional input features and the intermediate features. Further details about the U-Net are discussed in relation to… Fig. 8 described.

[0115] A diffusion process can also be modified based on conditional guidance. In some cases, a user provides a text prompt (for example, text prompt 760) that describes content to be included in a generated image. In some examples, guidance can be provided in a form other than text, such as an image, sketch, color, style, or layout. The system transforms the text prompt 760 (or other guidance) into a conditional guidance vector or other multidimensional representation. For example, text can be transformed into a vector or a set of vectors using a transformer model or a multimodal encoder. In some cases, the conditional guidance encoder is trained independently of the diffusion model.

[0116] A noise map containing random noise is initialized. The noise map can reside in a pixel space or a latent space. Initializing an image with random noise allows the generation of different versions of the image, each containing the content described by the conditional guidance. The Diffusion Model 700 then generates an image based on the noise map and the conditional guidance vector.

[0117] A diffusion process can include both a forward diffusion process 730 to add noise to an image (for example, the original image 705) or to features (for example, the original image feature 720) in a latent space 725, and a reverse diffusion process 740 to denoise the images (or features) to obtain a denoise-free image (for example, the output image 755). The forward diffusion process 730 can be represented as q(x t | x t-1) can be represented, and the reverse diffusion process 740 can be represented as p θ (x t-1 | x t ) will be presented. Further details on the diffusion process will be provided in relation to Fig. 9 described.

[0118] A diffusion model 700 can be trained using either a forward diffusion process 730 or a reverse diffusion process 740. In an example, the user initializes an untrained model. Initialization can include defining the model architecture and setting initial values ​​for the model parameters. In some cases, initialization can include defining hyperparameters, such as the number of layers, the resolution and number of channels in each layer block, the position of skip connections, and the like.

[0119] The system then adds noise to a training image in N steps using a forward diffusion process 730. In some cases, the forward diffusion process 730 is a fixed process in which Gaussian noise is successively added to an image. In latent diffusion models, the Gaussian noise can be successively added to features (for example, the original image feature 720) in a latent space 725.

[0120] In each stage n, starting with stage N, a reverse diffusion process 740 is used to predict the image or image features for stage n-1. For example, the reverse diffusion process 740 can predict the noise added by the forward diffusion process 730, and the predicted noise can be removed from the image to obtain the predicted image. In some cases, an original image 705 is predicted in each stage of the training process.

[0121] The training component (for example, the one relating to Fig. The training component described in section 5 compares the predicted image (or image features) at stage n-1 with an actual image (or image features), such as the image at stage n-1 or the original input image. For example, the diffusion model 700 can be trained using observed data x to determine the upper variational bound of the negative log-likelihood -log p. θ (x) of the training data. The training component then updates the parameters of the diffusion model 700 based on the comparison. For example, parameters of a U-Net can be updated using gradient descent. Time-dependent parameters of the Gaussian transitions can also be learned. Further details on training the diffusion model are given in relation to Fig. 12 described.

[0122] The original image 705 is an example of, or includes, aspects of the corresponding element, which relates to Fig. 9 is described. The forward diffusion process 730 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 9 is described. The reverse diffusion process 740 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 9 is described. Text prompt 760 is an example of, or includes aspects of, the corresponding element, which relates to the Fig. 3 and Fig. 6 is described.

[0123] Fig. Figure 8 shows an example of a U-Net-800 architecture according to aspects of the present disclosure. The example shown includes the U-Net 800, the input feature 805, the initial neural network layer 810, the intermediate feature 815, the down-sampling layer 820, the downsampled feature 825, the up-sampling process 830, the upsampled feature 835, the skip connection 840, the final neural network layer 845, and the output feature 850.

[0124] In some examples, the U-Net 800 is an example of the component that performs the reverse diffusion process 740 in relation to Fig. The diffusion model described in section 700 is executed and includes architectural elements of the model in relation to Fig. 5 described image generation model 525. The one in Fig. Figure 8 of the U-Net 800 shown is an example of, or encompasses, aspects of the architecture that are within the context of Fig. The reverse diffusion process described in section 7 is used.

[0125] In some examples, diffusion models are based on a neural network architecture known as a U-Net. The U-Net 800 takes the input feature 805 with an initial resolution and an initial number of channels and processes the input feature 805 using an initial neural network layer 810 (for example, a convolutional network layer) to generate the intermediate feature 815. The intermediate feature 815 is then downsampled using a downsampling layer 820, so that the downsampled feature 825 has a lower resolution than the initial resolution and a higher number of channels than the initial number of channels.

[0126] This process is repeated multiple times and then reversed. For example, the downsampled feature 825 is upsampled using the upsampling process 830 to obtain the upsampled feature 835. The upsampled feature 835 can be combined with the intermediate feature 815, which has the same resolution and the same number of channels, using a skip connection 840. These inputs are processed using a final neural network layer 845 to generate the output feature 850. In some cases, the output feature 850 has the same resolution and the same number of channels as the initial resolution.

[0127] In some cases, the U-Net 800 accepts an additional input feature to generate conditionally generated outputs. For example, the additional input feature can be a vector representation of an input prompt. Within the neural network, this additional input feature can be combined with the intermediate feature 815 in one or more layers. For example, a cross-attention module can be used to combine the additional input features and the intermediate feature 815. Diffusion processing

[0128] Fig. Figure 9 shows an example of a diffusion process 900 according to aspects of the present revelation. The example shown includes the diffusion process 900, a forward diffusion process 905, a reverse diffusion process 910, a noisy image 915, a first intermediate image 920, a second intermediate image 925, and an original image 930.

[0129] A diffusion process 900 can be used to perform a forward diffusion process 905 to add noise to the original image 930 (for example, the original image 705 in relation to Fig. 7) or features (for example, the original image feature 720 in relation to Fig. 7) in a latent space. In some aspects, the diffusion process 900 includes a reverse diffusion process 910 to denoise the noisy image 915 (or image features) to obtain a denoised image (or original image 930). The forward diffusion process 905 can be described as q(x t | x t-1 ) can be represented, and the reverse diffusion process 910 can be represented as p θ (x t-1 | x t). In some cases, the forward diffusion process 905 is used during training to generate images with successively increasing noise, and a neural network is trained to execute the reverse diffusion process 910 (for example, to successively remove the noise).

[0130] In an example of a forward diffusion process 905 for a latent diffusion model (for example, diffusion model 700 in relation to Fig. 7) The diffusion model maps an observed variable x0 (either in a pixel space or a latent space) to intermediate variables x1, ...,x T to obtain using a Markov chain. The Markov chain incrementally adds Gaussian noise to the data set to approximate the post-distribution q(x). 1:T | x0) to obtain, while the latent variables are processed by a neural network such as a U-Net, where x1, ..., x Texhibit the same dimensionality as x0.

[0131] The neural network can be trained to execute the reverse diffusion process 910. During the reverse diffusion process 910, the diffusion model starts with noisy data x. T , like a noisy image 915, and denoises the data to p θ (x t-1 | x t to obtain. In each step t - 1, the reverse diffusion process takes 910 x t (like the first intermediate image 920) and t as input. Here, t represents a step in the sequence of transitions associated with different noise levels. The reverse diffusion process 910 yields x t-1 , like the second intermediate image 925, iteratively, until x T which is reduced back to x0, the original image 930. The reverse diffusion process 910 can be represented as follows: pθ(xt−1|xt):=N(xt−1;μθ(xt,t),Σθ(xt,t)).

[0132] The joint probability of a sequence of samples in the Markov chain can be written as the product of conditional probabilities and the marginal probability: xT:pθ(x0:T):=p(xT)∏t=1Tpθ(xt−1|xt), where p(x T ) = N(x T ; 0, I) is the pure noise distribution, since the reverse diffusion process 910 takes the result of the forward diffusion process 905, a sample of pure noise, as input and ∏t=1T pθ(xt−1|xt) represents a sequence of Gaussian transitions, corresponding to a sequence of adding Gaussian noise to the sample.

[0133] At the time of interference, observed data x0 in a pixel space can be mapped as input to a latent space, and generated data x̃ are mapped from the latent space back to the pixel space as output. In some examples, x0 represents an original input image with low image quality, and latent variables x1, ..., x T x̃ represents noisy images, and x̃ represents the generated image with high image quality.

[0134] The forward diffusion process 905 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 7 is described. The reverse diffusion process 910 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 7 is described. The original image 930 is an example of, or includes aspects of, the corresponding element, which relates to Fig. 7 is described. Prompt modification

[0135] Fig. Figure 10 shows an example of a method 1000 for modifying a scene prompt according to aspects of this disclosure. In some examples, these operations are performed by a system comprising a processor that executes a set of codes to control functional elements of a device. Additionally or alternatively, certain processes are performed using specialized hardware. Generally, these operations are performed according to the methods and procedures described in accordance with aspects of this disclosure. In some cases, the operations described herein consist of various substeps or are performed in conjunction with other operations.

[0136] In step 1005, the system receives a modification command specifying an initial scene and a modified element. In some cases, the operations of this step relate to or can be performed by a language generation model, as in the case of Fig. 5 and Fig. 6 described. In some cases, a user (for example, user 100 in relation to Fig. 1) Provides the system with a modification command. For example, the modification command describes a change to an element (for example, room) in the scene prompt (for example, the first scene prompt 615 with respect to Fig. 6) to another element (for example, kitchen).

[0137] In step 1010, the system generates a modified scene prompt based on the modification command. In some cases, the operations of this step relate to or can be performed by a language generation model, as in the case of Fig. 5 and Fig. 6 described. In some cases, the speech generation model generates a modified scene prompt instead of the initial scene prompt. For example, as described in relation to Fig. As described in point 3, the modified scene prompt reads: “Blueberry is inspired by a book that depicts the outside world. Blueberry is in a kitchen and ready to go outside. Blueberry wants to go out into the world for an adventure.”

[0138] In step 1015, the system generates a modified synthetic image based on the initial scene and the modified element. In some cases, the operations of this step relate to or can be performed by an image generation model, as in the case of Fig. 5 and Fig. 6. In some embodiments, the modified scene prompt is provided to the image generation model to generate a modified synthetic image representing the change (for example, from room to kitchen). In some cases, one or more elements in one or more synthetic images can be modified based on a single modification command. Training and evaluation

[0139] Fig. Figure 11 shows an example of a flowchart that presents an algorithm as a step-by-step procedure in an example implementation of executable operations for training a machine learning model according to aspects of the present disclosure. In some embodiments, the method 1100 describes an operation of the training component for configuring the image generation model 525, as described in relation to Fig. 5 described. Procedure 1100 provides one or more examples of generating training data, using the training data to train a machine learning model, and using the trained machine learning model to perform a task.

[0140] At the beginning of this example, a machine learning system collects training data (block 1102), which serves as the basis for training a machine learning model that defines what is being modeled. The training data can be collected by the machine learning system from a variety of sources. Examples of training data sources include public datasets, system platforms of service providers that offer application programming interfaces (for example, social media platforms), user data collection systems (for example, digital surveys and online crowdsourcing systems), and the like. Training data collection can also involve data augmentation and synthetic data generation techniques to expand and diversify the available training data, as well as balancing techniques to even out the number of positive and negative examples, and so on.

[0141] The machine learning system can also be configured to identify features relevant to a specific task (Block 1104) for which the machine learning model is to be trained. Examples of such tasks include classification, natural language processing, generative artificial intelligence, recommendation systems, reinforcement learning, clustering, and so on. To this end, the machine learning system collects training data based on the identified features and / or filters the training data after collection based on these features. The training data is then used to train a machine learning model.

[0142] To train the machine learning model in the example shown, the machine learning model is first initialized (Block 1106). Initializing the machine learning model involves selecting a model architecture (Block 1108) to be trained. Examples of model architectures include neural networks, convolutional neural networks (CNNs), long short-term memory (LSTM) neural networks, generative adversarial networks (GANs), decision trees, support vector machines, linear regression, logistic regression, Bayesian networks, random forest learning, dimensionality reduction algorithms, boosting algorithms, deep learning neural networks, U-Net architecture, etc.

[0143] A loss function is also selected (Block 1110). The loss function is used to measure the difference between an output of the machine learning model (for example, the model predictions) and target values ​​(for example, as expressed by the training data), which will be used to train the machine learning model. Additionally, an optimization algorithm is selected (Block 1112) to be used in conjunction with the loss function to optimize the parameters of the machine learning model during training, with examples including gradient descent, stochastic gradient descent (SGD), etc.

[0144] The initialization of the machine learning model further includes setting initial values ​​for the model (Block 1116). Examples of this include initializing node weights and biases to increase training efficiency and reduce the consumption of computational resources during training. Additionally, hyperparameters are set (Block 1114) to control the training of the machine learning model. Examples of these include regularization parameters, model parameters (for example, the number of layers in a neural network), learning rate, batch sizes selected from the training data, and so on. The hyperparameters are set using a variety of techniques, including randomization, heuristics learned from other training scenarios, and so forth.

[0145] The machine learning model is then trained by the machine learning system using the training data (Block 1118). A machine learning model is a computer representation that can be adapted (for example, trained and retrained) based on inputs of the training data to approximate unknown functions. In particular, the term machine learning model can encompass a model that uses algorithms (for example, using the model architectures described above) to learn from known data and make predictions about it by analyzing training data to learn and relearn, producing outputs that reflect patterns and properties expressed by the training data.

[0146] Examples of training methods include supervised learning, which uses labeled data; unsupervised learning, which finds underlying structures or patterns in the training data; reinforcement learning based on optimization functions (for example, rewards and / or punishments); the use of nodes as part of deep learning; and so on. The machine learning model, for example, is configurable to include a multitude of nodes that together form a multitude of layers. The layers can be configured, for example, to include an input layer, an output layer, and one or more hidden layers.Computations are performed by the nodes within the layers via the hidden states using a system of weighted connections that are “learned” during training, for example by using the selected loss function and backpropagation to optimize the performance of the machine learning model in performing an associated task.

[0147] As part of the machine learning model training, a termination criterion (decision block 1120) is determined to be met. This criterion is used to validate the machine learning model. The termination criterion can be used to reduce overfitting of the machine learning model, decrease the consumption of computational resources, and promote the machine learning model's ability to respond to unseen data not included as examples in the training data. Examples of termination criteria include a predefined number of epochs, stabilization of validation loss, reaching a performance improvement threshold, whether a threshold level of accuracy has been reached, or based on performance metrics such as precision and recall.If the termination criterion is not met (“no” in decision block 1120), procedure 1100 in this example continues the training of the machine learning model using the training data (block 1118).

[0148] If the termination criterion is met (“yes” in decision block 1120), the trained machine learning model is used to generate an output based on subsequent data (block 1122). For example, the trained machine learning model is trained to perform a task described above and is therefore, once trained, configured to perform this task based on subsequent data received as input and processed by the machine learning model.

[0149] Fig. Figure 12 shows an example of a method for training a diffusion model according to aspects of this disclosure. In some examples, these operations are performed by a system comprising a processor that executes a set of codes to control functional elements of a device. Additionally or alternatively, certain processes are performed using specialized hardware. Generally, these operations are performed according to the methods and procedures described in accordance with aspects of this disclosure. In some cases, the operations described herein consist of various substeps or are performed in conjunction with other operations.

[0150] In some embodiments, method 1200 describes an operation of the training component for training the image generation model 525, as in relation to Fig. 5 described. Procedure 1200 represents an example of training a reverse diffusion process, as described above in relation to Fig. 9 described. In some examples, these operations are performed by a system that includes a processor which executes a set of codes to control functional elements of a device, such as the one described in Fig. 5 described image generation model.

[0151] In step 1205, the system initializes an untrained model. In some cases, the operations of this step relate to or can be performed by a training component, as in the case of Fig. As described in section 5, initialization can include defining the model's architecture and setting initial values ​​for the model parameters. In some cases, initialization may include defining hyperparameters such as the number of layers, the resolution and channels of each layer block, the position of skip connections, and the like.

[0152] In step 1210, the system adds noise to a media element in N steps using a forward diffusion process. In some cases, the operations of this step are related to, or can be performed by, a training component, as in the case of Fig. 5. In some cases, the media element is, for example, a training image. In other cases, the forward diffusion process is a fixed process in which Gaussian noise is successively added to the media element (such as an original image). In latent diffusion models, the Gaussian noise can be successively added to features in a latent space.

[0153] In step 1215, the system predicts a media element for stage n-1 at each stage n, starting with stage N. In some cases, the operations of this step relate to or can be performed by a training component, as in relation to Fig. 5 described. In some cases, the media element is a synthetic image generated using the image generation model. For example, the reverse diffusion process can predict the noise added by the forward diffusion process, and the predicted noise can be removed from the noise input to obtain the predicted output. In some cases, an original media element is predicted at each stage of the training process.

[0154] In step 1220, the system compares the predicted media element (or feature) in stage n-1 with the medium in stage n-1. In some cases, for example, the system compares the synthetic image (or predicted image feature) in state n-1 with the ground-truth image (or ground-truth feature) in state n-1. In some cases, the operations of this step are related to, or can be performed by, a training component, as in the case of Fig. 5 described. For example, the diffusion model, given observed data x, can be trained to find the upper variational limit of the negative log-likelihood -log p. θ (x) to minimize the training data.

[0155] In step 1225, the system updates the model parameters based on the comparison. In some cases, the operations of this step relate to, or can be performed by, a training component, as in the case of... Fig. Section 5 describes this. For example, parameters of a U-Net can be updated using gradient descent. Time-dependent parameters of the Gaussian transitions can also be learned. Computing device

[0156] Fig. Figure 13 shows an example of a computing device 1300 according to aspects of the present disclosure. The example shown comprises the computing device 1300, a processor 1305, a memory subsystem 1310, a communication interface 1315, an I / O interface 1320, a user interface component 1325, and a channel 1330.

[0157] In some embodiments, the computing device 1300 is an example of, or includes, aspects of, the following: Fig. 1 and Fig. 5 image processing device described. In some embodiments, the computing device 1300 comprises the processor 1305, which can execute instructions stored in the memory subsystem 1310 to capture a text prompt describing a story, generate a first scene prompt and a second scene prompt based on the text prompt, and generate a first synthetic image and a second synthetic image based on the first scene prompt and the second scene prompt, respectively.

[0158] In some embodiments, the Processor 1305 comprises one or more processors. In some cases, the Processor 1305 is an intelligent hardware device (for example, a general-purpose processing component, a DSP, a CPU, a GPU, a microcontroller, an ASIC, an FPGA, a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, or a combination thereof). In some cases, the Processor 1305 is configured to operate a memory arrangement using a memory controller. In other cases, a memory controller is integrated into the Processor 1305. In some cases, the Processor 1305 is configured to execute computer-readable instructions stored in memory to perform various functions.In some embodiments, the 1305 processor includes specialized components for modem processing, baseband processing, digital signal processing, or transmission processing. The 1305 processor is an example of, or includes, aspects of the processor unit that relate to... Fig. 5 is described.

[0159] According to some embodiments, the 1310 memory subsystem comprises one or more memory devices. Examples of a memory device include RAM, ROM, or a hard disk. Examples of memory devices include solid-state storage and a hard disk drive. In some examples, the memory is used to store computer-readable, computer-executable software, including instructions that, when executed, cause a processor to perform various functions described herein. In some cases, the memory includes, among other things, a BIOS that controls basic hardware or software operations, such as interaction with peripheral components or devices. In some cases, a memory controller manages memory cells. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, memory cells in the 1310 memory subsystem store information in the form of a logical state.

[0160] According to some embodiments, the communication interface 1315 operates at the boundary between communicating units (such as the computing device 1300, one or more user devices, a cloud, and one or more databases) and the channel, and can record and process communications. In some cases, the communication interface 1315 is provided to enable a processing system coupled with a transceiver (for example, a transmitter and / or a receiver). In some examples, the transceiver is configured to send and receive signals for a communication device via an antenna. In some cases, a bus is used in the communication interface 1315.

[0161] In some embodiments, the I / O interface 1320 is controlled by an I / O controller to manage input and output signals for the computing device 1300. In some cases, the I / O interface 1320 manages peripheral devices that are not integrated into the computing device 1300. In some cases, the I / O interface 1320 represents a physical connection or port to an external peripheral device. In some cases, the I / O controller uses an operating system such as iOS®, ANDROID®, MS-DOS®, MS-WINDOWS®, OS / 2®, UNIX®, or other well-known operating systems. In some cases, the I / O controller represents or interacts with a modem, keyboard, mouse, touchscreen, or similar device. In some cases, the I / O controller is implemented as a component of a processor.In some cases, a user interacts with the 1300 computing device via the I / O interface 1320 or via hardware components controlled by the I / O controller. The I / O interface 1320 is an example of, or encompasses, aspects of the I / O module that relate to... Fig. 5 is described.

[0162] According to some embodiments, the user interface component 1325 enables a user to interact with the computing device 1300. In some cases, the user interface component 1325 includes an audio device, such as an external speaker system, an external display device, such as a screen, an input device (for example, a remotely controllable device that is connected to the user interface directly or via the I / O controller), or a combination thereof.

[0163] The performance of the devices, systems, and methods of this disclosure has been evaluated, and the results show that embodiments of this disclosure have achieved improved performance compared to prior art technology (for example, prior art image generation models). Example experiments show that the image processing device based on this disclosure outperforms prior art image generation models. Details of example applications based on embodiments of this disclosure are provided in relation to Fig. 3 described.

[0164] The description and drawings represent described example configurations and do not represent all implementations within the scope of the claims. For example, the operations and steps can be rearranged, combined, or otherwise modified. Furthermore, structures and devices can be represented as block diagrams to illustrate the relationship between components and avoid cluttering the diagram. Similar components or features may have the same name but different reference numbers corresponding to different figures.

[0165] Some modifications to the disclosure may be immediately apparent to experts in the field, and the principles defined herein may be applied to other variations without departing from the scope of the disclosure. Thus, the disclosure is not limited to the examples and configurations described herein, but is intended to have the broadest scope consistent with the principles and novel features disclosed herein.

[0166] The described methods can be implemented or executed by devices comprising a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or another programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general-purpose processor can be a microprocessor, a processor known from the prior art, a controller, a microcontroller, or a state machine. A processor can also be implemented as a combination of computing devices (for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration).Thus, the functions described herein can be implemented in hardware or software and executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions can be stored as instructions or code on a computer-readable medium.

[0167] Computer-readable media include both non-transitory computer-readable storage media and communication media, including any media that facilitate the transmission of code or data. A non-transitory storage medium can be any available medium accessible to a computer. For example, non-transitory computer-readable media can include RAM, ROM, EEPROM, compact disc (CD) or other optical storage media, magnetic disk storage, or any other non-transitory medium for carrying or storing data or code.

[0168] Connecting components can also be accurately described as computer-readable media. For example, if code or data is transmitted from a website, server, or other remote source via a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology such as infrared, radio, or microwave signals, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology is included in the definition of a medium. Combinations of media are also included in the scope of computer-readable media.

[0169] In this disclosure and the subsequent claims, the word "or" indicates an inclusive list, so, for example, the list of X, Y, or Z means: X or Y or Z or XY or XZ or YZ or XYZ. Likewise, the expression "based on" is not used to represent a closed set of conditions. For example, a step described as "based on condition A" may rely on both condition A and condition B. In other words, the expression "based on" is to be interpreted as meaning "at least partially based on." Likewise, the words "a" or "an" mean "at least one."