Generating and managing personalized images using machine learning techniques

By generating personalized narrative images through a stable diffusion model and deep learning algorithms, this technology solves the problem of multi-user image generation and sharing in existing technologies, and enables efficient editing and sharing of personalized images.

CN121693757APending Publication Date: 2026-03-17SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Current systems face challenges in generating personalized narrative images that depict user portraits and story scenarios, particularly in generating multi-user images and allowing for editing, enhancement, and sharing.

Method used

A stable diffusion model combined with deep learning algorithms is used to generate personalized narrative images based on text prompts and conditions. Machine learning techniques are used to generate and manage AI-generated personalized images. Positive and negative text prompts control the image generation process, and the model is trained on a large dataset.

Benefits of technology

It enables the generation, editing, and sharing of personalized narrative images by multiple users, improving image stability and editing efficiency, and enhancing the convenience of user participation and resource sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121693757A_ABST
    Figure CN121693757A_ABST
Patent Text Reader

Abstract

Various examples described herein support or provide operations to generate and manage personalized images using machine learning techniques, the operations including receiving portrait images of an entity; generating an identity representing the entity using the machine learning model; identifying a template including the text description of the scene and the condition; and generating a personalized image based on the identity and the template using a machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CLAIM OF PRIORITY

[0002] This application claims the benefit of priority to U.S. Patent Application Serial No. 18 / 232,938, filed August 11, 2023, which is incorporated by reference herein in its entirety. TECHNICAL FIELD

[0003] The present disclosure relates generally to data management, and more specifically, the various examples described herein provide systems, methods, techniques, instruction sequences, and devices that facilitate generating and managing personalized images using machine learning techniques. BACKGROUND

[0004] Current systems face challenges in generating personalized narrative images that depict user portraits and story scenes using machine learning models. It is challenging to generate personalized narrative images that show more than one user in a story scene and allow such images to be edited, augmented, and shared among other users. BRIEF DESCRIPTION OF DRAWINGS

[0005] In the drawings, which are not necessarily drawn to scale, like numerals can describe similar components in different views. To easily identify the discussion of any particular element or act, the most significant digit or digits in a figure reference number can correspond to the figure number following the particular element or act with the most significant digit or digits being the left most digit or digits of the reference number. Some non-limiting examples are illustrated in the drawings, in which:

[0006] Figure 1 is a diagrammatic representation of a networked environment in which the present disclosure can be deployed according to some examples.

[0007] Figure 2 is a diagrammatic representation of a messaging system having both client-side and server-side functionality according to some examples.

[0008] Figure 3 is a diagrammatic representation of a data structure maintained in a database according to some examples.

[0009] Figure 4 is a diagrammatic representation of a message according to some examples.

[0010] Figure 5 is a flow diagram illustrating an example method for generating and managing personalized images using machine learning techniques according to some examples.

[0011] Figure 6 is a flow diagram illustrating an example method for generating and managing personalized images using machine learning techniques according to some examples.

[0012] Figure 7is a flowchart showing an example method for generating and managing personalized images using machine learning techniques, according to some examples.

[0013] Figure 8 An example template for generating personalized images, according to some examples, is shown.

[0014] Figure 9 An example graphical user interface displayed by an example system during operation, according to some examples, is shown.

[0015] Figure 10 An example graphical user interface displayed by an example system during operation, according to some examples, is shown.

[0016] Figure 11 An example graphical user interface displayed by an example system during operation, according to some examples, is shown.

[0017] Figure 12 An example graphical user interface displayed by an example system during operation, according to some examples, is shown.

[0018] Figure 13 An example graphical user interface displayed by an example system during operation, according to some examples, is shown.

[0019] Figure 14 is a diagrammatic representation of a machine in the form of a computer system, according to some examples, within which a set of instructions can be executed for causing the machine to perform any one or more of the methods discussed herein.

[0020] Figure 15 is a block diagram showing a software architecture, wherein an example can be implemented. DETAILED DESCRIPTION

[0021] The description that follows includes systems, methods, techniques, instruction sequences, and computer machine program products, amongst other things, that embody illustrative examples of the present disclosure. In the description

[0022] Reference throughout this specification to “one example” or “an example” means that a particular feature, structure, or characteristic described in connection with the example is included in at least one example of the present subject matter. Thus, appearances of the phrases “in one example” or “an example” in various places throughout this specification are not necessarily all referring to the same example.

[0023] For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the subject matter. However, it will be apparent to one skilled in the art that the described examples can be practiced without the specific details presented herein or in a variety of sub-combinations of the features presented herein. Moreover, well-known features can be omitted or simplified in order not to obscure the described examples. Throughout this description, various examples can be presented in terms of sequences of actions, but these examples can be performed in other orders or in parallel. Furthermore, examples can be presented with reference to certain terminology that is used to describe particular structures and actions. Those skilled in the art will appreciate that the described examples are not limited to the described or implied nomenclature and that the scope of the described examples encompasses numerous alternatives, modifications and equivalents.

[0024] Current systems face challenges in generating personalized narrative images that depict a user’s portrait and a story scene using machine learning models. Specifically, images generated by current systems can be avatar-facing, where the image only depicts the user’s portrait. In contrast, the personalized narrative images described herein depict a user along with a story scene associated with one or more themes. As such, such images are more meaningful in a way that can be used to tell a story and / or convey a message. Moreover, current systems face additional challenges in generating personalized narrative images that show more than one user based on AI-generated identities and that allow such images to be edited (e.g., with media overlays), augmented, and shared among other users to facilitate user engagement with various resources provided by the system.

[0025] Various examples described herein can use state-of-the-art machine learning (ML) and artificial intelligence (AI) techniques to generate and manage AI-generated personalized images (also referred to as personalized narrative images or personalized images). As used herein, a machine learning model (e.g., a stable diffusion model) can include any predictive model generated (or trained based on) training data. A stable diffusion model can be controlled by one or more deep learning algorithms (e.g., ControlNet) based on various conditions described herein. Specifically, a stable diffusion model can generate a synthetic image (e.g., a personalized narrative image) based on a text prompt and one or more conditions (e.g., a sketch, a pose map, a depth map, a normal map, or a canny edge). A condition can also be referred to as a visual resource or a condition, as described herein. A condition such as a sketch can be a control image that defines the general shape and location of an entity (e.g., a person or an animal) in an image. In various implementations, a text prompt can be a positive text prompt or a negative text prompt. A combination of both positive and negative text prompts can be used to generate one or more templates described herein. Negative text prompts can provide more stable results and fewer artifacts. In various implementations, a positive text prompt can direct diffusion toward an image associated therewith, while a negative text prompt can direct diffusion away from an image associated therewith. In various implementations, a text prompt (e.g., a positive text prompt or a negative text prompt) can be updated to generate a modified template (e.g., template 806).

[0026] In various examples, a stable diffusion model can be trained based on a large dataset including a large number of images. Further, a stable diffusion model can generate a representation of an identity of an entity based on images of the entity (e.g., a portrait image, such as a selfie image). Once generated and trained, a machine learning model (e.g., a stable diffusion model) can receive one or more inputs (e.g., a text prompt, a condition), extract one or more features, and generate an output (e.g., a personalized narrative image) for the input based on the training of the model. Different types of machine learning models can include, but are not limited to, models trained using supervised learning, unsupervised learning, reinforcement learning, or deep learning (e.g., complex neural networks).

[0027] Networked computing environment

[0028] Figure 1This is a block diagram illustrating an example interactive system 100 for facilitating interaction over a network, such as exchanging text messages, making text-to-audio and video calls, or playing games. Interactive system 100 includes multiple user systems 102, each hosting multiple applications, including interactive client 104 and other applications 106. Each interactive client 104 is communicatively coupled to (e.g., hosted on corresponding other user systems 102) other instances of interactive client 104, interactive server system 110, and third-party server 112 via one or more communication networks including network 108 (e.g., the Internet). Interactive client 104 can communicate with locally hosted applications 106 using application programming interfaces (APIs).

[0029] Each user system 102 may include multiple user devices, such as mobile devices 114, wearable devices, and computer client devices 116, that are communicatively connected to exchange data and messages.

[0030] Interactive client 104 interacts with other interactive clients 104 and with interactive server system 110 via network 108. The data exchanged between interactive clients 104 (e.g., interactive 118) and between interactive client 104 and interactive server system 110 includes functions (e.g., commands to invoke functions) and payload data (e.g., text, audio, video or other multimedia data).

[0031] Interactive server system 110 provides server-side functionality to interactive client 104 via network 108. While some functions of interactive system 100 are described herein as being performed by interactive client 104 or interactive server system 110, whether a function resides within interactive client 104 or interactive server system 110 can be a design choice. For example, it may be technically preferred that specific technologies and functions are initially deployed within interactive server system 110, but that technology and functions are later migrated to interactive client 104 of user system 102, which has sufficient processing power.

[0032] The interactive server system 110 supports various services and operations provided to the interactive client 104. Such operations include sending data to and receiving data from the interactive client 104, and processing data generated by the interactive client 104. This data may include message content, client device information, geolocation information, media enhancements and overlays, message content persistence conditions, entity relationship information, and live event information. Data exchange within the interactive system 100 is initiated and controlled via functions available through the user interface (UI) of the interactive client 104.

[0033] Specifically, the focus now shifts to interactive server system 110. Application Programming Interface (API) server 120 is coupled to interactive server 122 and provides a programming interface to interactive server 122, enabling interactive client 104, other applications 106, and third-party server 112 to access the functionality of interactive server 122. Interactive server 122 is communicatively coupled to database server 124, facilitating access to database 126, which stores data associated with the interactions processed by interactive server 122. Similarly, web server 128 is coupled to interactive server 122 and provides a web-based interface to interactive server 122. For this purpose, web server 128 handles incoming network requests via Hypertext Transfer Protocol (HTTP) and several other related protocols.

[0034] Application Programming Interface (API) server 120 receives and sends interactive data (e.g., command and message payloads) between interactive server 122 and user system 102 (as well as, for example, interactive client 104 and other applications 106) and third-party server 112. Specifically, API server 120 provides a set of interfaces (e.g., routines and protocols) that interactive client 104 and other applications 106 can invoke or query to activate the functionality of interactive server 122. Application Programming Interface (API) server 120 discloses various functions supported by interaction server 122, including: account registration; login functionality; sending interactive data via interaction server 122 from one interactive client 104 to another interactive client 104; transmission of media files (e.g., images or videos) from interactive client 104 to interaction server 122; setting up collections of media data (e.g., stories); retrieval of the user's friend list in user system 102; retrieval of messages and content; adding and deleting entities (e.g., friends) in entity relationship graphs (e.g., entity graph 310); locating friends within entity relationship graphs; and opening application events (e.g., related to interactive client 104).

[0035] Interactive server 122 hosts multiple systems and subsystems, as detailed below. Figure 2 Describe it.

[0036] Application of links

[0037] Returning to interactive client 104, the features and functionality of external resources (e.g., linked application 106 or applet) are available to the user via the interface of interactive client 104. In this context, "external" refers to the fact that application 106 or applet is outside of interactive client 104. External resources are typically provided by third parties, but may also be provided by the creator or provider of interactive client 104. Interactive client 104 receives the user's selection of options to launch or access the features of such external resources. External resources may be application 106 installed on user system 102 (e.g., a "local application"), or a smaller version (e.g., a "applet") of an application hosted on user system 102 or located remotely on user system 102 (e.g., on a third-party server 112). A smaller version of an application includes a subset of the features and functionality of the application (e.g., the full-scale, native version of the application) and is implemented using markup language documentation. In some examples, a smaller version of an application (e.g., a "applet") is a web-based markup language version of the application and is embedded in interactive client 104. In addition to using markup language documentation (e.g.,... In addition to ml files, mini-programs can also incorporate scripting languages ​​(e.g., ...). .js files or .json files) and stylesheets (e.g.) (SS file).

[0038] In response to receiving a user's selection of an option to launch or access an external resource, interactive client 104 determines whether the selected external resource is a web-based external resource or a locally installed application 106. In some cases, application 106, locally installed on user system 102, can be launched independently of and separately from interactive client 104, for example, by selecting an icon corresponding to application 106 on the home screen of user system 102. A smaller version of such an application can be launched or accessed via interactive client 104, and in some examples, no part of the smaller application can be accessed outside of interactive client 104, or only a limited portion of the smaller application can be accessed outside of interactive client 104. A smaller application can be launched by interactive client 104 by receiving, for example, a markup language document associated with the smaller application from third-party server 112 and processing such a document.

[0039] In response to determining that the external resource is a locally installed application 106, the interactive client 104 instructs the user system 102 to launch the external resource by executing locally stored code corresponding to the external resource. In response to determining that the external resource is a web-based resource, the interactive client 104 communicates with a third-party server 112 (e.g.) to obtain a markup language document corresponding to the selected external resource. The interactive client 104 then processes the obtained markup language document to render the web-based external resource within the user interface of the interactive client 104.

[0040] Interactive client 104 can notify users of user system 102 or other users (e.g., "friends") associated with such users of one or more external resources. For example, interactive client 104 can provide participants in a conversation (e.g., a chat session) within interactive client 104 with notifications related to the current or recent use of external resources by one or more members of a group of users. One or more users can be invited to join an active external resource or to activate (in a group of friends) a recently used but currently inactive external resource. External resources can provide participants in the conversation, each using the corresponding interactive client 104, with the ability to share items, conditions, states, or locations within the external resource with one or more members of a group of users during a chat session. Shared items can be interactive chat cards, which chat members can use to interact, for example, activate the corresponding external resource, view specific information within the external resource, or take chat members to a specific location or state within the external resource. Within a given external resource, response messages can be sent to users on interactive client 104. External resources can selectively include different media items in the response based on the current context of the external resource.

[0041] Interactive client 104 can present a list of available external resources (e.g., application 106 or mini-program) to the user to launch or access a given external resource. This list can be presented in a context-sensitive menu. For example, the icons representing different applications (or mini-programs) within application 106 (or mini-program) can change based on how the user launches the menu (e.g., from a conversational interface or a non-conversational interface).

[0042] System Architecture

[0043] Figure 2This is a block diagram illustrating further details of an interactive system 100 according to some examples. Specifically, the interactive system 100 is shown as including an interactive client 104 and an interactive server 122. The interactive system 100 comprises multiple subsystems, which are supported on the client side by the interactive client 104 and on the server side by the interactive server 122. In some examples, these subsystems are implemented as microservices. A microservice subsystem (e.g., a microservice application) may have components that enable it to operate independently and communicate with other services. Example components of a microservice subsystem may include:

[0044] Functional logic: Functional logic implements the functions of the microservice subsystem and represents the specific capabilities or functions provided by the microservice.

[0045] API Interface: Microservices can communicate with other components using lightweight protocols such as REST or messaging through well-defined APIs or interfaces. The API interface defines the inputs and outputs of a microservice subsystem and how it interacts with other microservice subsystems of the interactive system 100.

[0046] Data storage: The microservice subsystem can be responsible for its own data storage, which can be in the form of a database, cache, or other storage mechanisms (e.g., using database server 124 and database 126). This allows the microservice subsystem to operate independently of other microservices in the interactive system 100.

[0047] Service discovery: Microservice subsystems can find and communicate with other microservice subsystems in the interacting system 100. The service discovery mechanism enables microservice subsystems to locate and communicate with other microservice subsystems in a scalable and efficient manner.

[0048] Monitoring and logging: Microservice subsystems may need to be monitored and logged to ensure availability and performance. Monitoring and logging mechanisms enable the tracking of the health and performance of microservice subsystems.

[0049] In some examples, the interactive system 100 may employ a monolithic architecture, a service-oriented architecture (SOA), a function-as-a-service (FaaS) architecture, or a modular architecture:

[0050] The example subsystem is discussed below.

[0051] The image processing system 202 provides various functions that enable users to capture and enhance (e.g., annotate or otherwise modify or edit) media content associated with a message.

[0052] The camera device system 204 includes (e.g., in a camera device application) control software that interacts with and controls the hardware camera device hardware of the user system 102 (e.g., directly or via operating system controls) to modify and enhance real-time images captured and displayed via the interactive client 104.

[0053] Enhancement system 206 provides functionality related to generating and publishing enhancements (e.g., media overlays) for images captured in real-time by the camera device of user system 102 or retrieved from the memory of user system 102. For example, enhancement system 206 can operatively select, present, and display media overlays (e.g., image filters or image lenses) to interactive client 104 to enhance real-time images received via camera device system 204 or stored images retrieved from memory 502 of user system 102. These enhancements are selected and presented to the user of interactive client 104 by enhancement system 206 based on multiple inputs and data, such as:

[0054] The geographical location of user system 102; and

[0055] User entity relationship information of users in user system 102.

[0056] Enhancements may include audio and visual content as well as visual effects. Examples of audio and visual content include images, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects may be applied to media content items (e.g., photos or videos) at user system 102 to be transmitted in messages or applied to video content, such as a video content stream or feed transmitted from interactive client 104. Therefore, image processing system 202 can interact with and support various subsystems of communication system 208, such as messaging system 210 and video communication system 212.

[0057] Media overlays may include text or image data that can be superimposed on photographs taken by user system 102 or video streams generated by user system 102. In some examples, media overlays may be location overlays (e.g., Venice Beach), names of live events, or names of businesses (e.g., beach cafes). In other examples, image processing system 202 uses the geographic location of user system 102 to identify the media overlay, which includes the name of a business at the geographic location of user system 102. Media overlays may include other tags associated with businesses. Media overlays may be stored in database 126 and accessed through database server 124.

[0058] Image processing system 202 provides a user-based publishing platform that allows users to select geographic locations on a map and upload content associated with those locations. Users can also specify which media overlays should be provided to other users. Image processing system 202 generates a media overlay that includes the uploaded content and associates it with the selected geographic location.

[0059] The augmented creation system 214 supports augmented reality developer platforms and includes applications that enable content creators (such as artists and developers) to create and publish interactive clients 104, such as augmented reality experiences. The augmented creation system 214 provides content creators with a library of built-in features and tools, including, for example, custom shaders, tracking technologies, and templates.

[0060] In some examples, enhancement creation system 214 provides a merchant-based publishing platform that allows merchants to select specific enhancements associated with a geographic location through a bidding process. For example, enhancement creation system 214 associates the media overlay of the highest bidder with the corresponding geographic location for a predefined amount of time.

[0061] Communication system 208 is responsible for enabling and processing various forms of communication and interaction within interactive system 100, and includes messaging system 210, audio communication system 216, and video communication system 212. Messaging system 210 is responsible for enabling temporary or time-limited access to content by interactive client 104. Messaging system 210 incorporates multiple timers (e.g., within a short-lived timer system) that selectively enable access (e.g., for presentation and display) of messages and associated content via interactive client 104 based on duration and display parameters associated with a message or set of messages (e.g., a story). Audio communication system 216 enables and supports audio communication (e.g., real-time audio chat) between multiple interactive clients 104. Similarly, video communication system 212 enables and supports video communication (e.g., real-time video chat) between multiple interactive clients 104.

[0062] User management system 218 is operationally responsible for managing user data and profiles, and maintaining entity information about users and relationships between users in interactive system 100 (e.g., stored in entity table 308, entity diagram 310, and profile data 302).

[0063] The collection management system 220 is operationally responsible for managing collections or sets of media (e.g., collections of text, image, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into “event libraries” or “event stories.” Such collections can be made available for a specified time period (e.g., the duration of the event to which the content relates). For example, content related to a concert can be made available as a “story” for the duration of the concert. The collection management system 220 can also be responsible for publishing icons to the user interface of the interactive client 104, which provide notifications for specific collections. The collection management system 220 includes curation functionality that allows collection managers to manage and curate specific content collections. For example, the curation interface enables event organizers to curate collections of content related to a specific event (e.g., removing inappropriate content or redundant messages). Additionally, the collection management system 220 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, users may be paid to include user-generated content in the collection. In such a situation, the collection management system 220 operates to automatically pay such users for access to its content.

[0064] External resource system 222 provides interactive client 104 with an interface to communicate with remote servers (e.g., third-party server 112) to launch or access external resources, i.e., applications or applets. Each third-party server 112 hosts applications or smaller versions of applications (e.g., games, utilities, payment, or ride-sharing apps) based on markup languages ​​(e.g., HTML5). Interactive client 104 can launch web-based resources (e.g., applications) by accessing HTML5 files from the third-party server 112 associated with the web-based resource. The application hosted by third-party server 112 is programmed in JavaScript using a software development kit (SDK) provided by interactive server 122. The SDK includes application programming interfaces (APIs) with functions that can be called or invoked by the web-based application. Interactive server 122 hosts a JavaScript library that provides access to a given external resource for specific user data of interactive client 104. HTML5 is an example of a technology used for programming games, but applications and resources programmed using other technologies can be used.

[0065] Interactive client 104 presents a graphical user interface (GUI) for an external resource (e.g., a login page or title screen). During, before, or after presenting the login page or title screen, interactive client 104 determines whether the initiated external resource has previously been authorized to access the user data of interactive client 104. In response to determining that the initiated external resource has previously been authorized to access the user data of interactive client 104, interactive client 104 presents another GUI for the external resource, including its functionality and characteristics. In response to determining that the initiated external resource has not previously been authorized to access the user data of interactive client 104, after displaying the login page or title screen of the external resource for a threshold time period (e.g., 3 seconds), interactive client 104 slides up a menu for authorizing the external resource to access user data (e.g., animating the menu to appear from the bottom of the screen to the middle or other part of the screen). The menu identifies the type of user data that the external resource will be authorized to use. In response to receiving a user's selection of the accept option, interactive client 104 adds the external resource to the list of authorized external resources and allows the external resource to access user data from interactive client 104. The interactive client 104 authorizes external resources to access user data under the OAuth 2 framework.

[0066] Interactive client 104 controls the type of user data shared with external resources based on the type of authorized external resource. For example, access to a first type of user data (e.g., two-dimensional avatars of users with or without different avatar characteristics) is provided to external resources including full-scale applications (e.g., application 106). As another example, access to a second type of user data (e.g., payment information, two-dimensional avatars of users, three-dimensional avatars of users, and avatars with various avatar characteristics) is provided to external resources including smaller versions of applications (e.g., web-based versions of applications). Avatar characteristics include different ways of customizing the appearance and feel of an avatar, such as different poses, facial features, clothing, etc.

[0067] Artificial intelligence and machine learning system 224 (also referred to as AI and ML system 224 or system 224) provides various services to different subsystems within interactive system 100. For example, AI and machine learning system 224 operates in conjunction with image processing system 202 and camera device system 204 to analyze images and extract information such as objects, text, or faces. This information can then be used by image processing system 202 to enhance, filter, or manipulate the images. Enhancement system 206 can use AI and machine learning system 224 to generate enhanced content and augmented reality experiences, such as adding virtual objects or animations to real-world images. Communication system 208 and messaging system 210 can use AI and machine learning system 224 to analyze communication patterns and provide insights into how users interact with each other, and provide intelligent message classification and tagging, such as classifying messages based on sentiment or topic. AI and machine learning system 224 can also provide chatbot functionality for messaging interactions 118 between user systems 102 and between user systems 102 and interactive server system 110. The artificial intelligence and machine learning system 224 can also work with the audio communication system 216 to provide speech recognition and natural language processing capabilities, allowing users to interact with the interactive system 100 using voice commands.

[0068] In various examples, the artificial intelligence and machine learning system 224 (also referred to as AI and ML system 224 or system 224) uses one or more machine learning models (e.g., stable diffusion models) to generate and manage personalized images (e.g., AI-generated personalized narrative images). Personalized images, for example... Figure 10 The example personalized image 1006 shown can be a composite image generated based on text prompts and one or more conditions (e.g., sketch, pose map, depth map, normal map, Canney edge). Figure 8 Example pose diagram 802, example text prompt 804, and example template 806 are shown.

[0069] Data Architecture

[0070] Figure 3 This is a schematic diagram illustrating a data structure 300 that can be stored in a database 304 of an interactive server system 110, according to certain examples. Although the contents of the database 304 are shown as including multiple tables, it should be understood that the data can be stored in other types of data structures (e.g., as an object-oriented database).

[0071] Database 304 includes message data stored in message table 306. For any given message, this message data includes at least message sender data, message receiver (or recipient) data, and a payload. See below for reference. Figure 3Additional details describe information that can be included in a message and is contained within message data stored in message table 306.

[0072] Entity table 308 stores entity data and (e.g., by reference) links to entity diagram 310 and profile data 302. Entities for which records are maintained within entity table 308 can include individuals, company entities, organizations, objects, locations, events, etc. Regardless of entity type, any entity whose data is stored by the interactive server system 110 can be an identifiable entity. Each entity is provided with a unique identifier and an entity type identifier (not shown).

[0073] In various examples, one or more portrait images (e.g., selfies) of an entity (e.g., a person or animal) can be obtained to generate an identity representing the entity (e.g., an AI-generated identity).

[0074] Entity graph 310 stores information about relationships and associations (or connections) between entities. As an example only, such relationships can be social, professional (e.g., working in a common company or organization), interest-based, or activity-based. Some relationships between entities can be one-way, such as an individual user subscribing to digital content from a business or publishing user (e.g., a newspaper or other digital media channel or brand). Other relationships can be two-way, such as the "friendship" between individual users of interactive system 100.

[0075] Certain licenses and relationships can be attached to each relationship, and also to each direction of the relationship. For example, a two-way relationship (e.g., a friend relationship between individual users) can include authorization for the publication of digital content items between individual users, but certain restrictions or filters can be imposed on the publication of such digital content items (e.g., based on content characteristics, location data, or time of day data). Similarly, a subscription relationship between an individual user and a business user can impose varying degrees of restrictions on the publication of digital content from the business user to the individual user, and can significantly restrict or prevent the publication of digital content from the individual user to the business user. As an example of an entity, a specific user can record certain restrictions in a record for that entity within entity table 308 (e.g., via privacy settings). Such privacy settings can be applied to all types of relationships in the context of interaction system 100, or selectively applied to certain types of relationships.

[0076] Profile data 302 stores various types of profile data about a specific entity. Profile data 302 can be selectively used and presented to other users of the interactive system 100 based on privacy settings specified by the specific entity. In the case of an individual, profile data 302 includes, for example, a username, phone number, address, settings (such as notification and privacy settings), and an avatar representation (or a set of such avatar representations) selected by the user. A specific user can then selectively include one or more of these avatar representations within the content of messages transmitted via the interactive system 100 and on a map interface displayed to other users by the interactive client 104. The set of avatar representations may include “status avatars,” which present a graphical representation of a status or activity that the user can choose to convey at a specific time.

[0077] In various examples, profile data 302 associated with an entity (e.g., a person or animal) may include an AI-generated identity representing the entity and one or more personalized narrative images generated based on that identity and one or more templates described herein.

[0078] Database 304 also stores enhancement data, such as overlays or filters, in enhancement table 312. Enhancement data is associated with and applied to videos (whose data is stored in video table 314) and images (whose data is stored in image table 316), such as personalized images.

[0079] In some examples, filters are overlays displayed as superimposed on images (such as personalized images) or videos during presentation to the recipient user. Filters can be of various types, including those selected by the user from a set of filters presented to the sending user from the interactive client 104 when the sending user composes a message. Other types of filters include geolocation filters (also known as geographic filters), which can be presented to the sending user based on geographic location. For example, the interactive client 104 may present geolocation filters specific to nearby or particular locations within the user interface based on geolocation information determined by the Global Positioning System (GPS) unit of the user system 102.

[0080] Another type of filter is a data filter, which can be selectively presented to the sending user by the interactive client 104 based on other inputs or information collected by the user system 102 during the message creation process. Examples of data filters include the current temperature at a specific location, the sending user's current travel speed, the battery life of the user system 102, or the current time.

[0081] Other augmented data that can be stored within image table 316 includes, for example, augmented reality content items corresponding to the application of a "lens" or augmented reality experience. Augmented reality content items can be real-time effects and sounds that can be added to images or videos.

[0082] Collection table 318 stores data related to collections of messages and associated image, video, or audio data, compiled into collections (e.g., stories or galleries). The creation of a specific collection can be initiated by a specific user (e.g., each user for whom records are maintained in entity table 308). Users can create "personal stories," which are collections of content that the user has already created and sent / broadcast. For this purpose, the user interface of interactive client 104 may include user-selectable icons to allow the sending user to add specific content to his or her personal story.

[0083] Collections can also constitute "live stories," which are collections of content from multiple users created manually, automatically, or using a combination of manual and automatic technologies. For example, a "live story" can constitute a curated stream of user-submitted content from different locations and events. For instance, users with location services enabled on their client devices and who are at a common location event at a specific time can be presented with the option to contribute content to a specific live story via the user interface of interactive client 104. Users can be identified by interactive client 104 based on their location. The end result is a "live story" told from a community perspective.

[0084] Another type of content collection is called a "location story," which allows users of user system 102 located in a specific geographic location (e.g., within a college or university campus) to contribute to a specific collection. In some examples, contributions to a location story may employ a second level of verification to confirm that the end user belongs to a specific organization or other entity (e.g., is a student on a university campus).

[0085] As described above, video table 314 stores video data, which in some examples is associated with messages whose records are maintained within message table 306. Similarly, image table 316 stores image data associated with messages whose message data is stored in entity table 308. Entity table 308 can associate various enhancements from enhancement table 312 with various images and videos stored in image table 316 and video table 314.

[0086] Data communication architecture

[0087] Figure 4This is a schematic diagram illustrating the structure of message 400 according to some examples, generated by interactive client 104 for transmission to another interactive client 104 via interactive server 122. The content of a particular message 400 is used to populate message table 306 within database 304 accessible by interactive server 122. Similarly, the content of message 400 is stored in memory as "in-transit" or "in-flight" data for user system 102 or interactive server 122. Message 400 is shown to include the following example components:

[0088] Message Identifier 402: A unique identifier that identifies message 400.

[0089] Message text payload 404: The text to be generated by the user via the user interface of user system 102 and included in message 400.

[0090] Message image payload 406: Image data captured by the camera device component of user system 102 or retrieved from the memory component of user system 102 and included in message 400. The image data for the sent or received message 400 can be stored in image table 316.

[0091] Message video payload 408: Video data captured by the camera device component or retrieved from the memory component of the user system 102 and included in message 400. The video data for the sent or received message 400 can be stored in image table 316.

[0092] Message audio payload 410: Audio data captured by the microphone or retrieved from the memory component of the user system 102 and included in message 400.

[0093] Message enhancement data 412: Enhancement data (e.g., filters, labels, or other annotations or enhancements) represents enhancements to be applied to the message image payload 406, message video payload 408, or message audio payload 410 of message 400. Enhancement data for the sent or received message 400 can be stored in enhancement table 312.

[0094] Message duration parameter 414: Parameter value, in seconds, indicates the amount of time that the content of the message (e.g., message image payload 406, message video payload 408, message audio payload 410) is to be presented to the user via the interactive client 104 or is accessible to the user.

[0095] Message geolocation parameter 416: Geographic location data (e.g., latitude and longitude coordinates) associated with the message's content payload. Multiple message geolocation parameter 416 values ​​may be included in the payload, each of which is associated with a content item included in the content (e.g., a specific image within the message image payload 406 or a specific video within the message video payload 408).

[0096] Message Story Identifier 418: An identifier value that identifies one or more sets of content (e.g., "story" identified in set table 318) associated with a specific content item in the message image payload 406 of message 400. For example, multiple images within the message image payload 406 may each be associated with multiple sets of content using an identifier value.

[0097] Message Tag 420: Each message 400 can be labeled with multiple tags, each of which indicates the subject of the content included in the message payload. For example, in the case where a specific image depicts an animal (e.g., a lion) is included in the message image payload 406, a tag value can be included within the message tag 420 indicating the relevant animal. Tag values ​​can be manually generated based on user input, or can be automatically generated using, for example, image recognition.

[0098] Message sender identifier 422: An identifier (e.g., a message sending and receiving system identifier, email address, or device identifier) ​​indicating the user of the user system 102 on which message 400 is generated and from which message 400 is sent.

[0099] Message recipient identifier 424: An identifier (e.g., a message sending and receiving system identifier, email address, or device identifier) ​​indicating the user of the user system 102 to which message 400 is addressed.

[0100] The content (e.g., values) of each component of message 400 can be pointers to locations in tables where content data values ​​are stored. For example, image values ​​in message image payload 406 can be pointers to locations (or their addresses) within image table 316. Similarly, values ​​in message video payload 408 can point to data stored in image table 316, values ​​in message enhancement data 412 can point to data stored in enhancement table 312, values ​​in message story identifier 418 can point to data stored in set table 318, and values ​​in message sender identifier 422 and message receiver identifier 424 can point to user records stored in entity table 308.

[0101] In various examples, personalized images can also be edited (e.g., adding media overlays) and sent as one or more messages described herein to (or shared with) other users' client devices. For example, the content of message 400, which includes a personalized image, can be used to populate message table 306 stored in database 304 accessible by interactive server 122.

[0102] Figure 5 This is a flowchart illustrating an example method 500 for generating and managing personalized images using machine learning techniques, based on some examples. It will be understood that, based on some examples, a machine can execute the example methods described herein. For example, method 500 can be derived from... Figure 2 The artificial intelligence and machine learning system 224 described herein, or its individual components, may execute the operations. The various methods described herein may be executed by one or more hardware processors (e.g., a central processing unit or graphics processing unit) of a computing device (e.g., a desktop computer, server, laptop computer, mobile phone, tablet computer, etc.), which may be part of a cloud-based computing system. The example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of an electronic circuit system. For example, the operation of method 500 may be represented by executable instructions that, when executed by the processor of the computing device, cause the computing device to perform method 500. Depending on the example, the operations of the example methods described herein may be repeated in different ways or involve interventional operations not shown. While the operations of the example methods may be depicted and described in a specific order, the order in which the operations are performed may vary within the examples, including the parallel execution of specific operations.

[0103] At operation 502, the processor receives one or more portrait images of an entity from the device. The portrait images may be selfies (e.g., still images of an entity) captured by a camera device (e.g., a front-facing camera, a rear-facing camera) communicatively coupled to a user device associated with user system 102.

[0104] At operation 504, the processor uses one or more machine learning models (e.g., a stable diffusion model) to generate identities representing entities based on one or more portrait images.

[0105] At operation 506, the processor identifies one or more templates based on a user account (or user profile) associated with the device. A template may include one or more text descriptions of a scene (also referred to as text hints) and one or more conditions (e.g., sketches, pose maps, depth maps, normal maps, or Canney edges) that control the generation of personalized images (e.g., personalized narrative images) described herein. A scene may refer to a setting that tells a story and / or conveys a message. In various implementations, the text descriptions of the scene (also referred to as text hints) may depict entities (e.g., people) involved in actions and / or dialogue.

[0106] At operation 508, the processor uses one or more machine learning models (e.g., a stable diffusion model) to generate one or more personalized images based on identity and one or more templates.

[0107] Although not shown, method 500 may include operations that can be displayed (or cause to be displayed) a graphical user interface by a hardware processor. For example, operations may cause a client device (e.g., mobile device 114) to display a graphical user interface for generating and managing one or more personalized images. This operation for displaying the graphical user interface may be separate from operations 502 to 508, or alternatively, may form part of one or more of operations 502 to 508.

[0108] Figure 6 This is a flowchart illustrating an example method 600 for generating and managing personalized images using machine learning techniques, based on some examples. It will be understood that, based on some examples, the example methods described herein can be executed by a machine. For example, method 600 can be generated by... Figure 2The artificial intelligence and machine learning system 224 described herein, or its individual components, may execute the operations. The various methods described herein may be executed by one or more hardware processors (e.g., central processing unit or graphics processing unit) of a computing device (e.g., desktop computer, server, laptop computer, mobile phone, tablet computer, etc.), which may be part of a cloud-based computing system. The example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of an electronic circuit system. For example, the operation of method 600 may be represented by executable instructions that, when executed by the processor of the computing device, cause the computing device to perform method 600. Depending on the example, the operations of the example methods described herein may be repeated in different ways or involve interventional operations not shown. While the operations of the example methods may be depicted and described in a specific order, the order in which the operations are performed may vary within the examples, including the parallel execution of specific operations.

[0109] At operation 602, the processor causes a user interface to be displayed on a device (e.g., mobile device 114). The user interface (e.g., user interface 904) includes an invitation (e.g., invitation 906) for generating multiple personalized images associated with a theme (e.g., ancient mythology). Figure 9 As shown, Invitation 906 corresponds to a user-selectable user interface (UI) element associated with a text indicator declared as "Generate a pack, first 8 dreams free". The dream can be referred to as the personalized image described in this article.

[0110] At operation 604, the processor detects an indication that the user has selected an invitation (e.g., invitation 906).

[0111] At operation 606, the processor identifies a fixed number of templates from a plurality of templates associated with a theme. For example, a fixed number of templates (e.g., 8 templates) can be randomly selected from a plurality of templates (e.g., 16 templates) included in a set of templates. A set of templates may correspond to a single theme.

[0112] At operation 608, the processor uses one or more machine learning models (e.g., a stable diffusion model) to generate a fixed number of personalized images based on a fixed number of templates and the identities representing the entities. In various examples, each personalized image is generated based on a single template that includes one or more textual descriptions and one or more conditions. Different personalized images can be generated based on the same template for the same entity.

[0113] Although not shown, method 600 may include operations that can be displayed (or cause to be displayed) a graphical user interface by a hardware processor. For example, operations may cause a client device (e.g., mobile device 114) to display a graphical user interface for generating and managing one or more personalized images. This operation for displaying the graphical user interface may be separate from operations 602 to 608, or alternatively, may form part of one or more of operations 602 to 608.

[0114] Figure 7 This is a flowchart illustrating an example method for generating and managing personalized images using machine learning techniques, based on some examples. This is a flowchart illustrating an example method 700 for generating and managing personalized images using machine learning techniques, based on some examples. It will be understood that, based on some examples, a machine can execute the example methods described herein. For example, method 700 can be derived from... Figure 2 The artificial intelligence and machine learning system 224 described herein, or its individual components, may execute the operations. The various methods described herein may be executed by one or more hardware processors (e.g., a central processing unit or graphics processing unit) of a computing device (e.g., a desktop computer, server, laptop computer, mobile phone, tablet computer, etc.), which may be part of a cloud-based computing system. The example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of an electronic circuit system. For example, the operation of method 700 may be represented by executable instructions that, when executed by the processor of the computing device, cause the computing device to perform method 700. Depending on the example, the operations of the example methods described herein may be repeated in different ways or involve interventional operations not shown. While the operations of the example methods may be depicted and described in a specific order, the order in which the operations are performed may vary within the examples, including the parallel execution of specific operations.

[0115] At operation 702, the processor identifies the remaining template set from multiple templates associated with the theme (e.g., ancient mythology).

[0116] At operation 704, the processor causes a user interface (e.g., user interface 1002) to be displayed on a device (e.g., mobile device 114). The user interface may include an invitation (e.g., invitation 1008) for unlocking the availability of the remaining template set.

[0117] At operation 706, the processor detects an instruction from the user to select an invitation (e.g., invitation 1008) via a device (e.g., mobile device 114).

[0118] At operation 708, the processor initiates a transaction process to unlock the availability of the remaining template set in exchange for an element representing a cash payment, such as points (e.g., in-app currency or rewards). In various examples, the element could be a cash payment received via an external resource (e.g., an app or mini-program), such as... Figure 11 As shown.

[0119] At operation 710, in response to detecting the completion of a transaction, the processor uses one or more machine learning models to generate a personalized image based on the remaining template set and the identity of the represented entity. For example... Figure 10 As shown, the user interface 1002 includes multiple images. Some images, such as personalized images 1010 and 1012, are shown with special outlines (e.g., lighting outlines) to indicate that they were generated based on a pre-selected template. In various examples, personalized images can be distinguished from generic images (i.e., images not generated based on identity) based on color. Users can choose to unlock the availability of generic images by completing the transaction process described herein.

[0120] Although not shown, method 700 may include operations that can be displayed (or cause to be displayed) a graphical user interface by a hardware processor. For example, operations may cause a client device (e.g., mobile device 114) to display a graphical user interface for generating and managing one or more personalized images. This operation for displaying the graphical user interface may be separate from operations 702 to 710, or alternatively, may form part of one or more of operations 702 to 710.

[0121] Figure 8 Example templates for generating personalized images are shown, based on several examples. As shown, Figure 8 Example pose graph 802, example text cue 804, and example template 806 are illustrated. Template 806 may include one or more text descriptions of a scene and / or entities, as well as one or more conditions (e.g., sketches, pose graphs, depth graphs, normal graphs, or Canney edges). Conditions (e.g., pose graph 802) may control the generation of personalized images (e.g., personalized narrative images) described herein.

[0122] In various implementations, the text hints can be positive (e.g., text hint 804) or negative. A combination of both positive and negative text hints can be used to generate one or more templates described herein (e.g., template 806). Negative text hints can provide more stable results and fewer artifacts. In various implementations, positive text hints can guide diffusion toward the image associated with them, while negative text hints can guide diffusion away from it. Example negative text hints could be “distortion, warping, disfigurement, poor rendering, anatomical malformation, anatomical error, extra limbs, missing limbs, floating limbs, mutated hands and fingers, severed limbs, mutation, mutation, nude, nude breasts, ugly, disgusting, blurry, amputation, unclear”. In various implementations, the text hints (e.g., positive or negative text hints) can be modified to generate modified templates (e.g., template 806).

[0123] In various examples, pose graph 802 allows machine learning models (e.g., stable diffusion models) to replicate the same pose of an individual subject (e.g., a person) based on the locations of keypoints detected in the image (e.g., keypoints 808 and 810). When generating personalized images, the machine learning model can generate the texture and surface of the individual subject based on other conditions described herein.

[0124] In various examples, a depth map (not shown) can be an image (e.g., a grayscale image or a color image) that includes information related to the distance from the surface of a scene object to the viewpoint (e.g., illumination proportional to the distance from the camera device, brightness related to the distance from the nominal focal plane). When generating personalized images, a machine learning model (e.g., a stable diffusion model) can follow the structure of the depth map and fill in details as needed.

[0125] Figure 9 An example graphical user interface displayed by an example system (e.g., system 224) during operation is shown, based on some examples. Figure 9 This includes user interface 902 and user interface 904. User interface 902 displays two themes (i.e., ancient mythology and various parts of the world). Each theme corresponds to a set of multiple templates that can be set together by a common theme (e.g., theme 1018) or by setting a set of templates. A set of templates can refer to a collection of templates randomly selected from the set.

[0126] In various examples, a template may include an asset used to generate a personalized image (e.g., personalized image 1006). An unlocked template may refer to a template upon which a personalized image is generated.

[0127] In various implementations, a first set of templates may be randomly selected from a given theme (e.g., theme 1018) and offered free of charge to the user to generate personalized images. Any remaining templates from a given theme may be grouped into sets and offered to the user for purchase. One or more pre-selected templates may be determined by system 224 or an administrative user associated with system 224. There should be no duplicate templates between sets of images generated for the same theme.

[0128] In various implementations, a user can obtain a maximum number of template sets from each set. For example, a set may contain 32 templates. A user should be able to unlock all templates in the set by obtaining 4 sets of templates (e.g., 8 templates per set).

[0129] Figure 10 An example graphical user interface displayed by an example system (e.g., system 224) during operation is shown, based on some examples. Figure 10 This includes user interfaces 1002 and 1004. User interface 1002 includes multiple templates associated with theme 1018 (i.e., ancient mythology). Some templates can be randomly selected and are free to use for generating personalized images. In contrast, the rest of the templates (including pre-selected templates) are only available through the completion of a transaction process, such as... Figure 11 As shown, the user interface 1004 displays a personalized image 1006 generated based on a template associated with the theme "Ancient Mythology" and the identity of the representing entity (e.g., the user of mobile device 114). The personalized image 1006 can be edited (e.g., by adding media overlays), enhanced (e.g., by applying enhancements described herein), and sent (as a brief message) to other users by activating a user-selectable element 1014. The personalized image 1006 can also be generated as a story (e.g., a personal story or a live story) by activating a user-selectable element 1016.

[0130] Figure 11 Examples of graphical user interfaces displayed by an example system (e.g., system 224) during operation are shown, based on some examples. As shown, user interface 1102 illustrates an example user interface for completing a transaction process. Users of devices (e.g., mobile device 114) may choose to unlock the availability of certain templates in exchange for elements (e.g., element 1104) collected via external resources (e.g., applications or applets).

[0131] Figure 12An example graphical user interface (GUI) displayed by an example system (e.g., system 224) during operation is shown, based on some examples. As shown, the GUI 1202 includes a user-selectable element 1212 that allows the selection of a friend (e.g., friend 1206) with an identity generated by the artificial intelligence and machine learning system 224. After selection, the user can instruct the artificial intelligence and machine learning system 224 to generate a personalized image based on a template (e.g., template 1208) that includes multiple entities.

[0132] Figure 13 An example graphical user interface (GUI) displayed by an example system (e.g., system 224) during operation is shown, according to some examples. As shown, the GUI 1302 includes a user-selectable element 1306 that invites the user to generate a personalized image (e.g., personalized image 1006). After selecting the user-selectable element 1306, another GUI (not shown) may be generated to provide means for collecting a portrait image of the user (e.g., a selfie). The portrait image may be used to generate one or more AI-generated identities representing the user. In various examples, identities may be generated for animals or other objects.

[0133] As shown, the user interface 1304 includes user-selectable elements (e.g., user-selectable element 1308 and user-selectable element 1310) that allow users to change or remove personal identities (e.g., AI selfies). Users can also view personalized images of connected users and have artificial intelligence and machine learning systems 224 generate images based on templates from the same set of images.

[0134] In various examples, users can unlock the availability of a specific template and generate personalized images with a specific friend. In some examples, if a user wants to generate a personalized image with another friend using the same template, they must unlock (e.g., purchase) the availability of the specific template again.

[0135] Machine architecture

[0136] Figure 14This is a schematic representation of machine 1400, within which instructions 1402 (e.g., software, program, application, app, or other executable code) can be executed to cause machine 1400 to perform any or more of the methods discussed herein. For example, instructions 1402 can cause machine 1400 to perform any or more of the methods described herein. Instructions 1402 transform a general, unprogrammed machine 1400 into a specific machine 1400 programmed to perform the described and illustrated functions in the described manner. Machine 1400 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 1400 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1400 may include, but is not limited to, server computers, client computers, personal computers (PCs), tablets, laptops, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web appliances, network routers, network switches, bridges, or any machine capable of sequentially or otherwise executing instructions 1402 that specify actions to be taken by machine 1400. Furthermore, while a single machine 1400 is shown, the term "machine" should also be considered as a collection of machines that individually or jointly execute instructions 1402 to perform any one or more of the methods discussed herein. For example, machine 1400 may include user system 102 or any of a plurality of server devices forming part of interactive server system 110. In some examples, machine 1400 may also include both client and server systems, wherein certain operations of a particular method or algorithm are performed on the server side and certain operations of said particular method or algorithm are performed on the client side.

[0137] Machine 1400 may include a processor 1404, a memory 1406, and an input / output (I / O) unit 1408 that can be configured to communicate with each other via a bus 1410. In the example, processor 1404 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processors 1412 and 1414 that execute instruction 1402. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Figure 14 Multiple processors 1404 are shown, but machine 1400 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0138] Memory 1406 includes main memory 1416, static memory 1418, and memory cell 1420, all accessible by processor 1404 via bus 1410. Main memory 1406, static memory 1418, and memory cell 1420 store instructions 1402 embodying any one or more of the methods or functions described herein. Instructions 1402 may also reside wholly or partially in main memory 1416, static memory 1418, machine-readable medium 1422 within memory cell 1420, at least one processor of processor 1404 (e.g., the processor's cache memory), or any suitable combination thereof, during execution by machine 1400.

[0139] I / O component 1408 may include a wide variety of components for receiving input, providing output, generating output, transmitting information, exchanging information, capturing measurement results, etc. The specific I / O component 1408 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It will be understood that I / O component 1408 may include... Figure 14Many other components are not shown. In various examples, I / O component 1408 may include user output component 1424 and user input component 1426. User output component 1424 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tube (CRT) displays), acoustic components (e.g., speakers), tactile components (e.g., vibrating motors, resistive mechanisms), other signal generators, etc. User input component 1426 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), tactile input components (e.g., physical buttons, touchscreens or other tactile input components that provide touch gestures or the position and force of a touch), audio input components (e.g., microphones), etc.

[0140] In another example, I / O component 1408 may include biometric component 1428, motion component 1430, environmental component 1432, or position component 1434, as well as various other components. For example, biometric component 1428 includes components for detecting expressions (e.g., hand gestures, facial expressions, vocal expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brainwaves), and recognizing a person (e.g., speech recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Biometric component may include a brain-computer interface (BMI) system that allows communication between the brain and external devices or machines. This can be achieved by recording brain activity data, converting that data into a format that can be understood by a computer, and then using the resulting signals to control the device or machine.

[0141] Examples of BMI technology types include:

[0142] BMI based on electroencephalography (EEG) uses electrodes placed on the scalp to record electrical activity in the brain.

[0143] Invasive BMI, which uses electrodes surgically implanted into the brain.

[0144] Optogenetics BMI uses light to control the activity of specific nerve cells in the brain.

[0145] Any biometric data collected by biometric components is captured and stored only with user approval and deleted upon user request. Furthermore, such biometric data may be used for very limited purposes, such as authentication. To ensure the restricted and authorized use of biometric information and other personally identifiable information (PII), access to this data is limited to authorized personnel, if applicable. Any use of biometric data may be strictly limited to authentication purposes, and the data may not be shared or sold to any third party without the user's explicit consent. Additionally, appropriate technical and organizational measures are implemented to ensure the security and confidentiality of this sensitive information.

[0146] The moving part 1430 includes an acceleration sensor part (e.g., an accelerometer), a gravity sensor part, and a rotation sensor part (e.g., a gyroscope).

[0147] The environmental component 1432 includes, for example, one or more camera devices (with still image / photograph and video capabilities), illuminance sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting nearby objects), gas sensors (e.g., gas detection sensors for detecting hazardous gas concentrations or measuring pollutants in the atmosphere for safety purposes), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment.

[0148] Regarding the camera device, the user system 102 may have a camera device system, which includes, for example, a front-facing camera on the front surface of the user system 102 and a rear-facing camera on the rear surface of the user system 102. The front-facing camera may be used, for example, to capture still images and videos (e.g., “selfies”) of the user of the user system 102, and then enhance these still images and videos with enhancement data (e.g., filters) as described above. The rear-facing camera may be used, for example, to capture still images and videos in a more conventional camera mode, wherein these images are similarly enhanced with enhancement data. In addition to the front-facing and rear-facing cameras, the user system 102 may also include a 360° camera for capturing 360° photos and videos.

[0149] Furthermore, the camera system of user system 102 may include dual rear cameras (e.g., a main camera and a depth sensing camera) located on the front and rear sides of user system 102, or even triple, quadruple, or quintuple rear camera configurations. These multi-camera systems may include, for example, wide-angle cameras, ultra-wide-angle cameras, telephoto cameras, macro cameras, and depth sensors.

[0150] The position component 1434 includes a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure from which altitude can be obtained), an orientation sensor component (e.g., a magnetometer), etc.

[0151] Communication can be implemented using a variety of technologies. I / O component 1408 also includes communication component 1436, which is operable to couple machine 1400 to network 1438 or device 1440 via a suitable coupling or connection. For example, communication component 1436 may include a network interface component or another suitable device that interfaces with network 1438. In further examples, communication component 1436 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, Bluetooth components, etc. ® Components (e.g., Bluetooth) ® Low energy consumption), Wi-Fi ® Components and other communication components that provide communication via other modes. Device 1440 can be another machine or any peripheral device among various peripheral devices (e.g., a peripheral device coupled via USB).

[0152] Furthermore, communication component 1436 may detect identifiers or include components operable to detect identifiers. For example, communication component 1436 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes), multi-dimensional barcodes (e.g., Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyphs), etc. TM The device can utilize optical sensors (such as MaxiCode, PDF417, UltraCode, UCC RSS-2D barcodes, and other optical codes) or acoustic detection components (such as microphones for identifying audio signals of the tags). Additionally, various information can be obtained via the communication component 1436, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, or location by detecting NFC beacon signals that indicate a specific location.

[0153] Various memories (e.g., main memory 1416, static memory 1418, and memory of processor 1404) and storage units 1420 may store one or more sets of instructions and data structures (e.g., software) used or embodied by any one or more methods or functions described herein. These instructions (e.g., instruction 1402) cause various operations to implement the disclosed examples when executed by processor 1404.

[0154] Instructions 1402 can be sent or received via network 1438, using a transmission medium, via a network interface device (e.g., a network interface component included in communication component 1436), and using any of several known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 1402 can be sent or received via a transmission medium through a coupling to device 1440 (e.g., peer-to-peer coupling).

[0155] Software Architecture

[0156] Figure 15 This is a block diagram 1500 illustrating a software architecture 1502 that can be installed on any one or more of the devices described herein. The software architecture 1502 is supported by hardware such as a machine 1504 including a processor 1506, memory 1508, and I / O components 1510. In this example, the software architecture 1502 can be conceptualized as a stack of layers, each providing specific functionality. The software architecture 1502 includes layers such as an operating system 1512, libraries 1514, frameworks 1516, and applications 1518. Operationally, application 1518 invokes API calls 1520 via the software stack and receives messages 1522 in response to API calls 1520.

[0157] Operating system 1512 manages hardware resources and provides public services. Operating system 1512 includes, for example, a kernel 1524, services 1526, and drivers 1528. Kernel 1524 serves as an abstraction layer between hardware and other software layers. For example, kernel 1524 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. Services 1526 can provide other public services to other software layers. Drivers 1528 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 1528 may include display drivers, camera drivers, BLUETOOTH® or BLUETOOTH® low-power drivers, flash memory drivers, serial communication drivers (e.g., USB drivers), Wi-Fi® drivers, audio drivers, power management drivers, etc.

[0158] Library 1514 provides common low-level infrastructure used by application 1518. Library 1514 may include system library 1530 (e.g., the C standard library), which provides functions such as memory allocation, string manipulation, and mathematical functions. Additionally, library 1514 may include API library 1532, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Picture Experts Group (JPEG or JPG), or Portable Web Graphics (PNG)), graphics libraries (e.g., an OpenGL framework for rendering graphic content on a display in two-dimensional (2D) and three-dimensional (3D) formats), database libraries (e.g., SQLite for providing various relational database functions), web libraries (e.g., WebKit for providing web browsing functionality), etc. Library 1514 may also include a wide variety of other libraries 1534 to provide many other APIs to application 1518.

[0159] Framework 1516 provides common high-level infrastructure for use by application 1518. For example, framework 1516 provides various graphical user interface (GUI) functions, advanced resource management, and advanced location services. Framework 1516 can provide a wide range of other APIs that can be used by application 1518, some of which may be specific to a particular operating system or platform.

[0160] In the example, application 1518 may include homepage application 1536, contact application 1538, browser application 1540, book reader application 1542, location application 1544, media application 1546, messaging application 1548, game application 1550, and various other categories of applications, such as third-party application 1552. Application 1518 is a program that performs the functions defined in the program. One or more applications 1518 can be created using various programming languages, such as object-oriented programming languages ​​(e.g., Objective-C, Java, or C++) or procedural programming languages ​​(e.g., C or assembly language). In a particular example, third-party application 1552 (e.g., an application developed by an entity other than a platform vendor using the Android™ or iOS™ Software Development Kit (SDK)) may be mobile software that runs on a mobile operating system (e.g., iOS™, Android™, Windows® Phone, or another mobile operating system). In this example, a third-party application 1552 can invoke API call 1520 provided by the operating system 1512 to enable the functionality described herein.

[0161] Example

[0162] Example 1 is a method that includes: receiving a portrait image of an entity from a device; generating an identity representing the entity based on the portrait image using a machine learning model; identifying a template based on a user account associated with the device, the template including a textual description of the scene and conditions; and generating a personalized image based on the identity and the template using a machine learning model.

[0163] In Example 2 of the subject of Example 1, the machine learning model includes a stable diffusion model, and the personalized image includes a personalized narrative image that shows the entity in the story scene.

[0164] In Example 3 of the topic of Example 1, the conditions include a visual resource that controls the generation of the personalized image, which includes one of a sketch, pose map, depth map, normal map, and Canney edge.

[0165] In Example 4, the subject of Example 1 includes: displaying a user interface on a device, the user interface including an invitation to generate multiple personalized images associated with the subject; detecting an indication that a user has selected the invitation; identifying a fixed number of templates from multiple templates associated with the subject, the fixed number of templates including a first template; and generating a fixed number of personalized images based on the fixed number of templates and the identity of the represented entity using a machine learning model.

[0166] In Example 5 of the topic in Example 4, a fixed number of templates are identified based on random selection, and the multiple templates associated with the topic include pre-selected templates that are not subject to random selection.

[0167] In Example 6 of the theme of Example 4, a fixed number of templates are made available to generate a fixed number of personalized images without initiating a transaction process, wherein the user interface is a first user interface, and wherein the invitation is a first invitation. Example 6 includes: identifying a remaining set of templates from a plurality of templates associated with the theme; and causing a second user interface to be displayed on the device, the second user interface including a second invitation for unlocking the availability of the remaining set of templates.

[0168] In Example 7, the subject of Example 6 includes: detecting an instruction from a user to select a second invitation; initiating a transaction process to unlock the availability of the remaining template set in exchange for an element; and in response to detecting the completion of the transaction process, using a machine learning model to generate a remaining set of personalized images based on the remaining template set and the identity of the representing entity.

[0169] In Example 8, the subject of Example 1 includes: displaying a personalized image on a device’s user interface, the user interface including user-selectable elements that allow the personalized image to be sent to another device associated with another user account connected to the user account; detecting an indication that the user has selected a user-selectable element; and sending the personalized image as a brief message to the other device.

[0170] In Example 9, the subject of Example 1 includes: determining that a template corresponds to multiple entities; and identifying multiple user accounts connected to a first user account, each of which is associated with an identity generated based on one or more portrait images of the corresponding entity, wherein the user account is the first user account.

[0171] In Example 10, the topics of Example 9 include: detecting an instruction from a user to select one or more entities from a plurality of entities; and generating personalized images based on a template and one or more entities using a machine learning model.

[0172] Example 11 is a system for implementing any one of Examples 1 through 10.

[0173] Example 20 is an apparatus that includes means for implementing any one of Examples 1 to 10.

[0174] Glossary

[0175] "Carrier signal" refers to any intangible medium, such as a medium capable of storing, encoding, or carrying machine-executable instructions and including digital or analog communication signals, or other intangible medium facilitating the transmission of such instructions. Instructions can be sent or received over a network using a transmission medium via a network interface device.

[0176] "Client device" means any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. Client devices can be, but are not limited to, mobile phones, desktop computers, laptop computers, portable digital assistants (PDAs), smartphones, tablet computers, ultrabooks, netbooks, laptops, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, set-top boxes, or any other communication device that a user can use to access the network.

[0177] "Communications network" means, for example, one or more parts of a network, which can be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a Common Old-Style Telephone Service (POTS) network, a cellular telephone network, a wireless network, a Wi-Fi® network, other types of networks, or a combination of two or more such networks. For example, a network or part of a network may include a wireless network or a cellular network, and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or other types of cellular or wireless coupling. In this example, coupling can enable any data transmission technology of various types, such as single-carrier radio transmission technology (1xRTT), evolved data optimization (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rate GSM evolution (EDGE) technology, the 3rd Generation Partnership Project (3GPP) including 3G, fourth-generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High-Speed ​​Packet Access (HSPA), Global Microwave Access Interoperability (WiMAX), Long Term Evolution (LTE) standards, other data transmission technologies defined by various standards setting organizations, other long-distance protocols, or other data transmission technologies.

[0178] A “component” refers to, for example, a device, a physical entity, or logic having boundaries defined by functional or subroutine calls, branch points, APIs, or other technologies that provide partitioning or modularity for specific processing or control functions. A component can be combined with other components via its interface to perform machine processing. A component can be a packaged functional hardware unit designed to be used with other components, and part of a program that typically performs a related function. A component can constitute a software component (e.g., code embodied on a machine-readable medium) or a hardware component. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various examples, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components (e.g., processors or processor groups) of a computer system can be configured by software (e.g., an application or application portion) to operate to perform certain operations described herein. Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component can include a dedicated circuit system or logic permanently configured to perform certain operations. A hardware component can be a dedicated processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). Hardware components may also include programmable logic or circuitry systems temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific part of a machine) uniquely tailored to perform the configured function, and no longer a general-purpose processor. It will be understood that the decision to implement hardware components mechanically in a temporarily configured (e.g., software-configured) circuitry system or in a dedicated and permanently configured circuitry system may be driven by cost and time considerations. Therefore, the phrase “hardware component” (or “hardware-implemented component”) should be understood to encompass tangible entities that are physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain way or perform certain operations described herein. Consider examples of hardware components being temporarily configured (e.g., programmed), without needing to configure or instantiate each of the hardware components at any given time. For example, in cases where hardware components include a general-purpose processor that becomes a dedicated processor by software configuration, the general-purpose processor may be configured at different times as (e.g., including different hardware components) different dedicated processors. The software accordingly configures one or more specific processors to constitute a particular hardware component at one time and different hardware components at different times. Hardware components can provide information to and receive information from other hardware components. Therefore, the described hardware components can be considered communicatively coupled.In the presence of multiple hardware components, communication can be achieved through signal transmission between or among two or more hardware components (e.g., via appropriate circuitry and buses). In examples where multiple hardware components are configured or instantiated at different times, such communication between hardware components can be achieved, for example, through the storage and retrieval of information in a storage structure accessible to the multiple hardware components. For example, a hardware component can perform an operation and store the output of that operation in a storage device communicatively coupled to it. Another hardware component can then access the storage device at a later time to retrieve and process the stored output. Hardware components can also initiate communication with input or output devices and be able to operate on resources (e.g., collections of information). The various operations of the example methods described herein can be performed at least in part by one or more processors, temporarily or permanently configured (e.g., via software), to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented components that operate to perform one or more operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be implemented at least in part by processors, where one or more specific processors are examples of hardware. For example, at least some of the operations of the methods can be performed by one or more processors or processor-implemented components. Furthermore, one or more processors can also operate to support the execution of related operations in a “cloud computing” environment or as a “Software as a Service” (SaaS) operation. For example, at least some of the operations can be performed by a group of computers (as an example of machines including processors), where these operations are accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of some operations can be distributed across processors, not residing within a single machine, but deployed across multiple machines. In some examples, the processor or processor-implemented component may reside in a single geographic location (e.g., in a home environment, office environment, or server cluster). In other examples, the processor or processor-implemented component may be distributed across multiple geographic locations.

[0179] "Computer-readable storage medium" refers to both, for example, machine storage media and transmission media. Therefore, these terms include both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" refer to the same thing and can be used interchangeably in this disclosure.

[0180] A "brief message" is a message that is accessible for a limited time, such as a short period of time. Brief messages can be text, images, videos, etc. The access time for a brief message can be set by the message sender. Alternatively, this access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is temporary.

[0181] "Machine storage medium" refers to one or more storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. This term should be accordingly considered to include, but is not limited to, solid-state memory, as well as optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memory, including, by way of example, semiconductor storage devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGAs, and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The terms "machine storage medium," "device storage medium," and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage medium," "computer storage medium," and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium."

[0182] "Non-transitory computer-readable storage medium" means, for example, a tangible medium capable of storing, encoding, or carrying instructions that can be executed by a machine.

[0183] "Signal medium" means, for example, any intangible medium capable of storing, encoding, or carrying machine-executable instructions and including digital or analog communication signals, or other intangible medium facilitating the communication of software or data. The term "signal medium" should be considered to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal whose characteristics are set or altered in a manner that encodes information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.

[0184] "User equipment" means, for example, a device that is accessed, controlled, or owned by a user and that the user interacts with to perform interactions or actions thereon, including interactions with other users or computer systems.

Claims

1. A system comprising: at least one processor; at least one memory component having instructions stored therein that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: receiving, from a device, a portrait image of an entity; generating, using the portrait image and a machine learning model, a representation of an identity of the entity, the machine learning model comprising a stable diffusion model; identifying, based on a user account associated with the device, a template, the template comprising a textual description of a scene and a condition; and generating, using a machine learning model, a personalized image based on the identity and the template.

2. The system of claim 1, wherein, the personalized image comprises a personalized narrative image showing the entity in a story scene.

3. The system of claim 1, wherein, the condition comprises a visual resource that controls generation of the personalized image, the visual resource comprising one of a sketch, a pose map, a depth map, a normal map, and a canny edge.

4. The system of claim 1, wherein, the template is a first template, and wherein the operations further comprise: causing display of a user interface on the device, the user interface comprising an invitation to generate a plurality of personalized images associated with a theme; detecting an indication that a user selected the invitation; identifying, from a plurality of templates associated with the theme, a fixed number of templates, the fixed number of templates comprising the first template; and generating, using a machine learning model, a fixed number of personalized images based on the fixed number of templates and the representation of the entity.

5. The system of claim 4, wherein, the fixed number of templates is identified based on a random selection, and wherein the plurality of templates associated with the theme comprises preselected templates that are not subject to the random selection.

6. The system of claim 4, wherein, enabling the fixed number of templates to be used to generate the fixed number of personalized images without initiating a transaction process, wherein the user interface is a first user interface, wherein the invitation is a first invitation, and wherein the operations further comprise: identifying, from the plurality of templates associated with the theme, a remaining set of templates; and causing display of a second user interface on the device, the second user interface comprising a second invitation to unlock availability of the remaining set of templates.

7. The system of claim 6, wherein, the operations further comprise: detecting an indication that a user selected the second invitation; initiating the transaction process to unlock availability of the remaining set of templates in exchange for an element; and in response to detecting completion of the transaction process, generating, using a machine learning model, a remaining set of personalized images based on the remaining set of templates and the representation of the entity.

8. The system of claim 1, wherein, the operations further comprise: causing display of the personalized image on a user interface of the device, the user interface comprising a user-selectable element that allows the personalized image to be sent to other devices associated with other user accounts connected to the user account; detecting an indication that a user selected the user-selectable element; and sending the personalized image as an ephemeral message to the other devices.

9. The system of claim 1, wherein, the user account is a first user account, and wherein the operations further comprise: determining that the template corresponds to a plurality of entities; and identifying a plurality of user accounts connected to the first user account, each of the plurality of user accounts being associated with an identity generated based on one or more portrait images of a corresponding entity.

10. The system of claim 9, wherein, The operations further include: detecting an indication that a user selected one or more entities from the plurality of entities; and generating, using a machine learning model, the personalized image based on the template and the one or more entities.

11. A method comprising: receiving, from a device, a portrait image of an entity; generating, by using the portrait image and a machine learning model, an identity representing the entity, the machine learning model comprising a stable diffusion model; identifying, based on a user account associated with the device, a template, the template comprising a textual description of a scene and a condition; and generating, using a machine learning model, a personalized image based on the identity and the template.

12. The method of claim 11, wherein, The machine learning model comprises a stable diffusion model, and wherein the personalized image comprises a personalized narrative image showing the entity in a story scene.

13. The method of claim 11, wherein, The condition comprises a visual resource that controls the generation of the personalized image, the visual resource comprising one of a sketch, a pose map, a depth map, a normal map, and a canny edge.

14. The method of claim 11, wherein, The template is a first template, the method further comprising: causing display, on the device, of a user interface, the user interface comprising an invitation to generate a plurality of personalized images associated with a theme; detecting an indication that a user selected the invitation; identifying, from a plurality of templates associated with the theme, a fixed number of templates, the fixed number of templates comprising the first template; and generating, using a machine learning model, a fixed number of personalized images based on the fixed number of templates and the identity representing the entity.

15. The method of claim 14, wherein, The fixed number of templates is identified based on a random selection, and wherein the plurality of templates associated with the theme comprises preselected templates that are not subject to the random selection.

16. The method of claim 14, wherein, The fixed number of templates are made available for use in generating the fixed number of personalized images without initiating a transaction process, wherein the user interface is a first user interface, and wherein the invitation is a first invitation, the method further comprising: identifying, from the plurality of templates associated with the theme, a remaining set of templates; and causing display, on the device, of a second user interface, the second user interface comprising a second invitation to unlock availability of the remaining set of templates.

17. The method of claim 16, further comprising: detecting an indication that a user selected the second invitation; initiating the transaction process to unlock availability of the remaining set of templates in exchange for an element; and in response to detecting completion of the transaction process, generating, using a machine learning model, a remaining set of personalized images based on the remaining set of templates and the identity representing the entity.

18. The method of claim 11, further comprising: causing display of the personalized image on a user interface of the device, the user interface including a user-selectable element that allows sending of the personalized image to other devices associated with other user accounts connected to the user account; detecting an indication of a user selecting the user-selectable element; and sending the personalized image as a transient message to the other devices.

19. The method of claim 11, wherein, the user account is a first user account, the method further comprising: determining that the template corresponds to a plurality of entities; identifying a plurality of user accounts connected to the first user account, each of the plurality of user accounts being associated with an identity generated based on one or more portrait images of a corresponding entity; detecting an indication of a user selecting one or more entities from the plurality of entities; and generating, using a machine learning model, the personalized image based on the template and the one or more entities.

20. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: receiving, from a device, a portrait image of an entity; generating, by using the portrait image and a machine learning model, an identity representing the entity, the machine learning model comprising a stable diffusion model; identifying, based on a user account associated with the device, a template, the template comprising a textual description of a scene and a condition; and generating, using a machine learning model, a personalized image based on the identity and the template.