Facial identity retention using stable diffusion model
By combining a stable diffusion generative model with a segmentation network, we address the computationally intensive and high-latency issues of image processing under changing conditions, and achieve efficient image stylization with facial attribute preservation.
Patent Information
- Application Number
- CN202480010315.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-03
- Filing Date
- 2024-02-02
- Publication Date
- 2025-09-12
AI Technical Summary
Processing digital images captured under various changing conditions is computationally intensive and challenging. Existing technologies have difficulty effectively reducing latency and power consumption, while image stylization processing that keeps facial attributes unchanged is inefficient.
A stable diffusion generative model is adopted to obtain features from the input image and perform weighted combination using the mask of the segmentation network. The sampling technology is combined to maintain the consistency of facial attributes and achieve image stylization.
Improves image processing efficiency, reduces latency and capture device power consumption, while maintaining consistency in facial attributes.
Smart Images

Figure CN120641937A_ABST
Abstract
Description
[0001] Priority Declaration
[0002] This application claims the benefit of priority to U.S. patent application serial number 18 / 164,458, filed on February 3, 2023, which is incorporated herein by reference in its entirety. Background Art
[0003] With the increased use of digital images, the affordability of portable computing devices, the availability of increasing capacity digital storage media, and the increase in bandwidth and accessibility of network connections, digital images have become a part of daily life for more and more people. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] In the accompanying drawings (which are not necessarily drawn to scale), like reference numerals may describe similar components in different views. To easily identify the discussion of any particular element or action, the most significant digit or digits in a reference numeral refer to the figure number in which the element is first introduced. Some non-limiting examples are shown in the figures of the accompanying drawings, in which:
[0005] Figure 1 is a diagrammatic representation of a networked environment in which the present disclosure may be deployed, according to some examples.
[0006] Figure 2 is a diagrammatic representation of a messaging system having both client-side and server-side functionality, according to some examples.
[0007] Figure 3 is a diagrammatic representation of data structures as maintained in a database, according to some examples.
[0008] Figure 4 is a diagrammatic representation of messages according to some examples.
[0009] Figure 5 An example process flow of a computational architecture for facial identity preservation using an image-to-image model of a stable diffusion generative model is shown, in accordance with an embodiment of the subject technology.
[0010] Figure 6 An example of identity preservation through feature blending in accordance with an embodiment of the subject technology is shown.
[0011] Figure 7 An example of identity preservation through feature blending in accordance with an embodiment of the subject technology is shown.
[0012] Figure 8 is a flowchart illustrating a method according to certain example embodiments.
[0013] Figure 9is a diagrammatic representation of a machine in the form of a computer system, according to some examples, within which a set of instructions may be executed, causing the machine to perform any one or more of the methodologies discussed herein.
[0014] Figure 10 is a block diagram illustrating a software architecture within which examples may be implemented. DETAILED DESCRIPTION
[0015] Users from various locations and with various interests can capture digital images of various subjects and make the captured images available to others via a network (e.g., the Internet). To enhance the user's experience with digital images and provide various features, enabling a computing device to perform image processing operations on a variety of objects and / or features captured under various varying conditions (e.g., variations in image scale, noise, lighting, motion, or geometric distortion) can be challenging and computationally intensive.
[0016] Augmented reality technology aims to bridge the gap between virtual and real-world environments by providing an augmented real-world environment enhanced with electronic information. As a result, the electronic information appears to be part of the user's perceived real-world environment. In some examples, augmented reality technology also provides a user interface for interacting with the electronic information superimposed on the augmented real-world environment.
[0017] As mentioned above, with the increase in the use of digital images, the affordability of portable computing devices, the availability of increased capacity of digital storage media, and the bandwidth and accessibility of network connections, digital images have become part of the daily lives of more and more people. Users from various locations with various interests can capture digital images of various subjects and make the captured images available to others via a network (e.g., the Internet). In order to enhance the user's experience with digital images and provide various features, it can be challenging and computationally intensive to enable computing devices to perform image processing operations on various objects and / or features captured under various changing conditions (e.g., changes in image scale, noise, lighting, motion, or geometric distortion).
[0018] Users of mobile computing devices frequently use and increasingly utilize messaging systems in a variety of situations to provide different types of functionality in a convenient manner. As described herein, the subject messaging system includes practical applications that provide improvements in capturing image data and rendering AR content (e.g., images, videos, etc.) based on the captured image data by at least providing technical improvements in capturing image data using power- and resource-constrained electronic devices. Such improvements in capturing image data are achieved through the techniques provided by the subject technology, which reduce latency and improve the efficiency of processing captured image data, thereby also reducing the power consumption of the capture device.
[0019] As discussed further herein, the subject technology enables preserving facial attributes while stylizing an input image using a generative network (such as stable diffusion). This is achieved by modifying intermediate features of the neural network using features obtained from the input real image. The modification of features is achieved by weighted combination of values using a mask obtained from the segmentation network.
[0020] To produce consistent results with the same pose and facial attributes, the subject technology introduces a technique for sampling, which enables features for modification to be obtained by reconstructing the network's initial parameters (noise) similar to the real image.
[0021] As referred to herein, the phrases "augmented reality experience," "augmented reality content item," and "augmented reality content generator" include or refer to various image processing operations corresponding to image modifications, filters, AR content generators, media overlays, transformations, etc., and may also include playback of audio or music content during the presentation of AR content or media content, as further described herein.
[0022] Networked computing environment
[0023] Figure 1 1 is a block diagram illustrating an example interactive system 100 for facilitating interactions over a network (e.g., exchanging text messages, conducting text, audio, and video calls, or playing games). The interactive system 100 includes a plurality of client systems 102, each of which hosts a plurality of applications including an interactive client 104 and other applications 106. Each interactive client 104 is communicatively coupled to other instances of the interactive client 104 (e.g., hosted on respective other user systems 102), an interactive server system 110, and third-party servers 112 via one or more communication networks including a network 108 (e.g., the Internet). The interactive client 104 can also communicate with the locally hosted application 106 using an application programming interface (API).
[0024] Each user system 102 may include multiple user devices, such as a mobile device 114 , a head wearable device 116 , and a computer client device 118 , that are communicatively connected to exchange data and messages.
[0025] The interactive clients 104 interact with other interactive clients 104 and with the interactive server system 110 via the network 108. The data exchanged between the interactive clients 104 (e.g., interaction 120) and between the interactive clients 104 and the interactive server system 110 includes functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).
[0026] The interactive server system 110 provides server-side functionality to the interactive clients 104 via the network 108. Although certain functions of the interactive system 100 are described herein as being performed by either the interactive clients 104 or the interactive server system 110, whether certain functions are located within the interactive clients 104 or within the interactive server system 110 may be a design choice. For example, it may be technically preferable to initially deploy certain technologies and functions within the interactive server system 110, but later migrate the technologies and functions to the interactive clients 104 where the user system 102 has sufficient processing power.
[0027] The interactive server system 110 supports various services and operations provided to the interactive clients 104. Such operations include sending data to the interactive clients 104, receiving data from the interactive clients 104, and processing data generated by the interactive clients 104. The data may include message content, client device information, geographic location information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. The data exchange within the interactive system 100 is activated and controlled by functions available through the user interface (UI) of the interactive client 104.
[0028] Turning now specifically to the interaction server system 110, an application program interface (API) server 122 is coupled to the interaction server 124 and provides a programming interface thereto, making the functionality of the interaction server 124 accessible to the interaction clients 104, other applications 106, and third-party servers 112. The interaction server 124 is communicatively coupled to a database server 126, thereby facilitating access to a database 128 that stores data associated with interactions processed by the interaction server 124. Similarly, a web server 130 is coupled to the interaction server 124 and provides a web-based interface to the interaction server 124. To this end, the web server 130 processes incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.
[0029] The application program interface (API) server 122 receives and sends interaction data (e.g., commands and message payloads) between the interaction server 124 and the client system 102 (and, for example, the interaction client 104 and other applications 106), as well as the third-party server 112. Specifically, the application program interface (API) server 122 provides a set of interfaces (e.g., routines and protocols) that can be called or queried by the interaction client 104 and other applications 106 to activate the functionality of the interaction server 124. The application program interface (API) server 122 exposes various functions supported by the interaction server 124, including account registration; login functionality; sending interaction data from a particular interaction client 104 to another interaction client 104 via the interaction server 124; transferring media files (e.g., images or videos) from the interaction client 104 to the interaction server 124; setting up a collection of media data (e.g., a story); retrieving a friend list of a user of the user system 102; retrieving messages and content; adding and removing entities (e.g., friends) from an entity graph (e.g., a social graph); locating friends within a social graph; and opening application events (e.g., related to the interaction client 104).
[0030] Interactive server 124 hosts multiple systems and subsystems, see below Figure 2 Provide a description.
[0031] Linked Applications
[0032] Returning to the interactive client 104, the features and functionality of an external resource (e.g., a linked application 106 or applet) are made available to the user via the interface of the interactive client 104. In this context, "external" refers to the fact that the application 106 or applet is external to the interactive client 104. Although external resources are typically provided by a third party, they can also be provided by the creator or provider of the interactive client 104. The interactive client 104 receives a user selection of an option to launch or access features of such an external resource. The external resource can be an application 106 installed on the user system 102 (e.g., a "local app"), or a small-scale version of an application (e.g., a "mini-program") hosted on the user system 102 or located remotely from the user system 102 (e.g., on a third-party server 112). The small-scale version of an application includes a subset of the features and functionality of the application (e.g., the full-scale, local version of the application) and is implemented using a markup language document. In some examples, the small-scale version of an application (e.g., a "mini-program") is a web-based markup language version of the application and is embedded in the interactive client 104. In addition to using markup language documents (eg, .*ml files), applet may include scripting languages (eg, .*js files or .json files) and style sheets (eg, .*ss files).
[0033] In response to receiving a user selection of an option to launch or access a feature of an external resource, the interactive client 104 determines whether the selected external resource is a web-based external resource or a locally installed application 106. In some cases, an application 106 installed locally on the user system 102 can be independent of and launched separately from the interactive client 104, such as by selecting an icon corresponding to the application 106 on a home screen of the user system 102. A small-scale version of such an application can be launched or accessed via the interactive client 104, and in some examples, no portion of the small-scale application can be accessed outside of the interactive client 104 or only a limited portion of the small-scale application can be accessed outside of the interactive client 104. The small-scale application can be launched by the interactive client 104 receiving, for example, a markup language document associated with the small-scale application from a third-party server 112 and processing such a document.
[0034] In response to determining that the external resource is a locally installed application 106, the interactive client 104 instructs the user system 102 to launch the external resource by executing locally stored code corresponding to the external resource. In response to determining that the external resource is a web-based resource, the interactive client 104 communicates with (for example) a third-party server 112 to obtain a markup language document corresponding to the selected external resource. The interactive client 104 then processes the obtained markup language document to present the web-based external resource within the user interface of the interactive client 104.
[0035] The interactive client 104 can notify the user of the user system 102 or other users associated with such user (e.g., "friends") of activities occurring in one or more external resources. For example, the interactive client 104 can provide participants in a conversation (e.g., a chat session) within the interactive client 104 with notifications regarding the current or recent use of an external resource by one or more members of the user group. One or more users can be invited to join an active external resource or a recently used but currently inactive external resource (in a friend group) can be launched. The external resource can provide participants in the conversation, each using a corresponding interactive client 104, with the ability to share items, conditions, states, or locations in the external resource with one or more members of the user group in the chat session. The shared items can be interactive chat cards that members of the chat can use to interact, for example, to launch the corresponding external resource, view specific information within the external resource, or take members of the chat to a specific location or state within the external resource. Within a given external resource, a response message can be sent to the user on the interactive client 104. Based on the current context of the external resource, the external resource can selectively include different media items in the response.
[0036] The interactive client 104 can present a list of available external resources (e.g., applications 106 or applets) to the user to launch or access a given external resource. The list can be presented in the form of a context-sensitive menu. For example, the icons representing different applications (or applets) of the application 106 (or applets) can change based on how the user launches the menu (e.g., from a conversational interface or from a non-conversational interface).
[0037] System Architecture
[0038] Figure 2 is a block diagram illustrating additional details regarding the interactive system 100 according to some examples. Specifically, the interactive system 100 is shown as including an interactive client 104 and an interactive server 124. The interactive system 100 includes a plurality of subsystems supported on the client side by the interactive client 104 and on the server side by the interactive server 124. Example subsystems are discussed below.
[0039] Image processing system 202 provides various functions that enable a user to capture and enhance (eg, annotate or otherwise modify or edit) media content associated with a message.
[0040] The camera system 204 includes control software (e.g., in a camera application) that interacts with and controls the hardware camera hardware of the user system 102 (e.g., directly or via operating system control) to modify and enhance the real-time images captured and displayed via the interactive client 104.
[0041] The enhancement system 206 provides functionality related to the generation and publication of enhancements (e.g., media overlays) for images captured in real time by the camera of the user system 102 or retrieved from the memory of the user system 102. For example, the enhancement system 206 is operable to select, present, and display media overlays (e.g., image filters or image lenses) to the interactive client 104 for enhancing real-time images received via the camera system 204 or stored images retrieved from the memory of the user system 102. These enhancements are selected and presented to the user of the interactive client 104 by the enhancement system 206 based on a number of inputs and data, such as:
[0042] The geographic location of the user system 102; and
[0043] Social network information of users of the user system 102 .
[0044] Enhancement can include audio and visual content and visual effects. Examples of audio and visual content include pictures, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects can be applied to media content items (e.g., photos or videos) at user system 102 for transmitting in a message, or applied to video content such as a video content stream or feed sent from interactive client 104. Therefore, image processing system 202 can interact with the various subsystems of communication system 208 and support the various subsystems of communication system 208, such as messaging system 210 and video communication system 212.
[0045] Media overlays can include text or image data that can be superimposed on a photo taken by user system 102 or a video stream produced by user system 102. In some examples, the media overlay can be a location overlay (e.g., Venice Beach), the name of a live event, or a business name overlay (e.g., Beach Cafe). In other examples, image processing system 202 uses the geographic location of user system 102 to identify a media overlay that includes the name of a business at the geographic location of user system 102. The media overlay may include other tags associated with the business. The media overlay may be stored in database 128 and accessed by database server 126.
[0046] Image processing system 202 provides a user-based publishing platform that enables users to select a geographic location on a map and upload content associated with the selected geographic location. Users can also specify under what circumstances specific media overlays should be provided to other users. Image processing system 202 generates a media overlay that includes the uploaded content and associates it with the selected geographic location.
[0047] The augmented creation system 214 supports the augmented reality developer platform and includes applications for content creators (e.g., artists and developers) to create and publish augmentations (e.g., augmented reality experiences) for the interactive clients 104. The augmented creation system 214 provides content creators with a library of built-in features and tools, including, for example, custom shaders, tracking techniques, and templates.
[0048] In some examples, the enhancement creation system 214 provides a merchant-based publishing platform that enables merchants to select specific enhancements associated with a geographic location through a bidding process. For example, the enhancement creation system 214 associates the highest bidding merchant's media overlay with the corresponding geographic location for a predefined amount of time.
[0049] The communication system 208 is responsible for enabling and processing various forms of communication and interaction within the interactive system 100 and includes a messaging system 210, an audio communication system 216, and a video communication system 212. The messaging system 210 is responsible for enforcing temporary or time-limited access to content by the interactive clients 104. The messaging system 210 includes a plurality of timers (e.g., in a transient timer system 218) that selectively enable access (e.g., for presentation and display) of messages and associated content via the interactive clients 104 based on duration and display parameters associated with a message or a collection of messages (e.g., a story). Additional details regarding the operation of the transient timer system 218 are provided below. The audio communication system 216 enables and supports audio communication (e.g., real-time audio chat) between multiple interactive clients 104. Similarly, the video communication system 212 enables and supports video communication (e.g., real-time video chat) between multiple interactive clients 104.
[0050] The user management system 220 is operationally responsible for managing user data and profiles, and includes a social networking system 222 that maintains information about relationships between users of the interactive system 100 .
[0051] The collection management system 224 is operationally responsible for managing collections or collections of media (e.g., collections of text, images, video, and audio data). Collections of content (e.g., messages, including images, video, text, and audio) can be organized into "event galleries" or "event stories." Such collections can be made available for a specified time period (e.g., the duration of the event to which the content relates). For example, content related to a concert can be made available as a "story" for the duration of the concert. The collection management system 224 can also be responsible for publishing icons to the user interface of the interactive client 104 that provide notifications of specific collections. The collection management system 224 includes curation functionality that enables collection managers to manage and curate collections of specific content. For example, a curation interface enables event organizers to curate collections of content related to a specific event (e.g., removing inappropriate content or redundant messages). In addition, the collection management system 224 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, users can be compensated for including user-generated content in collections. In such cases, the collection management system 224 operates to automatically make payments to such users for use of their content.
[0052] The mapping system 226 provides various geolocation capabilities and supports the presentation of map-based media content and messages by the interactive client 104. For example, the mapping system 226 enables the display of user icons or avatars (e.g., stored in the profile data 302) on a map to indicate the current or past locations of the user's "friends" within the context of the map, as well as media content generated by such friends (e.g., a collection of messages including photos and videos). For example, on the map interface of the interactive client 104, messages posted by the user to the interactive system 100 from a particular geographic location can be displayed to the "friends" of a particular user within the context of that particular location on the map. A user can also share his or her location and status information with other users of the interactive system 100 via the interactive client 104 (e.g., using an appropriate status avatar), where the location and status information is similarly displayed to selected users within the context of the map interface of the interactive client 104.
[0053] The gaming system 228 provides various gaming functions within the context of the interactive client 104. The interactive client 104 provides a gaming interface that provides a list of available games that can be launched by a user within the context of the interactive client 104 and played with other users of the interactive system 100. The interactive system 100 also enables a particular user to invite other users to play a particular game by sending an invitation to such other users from the interactive client 104. The interactive client 104 also supports audio, video, and text messaging (e.g., chat) within the context of game play, provides leaderboards for games, and also supports the provision of in-game rewards (e.g., game coins and items).
[0054] The external resource system 230 provides an interface for the interactive client 104 to communicate with a remote server (e.g., a third-party server 112) to launch or access external resources (i.e., applications or applets). Each third-party server 112 hosts, for example, an application or a small-scale version of an application (e.g., a game application, a utility application, a payment application, or a ride-sharing application) based on a markup language (e.g., HTML5). The interactive client 104 can launch a web-based resource (e.g., an application) by accessing an HTML5 file from a third-party server 112 associated with the web-based resource. The applications hosted by the third-party server 112 are programmed in JavaScript using a software development kit (SDK) provided by the interactive server 124. The SDK includes an application program interface (API) with functions that can be called or activated by a web-based application. The interactive server 124 hosts a JavaScript library that provides access to a given external resource to specific user data of the interactive client 104. HTML5 is an example of a technology for programming games, but applications and resources programmed based on other technologies can be used.
[0055] To integrate the SDK's functionality into a web-based resource, the third-party server 112 downloads the SDK from the interaction server 124, or the third-party server 112 receives the SDK in some other manner. Once downloaded or received, the SDK is included as part of the application code of the external web-based resource. The code of the web-based resource can then call or activate certain functions of the SDK to integrate the features of the interaction client 104 into the web-based resource.
[0056] The SDK stored on the interactive server system 110 effectively provides a bridge between external resources (e.g., applications 106 or applets) and the interactive client 104. This gives users a seamless experience of communicating with other users on the interactive client 104 while also preserving the look and feel of the interactive client 104. In order to bridge the communication between the external resources and the interactive client 104, the SDK facilitates communication between the third-party server 112 and the interactive client 104. The WebViewJavaScriptBridge running on the user system 102 establishes two one-way communication channels between the external resources and the interactive client 104. Messages are sent asynchronously between the external resources and the interactive client 104 via these communication channels. Each SDK function activation is sent as a message and a callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message with the callback identifier.
[0057] By using the SDK, not all information from the interactive client 104 is shared with the third-party server 112. The SDK limits which information is shared based on the needs of the external resource. Each third-party server 112 provides an HTML5 file corresponding to the web-based external resource to the interactive server 124. The interactive server 124 can add a visual representation of the web-based external resource (e.g., a box design or other graphics) in the interactive client 104. Once the user selects the visual representation or instructs the interactive client 104 to access a feature of the web-based external resource through the GUI of the interactive client 104, the interactive client 104 obtains the HTML5 file and instantiates the resource for accessing the feature of the web-based external resource.
[0058] The interactive client 104 presents a graphical user interface (e.g., a login page or title screen) for the external resource. During, before, or after presenting the login page or title screen, the interactive client 104 determines whether the launched external resource has previously been authorized to access the user data of the interactive client 104. In response to determining that the launched external resource has previously been authorized to access the user data of the interactive client 104, the interactive client 104 presents another graphical user interface of the external resource, including the functions and features of the external resource. In response to determining that the launched external resource has not previously been authorized to access the user data of the interactive client 104, after displaying the login page or title screen of the external resource for a threshold period of time (e.g., 3 seconds), the interactive client 104 slides up a menu (e.g., animating the menu to emerge from the bottom of the screen to the middle or other portion of the screen) for authorizing the external resource to access the user data. The menu identifies the type of user data that the external resource is authorized to use. In response to receiving a user selection of the accept option, the interactive client 104 adds the external resource to the list of authorized external resources and enables the external resource to access the user data from the interactive client 104. External resources are authorized by the interactive client 104 to access user data under the OAuth 2 framework.
[0059] The interaction client 104 controls the type of user data shared with the external resource based on the type of external resource that is authorized. For example, an external resource comprising a full-scale application (e.g., application 106) is provided with access to a first type of user data (e.g., a two-dimensional avatar of the user with or without different avatar characteristics). As another example, an external resource comprising a small-scale version of an application (e.g., a web-based version of the application) is provided with access to a second type of user data (e.g., payment information, a two-dimensional avatar of the user, a three-dimensional avatar of the user, and an avatar with various avatar characteristics). Avatar characteristics include different ways to customize the look and feel of an avatar (e.g., different poses, facial features, clothing, etc.).
[0060] The advertising system 232 operatively enables third parties to purchase advertisements for presentation to end users via the interactive clients 104 and also handles the delivery and presentation of these advertisements.
[0061] Data Architecture
[0062] Figure 3 is a diagram illustrating a data structure 300 that may be stored in a database 304 of the interactive server system 110 according to certain examples. Although the contents of the database 304 are shown as including a plurality of tables, it should be understood that data may be stored in other types of data structures (e.g., as an object-oriented database).
[0063] The database 304 includes message data stored in a message table 306. For any particular message, the message data includes at least message sender data, message recipient (or receiver) data, and payload. Figure 3 Additional details regarding information that may be included in a message and within the message data stored in message table 306 are described.
[0064] The entity table 308 stores entity data and is linked (e.g., by reference) to the entity graph 310 and the profile data 302. Entities for which records are maintained within the entity table 308 may include individuals, corporate entities, organizations, objects, places, events, and the like. Regardless of the entity type, any entity for which the interactive server system 110 stores data may be an identified entity. Each entity is provided with a unique identifier and an entity type identifier (not shown).
[0065] The entity graph 310 stores information about relationships and associations between entities. By way of example only, such relationships may be social, professional (e.g., working at a common company or organization), interest-based, or activity-based. Some relationships between entities may be one-way, such as a subscription by an individual user to digital content from a business or publication (e.g., a newspaper or other digital media channel, or brand). Other relationships may be two-way, such as a "friend" relationship between individual users of the interactive system 100.
[0066] Certain permissions and relationships can be attached to each relationship, and also to each direction of the relationship. For example, a bidirectional relationship (e.g., a friend relationship between individual users) can include authorization for the publication of digital content items between the individual users, but can impose certain restrictions or filters on the publication of such digital content items (e.g., based on content characteristics, location data, or time of day data). Similarly, a subscription relationship between an individual user and a business user can impose varying degrees of restrictions on the publication of digital content from the business user to the individual user, and can significantly restrict or prevent the publication of digital content from the individual user to the business user. A particular user, as an example of an entity, can have certain restrictions recorded in the record for that entity within the entity table 308 (e.g., by way of privacy settings). Such privacy settings can apply to all types of relationships in the context of the interactive system 100, or can be selectively applied to certain types of relationships.
[0067] Profile data 302 stores various types of profile data about a particular entity. Based on the privacy settings specified by the particular entity, profile data 302 can be selectively used and presented to other users of the interactive system 100. In the case where the entity is a person, profile data 302 includes, for example, the user's name, phone number, address, settings (e.g., notification and privacy settings), and an avatar representation (or a collection of such avatar representations) selected by the user. The particular user can then selectively include one or more of these avatar representations within the content of messages transmitted via the interactive system 100 and on a map interface displayed to other users by the interactive client 104. The collection of avatar representations can include a "status avatar," which presents a graphical representation of the user's status or activity that they may choose to transmit at a particular time.
[0068] Where the entity is a group, the profile data 302 for the group may similarly include one or more avatar representations associated with the group, in addition to the group name, members, and various settings for the relevant group (eg, notifications).
[0069] Database 304 also stores enhancement data, such as overlays or filters, in enhancement table 312. Enhancement data is associated with and applied to videos (data for videos is stored in video table 314) and images (data for images is stored in image table 316).
[0070] In some examples, a filter is displayed as an overlay on an image or video during presentation to a recipient user. Filters can be of various types, including user-selected filters from a set of filters that the interactive client 104 presents to the sending user while the sending user is composing a message. Other types of filters include geolocation filters (also known as geofilters), which can be presented to the sending user based on geographic location. For example, a geolocation filter specific to a nearby or special location can be presented within the user interface by the interactive client 104 based on geographic location information determined by a global positioning system (GPS) unit of the user system 102.
[0071] Another type of filter is a data filter, which can be selectively presented to the sending user by the interaction client 104 based on other input or information collected during the message creation process by the user system 102. Examples of data filters include the current temperature at a particular location, the current speed the sending user is traveling, the battery life of the user system 102, or the current time.
[0072] Other augmented data that can be stored in the image table 316 include augmented reality content items (e.g., corresponding to application "lenses" or augmented reality experiences). Augmented reality content items can be real-time special effects and sounds that can be added to images or videos.
[0073] The story table 318 stores data about a collection of messages and associated images, video, or audio data that are compiled into a collection (e.g., a story or gallery). The creation of a particular collection can be initiated by a particular user (e.g., each user for whom a record is maintained in the entity table 308). A user can create a "personal story" in the form of a collection of content that has been created and sent / broadcasted by the user. To this end, the user interface of the interactive client 104 can include an icon that the user can select to enable the sending user to add specific content to his or her personal story.
[0074] The collection can also constitute a "live story," which is a collection of content from multiple users that is created manually, automatically, or using a combination of manual and automatic techniques. For example, a "live story" can constitute a curated stream of user-submitted content from various locations and events. Users whose client devices have location services enabled and who are at a common location event at a particular time can be presented with the option to contribute content to a particular live story, for example, via the user interface of the interactive client 104. Live stories can be identified to the user by the interactive client 104 based on his or her location. The end result is a "live story" told from the perspective of the group.
[0075] Another type of content collection is called a "location story," which enables users whose user systems 102 are located in a particular geographic location (e.g., on a college or university campus) to contribute to a particular collection. In some examples, contributions to location stories may employ secondary authentication to verify that the end user belongs to a particular organization or other entity (e.g., is a student on a university campus).
[0076] As mentioned above, video table 314 stores video data that, in some examples, is associated with messages for which records are maintained within message table 306. Similarly, image table 316 stores image data associated with messages whose message data is stored in entity table 308. Entity table 308 can associate various enhancements from enhancement table 312 with the various images and videos stored in image table 316 and video table 314.
[0077] Data communication architecture
[0078] Figure 4 is a schematic diagram illustrating the structure of a message 400 according to some examples, the message 400 being generated by an interaction client 104 for transmission to another interaction client 104 via an interaction server 124. The content of a particular message 400 is used to populate a message table 306 stored within a database 304 accessible by the interaction server 124. Similarly, the content of the message 400 is stored in memory as "in-transit" or "in-flight" data of the user system 102 or the interaction server 124. The message 400 is shown as including the following example components:
[0079] Message identifier 402 : A unique identifier that identifies the message 400 .
[0080] Message text payload 404 : Text to be generated by the user via the user interface of the user system 102 and included in the message 400 .
[0081] • Message image payload 406: Image data captured by the camera component of the user system 102 or retrieved from the memory component of the user system 102 and included in the message 400. The image data for a message 400 sent or received may be stored in the image table 316.
[0082] • Message video payload 408: Video data captured by the camera component or retrieved from the memory component of the user system 102 and included in the message 400. The video data for a message 400 sent or received may be stored in the image table 316.
[0083] • Message audio payload 410 : audio data captured by a microphone or retrieved from a memory component of the user system 102 and included in the message 400 .
[0084] Message enhancement data 412: Enhancement data (e.g., filters, stickers, or other annotations or enhancements) representing enhancements to be applied to the message image payload 406, message video payload 408, or message audio payload 410 of the message 400. Enhancement data for a sent or received message 400 may be stored in the enhancement table 312.
[0085] • Message duration parameter 414: A parameter value indicating the amount of time, in seconds, that the content of the message (e.g., message image payload 406, message video payload 408, message audio payload 410) is to be presented to or made accessible to the user via the interactive client 104.
[0086] Message geolocation parameters 416: Geolocation data (e.g., latitude and longitude coordinates) associated with the content payload of the message. Multiple message geolocation parameter 416 values may be included in the payload, with each of these parameter values being associated with a content item included in the content (e.g., a specific image within the message image payload 406 or a specific video within the message video payload 408).
[0087] Message story identifier 418: An identifier value that identifies one or more content collections (e.g., a "story" identified in stories table 318) associated with a particular content item in message image payload 406 of message 400. For example, the identifier value can be used to associate multiple images within message image payload 406 with multiple content collections.
[0088] Message tags 420: Each message 400 can be tagged with a plurality of tags, each of which indicates the subject matter of the content included in the message payload. For example, if a particular image included in the message image payload 406 depicts an animal (e.g., a lion), a tag value can be included within the message tags 420 indicating the relevant animal. The tag value can be manually generated based on user input, or can be automatically generated using, for example, image recognition.
[0089] • Message sender identifier 422: An identifier (eg, a messaging system identifier, an email address, or a device identifier) that indicates the user of the user system 102 on which the message 400 was generated and from which the message 400 was sent.
[0090] • Message recipient identifier 424: Indicates an identifier (eg, a messaging system identifier, email address, or device identifier) of the user of user system 102 to which message 400 is addressed.
[0091] The content (e.g., value) of each component of message 400 may be a pointer to a location in a table where the content data value is stored. For example, the image value in message image payload 406 may be a pointer to a location (or its address) within image table 316. Similarly, the value within message video payload 408 may point to data stored in image table 316, the value stored in message enhancement data 412 may point to data stored in enhancement table 312, the value stored in message story identifier 418 may point to data stored in story table 318, and the values stored in message sender identifier 422 and message recipient identifier 424 may point to user records stored in entity table 308.
[0092] Figure 5 An example process flow of a computational architecture for facial identity preservation using an image-to-image model using a stable diffusion generative model, in accordance with an embodiment of the subject technology, is shown.
[0093] Figure 5 The example shows that the input is projected into a smaller latent space and a diffusion model is applied. For example, the encoder network encodes the input into a latent representation, which can reduce the use of computational resources by processing the input in a lower-dimensional space. Next, a standard diffusion model (e.g., U-Net) can be applied to generate new data and upsampled by the decoder network.
[0094] The diffusion model can be understood as a probabilistic model designed to learn the data distribution p(x) by gradually denoising the normally distributed variables, which corresponds to the inverse process of learning a fixed Markov chain of length T.
[0095] Given the trained perceptual compression models shown as E and D, an efficient, low-dimensional latent space is accessible in which high-frequency, imperceptible details are abstracted out. This space is more suitable for likelihood-based generative models than a high-dimensional pixel space (e.g., pixel space 502) because such models can focus on the semantically important bits of the data and are trained in a lower-dimensional, computationally more efficient space.
[0096] Figure 5 The neural backbone of the model shown in is implemented as a temporal conditional U-Net model. In this example, the forward process is fixed and the noise z can be effectively obtained from the encoder E during training. t (e.g., a noise latent vector), and samples from p(z) can be decoded into image space with a single pass through the decoder D.
[0097] As shown, pixel space 502 includes an input image (x) received as input to an encoder (ε). The encoder projects the input image into a latent vector (z) included in latent space 504.
[0098] In the example, given an image x∈R in RGB space H×W×3 , the encoder E encodes x into a latent representation z = E(x), and the decoder D reconstructs the image based on the latent representation, giving where z∈R h×w×c In addition, the encoder can downsample the image by a factor of f = H / h = W / w, where the different downsampling factors are f = 2 m , where m∈N.
[0099] As part of the perceptual compression process, input data (e.g., input image) is converted into a different representation by removing high-frequency details, where a generative adversarial network (GAN) can perform such perceptual compression. In an example, such a GAN projects high-dimensional redundant data from pixel space 502 to a hyperspace called a latent space (latent space 504). It will be understood that the latent vectors in the latent space are compressed forms of the original pixel image (e.g., input image) and can be used in place of the original pixel image. The encoder shown in pixel space 502 can be implemented as an autoencoder (AE) structure that performs perceptual compression. In an embodiment, the encoder of the AE structure projects the high-dimensional data to the latent space, and the corresponding decoder converts the image back from the latent space 504.
[0100] In the latent space 504, a forward diffusion process is applied to the latent vector (z), where noise is gradually added to the latent vector, which generates a second latent vector (z) corresponding to the noisy latent vector. T ). This noisy latent vector is fed to a diffusion model (e.g., U-Net), where the vector undergoes a denoising process (e.g., back-diffusion) to generate the (original) latent vector (z).
[0101] As further shown, in the conditional space 506, a text / image converter (τ θ ), and the text / image converter (τ θ ) can be combined with a denoising network to generate the original latent vector (z).
[0102] For example, the text / image converter (τ θ ) determines the semantic structure in text and image. Therefore, the text / image transformer (τ θ The combination of the detail-preserving capabilities of the ) and diffusion models enables the generation of a fine-grained collection of highly detailed images while preserving the semantic structure in the images.
[0103] Diffusion models (e.g. Figure 5 (shown in ) can model a conditional distribution of the form p(z|y). This can be done by the conditional denoising autoencoder E θ (z t ,t,y) and enables the synthesis process to be controlled by the input y (such as text, semantic mapping or other image-to-image translation tasks).
[0104] To pre-process y from various modalities (e.g., language cues), a domain-specific encoder τ θ Project y to the intermediate representation τ θ (y)∈R M×dτ , this intermediate representation is then mapped to the intermediate layers of a U-Net (e.g., a convolutional neural network) via a cross-attention layer.
[0105] Based on the image-condition pair, we can calculate the θ and E θ to learn the conditional LDM (latent diffusion model). The conditional mechanism can be flexible because τ θ It can be parameterized using domain-specific experts, e.g., when y is a textual cue, a (masked) transformer can be used.
[0106] Returning to the pixel space 502, the decoder (D) upscales the generated (original) latent vector received from the latent space 504 to generate a second image (e.g., output image ).
[0107] In an embodiment, to obtain the initial noise from the original image (e.g., the reverse operation), a reverse sampling process may be performed:
[0108] 1. Input: Image
[0109] 2. Use an Euler sampler with an inverse objective function (e.g., a reverse diffusion process)
[0110] 3. Use text hints to provide descriptions of images
[0111] 4. Get the noise Z_t (e.g., noise latent vector)
[0112] If we utilize reconstruction noise and perform forward sampling again, we can obtain the output image generated by the generative network starting from this noise.
[0113] Figure 6 An example of identity preservation through feature blending in accordance with an embodiment of the subject technology is shown.
[0114] In an implementation, there are two parts of the generative network: an encoder and a decoder (e.g., as in Figure 5 ). The encoder iteratively samples the noise, and the decoder generates the image using an upsampling technique.
[0115] As discussed herein, the noise may have a size corresponding to a set of values of 64x64.
[0116] Regarding upsampling, the technique aims to generate a higher resolution image using convolutional filters (e.g., upsampling the image several times). In addition, the decoder is trained together with the encoder.
[0117] In an implementation, the decoder involves several upsampling layers. In an example, the stable diffusion decoder includes 4 upsampling layers.
[0118] In the example, activations are saved inside the decoder at specific layers and then combined with different model outputs.
[0119] The following are examples of operations performed with feature blending:
[0120] 1. Receive input noise Z_t.
[0121] 2. Provide input noise to the decoder network.
[0122] 3. Perform a partial forward pass and save the first intermediate result (e.g., input for the upsampling layer).
[0123] 4. Receive Z_t2 corresponding to the new input noise.
[0124] 5. Provide new input noise to the decoder network.
[0125] 6. Perform a partial forward pass and save the second intermediate features.
[0126] 7. Receive the segmentation mask M (from 0 to 1) of the input image.
[0127] 8. Apply the formula related to the weighted combination of features and coefficients to produce a new feature set A_new.
[0128] 9. Provide the new feature set A_new to the decoder to complete the forward pass.
[0129] In order to obtain intermediate features for the real image, the initial image is reconstructed and Z_t is obtained. Z_t corresponds to the initial noise used for the initial image reconstruction.
[0130] like Figure 6 As shown in FIG, an input image 602 (e.g., an original real image) is provided. Image 604 corresponds to the original image and the reconstruction result (e.g., a forward pass of the reconstructed original noise). In this example, image 604 includes mostly unchanged facial attributes. As further shown, an intermediate feature set is generated from a forward pass of the reconstructed original noise of the real image (e.g., image 602) in the decoder, and feature blending is performed to generate image 606.
[0131] Figure 7 An example of identity preservation through feature blending in accordance with an embodiment of the subject technology is shown.
[0132] As shown, the original (true) image is shown as image 702 , and a segmentation mask 704 of the image 702 is provided.
[0133] In image set 706, the effect of blending the fourth, third, second, and first upsampling layers (e.g., from right to left) is shown, respectively. In image set 706, the hair strokes become less visible when blending at deeper levels (e.g., from the leftmost image to the rightmost image). In this example, the upsampling layer sets correspond to resolutions of 64, 128, 256, and 512, respectively. In this example, feature blending works like image blending, except that the decoder makes the edges of the images invisible.
[0134] Figure 8 is a flow chart illustrating a method according to certain example embodiments. The method may be implemented in computer-readable instructions executed by one or more computer processors, such that the operations of the method may be performed in part or in whole by a computer client device 118. However, it should be understood that at least some of the operations of the method may be deployed on various other hardware configurations, and the method is not intended to be limited to the computer client device 118 or any of the components or systems mentioned above.
[0135] According to some examples, at operation 802 , the computer client device 118 receives an input image and a segmentation mask of the input image.
[0136] At operation 804 , the computer client device 118 obtains reconstructed noise of the input image using the input image and the segmentation mask.
[0137] At operation 806 , the computer client device 118 determines a first set of features by performing a first portion of a forward pass through the decoder to reconstruct the noise.
[0138] At operation 808 , the computer client device 118 determines a second set of features by processing the input image using an image-to-image (IMG2IMG) model for stable diffusion.
[0139] At operation 810 , the computer client device 118 generates a third set of features based on combining the first set of features and the second set of features with reconstruction noise using the segmentation mask.
[0140] At operation 812 , the computer client device 118 generates an output image by performing the remainder of the forward pass of the third feature set through the decoder.
[0141] In an embodiment, the computer client device 118 performs other operations including: applying an Euler sampler with an inverse objective function to the particular input image; receiving information related to a description of the particular input image, the description based on text input received from an input text prompt of a client application executed on the client device; and generating a noise latent vector for the particular input image based on the received information.
[0142] In an embodiment, the computer client device 118 performs other operations including: performing forward sampling on a noise latent vector for a particular input image; and generating the input image based at least in part on performing the forward sampling.
[0143] In an embodiment, the computer client device 118 performs other operations including: receiving a particular noise latent vector for a particular input image; providing the particular noise latent vector to a decoder network; and determining a first set of intermediate features by performing a partial forward pass of the particular noise latent vector through the decoder network.
[0144] In an embodiment, the computer client device 118 performs other operations including: receiving a second specific noise latent vector for the specific input image; providing the second specific noise latent vector to the decoder network; and determining a second set of intermediate features by performing a second partial forward pass of the second specific noise latent vector through the decoder network.
[0145] In an embodiment, the computer client device 118 performs other operations including performing feature blending on the first intermediate feature set, the second intermediate feature set, and the segmentation mask of the particular input image.
[0146] In embodiments, the computer client device 118 performs other operations including generating an output image based at least in part on performing feature blending.
[0147] In an embodiment, performing feature blending is based on a set of upsampling layers.
[0148] In an embodiment, the upsampling layer set corresponds to a resolution set comprising sizes of 64, 128, 256, and 512, respectively.
[0149] In an embodiment, a feature set is less visible in a first upsampling layer corresponding to a size of 512 than in a second upsampling layer corresponding to a size of 64, the feature set including details of hair on a person's face.
[0150] Machine Architecture
[0151] Figure 9900 within which instructions 902 (e.g., software, programs, applications, applet, apps, or other executable code) may be executed for causing the machine 900 to perform any one or more of the methodologies discussed herein. For example, the instructions 902 may cause the machine 900 to perform any one or more of the methodologies described herein. The instructions 902 transform a general-purpose, unprogrammed machine 900 into a specific machine 900 that is programmed to perform the functions described and illustrated in the manner described. The machine 900 may operate as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 900 may operate as a server or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 900 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing instructions 902 specifying actions to be taken by the machine 900. In addition, although only a single machine 900 is shown, the term "machine" should also be construed to include a collection of machines that individually or jointly execute instructions 902 to perform any one or more of the methods discussed herein. For example, the machine 900 may include the user system 102 or any server device of a plurality of server devices forming part of the interactive server system 110. In some examples, the machine 900 may also include both a client system and a server system, wherein certain operations of a particular method or algorithm are performed on the server side and certain operations of the particular method or algorithm are performed on the client side.
[0152] The machine 900 may include a processor 904, a memory 906, and input / output (I / O) components 908 that may be configured to communicate with each other via a bus 910. In an example, the processor 904 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 912 that executes instructions 902 and a processor 914. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that may execute instructions concurrently. Although Figure 9 Multiple processors 904 are shown, but the machine 900 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0153] The memory 906 includes a main memory 916, a static memory 918, and a storage unit 920, all of which are accessible to the processor 904 via the bus 910. The main memory 906, the static memory 918, and the storage unit 920 store instructions 902 that embody any one or more of the methodologies or functions described herein. The instructions 902 may also reside, completely or partially, within the main memory 916, within the static memory 918, within a machine-readable medium 922 within the storage unit 920, within at least one of the processors 904 (e.g., within a cache memory of a processor), or any suitable combination thereof during execution thereof by the machine 900.
[0154] The I / O components 908 may include various components for receiving input, providing output, generating output, transmitting information, exchanging information, capturing measurements, etc. The specific I / O components 908 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine will be less likely to include such a touch input device. It should be appreciated that the I / O components 908 may include Figure 9 Many other components are not shown in the figure. In various examples, the I / O components 908 may include user output components 924 and user input components 926. The user output components 924 may include visual components (e.g., displays such as plasma display panels (PDPs), light emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), acoustic components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. The user input components 926 may include alphanumeric input components (e.g., keyboards, touch screens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), pointing-based input components (e.g., mice, touch pads, trackballs, joysticks, motion sensors, or other pointing instruments), tactile input components (e.g., physical buttons, touch screens or other tactile input components that provide touch location and force or touch gestures), audio input components (e.g., microphones), etc.
[0155] In other examples, the I / O component 908 may include a biometric component 928, a motion component 930, an environmental component 932, or a positioning component 934, as well as a wide range of other components. For example, the biometric component 928 includes components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion component 930 includes an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, and a rotation sensor component (e.g., a gyroscope).
[0156] Environmental components 932 include, for example, one or more cameras (with still image / photo and video capabilities), an illumination sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor for detecting concentrations of hazardous gases for safety purposes or for measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment.
[0157] With respect to cameras, the user system 102 can have a camera system that includes, for example, a front-facing camera on the front surface of the user system 102 and a rear-facing camera on the rear surface of the user system 102. The front-facing camera can be used, for example, to capture still images and videos of the user of the user system 102 (e.g., “selfies”), which can then be enhanced with the enhancement data (e.g., filters) described above. The rear-facing camera can be used, for example, to capture still images and videos in a more conventional camera mode, where the images are similarly enhanced with the enhancement data. In addition to the front-facing camera and the rear-facing camera, the user system 102 can also include a 360° camera for capturing 360° photos and videos.
[0158] Additionally, the camera system of the user system 102 may include dual rear cameras (e.g., a main camera and a depth-sensing camera), or even triple, quad, or quintuple rear camera configurations on both the front and back sides of the user system 102. For example, these multi-camera systems may include a wide-angle camera, an ultra-wide-angle camera, a telephoto camera, a macro camera, and a depth sensor.
[0159] The positioning component 934 includes a position sensor component (for example, a GPS receiver component), an altitude sensor component (for example, an altimeter or barometer that detects air pressure, and the altitude can be obtained based on the air pressure), an orientation sensor component (for example, a magnetometer), etc.
[0160] Various technologies can be used to achieve communication. The I / O components 908 also include a communication component 936 that is operable to couple the machine 900 to a network 938 or device 940 via corresponding couplings or connections. For example, the communication component 936 may include a network interface component or another suitable device that interfaces with the network 938. In other examples, the communication component 936 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, Components (e.g. Low energy consumption), Components and other communication components for providing communication via other modalities. Device 940 can be another machine or any of a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0161] In addition, the communication component 936 can detect an identifier or include a component operable to detect an identifier. For example, the communication component 936 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional bar codes such as Universal Product Code (UPC) bar codes, multi-dimensional bar codes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, UltraCode, UCC RSS-2D bar codes, and other optical codes) or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tag). In addition, various information can be obtained via the communication component 936, such as location via Internet Protocol (IP) geolocation, location information via Signal triangulation to obtain location, location obtained by detecting NFC beacon signals that can indicate a specific location, etc.
[0162] Various memories (e.g., main memory 916, static memory 918, and memory of processor 904) and storage unit 920 may store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. These instructions (e.g., instructions 902) when executed by processor 904 cause various operations to implement the disclosed examples.
[0163] Instructions 902 may be sent or received over network 938 via a network interface device (e.g., a network interface component included in communications component 936) using a transmission medium and using any of several well-known transmission protocols, such as the Hypertext Transfer Protocol (HTTP). Similarly, instructions 902 may be sent or received via a coupling (e.g., a peer-to-peer coupling) with device 940 using a transmission medium.
[0164] Software Architecture
[0165] Figure 10 10 is a block diagram 1000 illustrating a software architecture 1002 that can be installed on any one or more of the devices described herein. The software architecture 1002 is supported by hardware such as a machine 1004 including a processor 1006, memory 1008, and I / O components 1010. In this example, the software architecture 1002 can be conceptualized as a stack of layers, where each layer provides specific functionality. The software architecture 1002 includes layers such as an operating system 1012, libraries 1014, frameworks 1016, and applications 1018. In operation, the application 1018 invokes an API call 1020 through the software stack and receives a message 1022 in response to the API call 1020.
[0166] The operating system 1012 manages hardware resources and provides common services. The operating system 1012 includes, for example, a kernel 1024, services 1026, and drivers 1028. The kernel 1024 acts as an abstraction layer between the hardware layer and other software layers. For example, the kernel 1024 provides memory management, processor management (e.g., scheduling), component management, networking and security settings, and other functions. Services 1026 can provide other common services to other software layers. Drivers 1028 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 1028 may include display drivers, camera drivers, or Low-power drivers, Flash drivers, serial communication drivers (e.g., USB drivers), drivers, audio drivers, power management drivers, etc.
[0167] The libraries 1014 provide a common low-level infrastructure used by the applications 1018. The libraries 1014 may include a system library 1030 (e.g., a C standard library) that provides functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, the libraries 1014 may include an API library 1032 such as a media library (e.g., a library for supporting the presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), a graphics library (e.g., an OpenGL framework for rendering graphical content on a display in two dimensions (2D) and three dimensions (3D), a database library (e.g., SQLite providing various relational database functions), a web library (e.g., WebKit providing web browsing functions), etc. The library 1014 may also include various other libraries 1034 to provide many other APIs to the application 1018 .
[0168] The framework 1016 provides a common high-level infrastructure used by applications 1018. For example, the framework 1016 provides various graphical user interface (GUI) functions, advanced resource management, and advanced positioning services. The framework 1016 can provide a wide range of other APIs that can be used by applications 1018, some of which may be specific to a particular operating system or platform.
[0169] In an example, applications 1018 may include a home application 1036, a contacts application 1038, a browser application 1040, a book reader application 1042, a location application 1044, a media application 1046, a messaging application 1048, a game application 1050, and a variety of other applications such as third-party applications 1052. Applications 1018 are programs that perform functions defined in the program. Various programming languages may be used to create one or more of the applications 1018 structured in various ways, such as an object-oriented programming language (e.g., Objective-C, Java, or C++) or a procedural programming language (e.g., C or assembly language). In a specific example, third-party applications 1052 (e.g., those developed by entities other than the vendor of a particular platform using ANDROID) may be used to create a third-party application 1052. TM or IOS TM Software Development Kit (SDK) can be used to develop applications on platforms such as IOS TM ANDROID TM 、 Mobile software running on the mobile operating system of the phone or another mobile operating system. In this example, third-party applications 1052 can activate API calls 1020 provided by the operating system 1012 to facilitate the functions described herein.
[0170] Glossary
[0171] "Carrier signal" refers to any intangible medium capable of storing, encoding, or carrying instructions for execution by a machine and includes, for example, digital or analog communication signals, or other intangible medium that facilitates the transmission of such instructions. Instructions may be transmitted or received over a network using a transmission medium via a network interface device.
[0172] "Client device" refers to any machine that interfaces with a communications network to obtain resources from one or more server systems or other client devices, for example. A client device may be, but is not limited to, a mobile phone, a desktop computer, a laptop computer, a portable digital assistant (PDA), a smartphone, a tablet computer, an ultrabook, a netbook, a laptop computer, a multiprocessor system, a microprocessor-based or programmable consumer electronics product, a game console, a set-top box, or any other communications device that a user can use to access a network.
[0173] "Communications network" refers to, for example, one or more parts of a network, which may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a Plain Old Telephone Service (POTS) network, a cellular telephone network, a wireless network, The coupling may be a network, another type of network, or a combination of two or more such networks. For example, the network or a portion of the network may include a wireless network or a cellular network, and the coupling may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, the coupling may implement any of various types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rates for GSM evolution (EDGE) technology, the third generation partnership project (3GPP) including 3G, fourth generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), world interoperability for microwave access (WiMAX), long term evolution (LTE) standards, other data transmission technologies defined by various standards development organizations, other long distance protocols, or other data transmission technologies.
[0174] "Component" refers to a logic or device, a physical entity, for example, having the following boundaries, which are defined by function or subroutine calls, branch points, APIs, or other technologies that provide partitioning or modularization for specific processing or control functions. A component can be combined with other components via its interface to perform machine processing. A component can be a packaged functional hardware unit designed for use with other components, and a part of a program that generally performs a specific function of a related function. A component can constitute a software component (for example, a code embodied on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit that can perform certain operations and can be configured or arranged in a certain physical manner. In various examples, one or more computer systems (for example, a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components (for example, a processor or a processor group) of a computer system can be configured by software (for example, an application or an application part) to be a hardware component that operates to perform certain operations as described herein. Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component can include a dedicated circuit or logic that is permanently configured to perform certain operations. The hardware component can be a special-purpose processor, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). The hardware component can also include a programmable logic or circuit that is temporarily configured to perform certain operations by software. For example, the hardware component can include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific component of a machine), which is uniquely customized to perform the configured function and is no longer a general-purpose processor. It will be appreciated that it can be decided whether to mechanically implement the hardware component in a dedicated and permanently configured circuit or in a temporarily configured (e.g., configured by software) circuit for cost and time considerations. Therefore, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, that is, an entity that is physically constructed, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in some way or perform certain operations described herein. Considering an example in which a hardware component is temporarily configured (e.g., programmed), it is not necessary to configure or instantiate each hardware component in the hardware component at any one time. For example, in the case where a hardware component includes a general-purpose processor that is configured by software to become a special-purpose processor, the general-purpose processor can be configured as a different special-purpose processor (e.g., including different hardware components) at different times. The software configures a specific one or more processors accordingly, such as to constitute a specific hardware component at one time and to constitute a different hardware component at a different time. A hardware component can provide information to other hardware components and receive information from other hardware components. Therefore, the described hardware components can be considered to be communicatively coupled.In the case of multiple hardware components being present at the same time, communication can be achieved by signal transmission between or among two or more hardware components (for example, by appropriate circuits and buses). In the example in which multiple hardware components are configured or instantiated at different times, communication between such hardware components can be achieved, for example, by storing information in a memory structure accessible to multiple hardware components and retrieving information in the memory structure. For example, a hardware component can perform an operation, and the output of the operation is stored in a memory device coupled to its communication ground. Then, other hardware components can access the memory device at a subsequent time to retrieve the stored output and process it. The hardware component can also initiate communication with an input device or an output device, and can operate on resources (for example, a collection of information). The various operations of the example methods described herein can be performed at least in part by temporarily configuring (for example, by software) or permanently configuring one or more processors to perform related operations. Whether it is temporarily configured or permanently configured, such a processor can constitute a processor-implemented component that operates to perform one or more operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the method described in this article can be implemented at least in part by a processor, wherein specific one or more processors are examples of hardware. For example, at least some operations in the operation of the method can be performed by one or more processors or the parts implemented by the processor. In addition, one or more processors can also operate to support the execution of related operations in a "cloud computing" environment or operate as "software as a service" (SaaS). For example, at least some operations in the operation can be performed by a computer group (as an example of a machine including a processor), wherein these operations can be accessed via a network (for example, the Internet) and via one or more appropriate interfaces (for example, API). The execution of some operations can be distributed between processors, not only resident in a single machine, but deployed across multiple machines. In some examples, a processor or the parts implemented by a processor can be located in a single geographical location (for example, in a home environment, an office environment or a server farm). In other examples, a processor or the parts implemented by a processor can be distributed across multiple geographical locations.
[0175] "Computer-readable storage media" refers to, for example, both machine storage media and transmission media. Thus, these terms encompass both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" mean the same thing and may be used interchangeably in this disclosure.
[0176] A "transient message" is a message that is accessible for a limited duration, for example. A transient message can be text, an image, a video, or the like. The access time for a transient message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is transient.
[0177] “Machine storage media” refers to, for example, a single or multiple storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Thus, the term should be taken to include, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGAs, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms “machine storage media,” “device storage media,” and “computer storage media” mean the same thing and are used interchangeably in this disclosure. The terms “machine storage media,” “computer storage media,” and “device storage media” expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are encompassed by the term “signal media.”
[0178] “Non-transitory computer-readable storage medium” refers to a tangible medium that is capable of storing, encoding, or carrying instructions for execution by a machine, for example.
[0179] "Signal medium" refers to any intangible medium that is capable of storing, encoding, or carrying instructions for execution by a machine and includes, for example, digital or analog communication signals, or other intangible media that facilitates the transmission of software or data. The term "signal medium" should be construed to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.
[0180] A "user device" refers to a device that is, for example, accessed, controlled, or owned by a user and with which the user interacts to perform actions or interact with other users or computer systems.
Claims
1. A method comprising: receiving an input image and a segmentation mask of the input image; obtaining a reconstructed noise of the input image using the input image and the segmentation mask; determining a first set of features by performing a first portion of a forward pass of said reconstructed noise through a decoder; determining a second set of features by processing the input image using an image-to-image (IMG2IMG) model for stable diffusion; generating a third set of features based on combining the first set of features and the second set of features with the reconstruction noise using the segmentation mask; as well as An output image is generated by performing the remainder of a forward pass of the third feature set through the decoder.
2. The method according to claim 1, further comprising: Apply an Euler sampler with an inverse objective function to a specific input image; receiving information related to a description of the particular input image, the description being based on text input received from an input text prompt of a client application executing on a client device; and A noise latent vector for the specific input image is generated based on the received information.
3. The method according to claim 2, further comprising: performing forward sampling on a noise latent vector of the specific input image; as well as The input image is generated based at least in part on performing the forward sampling.
4. The method according to claim 1, further comprising: Receives a specific noise latent vector for a specific input image; providing the specific noise latent vector to a decoder network; as well as A first set of intermediate features is determined by performing a partial forward pass of the particular noise latent vector through the decoder network.
5. The method according to claim 4, further comprising: receiving a second specific noise latent vector of the specific input image; providing the second specific noise latent vector to the decoder network; as well as A second set of intermediate features is determined by performing a second partial forward pass of the second specific noise latent vector through the decoder network.
6. The method according to claim 5, further comprising: Feature blending of the first set of intermediate features, the second set of intermediate features, and a segmentation mask of the specific input image is performed.
7. The method according to claim 6, further comprising: The output image is generated based at least in part on performing the feature blending.
8. The method according to claim 6, wherein: The feature mixing is performed based on a set of upsampling layers.
9. The method according to claim 8, wherein The upsampling layer set corresponds to a resolution set, which includes sizes of 64, 128, 256 and 512 respectively.
10. The method according to claim 9, wherein: A set of features, including details of hair on a person's face, is less visible in the first upsampling layer corresponding to a size of 512 than in the second upsampling layer corresponding to a size of 64.
11. A system comprising: processor; as well as a memory comprising instructions that, when executed by the processor, cause the processor to perform operations comprising: receiving an input image and a segmentation mask of the input image; obtaining a reconstructed noise of the input image using the input image and the segmentation mask; determining a first set of features by performing a first portion of a forward pass of said reconstructed noise through a decoder; determining a second set of features by processing the input image using an image-to-image (IMG2IMG) model for stable diffusion; generating a third set of features based on combining the first set of features and the second set of features with the reconstruction noise using the segmentation mask; and An output image is generated by performing the remainder of a forward pass of the third feature set through the decoder.
12. The system of claim 11, further comprising: Apply an Euler sampler with an inverse objective function to a specific input image; receiving information related to a description of the particular input image, the description being based on text input received from an input text prompt of a client application executing on a client device; and A noise latent vector for the specific input image is generated based on the received information.
13. The system of claim 12, further comprising: performing forward sampling on a noise latent vector of the specific input image; as well as The input image is generated based at least in part on performing the forward sampling.
14. The system of claim 11, further comprising: Receives a specific noise latent vector for a specific input image; providing the specific noise latent vector to a decoder network; as well as A first set of intermediate features is determined by performing a partial forward pass of the particular noise latent vector through the decoder network.
15. The system of claim 14, further comprising: receiving a second specific noise latent vector of the specific input image; providing the second specific noise latent vector to the decoder network; as well as A second set of intermediate features is determined by performing a second partial forward pass of the second specific noise latent vector through the decoder network.
16. The system of claim 15, further comprising: Feature blending of the first set of intermediate features, the second set of intermediate features, and a segmentation mask of the specific input image is performed.
17. The system of claim 16, further comprising: The output image is generated based at least in part on performing the feature blending.
18. The system according to claim 16, wherein: The feature mixing is performed based on a set of upsampling layers.
19. The system according to claim 18, wherein The upsampling layer set corresponds to a resolution set, and the resolution set includes sizes of 1614, 111218, 121516 and 151112 respectively.
20. A non-transitory computer-readable medium comprising instructions that, when executed by a computing device, cause the computing device to perform operations comprising: receiving an input image and a segmentation mask of the input image; obtaining a reconstructed noise of the input image using the input image and the segmentation mask; determining a first set of features by performing a first portion of a forward pass of said reconstructed noise through a decoder; determining a second set of features by processing the input image using an image-to-image (IMG2IMG) model for stable diffusion; generating a third set of features based on combining the first set of features and the second set of features with the reconstruction noise using the segmentation mask; as well as An output image is generated by performing the remainder of a forward pass of the third feature set through the decoder.