Generative model for suggesting image modifications
By using generative ML models and large language models in interactive applications, content items are automatically selected and modified, solving the problem of users creating high-quality shareable content and improving efficiency and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SNAP INC
- Filing Date
- 2024-10-22
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, users need to spend a lot of time and resources to create high-quality images or videos to share with other users, and using generative ML models to modify content and provide suggestions is time-consuming and expensive, which reduces the usability and enjoyment of the system.
By leveraging generative ML models and large language models through interactive applications, content items that users are interested in can be automatically selected and modified, personalized content modification prompts can be generated, and content items can be processed to improve the quality and efficiency of shared content.
It reduces the time and resources required to create high-quality, shareable content, improves the user experience, and makes the use of generative ML models more efficient and engaging.
Smart Images

Figure CN122095355A_ABST
Abstract
Description
[0001] Priority Statement
[0002] This application claims the benefit of priority to U.S. Provisional Application Serial No. 63 / 592,397, filed October 23, 2023, and U.S. Patent Application Serial No. 18 / 634,369, filed April 12, 2024, each of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure generally pertains to interactive applications used for sharing content items. Background Technology
[0004] Augmented reality (AR) is a virtual modification of the real environment. For example, in virtual reality (VR), the user is fully immersed in a virtual world, while in AR, the user is immersed in a world where virtual objects are combined or overlaid on top of real-world objects. AR systems are designed to generate and present virtual objects that interact realistically with the real-world environment. Examples of AR applications can include single-player or multiplayer video games, instant messaging systems, and more. Generally, these AR and / or VR systems are referred to as extended reality (XR) systems. Attached Figure Description
[0005] In the accompanying drawings (which are not necessarily drawn to scale), the same reference numerals may describe similar parts in different views. To facilitate identification of any discussion of a particular element or action, one or more of the most significant digits in the reference numerals indicate the drawing number in which the element was first introduced. Some non-limiting examples are shown in the accompanying drawings:
[0006] Figure 1 It is a diagrammatic representation of a networked environment in which the content of this disclosure can be deployed, based on some examples.
[0007] Figure 2 It is a graphical representation of a messaging system with both client-side and server-side functionalities, based on some examples.
[0008] Figure 3 It is a graphical representation based on examples such as data structures maintained in a database.
[0009] Figure 4 It is a graphical representation based on some example messages.
[0010] Figure 5 The example architecture is shown for generating shareable content items based on some examples of applying personalized personal artificial intelligence (AI) agents.
[0011] Figure 6A , Figure 6B, Figure 7A , Figure 7B and Figure 8 It is a graphical representation of example inputs and outputs of a system that feeds shareable content items, based on some examples.
[0012] Figure 9 , Figure 10A and Figure 10B This is a flowchart illustrating example operations and methods of a shareable content item feeding system based on some examples.
[0013] Figure 11 It is a graphical representation of a machine in the form of a computer system, based on some examples, within which a set of instructions can be executed to cause the machine to perform any or more of the methods discussed herein.
[0014] Figure 12 It is a block diagram showing an example of a software architecture that can be implemented therein.
[0015] Figure 13 The system shown is one of several examples in which a head-worn device can be implemented. Detailed Implementation
[0016] The following description includes systems, methods, techniques, instruction sequences, and computer program products that embody illustrative examples of the present disclosure. In the following description, numerous specific details are set forth for illustrative purposes in order to provide an understanding of various examples. However, it will be apparent to those skilled in the art that the examples can be practiced without these specific details. Generally, well-known examples of instructions, protocols, structures, and techniques are not necessarily shown in detail.
[0017] Typically, various communication platforms allow users to share content and create images to transmit to other users. These images can be used to promote products or services and / or simply represent different real-world objects in simulated or real-world environments. However, these systems require users to use expensive equipment and technology to create high-quality, engaging images. Furthermore, users may expend considerable effort carefully searching for images / videos to actually share. Once such images / videos are found, users spend even more time and effort placing the objects in different environments and manually adjusting lighting and other image properties to enhance their presentation. All these factors combined can make creating high-quality images (e.g., for sharing with other users) extremely costly and diminish the overall usability and enjoyment of the system. Additionally, users may miss opportunities to share and present objects with ideal settings because they may lack the resources needed to create high-quality images.
[0018] As generative machine learning (ML) models become more widespread, the ability to program new tasks using them has become easier. That said, generative ML models (e.g., AI that generates content) are currently a popular approach for programming specific operations and actions, and can potentially generate content from 2D models, 3D models, code, and more. However, to properly operate these generative ML models and obtain satisfactory results, specific prompts and specific types of input are required. Generating these prompts and selecting inputs to drive generative ML models is time-consuming and therefore expensive, further diminishing their overall usability and enjoyment. Consequently, users often avoid using these generative ML models to enhance content items intended for sharing with other users.
[0019] The disclosed technology seeks to improve the efficiency of using electronic devices by intelligently and automatically selecting content items that a user might be interested in sharing using a generative ML model and / or a large language model (LLM), and then automatically processing the selected content items to modify and improve various visual aspects of the selected content items. This can reduce the overall time and cost associated with identifying and generating content items to be shared with other users. Furthermore, the disclosed technology utilizes learned information about the user to dynamically generate context-sensitive cues for the generative ML model, enabling the modifications performed by the generative ML model on each individual content item among the selected content items to be personalized.
[0020] For example, the disclosed technology selects individual content items from multiple previously captured content items through an interactive application that match one or more criteria corresponding to shareable content, and generates a prompt including the individual content item and a request for multiple suggested modifications to the individual content item. The disclosed technology processes the prompt using LLM to generate multiple suggested modifications to the individual content item, and generates a modified individual content item corresponding to the individual suggested modification among the multiple suggested modifications. In this way, the disclosed technology improves the overall user experience when using electronic devices and reduces the total amount of resources required to complete the task of creating high-quality and unique shareable content items.
[0021] Networked computing environment
[0022] Figure 1This is a block diagram illustrating an example interactive system 100 for facilitating interactions on a network, such as exchanging text messages, making text audio and video calls, or playing games. Interactive system 100 includes multiple user systems 102, each hosting multiple applications, including interactive clients 104 and other applications 106. Each interactive client 104 is communicatively coupled to other instances of the interactive client 104 (e.g., hosted on corresponding other user systems 102), interactive server systems 110, and third-party servers 112 via one or more communication networks, including network 108 (e.g., the Internet). Interactive client 104 can also communicate with locally hosted applications 106 using application programming interfaces (APIs).
[0023] Each user system 102 may include multiple user devices, such as mobile devices 114, head-mounted wearable devices 116, and computer client devices 118, which are communicatively connected to exchange data and messages.
[0024] Interactive client 104 interacts with other interactive clients 104 and with interactive server system 110 via network 108. The data exchanged between interactive clients 104 (e.g., interaction 120) and between interactive client 104 and interactive server system 110 includes functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).
[0025] Interactive server system 110 provides server-side functionality to interactive client 104 via network 108. While some functions of interactive system 100 are described herein as being performed by interactive client 104 or interactive server system 110, the location of certain functions within interactive client 104 or interactive server system 110 may be a design choice. For example, it may be technically preferred that specific technologies and functions are initially deployed within interactive server system 110, but later migrated to interactive client 104 of user system 102 with sufficient processing power.
[0026] The interactive server system 110 supports various services and operations provided to the interactive client 104. Such operations include sending data to and receiving data from the interactive client 104, and processing data generated by the interactive client 104. This data may include message content, client device information, geolocation information, media enhancements and overlays, message content persistence conditions, entity relationship information, and live event information. Data exchange within the interactive system 100 is activated and controlled via functions available through the user interface (UI) of the interactive client 104.
[0027] Now, specifically, the focus shifts to interactive server system 110. API server 122 is coupled to interactive server 124 and provides it with a programming interface, making the functionality of interactive server 124 accessible to interactive client 104, other applications 106, and third-party server 112. Interactive server 124 is communicatively coupled to database server 126, thereby facilitating access to database 128, which stores data associated with the interactions processed by interactive server 124. Similarly, web server 130 is coupled to interactive server 124 and provides a web-based interface to interactive server 124. To this end, web server 130 handles incoming network requests via Hypertext Transfer Protocol (HTTP) and several other related protocols.
[0028] API server 122 receives and sends interactive data (e.g., command and message payloads) between interactive server 124 and user system 102 (as well as interactive client 104 and other applications 106) and third-party server 112. Specifically, API server 122 provides a set of interfaces (e.g., routines and protocols) that interactive client 104 and other applications 106 can call or query to activate the functionality of interactive server 124. API server 122 exposes various functions supported by interactive server 124, including account registration; login functionality; sending interactive data from one interactive client 104 to another interactive client 104 via interactive server 124; transferring media files (e.g., images or videos) from interactive client 104 to interactive server 124; setting media data sets (e.g., stories); retrieving the friend list of users in user system 102; retrieving messages and content; adding and deleting entities (e.g., friends) against an entity graph (e.g., entity graph 310); locating friends within the entity graph; and opening (e.g., application events associated with interactive client 104).
[0029] Interactive server 124 hosts multiple systems and subsystems, as shown below. Figure 2 Describe it.
[0030] Application of links
[0031] Returning to interactive client 104, the features and functionality of external resources (e.g., linked application 106 or applet) are available to the user via the interface of interactive client 104. In this context, "external" refers to the fact that application 106 or applet is outside of interactive client 104. External resources are typically provided by third parties, but may also be provided by the creator or provider of interactive client 104. Interactive client 104 receives user selections regarding options for launching or accessing the features of such external resources. External resources may be application 106 installed on user system 102 (e.g., a "local app"), or a smaller version (e.g., a "app") of an application hosted on user system 102 or located remotely on user system 102 (e.g., on a third-party server 112). A smaller version of an application includes a subset of the features and functionality of the application (e.g., a full-scale, local version of the application) and is implemented using markup language documentation. In some examples, a smaller version of an application (e.g., a "app") is a web-based markup language version of the application and is embedded in interactive client 104. In addition to using markup language documentation (e.g., ...), other applications may also use markup language documentation. In addition to ml files, mini-programs can incorporate scripting languages (e.g., ...). .js files or .json files) and stylesheets (e.g., ...). (ss file).
[0032] In response to receiving a user selection of an option for launching or accessing an external resource, interactive client 104 determines whether the selected external resource is a web-based external resource or a locally installed application 106. In some cases, application 106, locally installed on user system 102, can be launched independently of and separately from interactive client 104, for example, by selecting the icon corresponding to application 106 on the home screen of user system 102. A smaller version of such an application can be launched or accessed via interactive client 104, and in some examples, no part of the smaller application can be accessed outside of interactive client 104, or only a limited portion of the smaller application can be accessed outside of interactive client 104. A smaller application can be launched by interactive client 104 by receiving, for example, markup language documents associated with the smaller application from third-party server 112 and processing such documents.
[0033] In response to determining that the external resource is a locally installed application 106, the interactive client 104 instructs the user system 102 to launch the external resource by executing locally stored code corresponding to the external resource. In response to determining that the external resource is a web-based resource, the interactive client 104 communicates with a third-party server 112 (e.g.) to obtain a markup language document corresponding to the selected external resource. The interactive client 104 then processes the obtained markup language document to render the web-based external resource within the UI of the interactive client 104.
[0034] Interactive client 104 can notify users of user system 102 or other users (e.g., "friends") associated with such users of one or more external resources. For example, interactive client 104 can provide participants in a conversation (e.g., a chat session) within interactive client 104 with notifications related to external resources currently or recently used by one or more members of a group of users. One or more users can be invited to join an active external resource or to activate (in a group of friends) a recently used but currently inactive external resource. External resources can provide participants in the conversation, each using their respective interactive client 104, with the ability to share items, conditions, states, or locations within the external resource with one or more members of a group of users during the chat session. Shared items can be interactive chat cards that chat members can interact with to, for example, activate the corresponding external resource, view specific information within the external resource, or take a chat member to a specific location or state within the external resource. Within a given external resource, response messages can be sent to users on interactive client 104. External resources can selectively include different media items in the response based on the current context of the external resource.
[0035] Interactive client 104 can present a list of available external resources (e.g., application 106 or mini-program) to the user to launch or access a given external resource. This list can be presented in a context-sensitive menu. For example, the icons representing different applications (or mini-programs) of application 106 (or mini-program) can vary based on how the user launches the menu (e.g., from a conversational interface or from a non-conversational interface).
[0036] System Architecture
[0037] Figure 2 This is a block diagram illustrating further details of the interactive system 100 according to some examples. Specifically, the interactive system 100 is shown as including an interactive client 104 and an interactive server 124. The interactive system 100 includes multiple subsystems, which are supported on the client side by the interactive client 104 and on the server side by the interactive server 124.
[0038] In some examples, these subsystems are implemented as microservices. A microservice subsystem (e.g., a microservice application) can have components that enable it to operate independently and communicate with other services. Example components of a microservice subsystem may include:
[0039] Functional logic: Functional logic implements the functions of the microservice subsystem and represents the specific capabilities or functions provided by the microservice.
[0040] API Interface: Microservices can communicate with other components using lightweight protocols such as REST or messaging through well-defined APIs or interfaces. The API interface defines the inputs and outputs of a microservice subsystem and how it interacts with other microservice subsystems of the interactive system 100.
[0041] Data storage: The microservice subsystem can be responsible for its own data storage, which can be in the form of a database, cache, or other storage mechanisms (e.g., using database server 126 and database 128). This allows the microservice subsystem to operate independently of other microservices in the interactive system 100.
[0042] Service discovery: Microservice subsystems can find and communicate with other microservice subsystems in the interactive system 100. The service discovery mechanism enables microservice subsystems to locate and communicate with other microservice subsystems in a scalable and efficient manner.
[0043] Monitoring and logging: Microservice subsystems may need to be monitored and logged to ensure availability and performance. Monitoring and logging mechanisms enable the tracking of the health and performance of microservice subsystems.
[0044] In some examples, the interactive system 100 may employ a monolithic architecture, a service-oriented architecture (SOA), a function-as-a-service (FaaS) architecture, or a modular architecture:
[0045] The image processing system 202 provides various functions that enable users to capture and enhance (e.g., annotate or otherwise modify or edit) media content associated with a message.
[0046] The camera device system 204 includes (e.g., in a camera device application) control software that interacts with and controls the camera device hardware of the user system 102 (e.g., directly or via operating system controls) to modify and enhance real-time images captured and displayed via the interactive client 104.
[0047] Enhancement system 206 provides functionality related to the generation and distribution of enhancements (e.g., media overlays) of images captured in real time by the camera device of user system 102 or images retrieved from the memory of user system 102 (e.g., previously captured images). Some examples of images are discussed, but similar techniques are applied to any kind of content item, including animations, graphics, video, audio files, image collages, etc. For example, enhancement system 206 is operable to select, present, and display media overlays (e.g., image filters or image lenses) for interactive client 104 to enhance real-time images received via camera device system 204 or images retrieved from the memory of user system 102 (e.g., previously captured images). Figure 12 The stored images retrieved by memory 1208 (as shown) are enhancements. These enhancements are selected by enhancement system 206 based on some input and data, such as, for example:
[0048] The geographic location of user system 102; and
[0049] User entity relationship information of users in user system 102.
[0050] Enhancements may include audio and visual content and visual effects. Examples of audio and visual content include images, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects may be applied to media content items (e.g., photos and / or videos) at user system 102 for transmission in messages, or to video content such as video content streams or feeds sent from interactive client 104. Therefore, image processing system 202 can interact with and support various subsystems of communication system 208, such as messaging system 210 and video communication system 212.
[0051] Media overlays may include text or image data that can be superimposed on photographs taken by user system 102 or video streams produced by user system 102. In some examples, media overlays may be location overlays (e.g., Venice Beach), names of live events, or names of businesses (e.g., beach cafes). In other examples, image processing system 202 uses the geolocation of user system 102 to identify media overlays that include the name of a business at the geolocation of user system 102. Media overlays may include additional tags associated with the business. Media overlays may be stored in database 128 and accessed through database server 126.
[0052] Image processing system 202 provides a user-based publishing platform that allows users to select a geographic location on a map and upload content associated with that location. Users can also specify which media overlays should be provided to other users. Image processing system 202 generates a media overlay that includes the uploaded content and associates it with the selected geographic location.
[0053] The Enhanced Creation System 214 supports AR developer platforms and includes applications that enable content creators (e.g., artists and developers) to create and publish interactive clients 104, such as AR experiences. The Enhanced Creation System 214 provides content creators with a library of built-in features and tools, including, for example, custom shaders, tracking technologies, and templates.
[0054] In some examples, enhancement creation system 214 provides a merchant-based publishing platform that enables merchants to select specific enhancements associated with geolocation via a bidding process. For example, enhancement creation system 214 associates the media overlay of the highest bidder with a corresponding geolocation for a predefined amount of time.
[0055] Communication system 208 is responsible for enabling and processing various forms of communication and interaction within interactive system 100, and includes messaging system 210, audio communication system 216, and video communication system 212. Messaging system 210 is responsible for enabling temporary or time-limited access to content by interactive client 104. Messaging system 210 incorporates multiple timers (e.g., within user management system 218) that selectively enable access to messages and associated content (e.g., for presentation and display) via interactive client 104 based on duration and display parameters associated with a message or set of messages (e.g., a story). Audio communication system 216 enables and supports audio communication (e.g., real-time audio chat) between multiple interactive clients 104. Similarly, video communication system 212 enables and supports video communication (e.g., real-time video chat) between multiple interactive clients 104.
[0056] User management system 218 is operationally responsible for managing user data and profiles, and maintaining relationships between users and users of interaction system 100 (e.g., stored in...). Figure 3 (Entity information in Entity Table 308, Entity Diagram 310, and Profile Data 302).
[0057] The collection management system 220 is operationally responsible for managing collections or sets of media (e.g., collections of text, images, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into "event galleries" or "event stories." Such collections can be made available for a specified time period (e.g., the duration of the event to which the content relates). For example, content related to a concert can be made available as a "story" for the duration of the concert. The collection management system 220 can also be responsible for publishing icons that provide notifications of specific collections to the UI of the interactive client 104. The collection management system 220 includes curation functions that enable collection managers to manage and curate specific content collections. For example, a curation interface allows event organizers to curate collections of content related to a specific event (e.g., removing inappropriate content or redundant messages). Additionally, the collection management system 220 uses machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, users can be compensated for including user-generated content in a collection. In such cases, the collection management system 220 operates to automatically pay such users for using their content. Any item in a collection of content is sometimes referred to as a "memory".
[0058] Map system 222 provides various geolocation (e.g., geographic location) functions and supports the presentation of map-based media content and messages by interactive client 104. For example, map system 222 enables the display (e.g., stored on) maps. Figure 3 The user's profile data 302 (in which the user's icon or avatar is used) indicates the current or past location of the user's "friends" within the context of the map, as well as media content generated by such friends (e.g., a collection of messages including photos and videos). For example, on the map interface of interactive client 104, messages posted by the user from a specific geographic location to interactive system 100 can be displayed to the specific user's "friends" within the context of that specific location on the map. The user can also share his or her location and status information with other users of interactive system 100 via interactive client 104 (e.g., using an appropriate status avatar), where the location and status information is similarly displayed to selected users within the context of the map interface of interactive client 104.
[0059] Game system 224 provides various game functions within the context of interactive client 104. Interactive client 104 provides a game interface that offers a list of available games that can be initiated by a user within the context of interactive client 104 and played with other users of interactive system 100. Interactive system 100 also enables specific users to invite other users to participate in specific games by sending invitations from interactive client 104. Interactive client 104 also supports sending and receiving audio, video, and text messages (e.g., chat) within the context of playing the game, provides leaderboards for the game, and also supports providing in-game rewards (e.g., game currency and items).
[0060] External resource system 226 provides interactive client 104 with an interface to communicate with remote servers (e.g., third-party server 112) to launch or access external resources (i.e., applications or applets). Each third-party server 112 hosts applications or smaller versions of applications (e.g., game applications, utility applications, payment applications, or ride-sharing applications) based on markup languages (e.g., HTML5). Interactive client 104 can launch web-based resources (e.g., applications) by accessing HTML5 files from the third-party server 112 associated with the web-based resource. The application hosted by third-party server 112 is programmed in JavaScript using a software development kit (SDK) provided by interactive server 124. The SDK includes APIs with functionality that can be called or activated by the web-based application. Interactive server 124 hosts a JavaScript library that provides access to a given external resource for specific user data of interactive client 104. HTML5 is an example of a technology used for programming games, but applications and resources programmed using other technologies can be used.
[0061] To integrate the SDK's functionality into the web-based resource, the third-party server 112 downloads the SDK from the interactive server 124, or the third-party server 112 otherwise receives the SDK. Once downloaded or received, the SDK is included as part of the application code of the web-based external resource. The code of the web-based resource can then call or activate certain functions of the SDK to integrate the features of the interactive client 104 into the web-based resource.
[0062] The SDK stored on the interactive server system 110 effectively bridges the gap between external resources (e.g., application 106 or applet) and the interactive client 104. This provides users with a seamless experience communicating with other users on the interactive client 104 while preserving the look and feel of the interactive client 104. To bridge communication between the external resources and the interactive client 104, the SDK facilitates communication between a third-party server 112 and the interactive client 104. A bridging script running on the user system 102 establishes two unidirectional communication channels between the external resources and the interactive client 104. Messages are sent asynchronously between the external resources and the interactive client 104 via these communication channels. Each SDK function activation is sent as a message and callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message with that callback identifier.
[0063] By using the SDK, not all information from the interactive client 104 is shared with the third-party server 112. The SDK limits which information is shared based on the needs of the external resource. Each third-party server 112 provides the interactive server 124 with an HTML5 file corresponding to the web-based external resource. The interactive server 124 can add a visual representation (e.g., box design or other graphics) of the web-based external resource to the interactive client 104. Once the user selects the visual representation or instructs the interactive client 104 to access the features of the web-based external resource through the graphical user interface (GUI), the interactive client 104 obtains the HTML5 file and instantiates the resource for accessing the features of the web-based external resource.
[0064] Interactive client 104 presents a GUI (e.g., a login page or title screen) for an external resource. During, before, or after presenting the login page or title screen, interactive client 104 determines whether the initiated external resource has previously been authorized to access user data of interactive client 104. In response to determining that the initiated external resource has previously been authorized to access user data of interactive client 104, interactive client 104 presents another GUI for the external resource, including its functionality and characteristics. In response to determining that the initiated external resource has not previously been authorized to access user data of interactive client 104, after displaying the login page or title screen of the external resource for a threshold time period (e.g., 3 seconds), interactive client 104 slides up a menu (e.g., animates the menu to appear from the bottom of the screen to the middle or other part of the screen) to authorize the external resource to access user data. This menu identifies the type of user data that the external resource will be authorized to use. In response to receiving a user selection of the accept option, interactive client 104 adds the external resource to the list of authorized external resources and allows the external resource to access user data from interactive client 104. External resources are authorized by the interactive client 104 to access user data under the OAuth 2 framework.
[0065] Interactive client 104 controls the type of user data shared with external resources based on the type of authorized external resource. For example, access to a first type of user data (e.g., 2D avatars of users with or without different avatar characteristics) is provided to external resources including full-scale applications (e.g., application 106). As another example, access to a second type of user data (e.g., payment information, 2D avatars of users, 3D avatars of users, and avatars with various avatar characteristics) is provided to external resources including smaller versions of applications (e.g., web-based versions of applications). Avatar characteristics include different ways of customizing the appearance and feel of an avatar, such as different poses, facial features, clothing, etc.
[0066] The advertising system 228 is operationally designed to enable third parties to purchase advertisements to be presented to end users via the interactive client 104, and also handles the delivery and presentation of these advertisements.
[0067] Artificial intelligence and machine learning system 230 provides various services to different subsystems within interactive system 100. For example, artificial intelligence and machine learning system 230 operates in conjunction with image processing system 202 and camera device system 204 to analyze images and extract information such as objects, text, or faces. This information can then be used by image processing system 202 to enhance, filter, or manipulate the image. Artificial intelligence and machine learning system 230 can be used by enhancement system 206 to generate enhanced content, XR experiences, and AR experiences, such as adding virtual objects or animations to real-world images.
[0068] Communication system 208 and messaging system 210 can use artificial intelligence and machine learning system 230 to analyze communication patterns and provide insights into how users interact with each other, as well as intelligent message classification and tagging, such as classifying messages based on sentiment or topic. Artificial intelligence and machine learning system 230 can also provide chatbot functionality for message interactions 120 between user systems 102 and between user system 102 and interaction server system 110. Artificial intelligence and machine learning system 230 can also work with audio communication system 216 to provide speech recognition and natural language processing capabilities, allowing users to interact with interaction system 100 using voice commands. Artificial intelligence and machine learning system 230 can also provide personalized AI agent system 232 functionality for message interactions 120 between user systems 102 and between user system 102 and interaction server system 110.
[0069] In some cases, the AI and machine learning system 230 may implement one or more machine learning models that process a user profile to determine the user's sharing patterns or sharing criteria. The AI and machine learning system 230 uses sharing patterns to process a set of previously captured content items (e.g., memories) to select a subgroup of content items that the user is most likely to be interested in sharing. The AI and machine learning system 230 then automatically analyzes each content item in the subgroup to generate one or more enhancements or modifications unique to each content item. Modifications and / or enhancements may be selected based on obtained user preferences and / or based on enhancements and modifications that the user has already made to other content items previously shared with other users. In some cases, modifications are based on suggested modifications generated by one or more LLMs. After modifying the content items, the AI and machine learning system 230 may (automatically or in response to a specific request) generate a feed of shareable content items to be presented to the user. Each content item in the shareable content item feed may be presented individually and sequentially in full-screen or partial-screen mode, either automatically or semi-automatically. Once a specific content item of interest is found in the feed of shareable content items presented to the user, input can be received requesting to share that specific content item of interest with one or more other users.
[0070] In some cases, content items included in a content item feed may include indicators that are persistently or temporarily overlaid on top of the content item. These indicators may indicate to the user or consumer that the content item being viewed or consumed is at least partially generated automatically by the artificial intelligence and machine learning system 230. The indicators may be specific graphic elements or text.
[0071] In a broad sense, machine learning can involve using computer algorithms to automatically learn patterns and relationships in data, potentially without requiring explicit programming after the algorithms have been trained. Examples of machine learning algorithms can be categorized into three main types: supervised learning, unsupervised learning, and reinforcement learning.
[0072] Supervised learning involves training a model using labeled data to predict outputs for new, unseen inputs. Examples of supervised learning algorithms include linear regression, decision trees, and neural networks.
[0073] Unsupervised learning involves training a model on unlabeled data to find hidden patterns and relationships within the data. Examples of unsupervised learning algorithms include clustering, principal component analysis, and generative models such as autoencoders.
[0074] Reinforcement learning involves training a model to make decisions in dynamic environments by receiving feedback in the form of rewards or penalties. Examples of reinforcement learning algorithms include Q-learning and policy gradient methods.
[0075] Examples of specific machine learning algorithms that can be deployed include logistic regression, a type of supervised learning algorithm for binary classification tasks. Logistic regression models the probability of a binary response variable based on one or more predictor variables. Another example type of machine learning algorithm is Naive Bayes, another supervised learning algorithm for classification tasks. Naive Bayes is based on Bayes' theorem and assumes that the predictor variables are independent of each other. Random forests are another type of supervised learning algorithm used for classification, regression, and other tasks. Random forests build an ensemble of decision trees and combine their outputs to make predictions. Further examples include neural networks, which consist of interconnected layers of nodes (or neurons) that process information and make predictions based on input data. Matrix factorization is another type of machine learning algorithm used for recommendation systems and other tasks. Matrix factorization decomposes a matrix into two or more matrices to reveal hidden patterns or relationships in the data. Support Vector Machines (SVMs) are a type of supervised learning algorithm used for classification, regression, and other tasks. SVMs find hyperplanes that separate different classes in the data. Other types of machine learning algorithms include decision trees, k-nearest neighbors, clustering algorithms, and deep learning algorithms such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Transformer models. The choice of algorithm depends on the nature of the data, the complexity of the problem, and the performance requirements of the application.
[0076] While this paper discusses several specific examples of machine learning algorithms, the principles discussed herein can also be applied to other machine learning algorithms. Deep learning algorithms such as convolutional neural networks, recurrent neural networks, and Transformers, as well as more traditional machine learning algorithms such as decision trees, random forests, and gradient boosting, can be used in a variety of machine learning applications. Generating trained machine learning programs, such as artificial intelligence and machine learning systems 230, can include various types of stages that form part of a machine learning pipeline, such as the following stages:
[0077] Data collection and preprocessing: This can include acquiring and cleaning data to ensure it is suitable for use in machine learning models. Data can be collected from user content creation and labeled using machine learning algorithms trained to label data. Data can also be generated by applying machine learning algorithms to identify or generate similar data. This can also include removing duplicates, handling missing values, and transforming data into a suitable format.
[0078] Feature engineering: This can include selecting and transforming training data to create features that are useful for predicting the target variable. Feature engineering can include (1) receiving features (e.g., as structured or labeled data in supervised learning) and / or (2) identifying features in the training data (e.g., unstructured or unlabeled data for unsupervised learning).
[0079] Model selection and training: This can include specifying a particular problem or desired response from the input data, selecting an appropriate machine learning algorithm, and training it on preprocessed data. It can also involve splitting the data into training and test sets, using cross-validation to evaluate the model, and tuning hyperparameters to improve performance. Model selection can be based on factors such as data type, problem complexity, computational resources, or desired performance.
[0080] Model evaluation: This can include evaluating the performance of a trained model (e.g., a trained machine learning program) on a separate test dataset. This can help determine whether the model is overfitting or underfitting and whether it is suitable for deployment.
[0081] Prediction: This involves using a trained model (e.g., a trained machine learning program) to generate predictions for new, unseen data.
[0082] Validation, refinement, or retraining: This can include updating the model based on feedback generated from the prediction phase, such as new data or user feedback.
[0083] Deployment: This can include integrating the trained model (e.g., a trained machine learning program) into a larger system or application, such as a web service, mobile app, or Internet of Things (IoT) device. This may involve installing the API, building the UI, and ensuring the model is scalable and can handle large amounts of data.
[0084] Prior to the training phase, feature engineering is used to identify features. This can include identifying informative, distinctive, and independent features that will enable the trained machine learning program to operate effectively in pattern recognition, classification, and regression. In some examples, the training data includes labeled data known to the pre-identified features and one or more outcomes.
[0085] Each feature can be a variable or attribute, such as a measurable characteristic of a process, item, system, or phenomenon represented by a dataset (e.g., training data). Features can also be of different types, such as numerical features, strings, vectors, matrices, codes, and graphs, and can include features from... Figure 5 The data obtained by the multimodal memory 508 shown herein. Conceptual features may include abstract relationships or patterns in the data, such as determining the topic of a document or a discussion in a chat window between users. Content features include determining the context based on input information, such as determining the user's context based on user interaction or surrounding environmental factors. Contextual features may include: text features, such as the frequency or preference of words or phrases; image features, such as pixel, texture, or pattern recognition; audio classification, such as spectrograms, etc. Attribute features include intrinsic attributes (directly observable) or extrinsic features (derived), such as identifying the square footage, location, or age of real estate identified in a camera feed. User data features include data relating to a specific individual or, for example, a group of individuals in a geographic location or sharing demographic characteristics. User data may include demographic data (e.g., age, gender, location, or occupation), user behavior (e.g., browsing history, purchase history, conversion rates, click-through rates, or engagement metrics), or user preferences (e.g., preferences for certain video, text, or digital content items). Historical data includes past events or trends that can help identify patterns or relationships over time.
[0086] During the training phase, the machine learning pipeline uses training data to identify correlations between features that influence predictions or prediction / inference data. The trained machine learning program is trained during the training phase, utilizing the training data and the identified features. The machine learning program evaluates the values of features as they correlate with the training data. The result of the training is the trained machine learning program (e.g., a trained or learned model).
[0087] Furthermore, the training phase can involve machine learning, where the training data is structured (e.g., labeled during preprocessing), and the trained machine learning program implements a relatively simple neural network capable of performing operations such as classification and clustering. In other examples, the training phase can involve deep learning, where the training data is unstructured, and the trained machine learning program implements a deep neural network capable of performing both feature extraction and classification / clustering operations. Neural networks comprise hierarchical (e.g., layered) organization of neurons, where each layer includes multiple neurons or nodes. Neurons in the input layer receive input data, while neurons in the output layer produce the network's final output. Between the input and output layers, there can be one or more hidden layers, each containing multiple neurons.
[0088] Each neuron in a neural network operationally computes a small function, such as an activation function, which takes as input a weighted sum of the outputs of neurons in the previous layer and a bias term. The output of this function is then passed as input to neurons in the next layer. If the output of the activation function exceeds a certain threshold, the output is passed from that neuron (e.g., the sending neuron) to connected neurons (e.g., the receiving neuron) in the next layer. Connections between neurons have associated weights that define the effect of an input from the sending neuron on the receiving neuron. During the training phase, these weights are adjusted by a learning algorithm to optimize the network's performance. Different types of neural networks can use different activation functions and learning algorithms, which can affect their performance on different tasks. In summary, the hierarchical organization of neurons and the use of activation functions and weights enable neural networks to model complex relationships between inputs and outputs and generalize to new inputs not seen during training.
[0089] In some examples, and by way of example only, a neural network can also be one or a combination of several different types of neural networks, such as a single-layer feedforward network, a multilayer perceptron (MLP), an artificial neural network (ANN), a recurrent neural network (RNN), a long short-term memory network (LSTM), a bidirectional neural network, a symmetric connection neural network, a deep belief network (DBN), a convolutional neural network (CNN), a generative adversarial network (GAN), an autoencoder neural network (AE), a restricted Boltzmann machine (RBM), a Hopfield network, a self-organizing map (SOM), a radial basis function network (RBFN), a spiking neural network (SNN), a liquid state machine (LSM), an echo state network (ESN), a neural Turing machine (NTM), or a Transformer network.
[0090] Neural networks can be iteratively trained by adjusting model parameters to minimize a specific loss function or maximize a certain objective. The system can continue training the neural network by adjusting parameters based on validation and refined output, or by retraining blocks, and by rerunning predictions on new or previously run training data. The system can employ optimization techniques for these adjustments, such as gradient descent, momentum algorithms, Nesterov accelerated gradient (NAG) algorithms, etc. Even after deploying the neural network, the system can continue to iteratively train it. As new data becomes available, the neural network can be continuously trained, for example, based on user-created or system-generated training data.
[0091] In some examples, a trained machine learning program, such as a personalized AI agent system 232, can be a generative AI model. Generative AI is a term that can refer to any type of artificial intelligence that can create new content from training data. For example, generative AI can produce text, images, videos, audio, code, or synthetic data that are similar to but not identical to the original data. In some cases, generative AI may include or implement a large language model (LLM). The generative AI and / or LLM receives a prompt (including instructions) and a set of data to be processed based on the prompt. The generative AI and / or LLM processes the data according to the instructions in the prompt and generates an output that includes modifications to the set of data based on the prior knowledge of the generative AI and / or LLM.
[0092] Some techniques that can be used in generative AI are:
[0093] Convolutional Neural Networks (CNNs): CNNs are commonly used for image recognition and computer vision tasks. They are designed to extract features from images by scanning the input image and highlighting important patterns using filters or kernels. CNNs can be used in applications such as object detection, face recognition, and autonomous driving.
[0094] Recurrent Neural Networks (RNNs): RNNs are designed to process sequential data such as speech, text, and time-series data. They have feedback loops that allow them to capture temporal dependencies and remember past inputs. RNNs can be used in applications such as speech recognition, machine translation, and sentiment analysis.
[0095] Generative Adversarial Networks (GANs): These are models consisting of two neural networks: a generator and a discriminator. The generator attempts to create realistic content that can fool the discriminator, which tries to distinguish between real and fake content. The two networks compete with each other and improve over time. GANs can be used in applications such as image synthesis, video prediction, and style transfer.
[0096] Variational autoencoders (VAEs): These are models that encode input data into a latent space (compressed representation) and then decode it back into output data. The latent space can be manipulated to generate new variations in the output data. They can use self-attention mechanisms to process the input data, allowing them to handle long sequences of text and capture complex dependencies.
[0097] Transformer models: These are models that use attention mechanisms to learn the relationships between different parts of input data (such as words or pixels) and generate output data based on these relationships. Transformer models can handle sequential data such as text or speech, as well as non-sequential data such as images or code.
[0098] In generative AI examples, the output prediction / inference data includes trend assessment and prediction, translation, summarization, image or video recognition and classification, natural language processing, facial recognition, user sentiment assessment, ad targeting and optimization, speech recognition or media content generation, recommendation and personalization.
[0099] Personalized AI agent system 232 analyzes user data and behavior to understand users' preferences and interests, providing personalized features to users of interactive client 104. By leveraging ML algorithms and data analysis, personalized AI agent system 232 can learn and adapt to user inferences, and then generatively suggest user-relevant, user-specific, and user-customized content. For example, personalized AI agent system 232 can use the learned user preferences and inferences to identify which content items in a set of content items the user might be most interested in sharing. Personalized AI agent system 232 can also select and / or generate enhanced content to augment the identified content items, creating unique experiences and shareable content items. Personalized AI agent system 232 can analyze data from multiple sources such as various user systems 102, messages, profile information, external data sources, image data captured in real-time by camera devices of user system 102, and / or any combination thereof to generate content items (e.g., codes and / or prompts) in real-time, provide such content items to the user, and select and modify content items to be shared based on this.
[0100] Personalized AI agent system 232 tracks user activities, such as posts (content items) that the user likes, shares, or comments on, topics the user follows, people the user contacts, and time the user spends on the platform. Tracking performed by personalized AI agent system 232 is only enabled when the user chooses to join an experience receiving real-time generated content item modifications. Personalized AI agent system 232 can present the user with a complete list of all activities and information that will be tracked and used to generate real-time content item modifications. Personalized AI agent system 232 only begins collecting such data and using it to provide and generate real-time and immediate content item modifications for presentation to the user after receiving confirmation from the user that they have approved the tracking of such activities and information.
[0101] Personalized AI agent system 232 can retrieve data from multiple data sources, such as activity on a user's mobile phone, AR / VR device, smartwatch, laptop, or other user devices. Based on this information, personalized AI agent system 232 can identify patterns and predict user interests to generate multimodal memories tailored to specific users. Personalized AI agent system 232 analyzes user profile information, such as their age, gender, location, exchanged messages, and / or interactions performed on user system 102, to provide personalized features. In some examples, personalized AI agent system 232 suggests nearby events and groups, or recommends job opportunities that match the user's qualifications. Personalized AI agent system 232 can generate real-time AR experiences and / or messaging content relevant to the current situation and / or the real-world environment perceived by the user.
[0102] Furthermore, the personalized AI agent system 232 analyzes user-created content and suggests the optimal time to post, the best tags to use, and the types of content to garner the most engagement. By doing so, the personalized AI agent system 232 helps users increase their visibility and reach a wider audience. In this way, the personalized AI agent system 232 can evaluate data from different devices to deliver personalized features across a variety of devices. The personalized AI agent system 232 automatically delivers such personalized features in real time based on multimodal memories associated with the user and in communication channels containing specific content preferred by the user. Analyzing user data and behavior to understand their preferences and interests and then suggesting and generating content items relevant to the user not only enhances the user experience but also increases engagement and retention on the platform.
[0103] In some examples, the personalized AI agent system 232 selects individual content items from a plurality of previously captured content items that match one or more criteria corresponding to shareable content, and generates a prompt that includes the individual content item and a request for multiple suggested modifications to the individual content item. The personalized AI agent system 232 processes the prompt, for example, via LLM to generate multiple suggested modifications to the individual content item, and generates a modified individual content item corresponding to the individual suggested modification among the multiple suggested modifications, for example, to be included in the content item feed.
[0104] Data Architecture
[0105] Figure 3 This is a schematic diagram illustrating a data structure 300 that can be stored in a database 304 of an interactive server system 110, according to certain examples. Although the contents of the database 304 are shown as including multiple tables, it will be appreciated that data can be stored in other types of data structures, such as an object-oriented database.
[0106] Database 304 includes message data stored in message table 306. For any given message, this message data includes at least message sender data, message receiver (or recipient) data, and a payload. See below for reference. Figure 4 Further details are provided regarding information that can be included in the message and within the message data stored in message table 306.
[0107] Entity table 308 stores entity data and (for example, links to entity diagram 310 and profile data 302). Entities for which records are maintained in entity table 308 can include individuals, company entities, organizations, objects, locations, events, etc. Regardless of entity type, any entity for which the interactive server system 110 stores data can be an identifiable entity. Each entity is assigned a unique identifier and an entity type identifier (not shown).
[0108] Entity graph 310 stores information about relationships and associations between entities. As an example only, such relationships can be social, professional (e.g., working in a common company or organization), interest-based, or activity-based. Some relationships between entities can be one-way, such as an individual user subscribing to digital content from a business or publishing user (e.g., a newspaper or other digital media export or brand). Other relationships can be two-way, such as the "friend" relationships between the various users of interactive system 100.
[0109] Certain permissions and relationships can be attached to each relationship, and also to each direction of the relationship. For example, a two-way relationship (e.g., a friend relationship between individual users) can include authorization for the posting of digital content items between the individual users, but certain restrictions or filters can be imposed on the posting of such digital content items (e.g., based on content characteristics, location data, or time of day data). Similarly, a subscription relationship between an individual user and a business user can impose varying degrees of restrictions on the posting of digital content from the business user to the individual user, and can significantly restrict or prevent the posting of digital content from the individual user to the business user. A specific user, as an example of an entity, can (e.g., through privacy settings) record certain restrictions in the records of that entity within entity table 308. Such privacy settings can be applied to all types of relationships within the context of interaction system 100, or selectively applied to certain types of relationships.
[0110] Profile data 302 stores various types of profile data about a specific entity. Based on privacy settings specified by the specific entity, profile data 302 can be selectively used and presented to other users of the interaction system 100. In the case of an individual, profile data 302 includes, for example, a username, phone number, address, settings (e.g., notification and privacy settings), and an avatar representation (or a set of such avatar representations) selected by the user. The specific user can then selectively include one or more of these avatar representations within the content of messages transmitted via the interaction system 100 and on a map interface displayed to other users by the interaction client 104. The set of avatar representations may include “status avatars,” which present a graphical representation of a status or activity that the user can choose to transmit at a specific time.
[0111] In the case that the entity is a group, in addition to the group name, members and various settings for the relevant group (e.g., notifications), the profile data 302 for the group may similarly include one or more avatars associated with the group.
[0112] Database 304 also stores enhancement data, such as overlays or filters, in enhancement table 312. Enhancement data is associated with and applied to videos (video data is stored in video table 314) and images (image data is stored in image table 316).
[0113] In some examples, filters are displayed as overlays on images or videos during presentation to the recipient user. Filters can be of various types, including user-selected filters from a set of filters presented to the sending user by the interactive client 104 while the sending user is composing a message. Other types of filters include geolocation filters (also known as geographic filters), which can be presented to the sending user based on geographic location. For example, geolocation filters specific to nearby or particular locations can be presented by the interactive client 104 within the user interface based on geolocation information determined by the Global Positioning System (GPS) unit of the user system 102.
[0114] Another type of filter is a data filter, which can be selectively presented to the sending user by the interactive client 104 based on other input or information collected by the user system 102 during the message creation process. Examples of data filters include the current temperature at a specific location, the current speed at which the sending user is traveling, the battery life of the user system 102, or the current time. Each filter can be associated with various metadata describing the filter. Such metadata may include one or more keywords describing the filter.
[0115] Other augmented data that may be stored within image table 316 includes (for example, AR content items corresponding to an application "lens" or XR experience). XR content items (e.g., XR objects) can be real-time special effects and sounds that can be added to images or videos. Any discussion of XR content and / or XR experiences and XR applications can be similarly applied to AR and VR content, experiences, and / or applications.
[0116] Collection table 318 stores data about collections of messages and associated image, video, or audio data, compiled into collections (e.g., stories or galleries). The creation of a specific collection can be initiated by a specific user (e.g., each user whose records are maintained in entity table 308). A user can create a "personal story" in the form of a collection of content that has already been created and sent / broadcast by that user. For this purpose, the user interface of interactive client 104 may include user-selectable icons that allow the sending user to add specific content to his or her personal story.
[0117] The collection can also constitute a "live story," which is a collection of content from multiple users created manually, automatically, or using a combination of manual and automatic technologies. For example, a "live story" can constitute a curated stream of user-submitted content from different locations and events. Users whose client devices have location services enabled and who are at a co-located event at a specific time can be presented with the option to contribute content to a specific live story, for example, via the user interface of interactive client 104. Live stories can be identified to a user by interactive client 104 based on their location. The end result is a "live story" told from a collective perspective.
[0118] Another type of content collection is called a "location story," which allows users of user system 102 located in a specific geographic location (e.g., on a college or university campus) to contribute to a specific collection. In some examples, contributions to a location story may employ secondary authentication to verify that the end user belongs to a specific organization or other entity (e.g., is a student on a university campus).
[0119] As mentioned above, video table 314 stores video data, which in some examples is associated with messages for which records are maintained within message table 306. Similarly, image table 316 stores image data associated with messages whose message data is stored in entity table 308. Entity table 308 can associate various enhancements from enhancement table 312 with various images and videos stored in image table 316 and video table 314.
[0120] Database 304 also includes trained machine learning techniques 307, which are stored in the shareable content item feeding system 590. Figure 5 ) and / or personalized AI agent systems 232 ( Figure 2 The parameters of one or more machine learning models that have been trained during the training period. For example, trained machine learning technique 307 stores the training parameters of one or more artificial neural network machine learning models or techniques.
[0121] Data communication architecture
[0122] Figure 4This is a schematic diagram illustrating the structure of message 400 according to some examples, generated by interactive client 104 for transmission to another interactive client 104 via interactive server 124. The content of a particular message 400 is used to populate message table 306 within database 304 accessible by interactive server 124. Similarly, the content of message 400 is stored in memory as "in-transit" or "in-flight" data for user system 102 or interactive server 124. Message 400 is shown to include the following example components:
[0123] Message Identifier 402: A unique identifier that identifies message 400.
[0124] Message text payload 404: The text to be generated by the user via the user interface of user system 102 and included in message 400.
[0125] Message image payload 406: Image data captured by the camera device component of user system 102 or retrieved from the memory component of user system 102 and included in message 400. The image data for the sent or received message 400 can be stored in image table 316.
[0126] Message video payload 408: Video data captured by the camera device component or retrieved from the memory component of the user system 102 and included in message 400. The video data for the sent or received message 400 can be stored in image table 316.
[0127] Message audio payload 410: Audio data captured by the microphone or retrieved from the memory component of the user system 102 and included in message 400.
[0128] Message enhancement data 412: This represents enhancement data (e.g., filters, labels, or other annotations or enhancements) to be applied to the message image payload 406, message video payload 408, or message audio payload 410 of message 400. Enhancement data for the sent or received message 400 can be stored in enhancement table 312.
[0129] Message duration parameter 414: A parameter value, in seconds, indicating the amount of time that the content of the message (e.g., message image payload 406, message video payload 408, message audio payload 410) will be presented to the user via the interactive client 104 or made accessible to the user.
[0130] Message geolocation parameter 416: Geolocation data (e.g., latitude and longitude coordinates) associated with the message's content payload. Multiple message geolocation parameter 416 values may be included in the payload, each of which is associated with a content item included in the content (e.g., a specific image within the message image payload 406 or a specific video within the message video payload 408).
[0131] Message Story Identifier 418: An identifier value that identifies one or more sets of content (e.g., “story” identified in set table 318) associated with a specific content item in the message image payload 406 of message 400. For example, the identifier value can be used to associate multiple images within the message image payload 406 with multiple sets of content, respectively.
[0132] Message Tag 420: Each message 400 can be labeled with multiple tags, each of which indicates the subject of the content included in the message payload. For example, in the case where a specific image depicts an animal (e.g., a lion) is included in the message image payload 406, a tag value can be included within the message tag 420 indicating the relevant animal. Tag values can be manually generated based on user input, or can be automatically generated using, for example, image recognition.
[0133] Message sender identifier 422: An identifier (e.g., message sending system identifier, email address, or device identifier) indicating the user of the user system 102 on which message 400 is generated and from which message 400 is sent.
[0134] Message receiver identifier 424: An identifier (e.g., message sending and receiving system identifier, email address, or device identifier) indicating the user of the user system 102 to which message 400 is addressed.
[0135] The content (e.g., values) of each component of message 400 can be pointers to locations in tables where content data values are stored. For example, image values in message image payload 406 can be pointers to locations (or their addresses) within image table 316. Similarly, values in message video payload 408 can point to data stored in image table 316, values in message enhancement data 412 can point to data stored in enhancement table 312, values in message story identifier 418 can point to data stored in set table 318, and values in message sender identifier 422 and message receiver identifier 424 can point to user records stored in entity table 308.
[0136] Personal AI Agent System
[0137] Figure 5 An example architecture 500 for applying a personal AI agent 502 to identify relevant features for user personalization is shown. The example architecture 500 may include a personal AI agent 502, a user database 504, a tool component 512, a shareable content item feed system 590, and a UI component 520. The personal AI agent 502 communicates with the user database 504, UI component 520, and tool component 512 to selectively and intelligently generate modified content items for user sharing or to be included in one or more machine learning models in the shareable content item feed.
[0138] User database 504 includes user-defined database 506, multimodal memory 508, and user-specific model 510. In some cases, personal AI agent 502 collects data from various sources and generates a multimodal memory 508 specific to a particular user. The personal AI agent 502 then provides personalized features to the user based on the identity model captured in the multimodal memory 508.
[0139] Multimodal memory 508 stores user-related information. Any data collected and stored in multimodal memory 508 is collected and stored with the explicit permission of the user on an opt-in basis. Non-limiting examples of multimodal memory types include:
[0140] Demographic data: such as age, gender, location, income, education, and occupation.
[0141] Behavioral data: Information about an individual's behavior and interactions with websites, apps, VR devices, or other digital touchpoints. This data can include website visits, clicks, downloads, purchases, interactions between users, and content items shared by users with other users.
[0142] Psychological data: Information about an individual's personality.
[0143] Contextual data: Information about the time, location, and device used by an individual when interacting with digital touchpoints. This data can help the personal AI agent 502 understand the context of the interaction and personalize the experience accordingly.
[0144] Purchase history data: Information about an individual's past purchases, such as the products bought, the frequency of purchases, and the amount spent. This data can be used by a personal AI agent (502) to create personalized recommendations and offers. The data may also include ad interaction data, which includes user interactions with ads such as clicks, impressions, and conversions.
[0145] Interest data: Information about an individual's hobbies, interests, and passions, which can be collected through online posts and activities.
[0146] Communication data: Information about how individuals prefer to be contacted and their communication history with other users. This data can be used to personalize communication channels and messaging (e.g., SMS messages on a phone or pop-up messages on an AR device).
[0147] Data from augmentation devices (e.g., AR / VR devices): information from camera feeds used by the user to capture images or videos of their surroundings; selection of digital content items (e.g., augmentations or overlays used on camera feeds); biometric data such as heart rate, body temperature, facial expressions, etc.; the user's preferred interaction method on the AR device; the type / duration of the AR interaction; eye tracking, eye focus, and the direction and duration of the user's gaze; body movements such as head or body movements, gestures, or interactions with the virtual environment; hand and finger movements using controllers or tracking devices; audio data from the microphone; user emotional responses such as heart rate, skin conductance, or facial expressions; user cognitive performance such as attention, memory, or problem-solving abilities; and so on. Examples of augmentation types applied to content items shared by the user with other users include the following.
[0148] Contact and Connection Data: Data about a user's contacts and connections, including their friends, followers, and groups.
[0149] Device data: Information about the user's device, including the device, operating system, and browser type. This information can be used to optimize platform performance and provide a seamless experience across different devices.
[0150] Personal AI agent 502 links different aspects of a user profile to multimodal memory 508. Multimodal memory 508 stores various aspects of the user profile (e.g., user preferences, lifestyle, interests, friends), and this contextual information can later be used to create tailored content. The value of such tailored content increases over time and across other factors as personal AI agent 502 collects more data related to the user. Over time, with an increasing number of digital touchpoints from the user, personal AI agent 502 gains a deeper understanding of the user and can use historical context to find relevance in any current user activity.
[0151] Multimodal memory 508 includes entity-based memory, which refers to memory focused on specific things such as objects, people, places, events, and experiences. This aspect of multimodal memory 508 stores information about the attributes, characteristics, and traits of these specific objects or people, such as remembering people's names, their appearance, their occupations, or their interests. Multimodal memory 508 also includes knowledge graph memory, which refers to memory focused on how things are related to or connected to each other. This aspect of multimodal memory 508 organizes information in a structured way, where entities are linked together in a weighted manner based on their relationships and attributes. Knowledge graph memory helps the personal AI agent 502 understand the context and meaning of information by showing how different entities and concepts are related. For example, a knowledge graph can store associations indicating what objects a user likes (e.g., animals) and what objects a user dislikes (e.g., food). A knowledge graph can store associations indicating what words / phrases / enhancements a user applies to images containing or depicting items the user likes and what words / phrases / enhancements a user applies to images containing or depicting items the user dislikes.
[0152] Each entity in the multimodal memory 508 can be linked to every other entity. These links between entities can be weighted in different ways to represent how closely related the entities are to each other. For example, an entity representing a dog's name can be linked to an entity representing a user and an entity representing an animal, while another entity representing a human with the same name can be linked to an entity representing a user and another entity representing a user's contact. The entity representing a dog's name can be linked to the entity representing an animal with a greater weight than the link between the entity representing a human with the same name and the entity representing an animal. Similarly, the entity representing a human with the same name as a dog can be linked to the entity representing a contact with a greater weight than the link between the entity representing a dog's name and the entity representing a human. In this way, various information about a given user, individual, or organization can be collected and correlated to establish links between the information.
[0153] In some cases, entity-based memories focus on specific things, while knowledge graph memories focus on how things are related. Entity-based memories store information about individual entities, while knowledge graph memories organize information based on the relationships and connections between entities. Both types of memories can be used to learn and understand the current context associated with one or more users.
[0154] Personal AI agent 502 utilizes embeddings from multimodal memory 508, which refers to techniques in machine learning used to represent and store data from multiple modalities (e.g., images, text, and audio) in a common vector space. The purpose of embeddings is to capture the semantic meaning and relationships between different modalities, enabling more efficient and accurate processing of multimodal data. In some examples, personal AI agent 502 feeds knowledge graphs to external or internal processes to generate latent embeddings. Personal AI agent 502 stores these latent embeddings in multimodal memory 508 for use in generating on-demand content recommendations and analysis. In some cases, content recommendations are provided to user system 102 without the user issuing a specific request for the recommendations. For example, user system 102 may be used to capture or access images, and personal AI agent 502 may detect the image that user system 102 is currently viewing. Personal AI agent 502 leverages user database 504 to generate specific AR experiences, and these generated AR experiences can be automatically applied to the image that user system 102 is currently accessing or viewing. In some cases, the personal AI agent 502 uses tool component 512 (discussed below) to generate unique AR experiences to deliver and activate AR experiences on the user system 102.
[0155] The personal AI agent 502 uses these embeddings for cross-modal comparison and analysis. In some examples, the embedding of an image is compared with the embedding of its corresponding text description to identify the semantic relationship between them. In the context of the multimodal memory 508, embeddings are used to store and retrieve information from different modalities in a more efficient and effective manner. In some examples, if the user has already stored entities including images, text descriptions, and audio recordings (e.g., memory objects), embeddings can be used to represent each of these modalities in a common vector space. This enables efficient retrieval and integration of information from multiple modalities when accessing memory objects.
[0156] The personal AI agent 502 uses one or more techniques, such as neural networks, to generate embeddings. These embeddings can be fine-tuned and optimized for specific applications and tasks. The personal AI agent creates embeddings for multimodal memories, which include information about individuals and their relationships with other individuals, entities, and devices, and are generated using various data sources as described herein by identifying patterns and connections between entities. Therefore, the personal AI agent 502 can apply multimodal memories to users in a variety of different ways to provide personalized features.
[0157] Past users may have already selected certain customization options, such as content enhancements, graphics, or features within the app or AR device. The interaction system 100 stores such user customizations in a user customization database 506. The personal AI agent 502 uses such customizations when delivering relevant content, for example, by recognizing user customization preferences. In some examples, the user may have already used some type of enhancement (e.g., adding pizza to a camera feed) or a sticker the user sent to a friend.
[0158] User-defined database 506 includes customizations made by users within interactive system 100. User-defined database 506 stores profile customizations (e.g., profile pictures, cover photos, and introductions), news feed preferences (e.g., prioritizing certain friends, pages, and groups), and privacy settings (e.g., who can see posts, profile information, and events). User-defined database 506 stores personalized avatar and tag selections, customized content enhancements and how these are applied to camera feeds, sound and music preferences used for content creation, channel and creator subscription preferences, and other types of user customization based on the user's viewing history and engagement.
[0159] User-specific model 510 includes generative models for generating graphics, text, and images for users. For example, user-specific model 510 can generate tags that include photos, graphics, or animations. User-specific model 510 generates avatars representing users on the interactive system platform. Users can customize their avatars to reflect their appearance, personality, and interests. User-specific model 510 generates filters and content enhancements, such as XR effects that can be applied in real time to enhance photos and videos. User-specific model 510 generates memes to share humorous images, videos, and descriptive text that convey specific cultural ideas or trends. User-specific model 510 generates hashtags to categorize and organize user posts based on specific topics or themes, enabling users to connect with other users sharing similar interests or participate in trending conversations. User-specific model 510 generates short-lived photos and videos that their followers can view for a limited time, including filters, tags, and text overlays to make them more engaging.
[0160] Tool component 512 includes one or more neural network engines 514, one or more external data sources 516, and one or more feature APIs 518. Neural network engines 514 include one or more generative machine learning models. These models can be trained to generate a variety of different content. For example, a generative machine learning model is trained to receive cues as input (which may include any combination of text, images, audio, and / or video) and generate output in response to the cues. In some cases, the generative machine learning model generates artificial images / videos, code snippets, and / or text in response to cues. In other cases, the generative machine learning model generates content enhancements, such as overlaying, modifying, or enhancing filters fed from a real-world camera device using digital content items or previously captured content items. Tool component 512 may include filter tools, explanatory text tools, and / or image-to-image tools. The descriptive text tool (or add descriptive text tool) can be used to add descriptive text to individual content items, the image-to-image tool can be used to process individual content items through a generative model to generate new content items based on the description, and the filter tool can be used to select predefined filters that match one or more keywords to apply to individual content items.
[0161] In some cases, the personal AI agent 502 provides a cue as input to the neural network engine 514. This cue is generated by the personal AI agent 502 based on information collected from the UI component 520 (representing the current real-world environment) and the user database 504. That is, the personal AI agent 502 generates a cue that includes an image captured by the user system 102 and one or more vectors obtained from the multimodal memory 508. This cue can then be provided as input to the neural network engine 514. The neural network engine 514 then accesses additional data sources, such as external data source 516 and / or feature API 518, to generate content that matches the cue input. In some examples, the personal AI agent 502 collects contextual data of the conversation, hears audio from the user, and / or receives other input from the user or the user environment, and generates a request to the neural network engine 514.
[0162] External data source 516 may include various search engines, chatbots, email applications, calendar applications, messaging applications, social networking applications, news sources, live media sources, and / or any combination thereof. Personal AI agent 502 accesses external data from external data source 516 to generate multimodal memory 508 for the user and / or apply multimodal memory 508 to the user. In some examples, personal AI agent 502 collects user data from different media types and creates embeddings for the user's multimodal memory 508. Personal AI agent 502 may retrieve data from external data source 516 to apply to multimodal memory 508 and / or to generate input for a neural network engine. External data sources 516 may include any combination of scientific paper repositories, which include: content information about papers, titles, authors, abstracts, keywords, and citations; data from email, such as email addresses, contact lists, email content, attachments, and metadata (such as timestamps and IP addresses); search engine data, such as search queries, search history, location data, device information, and web activity; and / or communication data, such as messages, voice and video calls, roles in group messages, posts, comments, polls, user subscriptions, likes, followers, and hashtags.
[0163] Neural engine 514 can intelligently select one or more external data sources 516 to populate data in response to prompts received from personal AI agent 502. Feature API 518 can provide access to a variety of additional tools, which may be proprietary. Using a given API from Feature API 518, neural engine 514 can access additional machine learning tools to generate additional content. For example, one of the proprietary tools may include an AR experience generation tool. Neural engine 514 can access the tool's API to prompt the tool to generate an AR experience with specific features generated by neural engine 514 based on the prompts received from personal AI agent 502. Neural engine 514 can receive a specific AR experience and can further apply various effects and modifications to the AR experience to generate a response to personal AI agent 502, including an AR experience that matches the prompts. Personal AI agent 502 can then automatically enhance and / or modify the image currently being presented on user system 102 using the unique AR experience generated by neural engine 514.
[0164] Personal AI agent 502 retrieves personalized content for users and applies that content via various feature APIs 518. Personal AI agent 502 applies personalized and / or recommended content to interactive client features via feature APIs 518, including photos, videos, or podcasts captured and shared by the user with friends. Personal AI agent 502 recommends content applied to camera feeds, which can then be recorded and sent to individual users as filters, tags, text overlays, or content enhancements. Personal AI agent 502 applies personalized and / or recommended content to interactive client features via Feature API 518, said personalized and / or recommended content including: messages, photos, and videos for individual friends or groups; collections of photos or videos, such as collections that can be viewed by friends over a maximum specific time period (e.g., 24 hours); content from media partners, such as news articles, videos, and programs curated for users; map-related features, such as when and with whom location is shared, how the system shares it, and what information is displayed on the map UI; filters and / or XR experiences, including visual overlays applied to photos or videos, such as adding location-based information, temperature, time, and other graphics; and customized avatars that can be used in photos / videos, chat messages, and other features.
[0165] In many cases, the personal AI agent 502 can access multiple input devices to generate cues for the neural network engine 514. These input devices or combinations of input devices may include wearable devices 522, such as AR / VR devices 526, devices with monitors 524, such as user system 102, and / or smartwatches 528. The output generated by the personal AI agent 502 can be provided to any of the same or different combinations of AR / VR devices 526, user system 102, and / or smartwatches 528. The personal AI agent 502 identifies preferred communication methods using multimodal memory 508. In some examples, the personal AI agent 502 recognizes that the user prefers to be alerted via smartwatch 528, given instructions on AR / VR devices 526, and receive text messages on the user's mobile phone.
[0166] The process for training the neural network engine 514 can be the same as previously discussed. The neural network engine 514 can receive a set of training data, which includes various cues and real responses. These real responses can include identifiers of data sources that can be used to respond to cues and / or image or text content generated in response to cues. The neural network engine 514 can iterate over the training data until the loss function meets a stopping criterion, wherein the parameters of the neural network engine 514 are updated at each iteration. Similarly, a personal AI agent 502 can be trained based on training data to generate cues to be provided to the neural network engine 514. The training data can include various latent vectors, images, and information obtained from the user database 504, along with corresponding real cues. The personal AI agent 502 can iterate over the training data until the loss function meets a stopping criterion, wherein the parameters of the personal AI agent 502 are updated at each iteration. In this way, the personal AI agent 502 can generate real-time cues based on current information accessed or obtained by the UI component 520, to be provided to the tool component 512 for generating content to be provided back to the UI component 520.
[0167] The shareable content item feeding system 590 can continuously or periodically analyze images captured by one or more interactive clients 104, such as UI components 520. The system can apply one or more ML models to a set of previously captured content items to select those that match criteria indicating potential sharing interest. In response to selecting content items that match the criteria, the system can generate prompts to individually modify each of the selected content items using different enhancements selected based on preferences provided by a personal AI agent 502 and / or based on objects depicted in the corresponding images. In some cases, prompts are generated dynamically based on contextual cues gathered from the images and / or information obtained from multimodal memory 508.
[0168] The shareable content item feed system 590 accesses the XR application via interactive client 104. The shareable content item feed system 590 can provide prompts to tool component 512, either together with or separately from personal AI agent 502, to process the prompts using a generative ML model (e.g., neural network engine 514), thereby generating revisions or enhancements for each selected content item. The modified content items are then ranked, for example, based on the recentity of capture, and placed in the shareable content item feed.
[0169] The shareable content item feeding system 590 can be part of a personal AI agent 502 and perform functions similar to the personal AI agent 502. The shareable content item feeding system 590 can then automatically and / or in response to input received from the user to present the shareable content item feed. For example, it can receive a user request to share content. In response, a first content item from the shareable content item feed is presented in full screen. Gestures (e.g., an up swipe gesture) can be detected, and in response, the shareable content item feeding system 590 retrieves a second content item from the shareable content item feed and presents it in full screen. The shareable content item feeding system 590 monitors user interaction, such as dwell time, for each content item presented in the shareable content item feed. The shareable content item feeding system 590 can dynamically change the priority and re-rank the content items in the shareable content item feed and / or retrieve and generate new modified content to include in the shareable content item feed based on the monitored interactions. Once a user selects a content item of interest, the shareable content item feeding system 590 detects the input of an option to share the content item. In this case, the shareable content item feeding system 590 sends the content item from the shareable content item feed to one or more designated recipients.
[0170] The personalized AI agent system 232 and the shareable content item feed system 590 provide users with an opt-in option / an opt-out option for opting out of the use of their personal data. The personalized AI agent system 232 and the shareable content item feed system 590 can operate as opt-in systems by default. Opt-in and opt-out options are mechanisms used by the system to give users control over the use of their personal data. These options are particularly important in an era of data privacy regulations such as the General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA). Below are some examples of opt-in mechanisms:
[0171] 1. Opt-in: The opt-in method requires users to actively give their consent before their personal data can be collected, processed, or shared by an app or website. This method is considered more privacy-friendly because it ensures users are fully aware of data practices and intentionally choose to participate. Such opt-in options may appear as banners or pop-ups. These opt-in options appear when a user first visits a website or launches an app, requesting permission to collect and process personal data for specific purposes (e.g., targeted advertising, analytics, or personalization). Other opt-in options may be displayed as checkboxes or toggle switches, allowing users to individually enable or disable data collection for specific purposes. In some examples, opt-in options are presented as contextual prompts that users may encounter when accessing specific features or functions within an app or website that rely on data collection (e.g., location-based services).
[0172] 2. Opt-out: In the opt-out method, the system assumes by default that the user agrees to data collection and processing. However, the user is provided with the option to withdraw their consent at any time. The system applies the opt-out method in a limited number of environments.
[0173] The system provides a privacy policy or settings where users can access the application's or website's privacy policy, which includes information on how to opt out of data collection and processing, or preference options that allow users to manage their privacy settings and disable specific data collection and sharing practices. To enhance transparency and user control, the system described herein clearly communicates data collection and processing practices and provides easy-to-use opt-in and opt-out options.
[0174] Shareable content item feed system
[0175] In some examples, the shareable content item feed system 590 accesses a list of previously captured content items, such as images, pictures, videos, audio files, etc. The shareable content item feed system 590 can search only those previously captured content items associated with timestamps prior to the current time threshold period. For example, the shareable content item feed system 590 can access only those content items captured last week or in the previous two days. In some cases, the shareable content item feed system 590 can access all previously captured content items associated with a user and stored in user database 504 and / or user system 102.
[0176] The shareable content item feed system 590 uses one or more machine learning models, such as a personal AI agent 502, to process previously captured content items that have been retrieved or accessed. The shareable content item feed system 590 can determine which of the previously captured content items match or are associated with attributes corresponding to one or more shareable criteria. For example, the shareable content item feed system 590 can access criteria that define good candidates for content items to be shared with other users. The criteria can be general to the user group of interactive client 104 and / or specific to the users of interactive client 104.
[0177] For example, in some cases, shareability criteria may initially be defined as including a category or theme of content and certain depictions of items or objects in an image. In such cases, the shareable content item feed system 590 extracts features from each of the previously captured content items that have been retrieved or accessed, and determines whether the features correspond to the initially defined shareability criteria. Any content item that matches the initially defined shareability criteria may be included in a set of content items to be modified. This results in a set of content items selected from the previously captured content items to be included in the initial group of content items to be modified. The initial group of content items can then be processed by the shareable content item feed system 590, as discussed below, to apply one or more different modifications to each corresponding content item for generating a shareable content item feed.
[0178] In some cases, the shareable content item feed system 590 generates an embedding for each content item previously captured by the user system 102 (e.g., captured over a past time interval, such as the past week). The embedding may represent tags, visual cues, and / or descriptions for each content item. In some cases, these cues, tags, and descriptions are generated by processing each content item using prompts from the LLM requesting the generation of descriptive tags for each item. The shareable content item feed system 590 can then access a list of predefined descriptions. The shareable content item feed system 590 can compare the visual cues or tags and descriptions to the list of predefined descriptions. If a content item is associated with a cue or tag and description on the list of predefined descriptions, the shareable content item feed system 590 adds that content item to the list of shareable candidate content items. The shareable content item feed system 590 can also determine whether any content item in the list of candidate content items has been designated as private by the user. The shareable content item feed system 590 can remove such privately designated content items from the list of candidate content items.
[0179] The shareable content item feed system 590 can then determine whether the lighting conditions meet a certain darkness threshold based on visual cues and labels of candidate content items. If so, the shareable content item feed system 590 determines that the content item has poor lighting and removes such content items from the list of candidate content items. The filtered list of candidate content items can then be used to select individual content items for automatic modification by the shareable content item feed system 590.
[0180] In some examples, the personal AI agent 502 can process previously captured content items based on user preferences (e.g., based on the history of interactions associated with the user) to identify those content items the user is most likely to be interested in sharing. Specifically, after initially generating a feed of shareable content items, the personal AI agent 502 can build profiles representing the content items shared by the user of the interacting client 104 over time, such as over days, weeks, and / or months. For example, after a certain number of content items have been shared by the user, the personal AI agent 502 can process these content items to extract features representing the content items shared by the user. The personal AI agent 502 can then update the initially defined criteria to be more user-specific. The shareable content item feed system 590 can then use the updated criteria to select a new group or set of content items from the previously captured content items (which can now include new content items or other previously captured content items). The shareable content item feed system 590 can then add the new group or set of selected content items to an updated set of content items to be automatically modified. As discussed below, the updated set of content items can then be processed by the shareable content item feed system 590 to apply one or more different modifications to each corresponding content item in order to generate a new or updated shareable content item feed.
[0181] In some examples, the criteria used to select content items to include in a set of content items to be modified can be automatically obtained or determined by a personal AI agent 502. That is, the personal AI agent 502 can generate criteria using one or more generative machine learning models and prompts. The prompts can instruct the generative machine learning models to process data including the entire set of content items that a given user has previously captured and / or accessed, and to obtain the user's interests or preferences. The prompts can also instruct the generative machine learning models to obtain criteria based on data representing the most likely attributes of content items that the user would like to share with other users. This criterion can then be used by a shareable content item feeder system 590 to select a set of content items to modify from the previously captured content items.
[0182] After generating the set of content items to be modified, the shareable content item feed system 590 can rank the content items in the set. Ranking can be based on user preferences and / or the likelihood that a content item is something the user is interested in sharing. That is, the shareable content item feed system 590 can generate a score for each content item in the set. The score can represent the likelihood that a content item is something the user is interested in sharing. The content item with the highest score can be presented at the top of the list of content items in the set. In some cases, the set of content items is ranked based on the capture timestamp associated with each content item. In such cases, the more recently captured content item can be placed before other content items captured earlier. This ensures that the more recently captured content item is modified and presented to the user first in the shareable content item feed.
[0183] Shareable content item feed system 590 processes the set of content items and analyzes the set of content items with the help of personal AI agent 502 to apply one or more modifications to each content item. For example, shareable content item feed system 590 can retrieve the first content item in the list. Shareable content item feed system 590 can obtain one or more attributes of the first content item. One or more attributes may indicate what real-world or virtual objects are depicted in the first content item, the location where the first content item was captured, a timestamp indicating when the first content item was captured, and various other information associated with the first content item. Personal AI agent 502 can analyze the first content item and predict one or more modifications to be applied to the first content item based on user preferences (e.g., data stored in user database 504). In some cases, the personal AI agent 502 can receive prompts from the shareable content item feed system 590, which include instructions to analyze the first content item based on knowledge about the user to predict or identify one or more modifications to be applied to or associated with the first content item (e.g., descriptive text, augmented reality elements or experiences, filters, annotations, etc.).
[0184] A personal AI agent 502 provides one or more identified or predicted modifications to a shareable content item feed system 590. The shareable content item feed system 590 then modifies a first content item based on the one or more modifications and adds the modified first content item to the shareable content item feed. For example, the shareable content item feed system 590 overlays a first set of augmented reality elements onto the first content item to generate the modified first content item. This process is repeated for each other content item in that set.
[0185] For example, the shareable content item feed system 590 can access prompt 600, such as... Figures 6A to 6BAs shown in the illustration, Tip 600 may include visual cues or descriptive labels associated with the selected individual content item, as well as a request for multiple suggested modifications to be performed for that individual content item. Specifically, Tip 600 may include Instruction 610 to inform or request the LLM to create shareable modifications (edits) given the current date, time, location, and information about items that are good candidates for modification. Instruction 610 may specify the presence of various tools that can be used to perform the suggested modifications, such as filter tools, image-to-image editing tools, and / or descriptive text tools. In some cases, Instruction 610 may specify that image-to-image editing tools should always be selected or may be selected optionally. Instruction 610 may instruct the LLM to avoid generating content that does not meet safety standards (e.g., preventing the generation of suggested modifications that are unsafe for work (NSFW)).
[0186] In some cases, prompt 600 includes a descriptive section 620 that defines the operation of each available tool. For example, the descriptive section 620 may include a description of the Add Explanatory Text tool. The description may specify that, for the Add Explanatory Text tool, the LLM is instructed to provide plain text for adding to the content item, which is clever, funny, ironic, sarcastic, and / or resonates with someone within a specific age range (which may correspond to the age of the user of user system 102). The description may indicate that the explanatory text should be kept short and funny and should be written in English. Instructions for describing the Add Explanatory Text tool may instruct the LLM to add specified graphic elements before and after the text of the explanatory text. The specified graphic elements indicate to the user that the explanatory text is automatically generated.
[0187] As another example, description section 620 includes an image-to-image tool description. This description can instruct the LLM to generate suggested modifications that reimagine the content item in a different style, highlighting relevant details to be considered. Description section 620 also includes a filter tool description. This description can instruct the LLM to generate suggested modifications by adding colors or other adjustments to alter the look and feel of the input content item, and provides two or three keywords describing the suggested filter. Such filters change the atmosphere or visual style without altering the scene being depicted in the content item.
[0188] Tip 600 may include instruction 630 that the LLM uses any combination of available tools described in description section 620 to return a predetermined number of suggested modifications (e.g., three suggested modifications). Instruction 630 may instruct the LLM to construct each suggested modification according to a specified format 640. The specified format 640 may be JSON or HTML or other markup language format and may include fields for the name of the suggested modification, a field for a summary of the suggested modification, a field for a description of the suggested modification in a single word, a field for filter keywords (if a filter tool is selected), a field for an image-to-image modification description (if an image-to-image tool is selected), and a descriptive text field for descriptive text (if a descriptive text tool is selected).
[0189] For example, the shareable content item feed system 590 can retrieve a second content item from the list. The shareable content item feed system 590 can obtain one or more attributes of the second content item. One or more attributes may indicate what real-world or virtual objects are depicted in the second content item, the location where the second content item was captured, a timestamp indicating when the second content item was captured, and various other information associated with the second content item. A personal AI agent 502 can analyze the second content item and, based on user preferences (e.g., data stored in a user database 504), predict additional sets of modifications to be applied to the second content item. In some cases, the personal AI agent 502 may receive prompts from the shareable content item feed system 590 with the instruction to analyze the second content item based on knowledge about the user to predict or identify one or more modifications to be applied to or associated with the second content item (e.g., descriptive text, augmented reality elements or experiences, filters, annotations, etc.). The personal AI agent 502 provides the shareable content item feed system 590 with additional sets of the identified or predicted modifications. For example, the shareable content item feed system 590 overlays a second set of augmented reality elements onto a first content item and adds specific descriptive text to generate a modified second content item. Then, the shareable content item feed system 590 modifies the second content item based on one or more additional modified sets and adds the modified second content item to the shareable content item feed. In this way, each content item in the shareable content item feed can be modified with different augmentations.
[0190] In some examples, the shareable content item feed system 590 generates a notification to the user of the interactive client 104 instructing the shareable content item feed to include a new shareable content item. In response to receiving input from a selection notification, the shareable content item feed system 590 presents the shareable content item feed. That is, the shareable content item feed system 590 retrieves a first content item that is the first in the shareable content item feed and presents that content item to the user in full screen or a portion of the screen. The shareable content item feed system 590 presents an option to share the first content item with one or more specified recipients. The shareable content item feed system 590 receives a gesture (e.g., an up swipe gesture). In response, the shareable content item feed system 590 retrieves a second content item that is the second in the shareable content item feed and presents that second content item instead of the first content item. This process of navigating through the content items in the shareable content item feed can continue indefinitely or until no more content items remain in the shareable content item feed.
[0191] In some cases, the shareable content item feed system 590 receives input from a live or real-time camera feed that activates the user system 102. In response, the shareable content item feed system 590 presents the image being received from the live or real-time camera feed from the front or rear camera of the user system 102 in full screen. The shareable content item feed system 590 can detect gestures performed by the user. These gestures may include swiping up along the screen on which the live or real-time camera feed is being presented. In response, the shareable content item feed system 590 retrieves a first content item that is the first in the shareable content item feed and presents that content item to the user in full screen or a portion of the screen, instead of the live or real-time camera feed. The shareable content item feed system 590 presents an option to share the first content item with one or more designated recipients. The shareable content item feed system 590 receives another instance of a gesture (e.g., an up swipe gesture). In response, the shareable content item feed system 590 retrieves the second content item that is second in the shareable content item feed and presents that second content item instead of the first content item. Navigation through the content items in the shareable content item feed can continue indefinitely, or until no more content items remain in the shareable content item feed.
[0192] In some examples, the shareable content item feed system 590 receives input from a dialogue or chat interface to access the shareable content item feed. That is, the interactive client 104 can present a UI in which one or more messages are exchanged between multiple participants. The interactive client 104 can receive input from the user to share content with participants in the dialogue. In response, the interactive client 104 can retrieve the identities of multiple participants and provide the participants' identifiers to the shareable content item feed system 590. The shareable content item feed system 590 can generate prompts that include instructions for the personal AI agent 502 to process the set of previously captured content items and retrieve or select content items that match the preferences of the participants' identifiers in the dialogue and / or depict the faces associated with the identifiers. The prompts can also instruct the personal AI agent 502 to further retrieve or select content items based on those that the user is most interested in sharing. Once the selection of content items is retrieved, the shareable content item feed system 590 can instruct the personal AI agent 502 to modify the selection of content items and add the modified selection of content items to the shareable content item feed specific to the participants in the dialogue. The shareable content item feeding system 590 can display the first content item that is the first in the shareable content item feeding in full screen. The shareable content item feeding system 590 presents an option to share this first content item with the dialogue participant. The shareable content item feeding system 590 receives another instance of a gesture (e.g., an up swipe gesture). In response, the shareable content item feeding system 590 retrieves the second content item that is the second in the shareable content item feeding and displays the second content item instead of the first content item. This process of navigating through the content items in the shareable content item feeding can continue indefinitely, or until no more content items remain in the shareable content item feeding.
[0193] In some examples, the shareable content item feed system 590 periodically, continuously, and / or in response to detecting a specified number of new content items being added to a previously captured list of content items updates the content items included in the shareable content item feed. In such cases, the shareable content item feed system 590 reprocesses the updated list of previously captured content items to generate newly modified content items for inclusion in the shareable content item feed. That is, the shareable content item feed system 590 can modify one or more content items that have been newly added to the previously captured list and meet (match) the shareability criteria, and add the modified one or more content items as new items to the shareable content item feed. The shareable content item feed system 590 can place such newly modified content items at the top of the shareable content item feed so that they are presented to the user first when the shareable content item feed is accessed.
[0194] In some examples, while a shareable content item feed is being presented to a user, the shareable content item feed system 590 analyzes the user's interaction with the feed. The system can determine the dwell time for each content item in the presented feed (the amount of time a user spends viewing each item before navigating to a different item). The system can also determine that a particular item is associated with a dwell time exceeding a threshold or greater than the determined dwell time for other items. In response, the system can extract one or more features from the specific item. The system can then search the previously captured set of items (or feeds) for items with features similar to the extracted features. For example, the system might determine that a particular item depicts two faces of two people. The shareable content item feed system 590 can search among the previously captured content items in the group for additional content items that exclusively or inclusively depict the same two faces of two people. The shareable content item feed system 590 can then (in a manner similar to that previously discussed) modify the additional content items to enhance them using one or more modifications. The shareable content item feed system 590 can then place the enhanced / modified additional content items in the shareable content item feed, so that they are presented before all other content items that have not yet been presented to the user.
[0195] The shareable content item feed system 590 can process the prompt 600 and the selected content item, for example, via LLM. The shareable content item feed system 590 can generate output 700 based on the prompt 600. For example, output 700 may include multiple suggested modifications for the content item. Suggested modifications may include a first suggested modification 710, a second suggested modification 720, and a third suggested modification 730. Each suggested modification may include different combinations of available tools. For example, the second suggested modification 720 may include a field 722 for image-to-image tools, while the third suggested modification 730 does not include a field 722 for image-to-image tools. The third suggested modification 730 may include a filter tool field 732 with keywords that can be used to select filters to apply to the content item, while the second suggested modification 720 does not include a filter tool field. The first suggested modification 710, the second suggested modification 720, and the third suggested modification 730 may each include a descriptive text field 734, which includes a text string for suggested descriptive text generated by the LLM. In some cases, fewer than all suggested modifications are included in the description text field in the output 700.
[0196] The shareable content item feed system 590 can process output 700 to select a specific suggested modification to apply to the content item. For example, the shareable content item feed system 590 can randomly select one of the suggested modifications, such as a second suggested modification 720.
[0197] The shareable content item feed system 590 can process the selected suggested modification fields to automatically or semi-automatically modify the content item. For example, the shareable content item feed system 590 can determine that the selected suggested modification includes the description text field 734. In such a case, the shareable content item feed system 590 retrieves the text string specified in the description text field 734 and overlays the text string onto the content item. For example, as... Figure 8 As shown, the shareable content item feed system 590 modifies the content item 800 by overlaying the text of the description text 810 15% below the center of the content item. The shareable content item feed system 590 can attach a graphic indicator 812 at the beginning of the text (to indicate to the consumer that the description text is automatically generated by LLM) and the same graphic indicator 812 at the end of the text.
[0198] In some cases, the shareable content item feed system 590 may determine that the selected suggested modification includes a descriptive text field 732, which includes keywords describing the filter to be applied to the content item. In response, the shareable content item feed system 590 accesses a list of pre-defined filters available in the interactive server system 110. In some cases, the pre-defined filters may be provided by a third-party server 112. The shareable content item feed system 590 obtains metadata for the pre-defined filters and searches the metadata to identify a set of filters with metadata matching the keywords provided in the descriptive text field 734. The shareable content item feed system 590 may then rank this set of filters and select the filter associated with the highest ranking. The shareable content item feed system 590 automatically modifies the content item using the selected filter.
[0199] In some cases, the shareable content item feed system 590 may determine that the selected suggested modification includes field 722 for the image-to-image tool. In such cases, the shareable content item feed system 590 accesses the image-to-image cue template 701 (FIG. 7) to refine the description of the image modification provided by field 722 for the image-to-image tool. The template may include a set of instructions for the cue that instruct the LLM that it is being given a cue to modify or enhance an existing content item, and instruct the LLM to rewrite the cue to be more useful and safer by expanding the given concept by adding more helpful details of clothing, accessories, or environment. Specifically, the shareable content item feed system 590 may process template 701 with the description of the image modification to generate an improved or new description for modifying the image by generating a new image. The shareable content item feed system 590 processes the new cue through the LLM to generate a revised description of the image modification. In some cases, the new cue includes one or more negative revisions indicating modifications that are not permitted. For example, LLM is instructed to generate new suggestions that are not complete sentences, but rather a series of 5 to 15 comma-separated keywords in English.
[0200] After the LLM processes the content of field 722 for the image-to-image tool using image-to-image hint template 701, a new hint is generated. This new hint, along with the content item, is processed by shareable content item feed system 590 to generate a new content item that matches the description provided by the new hint generated by the LLM. For example, as... Figure 8 The set of images 801 shown, including the first image 811 and the second image 813, can be processed together with various prompts 830, 840, and 850 to render or generate new images 832, 842, and 852. One of these new images 832, 842, and 852 can be provided to the user in a content item feed, wherein an indicator superimposed on images 832, 842, and 852 informs the consumer that images 832, 842, and 852 are automatically generated. In some cases, if the selected suggested modification includes a descriptive text field, the text of the descriptive text is superimposed on the newly generated image. In some cases, one or more security checks can be performed by processing the automatically modified image through a security check model before making the automatically modified content item available to the consumer. The automatically modified content item is only made available if the modified content item passes the security check model.
[0201] For example, the first image 811 can be processed by a generative model using prompt 830 to generate a new image 832. The second image 813 can be processed by a generative model using prompt 840 to generate a new image 842. The first image 811 can be processed by a generative model using prompt 850 to generate a new image 852, and the new image 852 is different from image 832 due to the difference between the description of prompt 850 and the description of prompt 830.
[0202] Figure 9 This is a flowchart of a process or method 900 performed by a shareable content item feeding system 590, based on some examples. While a flowchart can describe operations as a sequential process, many operations can be performed in parallel or simultaneously. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed. A process can correspond to a method, procedure, etc. The steps of a method can be performed in whole or in part, in combination with some or all of the steps in other methods, and can be performed by any number of different systems or any part thereof, such as a processor in any system included in the system.
[0203] At operation 901, as discussed above, the shareable content item feeding system 590 (e.g., user system 102 or server) selects, through an interactive application, a single content item that matches one or more criteria corresponding to the shareable content from a plurality of previously captured content items.
[0204] At operation 902, as discussed above, the shareable content item feed system 590 generates a prompt that includes individual content items and requests for multiple suggested modifications to those individual content items.
[0205] At operation 903, as discussed above, the shareable content item feed system 590 processes prompts via LLM to generate multiple suggested modifications for individual content items.
[0206] At operation 904, as discussed above, the shareable content item feed system 590 generates a modified individual content item corresponding to the individual suggested modification among multiple suggested modifications.
[0207] Figures 10A to 10BThis is a flowchart of a process or method 1000 performed by a shareable content item feed system 590, based on some examples. While a flowchart can describe operations as a sequential process, many operations can be performed in parallel or simultaneously. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed. A process can correspond to a method, procedure, etc. The steps of a method can be performed in whole or in part, in combination with some or all of the steps in other methods, and can be performed by any number of different systems or any part thereof, such as a processor in any system included in the system. The process or method 1000 is repeated for each additional content item included in the candidate group of content items generated by the shareable content item feed system 590.
[0208] At operation 1001, as discussed above, the shareable content item feed system 590 (e.g., user system 102 or server) detects that a user has saved or created a new content item. In response, at operation 1002, the shareable content item feed system 590 generates embeddings, clues, or descriptive tags that describe the visual attributes of the content item.
[0209] At operation 1010, the shareable content item feed system 590 retrieves a set of criteria, such as a pre-selected set of visual tags for identifying shareable content. Then, at operation 1011, the shareable content item feed system 590 compares this set of criteria with the embedding of the content item to determine whether the content item meets the shareability criteria. If so, at operation 1012, the shareable content item feed system 590 determines whether the user who captured the content item has authorized the shareable content item feed system 590 to automatically perform editing.
[0210] At operation 1020, the shareable content item feed system 590 determines whether the user's credentials qualify the user or authorize the user to automatically modify the content item. In response to determining that the user is not authorized to make automatically generated modifications, the shareable content item feed system 590 proceeds to operation 1022, where another content item is analyzed at operations 1001 and / or 1002. In response to determining that the user is authorized to make automatically generated modifications, the shareable content item feed system 590 executes operation 1024.
[0211] At operation 1024, the security of the content item is scanned using a security model. If the content item is determined to meet security criteria at operation 1030, the shareable content item feed system 590 proceeds to operation 1040. Otherwise, the shareable content item feed system 590 proceeds to operation 1022. At operation 1040, the description of the content item, along with a hint 600, is provided to the LLM. At operation 1042, the LLM processes the content item and / or its description to generate multiple suggested modifications. The LLM may receive one or more positive and negative hints 1043 to ensure that the generated suggested modifications meet security criteria and pass the security model / security check. In some cases, the shareable content item feed system 590 stores the hints in cold storage 1044.
[0212] At operation 1050, the shareable content item feed system 590 selects a random suggested modification from a plurality of suggested modifications. The shareable content item feed system 590 determines whether the selected suggested modification includes explanatory text editing (e.g., including fields that provide descriptions for explanatory text tools). If so, the shareable content item feed system 590 performs operation 1052 to check the security of the explanatory text's text using a security model. If it is determined in operation 1054 that the suggested modification does not include explanatory text editing, the shareable content item feed system 590 performs operation 1056. The shareable content item feed system 590 performs operation 1082, wherein if the explanatory text's text fails the security check, editing is not performed or not allowed, and the shareable content item feed system 590 selects another candidate content item at operations 1001 and / or 1002.
[0213] In response to the determination that the descriptive text passes a security check, the shareable content item feed system 590 performs operation 1056 to determine whether the suggested modifications include image-to-image modifications. If not, the shareable content item feed system 590 proceeds to operation 1090, where the content item is modified according to the suggested modifications (e.g., adding descriptive text, adding filters, and / or generating a new image with filters / descriptive text) and presented to the consumer / user.
[0214] In response to determining that the suggested modification includes an image-to-image modification in the field specifying the modification tool, the shareable content item feed system 590 performs operation 1060. At operation 1060, the shareable content item feed system 590 uses LLM processing based on prompt template 701 to process the descriptive text in the field used for the image-to-image modification. The LLM revises the prompt and generates a new prompt, which is then analyzed for security at operation 1062. If the new prompt passes security at operation 1064, a generative model is applied to generate a new content item based on the new prompt at operation 1070. The security of the new image is analyzed at operation 1080, and if the new content item passes the security check, the shareable content item feed system 590 proceeds to operation 1090, where the new content item is presented and / or modified according to the selected suggested modification. If the new content item fails the security check, the shareable content item feed system 590 performs operation 1082, where the content item is discarded and another content item is selected for analysis.
[0215] Example
[0216] Example 1. A method comprising: selecting, via an interactive application, a single content item that matches one or more criteria corresponding to shareable content from a plurality of previously captured content items; generating a prompt including the single content item and a request for a plurality of suggested modifications to the single content item; processing the prompt via a large language model (LLM) to generate the plurality of suggested modifications to the single content item; and generating a modified single content item corresponding to the single suggested modification among the plurality of suggested modifications.
[0217] Example 2. According to the method of Example 1, wherein one or more criteria include a list of predetermined descriptions, wherein one or more criteria exclude content items that are specified as private by the user of the interactive application, and wherein one or more criteria exclude content items with lighting that meets a darkness threshold.
[0218] Example 3. A method according to any one of Examples 1 to 2, wherein each of the previously captured content items is processed to generate a visual label and the visual label is compared with one or more criteria to identify shareable content.
[0219] Example 4. Following the method of Example 3, the prompt includes a visual label associated with an individual content item, a timestamp indicating when the individual content item was captured, the location where the individual content item was captured, the current time and date, and the language associated with the user.
[0220] Example 5. Following the method of Example 4, the prompt identifies multiple creative tools that the LLM can use when generating multiple suggested modifications.
[0221] Example 6. According to the method of Example 5, the multiple creative tools include an Add Description Text tool for adding description text to individual content items, an Image-to-Image tool for processing individual content items through a generative model to generate new content items based on the description, and a Filter tool for selecting predefined filters that match one or more keywords.
[0222] Example 7. According to the method of Example 6, wherein the plurality of suggested modifications include: a first suggested modification, the first suggested modification including a first description of the first suggested modification, a first atmosphere representing the first suggested modification, and a first combination of one or more creative tools; and a second suggested modification, the second suggested modification including a second description of the second suggested modification, a second atmosphere representing the second suggested modification, and a second combination of one or more creative tools, the second combination including a subset of the creative tools that is different from the first combination.
[0223] Example 8. According to the method of Example 7, wherein a first combination of one or more creative tools includes an add descriptive text tool containing a first descriptive text and a filter tool containing a first set of keywords.
[0224] Example 9. The method according to any one of Examples 7 to 8 further includes: randomly selecting a first suggested modification from a plurality of suggested modifications.
[0225] Example 10. The method according to Example 9 further includes: determining that the first suggested modification includes an add description text tool containing the first description text; and in response to determining that the first suggested modification includes an add description text tool containing the first description text, overlaying the text including the first description text onto a separate content item at a predetermined location.
[0226] Example 11. The method of Example 10 further includes: attaching a graphic element indicating that the text was generated by an LLM to the beginning of the text; and attaching a graphic element indicating that the text was generated by an LLM to the end of the text.
[0227] Example 12. A method according to any of Examples 10 to 11, wherein, in response to user input, the text can be removed from an overlay on a separate content item.
[0228] Example 13. The method according to any one of Examples 9 to 12 further includes: determining that the first suggested modification includes a filter tool containing the first set of keywords; and in response to determining that the first suggested modification includes a filter tool containing the first set of keywords, searching among a plurality of predetermined filters for individual filters that match the first set of keywords.
[0229] Example 14. The method according to Example 13 further includes: ranking the multiple predefined filters based on comparing metadata associated with each of the multiple predefined filters with a first set of keywords; and selecting the individual filter associated with the highest ranking in response to the ranking of the multiple predefined filters.
[0230] Example 15. According to the method of any one of Examples 7 to 14, wherein a second combination of one or more creative tools includes an image-to-image tool containing a first descriptive modification.
[0231] Example 16. The method according to Example 15 further includes: generating a new prompt that includes a first descriptive modification and a request to expand and refine the first descriptive modification; and processing the new prompt via LLM to generate a revised image modification description, the new prompt including one or more negative revisions indicating that the modification is not allowed to be performed.
[0232] Example 17. The method according to Example 16 further includes: processing individual content items and new hints through a generative model to generate new content items corresponding to the revised image modification description, wherein the modified individual content items include a portion of the new content items.
[0233] Example 18. The method according to any one of Examples 1 to 17 further includes: presenting a modified individual content item with an indicator via an interactive application, the indicator specifying that the modified individual content item is automatically generated.
[0234] Example 19. A system comprising: at least one processor; and at least one memory unit storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations including: selecting, via an interactive application, a single content item from a plurality of previously captured content items that matches one or more criteria corresponding to shareable content; generating a prompt including the single content item and a request for a plurality of suggested modifications to the single content item; processing the prompt via a large language model (LLM) to generate the plurality of suggested modifications to the single content item; and generating a modified single content item corresponding to the single suggested modification among the plurality of suggested modifications.
[0235] Example 20. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations including: selecting, via an interactive application, a single content item from a plurality of previously captured content items that matches one or more criteria corresponding to shareable content; generating a prompt including the single content item and a request for a plurality of suggested modifications to the single content item; processing the prompt via a large language model (LLM) to generate the plurality of suggested modifications to the single content item; and generating a modified single content item corresponding to the single suggested modification among the plurality of suggested modifications.
[0236] Machine architecture
[0237] Figure 11 This is a schematic representation of machine 1100, within which instructions 1102 (e.g., software, program, application, app, or other executable code) can be executed to cause machine 1100 to perform any or more of the methods discussed herein. For example, instructions 1102 can cause machine 1100 to perform any or more of the methods described herein. Instructions 1102 transform the general, unprogrammed machine 1100 into a specific machine 1100 programmed to perform the described and illustrated functions in the described manner. Machine 1100 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 1100 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1100 may include, but is not limited to, server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing instructions 1102 specifying actions to be taken by machine 1100. Furthermore, while a single machine 1100 is shown, the term "machine" should also be considered as a collection of machines that individually or jointly execute instructions 1102 to perform any or more of the methods discussed herein. For example, machine 1100 may include user system 102 or any of a plurality of server devices forming part of interactive server system 110. In some examples, machine 1100 may also include both client and server systems, wherein some operations of a particular method or algorithm are performed on the server side, and wherein some operations of a particular method or algorithm are performed on the client side.
[0238] Machine 1100 may include a processor 1104, a memory 1106, and an input / output (I / O) unit 1108 that can be configured to communicate with each other via a bus 1110. In the example, processor 1104 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processors 1112 and 1114 that execute instruction 1102. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Figure 11 Multiple processors 1104 are shown, but machine 1100 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0239] Memory 1106 includes main memory 1116, static memory 1118, and storage cells 1120, all of which are accessible by processor 1104 via bus 1110. Main memory 1106, static memory 1118, and storage cells 1120 store instructions 1102 embodying any one or more of the methods or functions described herein. Instructions 1102 may also reside wholly or partially in main memory 1116, in static memory 1118, in machine-readable medium 1122 within storage cell 1120, in at least one processor of processor 1104 (e.g., in the processor's cache memory), or in any suitable combination thereof during execution by machine 1100.
[0240] I / O component 1108 may include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurements, etc. The specific I / O component 1108 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It will be understood that I / O component 1108 may include... Figure 11Many other components are not shown. In various examples, I / O component 1108 may include user output component 1124 and user input component 1126. User output component 1124 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tube (CRT) displays), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. User input component 1126 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens providing position and force for touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc. Any biometrics collected by the biometric component are captured and stored with the user's consent and deleted upon the user's request.
[0241] Furthermore, such biometric data can be used for very limited purposes, such as identity verification. To ensure the limited and authorized use of biometric information and other personally identifiable information (PII), access to this data is restricted to authorized personnel, if permitted. Any use of biometric data may be strictly limited to identity verification purposes, and the data may not be shared or sold to any third party without the user's explicit consent. In addition, appropriate technical and organizational measures have been implemented to ensure the security and confidentiality of this sensitive information.
[0242] In another example, I / O component 1108 may include biometric component 1128, motion component 1130, environmental component 1132, or position component 1134, and various other components. For example, biometric component 1128 includes components for detecting expressions (e.g., hand gestures, facial expressions, vocal expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and recognizing a person (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). The biometric component may include a brain-computer interface (BMI) system that allows communication between the brain and external devices or machines. This can be achieved by recording brain activity data, converting that data into a format that can be understood by a computer, and then using the resulting signals to control the device or machine.
[0243] Examples of BMI technology types include:
[0244] Brain-based brain-computer interfaces (BMIs) use electrodes placed on the scalp to record electrical activity in the brain.
[0245] Invasive BMI, which uses electrodes that are surgically implanted in the brain.
[0246] Optogenetics BMI uses light to control the activity of specific nerve cells in the brain.
[0247] The moving part 1130 includes an acceleration sensor part (e.g., an accelerometer), a gravity sensor part, and a rotation sensor part (e.g., a gyroscope).
[0248] The environmental component 1132 includes, for example, one or more camera devices (with still image / photograph and video capabilities), lighting sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting nearby objects), gas sensors (e.g., gas detection sensors for detecting the concentration of hazardous gases for safety purposes or for measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment.
[0249] Regarding the camera device, user system 102 may have a camera device system including, for example, a front-facing camera on the front surface of user system 102 and a rear-facing camera on the rear surface of user system 102. The front-facing camera may be used, for example, to capture still images and videos (e.g., “selfies”) of the user of user system 102, which can then be enhanced with the enhancement data (e.g., filters) described above. For example, the rear-facing camera may be used to capture still images and videos in a more conventional camera device mode, wherein these images are similarly enhanced with enhancement data. In addition to the front-facing and rear-facing cameras, user system 102 may also include features for capturing 360°... 360 photos and videos Camera device.
[0250] Furthermore, the camera system of user system 102 may include dual rear cameras (e.g., a main camera and a depth-sensing camera), or even triple, quadruple, or quintuple rear camera configurations on the front and rear sides of user system 102. For example, these multiple camera systems may include wide-angle cameras, ultra-wide-angle cameras, telephoto cameras, macro cameras, and depth sensors.
[0251] The position component 1134 includes a positioning sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure to obtain altitude), an orientation sensor component (e.g., a magnetometer), etc.
[0252] Various technologies can be used to achieve communication. I / O component 1108 also includes a communication component 1136 operable to couple machine 1100 to network 1138 or device 1140 via a corresponding coupling or connection. For example, communication component 1136 may include a network interface component or another suitable device that interfaces with network 1138. In further examples, communication component 1136 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, or Bluetooth. Components (e.g., Bluetooth) Low energy consumption), Wi-Fi Components and other communication components that provide communication via other modes. Device 1140 can be any peripheral device from another machine or various peripheral devices (e.g., a peripheral device coupled via a universal serial bus (USB)).
[0253] Furthermore, communication component 1136 may detect identifiers or include components operable to detect identifiers. For example, communication component 1136 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, QR codes such as Quick Response (QR) codes, Aztec codes, data matrices, and data symbols). The system can utilize optical sensors for multidimensional barcodes and other optical codes, such as MaxiCode, PDF417, UltraCode, and UCC RSS-2D barcodes, or acoustic detection components (e.g., microphones for identifying audio signals of the tags). Additionally, various information can be obtained via communication component 1136, such as location derived from Internet Protocol (IP) geolocation, or information via Wi-Fi. Location can be obtained through signal triangulation or by detecting NFC beacon signals that indicate a specific location.
[0254] Various memories (e.g., main memory 1116, static memory 1118, and the memory of processor 1104) and storage units 1120 may store one or more sets of instructions and data structures (e.g., software) embodied or used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 1102) cause various operations to implement the disclosed examples when executed by processor 1104.
[0255] Instructions 1102 can be sent or received over network 1138 via a transmission medium using a network interface device (e.g., a network interface component included in communication component 1136) and using any of several known transmission protocols (e.g., HTTP). Similarly, instructions 1102 can be sent or received via a transmission medium through a coupling to device 1140 (e.g., a peer-to-peer coupling).
[0256] Software Architecture
[0257] Figure 12 This is a block diagram 1200 illustrating a software architecture 1202 that can be installed on any or more of the devices described herein. The software architecture 1202 is supported by hardware such as a machine 1204 including a processor 1206, memory 1208, and I / O components 1210. In this example, the software architecture 1202 can be conceptualized as a stack of layers, where each layer provides specific functionality. The software architecture 1202 includes layers such as an operating system 1212, libraries 1214, frameworks 1216, and applications 1218. Operationally, application 1218 activates API calls 1220 via the software stack and receives messages 1222 in response to API calls 1220.
[0258] Operating system 1212 manages hardware resources and provides public services. Operating system 1212 includes, for example, a kernel 1224, services 1226, and drivers 1228. Kernel 1224 serves as an abstraction layer between hardware and other software layers. For example, kernel 1224 provides memory management, processor management (e.g., scheduling), component management, network and security settings, and other functions. Services 1226 can provide other public services to other software layers. Driver 1228 is responsible for controlling or interfacing with the underlying hardware. For example, driver 1228 may include a display driver, a camera driver, or a Bluetooth driver. or Bluetooth Low-power drives, flash drives, serial communication drives (e.g., USB drives), Wi-Fi Drivers, audio drivers, power management drivers, etc.
[0259] Library 1214 provides common low-level infrastructure used by application 1218. Library 1214 may include system library 1230 (e.g., the C standard library), which provides functions such as memory allocation, string manipulation, and mathematical functions. Additionally, library 1214 may include API library 1232, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Picture Experts Group (JPEG or JPG), or Portable Web Graphics (PNG)), graphics libraries (e.g., the OpenGL framework for rendering graphic content in 2D and 3D on a display), database libraries (e.g., SQLite, which provides various relational database functions), web libraries (e.g., WebKit, which provides web browsing capabilities), and so on. Library 1214 may also include various other libraries 1234 to provide many other APIs to application 1218.
[0260] Framework 1216 provides common high-level infrastructure for use by application 1218. For example, framework 1216 provides various GUI functions, advanced resource management, and advanced location services. Framework 1216 can provide a wide range of other APIs that can be used by application 1218, some of which may be specific to a particular operating system or platform.
[0261] In the example, application 1218 may include home application 1236, contact application 1238, browser application 1240, book reader application 1242, location application 1244, media application 1246, messaging application 1248, game application 1250, and a wide variety of other applications such as third-party application 1252. Application 1218 is a program that performs the functions defined in the program. One or more applications 1218 can be created using various programming languages, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a particular example, third-party application 1252 (e.g., used by an entity other than the vendor of a particular platform using Android) or iOS Applications developed using the SDK can be used in systems such as iOS. ANDROID WINDOWS Mobile software running on the phone's mobile operating system or other mobile operating systems. In this example, a third-party application 1252 can activate API call 1220 provided by the operating system 1212 to facilitate the functionality described herein.
[0262] Systems with head-worn devices
[0263] Figure 13 A system 1300 including a head-worn wearable device 116 with a selector input device is shown according to some examples. Figure 13 This is a high-level functional block diagram of an example head-mounted wearable device 116 that is communicatively coupled to mobile devices 114 and various server systems 1304 (e.g., interactive server system 110) via various networks 1316.
[0264] The head-mounted wearable device 116 includes one or more camera devices, each of which may be, for example, a visible light camera 1306, an infrared emitter 1308, and an infrared camera 1310.
[0265] Mobile device 114 connects to head-mounted wearable device 116 using both low-power wireless connection 1312 and high-speed wireless connection 1314. Mobile device 114 also connects to server system 1304 and network 1316.
[0266] The head-mounted wearable device 116 also includes two image displays 1318 of an optical assembly. The two image displays 1318 of the optical assembly include an image display associated with the left lateral side of the head-mounted wearable device 116 and an image display associated with the right lateral side of the head-mounted wearable device 116. The head-mounted wearable device 116 also includes an image display driver 1320, an image processor 1322, a low-power circuitry system 1324, and a high-speed circuitry system 1326. The image displays 1318 of the optical assembly are used to present images and videos including the images to the user of the head-mounted wearable device 116, the images of which may include a GUI.
[0267] The image display driver 1320 commands and controls the image display 1318 of the optical components. The image display driver 1320 can directly transmit image data to the image display 1318 of the optical components for presentation, or it can convert the image data into a signal or data format suitable for transmission to the image display device. For example, the image data can be video data formatted according to compression formats such as H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, etc., and still image data can be formatted according to compression formats such as PNG, JPEG, Tagged Image File Format (TIFF), or Exchangeable Image File Format (EXIF).
[0268] The head-worn device 116 includes a frame and a stem (or temple) extending laterally from the frame. The head-worn device 116 also includes a user input device 1328 (e.g., a touch sensor or button) that includes an input surface on the head-worn device 116. The user input device 1328 (e.g., a touch sensor or button) receives input selections from the user to manipulate a GUI of the presented image.
[0269] Figure 13 The components shown for the head-mounted wearable device 116 are located on one or more circuit boards (e.g., PCBs or flexible PCBs) in the frame or temples. Alternatively or additionally, the depicted components may be located in blocks, frames, hinges, or bridges of the head-mounted wearable device 116. The left and right visible light camera devices 1306 may include digital camera elements, such as complementary metal-oxide-semiconductor (CMOS) image sensors, charge-coupled devices, camera lenses, or any other corresponding visible or light-capturing elements that can be used to capture data including images of scenes with unknown objects.
[0270] The head-mounted wearable device 116 includes a memory 1302 that stores instructions for performing a subset or all of the functions described herein. The memory 1302 may also include a storage device.
[0271] like Figure 13As shown, the high-speed circuit system 1326 includes a high-speed processor 1330, a memory 1302, and a high-speed wireless circuit system 1332. In some examples, an image display driver 1320 is coupled to the high-speed circuit system 1326 and operated by the high-speed processor 1330 to drive the left and right image displays of the image display 1318 of the optical components. The high-speed processor 1330 can be any processor capable of managing high-speed communication and operation of any general-purpose computing system required for the head-mounted wearable device 116. The high-speed processor 1330 includes the processing resources required to manage high-speed data transmission over a high-speed wireless connection 1314 to a wireless local area network (WLAN) using the high-speed wireless circuit system 1332. In some examples, the high-speed processor 1330 executes an operating system such as the LINUX operating system or another such operating system for the head-mounted wearable device 116, and the operating system is stored in the memory 1302 for execution. Among other duties, the high-speed processor 1330, which executes the software architecture for the head-mounted wearable device 116, also manages data transmission with the high-speed wireless circuit system 1332. In some examples, the high-speed wireless circuit system 1332 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, also referred to herein as WiFi. In some examples, other high-speed communication standards may be implemented by the high-speed wireless circuit system 1332.
[0272] The low-power wireless circuit system 1334 and high-speed wireless circuit system 1332 of the head-mounted wearable device 116 may include a short-range transceiver (Bluetooth). The device 114 includes a wireless wide area network transceiver, a local area network transceiver, or a wide area network transceiver (e.g., cellular or WiFi). The mobile device 114 (including transceivers communicating via low-power wireless connection 1312 and high-speed wireless connection 1314) can be implemented using the architectural details of the head-mounted wearable device 116, as can other components of the network 1316.
[0273] like Figure 13 As shown, the low-power processor 1336 or high-speed processor 1330 of the head-mounted wearable device 116 may be coupled to a camera device (visible light camera 1306, infrared emitter 1308 or infrared camera 1310), an image display driver 1320, a user input device 1328 (e.g., a touch sensor or button) and a memory 1302.
[0274] The head-mounted wearable device 116 is connected to a host computer. For example, the head-mounted wearable device 116 is paired with the mobile device 114 via a high-speed wireless connection 1314, or connected to the server system 1304 via a network 1316. The server system 1304 may be one or more computing devices as part of a service or network computing system, including, for example, a processor, memory, and a network communication interface for communicating with the mobile device 114 and the head-mounted wearable device 116 via the network 1316.
[0275] Mobile device 114 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via network 1316, low-power wireless connection 1312, or high-speed wireless connection 1314. Mobile device 114 may also store in its memory at least a portion of instructions for generating binaural audio content to achieve the functions described herein.
[0276] The output components of the head-worn wearable device 116 include visual components, such as displays (e.g., LCD, PDP, LED displays), projectors, or waveguides. The image display of the optical components is driven by an image display driver 1320. The output components of the head-worn wearable device 116 also include acoustic components (e.g., speakers), haptic components (e.g., vibration motors), other signal generators, etc. The input components (e.g., user input devices 1328) of the head-worn wearable device 116, mobile device 114, and server system 1304 may include alphanumeric input components (e.g., keyboards, touchscreens, photoelectric keyboards, or other alphanumeric input components configured to receive alphanumeric input), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens that provide position and force for touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc.
[0277] The head-mounted wearable device 116 may also include additional peripheral device elements. Such peripheral device elements may include biometric sensors, additional sensors, or display elements integrated with the head-mounted wearable device 116. For example, peripheral device elements may include any I / O components, including output components, motion components, position components, or any other such elements described herein.
[0278] For example, biometric components include those for detecting expressions (e.g., gestures, facial expressions, vocalizations, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brainwaves), and identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Biometric components may include a BMI system that allows communication between the brain and external devices or machines. This can be achieved by recording brain activity data, converting that data into a format that can be understood by a computer, and then using the resulting signals to control the device or machine.
[0279] Moving components include accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Positioning components include position sensor components (e.g., GPS receiver components) for generating position coordinates, and Wi-Fi or Bluetooth for generating positioning system coordinates. Transceivers, altitude sensor components (e.g., altimeters or barometers that detect air pressure to obtain altitude), orientation sensor components (e.g., magnetometers), etc. The coordinates of such a positioning system can also be received from the mobile device 114 via a low-power wireless circuit system 1334 or a high-speed wireless circuit system 1332 through a low-power wireless connection 1312 and a high-speed wireless connection 1314.
[0280] Glossary
[0281] "Carrier signal" refers to any intangible medium, such as a medium capable of storing, encoding, or carrying machine-executable instructions and including digital or analog communication signals, or other intangible medium facilitating the transmission of such instructions. Instructions can be sent or received over a network using a transmission medium via a network interface device.
[0282] "Client device" means any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. Client devices can be, but are not limited to, mobile phones, desktop computers, laptop computers, portable digital assistants (PDAs), smartphones, tablet computers, ultrabooks, netbooks, laptop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, STBs, or any other communication device that a user can use to access the network.
[0283] "Communication network" refers to one or more parts of a network, such as an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), WLAN, wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a part of the Internet, a part of the public switched telephone network (PSTN), a common old-style telephone service (POTS) network, a cellular telephone network, a wireless network, or Wi-Fi. A network, other types of networks, or a combination of two or more such networks. For example, a network or part of a network may include a wireless network or a cellular network, and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or other types of cellular or wireless coupling. In this example, the coupling can implement any data transmission technology of various types, such as Single Carrier Radio Transmission (1xRTT), Evolved Data Optimization (EVDO), General Packet Radio Service (GPRS), Enhanced Data Rate Evolution of GSM (EDGE), the 3rd Generation Partnership Project (3GPP) including 3G, fourth-generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Global Microwave Access Interoperability (WiMAX), Long Term Evolution (LTE) standards, other data transmission technologies defined by various standards setting organizations, other long-distance protocols, or other data transmission technologies.
[0284] A "component" can refer, for example, to a logical or physical entity having boundaries defined by functional or subroutine calls, branch points, APIs, or other technologies that partition or modularize specific processing or control functions. A component can be combined with other components via its interface to perform machine processing. A component can be a packaged functional hardware unit designed for use with other components, and part of a program that typically performs related functions. A component can constitute a software component (e.g., code embodied on a machine-readable medium) or a hardware component.
[0285] A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in a physical manner. In various examples, one or more hardware components (e.g., processors or processor groups) of a computer system (e.g., a standalone computer system, a client computer system, or a server computer system) or a computer system can be configured by software (e.g., an application or an application portion) to operate to perform certain operations described herein.
[0286] Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component may include a dedicated circuit system or logic permanently configured to perform certain operations. A hardware component may be a dedicated processor, such as a field-programmable gate array (FPGA) or an ASIC. A hardware component may also include programmable logic or a circuit system temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific part of a machine) uniquely tailored to perform the configured function, and is no longer a general-purpose processor. It will be understood that the decision to implement a hardware component mechanically in a dedicated and permanently configured circuit system or in a temporarily configured (e.g., software-configured) circuit system may be driven by cost and time considerations. Therefore, the phrase “hardware component” (or “hardware-implemented component”) should be understood to encompass tangible entities that are physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate or perform certain operations described herein.
[0287] Consider an example where hardware components are temporarily configured (e.g., programmed), without requiring each of the hardware components to be configured or instantiated at any given time. For example, in cases where the hardware components include a general-purpose processor that becomes a dedicated processor through software configuration, this general-purpose processor can be configured at different times as (e.g., including different hardware components) different dedicated processors. The software accordingly configures one or more specific processors to constitute a specific hardware component at one time and different hardware components at different times. Hardware components can provide information to and receive information from other hardware components. Thus, the described hardware components can be considered communicatively coupled. In cases where multiple hardware components exist simultaneously, communication can be achieved through signal transmission between or among two or more hardware components (e.g., via appropriate circuitry and buses). In examples where multiple hardware components are configured or instantiated at different times, such communication between hardware components can be achieved, for example, through the storage and retrieval of information in a memory structure accessible to the multiple hardware components. For example, a hardware component can perform an operation and store the output of that operation in a memory device communicatively coupled to it. Another hardware component can then access the memory device at a later time to retrieve and process the stored output. Hardware components can also initiate communication with input or output devices and can operate on resources (e.g., collections of information). The various operations of the example methods described herein can be performed, at least in part, by (e.g., via software) one or more processors temporarily or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute components of a processor implementation that operate to perform one or more operations or functions described herein.
[0288] As used herein, a “processor-implemented component” refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be implemented at least in part by processors, where one or more specific processors are examples of hardware. For example, at least some of the operations of the methods can be performed by one or more processors or processor-implemented components. Furthermore, one or more processors can also operate to support the execution of related operations in a “cloud computing” environment or as a “Software as a Service” (SaaS) operation. For example, at least some of the operations can be performed by a group of computers (as an example of machines including processors), where these operations are accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of some operations can be distributed across processors, not residing within a single machine, but deployed across multiple machines. In some examples, the processor or processor-implemented component may reside in a single geographic location (e.g., in a home environment, office environment, or server cluster). In other examples, the processor or processor-implemented component may be distributed across multiple geographic locations.
[0289] "Computer-readable storage medium" refers to both, for example, machine storage media and transmission media. Therefore, these terms include both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" refer to the same thing and can be used interchangeably in this disclosure. "Temporary message" refers to a message that is accessible for a limited time period, for example. A temporary message can be text, an image, video, etc. The access time of a temporary message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is temporary.
[0290] "Machine storage medium" refers to one or more storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. This term should be accordingly considered to include, but is not limited to, solid-state memory, as well as optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memory, including, by way of example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGAs, and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium," "device storage medium," and "computer storage medium" refer to the same thing and may be used interchangeably in this disclosure.
[0291] The terms “machine storage medium,” “computer storage medium,” and “device storage medium” explicitly exclude carrier waves, modulated data signals, and other such media, at least some of which are encompassed within the term “signal medium.” “Non-transitory computer-readable storage medium” refers to, for example, a tangible medium capable of storing, encoding, or carrying instructions executable by a machine. “Signal medium” refers to, for example, any intangible medium capable of storing, encoding, or carrying instructions executable by a machine and comprising digital or analog communication signals, or other intangible medium facilitating the transmission of software or data. The term “signal medium” should be considered to include any form of modulated data signal, carrier wave, etc. The term “modulated data signal” means a signal whose characteristics are set or altered in a manner that encodes information in the signal. The terms “transmission medium” and “signal medium” refer to the same thing and may be used interchangeably in this disclosure.
[0292] "User equipment" means, for example, a device that is accessed, controlled, or owned by a user and that the user interacts with to perform actions or interactions, including interactions with other users or computer systems. "Carrier signal" means any intangible medium or other intangible medium capable of storing, encoding, or carrying machine-executable instructions and comprising digital or analog communication signals. Instructions can be sent or received over a network using a transmission medium via a network interface device. "Client device" means any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. Client devices can be, but are not limited to, mobile phones, desktop computers, laptop computers, PDAs, smartphones, tablet computers, ultrabooks, netbooks, laptop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, STBs, or any other communication device that a user can use to access the network.
[0293] "Communication network" refers to one or more parts of a network, which can be an ad hoc network, intranet, extranet, VPN, LAN, WLAN, WAN, WWAN, MAN, the Internet, a part of the Internet, a part of PSTN, POTS network, cellular telephone network, wireless network, Wi-Fi. A network, other types of networks, or a combination of two or more such networks. For example, a network or part of a network may include a wireless network or a cellular network, and coupling may be a CDMA connection, a GSM connection, or other types of cellular or wireless coupling. In this example, coupling can implement any data transmission technology of various types, such as 1xRTT, EVDO, GPRS, EDGE, 3GPP (including 3G), 4G networks, UMTS, HSPA, WiMAX, LTE standards, other data transmission technologies defined by various standards setting organizations, other long-distance protocols, or other data transmission technologies.
[0294] Components can constitute software components (e.g., code embodied on a machine-readable medium) or hardware components. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various examples, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components (e.g., processors or processor groups) of a computer system can be configured by software (e.g., an application or application portion) to operate to perform some of the operations described herein.
[0295] Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component may include a dedicated circuit system or logic permanently configured to perform certain operations. A hardware component may be a dedicated processor, such as an FPGA or ASIC. A hardware component may also include programmable logic or a circuit system temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific part of a machine) uniquely tailored to perform the configured function, and is no longer a general-purpose processor. It will be understood that the decision to implement a hardware component mechanically in a dedicated and permanently configured circuit system or in a temporarily configured (e.g., software-configured) circuit system may be driven by cost and time considerations. Therefore, the phrase “hardware component” (or “hardware-implemented component”) should be understood to encompass tangible entities that are physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate or perform certain operations described herein.
[0296] The various operations of the example methods described herein can be performed at least in part by one or more processors, which are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute components of a processor implementation that performs operations to execute one or more of the operations or functions described herein.
[0297] Changes and modifications may be made to the disclosed examples without departing from the scope of this disclosure. Such and other changes or modifications are intended to be included within the scope of this disclosure as set forth in the appended claims.
Claims
1. A method comprising: Select individual content items from multiple previously captured content items that match one or more criteria corresponding to shareable content through interactive applications; Generate a prompt that includes the individual content item and a request for multiple suggested modifications to the individual content item; The prompts are processed using a large language model (LLM) to generate the multiple suggested modifications for the individual content items; as well as Generate a modified individual content item corresponding to the individual suggested modification among the plurality of suggested modifications.
2. The method according to claim 1, wherein, The one or more criteria include a list of predetermined descriptions, wherein the one or more criteria exclude content items designated as private by the user of the interactive application, and wherein the one or more criteria exclude content items with lighting that meets a darkness threshold.
3. The method according to any one of claims 1 to 2, wherein, Each of the previously captured content items is processed to generate a visual tag, and the visual tag is compared with one or more criteria to identify shareable content.
4. The method according to claim 3, wherein, The prompt includes the visual label associated with the individual content item, a timestamp indicating when the individual content item was captured, the location where the individual content item was captured, the current time and date, and the language associated with the user.
5. The method according to any one of claims 1 to 4, wherein, The prompts indicate the multiple creative tools that the LLM can use when generating the multiple suggested modifications.
6. The method according to claim 5, wherein, The plurality of creative tools include an add description text tool for adding description text to the individual content items, an image-to-image tool for processing the individual content items through a generative model to generate new content items based on the description, and a filter tool for selecting predefined filters that match one or more keywords.
7. The method according to any one of claims 1 to 6, wherein, The proposed modifications include: The first proposed modification includes a first description of the first proposed modification, a first atmosphere representing the first proposed modification, and a first combination of one or more of the plurality of creative tools; and The second proposed modification includes a second description of the second proposed modification, a second atmosphere representing the second proposed modification, and a second combination of one or more of the plurality of creative tools, the second combination including a subset of the plurality of creative tools that is different from the first combination.
8. The method according to claim 7, wherein, The first combination of one or more of the plurality of creative tools includes the add descriptive text tool containing a first descriptive text and the filter tool containing a first set of keywords.
9. The method according to any one of claims 1 to 8, further comprising: The first suggested modification is randomly selected from the plurality of suggested modifications.
10. The method of claim 9, further comprising: It is determined that the first suggested modification includes the tool for adding explanatory text containing the first explanatory text; as well as In response to determining that the first suggested modification includes the add explanatory text tool containing the first explanatory text, the text including the first explanatory text is overlaid on the separate content item at a predetermined location.
11. The method according to any one of claims 1 to 10, further comprising: Add a graphic element to the front portion of the text to indicate that the text was generated by the LLM; as well as Append the graphic element indicating that the text was generated by the LLM to the end of the text.
12. The method according to any one of claims 1 to 11, wherein, In response to user input, the text can be removed from the overlay on the individual content item.
13. The method according to any one of claims 1 to 12, further comprising: It is determined that the first suggested modification includes the filter tool containing the first set of keywords; as well as In response to determining that the first suggested modification includes the filter tool containing the first set of keywords, a search is conducted among a plurality of predetermined filters for individual filters that match the first set of keywords.
14. The method of claim 13, further comprising: The multiple predetermined filters are ranked by comparing the metadata associated with each of the multiple predetermined filters with the first set of keywords; as well as In response to the ranking of the plurality of predetermined filters, the individual filter associated with the highest ranking is selected.
15. The method according to any one of claims 1 to 14, wherein, The second combination of one or more of the plurality of creative tools includes the image-to-image tool containing a first descriptive modification.
16. The method of claim 15, further comprising: Generate a new prompt that includes the first descriptive modification and requests to expand and refine the first descriptive modification; as well as The new prompt is processed by the LLM to generate a revised image modification description, the new prompt including one or more negative revisions indicating that modifications are not allowed.
17. The method according to any one of claims 1 to 16, further comprising: The individual content items and the new hints are processed by a generative model to generate new content items corresponding to the revised image modification description, wherein the revised individual content items include a portion of the new content items.
18. The method according to any one of claims 1 to 17, further comprising: The modified individual content item is presented via the interactive application with an indicator that specifies that the modified individual content item is automatically generated.
19. A system comprising: At least one processor; as well as At least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform an operation, the operation including: Select individual content items from multiple previously captured content items that match one or more criteria corresponding to shareable content through interactive applications; Generate a prompt that includes the individual content item and a request for multiple suggested modifications to the individual content item; The prompts are processed using a large language model (LLM) to generate the multiple suggested modifications for the individual content items; and Generate a modified individual content item corresponding to the individual suggested modification among the plurality of suggested modifications.
20. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform an operation, the operation comprising: Select individual content items from multiple previously captured content items that match one or more criteria corresponding to shareable content through interactive applications; Generate a prompt that includes the individual content item and a request for multiple suggested modifications to the individual content item; The prompts are processed using a large language model (LLM) to generate the multiple suggested modifications for the individual content items; as well as Generate a modified individual content item corresponding to the individual suggested modification among the plurality of suggested modifications.