Storing entries in and retrieving information from episode object memory

The landmark memory system within an episodic object memory addresses the challenge of efficiently encoding and retrieving relevant memories by generating and ranking embeddings, enhancing user and computational efficiency in accessing contextual information.

JP2026507396APending Publication Date: 2026-03-04MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing systems face challenges in efficiently encoding and retrieving relevant memories due to limited cognitive resources and the overwhelming complexity of human and machine experiences, leading to difficulties in recalling important details from various everyday moments.

Method used

The implementation of landmark memories within an episodic object memory system, where embeddings are generated and ranked to identify and store salient memories, allowing for efficient retrieval and access based on contextual relationships.

Benefits of technology

Improves user efficiency and computational resources by enabling context-sensitive access to relevant content, reducing storage requirements and search time for information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026507396000001_ABST
    Figure 2026507396000001_ABST
Patent Text Reader

Abstract

A method is provided for identifying and storing landmark memories within an episodic object memory, including receiving one or more content items. The content items each have one or more content data. The one or more content data associated with the one or more content items may be provided and integrated into one or more embedding models representing a growing set of episodic memories. Episodic memory relates to the ability to recall content from one's personal past, such as in the form of landmark memories, which may be filtered from multiple memories based on their degree of salience. References to the one or more landmark memories and associated content items are inserted into the episodic object memory for later recall and use in the course of the context and task at hand.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] background Human and machine-based intelligences are typified as operating under constraints on a set of key cognitive resources, such as memory and the computational power required for recall and reasoning. The limited efficiency with which humans or AI-based systems can selectively encode experiences into short-term and long-term memory and then later recall the information most relevant to the current context and task at hand drives the need for mechanisms that can identify the most salient memories for storage and retrieval.

[0002] Agents, whether human or artificial, can be exposed to a wealth of information about the experiences they encounter, including sensed information, problems encountered and solved, successes and failures in achieving goals, and encodings of reasoning processes and strategies. Experiences can include specific occurrences or events, important people, relationships, and locations, signals about rewards for actions or situations, and locations, as well as the learning of new knowledge and strategies, whether custom-tailored or not, that can be applied many times in the future to solve problems more efficiently.

[0003] The process for selectively encoding and accessing important memories, referred to herein as "landmark memories," must select landmark memories from an overwhelming amount of experience. Selectivity is important for endowing the cognitive system with the ability to efficiently store and retrieve the most valuable information when needed to maximize its goals.

[0004] Some systems have efficient means for storing large amounts of experience, but may have limited ability to recall and infer related content in specific contexts. Landmark memories can be used as specific pointers to larger memory stores, allowing for efficient organization of large amounts of stored information for efficient retrieval, such as landmarks serving as easily accessible and recognizable "handles" to progressively more detailed memories. For example, people often want to recall past events, but because memory is known to fade, it can be difficult for users to remember the details of various moments. When attempting to recall details, users may recall at least some of the various types of contextual information relevant to their daily lives. For example, after a meeting where many different topics were covered, a user may not remember all of them. However, the user may remember other aspects of the meeting, such as the location, one or more of the attendees, etc. In fact, many events in daily life include many different subsets of information, such as weather, location, people, news, sights, and sounds. While a computing device can collect weather data, location data, attendee information, related news data, and other audio / video data, the unstructured nature of such data makes this information useless for user recall.

[0005] It is with respect to these and other general considerations that the embodiments are described. Moreover, while relatively specific problems are discussed, it should be understood that the embodiments should not be limited to solving the specific problems identified in the background. Summary of the Invention [Means for solving the problem]

[0006] overview Aspects of the present disclosure relate to methods, systems, and media for storing entries in an episodic object memory and / or receiving information from an episodic object memory.

[0007] In some examples, one or more experienced content items, each having one or more content data, are received. The one or more content data may be provided to a memory processor, which creates or accesses one or more embedding models and continues to integrate the content within an organized metric space that captures regularized distances between different content data based on attributes of the data. A set of embeddings may be received from one or more of the embedding models. In examples, the embeddings may be an embedding object that includes multiple embeddings. Each embedding in the set of embeddings may correspond to at least one content data from an individual content item. One or more embeddings in the set of embeddings may be determined by the memory processing device to be landmark memory embeddings. For example, one or more of the embeddings may be ranked, weighted, and / or scored based on a machine learning model trained by supervised learning methods and / or a calculated dissimilarity to at least one of the other embeddings in the set of embeddings, etc. A landmark memory embedding can be inserted into the episodic object memory linking memories into other landmark memories based on relationships between one or more attributes of the landmarks, such temporal and / or spatial relationships between content linked to the landmarks, or other relationships, such as social relationships between people or similarity in the context of the application of content for problem solving. In this way, the landmark memory embedding can group together experienced content that is useful to have joint access to in real-time recognition and problem solving. Additionally, in some examples, input (e.g., user input) can be received corresponding to the degree of salience of content to be stored and retrieved from the episodic memory.

[0008] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of examples will be set forth in part in the description that follows, and in part will be obvious from the description, or may be learned by practice of the disclosure.

[0009] BRIEF DESCRIPTION OF THE DRAWINGS Non-limiting and non-exhaustive examples are described with respect to the following figures: [Brief explanation of the drawings]

[0010] [Figure 1] 1 illustrates an overview of an exemplary system according to some aspects described herein. [Figure 2] 1 illustrates exemplary content data according to some aspects described herein. [Figure 3] 1 illustrates an example flow for storing entries in an episode object memory according to some aspects described herein. [Figure 4] 1 illustrates an exemplary vector space according to some aspects described herein. [Figure 5A] 1 illustrates an example method for storing a landmark memory within an episode object memory according to some aspects described herein. [Figure 5B] 1 illustrates an example method for determining that an embedding is a landmark memory embedding according to some aspects described herein. [Figure 6A] 1 illustrates an example system for retrieving information from an episode object memory according to some aspects described herein. [Figure 6B] 1 illustrates an example system for retrieving information from an episode object memory according to some aspects described herein. [Figure 7]1 illustrates an example method for retrieving information from an episode object memory according to some aspects described herein. [Figure 8A] 1 illustrates an overview of an exemplary generative machine learning model that may be used in accordance with aspects described herein. [Figure 8B] 1 illustrates an overview of an exemplary generative machine learning model that may be used in accordance with aspects described herein. [Figure 9] FIG. 1 is a block diagram illustrating exemplary physical components of a computing device in which aspects of the present disclosure may be practiced. [Figure 10] FIG. 1 shows a simplified block diagram of a computing device in which aspects of the present disclosure may be practiced. [Figure 11] FIG. 1 is a simplified block diagram of a distributed computing system in which aspects of the present disclosure may be practiced. DETAILED DESCRIPTION OF THE INVENTION

[0011] Detailed Description In the following detailed description, reference is made to the accompanying drawings which form a part hereof, and which show, by way of illustration, specific embodiments or examples. These aspects may be combined, other aspects may be utilized, and structural changes may be made without departing from the disclosure. The embodiments may be embodied as a method, system, or apparatus. Thus, the embodiments may take the form of a hardware implementation, an entirely software implementation, or an implementation combining software and hardware aspects. Therefore, the following detailed description is not to be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims and their equivalents.

[0012] Eric J. Horvitz, the lead inventor of the present disclosure, is also the lead inventor of U.S. patent application Ser. No. 11 / 172,467, filed Jun. 30, 2005, and entitled "Methods for Selecting, Organizing, and Displaying Images from Large Stores of Personal Content," which describes a system and method for automatically organizing, managing, and presenting information content in a manner that is relevant or personal to the user.

[0013] Due to limited cognitive resources of machines or humans and the overwhelming complexity of the world in which they operate, it can be difficult to remember and retrieve important details about experienced content from various everyday moments. A computing device can receive multiple different types of contextual information related to a user's everyday life, such as audio data, video data, gaze data, weather data, location data, news data, or other types of ambient data that may be recognized by those skilled in the art. A user can associate some of these different types of contextual or ambient information with details of their everyday life, such as documents presented in a meeting, emails sent on a given day, conversations about a particular topic, etc.

[0014] As humans, we often associate memories with contextual information related to the memory we are trying to recall. For example, when searching for a document on a computer, a user may remember that an important news article was published the last time they opened the document, or that there was a storm outside the last time they opened the document, or that they last accessed the document via their smartphone while traveling to a particular park. Additionally or alternatively, a user may remember that they were on a phone or video call with a particular person the last time they opened the document. However, the process of tracing a user's own memory associations to determine where to search for a document on one or more of their computing devices can be time-consuming, frustrating, and unnecessarily consume computing resources (e.g., processor and / or memory).

[0015] Furthermore, in some examples, while a user attempts to perform an action on their computing device (e.g., searching for a document, opening a document, sending an email, scheduling a calendar event, etc.), the user's computing device may, with the user's permission, record only the computing device's screen, or only the computing device's audio, or only the computing device's location, so that the record can be searched for the association the user is attempting to make. However, storing a recording (e.g., of video, audio, or location) can occupy a relatively large amount of memory. Furthermore, searching a recording can be relatively time-consuming and computationally expensive. Thus, there is a need to improve user efficiency in performing actions on a computing device based on landmark memory and associated context information.

[0016] A landmark memory, as described herein, may be a memory that a system and / or user distinguishes as important, unique, and / or different from other memories that may be stored. For example, a landmark memory may be distinguished because it is a memorable person, an important transaction, a meaningful portion of a file, or a distinct piece of content that is useful to access to recall past events. A landmark memory may have a structure of associated attributes that include dimensions of data related to the memory determined to be a landmark memory. Landmark memories and their attributes can enable efficient, context-sensitive access of relevant experienced and encoded content using the mechanisms provided herein.

[0017] A landmark memory can be a parent structure with substructures, with relationships to and / or between the substructures, such as the lower landmark memories. Thus, landmark memories can have relationships to one another in terms of location, people, task type, and / or time (as described with respect to the term "episodic memory" used throughout this specification, which those skilled in the art will recognize originates from psychology and was originally described by the teachings of psychologist Endel Tulving). Landmark memories may be proximate to one another, such as the Fourth of July and the fireworks explosion, and a dinner party the night before, including salient aspects of the key attendees and conversations. Additional and / or alternative examples may be recognized by those skilled in the art.

[0018] Accordingly, some aspects of the present disclosure relate to methods, systems, and media for storing landmark memories within an episodic object memory. Generally, a content item having one or more content data (e.g., emails, audio data, video data, messages, internet encyclopedia data, skills, commands, source code, program ratings, etc.) may be received. The one or more content data may be provided to a model to generate a set of embeddings (e.g., semantic embeddings). One or more embeddings from the set of embeddings may be determined to be landmark memory embeddings. For example, to determine that one or more embeddings are landmark memory embeddings, the set of embeddings may be ranked, weighted, scored, and / or provided to a trained machine learning model. Such models may potentially include models that can make this determination without being directly trained to perform this task (i.e., by skill transfer). One or more landmark memory embeddings may be inserted into the episodic object memory, such as with instructions corresponding to the ranking, weighting, scoring, etc., and / or with references to related data.

[0019] Additionally or alternatively, some aspects of the present disclosure relate to methods, systems, and media for retrieving landmark memories from an episodic object memory. Generally, a user interface can be generated that accepts input corresponding to the degree of the episodic memory. Based on the degree of the episodic memory, an indication of one or more landmark memories can be received from the episodic memory.

[0020] Benefits of the mechanisms disclosed herein may include improved user efficiency for performing actions (e.g., retrieving information, searching virtual documents, generating draft emails, generating draft calendar events, providing content information related to virtual documents, etc.) with a computing device based on content that has been summarized (e.g., by text and / or embeddings) and analyzed based on the summary (e.g., ranked, weighted, scored, and provided to a trained machine learning model). Additionally, the mechanisms disclosed herein for generating episodic content can improve computational efficiency, e.g., by reducing the amount of storage required to track content (e.g., by feature vectors or labels, as opposed to audio / video recordings or relatively large amounts of text). Still further, the mechanisms disclosed herein can improve computational efficiency for receiving content from an episodic object memory, such as by comparing feature vectors, as opposed to searching a relatively large record or repository stored in memory.

[0021] 1 illustrates an example of a system 100 according to some aspects of the disclosed subject matter. System 100 may be a system for storing landmark memories within an episodic object memory. Additionally or alternatively, system 100 may be a system for using an episodic object memory, such as by retrieving information from the episodic object memory. System 100 includes one or more computing devices 102, one or more servers 104, a content data source 106, an input data source 107, and a communication network or networks 108.

[0022] The computing device 102 can receive content data 110 from a content data source 106, which can be, for example, a microphone, a camera, a global positioning system (GPS), etc., that transmits content data, a computer-executable program that generates the content data, and / or a memory having stored thereon data corresponding to the content data. The content data 110 can include visual content data, audio content data (e.g., conversations or ambient noise), gaze content data, calendar entries, emails, document data (e.g., virtual documents), weather data, news data, blog data, encyclopedia data, and / or other types of virtual and / or real content data that can be recognized by those skilled in the art. In some examples, the content data can include text, images, source code, commands, skills, and / or program ratings.

[0023] The computing device 102 may further receive input data 111 from an input data source 107, which may be, for example, a camera, a microphone, a computer-executed program that generates the input data, and / or a memory having data stored therein corresponding to the input data. The content data 111 may be, for example, user input such as a voice query, a text query, an image, an action performed by a user and / or a device, a computer command, a program evaluation, or any other input data that may be recognized by one of ordinary skill in the art.

[0024] Additionally or alternatively, the network 108 may receive content data 110 from a content data source 106. Additionally or alternatively, the network 108 may receive input data 111 from an input data source 107.

[0025] The computing device 102 may include a communication system 112, an episodic object memory insertion engine or component 114, and / or an episodic object memory retrieval engine or component 116. In some examples, the computing device 102 may execute at least a portion of the episodic object memory insertion component 114 to generate a set of embeddings corresponding to one or more subsets of the received content data 110 to be inserted into the episodic object memory. For example, each of the subsets of content data may be provided to a machine learning model, such as a natural language processor and / or a vision processor, to generate the set of embeddings. In some examples, the subsets of content data may be provided to another type of model, such as a generative large-scale language model (LLM).

[0026] Further, in some examples, the computing device 102 can execute at least a portion of the episodic object memory retrieval component 116 to retrieve one or more landmark memories or related information from the episodic object memory based on, for example, an input (e.g., generated based on the input data 111). In some examples, the episodic object memory retrieval component 116 can further determine an action. For example, the action can be determined based on the input and one or more embeddings (e.g., one or more landmark memory embeddings) corresponding to the landmark memories.

[0027] The server 104 may include a communication system 112, an episodic object memory insertion engine or component 114, and / or an episodic object memory retrieval engine or component 116. In some examples, the server 104 may execute at least a portion of the episodic object memory insertion component 114 to generate a set of embeddings corresponding to one or more subsets of the received content data 110 to be inserted into the episodic object memory. For example, each of the subsets of content data may be provided to a machine learning model, such as a natural language processor and / or a vision processor, to generate the set of embeddings. In some examples, the subsets of content data may be provided to another type of model, such as a generative large-scale language model (LLM).

[0028] Further, in some examples, the server 104 can execute at least a portion of the episode object memory retrieval component 116 to retrieve one or more landmark memories or related information from the episode object memory based on, for example, an input (e.g., generated based on the input data 111). In some examples, the episode object memory retrieval component 116 can further determine an action. For example, the action can be determined based on the input and one or more embeddings (e.g., one or more landmark memory embeddings) corresponding to the landmark memory.

[0029] Additionally or alternatively, in some examples, computing device 102 can communicate data received from content data source 106 and / or input data source 107 over communication network 108 to server 104, which can execute at least a portion of episode object memory insertion component 114 and / or episode object memory retrieval engine 116. In some examples, episode object memory insertion component 114 can execute one or more portions of methods / processes 400 and / or 700 described below with respect to Figures 4 and 7, respectively. Further, in some examples, episode object memory retrieval component 116 can execute one or more portions of methods / processes 500 and / or 700 described below with respect to Figures 5 and 7, respectively.

[0030] In some examples, the computing device 102 and / or server 104 may be any suitable computing device or combination of devices, such as a desktop computer, a car computer, a mobile computing device (e.g., a laptop computer, a smartphone, a tablet computer, a wearable computer, etc.), a server computer, a virtual machine executed by a physical computing device, a web server, etc. Further, in some examples, there may be multiple computing devices 102 and / or multiple servers 104. Those skilled in the art should understand that content data 110 and / or input data 111 may be received at one or more of the multiple computing devices 102 and / or one or more of the multiple servers 104, such that the mechanisms described herein can insert entries into and / or use embedded object memory based on aggregation of content data 110 and / or input data 111 received across the computing devices 102 and / or servers 104.

[0031] In some examples, the content data source 106 may be any suitable source of content data (e.g., a microphone, a camera, a GPS, a sensor, etc.). In more specific examples, the content data source 106 may include memory that stores content data (e.g., local memory of the computing device 102, local memory of the server 104, cloud storage, portable memory connected to the computing device 102, portable memory connected to the server 104, etc.). In another more specific example, the content data source 106 may include an application configured to generate content data. In some examples, the content data source 106 may be local to the computing device 102. Additionally or alternatively, the content data source 106 may be remote from the computing device 102 and may communicate the content data 110 to the computing device 102 (and / or the server 104) via a communications network (e.g., communications network 108).

[0032] In some examples, the input data source 107 may be any suitable source of input data (e.g., a microphone, a camera, a sensor, etc.). In more specific examples, the input data source 107 may include memory that stores input data (e.g., local memory of the computing device 102, local memory of the server 104, cloud storage, portable memory connected to the computing device 102, portable memory connected to the server 104, privately accessible memory, publicly accessible memory, etc.). In another more specific example, the input data source 107 may include an application configured to generate the input data. In some examples, the input data source 107 may be local to the computing device 102. Additionally or alternatively, the input data source 107 may be remote from the computing device 102 and may communicate the input data 111 to the computing device 102 (and / or the server 104) via a communications network (e.g., communications network 108).

[0033] In some examples, communication network 108 may be any suitable communication network or combination of communication networks. For example, communication network 108 may include a Wi-Fi network (which may include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, a 5G network, etc. conforming to any suitable standards), a wired network, etc. In some examples, communication network 108 may be a local area network (LAN), a wide area network (WAN), a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Each of the communication links (arrows) shown in FIG. 1 may be any suitable communication link or combination of communication links, such as a wired link, an optical fiber link, a Wi-Fi link, a Bluetooth link, a cellular link, etc.

[0034] 2 illustrates examples of virtual content 200 and real content 250, according to some aspects described herein. As discussed with respect to system 100, the mechanisms described herein may include receiving content data (e.g., content data 110) from a content data source. The content data may be virtual content 200 and / or real content 250.

[0035] Generally, when a user interacts with a computing device (e.g., computing device 102), the user is physically present in reality (e.g., a physical environment) while interacting with a virtual environment. Thus, the contextual information that may be invoked when the user interacts with the computing device may be virtual content (e.g., virtual content 200) and / or real content (e.g., real content 250).

[0036] Virtual content 200 includes virtual person 202, audio content 204, virtual document 206, and / or visual content 208. Virtual person 202 may include data corresponding to a virtual image of an individual generated by a video stream, a still image, a virtual avatar corresponding to the person, etc. Additionally or alternatively, virtual person 202 may include data corresponding to the person, such as an icon corresponding to the person, or other indicator corresponding to a particular person that would be recognizable to one skilled in the art.

[0037] Audio content 204 may include data corresponding to speech data generated within the virtual environment. For example, audio content 204 within the virtual environment may be generated by computing device 102 to correspond to audio received from a user (e.g., when a user speaks into a microphone on a computing device that may be separate from computing device 102). Additionally or alternatively, audio content 202 may correspond to other types of audio data that may be generated within the virtual environment, such as animal sounds, beeps, buzzers, or other types of audio indicators.

[0038] Virtual document 206 may include any type of document found in a virtual environment, such as a text edit document, a presentation, an image, a spreadsheet, an animated series of images, a calendar invitation, an email, a notification, or any other type of virtual document recognized by those skilled in the art.

[0039] Visual content 208 may include data corresponding to graphical content that may be displayed or generated by a computing device. For example, visual content 208 may be content generated via an application (e.g., a web browser, a presentation application, a teleconferencing application, a business management application, etc.) running on computing device 102. Visual content 208 may include data scraped from a screen display of computing device 102. For example, any visual indication displayed on computing device 102 may be included in visual content 208.

[0040] Each of the multiple types of virtual content 200 may be a subset of virtual content 200 that may be received by the mechanisms described herein as a subset of content data 110. Additionally, while specific examples of types of virtual content are discussed above, additional and / or alternative types of virtual content may be recognized by those skilled in the art as relating to virtual components of virtual environments and / or augmented reality environments.

[0041] Reality content 250 includes visual content 252, audio content 254, device used 256, location 258, weather 260, news 262, time 264, people 268, and / or gaze content 270. Visual content 252 may include data received from a camera or optical sensor. For example, visual content 252 may include information about a user's biometric data collected with the user's permission, information about the user's physical environment, information about clothing worn by the user, etc. Additional examples of visual content 252 may be recognized by those skilled in the art.

[0042] Audio content 254 includes audio data originating from a real or physical environment. For example, audio content 254 may include data corresponding to audio data (e.g., speech spoken by a user or speech that can be otherwise converted into text related to a conversation). Additionally or alternatively, audio content 204 may include data corresponding to ambient noise data. For example, audio content 204 may include animal sounds (e.g., dogs barking, cats meowing, etc.), traffic sounds (e.g., airplanes, cars, sirens, etc.), and nature sounds (e.g., waves, wind, etc.). Additional examples of audio content 254 may be recognized by those skilled in the art.

[0043] Device in use 256 may include information about which device the user is using. For example, when a user is attempting to find a virtual document (e.g., virtual document 206), the user may remember that they last accessed the virtual document on a first computing device (e.g., a cell phone) compared to a second computing device (e.g., a laptop). Thus, based on the user's memory of associating a particular device with the element (e.g., a virtual document or computer application) being accessed or discovered, the mechanisms disclosed herein can locate or open the desired element being accessed or discovered.

[0044] Location 258 may include information about the location of the user. For example, location data may be received from a global positioning system (GPS), a satellite positioning system, a cellular positioning system, or other types of location determination systems. A user may associate their physical location with actions performed on a computing device. Thus, for example, using the mechanisms disclosed herein, a user may provide a query such as, "While I was at the beach and videoconferenced with my boss, what documents was I working on?" and the mechanisms disclosed herein may determine one or more documents the user may have been viewing based on information stored regarding which documents the user was working on when they were at the beach and when they were videoconferenced with their boss.

[0045] Weather 260 may include information about the weather surrounding the user. For example, for a given time determined by time content 264, weather information (e.g., precipitation, temperature, humidity, etc.) may be received or otherwise obtained for where the user is located (e.g., based on location content 258). Thus, for example, using the mechanisms disclosed herein, a user may provide a query such as, "Who was I video calling with on my cell phone when it was freezing and snowing outside?" The mechanisms disclosed herein may determine the virtual person (e.g., from virtual person content 202) that the user is trying to remember based on when the user's cell phone was used (e.g., from device used content 256) and when it was freezing and snowing outside (e.g., from weather content 260).

[0046] News 262 may include information about recent news articles that the user may be aware of. For example, for a given time determined by time content 264, relatively recent news articles covering important events may have been published. Thus, for example, using the mechanisms disclosed herein, a user may provide a query such as, "Who was I video calling with on my laptop the day the Chicago basketball team won the national championship?" The mechanisms disclosed herein may determine the virtual person (e.g., from virtual person content 202) that the user is trying to remember based on when the user's laptop (e.g., from device in use content 256) was used and when the Chicago basketball team won the national championship (e.g., from news content 262). Additional or alternative types of news articles may include holidays, birthdays, local events, national events, natural disasters, celebrity updates, scientific discoveries, sports updates, or any other type of news recognized by those skilled in the art.

[0047] Time 264 may include one or more timestamps at which various content types are received. For example, each content type (e.g., virtual content 200 and / or real content 250, or a particular content type therein) may be time-stamped when it is obtained or received. Generally, the mechanisms described herein may be temporal in that actions performed by a computing device based on multiple stored content types may depend on the timestamps assigned to each of the content types.

[0048] Person 268 may include information about a person in the physical environment surrounding one or more computing devices (e.g., computing device 102). A person may be identified using biometric recognition (e.g., facial recognition, voice recognition, fingerprint recognition, etc.) with the user's permission, after receiving and storing any relevant biometric data. Additionally or alternatively, a person may be identified by engaging in performing a particular action on a computing device (e.g., engaging in particular software). Further, some people may be identified by logging into one or more computing devices. For example, a person may be the owner of a computing device, and the computing device may be linked to that person (e.g., by a passcode, biometric entry, etc.). Thus, when a computing device is logged in, the person is identified thereby. Similarly, a person may be identified by logging into a particular application (e.g., by a passcode, biometric entry, etc.). Thus, when a particular application is logged in, the person is identified thereby. Additionally or alternatively, one or more people may be identified using a radio frequency identification tag (RFID), an ID badge, a barcode, a QR code, or any other means of identification capable of identifying a person through some technological interface. Additionally or alternatively, two or more people may be identified using the mechanisms described herein based on which the proximity of particular people to each other may be identified.

[0049] Gaze content 270 may include information corresponding to where a user is looking on one or more computing devices. Thus, if a first user is in a video conference call and, while viewing a document on their computing device, communicates to a second user that they will send the second user a document they discussed, the mechanisms described herein can generate a draft email for sending the draft document from the first user to the second user based on the audio content received from the video conference call, the received gaze content, and an understanding of who the second user is, such that an email is drafted to be sent to the current person. Similarly, the draft document can be saved to a user's clipboard so that it can be pasted into an email, message, or other form of virtual communication.

[0050] Additional types of virtual content 200 and / or real content 250 may be recognized by those skilled in the art. Furthermore, while certain content types have been illustrated and described as types of virtual content 200 (e.g., virtual person 202, audio content 204, virtual document 206, and visual content 208), while other content types have been illustrated and described as real content (e.g., visual content 252, audio content 254, device used 256, location 258, weather 260, news 262, time 264, person 268, and gaze content 270), it should be understood that in some instances, such categorizations of virtual content 200 and real content 250 are interchangeable with the types of content disclosed herein (e.g., where location can refer to a virtual location, such as a virtual meeting space, as well as a physical location for a device being used to access the virtual meeting space), while in other instances, such categorizations of virtual content 200 and real content 250 are fixed (e.g., where location refers only to a physical location, such as a geographic coordinate location in the physical world).

[0051] The various content types generally discussed with respect to FIG. 2 provide various contexts with which a user associates information that the user is attempting to recall and / or use to perform an action with a computing device. In some examples, a user may provide real content (e.g., real content 250) information to receive information related to the virtual content (e.g., virtual content 200) from a computing device or to perform an action related to the virtual content (e.g., virtual content 200) on a computing device. Additionally or alternatively, a user may provide virtual content (e.g., virtual content 200) information to receive information related to real or physical content (e.g., real content 250) from a computing device. Additionally or alternatively, a user may provide a combination of virtual and real content to receive information corresponding to the associated real content, receive information corresponding to the virtual content, and / or perform an action related to the virtual content with a computing device.

[0052] Those skilled in the art will recognize various contexts in which examples of virtual content 200 and real content 250 may be collected. For example, content 200, 250 may be collected for business operations, coaching history, medical experiences, project management, event planning and execution. In healthcare, for example, landmark memories may correspond to content encoded based on landmark medical experiences, events, and / or content over time (e.g., the onset of major illnesses, notable test and lab results, diagnoses, hospitalizations, surgeries, doctor visits, family health challenges, etc.). These landmark medical items may be stored as landmark memories using the mechanisms provided herein and utilized in interactions to connect doctors or patients with artificial intelligence (AI) systems that incorporate aspects of the systems provided herein.

[0053] In general, users may be more likely to remember atypical content information than typical content information. Thus, using the mechanisms described herein, it may be determined that atypical data is associated with a landmark memory. Atypical data is data that is irregular or unusual given a category into which the data may be classified. For example, a piece of data may be determined to be atypical if it occurs less than 1% of the time within a category of data into which the piece of data may be classified, such as a category related to people, events, locations, etc. In alternative examples, the 1% threshold may be 5%, or 10%, or 20%, or any value or range therebetween.

[0054] 3 illustrates an example flow 300 for storing an entry in an episode object memory. The example flow begins with receiving a content item 302. The content item 302 may be received from a content data source, such as content data source 110 described herein above with respect to FIG. 1. Furthermore, the content item 302 may include virtual content and / or real content, such as virtual content 200 and / or real content 250 described herein above with respect to FIG. 2. The content item 302 may be one or more content items or objects, each including one or more content data.

[0055] The content data of the content item 302 can be input into the model 304. For example, the model 304 can include a machine learning model, such as a machine learning model trained to generate one or more embeddings based on the received content data. In some examples, the model 304 includes a generative model, such as a generative large-scale language model (LLM). In some examples, the model 304 includes a natural language processor. Additionally or alternatively, in some examples, the model 304 includes a vision processor. The model 304 can be trained with one or more datasets compiled by an individual and / or a system. Additionally or alternatively, the model 304 can be trained based on a dataset that includes information obtained from the internet. Furthermore, the model 304 can include versions. One skilled in the art will appreciate that any type of model can be employed as part of the aspects disclosed herein without departing from the scope of the present disclosure.

[0056] The model 304 can output a set of embeddings 306. The set of embeddings 306 can include one or more embeddings (e.g., one or more semantic embeddings). The set of embeddings 306 can be specific to the model 304. Additionally or alternatively, the set of embeddings 306 can be specific to a particular version of the model 304. Each embedding in the set of embeddings 306 can correspond to an individual content object and / or individual content data from the content item 302. Thus, for each object and / or content data in the content item 302, the model 304 can generate an embedding.

[0057] Each embedding in the set of embeddings 306 may be associated with an individual instruction (e.g., a byte address) that corresponds to the location of the source data associated with the embedding. For example, if the content item 302 corresponds to a set of emails, the set of embeddings 306 may be hash values ​​generated for each of the emails based on the content within the emails (e.g., the abstract meaning of words or images contained within the emails). An instruction corresponding to the location (e.g., in memory) of the emails may then be associated with the embedding generated based on the emails. In this regard, the actual source data of the content object may be stored separately from the corresponding embedding, which may occupy less memory. Specific examples, additional, and / or alternative examples of where the content item 302 corresponds to a set of emails will be recognized by those skilled in the art, at least in light of the teachings provided herein.

[0058] The embeddings 306 may be provided to a ranker engine 308. The ranker engine 308 may rank, score, and / or assign weights to one or more of the embeddings 306. The ranking and / or weighting may be based on the dissimilarity of one or more of the embeddings from the set of embeddings 306 to one or more of the other embeddings from the set of embeddings 306. One or more of the ranked, scored, and / or weighted embeddings may correspond to a landmark memory. One or more of the ranked, scored, and / or weighted embeddings corresponding to a landmark memory may be determined to be a landmark memory embedding.

[0059] In general, users may be more likely to remember atypical content information than typical content information. Therefore, using the mechanisms described herein, it may be determined that atypical data is associated with and / or labeled as a landmark memory. Based on determining the landmark memory, data related to the landmark memory (e.g., related by time, location, person, etc.) may also be determined, such that the landmark memory may become a pivot point for searching for data related thereto. Atypical data is irregular or abnormal data. For example, a piece of data that occurs less than 1% of the time within a category of data into which the piece of data may be classified, such as a person, event, location, email, etc., may be determined to be atypical data. In some examples, the 1% threshold may instead be 5%, or 10%, or 20%, or any value or range therebetween.

[0060] One or more embeddings that have been ranked, scored, and / or weighted (e.g., with respect to the degree to which they are landmark memory embeddings) may be inserted into the episode object memory 310. Additionally or alternatively, a subset of the ranked, scored, and / or weighted embeddings may be inserted into the episode object memory 310 based on comparing the ranking, scoring, and / or weighting to a predetermined threshold, thereby determining which of the embeddings are landmark memory embeddings. One or more landmark memory embeddings may be stored in the episode object memory 310 along with an indication of their corresponding ranking, scoring, and / or weighting. Additionally, a landmark memory embedding may be associated with references (e.g., pointers) to related data related to the one or more landmark memory embeddings (e.g., embeddings similar to the landmark memory embedding, source data related to the landmark memory embedding). A landmark memory as described herein may include a set of properties that define a schema. The set of properties may include a summary of the landmark memory (e.g., a natural language summary of the landmark memory) and / or a reference (e.g., pointer) to data related to the landmark memory.

[0061] In some examples, episode object memory 310 stores only landmark memory embeddings. Additionally or alternatively, in some examples, episode object memory 310 stores embeddings that are landmark memory embeddings and embeddings that are not landmark memory embeddings, such as along with an indication corresponding to which embeddings are landmark memory embeddings and which embeddings are not landmark memory embeddings.

[0062] 4 illustrates an example vector space 400 according to some aspects described herein. The vector space 400 includes a plurality of feature vectors, such as a first feature vector 402, a second feature vector 404, a third feature vector 406, a fourth feature vector 408, and a fifth feature vector 410. Each of the plurality of feature vectors 402, 404, 406, and 408 corresponds to an individual embedding 403, 405, 407, and 409 generated based on a plurality of subsets of content data (e.g., subsets of content data 110, virtual content 200, and / or real content 250). The embeddings 403, 405, 407, and 409 may be semantic embeddings. The fifth feature vector 410 is generated based on an input embedding 411 (e.g., an embedding that is ranked or weighted relative to other embeddings). The input embedding may be generated based on the content data. Alternatively, the input embedding may be provided separately (e.g., based on user input).

[0063] The feature vectors 402, 404, 406, 408, and 410 each have a measurable distance from one another. For example, the distance between the feature vectors 402, 404, 406, and 408 and the fifth feature vector 410 corresponding to the input embedding 411 may be measured using cosine similarity. Alternatively, the distance between the feature vectors 402, 404, 406, and 408 and the fifth feature vector 410 may be measured using another distance measurement technique (e.g., an n-dimensional distance function) that would be recognized by one skilled in the art.

[0064] The similarity of each of the feature vectors 402, 404, 406, 408 to the feature vector 410 corresponding to the input embedding 411 may be determined based on, for example, a distance measurement between the feature vectors 402, 404, 406, 408 and the feature vector 410. The similarity between the feature vectors 402, 404, 406, 408 and the feature vector 410 may be used to rank, weight, group, or cluster the feature vectors 402, 404, 406, and 408 (e.g., based on how dissimilar one vector is to another).

[0065] Embeddings 403 and 405, corresponding to feature vectors 402 and 404, respectively, may be included in the same content group. For example, embedding 403 may relate to a first email, and embedding 405 may relate to a second email. Additional and / or alternative examples of content groups into which embeddings may be classified may be recognized by those skilled in the art.

[0066] An exemplary cluster of embeddings, including one or more embeddings selected from embeddings 403, 405, 407, 409, and 411, may be a cluster of embeddings corresponding to a person, a geographic location, an image, and / or other modalities, such as olfaction, touch, etc. In some examples, the embeddings described herein include emotions (e.g., photos, videos, etc.), cognitive states, etc. Meta-information about the content corresponding to the embeddings disclosed herein may be established. For example, the meta-information may include patterns of output, number of queries from users, etc. In some examples, the embeddings may be filtered based on the salience of the memories to which the embeddings correspond. Thus, the importance or salience of one or more embeddings may be determined using the mechanisms provided herein based on the memories to which the one or more embeddings correspond.

[0067] One or more of the embeddings 403, 405, 407, 409, and 411 may be stored in a data structure such as an ANN tree, a kd-tree, an octree, another n-dimensional tree, or another data structure that would be recognized by one skilled in the art that is capable of storing vector space representations. Additionally, the memory corresponding to the data structure in which one or more of the embeddings 403, 405, 407, 409, and 11 are stored may be arranged or stored within the data structure in a way that groups the embeddings together, such as to improve the efficiency of subsequent search processes.

[0068] In some examples, feature vectors and their corresponding embeddings generated according to the mechanisms described herein may be stored indefinitely. Additionally or alternatively, in some examples, as new feature vectors and / or embeddings are generated and stored, the new feature vectors and / or embeddings may overwrite older feature vectors and / or embeddings stored in memory (e.g., based on metadata in the embeddings indicating their version), such as to improve memory capacity. Additionally or alternatively, in some examples, feature vectors and / or embeddings may be deleted from memory at specified time intervals to improve memory capacity and / or based on the amount of available memory (e.g., in embedded object memory 310).

[0069] Generally, the ability to store embeddings corresponding to received content data allows users to associate and locate data in novel ways that have the advantage of being computationally efficient. For example, instead of storing a video recording of a computing device's screen or a web page on the Internet, a user can instead store an embedding corresponding to the content object using the mechanisms described herein. The embedding may be a hash, as opposed to, for example, a video recording, which may be hundreds of thousands of pixels per frame. Thus, the mechanisms described herein are efficient not only for reducing the use of processing resources to search for stored content, but also for reducing memory usage. Additional and / or alternative advantages may be recognized by those skilled in the art.

[0070] 5A illustrates an example method 500 for storing landmark memories (e.g., landmark memory embeddings and / or landmark memory text) within an episode object memory according to some aspects described herein. In examples, aspects of method 500 are performed by devices such as computing device 102 and / or server 104 described above with respect to FIG. 1.

[0071] Method 500 begins at operation 502 of receiving one or more content items or objects. The content items may have one or more content data. The one or more content data may include at least one real content data and / or at least one virtual content data. The content data may be similar to content data 110 discussed with respect to FIG. 1. Additionally or alternatively, the content data may be similar to virtual content data 200 and / or real content data 250 discussed with respect to FIG. 2.

[0072] The content data may include one or more of audio content data, visual content data, gaze content data, calendar content data, email content data, virtual documents, data generated by a particular software application, weather content data, news content data, encyclopedia content data, location content data, and / or blog content data. In some examples, the content data may include at least one of skills, commands, or program ratings. Additional and / or alternative types of content data may be recognized by those skilled in the art.

[0073] At operation 504, it is determined whether at least one of the subsets of content data has an associated embedding model (e.g., a semantic embedding model). For example, the embedding model may be similar to model 304 discussed with respect to FIG. 3. The embedding model may be trained to generate one or more embeddings based on at least one of the subsets of content data. In some examples, the embedding model may include a natural language processor. In some examples, the embedding model may include a vision processor. In some examples, the embedding model may include a machine learning model. Still further, in some examples, the embedding model may include a generative large-scale language model. The embedding model may be trained on one or more datasets compiled by an individual and / or a system. Additionally or alternatively, the embedding model may be trained based on a dataset including information obtained from the internet.

[0074] If it is determined that at least one of the content data does not have an associated embedding model, flow branches "NO" to operation 506, where a default action is performed. For example, the content data and / or content item may have an associated pre-configured action. In other examples, method 500 may include determining whether the content data and / or content item has an associated default action, such that in some cases, no action may be performed as a result of the received content item. Method 500 may end at operation 506. Alternatively, method 500 may return to operation 502, providing an iterative loop that receives one or more content items and determines whether at least one of the content data of the content items has an associated embedding model.

[0075] However, if it is determined that at least one of the content data has an embedding model, the flow instead branches "yes" to operation 508, where one of the content data associated with the content item is provided to one or more embedding models. The embedding model generates one or more embeddings. Furthermore, the one or more embedding models may include versions. Each of the embeddings generated by an individual embedding model may include metadata corresponding to the version of the embedding model that generated the embedding.

[0076] In some examples, content data is provided locally to one or more embedded models. In some examples, content data is provided to one or more embedded models by an application programming interface (API). For example, a first device (e.g., computing device 102 and / or server 104) can interface with a second device (e.g., computing device 102 and / or server 104) by an API configured to embed or provide the content data or an indication thereof in response to receiving the content data or an indication thereof.

[0077] At operation 510, one or more embeddings are received from one or more of the embedding models. In some examples, the one or more embeddings are a set of embeddings or multiple embeddings. For example, a set of embeddings (such as set of embeddings 306) may be associated with a first embedding model (e.g., model 304). In some examples, a set of embeddings may be uniquely associated with a first embedding model, such that a second embedding model produces a different set of embeddings than the first embedding model. Furthermore, in some examples, a set of embeddings may be uniquely associated with a version of the first embedding model, such that different versions of the first embedding model produce different sets of embeddings.

[0078] The set of embeddings may include embeddings generated by the first embedding model for at least one of the plurality of content data. For example, the content data may correspond to an email, an audio file, a message, a website page, etc. Additional and / or alternative types of content objects or items corresponding to the content data may be recognized by those skilled in the art.

[0079] At operation 512, one or more of the embeddings in the set of embeddings are determined to be landmark memory embeddings. In some examples, one or more of the embeddings are ranked, weighted, and / or scored to determine whether they are landmark memory embeddings. The ranking, weighting, and / or scoring may be based on a dissimilarity of the embedding to at least one of the other embeddings in the set of embeddings. In some examples, the dissimilarity may be calculated based on a distance formula (e.g., in a vector space such as vector space 400), a text comparison, a visual comparison, and / or other comparison techniques that may be recognized by one of ordinary skill in the art.

[0080] In some examples, the determining in operation 512 includes obtaining a machine learning model previously trained to identify landmark memories. The set of embeddings may be provided to the machine learning model. A score corresponding to the degree to which the embeddings are landmark memory embeddings may be received from the machine learning model. The score may be used to rank and / or weight the embeddings (e.g., using mechanisms described herein). Additionally or alternatively, the score may be compared to a threshold to identify one or more embeddings in the set of embeddings as landmark memory embeddings.

[0081] In general, users may be more likely to remember atypical content information than typical content information. Therefore, using the mechanisms described herein, embeddings generated based on atypical data may be determined and / or labeled as landmark memory embeddings. Based on determining the landmark embedding memory, data related to the landmark embedding memory (e.g., based on time, location, person, or other content data corresponding to the landmark embedding memory) may also be determined, such that the landmark memory embedding may become a pivot point for searching for data related thereto. Atypical data is irregular or abnormal data. For example, a piece of data that occurs less than 1% of the time within a category of data into which the piece of data may be classified, such as a person, event, location, email, etc., may be determined to be atypical data. In some examples, the 1% threshold may instead be 5%, or 10%, or 20%, or any value or range therebetween.

[0082] Some examples include accepting input (e.g., user input) corresponding to the degree or threshold of episodic memory. The input can be accepted from any of a variety of inputs, such as a button, gaze input, voice input, text input, gesture input, touchpad, mouse, slider, etc. Additionally, the user input can be accepted from a graphical user interface. In some examples, the ranking, weighting, and / or score of the embeddings can be compared to an episodic memory threshold to determine which of one or more embeddings from a set of embeddings is a landmark memory embedding. For example, if the ranking of a particular embedding is higher than the episodic memory threshold, the particular embedding may be a landmark memory embedding. Alternatively, in some examples, if the ranking of a particular embedding is lower than the degree of episodic memory, the particular embedding may be a landmark memory embedding. Additional and / or alternative examples for comparing a landmark memory embedding metric to a threshold may be recognized by those skilled in the art.

[0083] At operation 514, one or more landmark memory embeddings are inserted into the episodic object memory. In some examples, the episodic object memory may be an index, a database, a tree (e.g., an ANN tree, a kd tree, etc.), or another type of memory that would be recognized by one of ordinary skill in the art. The episodic object memory includes one or more embeddings (e.g., landmark memory embeddings) from a collection of embeddings. Furthermore, one or more embeddings may be associated with individual instructions corresponding to the location of source data associated with the one or more embeddings. The source data may include one or more of an audio file, a text file, an image file, a video file, and / or a website page. Additional and / or alternative types of source data may be recognized by one of ordinary skill in the art. Embeddings may occupy relatively little memory. For example, each embedding may be a 64-bit hash. In comparison, source data may occupy relatively large amounts of memory. Thus, by storing instructions to the location of source data in the embedded object memory, as opposed to the raw source data, the mechanisms described herein may be relatively efficient with respect to memory usage of one or more computing devices.

[0084] The indication corresponding to the location of the source data may be a byte address, a uniform resource indicator (e.g., a uniform resource link), or another form of data capable of identifying the location of the source data. Furthermore, the episode object memory may be stored in a location different from the location of the source data. Additionally or alternatively, the source data may be stored in memory of the computing device or server on which the episode object memory (e.g., episode object memory 310) is located. Additionally or alternatively, the source data may be stored in memory remote from the computing device or server on which the episode object memory is located.

[0085] In some examples, a subset of the ranked embeddings may be inserted into the episode object memory based on comparing their rankings (or weights) to a predetermined threshold. One or more ranked embeddings may be stored in the episode object memory along with an indication of their corresponding rankings (or weights). Additionally, the ranked embeddings may be associated with references (e.g., pointers) to related data related to the one or more ranked embeddings (e.g., embeddings similar to the ranked embeddings, source data related to the ranked embeddings). The landmark memory embeddings described herein may include a set of properties that define a schema. The set of properties may include a summary of the landmark memory embedding (e.g., a natural language summary of the landmarks) and / or references (e.g., pointers) to data related to the landmark memory embedding.

[0086] In some examples, the episode object memory stores only landmark memory embeddings. Additionally or alternatively, in some examples, the episode object memory stores embeddings that are landmark memory embeddings and embeddings that are not landmark memory embeddings, along with corresponding indications of which embeddings are landmark memory embeddings and which are not landmark memory embeddings, etc.

[0087] In some examples, the insertion in operation 514 triggers a spatial storage operation to store a vector representation of one or more landmark memory embeddings. The vector representation may be stored in a data structure such as an ANN tree, a kd tree, an n-dimensional tree, an octree, or another data structure that may be recognized by one of ordinary skill in the art in light of the teachings herein. Additional and / or alternative types of storage mechanisms that may store the vector spatial representation may be recognized by one of ordinary skill in the art.

[0088] At operation 514, an episode object memory is provided. For example, the episode object memory may be provided as an output for further processing. In some examples, the episode object memory may be used to obtain information and / or generate actions, as described in some examples in further detail herein.

[0089] In general, embeddings can be used as quantitative representations of abstract meanings perceived from content data. Thus, different embedding models may generate different embeddings for the same received content data, depending on how they are trained or configured (e.g., based on different datasets and / or training methods). In some examples, some embedding models may be trained to provide a relatively broader interpretation of the received content than other embedding models trained to provide a relatively narrower interpretation of the content data. Thus, such embedding models may generate different embeddings.

[0090] Method 500 may end at operation 516. Alternatively, method 400 may return to operation 502 (or any other operation from method 500) to provide an iterative loop, such as receiving one or more content items and inserting one or more landmark memory embeddings in the episode object memory based thereon.

[0091] The operations provided in method 500 provide the ability to selectively formulate, store, and later retrieve (e.g., see method 700 discussed with respect to FIG. 7) landmark memories drawn from multiple experienced content as a means to provide humans and machines (given their limited memory and processing capabilities) with efficient options for storing and subsequently having powerful handles or indexes that enable constrained reasoning and recall systems to efficiently retrieve context-dependent, relevant content for use in real time (e.g., to maximize human or machine success). Such mechanisms may be particularly important for cognitive systems (whether human or machine) that continue to accumulate and leverage experience over a continuous, lifelong period of engagement with a complex world.

[0092] 5B illustrates a detailed example of operation 512 according to some aspects described herein. For example, a detailed example of operation 512 may be a method for determining that one or more embeddings of a set of embeddings are landmark memory embeddings. Although operations 550, 552, 554, and 556 described below are discussed as sub-operations relative to operation 512 of method 500, it should be recognized by those skilled in the art that operations 550, 552, 554, and / or 556 may be performed independently of method 500 and / or in conjunction with methods other than method 500.

[0093] At operation 550, a machine learning model that has been previously trained to identify landmark memories is obtained. For example, the machine learning model may be trained based on a dataset of embeddings and a metric that corresponds to the degree to which a particular embedding from the dataset of embeddings is a landmark memory.

[0094] At operation 552, the machine learning model is provided with a set of embeddings. The set of embeddings may include the embeddings received at operation 510 of FIG. 5A, and additionally or alternatively, the set of embeddings may be a set of embeddings otherwise received using mechanisms recognized by those skilled in the art.

[0095] At operation 554, a metric is received from the machine learning model. The metric corresponds to the extent to which one or more embeddings from the collection of embeddings are landmark memory embeddings. The metric may be a score, weight, and / or ranking for each of the one or more embeddings.

[0096] At operation 556, the metric may be compared to a threshold to identify one or more embeddings in the set of embeddings as landmark memory embeddings. For example, the threshold may be provided by user input and / or otherwise pre-configured. In some examples, the metric must be greater than the threshold. In some examples, the metric must be less than the threshold. In some examples, the metric must be equal to the threshold. Those skilled in the art will recognize that a combination of equality constraints may be used when comparing the metric to a threshold to identify one or more embeddings in the set of embeddings as landmark memory embeddings. In some examples, further processing may be performed on the metric, such as using a derivative of the metric to identify one or more embeddings in the set of embeddings as landmark memory embeddings.

[0097] Identifying, storing, and manipulating landmark memories corresponding to experiences, learning, and / or other content can be particularly important when two or more agents (such as humans and assistive AI systems) collaborate over time to efficiently base references to shared experiences. Long-term or lifelong agents coordinating or communicating over time on a project may need efficient mechanisms for storing, recognizing, co-referencing, and / or retrieving experiences, learning, and / or other content accumulated over the course of ongoing, long-term (e.g., lifelong) engagement in a complex world. Examples of content that landmark memories may correspond to are discussed earlier in this specification with respect to FIG. 2.

[0098] As an example, humans may conceivably be able to refer to events they naturally encode and capture as landmark memories in episodic memory when working with an AI system, with the same efficiency of communication with the machine as when working with a human collaborator. To do this successfully, the machine may have a representation of the nature of how humans encode landmark memories, allowing for machine-to-human and / or human-to-machine co-referencing and / or co-basing. Thus, part of striving for mutual grounding between machines and humans in long-term, effective collaboration may be for the machine to learn to understand and encode landmark memories the way humans encode them (e.g., in episodic object memory 310). This understanding may enable more deeply rooted interaction and problem-solving, particularly in long-term human-AI collaboration settings where AI systems are dedicated to supporting specific users and can engage in life-sharing sensing, problem-solving, and more broadly, life experiences with those users.

[0099] Operation 512 may end at operation 556. Alternatively, operation 512 may continue from operation 556 to operation 514 so that the episode object memory may be updated based on the determination of operation 512. In some examples, operation 512 may continue from operation 556 to a different method and / or process that may be recognized by one of ordinary skill in the art in light of at least the teachings provided herein.

[0100] 6A and 6B illustrate an exemplary system 600 for retrieving information from an episode object memory. The exemplary system 600 includes a computing device 602. The computing device 602 may be similar to the computing device 102 described herein above with respect to FIG. 1. The system 600 further includes a user interface 604. The user interface 604 may be a graphical user interface (GUI) displayed on a screen of the computing device 602. The user interface 604 may be configured to accept input from a user.

[0101] For example, GUI 604 includes a slider 606 that can be adjusted based on user input. The input can correspond to a degree or threshold of episodic memory. For example, in Figure 6A, slider 606 is shown in a first position corresponding to a first degree or threshold of episodic memory. In comparison, in Figure 6B, the slider is shown in a second position corresponding to a second degree or threshold of episodic memory.

[0102] 6A , one or more first indications 608 of one or more landmark memories may be received (e.g., based on the first extent of the episodic memory). In some examples, one or more second indications 610 of stored data not corresponding to a landmark memory may also be received. The one or more first indications 608 and / or the one or more second indications 610 may be presented (e.g., on a display screen of a computing device). In some examples, the one or more first indications 608 and / or the one or more second indications 610 may be symbols, text, images, sounds, and / or expressions related to the data to which the indications 608, 610 correspond. Furthermore, in some examples, the one or more first indications 608 and / or the one or more second indications 610 may comprise temporal information corresponding to when source data associated with the indications 608, 610 was generated and / or stored (e.g., when an email was sent, when a phone call was made, when a person was at a given location, etc.).

[0103] 6B, the number of one or more first indications 608 and / or second indications 610 may be different than when the slider 606 is at the first degree of the episodic memory. Thus, the number of content data determined to be landmark memories may change based on when the designated degree of the episodic memory changes.

[0104] The one or more first instructions 608 and / or second instructions 610 may be received from an episodic memory, such as the episodic memory 310 described with respect to Figure 3. For example, the episodic memory may be generated using content data that has been embedded, ranked, and / or weighted based on semantic context, etc. The content data may be ranked based on dissimilarity between the content data, such as by distance formulas, text comparisons, visual comparisons, and / or other comparison techniques that may be recognized by those skilled in the art.

[0105] Memories can be organized to provide an efficient handle within a large amount of stored content that can be efficiently utilized by being pointed to by a memory (e.g., pointed to by a first instruction 608) that is a landmark memory. Such a rich memory (e.g., a landmark memory) can be progressively refined by the less efficient process of drilling down into larger memory encodings, such as by changing a slider 606. In this role, limited retrieval and inference systems can operate efficiently at the level of the landmark memory and use the landmark memory in identifying relevant content and in making decisions to perform laborious drilling down into the less structured, more difficult-to-retrieve details to which the landmark memory points. In this regard, by way of example, an instruction 608 corresponding to a landmark memory can include information corresponding to a relationship with a second instruction 610 that does not correspond to the landmark memory, so that a desired memory associated with the landmark memory can be retrieved or used to perform an action.

[0106] 7 illustrates an exemplary method for retrieving information, such as a landmark memory, from an episode object memory according to some aspects described herein. In an example, aspects of method 700 are performed by devices such as computing device 102 and / or server 104 described above with respect to FIG.

[0107] Method 700 begins at operation 702, which generates a user interface (such as user interface 602). The user interface may be a graphical user interface (GUI) displayed on a screen of a computing device (e.g., computing device 602). Alternatively, the user interface may be a touchpad, mechanical buttons, a camera, a microphone, or any other interface configured to accept user input.

[0108] At operation 704, an input is accepted by a user interface. The input corresponds to a degree of episodic memory. For example, the user interface can include a slider, and the input can correspond to a position or movement of the slider. The position or movement of the slider can be related to the degree of episodic memory. In other examples, the input can be a voice command, text, an automated skill, a gaze command, a gesture, etc. related to the degree of episodic memory.

[0109] At operation 706, it is determined whether there are search results from the episodic object memory that correspond to the input. If it is determined that there are no search results from the episodic object memory that correspond to the input, the flow branches "NO" to operation 708, where a default action is performed. For example, the input and / or the episodic object memory may have a pre-configured action associated with it. In other examples, method 700 may include determining whether the input and / or the episodic object memory have a default action associated with it, such that in some examples, no action may be performed as a result of the received query. Method 700 may end at operation 708. Alternatively, method 700 may return to operations 702 or 704 to provide an iterative loop.

[0110] However, if it is determined that there are search results from the episodic object memory that correspond to the input, then the flow instead branches "yes" to operation 710, where an indication of one or more landmark memories is received based on the extent of the episodic object memory. For example, the indication of the one or more landmark memories may be similar to first display 608 discussed with respect to FIG. 6. Additionally or alternatively, the indication may be provided as an output, such as a report, which may be used for further processing.

[0111] In some examples, after receiving an indication of one or more landmark memories, data associated with the landmark memories may be searched for using the mechanisms provided herein. For example, if a particular landmark memory is associated with content data at a particular moment in time, other content data from within a specified time range of that particular landmark memory may be searched for (e.g., based on a query that may be provided by a user). Generally, once a landmark memory is searched, that landmark memory may serve as a pivot point based on which searches may be made to other data referenced in storage locations of that landmark memory and / or other data otherwise determined to be associated with the landmark memory.

[0112] In some examples, the episode object memory is generated using content data that has been embedded, ranked, weighted, and / or scored based on semantic context, etc. Embeddings generated based on the content data may be ranked, weighted, and / or scored based on the dissimilarity of the embedding to at least one other embedding. The content data may include one or more of audio content data, visual content data, gaze content data, weather content data, news content data, calendar content data, email content data, or location content data. Furthermore, the episode object memory may be stored in a location different from the location of the source data corresponding to the content item.

[0113] In some examples, the input is a first input, the degree of episodic memory is a first degree of episodic memory, and the instruction is a first instruction, and method 700 further includes accepting a second input corresponding to a second degree of episodic memory. Thereafter, based on the second degree of episodic memory, second instructions of one or more landmark memories corresponding to the second input can be received from the episodic object memory.

[0114] In some examples, landmark memories provide a system with the ability to interact (e.g., with a human or another system) in a conversational chat setting that helps narrow down which landmark memory of multiple landmark memories the human or system is referring to. In some examples, dialogue can help clarify ambiguity in obtaining information related to a landmark memory. For example, when asked about a restaurant experience, the system can provide a prompt stating, "Are you talking about the last time you went to Restaurant X in September, or the time you went with Person A and Person B?" Thus, the systems provided herein can request feedback to obtain information to determine which landmark memory (e.g., within an episodic object memory) it intends to explore. In some examples, a mutually shared handle of a library of salient landmark memories enables interaction with an AI system that has access to the agent's (e.g., human or system's) own model of episodic memory. In such examples, the same language can be used to refer to an agent's episodic object memory that is shared (at least partially or largely) with the AI ​​system.

[0115] Method 700 may end at operation 710. Alternatively, method 700 may return to operation 702 (or any other operation from method 700) to provide an iterative loop, such as generating a user interface, receiving a query having information corresponding to at least two different content types, receiving search results corresponding to the query from an episode object memory, etc.

[0116] 8A and 8B illustrate an overview of an exemplary generative machine learning model that may be used in accordance with aspects described herein. Referring initially to FIG. 8A, a conceptual diagram 800 illustrates an overview of a pre-trained generative model package 804 that processes an input 802 and generates a model output for storing entries (e.g., landmark memories) in an episodic object memory 806 and / or retrieving information (e.g., landmark memories) from the episodic object memory 806 in accordance with aspects described herein. Examples of pre-trained generative model packages 604 include, but are not limited to, the Megatron-Turing Natural Language Generation Model (MT-NLG), Generative Pre-trained Transformer 3 (GPT-3), Generative Pre-trained Transformer 4 (GPT-4), BigScience BLOOM (a large-scale open science open access multilingual model), DALL-E, DALL-E 2, Stable Diffusion, or Jukebox.

[0117] In examples, the generative model package 804 is pre-trained according to various inputs (e.g., various human languages, various programming languages, and / or various content types) and thus does not need to be fine-tuned or trained for a particular scenario. Rather, the generative model package 804 may be more generally pre-trained such that the input 802 includes prompts that are generated, selected, or engineered to guide the generative model package 804 to produce a certain generative model output 806. For example, the prompts may include a context and / or one or more complementary prefixes, thereby pre-loading the generative model package 804 accordingly. As a result, the generative model package 804 is guided to generate a prompt-based output that includes a predicted sequence of tokens related to the prompt (e.g., up to the token limit of the generative model package 804). In examples, the predicted sequence of tokens is further processed (e.g., by output decoding 816) to produce the output 806. For example, each token is processed to identify a corresponding word, word fragment, or other content that forms at least a portion of the output 806. It will be appreciated that the input 802 and the generative model output 806 may each include any of a variety of content types, including, but not limited to, text output, image output, audio output, video output, programmatic output, and / or binary output, among other examples. In an example, the input 802 and the generative model output 806 may have different content types, such as when the generative model package 804 includes a generative multimodal machine learning model.

[0118] In this manner, generative models package 804 may be used in any of a variety of scenarios, and further, a different generative models package may be substituted for generative models package 804 without significantly modifying other aspects involved (e.g., similar to those described herein with respect to Figures 1-7). Generative models package 804 thus operates as a tool upon which machine learning processes are performed, and certain inputs 802 to generative models package 804 are programmatically generated or determined, thereby causing generative models package 804 to produce model outputs 806 that can later be used for further processing.

[0119] The generative models package 804 may be provided or used according to any of a variety of paradigms. For example, the generative models package 804 may be used locally to a computing device (e.g., computing device 102 of FIG. 1 ) or may be accessed remotely from a machine learning service. In other examples, aspects of the generative models package 804 are distributed across multiple computing devices. In some examples, the generative models package 804 is accessible via an application programming interface (API), which may be provided by the operating system of the computing device and / or by the machine learning service, among other examples.

[0120] Referring now to the illustrated aspects of generative models package 804, generative models package 804 includes input tokenization 808, input embeddings 810, model layer 812, output layer 814, and output decoding 816. In the example, input tokenization 808 processes input 802 to generate input embeddings 810, which include a sequence of symbolic representations corresponding to input 802. Input embeddings 810 are then processed by model layer 812, output layer 814, and output decoding 816 to result in model output 806. An exemplary architecture corresponding to generative models package 804 is shown in FIG. 8B and is discussed in more detail below. Even so, it will be understood that the architectures shown and described herein should not be construed in a limiting sense, and that any of a variety of other architectures may be used in other examples.

[0121] 8B is a conceptual diagram illustrating an example architecture 850 of a pre-trained generative machine learning model that may be used in accordance with aspects described herein. As noted above, any of a variety of alternative architectures and corresponding ML models may be used in other examples without departing from aspects described herein.

[0122] As shown, architecture 850 processes input 802 to produce generative model output 806, aspects of which were discussed above with respect to FIG. 8A. Architecture 850 is depicted as a Transformer model including an encoder 852 and a decoder 854. Encoder 852 processes input embedding 858 (which may be similar in aspects to input embedding 810 of FIG. 8A) including a sequence of symbolic representations corresponding to input 856. In an example, input 856 includes input content 802 corresponding to a content type, aspects of which may be similar to input data 111, virtual content 200, and / or real content 250.

[0123] Additionally, positional encoding 860 may introduce information regarding the relative and / or absolute positions of tokens in input embedding 858. Similarly, while output embedding 874 includes a sequence of symbolic representations corresponding to output 872, positional encoding 876 may similarly introduce information regarding the relative and / or absolute positions of tokens in output embedding 874.

[0124] As shown, the encoder 852 includes an exemplary layer 870. It will be understood that any number of such layers may be used, and that the depicted architecture is simplified for illustrative purposes. The exemplary layer 870 includes two sublayers: a multi-head attention layer 862 and a feedforward layer 866. In the example, residual connections are included around each layer 862, 866, followed by normalization layers 864 and 868, respectively.

[0125] The decoder 854 includes an exemplary layer 890. As with the encoder 852, any number of such layers may be used in other examples, and the depicted architecture of the decoder 854 is simplified for illustrative purposes. As shown, the exemplary layer 890 includes three sublayers: a masked multi-head attention layer 878, a multi-head attention layer 882, and a feedforward layer 886. Aspects of the multi-head attention layer 882 and the feedforward layer 886 may be similar to those discussed above with respect to the multi-head attention layer 862 and the feedforward layer 866, respectively. Additionally, the masked multi-head attention layer 878 performs multi-head attention on the output of the encoder 852 (e.g., output 872). In the example, the masked multi-head attention layer 878 prevents one location from directing attention to a subsequent location. Such masking, in combination with offsetting the embedding (e.g., one position as shown by multi-head attention layer 882), may ensure that predictions for a given position depend on known outputs for one or more positions smaller than the given position. As shown, residual connections are also included around layers 878, 882, and 886, followed by normalization layers 880, 884, and 888, respectively. Multi-head attention layers 862, 888, and 882 can each linearly project the query, key, and value using a set of linear projections onto the corresponding dimensions. Each linear projection can be processed using an attention function (e.g., dot-product attention or additive attention), thereby producing an n-dimensional output value for each linear projection. The resulting values ​​can be concatenated and re-projected so that the values ​​can be subsequently processed (e.g., by the corresponding normalization layer 864, 880, or 884) as shown in FIG. 8B.

[0126] Feedforward layers 866 and 886 may each be a fully connected feedforward network applied to each location. In an example, feedforward layers 866 and 886 each include multiple linear transformations with rectified linear unit activations between them. In an example, each linear transformation may be the same across different locations, while different parameters may be used compared to other linear transformations in the feedforward network.

[0127] Additionally, aspects of the linear transform 892 may be similar to the linear transforms discussed above with respect to the multi-head attention layers 862, 888, and 882 and the feedforward layers 866 and 886. A softmax 894 may further convert the output of the linear transform 892 into a probability of a predicted next token, as indicated by output probability 896. It will be understood that the illustrated architecture is provided by way of example, and that in other examples, any of a variety of other model architectures may be used in accordance with the disclosed aspects. In some examples, multiple iterations of processing are performed in accordance with the above-described aspects (e.g., using the generative model package 804 of FIG. 8A or the encoder 852 and decoder 854 of FIG. 8B ) to generate a sequence of output tokens (e.g., words) that are combined to result in, for example, a complete sentence (and / or any of a variety of other content). It will be understood that other generative models may generate multiple output tokens in a single iteration and, therefore, may use a reduced number of iterations or a single iteration.

[0128] The output probabilities 896 may thus form the embedded outputs 806 in accordance with aspects described herein, such that the outputs of the generative ML model (which may include, for example, structured outputs) are used as inputs for determining actions in accordance with aspects described herein. In other examples, the embedded outputs 806 are provided as generative outputs for updating an episodic object memory.

[0129] 9-11 and the associated description provide an illustration of various operating environments in which aspects of the present disclosure may be practiced, although the devices and systems shown and discussed with respect to Figures 9-11 are for purposes of example and illustration and are not intended to limit the vast number of computing device configurations that may be utilized to implement aspects of the disclosure described herein.

[0130] 9 is a block diagram illustrating the physical components (e.g., hardware) of a computing device 900 in which aspects of the present disclosure may be practiced. The computing device components described below may be suitable for the computing devices described above, including computing device 102 of FIG. 1. In a basic configuration, computing device 900 may include at least one processing unit 902 and system memory 904. Depending on the configuration and type of computing device, system memory 904 may include, but is not limited to, volatile storage (e.g., random access memory), non-volatile storage (e.g., read-only memory), flash memory, or any combination of such memory.

[0131] The system memory 904 may include an operating system 905 and one or more program modules 906 suitable for executing software applications 920, such as one or more components supported by the system described herein. By way of example, the system memory 904 may store an episodic object memory insertion engine or component 924 and / or an episodic object memory retrieval engine or component 926. The operating system 905 may be suitable for controlling the operation of the computing device 900, for example.

[0132] Furthermore, aspects of the present disclosure may be practiced with graphics libraries, other operating systems, or any other application programs and are not limited to any particular application or system. This basic configuration is illustrated in FIG. 9 by the components within dashed line 908. Computing device 900 may have additional features or functionality. For example, computing device 900 may also include additional data storage devices (removable and / or non-removable), such as, for example, magnetic disks, optical disks, or tape. Such additional storage is illustrated in FIG. 9 by removable storage 909 and non-removable storage 910.

[0133] As mentioned above, several program modules and data files can be stored within the system memory 904. While executing on the processing unit 902, the program modules 906 (e.g., applications 920) can perform processes including, but not limited to, the aspects described herein. Other program modules that can be used in accordance with aspects of the present disclosure can include email and contact applications, word processing applications, spreadsheet applications, database applications, slide presentation applications, drawing or computer-aided application programs, etc.

[0134] Furthermore, aspects of the present disclosure may be practiced in electrical circuits including discrete electronic elements, packaged or integrated electronic chips including logic gates, circuits utilizing a microprocessor, or on a single chip including electronic elements or a microprocessor. For example, aspects of the present disclosure may be practiced by a system-on-chip (SOC), in which each or many of the components shown in FIG. 9 may be integrated onto a single integrated circuit. Such an SOC device may include one or more processing units, graphics units, communications units, system virtualization units, and various application functions, all of which are integrated (or "burned") onto the chip substrate as a single integrated circuit. When operated by a SOC, the functionality described herein with respect to the client's ability to switch protocols may be operated by application-specific logic integrated with other components of the computing device 900 on a single integrated circuit (chip). Some aspects of the present disclosure may also be practiced using other technologies capable of performing logical operations such as AND, OR, and NOT, including, but not limited to, mechanical, optical, fluidic, and quantum technologies. Additionally, some aspects of the present disclosure may be practiced within a general-purpose computer or any other circuit or system.

[0135] The computing device 900 may also have one or more input devices 912, such as a keyboard, mouse, pen, voice or audio input device, touch or swipe input device, etc. Output devices 914, such as a display, speakers, printer, etc., may also be included. The aforementioned devices are examples, and other devices may be used. The computing device 900 may include one or more communication connections 916 that enable communication with other computing devices 950. Examples of suitable communication connections 916 include, but are not limited to, radio frequency (RF) transmitter, receiver, and / or transceiver circuitry, universal serial bus (USB), parallel and / or serial ports.

[0136] As used herein, the term computer-readable medium may include computer storage media. Computer storage media may include volatile and nonvolatile, removable and non-removable media implemented by any method or technology for storage of information, such as computer-readable instructions, data structures, or program modules. System memory 904, removable storage 909, and non-removable storage 910 are all examples of computer storage media (e.g., memory storage). Computer storage media may include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article of manufacture that can be used to store information and that can be accessed by computing device 900. Any such computer storage media may be part of computing device 900. Computer storage media do not include carrier waves or other propagated or modulated data signals.

[0137] Communication media may be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term "modulated data signal" may refer to a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared and other wireless media.

[0138] 10 is a block diagram illustrating the architecture of one aspect of a computing device. That is, a computing device can incorporate a system (e.g., architecture) 1002 for implementing some aspects. In some examples, the system 1002 is implemented as a "smartphone" capable of running one or more applications (e.g., a browser, email, calendar, contact manager, messaging client, games, and media client / player). In some aspects, the system 1002 is integrated into computing devices such as an integrated personal digital assistant (PDA) and wireless telephone.

[0139] One or more application programs 1066 may be loaded into memory 1062 and execute on or in association with operating system 1064. Examples of application programs include a phone dialer program, an email program, a personal information manager (PIM) program, a word processing program, a spreadsheet program, an internet browser program, a messaging program, etc. System 1002 also includes a non-volatile storage area 1068 within memory 1062. Non-volatile storage area 1068 may be used to store persistent information that should not be lost even if system 1002 is powered down. Application programs 1066 may use and store information in non-volatile storage area 1068, such as emails or other messages used by an email application. A synchronization application (not shown) is also present on system 1002 and is programmed to interact with a corresponding synchronization application on the host computer to keep information stored in non-volatile storage area 1068 synchronized with corresponding information stored at the host computer. It should be understood that other applications may be loaded into memory 1062 and executed on the mobile computing device 1000 described herein (e.g., an embedded object memory insertion engine, an embedded object memory retrieval engine, etc.).

[0140] The system 1002 includes a power supply 1070, which may be implemented as one or more batteries. The power supply 1070 may also include an external power source, such as an AC adapter or a powered base, that supplements or recharges the batteries.

[0141] The system 1002 may also include a radio interface layer 1072 that performs the function of transmitting and receiving radio frequency communications. The radio interface layer 1072 facilitates wireless connectivity between the system 1002 and the "outside world" via a communications carrier or service provider. Transmissions to and from the radio interface layer 1072 are performed under the control of the operating system 1064. In other words, communications received by the radio interface layer 1072 may be disseminated by the operating system 1064 to the application programs 1066, and vice versa.

[0142] The visual indicator 1020 can be used to provide a visual notification, and / or the audio interface 1074 can be used to produce an audible notification via the audio transducer 1025. In the illustrated example, the visual indicator 1020 is a light-emitting diode (LED), and the audio transducer 1025 is a speaker. These devices can be directly coupled to the power source 1070 so that upon activation, they remain on for a duration dictated by the notification mechanism, even though the processor 1060 and / or dedicated processor 1061 and other components may shut down to conserve battery power. The LED can be programmed to remain on indefinitely to indicate a powered-on state of the device until the user takes action. The audio interface 1074 is used to provide audible signals to and receive audible signals from the user. For example, in addition to being coupled to the audio transducer 1025, the audio interface 1074 can also be coupled to a microphone for receiving audible input to facilitate telephone conversations. According to aspects of the present disclosure, the microphone can also function as an audio sensor to facilitate control of notifications, as described below. The system 1002 may further include a video interface 1076 that enables operation of an on-board camera 1030 for recording still images, video streams, and the like.

[0143] A computing device implementing system 1002 may have additional features or functionality. For example, a computing device may also include additional data storage devices (removable and / or non-removable) such as magnetic disks, optical disks, or tape. Such additional storage is illustrated in Figure 10 by non-volatile storage area 1068.

[0144] The data / information generated or captured by a computing device and stored by system 1002 may be stored locally on the computing device, as described above, or the data may be stored on any number of storage media that may be accessed by the device via wireless interface layer 1072 or via a wired connection between the computing device and another computing device associated with the computing device, such as a server computer in a distributed computing network such as the Internet. As should be understood, such data / information may be accessed by a computing device via wireless interface layer 1072 or via a distributed computing network. Similarly, such data / information may be readily transferred between computing devices for storage and use in accordance with known data / information transfer and storage means, including email and collaborative data / information sharing systems.

[0145] 11 illustrates one aspect of a system architecture for processing data received at a computing system from a remote source, such as a personal computer 1104, a tablet computing device 1106, or a mobile computing device 1108, as described above. Content displayed on the server device 1102 may be stored via different communication channels or other storage types. For example, various documents may be stored using a directory service 1124, a web portal 1125, a mailbox service 1126, an instant messaging store 1128, or a social networking site 1130.

[0146] An application 1120 (e.g., similar to application 920) may be employed by a client in communication with server device 1102. Additionally or alternatively, an episodic object memory insertion engine 1121 and / or an episodic object memory retrieval engine 1122 may be employed by server device 1102. Server device 1102 can provide data to and from client computing devices, such as personal computer 1104, tablet computing device 1106, and / or mobile computing device 1108 (e.g., smartphone) over network 1115. By way of example, the computing system described above can be embodied by personal computer 1104, tablet computing device 1106, and / or mobile computing device 1108 (e.g., smartphone). Any of these examples of computing devices can obtain content from store 1116 in addition to receiving graphical data that can be used for pre-processing in a graphics generating system or post-processing in a receiving computing system.

[0147] Aspects of the present disclosure are described above with reference to block diagrams and / or operational illustrations of, for example, methods, systems, and computer program products according to aspects of the present disclosure. The functions / acts noted in the blocks may occur out of the order noted in any flowchart. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functions / acts involved.

[0148] The description and illustrations of one or more aspects set forth herein are not intended to in any way limit or restrict the scope of the present disclosure as set forth in the claims. The aspects, examples, and details provided herein are believed to be sufficient to convey ownership and enable others to make and use aspects of the present disclosure as set forth in the claims. The disclosure as set forth in the claims should not be construed as limited to any aspect, example, or detail provided herein. It is intended that various features (both structural and methodological) be selectively included or omitted to create embodiments having a particular set of characteristics, whether shown and described in combination or separately. Having provided the description and illustrations herein, those skilled in the art will be able to devise variations, modifications, and alternative embodiments that fall within the spirit of the broader aspects of the general inventive concepts embodied herein without departing from the broader scope of the disclosure as set forth in the claims.

Claims

1. 1. A method for storing landmark memory embeddings in an episodic object memory, comprising: receiving one or more content items, each of the content items having one or more content data; providing the one or more content data associated with the one or more content items to one or more embedding models, the one or more embedding models generating one or more embeddings; receiving a set of embeddings from one or more of the embedding models, each embedding in the set of embeddings corresponding to at least one piece of content data from a respective content item; determining that one or more embeddings of the set of embeddings are landmark memory embeddings; inserting the one or more landmark memory embeddings into the episode object memory, the one or more landmark memory embeddings being associated with references to associated data related to the one or more landmark memory embeddings; and providing said episode object memory; A method comprising:

2. 2. The method of claim 1, wherein the insertion triggers a spatial storage operation to store a vector representation of the landmark memory embedding, the vector representation being stored in at least one of an approximate nearest neighbor (ANN) tree, a kd tree, or a multidimensional tree.

3. 2. The method of claim 1 , wherein the determining comprises ranking the one or more of the embeddings based on a dissimilarity to at least one other embedding in the set of embeddings, and the inserting comprises storing an indication of the ranking corresponding to the landmark memory embedding.

4. The determining step comprises: Obtaining a machine learning model previously trained to identify landmark memories; providing the set of embeddings to the machine learning model; receiving a score from the machine learning model corresponding to the extent to which the embedding is a landmark memory embedding; and comparing the score to a threshold to identify one or more embeddings in the set of embeddings as landmark memory embeddings; The method of claim 1 , comprising:

5. Before the comparison, accepting a user input corresponding to the threshold, the threshold being an episodic memory threshold; The method of claim 4 further comprising:

6. The method of claim 5 , wherein the user input is received from a slider in a graphical user interface.

7. The method of claim 1 , wherein the content data is one or more of audio content data, visual content data, gaze content data, weather content data, news content data, calendar content data, email content data, or location content data.

8. The method of claim 1 , wherein the episode object memory is stored in a location that is different from the location of source data corresponding to the content item.

9. The method of claim 1 , wherein the one or more landmark memory embeddings include a set of properties that define a schema.

10. 1. A method for retrieving a landmark memory from an episode object memory, comprising: generating a user interface; accepting an input by the user interface, the input corresponding to a measure of episodic memory; receiving from the episodic memory an indication of one or more landmark memories based on the extent of the episodic memory; Including, the episodic object memory is generated using content data embedded and ranked based on semantic context; method.

11. The method of claim 10 , wherein the embeddings generated based on the content data are ranked based on the dissimilarity of the embedding to at least one other embedding.

12. the input is a first input, the degree is a first degree, the indication is a first indication, and the method comprises: receiving a second input, the second input corresponding to a second degree of episodic memory; and receiving a second indication of one or more landmark memories corresponding to the second input from the episodic object memory based on a second degree of the episodic memory; The method of claim 10 further comprising:

13. The method of claim 12 , wherein the user interface includes a slider, and the first and second inputs correspond to movement of the slider.

14. 1. A method for storing a landmark memory in an episode object memory and for retrieving a landmark memory from said episode object memory, comprising: receiving one or more content items, each of the content items having one or more content data; providing the one or more content data associated with the one or more content items to one or more embedding models, the one or more embedding models generating one or more embeddings; receiving a set of embeddings from one or more of the embedding models, each embedding in the set of embeddings corresponding to at least one piece of content data from a respective content item; determining that one or more embeddings of the set of embeddings are landmark memory embeddings; inserting the one or more landmark memory embeddings into the episode object memory; and accepting an input corresponding to a degree of an episodic memory and receiving an indication of one or more landmark memories from the episodic object memory based on the degree of the episodic memory; A method comprising:

15. 15. The method of claim 14, wherein the insertion triggers a spatial storage operation to store a vector representation of the landmark memory embedding, the vector representation being stored in at least one of an approximate nearest neighbor (ANN) tree, a kd tree, or a multidimensional tree.