Storing entries in contextual object memory and retrieving information therefrom
By generating an embedding model and inserting a landmark memory embedding into the context object memory, the problem of low efficiency in storing and retrieving memory information in the existing technology is solved, efficient information storage and retrieval is achieved, and user operation efficiency is improved.
Patent Information
- Application Number
- CN202380092931.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-01
- Filing Date
- 2023-12-26
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies have difficulty in efficiently storing and retrieving memory information related to a specific context by humans or AI-based systems, resulting in inefficient information retrieval.
By generating an embedding model from content data, using a machine learning model to rank, weight, and score the embeddings, the landmark memory embedding is determined and inserted into the context object memory, and retrieved in combination with user input.
It improves the efficiency of information storage and retrieval, reduces the consumption of computing resources, and enhances the operating efficiency of users on computing devices.
Smart Images

Figure CN120641909A_ABST
Abstract
Description
Background Art
[0001] Human and machine-based intelligence is characterized by operating under the constraints of a set of significant cognitive resources, such as memory and computational power required for recall and reasoning. Limitations in the efficiency with which humans or AI-based systems can selectively encode experiences into short-term and long-term memory and then later recall the information most relevant to the current context and task at hand drive the need for mechanisms that can identify the most salient memories for storage and retrieval.
[0002] Whether human or artificial, agents are potentially exposed to a vast amount of information about the experiences they encounter, including sensory information, problems encountered and solved, successes and failures in achieving goals, and encodings of reasoning processes and strategies. Experiences can include specific occurrences or events, important people, relationships, and locations, signals about rewards for actions or situations, and locations, as well as the learning of new knowledge and strategies that can be applied multiple times in the future, with or without customization, to more effectively solve problems.
[0003] The process used to selectively encode and access important memories (referred to herein as "iconic memories") must select iconic memories from a large body of experience. Selectivity is important for giving the cognitive system the ability to efficiently store and retrieve information when needed, and when this information will be most valuable in maximizing its goals.
[0004] Some systems may have efficient means for storing vast amounts of experience, but their ability to recall and reason about relevant content within a specific context is limited. Landmark memory can be used as specific pointers to a larger memory repository, enabling efficient organization of large amounts of stored information for retrieval, such as with landmarks acting as easily accessible, recognizable "handles" for accessing progressively more detailed memories. For example, people often wish to recall past events, but because memory is notoriously susceptible to failure, users may struggle to remember the details of various moments. When attempting to recall details, users can at least remember some different types of contextual information relevant to their daily lives. For example, after a meeting covering many different topics, a user may not be able to remember all of the topics. However, the user may remember other aspects of the meeting, such as the location, one or more attendees, and so on. In fact, many events in our daily lives include many different subsets of information, such as weather, locations, people, news, sights, and sounds. While computing devices can collect weather data, location data, attendee information, relevant news data, and other audio / video data, the unstructured nature of such data does not aid user recall.
[0005] It is with these and other general considerations that the embodiments are described.Furthermore, although relatively specific problems have been discussed, it should be understood that the embodiments should not be limited to solving the specific problems identified in the background. Summary of the Invention
[0006] Aspects of the present disclosure relate to methods, systems, and media for storing entries in and / or receiving information from a context object store.
[0007] In some examples, one or more experience content items are received, each having one or more content data. The one or more content data may be provided to a memory processor, which creates or accesses one or more embedding models, and the one or more embedding models proceed to integrate the content into an organized metric space that represents regularized distances between different content data based on attributes of the data. A set of embeddings may be received from one or more embedding models in the embedding model. In an example, an embedding may also be an embedding object comprising multiple embeddings. Each embedding in the embedding set may correspond to at least one content data from a corresponding content item. One or more embeddings in the embedding set may be determined by the memory processing device as a landmark memory embedding. For example, one or more of the embeddings may be ranked, weighted, and / or scored based on, for example, a machine learning model trained via a supervised learning method and / or a calculated dissimilarity to at least one other embedding in the embedding set. The landmark memory embeddings may be inserted into a context object memory, which links the memory to other landmark memories based on relationships between one or more attributes of the landmark (such as temporal and / or spatial relationships between content linked to the landmark) or other relationships (such as social relationships between people or similarities in the context of applying content to solve a problem). In this way, the embedding of iconic memories can integrate useful experiential content that needs to be jointly accessed in real-time recognition and problem solving. In addition, in some examples, input (e.g., user input) corresponding to the degree of salience of content stored and retrieved from the contextual memory can be received.
[0008] This summary is provided to introduce some concepts in a simplified form that will be further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of the examples will be set forth in part in the following description and in part will be apparent from the description or may be learned through practice of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Non-limiting and non-exhaustive examples are described with reference to the following figures.
[0010] Figure 1 An overview of an example system according to some aspects described herein is shown.
[0011] Figure 2 Example content data is shown according to some aspects described herein.
[0012] Figure 3 An example process for storing entries in a context object store according to some aspects described herein is shown.
[0013] Figure 4 An example vector space is shown according to some aspects described herein.
[0014] Figure 5A An example method for storing landmark memories in a contextual object memory according to some aspects described herein is shown.
[0015] Figure 5B
[0014] An example method for determining that an embedding is a landmark memory embedding according to some aspects described herein is shown.
[0016] Figure 6A
[0014] An example system for retrieving information from a context object store according to some aspects described herein is shown.
[0017] Figure 6B
[0014] An example system for retrieving information from a context object store according to some aspects described herein is shown.
[0018] Figure 7 An example method for retrieving information from a context object store according to some aspects described herein is shown.
[0019] Figure 8A and Figure 8B An overview of an example generative machine learning model that can be used in accordance with various aspects described herein is shown.
[0020] Figure 9 A block diagram illustrating example physical components of a computing device in which aspects of the present disclosure may be practiced.
[0021] Figure 10 Shown is a simplified block diagram of a computing device that can be used to practice aspects of the present disclosure.
[0022] Figure 11 is a simplified block diagram of a distributed computing system in which aspects of the present disclosure may be practiced. DETAILED DESCRIPTION
[0023] In the following detailed description, reference will be made to the accompanying drawings, which form a part of this document and in which specific embodiments or examples are shown by way of illustration. These aspects can be combined, other aspects can be utilized, and structural changes can be made without departing from the present disclosure. The embodiments can be practiced as methods, systems or devices. Therefore, the embodiments can take the form of hardware implementation, complete software implementation or a combination of software and hardware. Therefore, the following detailed description should not be regarded as restrictive, and the scope of the present disclosure is defined by the appended claims and their equivalents.
[0024] Eric J. Horvitz, the first named inventor of the present disclosure, is also the first named inventor of U.S. patent application No. 11 / 172,467, filed on June 30, 2005, entitled “Methods for Selecting, Organizing, and Display Images from Large Stores of Personal Content,” which describes systems and methods for automatically organizing, managing, and presenting information content in a manner that is relevant or personal to a user.
[0025] Given the limitations of a machine or human's cognitive resources and the immense complexity of the world in which they operate, it can be difficult to store and retrieve important details about the content of experience at various moments in daily life. Computing devices can receive many different types of contextual information related to a user's daily life, such as audio data, video data, gaze data, weather data, location data, news data, or other types of environmental data that will be recognized by those skilled in the art. Users can associate some of these different types of contextual or environmental information with details of their daily lives, such as documents presented in meetings, emails sent on a particular date, conversations about a particular topic, and so on.
[0026] As humans, we often associate memories with contextual information related to the memory we are trying to recall. For example, when looking for a document on a computer, a user may remember that an important news story was released when they last opened the document, or that a storm was happening outside when they last opened the document, or that they were heading to a particular park when they last accessed the document via their smartphone. Additionally or alternatively, the user may remember that they were on a phone or video call with a particular person when they last opened the document. However, the process of the user determining where to locate the document on one or more of their computing devices by tracing back their own memory associations can be time-consuming, frustrating, and unnecessarily consumes computing resources (e.g., processor and / or memory).
[0027] In addition, in some examples, the user's computing device can record only the screen of the computing device, or only the audio of the computing device, or only the location of the computing device, with the user's permission, so that when the user attempts to perform an action on their computing device (e.g., locate a document, open a document, send an email, schedule a calendar event, etc.), the recording can be searched for the association the user is trying to establish. However, storing recordings (e.g., video, audio, or location) may take up a relatively large amount of memory. In addition, searching through recordings may be relatively time-consuming and computationally expensive. Therefore, there is a need to improve user efficiency in performing actions on a computing device based on landmark memory and associated contextual information.
[0028] The landmark memories described herein can be memories that are significant, unique, and / or distinct that the system and / or user distinguish from other memories that may be stored. For example, a landmark memory can be distinguished because it is a memo person, an important transaction, a meaningful portion of a file, or another distinguishable piece of content that is useful for meeting past events. A landmark memory can have an associated attribute structure, such as a dimension that includes data about the memory identified as a landmark memory. Landmark memories and their attributes can use the mechanisms provided herein to enable efficient, context-sensitive access to relevant experiences and encoded content.
[0029] A landmark memory can be a parent structure with child structures (such as child landmark memories) in which relationships to and / or relationships between child structures are stored. Thus, landmark memories can have relationships to each other, such as places, people, task types, and / or times (as described with respect to the term "episodic memory" used throughout, which has its roots in psychology and which one of ordinary skill in the art will recognize has been originally described by the teachings of psychologist Endel Tulving). For example, landmark memories can be close to each other, such as the 4th of July and the explosion of fireworks, and the dinner party the previous night, including key participants and salient aspects of the conversation. One of ordinary skill in the art will recognize additional and / or alternative examples.
[0030] Thus, some aspects of the present disclosure relate to methods, systems, and media for storing landmark memories in a context object memory. Typically, a content item having one or more content data (e.g., an email, audio data, video data, a message, internet encyclopedia data, a skill, a command, source code, a programming assessment, etc.) can be received. The one or more content data can be provided to a model to generate a set of embeddings (e.g., semantic embeddings). One or more embeddings from a set of embeddings can be determined to be landmark memory embeddings. For example, the set of embeddings can be ranked, weighted, scored, and / or provided to a trained machine learning model to determine that one or more embeddings are landmark memory embeddings. This can include models that are capable of making this determination, which models are potentially not directly trained to perform the task (i.e., via transfer of skills). One or more landmark memory embeddings can be inserted into the context object memory with an indication, such as corresponding to a ranking, weighting, scoring, etc., and / or with a reference to related data.
[0031] Additionally or alternatively, some aspects of the present disclosure relate to methods, systems, and media for retrieving landmark memories from a contextual object memory. Generally speaking, a user interface can be generated through which input corresponding to a degree of contextual memory is received. Based on the degree of contextual memory, an indication of one or more landmark memories can be received from the contextual memory.
[0032] Advantages of the mechanisms disclosed herein can include improved user efficiency for performing actions (e.g., retrieving information, locating a virtual document, generating a draft email, generating a draft calendar event, providing content information related to a virtual document, etc.) via a computing device based on content that has been summarized (e.g., via text and / or embeddings) and analyzed (e.g., ranked, weighted, scored, provided to a trained machine learning model). Additionally, the mechanisms disclosed herein for generating context can improve computational efficiency by, for example, reducing the amount of memory required to track content (e.g., via feature vectors or tags, rather than audio / video recordings, relatively large amounts of text). Still further, the mechanisms disclosed herein can improve computational efficiency for receiving content from a context object memory, such as by comparing feature vectors, as opposed to searching a relatively large record or repository stored in memory.
[0033] Figure 1 An example of a system 100 according to some aspects of the disclosed subject matter is shown. System 100 can be a system for storing landmark memories in a context object memory. Additionally or alternatively, system 100 can be a system for using the context object memory, such as by retrieving information from the context object memory. System 100 includes one or more computing devices 102, one or more servers 104, a content data source 106, an input data source 107, and a communication network or networks 108.
[0034] The computing device 102 may receive content data 110 from a content data source 106, which may be, for example, a microphone, a camera, a global positioning system (GPS), or the like that transmits the content data, a computer-executed program that generates the content data, and / or a memory having data corresponding to the content data stored therein. The content data 110 may include visual content data, audio content data (e.g., speech or ambient noise), gaze content data, calendar entries, emails, document data (e.g., virtual documents), weather data, news data, blog data, encyclopedia data, and / or other types of virtual and / or real content data that one of ordinary skill in the art would recognize. In some examples, the content data may include text, images, source code, commands, skills, and / or programming assessments.
[0035] The computing device 102 may also receive input data 111 from an input data source 107, which may be, for example, a camera, a microphone, a program executed by a computer, and / or a memory having data corresponding to the input data stored therein. The content data 111 may be, for example, user input (such as a voice query, a text query, etc.), an image, an action performed by a user and / or a device, a computer command, a program evaluation, or some other input data recognizable by one of ordinary skill in the art.
[0036] Additionally or alternatively, the network 108 may receive content data 110 from the content data source 106. Additionally or alternatively, the network 108 may receive input data 111 from the input data source 107.
[0037] Computing device 102 may include a communication system 112, a context object store insertion engine or component 114, and / or a context object store retrieval engine or component 116. In some examples, computing device 102 may execute at least a portion of context object store insertion component 114 to generate a set of embeddings corresponding to one or more subsets of received content data 110 to be inserted into the context object store. For example, each subset of the content data may be provided to a machine learning model, such as a natural language processor and / or a vision processor, to generate the set of embeddings. In some examples, the subset of the content data may be provided to another type of model, such as a generative large language model (LLM).
[0038] Furthermore, in some examples, computing device 102 can execute at least a portion of context object memory retrieval component 116 to retrieve one or more landmark memories or associated information from a context object memory, such as based on an input (e.g., generated based on input data 111). In some examples, context object memory retrieval component 116 can also determine an action. For example, an action can be determined based on the input and one or more embeddings corresponding to the landmark memories (e.g., one or more landmark memory embeddings).
[0039] The server 104 may include a communication system 112, a context object store insertion engine or component 114, and / or a context object store retrieval engine or component 116. In some examples, the server 104 may execute at least a portion of the context object store insertion component 114 to generate a set of embeddings corresponding to one or more subsets of the received content data 110 to be inserted into the context object store. For example, each subset of the content data may be provided to a machine learning model, such as a natural language processor and / or a vision processor, to generate the set of embeddings. In some examples, the subset of the content data may be provided to another type of model, such as a generative large language model (LLM).
[0040] Furthermore, in some examples, server 104 can execute at least a portion of context object memory retrieval component 116 to retrieve one or more landmark memories or associated information from a context object memory, such as based on an input (e.g., generated based on input data 111). In some examples, context object memory retrieval component 116 can also determine an action. For example, an action can be determined based on the input and one or more embeddings corresponding to the landmark memories (e.g., one or more landmark memory embeddings).
[0041] Additionally or alternatively, in some examples, computing device 102 may transmit data received from content data source 106 and / or input data source 107 to server 104 via communication network 108, which may execute at least a portion of context object memory insertion component 114 and / or context object memory retrieval engine 116. In some examples, context object memory insertion component 114 may execute the following, respectively, in conjunction with: Figure 4 and Figure 7 Furthermore, in some examples, the context object memory retrieval component 116 may perform the following steps in conjunction with FIG. 5 and FIG. Figure 7 One or more portions of the method / process 500 and / or 700 are described.
[0042] In some examples, computing device 102 and / or server 104 can be any suitable computing device or combination of devices, such as a desktop computer, a vehicle computer, a mobile computing device (e.g., a laptop, a smartphone, a tablet computer, a wearable computer, etc.), a server computer, a virtual machine executed by a physical computing device, a network server, etc. Furthermore, in some examples, there can be multiple computing devices 102 and / or multiple servers 104. One of ordinary skill in the art will recognize that content data 110 and / or input data 111 can be received at one or more of the multiple computing devices 102 and / or one or more of the multiple servers 104, such that the mechanisms described herein can insert entries into and / or use the embedded object memory based on an aggregation of content data 110 and / or input data 111 received across computing devices 102 and / or servers 104.
[0043] In some examples, content data source 106 can be any suitable source of content data (e.g., a microphone, a camera, a GPS, a sensor, etc.). In a more specific example, content data source 106 can include a memory that stores content data (e.g., local memory of computing device 102, local memory of server 104, cloud storage, a portable memory connected to computing device 102, a portable memory connected to server 104, etc.). In another more specific example, content data source 106 can include an application configured to generate content data. In some examples, content data source 106 can be local to computing device 102. Additionally or alternatively, content data source 106 can be remote from computing device 102 and can transmit content data 110 to computing device 102 (and / or server 104) via a communication network (e.g., communication network 108).
[0044] In some examples, input data source 107 can be any suitable input data source (e.g., a microphone, a camera, a sensor, etc.). In a more specific example, input data source 107 can include a memory that stores input data (e.g., local memory of computing device 102, local memory of server 104, cloud storage, portable memory connected to computing device 102, portable memory connected to server 104, privately accessible memory, publicly accessible memory, etc.). In another more specific example, input data source 107 can include an application configured to generate input data. In some examples, input data source 107 can be located locally on computing device 102. Additionally or alternatively, input data source 107 can be remote from computing device 102 and can transmit input data 111 to computing device 102 (and / or server 104) via a communication network (e.g., communication network 108).
[0045] In some examples, communication network 108 can be any suitable communication network or combination of communication networks. For example, communication network 108 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, a 5G network, etc., conforming to any suitable standard), a wired network, etc. In some examples, communication network 108 can be a local area network (LAN), a wide area network (WAN), a public network (e.g., the Internet), a private or semi-private network (e.g., a company or university intranet), any other suitable type of network, or any suitable combination of networks. Figure 1 The communication links (arrows) shown may each be any suitable communication link or combination of communication links, such as a wired link, a fiber optic link, a Wi-Fi link, a Bluetooth link, a cellular link, or the like.
[0046] Figure 2 An example of virtual content 200 and real content 250 according to some aspects described herein is shown. As discussed with respect to system 100, the mechanisms described herein can include receiving content data (e.g., content data 110) from a content data source. The content data can be virtual content 200 and / or real content 250.
[0047] Typically, when a user is interacting with a computing device (e.g., computing device 102), they are interacting with a virtual environment while physically being in a real environment (e.g., a physical environment). Therefore, the contextual information that a user can invoke when interacting with the computing device can be virtual content (e.g., virtual content 200) and / or real content (e.g., real content 250).
[0048] Virtual content 200 includes a virtual person 202, audio content 204, a virtual document 206, and / or visual content 208. Virtual person 202 may include data corresponding to a virtual image of an individual, such as generated via a video stream, a still image, an avatar corresponding to the person, etc. Additionally or alternatively, virtual person 202 may include data corresponding to the person, such as an icon corresponding to the person or other indicators corresponding to a particular person as would be recognized by one of ordinary skill in the art.
[0049] Audio content 204 may include data corresponding to speech data generated in the virtual environment. For example, audio content 204 in the virtual environment may be generated by computing device 102 to correspond to audio received from a user (e.g., where the user is speaking into a microphone of a computing device that may be separate from computing device 102). Additionally or alternatively, audio content 204 may correspond to other types of audio data that may be generated in the virtual environment, such as animal sounds, beeps, buzzers, or another type of audio indicator.
[0050] Virtual document 206 can include the type of document found in the virtual environment. For example, virtual document 206 can be a text editing document, a presentation, an image, a spreadsheet, an animated sequence of images, a calendar invitation, an email, a notification, or any other type of virtual document that one of ordinary skill in the art can recognize.
[0051] Visual content 208 may include data corresponding to graphical content that may be displayed or generated by a computing device. For example, visual content 208 may be content generated via an application (e.g., a web browser, a presentation application, a teleconferencing application, a business management application, etc.) running on computing device 102. Visual content 208 may include data scraped from the screen display of computing device 102. For example, any visual indication displayed on computing device 102 may be included in visual content 208.
[0052] Each of the multiple types of virtual content 200 may be a subset of the virtual content 200 that may be received by the mechanisms described herein as a subset of the content data 110. Furthermore, while specific examples of virtual content types have been discussed above, persons of ordinary skill in the art may recognize additional and / or alternative types of virtual content as they relate to virtual components of a virtual environment and / or augmented reality environment.
[0053] Real content 250 includes visual content 252, audio content 254, devices used 256, location 258, weather 260, news 262, time 264, people 268, and / or gaze content 270. Visual content 252 may include data received from a camera or optical sensor. For example, visual content 252 may include information about a user's biometric data collected with the user's permission, information about the user's physical environment, information about clothing worn by the user, and the like. Persons of ordinary skill in the art will recognize additional examples of visual content 252.
[0054] Audio content 254 includes audio data originating from a real or physical environment. For example, audio content 254 may include data corresponding to speech data (e.g., audio spoken by a user or audio that can be otherwise converted into text associated with speech). Additionally or alternatively, audio content 204 may include data corresponding to ambient noise data. For example, audio content 204 may include animal sounds (e.g., a dog barking, a cat meowing, etc.), traffic sounds (e.g., airplanes, cars, sirens, etc.), and natural sounds (e.g., waves, wind, etc.). Those of ordinary skill in the art will recognize additional examples of audio content 254.
[0055] Device used 256 may include information about which device the user is currently using. For example, when a user attempts to locate a virtual document (e.g., virtual document 206), they may remember that they last accessed the virtual document on a first computing device (e.g., a mobile phone) as compared to a second computing device (e.g., a laptop). Thus, based on the user's memory of associating a particular device with an element (e.g., a virtual document or computer application) being attempted to be accessed or discovered, the mechanisms disclosed herein may locate or open the desired element being attempted to be accessed or discovered.
[0056] Location 258 may include information about the location of the user. For example, location data may be received from a global positioning system (GPS), a satellite positioning system, a cellular positioning system, or other type of location determination system. A user may associate the location where they are physically located with actions performed on a computing device. Thus, for example, using the mechanisms disclosed herein, a user may provide a query such as "While I was at the beach and in a video conference with my boss, what document was I working on?" The mechanisms disclosed herein may determine the document or documents that the user may be referring to based on stored information about which document they were working on, when they were at the beach, and when they were in a video conference with their boss.
[0057] Weather 260 may include information about the weather around the user. For example, for a given time, as determined by time content 264, weather information (e.g., precipitation, temperature, humidity, etc.) for the user's location (e.g., based on location content 258) may be received or otherwise obtained. Thus, for example, using the mechanisms disclosed herein, a user may provide a query such as "Who do I video call on my phone when it's cold and snowing outside?" The mechanisms disclosed herein may determine what avatar (e.g., from avatar content 202) the user is attempting to meet based on when the user is using a cell phone (e.g., from device used content 256) and when it's cold and snowing (e.g., from weather content 260).
[0058] News 262 may include information about recent news stories of which the user may be aware. For example, for a given time, as determined by time content 264, relatively recent news stories covering important events may have been published. Thus, for example, using the mechanisms disclosed herein, a user may provide a query such as "Who was I video chatting with on my laptop computer the day the Chicago basketball team won the national championship?" The mechanisms disclosed herein may determine what avatar (e.g., from avatar content 202) the user is trying to recall based on when the user used the laptop computer (e.g., from device content 256 used) and when the Chicago basketball team won the national championship (e.g., from news content 262). Additional or alternative types of news stories may include holidays, birthdays, local events, national events, natural disasters, celebrity updates, scientific discoveries, sports updates, or any other type of news that one of ordinary skill in the art would recognize.
[0059] Time 264 may include one or more timestamps for receiving various content types. For example, each content type (e.g., virtual content 200 and / or real content 250, or a specific content type therein) may be timestamped when it is acquired or received. In general, the mechanisms described herein may be transient, as actions performed by a computing device based on multiple stored content types may depend on the timestamps assigned to each content type.
[0060] People 268 may include information about people in the physical environment surrounding one or more computing devices (e.g., computing device 102). With the user's permission, after receiving and storing any relevant biometric data, biometric recognition (e.g., facial recognition, voice recognition, fingerprint recognition, etc.) may be used to identify a person. Additionally or alternatively, a person may be identified by participating in a specific action on a computing device (e.g., engaging with specific software). Furthermore, some individuals may be identified by logging into one or more computing devices. For example, a person may be the owner of a computing device, and the computing device may be linked to the person (e.g., via a password, biometric input, etc.). Thus, when the computing device is logged into, the person is thereby identified. Similarly, a person may be identified by logging into a specific application (e.g., via a password, biometric input, etc.). Thus, when the specific application is logged into, the person is thereby identified. Additionally or alternatively, one or more individuals may be identified using a radio frequency identification tag (RFID), an ID badge, a barcode, a QR code, or some other identification means capable of identifying a person via some technical interface. Additionally or alternatively, the proximity of certain persons relative to each other may be identified based on identifying two or more persons using the mechanisms described herein.
[0061] Gaze content 270 may include information corresponding to the location on one or more computing devices that a user is looking at. Thus, if a first user is on a video conference call and tells a second user that they will send the second user a document they are discussing while viewing the document on their computing device, the mechanisms described herein may generate a draft email from the first user to the second user based on the audio content received from the video conference call, the received gaze content, and an identification of who the second user is, such that the email is drafted to be sent to the current person. Similarly, the draft document may be saved to the user's clipboard so that the draft document can be pasted into an email, message, or other form of virtual communication.
[0062] One of ordinary skill in the art may recognize additional types of virtual content 200 and / or real content 250. Furthermore, while certain content types have been illustrated and described as types of virtual content 200 (e.g., virtual person 202, audio content 204, virtual document 206, and visual content 208), and other content types have been illustrated and described as types of real content (e.g., visual content 252, audio content 254, device used 256, location 258, weather 260, news 262, time 264, person 268, and gaze content 270), it should be recognized that in some examples, such categorization of virtual content 200 and real content 250 may be interchangeable for the content types disclosed herein (e.g., when location may refer to both a virtual location (such as a virtual meeting space) and a physical location of a device being used to access the virtual meeting space), while in other cases, such categorization of virtual content 200 and real content 250 is fixed (e.g., in examples where location refers only to a physical location (such as a geographic coordinate location in the physical world)).
[0063] Usually, about Figure 2 The different content types discussed provide various contexts through which users associate information they are attempting to recall and / or use to perform actions via a computing device. In some examples, a user can provide real content (e.g., real content 250) information to receive information from a computing device, or perform an action related to virtual content (e.g., virtual content 200) on a computing device. Additionally or alternatively, a user can provide virtual content (e.g., virtual content 200) information to receive information related to real or physical content (e.g., real content 250) from a computing device. Additionally or alternatively, a user can provide a combination of virtual content and real content to receive information corresponding to the related real content, receive information corresponding to the virtual content, and / or perform an action related to the virtual content via a computing device.
[0064] Those skilled in the art will recognize the various contexts in which examples of virtual content 200 and real content 250 may be collected. For example, content 200, 250 may be collected for business operations, instructional history, medical experience, project management, event planning and execution. For example, in healthcare, landmark memories may correspond to content encoded over time based on landmark medical experiences, events, and / or content (e.g., the onset of a critical illness, significant test and lab results, diagnoses, hospitalizations, surgeries, visits to the doctor, family health challenges, etc.). These landmark medical items may be stored as landmark memories using the mechanisms provided herein and utilized in interactions that connect a physician or patient with an artificial intelligence (AI) system, such as those incorporating aspects of the systems provided herein.
[0065] Typically, users are more likely to remember atypical content information than regular content information. Therefore, the mechanisms described herein can be used to determine that atypical data is associated with a signature memory. Atypical data is data that is irregular or abnormal given a category that can be used to classify data. For example, a data segment that appears less than 1% of the time within a data category (such as a person, event, location, etc.) that can be used to classify data segments can be determined to be abnormal data. In alternative examples, the 1% threshold can be 5%, or 10%, or 20%, or any value or range of values therebetween.
[0066] Figure 3 An example process 300 for storing entries in a context object store is shown. The example process begins by receiving a content item 302. The content item 302 may be received from a content data source, such as the one previously described herein. Figure 1 Furthermore, the content items 302 may include virtual content and / or real content, such as those previously described herein. Figure 2 Depicted virtual content 200 and / or real content 250. Content item 302 may be one or more content items or objects each including one or more content data.
[0067] The content data of content item 302 can be input into model 304. For example, model 304 may include a machine learning model, such as a machine learning model trained to generate one or more embeddings based on the received content data. In some examples, model 304 includes a generative model, such as a large language model (LLM). In some examples, model 304 includes a natural language processor. Additionally or alternatively, in some examples, model 304 includes a visual processor. Model 304 can be trained on one or more data sets compiled by an individual and / or system. Additionally or alternatively, model 304 can be trained based on a data set including information obtained from the Internet. In addition, model 304 can include a version. Those skilled in the art will understand that any type of model can be employed as part of the aspects disclosed herein without departing from the scope of the present disclosure.
[0068] Model 304 can output an embedding set 306. Embedding set 306 can include one or more embeddings (e.g., one or more semantic embeddings). Embedding set 306 can be unique to model 304. Additionally or alternatively, embedding set 306 can be unique to a specific version of model 304. Each embedding in embedding set 306 can correspond to a corresponding content object and / or corresponding content data from content item 302. Thus, for each object and / or content data in content item 302, model 304 can generate an embedding.
[0069] Each embedding in embedding set 306 can be associated with a corresponding indication (e.g., a byte address) corresponding to the location of the source data associated with the embedding. For example, if content item 302 corresponds to a collection of emails, embedding set 306 can be a hash value generated for each email based on the content within the email (e.g., the abstract meaning of the words or images included in the email). An indication corresponding to the location (e.g., in memory) of the email can then be associated with the embedding generated based on the email. In this regard, the actual source data of the content object can be stored separately from the corresponding embedding, which may occupy less memory. Although content item 302 corresponds to a specific example of a collection of emails, those of ordinary skill in the art will recognize additional and / or alternative examples, at least in light of the teachings provided herein.
[0070] Embeddings 306 can be provided to a ranking engine 308. Ranking engine 308 can rank, score, and / or assign weights to one or more of embeddings 306. The ranking and / or weighting can be based on the dissimilarity of one or more embeddings from set of embeddings 306 to one or more other embeddings from set of embeddings 306. One or more of the ranked, scored, and / or weighted embeddings can correspond to a landmark memory. One or more of the ranked, scored, and / or weighted embeddings that correspond to a landmark memory can be determined as a landmark memory embedding.
[0071] Typically, users are more likely to remember atypical content information than regular content information. Therefore, the mechanisms described herein can be used to determine that atypical data is associated with and / or marked as a landmark memory. Based on determining the landmark memory, data related to the landmark memory can also be determined (e.g., related by time, location, person, etc.), so that the landmark memory can be a pivot point for searching for data related thereto. Atypical data is irregular or abnormal data. For example, a data segment that appears less than 1% of the time within a data category (such as a person, event, location, email, etc.) that can classify the data segment can be determined to be abnormal data. In some examples, the 1% threshold can alternatively be 5%, or 10%, or 20%, or any value or range of values therebetween.
[0072] One or more embeddings that are ranked, scored, and / or weighted (e.g., relative to the extent to which they are landmark memory embeddings) may be inserted into the context object memory 310. Additionally or alternatively, a subset of the ranked, scored, and / or weighted embeddings may be inserted into the context object memory 310 based on comparing the ranked, scored, and / or weighted embeddings to a predetermined threshold, thereby determining which of the embeddings are landmark memory embeddings. One or more landmark memory embeddings may be stored in the context object memory 310 along with an indication of the corresponding ranking, score, and / or weighting. Furthermore, the landmark memory embeddings may be associated with references (e.g., pointers) to related data associated with the one or more landmark memory embeddings (e.g., embeddings similar to the landmark memory embeddings, source data associated with the landmark memory embeddings). The landmark memories described herein may include a set of attributes that define a pattern. The set of attributes may include a summary of the landmark memory (e.g., a natural language summary of the landmark memory) and / or a reference (e.g., a pointer) to data related to the landmark memory.
[0073] In some examples, context object storage 310 stores only landmark memory embeddings. Additionally or alternatively, in some examples, context object storage 310 stores embeddings that are landmark memory embeddings and embeddings that are not landmark memory embeddings, such as with indications corresponding to which embeddings are landmark memory embeddings and which embeddings are not landmark memory embeddings.
[0074] Figure 4 An example vector space 400 according to some aspects described herein is shown. Vector space 400 includes multiple feature vectors, such as a first feature vector 402, a second feature vector 404, a third feature vector 406, a fourth feature vector 408, and a fifth feature vector 410. Each of the multiple feature vectors 402, 404, 406, and 408 corresponds to a corresponding embedding 403, 405, 407, 409 generated based on multiple content data subsets (e.g., content data subset 110, virtual content 200, and / or real content 250). Embeddings 403, 405, 407, and 409 can be semantic embeddings. Fifth feature vector 410 is generated based on input embedding 411 (e.g., an embedding ranked or weighted against other embeddings). Input embeddings can also be generated based on content data. Alternatively, input embeddings can be provided separately (e.g., based on user input).
[0075] Each of the feature vectors 402, 404, 406, 408, and 410 has a measurable distance from one another. For example, the distance between the feature vectors 402, 404, 406, and 408 and the fifth feature vector 410 corresponding to the input embedding 411 can be measured using cosine similarity. Alternatively, the distance between the feature vectors 402, 404, 406, and 408 and the fifth feature vector 410 can be measured using another distance measurement technique (e.g., an n-dimensional distance function) that one of ordinary skill in the art will recognize.
[0076] The similarity of each of the feature vectors 402, 404, 406, 408 to the feature vector 410 corresponding to the input embedding 411 can be determined, for example, based on a measured distance between the feature vectors 402, 404, 406, 408 and the feature vector 410. The similarity between the feature vectors 402, 404, 406, 408 and the feature vector 410 can be used to rank, weight, group, or cluster the feature vectors 402, 404, 406, and 408 (e.g., based on how dissimilar one vector is to the other vectors).
[0077] Embeddings 403 and 405, corresponding to feature vectors 402 and 404, respectively, may fall into the same content group. For example, embedding 403 may be associated with a first email, and embedding 405 may be associated with a second email. One of ordinary skill in the art will recognize additional and / or alternative examples of content groups in which embeddings may be categorized.
[0078] Example clusters of embeddings (such as including one or more embeddings selected from embeddings 403, 405, 407, 409, and 411) can be clusters of embeddings corresponding to people, geographic locations, images, and / or other modalities (such as smell, touch, etc.). In some examples, the embeddings described herein include emotions (e.g., of photos, videos, etc.), cognitive states, etc. Meta information about content corresponding to the embeddings disclosed herein can be established. For example, the meta information can include the pattern of output, the number of queries from users, etc. In some examples, embeddings can be filtered based on the significance of the memories to which the embeddings correspond. Therefore, the importance or significance of one or more embeddings can be determined based on the memories to which the one or more embeddings correspond using the mechanisms provided herein.
[0079] One or more of embeddings 403, 405, 407, 409, and 411 may be stored in a data structure, such as an ANN tree, a kd-tree, an octree, another n-dimensional tree, or another data structure capable of storing a vector space representation as would be appreciated by one of ordinary skill in the art. Furthermore, the memory corresponding to the data structure in which one or more of embeddings 403, 405, 407, 409, and 411 are stored may be arranged or stored within the data structure in a manner that groups the embeddings together, such as to improve the efficiency of a subsequent search process.
[0080] In some examples, feature vectors and their corresponding embeddings generated according to the mechanisms described herein can be stored for an indefinite period of time. Additionally or alternatively, in some examples, as new feature vectors and / or embeddings are generated and stored, the new feature vectors and / or embeddings can overwrite older feature vectors and / or embeddings stored in memory (e.g., based on metadata indicating the version of the embedding) to increase memory capacity. Additionally or alternatively, in some examples, feature vectors and / or embeddings can be deleted from memory at specified time intervals and / or based on the amount of available memory (e.g., in the embedding object memory 310) to increase memory capacity.
[0081] In general, the ability to store embeddings corresponding to received content data allows users to associate and locate data in novel ways with the benefit of computational efficiency. For example, instead of storing a video recording of a computing device's screen or a web page on the Internet, a user can instead use the mechanisms described herein to store embeddings corresponding to content objects. The embeddings can be hashes, rather than, for example, video recordings of hundreds of thousands of pixels per frame. Thus, the mechanisms described herein are useful for reducing memory usage and for reducing the use of processing resources for searching stored content. One of ordinary skill in the art will recognize additional and / or alternative advantages.
[0082] Figure 5A An example method 500 is shown for storing a landmark memory (e.g., a landmark memory embed and / or a landmark memory text) in a context object memory according to some aspects described herein. In an example, aspects of the method 500 are performed by a device, such as described above with respect to Figure 1 Computing device 102 and / or server 104 are discussed.
[0083] Method 500 begins at operation 502, where one or more content items or objects are received. The content items may have one or more content data. The one or more content data may include at least one real content data and / or at least one virtual content data. The content data may be similar to Figure 1 The content data 110 discussed. Additionally or alternatively, the content data may be similar to the content data 110 discussed. Figure 2The virtual content data 200 and / or the real content data 250 are discussed.
[0084] The content data may include one or more of audio content data, visual content data, gaze content data, calendar content data, email content data, virtual documents, data generated by a specific software application, weather content data, news content data, encyclopedia content data, location content data, and / or blog content data. In some examples, the content data may include at least one of a skill, a command, or a programming assessment. Those skilled in the art will recognize additional and / or alternative types of content data.
[0085] At operation 504, it is determined whether at least one of the subsets of content data has an associated embedding model (e.g., a semantic embedding model). For example, the embedding model may be similar to the one regarding Figure 3 Model 304 discussed. An embedding model can be trained to generate one or more embeddings based on at least one of the subsets of content data. In some examples, the embedding model can include a natural language processor. In some examples, the embedding model can include a vision processor. In some examples, the embedding model can include a machine learning model. Furthermore, in some examples, the embedding model can include a generative large-scale language model. The embedding model can be trained on one or more datasets compiled by an individual and / or the system. Additionally or alternatively, the embedding model can be trained based on a dataset that includes information obtained from the internet.
[0086] If it is determined that at least one item of the content data does not have an associated embedding model, the process branches to "No" to proceed to operation 506, where a default action is performed. For example, the content data and / or content item may have an associated preconfigured action. In other examples, method 500 may include determining whether the content data and / or content item has an associated default action so that, in some cases, no action may be performed as a result of the received content item. Method 500 may terminate at operation 506. Alternatively, method 500 may return to operation 502 to provide an iterative loop to receive one or more content items and determine whether at least one item of the content data of the content item has an associated embedding model.
[0087] However, if it is determined that at least one of the content data items has an embedding model, the flow branches "yes" to operation 508, where one of the content data items associated with the content item is provided to one or more embedding models. The embedding models generate one or more embeddings. Furthermore, one or more embedding models may include a version. Each embedding generated by a corresponding embedding model may include metadata corresponding to the version of the embedding model that generated the embedding.
[0088] In some examples, the content data is provided locally to the one or more embedding models. In some examples, the content data is provided to the one or more embedding models via an application programming interface (API). For example, a first device (e.g., computing device 102 and / or server 104) can interface with a second device (e.g., computing device 102 and / or server 104) via an API configured to provide embedding or an indication thereof in response to receiving the content data or an indication thereof.
[0089] At operation 510, one or more embeddings are received from one or more of the embedding models. In some examples, the one or more embeddings are a set of embeddings or a plurality of embeddings. For example, a set of embeddings (such as embedding set 306) can be associated with a first embedding model (e.g., model 304). In some examples, the set of embeddings can be uniquely associated with the first embedding model, such that a second embedding model generates a different set of embeddings than the first embedding model. Furthermore, in some examples, the set of embeddings can be uniquely associated with a version of the first embedding model, such that different versions of the first embedding model generate different sets of embeddings.
[0090] The embedding set may include embeddings generated by the first embedding model for at least one of the plurality of content data. For example, the content data may correspond to an email, an audio file, a message, a website page, etc. Persons of ordinary skill in the art will recognize additional and / or alternative types of content objects or items corresponding to the content data.
[0091] At operation 512, one or more embeddings in the set of embeddings are determined to be landmark memory embeddings. In some examples, one or more of the embeddings are ranked, weighted, and / or scored to determine whether they are landmark memory embeddings. The ranking, weighting, and / or scoring can be based on dissimilarity to at least one other embedding in the set of embeddings. In some examples, dissimilarity can be calculated based on a distance formula (e.g., in a vector space such as vector space 400), textual comparison, visual comparison, and / or other comparison techniques that will be recognized by one of ordinary skill in the art.
[0092] In some examples, the determination at operation 512 includes obtaining a machine learning model previously trained to identify a landmark memory. The set of embeddings can be provided to the machine learning model. A score corresponding to the degree to which the embedding is a landmark memory embedding can be received from the machine learning model. The score can be used to rank and / or weight the embeddings (e.g., using the mechanisms described herein). Additionally or alternatively, the score can be compared to a threshold to identify one or more embeddings in the set of embeddings as being a landmark memory embedding.
[0093] Typically, users are more likely to remember atypical content information than regular content information. Therefore, the mechanisms described herein can be used to determine and / or label embeddings generated based on atypical data as landmark memory embeddings. Based on determining the landmark embedded memory, data related to the landmark embedded memory can also be determined (e.g., based on time, location, person, or other content data corresponding to the landmark embedded memory), so that the landmark memory embedding can be a pivot point for searching for data related thereto. Atypical data is irregular or abnormal data. For example, a data segment that appears less than 1% of the time within a data category (such as a person, event, location, email, etc.) that can classify the data segment can be determined to be abnormal data. In some examples, the 1% threshold can alternatively be 5%, or 10%, or 20%, or any value or range of values therebetween.
[0094] Some examples include receiving an input (e.g., user input) corresponding to a degree or threshold of episodic memory. The input can be received from any of a variety of inputs, such as a button, gaze input, voice input, text input, gesture input, touchpad, mouse, slider, etc. In addition, the user input can be received from a graphical user interface. In some examples, the ranking, weighting, and / or score of the embedding can be compared with a threshold of episodic memory to determine which embedding of one or more embeddings from the set of embeddings is a landmark memory embedding. For example, if the ranking of a particular embedding is above the threshold of episodic memory, the particular embedding can be a landmark memory embedding. Alternatively, in some examples, if the ranking of a particular embedding is below the degree of episodic memory, the particular embedding can be a landmark memory embedding. One of ordinary skill in the art will recognize additional and / or alternative examples for comparing a metric of a landmark memory embedding to a threshold.
[0095] At operation 514, one or more marker memory embeddings are inserted into the context object memory. In some examples, the context object memory can be an index, or a database, or a tree (e.g., an ANN tree, a kd tree, etc.), or other types of memory that one of ordinary skill in the art will recognize. The context object memory includes one or more embeddings (e.g., marker memory embeddings) from a set of embeddings. In addition, one or more embeddings can be associated with corresponding indications that correspond to the location of source data associated with the one or more embeddings. The source data can include one or more of an audio file, a text file, an image file, a video file, and / or a website page. One of ordinary skill in the art will recognize additional and / or alternative types of source data. The embeddings can take up relatively little memory. For example, each embedding can be a 64-bit hash. In contrast, the source data can take up a relatively large amount of memory. Therefore, by storing an indication of the location of the source data in the embedding object memory, as opposed to the original source data, the mechanism described herein can be relatively efficient for memory usage of one or more computing devices.
[0096] The indication corresponding to the location of the source data can be a byte address, a uniform resource indicator (e.g., a uniform resource link), or other form of data capable of identifying the location of the source data. In addition, the context object memory can be stored at a location different from the location of the source data. Additionally or alternatively, the source data can be stored in a memory of a computing device or server where the context object memory (e.g., context object memory 310) is located. Additionally or alternatively, the source data can be stored in a memory remote from the computing device or server where the context object memory is located.
[0097] In some examples, a subset of the ranked embeddings can be inserted into the context object memory based on a comparison of their ranking (or weighting) to a predetermined threshold. One or more ranked embeddings can be stored in the context object memory along with an indication of the corresponding ranking (or weighting). In addition, the ranked embeddings can be associated with references (e.g., pointers) to related data that are associated with the one or more ranked embeddings (e.g., embeddings similar to the ranked embeddings, source data associated with the ranked embeddings). The token memory embeddings described herein can include a set of attributes that define a pattern. The set of attributes can include a summary of the token memory embedding (e.g., a natural language summary of the token) and / or a reference (e.g., a pointer) to data related to the token memory embedding.
[0098] In some examples, the context object memory stores only landmark memory embeddings. Additionally or alternatively, in some examples, the context object memory stores embeddings that are landmark memory embeddings and embeddings that are not landmark memory embeddings, such as with indications corresponding to which embeddings are landmark memory embeddings and which are not landmark memory embeddings.
[0099] In some examples, the insertion at operation 514 triggers a spatial storage operation to store a vector representation of the one or more marker memory embeddings. The vector representation can be stored in a data structure such as an ANN tree, a kd-tree, an n-dimensional tree, an octree, or other data structures that one of ordinary skill in the art would recognize based on the teachings described herein. One of ordinary skill in the art would recognize additional and / or alternative types of storage mechanisms capable of storing vector space representations.
[0100] At operation 514, a context object memory is provided. For example, the context object memory can be provided as an output for further processing to occur. In some examples, the context object memory can be used to retrieve information and / or generate actions, as will be described in further detail in some examples herein.
[0101] In general, embeddings can be used as quantitative representations of abstract meaning discerned from content data. Thus, different embedding models can generate different embeddings for the same received content data, such as depending on how the embedding model is trained (e.g., based on different datasets and / or training methods) or how the embedding model is configured. In some examples, some embedding models can be trained to provide a relatively broad interpretation of the received content, while other embedding models are trained to provide a relatively narrow interpretation of the content data. Thus, such embedding models can generate different embeddings.
[0102] Method 500 may terminate at operation 516. Alternatively, method 500 may return to operation 502 (or any other operation of method 500) to provide an iterative loop, such as receiving one or more content items and inserting one or more landmark memory embeds into the context object memory based thereon.
[0103] The operations provided in method 500 provide for selective construction, storage, and subsequent retrieval (e.g., see Figure 7 The discussed method 700) is a means for providing humans and machines with efficient options (given the limited storage and processing capabilities of cognitive systems (human or machine)) for storing and then implementing powerful handles or indexes that enable bounded reasoning and recall systems to perform efficient context-sensitive retrieval of relevant content for real-time use (e.g., to maximize the success rate of the human or machine). Such mechanisms are particularly important for cognitive systems (whether human or machine) that interact with the complex world in an ongoing, lifelong manner, continuously accumulating and leveraging experience.
[0104] Figure 5BDetailed examples of operation 512 according to some aspects described herein are shown. For example, the detailed example of operation 512 can be a method for determining that one or more embeddings in a set of embeddings are marker memory embeddings. Although operations 550, 552, 554, and 556 described below are discussed as sub-operations of operation 512 of method 500, those skilled in the art will recognize that operations 550, 552, 554, and / or 556 can be performed independently of method 500 and / or in conjunction with methods other than method 500.
[0105] At operation 550, a machine learning model is obtained that was previously trained to identify landmark memories. For example, the machine learning model can be trained based on a dataset of embeddings and a metric corresponding to the extent to which embeddings from the dataset of embeddings are landmark memories.
[0106] At operation 552, the embedding set is provided to the machine learning model. The embedding set may include Figure 5A Additionally or alternatively to the embeddings received at operation 510, the set of embeddings may be a set of embeddings received using mechanisms that will be recognized by one of ordinary skill in the art.
[0107] At operation 554, a metric is received from the machine learning model. The metric corresponds to the degree to which one or more embeddings from the set of embeddings are signature memory embeddings. The metric can be a score, weight, and / or ranking for each of the one or more embeddings.
[0108] At operation 556, the metric can be compared to a threshold to identify one or more embeddings in the set of embeddings as being a marker memory embedding. For example, the threshold can be a threshold provided via user input and / or otherwise preconfigured. In some examples, the metric must be greater than the threshold. In some examples, the metric must be less than the threshold. In some examples, the metric must be equal to the threshold. When comparing the metric to the threshold to identify one or more embeddings in the set of embeddings as being a marker memory embedding, one of ordinary skill in the art will recognize the use of a combination of equality constraints. In some examples, further processing can be performed on the metric such that a derivative of the metric is used to identify one or more embeddings in the set of embeddings as being a marker memory embedding.
[0109] When two or more agents (such as humans and supporting AI systems) work together for an extended period of time, identifying, storing, and manipulating iconic memories corresponding to experiences, learning, and / or other content can be particularly important for the efficient establishment of references to shared experiences. In extended duration or lifelong agents that coordinate or communicate on projects over time, efficient mechanisms may be needed for storing, identifying, jointly referencing, and / or retrieving experiences, learning, and / or other content that arises during sustained, long-term (e.g., lifelong) engagement in a complex world. Examples of content to which iconic memories can correspond are previously referenced herein. Figure 2 discuss.
[0110] As an example, humans can assume that when they are working with an AI system, just as they assume when they are working with a human collaborator, they can leverage the efficiency of communication with the machine to refer to events that humans have naturally encoded and retrieved as iconic memories in episodic memory. To this end, the machine can have a representation of the properties of how humans encode iconic memories so that there can be joint references and / or relationship building from machine to human and / or from human to machine. Therefore, part of finding mutual understanding between machines and humans in long-term effective collaboration can be machine learning to understand and encode iconic memories (e.g., in episodic object memory 310) in the same way that humans encode them. This understanding can enable deeper conversations and problem solving in long-term human-AI collaborative settings, particularly where the AI system is dedicated to supporting a specific user and can experience lived shared sensing, problem solving, and more generally, lived experiences of the specific user.
[0111] Operation 512 may terminate at operation 556. Alternatively, operation 512 may continue from operation 556 to operation 514, whereby the context object memory may be updated based on the determination of operation 512. In some examples, operation 512 may continue from operation 556 to different methods and / or processes as may be appreciated by one of ordinary skill in the art, based at least on the teachings provided herein.
[0112] Figure 6A and Figure 6B An example system 600 for retrieving information from a context object memory is shown. The example system 600 includes a computing device 602. The computing device 602 may be similar to the computing device 602 described previously herein. Figure 1 The system 600 also includes a user interface 604. The user interface 604 may be a graphical user interface (GUI) displayed on a screen of the computing device 602. The user interface 604 may be configured to receive input from a user.
[0113] For example, the GUI 604 includes a slider 606 that can be adjusted based on user input. The input can correspond to a level or threshold of episodic memory. Figure 6A , slider 606 is shown at a first position corresponding to a first degree or threshold of episodic memory. Figure 6B , the slider is shown at a second position corresponding to a second degree or threshold of episodic memory.
[0114] exist Figure 6A At the first degree of episodic memory shown, one or more first indications 608 of one or more landmark memories (e.g., based on the first degree of episodic memory) may be received. In some examples, one or more second indications 610 of stored data that do not correspond to landmark memories may also be received. The one or more first indications 608 and / or the one or more second indications 610 may be presented (e.g., on a display screen of a computing device). In some examples, the one or more first indications 608 and / or the one or more second indications 610 may be symbols, text, images, audio, and / or representations associated with the data corresponding to the indications 608, 610. Additionally, in some examples, time information may be provided to the one or more first indications 608 and / or the one or more second indications 610, corresponding to when the source data associated with the indications 608, 610 was generated and / or stored (e.g., when an email was sent, when a call occurred, when a person was at a given location, etc.).
[0115] exist Figure 6B At the second level of episodic memory shown, the number of one or more first indications 608 and / or second indications 610 may be different than when the slider 606 is at the first level of episodic memory. Thus, the amount of content data determined to be a marker memory may change based on when the designated level of episodic memory is changed.
[0116] can be drawn from episodic memory (such as Figure 3 The contextual memory 310 described above receives one or more first indications 608 and / or second indications 610. For example, contextual memory can be generated using content data that is embedded, ranked, and / or weighted based on semantic context. Content data can be ranked based on dissimilarity between the content data, such as via a distance formula, text comparison, visual comparison, and / or other comparison techniques that will be recognized by one of ordinary skill in the art.
[0117] Memories can be organized to provide efficient handles to large amounts of stored content that are efficiently available by pointing to memories that are landmark memories (e.g., pointing to via first indication 608). Such richer memories (e.g., landmark memories) can be made progressively more detailed by an inefficient process of delving into more extensive memory encodings, such as by changing slider 606. In this case, limited retrieval and reasoning systems can effectively operate at the level of landmark memories and use landmark memories to identify relevant content and make decisions to delve deeper into the less structured and more difficult to retrieve detailed information pointed to by those landmark memories. In this regard, as an example, an indication 608 corresponding to a landmark memory can include information corresponding to a relationship with a second indication 610 that does not correspond to a landmark memory, such that a desired memory associated with the landmark memory can be retrieved or otherwise used to perform an action.
[0118] Figure 7 An example method for retrieving information (such as a landmark memory) from a context object memory according to some aspects described herein is shown. In an example, aspects of method 700 are performed by a device, such as described above with respect to Figure 1 Computing device 102 and / or server 104 are discussed.
[0119] Method 700 begins at operation 702, where a user interface (such as user interface 602) is generated. The user interface may be a graphical user interface (GUI) displayed on a screen of a computing device (e.g., computing device) 602. Alternatively, the user interface may be a touchpad, mechanical buttons, a camera, a microphone, or any other interface configured to receive user input.
[0120] At operation 704, input is received via the user interface. The input corresponds to the level of episodic memory. For example, the user interface may include a slider, and the input may correspond to the position or movement of the slider. The position or movement of the slider may be associated with the level of episodic memory. In other examples, the input may be a voice command, text, an automated skill, a gaze command, a gesture, etc. associated with the level of episodic memory.
[0121] At operation 706, a determination is made as to whether there are search results corresponding to the input from the context object memory. If it is determined that there are no search results corresponding to the input in the context object memory, the process branches "NO" to operation 708, where a default action is performed. For example, the input and / or the context object memory may have an associated pre-configured action. In other examples, method 700 may include determining whether the input and / or the context object memory has an associated default action such that, in some cases, no action may be performed as a result of the received query. Method 700 may terminate at operation 708. Alternatively, method 700 may return to operation 702 or 704 to provide an iterative loop.
[0122] However, if it is determined that a search result corresponding to the input exists in the context object memory, the flow instead branches "yes" to operation 710, where one or more indications of landmark memories are received based on the extent of the context object memory. For example, the one or more indications of landmark memories may be similar to the first indication 608 discussed with respect to FIG. 6. Additionally or alternatively, these indications may be provided as output (such as in a report), which may be used for further processing.
[0123] In some examples, after receiving an indication of one or more landmark memories, the mechanisms provided herein can be used to search for data associated with the landmark memory. For example, if a particular landmark memory is associated with content data at a particular moment in time, other content data within a specified time range from that particular landmark memory can be searched (e.g., based on a query that can be provided by a user). Typically, once a landmark memory is located, it can act as a pivot point from which other data referenced in the memory location of the landmark memory and / or otherwise determined to be associated with the landmark memory can be searched.
[0124] In some examples, a context object store is generated using content data that is embedded, ranked, weighted, and / or scored based on a semantic context. Embeddings generated based on the content data can be ranked, weighted, and / or scored based on the dissimilarity of the embedding to at least one other embedding. The content data can include one or more of audio content data, visual content data, gaze content data, weather content data, news content data, calendar content data, email content data, or location content data. Furthermore, the context object store can be stored at a location different from the location of the source data corresponding to the content item.
[0125] In some examples, the input is a first input, the level of contextual memory is a first level of contextual memory, the indication is a first indication, and method 700 further includes receiving a second input corresponding to a second level of contextual memory. Subsequently, a second indication of one or more landmark memories corresponding to the second input can be received from the contextual object memory based on the second level of contextual memory.
[0126] In some examples, landmark memories give the system the ability to establish a dialogue (e.g., with a person or another system) in a conversational chat setting that will help narrow down which landmark memory among multiple landmark memories the person or system is referencing. In some examples, the dialogue can help clarify ambiguity when retrieving information associated with landmark memories. For example, when asked about a restaurant experience, the system can provide a prompt such as "Do you mean you were last at restaurant S on September 12, or when you were with person A and person B?" Therefore, the system provided herein can request feedback to obtain information for determining which landmark memory (e.g., in a context object memory) is intended to be located. In some examples, the mutually shared handles of the significant landmark memory library enable a dialogue with an AI system that can access a model of the context memory of an agent (e.g., a person or system). In such an example, the same language can be used to refer to the context object memory of an agent shared (at least partially or approximately) with the AI system.
[0127] Method 700 may terminate at operation 710. Alternatively, method 700 may return to operation 702 (or any other operation of method 700) to provide an iterative loop, such as generating a user interface, receiving a query having information corresponding to at least two different content types, and receiving search results corresponding to the query, such as from a context object store.
[0128] Figure 8A and Figure 8B An overview of an example generative machine learning model that can be used in accordance with aspects described herein is shown. Figure 8A , conceptual diagram 800 depicts an overview of a pre-trained generative model package 804 according to aspects described herein, which processes input 802 to generate model output for storing entries (e.g., token memory) in context object memory 806 and / or retrieving information (e.g., token memory) from context object memory 806. Examples of pre-trained generative model packages 604 include, but are not limited to, the Megatron-Turing Natural Language Generation model (MT-NLG), Generative Pre-Trained Transformer 3 (GPT-3), Generative Pre-Trained Transformer 4 (GPT-4), BigScience BLOOM (Big Open Science Open Access Multilingual Language Model), DALL-E, DALL-E2, StableDiffusion, or Jukebox.
[0129] In the example, the generative model package 804 is pre-trained based on a variety of inputs (e.g., various human languages, various programming languages, and / or various content types) and therefore does not require fine-tuning or training for specific scenarios. Instead, the generative model package 804 can be pre-trained more generally such that the input 802 includes prompt words that are generated, selected, or otherwise designed to induce the generative model package 804 to produce certain generative model outputs 806. For example, the prompt word includes context and / or one or more completion prefixes and the generative model package 804 is pre-loaded accordingly. Thus, the generative model package 804 is induced to generate an output based on the prompt word that includes a predicted sequence of word segments related to the prompt word (e.g., up to the word segmentation limit of the generative model package 804). In the example, the predicted word segmentation sequence is further processed (e.g., by output decoding 816) to produce the output 806. For example, each word segmentation is processed to identify a corresponding word, word fragment, or other content that forms at least a portion of the output 806. It should be understood that the input 802 and the generative model output 806 can each include any of a variety of content types, including but not limited to text output, image output, audio output, video output, programmatic output, and / or binary output, etc. In an example, the input 802 and the generative model output 806 can have different content types, such as may be the case when the generative model package 804 includes generating a multimodal machine learning model.
[0130] Thus, the generative model package 804 can be used in any of a variety of contexts and, further, can be used without substantially modifying other associated aspects (e.g., similar to those described herein with respect to Figures 1 to 7 In the absence of those aspects described herein, a different generative model package may be used in place of generative model package 804. Thus, generative model package 804 operates as a tool for performing machine learning processing, wherein certain inputs 802 to generative model package 804 are programmatically generated or otherwise determined, causing generative model package 804 to produce model output 806, which may then be used for further processing.
[0131] The generative model package 804 can be provided or otherwise used according to any of a variety of paradigms. For example, the generative model package 804 can be provided or otherwise used on a computing device (e.g., Figure 1 In some examples, generative model package 804 may be used locally on computing device 102 or may be accessed remotely from a machine learning service. In other examples, aspects of generative model package 804 are distributed across multiple computing devices. In some examples, generative model package 804 may be accessed via an application programming interface (API), such as may be provided by an operating system of a computing device and / or by a machine learning service, among other examples.
[0132] Referring now to the illustrated aspects of the generative model package 804, the generative model package 804 includes input tokenization 808, input embedding 810, model layer 812, output layer 814, and output decoding 816. In the example, input tokenization 808 processes input 802 to generate input embedding 810, which includes a sequence of symbolic representations corresponding to input 802. Thus, input embedding 810 is processed by model layer 812, output layer 814, and output decoding 816 to produce model output 806. Figure 8B An example architecture corresponding to the generative model package 804 is depicted in , which is discussed in further detail below. Even so, it should be understood that the architecture shown and described herein should not be considered limiting, and in other examples, any of a variety of other architectures may be used.
[0133] Figure 8B 850 is a conceptual diagram depicting an example architecture 850 of a pre-trained generative machine learning model that can be used in accordance with aspects described herein. As described above, any of a variety of alternative architectures and corresponding ML models can be used in other examples without departing from aspects described herein.
[0134] As shown, architecture 850 processes input 802 to produce generative model output 806, aspects of which are described above with respect to Figure 8A Discussion. Architecture 850 is depicted as a transformer model comprising an encoder 852 and a decoder 854. The encoder 852 processes an input embedding 858 (which may be similar in aspects to Figure 8A ), which includes a sequence of symbolic representations corresponding to input 856. In an example, input 856 includes input content 802 corresponding to a content type, aspects of which can be similar to input data 111, virtual content 200, and / or real content 250.
[0135] Furthermore, positional encoding 860 can incorporate information about the relative and / or absolute positions of the word segments of input embedding 858. Similarly, output embedding 874 includes a sequence of symbolic representations corresponding to output 872, and positional encoding 876 can similarly incorporate information about the relative and / or absolute positions of the word segments of output embedding 874.
[0136] As shown, the encoder 852 includes an example layer 870. It should be understood that any number of such layers can be used, and the depicted architecture is simplified for illustrative purposes. The example layer 870 includes two sublayers: a multi-head attention layer 862 and a feed-forward layer 866. In the example, residual connections are included around each layer 862, 866, followed by normalization layers 864 and 868, respectively.
[0137] The decoder 854 includes an example layer 890. Similar to the encoder 852, any number of such layers may be used in other examples, and the architecture of the decoder 854 depicted is simplified for illustrative purposes. As shown, the example layer 890 includes three sublayers: a masked multi-headed attention layer 878, a multi-headed attention layer 882, and a feedforward layer 886. The aspects of the multi-headed attention layer 882 and the feedforward layer 886 can be similar to those discussed above with respect to the multi-headed attention layer 862 and the feedforward layer 866, respectively. In addition, the masked multi-headed attention layer 878 performs multi-headed attention on the output of the encoder 852 (e.g., output 872). In the example, the masked multi-headed attention layer 878 prevents certain positions from noticing subsequent positions. This mask combined with an offset embedding (e.g., offsetting one position, as shown in the multi-headed attention layer 882) can ensure that the prediction for a given position depends on the known output of one or more positions that are smaller than the given position. As shown, residual connections are also included around layers 878, 882, and 886, followed by normalization layers 880, 884, and 888, respectively.
[0138] The multi-head attention layers 862, 878, and 882 can each linearly project the query, key, and value to the corresponding dimension using a set of linear projections. Each linear projection can be processed using an attention function (e.g., dot product or additive attention), thereby producing an n-dimensional output value for each linear projection. The resulting values can be concatenated and projected again, so that the value is then as follows Figure 8B , as shown, is processed (e.g., by a corresponding normalization layer 864, 880, or 884).
[0139] Feedforward layers 866 and 886 can each be a fully connected feedforward network applied to each position. In an example, feedforward layers 866 and 886 each include multiple linear transformations with rectified linear unit activations therebetween. In an example, each linear transformation is the same at different positions, and each linear transformation can use different parameters compared to other linear transformations in the feedforward network.
[0140] Additionally, aspects of the linear transformation 892 can be similar to the linear transformations discussed above with respect to the multi-head attention layers 862, 878, and 882 and the feed-forward layers 866 and 886. Softmax 894 can also convert the output of the linear transformation 892 into a predicted probability of the next word segmentation, as shown by output probability 896. It should be understood that the illustrated architecture is provided as an example, and in other examples, any of a variety of other model architectures can be used in accordance with the disclosed aspects. In some cases, multiple iterations of processing (e.g., using Figure 8A Generative model package 804 or Figure 8B852 and decoder 854 in ) to generate a series of output tokens (e.g., words), which are then combined to produce a complete sentence (and / or any of a variety of other contents). It should be understood that other generative models can generate multiple output tokens in a single iteration and thus can use a reduced number of iterations or a single iteration.
[0141] Thus, the output probabilities 896 can form the embedding output 806 according to aspects described herein, such that the output of the generative ML model (e.g., which can include structured output) is used as input for determining an action according to aspects described herein. In other examples, the embedding output 806 is provided as a generated output for updating the context object memory.
[0142] Figures 9-11 and the associated descriptions provide a discussion of various operating environments in which aspects of the present disclosure may be practiced. Figures 9-11 The devices and systems shown and discussed are for purposes of example and explanation and are not limiting of the wide variety of computing device configurations that can be used to practice the aspects of the present disclosure described herein.
[0143] Figure 9 is a block diagram illustrating the physical components (eg, hardware) of a computing device 900 that can practice aspects of the present disclosure. The computing device components described below may be applicable to the computing devices described above, including Figure 1 1. In a basic configuration, computing device 900 may include at least one processing unit 902 and system memory 904. Depending on the configuration and type of computing device, system memory 904 may include, but is not limited to, volatile storage (e.g., random access memory), non-volatile storage (e.g., read-only memory), flash memory, or any combination of these memories.
[0144] System memory 904 may include an operating system 905 and one or more program modules 906 suitable for running software applications 920, such as one or more components supported by the system described herein. As an example, system memory 904 may store a context object memory insertion engine or component 924 and / or a context object memory retrieval engine or component 926. Operating system 905 may, for example, be suitable for controlling the operation of computing device 900.
[0145] Furthermore, aspects of the present disclosure may be practiced in conjunction with graphics libraries, other operating systems, or any other application programs and are not limited to any particular application or system. Figure 9908. The computing device 900 may have additional features or functionality. For example, the computing device 900 may also include additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or tapes. Such additional storage Figure 9 909 and non-removable storage device 910.
[0146] As described above, a number of program modules and data files can be stored in system memory 904. When executed on processing unit 902, program modules 906 (e.g., applications 920) can perform processes including, but not limited to, the aspects described herein. Other program modules that can be used in accordance with aspects of the present disclosure can include email and contact applications, word processing applications, spreadsheet applications, database applications, slide presentation applications, drawing or computer-assisted applications, and the like.
[0147] Furthermore, aspects of the present disclosure may be practiced on circuits comprising discrete electronic components, packaged or integrated electronic chips containing logic gates, circuits utilizing a microprocessor, or a single chip containing electronic components or a microprocessor. For example, aspects of the present disclosure may be practiced via a system on a chip (SOC) wherein Figure 9 Each or many components shown in can be integrated onto a single integrated circuit. Such a SOC device may include one or more processing units, a graphics unit, a communication unit, a system virtualization unit, and various application functions, all of which are integrated (or "burned") onto a chip substrate as a single integrated circuit. When operated via the SOC, the functions described herein regarding the ability to switch protocols for the client can be operated via dedicated logic integrated with other components of the computing device 900 on a single integrated circuit (chip). Some aspects of the present disclosure may also be practiced using other technologies capable of performing logical operations (e.g., AND, OR, and NOT), including but not limited to mechanical, optical, fluid, and quantum technologies. In addition, some aspects of the present disclosure may be practiced within a general-purpose computer or in any other circuit or system.
[0148] The computing device 900 may also have one or more input devices 912, such as a keyboard, a mouse, a pen, an audio or voice input device, a touch or slide input device, and the like. Output device(s) 914 may also be included, such as a display, speakers, a printer, and the like. The above devices are examples, and other devices may also be used. The computing device 900 may include one or more communication connections 916 that allow communication with other computing devices 950. Examples of suitable communication connections 916 include, but are not limited to, radio frequency (RF) transmitters, receivers, and / or transceiver circuits; universal serial bus (USB), parallel, and / or serial ports.
[0149] As used herein, the term "computer-readable media" may include computer storage media. Computer storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as, computer-readable instructions, data structures, or program modules). System memory 904, removable storage device 909, and non-removable storage device 910 are all examples of computer storage media (e.g., memory storage). Computer storage media may include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article that can be used to store information and can be accessed by the computing device 900. Any such computer storage media may be part of the computing device 900. Computer storage media does not include carrier waves or other propagated or modulated data signals.
[0150] Communication media may be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" may describe a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0151] Figure 10 10 is a block diagram illustrating an architecture of one aspect of a computing device. That is, a computing device can incorporate system 1002 (e.g., architecture) to implement some aspects. In some examples, system 1002 is implemented as a "smartphone" capable of running one or more applications (e.g., a browser, email, calendar, contact manager, messaging client, games, and media client / player). In some aspects, system 1002 is integrated into a computing device, such as an integrated personal digital assistant (PDA) and wireless phone.
[0152] One or more application programs 1066 can be loaded into memory 1062 and run on or in association with operating system 1064. Examples of application programs include a phone dialer, an email program, a personal information management (PIM) program, a word processing program, a spreadsheet program, an Internet browser program, a messaging program, and the like. System 1002 also includes a non-volatile storage area 1068 within memory 1062. Non-volatile storage area 1068 can be used to store persistent information that should not be lost if system 1002 loses power. Application programs 1066 can use and store information in non-volatile storage area 1068, such as emails or other messages used by an email application. A synchronization application (not shown) also resides on system 1002 and is programmed to interact with a corresponding synchronization application residing on a host computer to keep information stored in non-volatile storage area 1068 synchronized with corresponding information stored on the host computer. It should be understood that other applications may be loaded into the memory 1062 and executed on the mobile computing device 1000 described herein (eg, an embedded object memory insertion engine, an embedded object memory retrieval engine, etc.).
[0153] The system 1002 has a power source 1070, which can be implemented as one or more batteries. The power source 1070 can also include an external power source, such as an AC adapter or a powered docking station to replenish or recharge the batteries.
[0154] System 1002 may also include a radio interface layer 1072 that performs the functions of transmitting and receiving radio frequency communications. Radio interface layer 1072 supports wireless connectivity between system 1002 and the "outside world" via a communications carrier or service provider. Transmissions to and from radio interface layer 1072 are controlled by operating system 1064. In other words, communications received by radio interface layer 1072 can be passed to application programs 1066 via operating system 1064, and vice versa.
[0155] The visual indicator 1020 can be used to provide a visual notification, and / or the audio interface 1074 can be used to generate an audible notification via the audio transducer 1025. In the example shown, the visual indicator 1020 is a light emitting diode (LED), and the audio transducer 1025 is a speaker. These devices can be directly coupled to the power supply 1070 so that when activated, they remain on for a duration indicated by the notification mechanism, even if the processor 1060 and / or the dedicated processor 1061 and other components may be turned off to save battery power. The LED can be programmed to remain on indefinitely until the user takes action to indicate the power-on status of the device. The audio interface 1074 is used to provide audible signals to the user and receive audible signals from the user. For example, in addition to being coupled to the audio transducer 1025, the audio interface 1074 can also be coupled to a microphone to receive audible input, such as to support a telephone conversation. According to aspects of the present disclosure, the microphone can also be used as an audio sensor to support control of notifications, as will be described below. The system 1002 may also include a video interface 1076 that enables operation of the onboard camera 1030 to record still images, video streams, and the like.
[0156] The computing device implementing system 1002 may have additional features or functionality. For example, the computing device may also include additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or tapes. Such additional storage Figure 10 denoted by non-volatile storage area 1068.
[0157] As described above, data / information generated or captured by a computing device and stored via system 1002 can be stored locally on the computing device, or the data can be stored on any number of storage media that can be accessed by the device via the radio interface layer 1072 or via a wired connection between the computing device and an independent computing device associated with the computing device (e.g., a server computer in a distributed computing network such as the Internet). It should be understood that such data / information can be accessed by the computing device via the radio interface layer 1072 or via a distributed computing network. Similarly, such data / information can be easily transferred between computing devices for storage and use according to well-known data / information transmission and storage means (including email and collaborative data / information sharing systems).
[0158] Figure 11One aspect of the architecture of a system as described above for processing data received at a computing system from a remote source, such as a personal computer 1104, a tablet computing device 1106, or a mobile computing device 1108, is shown. The content displayed at the server device 1102 can be stored in different communication channels or other storage types. For example, a directory service 1124, a web portal 1125, a mailbox service 1126, an instant messaging repository 1128, or a social networking site 1130 can be used to store various documents.
[0159] Application 1120 (e.g., similar to application 920) can be employed by a client communicating with server device 1102. Additionally or alternatively, a context object storage insertion engine 1121 and / or a context object storage retrieval engine 1122 can be employed by server device 1102. Server device 1102 can provide data to and from client computing devices such as personal computers 1104, tablet computing devices 1106, and / or mobile computing devices 1108 (e.g., smartphones) via network 1115. As examples, the aforementioned computer system can be embodied in personal computers 1104, tablet computing devices 1106, and / or mobile computing devices 1108 (e.g., smartphones). Any of these examples of computing devices can retrieve content from repository 1116 and can also receive graphics data that can be used for pre-processing at a graphics generation system or post-processing at a receiving computing system.
[0160] For example, various aspects of the present disclosure are described above with reference to block diagrams and / or operational diagrams of methods, systems, and computer program products according to various aspects of the present disclosure. The functions / actions indicated in the blocks may not occur in the order shown in any flowchart. For example, depending on the functions / actions involved, two blocks shown in succession may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order.
[0161] The description and explanation of one or more aspects provided in this application are not intended to limit or restrict the scope of this disclosure in any way. The aspects, examples and details provided in this application are considered to be sufficient to convey ownership and enable others to make and use the aspects claimed in this disclosure. The disclosure claimed should not be interpreted as being limited to any aspect, example or details provided in this application. Whether shown and described in combination or individually, various features (structure and method) are intended to be selectively included or omitted, to produce an embodiment with a specific feature set. In the case of providing the description and explanation of this application, those skilled in the art can envision the variations, modifications and alternative aspects of the spirit that fall within the broader aspects of the overall inventive concept embodied in this application, and these variations, modifications and alternative aspects do not depart from the broader scope of the disclosure claimed.
Claims
1. A method for embedding a landmark memory in a context object memory, the method comprising: receiving one or more content items, each of the content items having one or more content data; providing the one or more content data associated with the one or more content items to one or more embedding models, wherein the one or more embedding models generate one or more embeddings; receiving a set of embeddings from one or more of the embedding models, wherein each embedding in the set of embeddings corresponds to at least one content data from a corresponding content item; determining that one or more embeddings in the set of embeddings are landmark memory embeddings; inserting the one or more landmark memory embeddings into the context object memory, wherein the one or more landmark memory embeddings are associated with references to associated data associated with the one or more landmark memory embeddings; and The context object storage is provided.
2. The method of claim 1 , wherein the inserting triggers a spatial storage operation to store a vector representation of the landmark memory embedding, and wherein the vector representation is stored in at least one of an approximate nearest neighbor (ANN) tree, a k-d tree, or a multidimensional tree.
3. The method of claim 1 , wherein the determining comprises ranking the one or more of the embeddings based on dissimilarity to at least one other of the other embeddings of the set of embeddings, and wherein the inserting comprises storing an indication of the ranking corresponding to the landmark memory embedding.
4. The method of claim 1 , wherein the determining comprises: Obtaining a previously trained machine learning model to identify landmark memories; providing the set of embeddings to the machine learning model; receiving, from the machine learning model, a score corresponding to the extent to which the embedding is a landmark memory embedding; The score is compared to a threshold to identify one or more embeddings in the set of embeddings as a landmark memory embedding.
5. The method according to claim 4, further comprising, before the comparing: A user input corresponding to the threshold is received, wherein the threshold is a threshold of the context memory. The method of claim 5 , wherein the user input is received from a slider of a graphical user interface.
7. The method of claim 1, wherein the content data is one or more of: audio content data, visual content data, gaze content data, weather content data, news content data, calendar content data, email content data, or location content data.
8. The method of claim 1, wherein the context object store is stored at a location different from a location of source data corresponding to the content item.
9. The method of claim 1, wherein the one or more landmark memory embeddings comprise a set of attributes defining a pattern.
10. A method for retrieving a landmark memory from a context object memory, the method comprising: Generate user interface; receiving an input via the user interface, wherein the input corresponds to a degree of episodic memory, receiving, from said contextual memory, an indication of one or more landmark memories based on said degree of contextual memory, The context object memory is generated based on semantic context using embedded and ranked content data.
11. The method of claim 10, wherein the embeddings generated based on the content data are ranked based on a dissimilarity of the embedding to at least one other embedding.
12. The method of claim 10, wherein the input is a first input, wherein the degree is a first degree, wherein the indication is a first indication, and wherein the method further comprises: receiving a second input, wherein the second input corresponds to a second degree of episodic memory; as well as Based on the second degree of contextual memory, a second indication of one or more landmark memories corresponding to the second input is received from the contextual object memory.
13. The method of claim 12, wherein the user interface comprises a slider, and wherein the first input and the second input correspond to movement of the slider.
14. A method for storing and retrieving a landmark memory in and from a context object memory, the method comprising: receiving one or more content items, each of the content items having one or more content data; providing the one or more content data associated with the one or more content items to one or more embedding models, wherein the one or more embedding models generate one or more embeddings; receiving a set of embeddings from one or more of the embedding models, wherein each embedding in the set of embeddings corresponds to at least one content data from a corresponding content item; determining that one or more embeddings in the set of embeddings are landmark memory embeddings; inserting the one or more landmark memory embeds into the context object memory; as well as receiving an input, wherein the input corresponds to a degree of episodic memory, Based on the degree of contextual memory, an indication of one or more landmark memories is received from the contextual object memory.
15. The method of claim 14, wherein the inserting triggers a spatial storage operation to store a vector representation of the landmark memory embedding, and wherein the vector representation is stored in at least one of an approximate nearest neighbor (ANN) tree, a kd-tree, or a multidimensional tree.
Citation Information
Cited By
Memory data processing method and device, storage medium and computer program product
CN121562658A