Artificial intelligence memory architecture, system and method

WO2026199042A1PCT designated stage Publication Date: 2026-10-01ANDON TECH PTY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/AU2026/050350
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-06-26
Filing Date
2026-04-12
Publication Date
2026-10-01

Smart Images

  • Figure AU2026050350_01102026_PF_FP_ABST
    Figure AU2026050350_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented system (100) is disclosed for constructing artificial-intelligence prompts using semantically organised memory derived from a plurality of source texts. An ingestion engine (120) receives the source texts and segments them into discrete memory units using NLP. An embedding engine (128) generates a respective vector representation for each discrete memory unit, and a tagging engine (124) applies metadata tags to each discrete memory unit. A memory store (130), comprising a vector index (134) and a metadata storage (132), stores the discrete memory units, the respective vector representations, and the metadata tags in a memory space organised according to semantic relationship among the discrete memory units rather than according to grouping by source text, such that semantically related discrete memory units derived from different source texts are stored closer to one another than semantically unrelated discrete memory units derived from a same source text, independently of source-document adjacency.
Need to check novelty before this filing date? Find Prior Art

Description

Artificial Intelligence Memory Architecture, System and Method Technical Field

[0001] The present disclosure broadly relates to interacting with artificial intelligence and, more particularly, to a system for, and a method of, enhancing prompts provided to artificial intelligence models.Background

[0002] Large Language Models (LLMs), such as those based on Transformer architectures, have enabled significant advancements in natural language processing (NLP) and conversational Al. These models are capable of generating coherent and contextually relevant text by utilizing large-scale statistical representations of human language. However, they are inherently stateless and do not retain long-term memory between sessions. As a result, each interaction with an LLM must either include the full conversational history within the model’s context window or forgo continuity altogether. This constraint limits the model’s ability to maintain persistent understanding or reference prior conversations beyond what can be temporarily provided at inference time.

[0003] To address context limitations, current systems typically store past conversations in session logs or databases external to the LLM. When attempting to simulate memory or continuity, these systems often rely on reloading text-based history into the model's prompt, constrained by maximum token limits. However, this approach can lack semantic prioritization, and is generally sensitive to prompt length constraints, and also cannot effectively incorporate deeper understanding such as emotional nuance, decision history, or cross-session continuity. LLMs have no built-in ability to organize, recall, or learn from past user interactions in a structured or adaptive manner, which limits their effectiveness in applications requiring personalization, context retention, or multi-session reasoning.

[0004] Any discussion of documents, acts, materials, devices, articles or the like which has been included in the present specification is not to be taken as an admission that any or all of these matters form part of the prior art base or were common general knowledge in the field relevant to the present disclosure as it existed before the priority date of each claim of this application.Summary

[0005] Current LLM interfaces have limitations in long-term memory and often rely on model-native memory systems. These interfaces do not personalize outputs based on a user’s full interaction history or evolving emotional and cognitive patterns, which affects their effectiveness in long-term productivity, decision-making, and emotionally supportive applications.

[0006] Described herein are systems for external memory management that enhance the personalization and continuity of interactions with LLMs. A persistent memory system enhances interaction with large language models (LLMs) through structured memory storage, dynamic retrieval, and intelligent context injection. The system records and categorizes prior user interactions, emotional states, decisions, semantic content, and the like, allowing future prompts to be informed by historical data. This enables context-aware Al behaviour with minimal model-side memory dependence.

[0007] In one aspect, the invention provides a computer-implemented system comprising: an ingestion engine configured to receive a plurality of source texts and to segment the source texts into discrete memory units using natural language processing; an embedding engine configured to generate, for each discrete memory unit, a respective vector representation; a tagging engine configured to apply one or more metadata tags to each discrete memory unit; a memory store comprising a vector index and a metadata storage, the memory store being configured to store the discrete memory units, the respective vector representations, and the metadata tags in a memory space organised according to semantic relationship among the discrete memory units rather than according to grouping by source text, such that semantically related discretememory units derived from different source texts are stored closer to one another in the memory space than semantically unrelated discrete memory units derived from a same source text, independently of source-document adjacency; a retrieval engine configured, in response to a user query, to retrieve from the memory store a subset of the discrete memory units, the subset being assembled across a plurality of source texts for inclusion in a single artificial intelligence prompt; and a prompt construction module configured to construct the artificial intelligence prompt using the user query and the retrieved subset of discrete memory units.

[0008] In some embodiments, the memory store is configured to disregard original document boundaries when storing the discrete memory units, such that discrete memory units derived from a same source text are stored in different regions of the memory space when they are semantically unrelated.

[0009] In some embodiments, the source texts comprise one or more of emails, notes, files, and chat messages. In some embodiments, the natural language processing comprises one or more of syntactic parsing, topic modelling, entity recognition, and sentence boundary detection. In some embodiments, the metadata tags comprise one or more of timestamps, detected entities, topic tags, emotional tone, decision points, action items, and temporal markers.

[0010] In some embodiments, the retrieval engine is configured to rank candidate discrete memory units for retrieval based on one or more of semantic similarity to the user query, recency of the associated source text, metadata, historical usage, and user feedback.

[0011] In some embodiments, the prompt construction module comprises a token budgeting submodule configured to estimate a total token count for the artificial intelligence prompt and to truncate or compress lower-priority discrete memory units when a maximum prompt size is exceeded.

[0012] In some embodiments, the system further comprises a memory review interface configured to permit one or more of annotation, modification, deletion, and reorganisation of stored discrete memory units, wherein information representative of such user actions is stored in the memory store and is usable by the retrieval engine in retrieval or ranking of discrete memory units.

[0013] In some embodiments, the system further comprises a learning module configured to adjust a priority ranking of retrieved discrete memory units based on one or more of historical usage and user feedback.

[0014] In some embodiments, the memory store is further configured to store a summary memory entry as an independent retrievable memory unit in the memory space, the summary memory entry being a newly derived memory unit representing a compressed synthesis of a plurality of prior conversational exchanges, stored with one or more metadata tags and retrievable alongside semantically related discrete memory units derived from other source texts.

[0015] In another aspect, the invention provides a computer-implemented method comprising: receiving, at a conversation interface, a sequence of messages in a conversation; monitoring whether the sequence of messages has reached a predefined threshold number of exchanges; upon the threshold number of exchanges being reached, generating, by a summarisation module, a summary memory entry comprising a newly derived memory unit representing a compressed synthesis of content of the sequence of messages, the summary memory entry being distinct from a transcript of the underlying message sequence; storing the summary memory entry in a memory store as an independent retrievable memory item, the memory store being organised according to semantic relationship among stored memory units rather than according to source-text grouping, such that the summary memory entry is retrievable alongside semantically related discrete memory units derived from other source texts; and, in response to a later user query, retrieving the summary memory entry from the memory store based on semantic similarity between the user query and the summary memory entry, without reprocessing the underlying message sequence.

[0016] In some embodiments, the predefined threshold number of exchanges is four or more messages. In some embodiments, the method further comprises storing one or more metadata tags with the summary memory entry, the metadata tags comprising one or more of topic tags, detected entities, decision points, action items, emotional states, and temporal markers represented in the sequence of messages.

[0017] In some embodiments, the method further comprises constructing an artificial intelligence prompt using the user query and the retrieved summary memory entry, and submitting the artificial intelligence prompt to an artificial intelligence model.

[0018] In some embodiments, the method further comprises: receiving a plurality of source texts; segmenting the source texts into discrete memory units using natural language processing; generating, for each discrete memory unit, a respective vector representation and one or more metadata tags; and storing the discrete memory units in the memory store such that semantically related discrete memory units derived from different source texts are stored closer to one another than semantically unrelated discrete memory units derived from a same source text, independently of sourcedocument adjacency.

[0019] In some embodiments, the memory store disregards original document boundaries when storing the discrete memory units. In some embodiments, retrieving the summary memory entry comprises assembling a retrieval subset comprising the summary memory entry and discrete memory units derived from different source texts for inclusion in a single artificial intelligence prompt.

[0020] In some embodiments, candidate memory units are ranked for retrieval based on one or more of semantic similarity to the user query, recency, metadata, historical usage, and user feedback. In some embodiments, the method further comprises adjusting a priority ranking of retrieved memory units based on one or more of historical usage and user feedback received via a memory review interface.

[0021] In some embodiments, the method further comprises estimating a token count for an artificial intelligence prompt constructed from retrieved memory units and truncating or compressing lower-priority memory units when a maximum prompt size is exceeded.

[0022] In another aspect, the invention provides a computer-implemented system comprising a memory store configured to store a plurality of discrete memory units derived from a plurality of source texts, each discrete memory unit being associated with a respective vector representation, wherein the discrete memory units are stored in the memory store such that semantically related discrete memory units from different source texts are arranged according to semantic closeness rather than according to source-text grouping.

[0023] In some embodiments, the system further comprises an ingestion engine configured to derive the discrete memory units by segmenting the source texts into granular, meaning-preserving units.

[0024] In some embodiments, semantically related discrete memory units from different source texts are stored closer to one another than semantically unrelated discrete memory units derived from the same source text. In some embodiments, the source texts comprise one or more of emails, notes, files, and chat messages. In some embodiments, the system further comprises an ingestion engine configured to generate the discrete memory units from the source texts. In some embodiments, the ingestion engine is configured to generate vector embeddings associated with the discrete memory units. In some embodiments, the discrete memory units are stored with metadata tags.

[0025] In some embodiments, the metadata tags comprise one or more of timestamps, detected entities, emotional tone, topic tags, decision points, and action items. In some embodiments, the system further comprises a retrieval engine configured to retrieve, in response to a query, a subset of the discrete memory units. In some embodiments, the retrieved subset comprises discrete memory units originating from different sourcetexts. In some embodiments, the retrieval engine retrieves compact, context-specific snippets for use in constructing an artificial intelligence prompt. In some embodiments, the system further comprises a prompt construction module configured to construct the artificial intelligence prompt using the retrieved discrete memory units.

[0026] In another aspect, the invention provides a computer-implemented system comprising a conversation interface configured to receive a sequence of messages in a conversation, a monitoring module configured to determine whether the sequence of messages has reached a threshold number of exchanges, a summarisation module configured, upon the threshold number of exchanges being reached, to generate a summary memory entry representing content of the conversation, and a memory store configured to store the summary memory entry as a retrievable memory item.

[0027] In some embodiments, the threshold number of exchanges is a predefined threshold number of exchanges. In some embodiments, the threshold number of exchanges is four or more messages. In some embodiments, the summary memory entry is stored together with one or more metadata tags. In some embodiments, the system further comprises a retrieval engine configured to retrieve the summary memory entry in response to a query.

[0028] In another aspect, the invention provides a computer-implemented method comprising receiving a sequence of messages in a conversation, determining whether the sequence of messages has reached a threshold number of exchanges, upon the threshold number of exchanges being reached, generating a summary memory entry representing content of the conversation, and storing the summary memory entry in a memory store as a retrievable memory item.

[0029] In some embodiments, the threshold number of exchanges is a predefined threshold number of exchanges. In some embodiments, the threshold number of exchanges is four or more messages. In some embodiments, the method further comprises storing the summary memory entry together with one or more metadata tags.

[0030] In some embodiments, the method further comprises retrieving the summary memory entry from the memory store in response to a query. In another aspect there is provided a computer-implemented system comprising a memory store configured to store a plurality of memory units derived from a plurality of source texts, wherein the memory units are stored according to semantic relationship between the memory units rather than according to grouping by source text.

[0031] In another aspect there is provided a computer-implemented system comprising a memory store configured to store a plurality of memory units derived from source text, wherein semantically related memory units are stored together independently of the source text from which the memory units were derived.

[0032] In some embodiments, the memory units are generated by segmenting source text into granular, meaning-preserving units. In some embodiments, semantically related memory units from different source texts are stored closer to one another than memory units from the same source text that are semantically unrelated. In some embodiments, the source texts comprise emails, notes, files, and / or chat messages. In some embodiments, the memory units are associated with metadata and / or vector embeddings. In some embodiments, a retrieval engine retrieves a subset of the memory units responsive to a query. In some embodiments, retrieved memory units from different source texts are combined for prompt construction for an Al model.

[0033] In another aspect there is provided a system comprising: a structured memory system comprising: an ingestion engine that receives input data, and is configured to convert the received input data to converted data for processing, and to determine and apply metadata tags, thereby creating tagged converted data, each tagged converted datum associated with respective received input data; and a memory store comprising: a metadata database that stores the tagged converted data; and a vector database that indexes vector embeddings associated with the converted data; a retrieval engine configured to: receive a user query; convert the received query to converted query data; query the vector database based on the converted query data; and retrieve converted data from the metadata database; a prompt construction module that receives theretrieved tagged converted data from the metadata database via the retrieval engine, and creates an artificial intelligence prompt based on both the received user query and the retrieved converted data.

[0034] The retrieval engine may be configured to rank retrieved converted data based on one or more of: a similarity of the retrieved converted data to the converted query data, recency of the associated respective received input data, and user feedback.

[0035] The prompt construction module may comprise a token budgeting submodule that estimates total token count and truncates or compresses lower-priority memories if a maximum prompt size is exceeded, thereby balancing token limits with relevance scoring to optimize depth of recall.

[0036] The ingestion engine may comprise: a data conversion engine that converts the received input data to converted data for processing; and a tagging engine that processes the converted data to determine and apply metadata tags, thereby creating the tagged converted data.

[0037] The ingestion engine converting the received input data may comprise segmenting the received input data into discrete memory units using natural language processing, so that each converted datum comprises a discrete memory unit.

[0038] The ingestion engine may comprise an embedding engine that processes the converted data to create vector embeddings associated with the converted data.

[0039] The system may further comprise a memory review user interface configured for user actions comprising one or more of: annotate, modify, delete, and / or reorganize stored converted data, and wherein said user actions are stored in the memory store, retrieved by the retrieval engine, and used by the retrieval engine for ranking retrieved converted data.

[0040] The system may further comprise a learning module configured to adjust a priority ranking of retrieved converted data based on historical usage and / or feedback.

[0041] In another aspect there is provided a computer-implemented method comprising: at an ingestion engine: receiving input data; converting the input data to converted data comprising discrete memory units; and applying metadata tags to form tagged converted data associated with respective received input data; storing, at a memory store, tagged converted data and vector embeddings associated with the converted data; in response to a user query, at a retrieval engine: determining a relevant data subset of the converted data based on one or more of: a semantic match between the user query and the converted data, temporal distance, and user behaviour; and retrieving the relevant data subset from the memory store; creating, by a prompt construction module, an artificial intelligence prompt based on both the user query and the relevant data subset; and submitting said prompt to an artificial intelligence model.

[0042] The method may comprise receiving a user feedback input and applying the user feedback input to determine the relevant data subset.

[0043] Determining the relevant data subset may comprise determining a priority ranking of the converted data within the relevant data subset based on one or more of: a similarity of the converted data to the user query, recency of the converted data, and user feedback.

[0044] The method may comprise adjusting the priority ranking based on historical usage and / or feedback.

[0045]

[0046] These and other aspects and features will now become apparent to those skilled in the art upon review of the following description of specific non-limiting embodiments in conjunction with the accompanying drawings.Brief Description of Drawings

[0047] The detailed description of illustrative and non-limiting embodiments will be more fully appreciated when taken in conjunction with the accompanying drawings in which:

[0048] Figure l is a schematic representation of an embodiment of a system for enhancing conversational interactions with a large language model.

[0049] Figure 2 is a schematic representation of an embodiment of a memory store forming part of a structured memory system.

[0050] Figure 3 is a schematic representation of an embodiment of an ingestion engine.

[0051] Figure 4 is a flow diagram of an embodiment of a method enhancing conversational interactions with a large language model.

[0052] In the drawings, like reference numerals designate similar parts.

[0053] The drawings are not necessarily to scale and may be illustrated by phantom lines, diagrammatic representations and fragmentary views. In certain instances, details that are not necessary for an understanding of the embodiments or that render other details difficult to perceive may have been omitted.Detailed Description

[0054] The following detailed description sets forth illustrative implementations of a system and method for augmenting large language model (LLM) interactions with persistent, structured, and semantically indexed memory. The embodiments described are not intended to limit the invention, but rather to illustrate specific ways in which the invention may be implemented using concrete technical means.

[0055] A Large Language Model (LLM) is a type of Al model, specifically a statistical language model, trained on massive amounts of data. It can understand, generate, and translate human language and perform other natural language processing tasks. LLMs are typically based on deep learning architectures, such as the Transformer model.

[0056] A chatbot is a computer program that simulates human conversation through voice commands or text chats or both. LLMs are used to power conversational Al and chatbots. Some well-known LLMs include GPT models (like those behind ChatGPT), LaMDA, and others. A GPT, or Generative Pre-trained Transformer, is a specific type of LLM. LLMs are a broader category of Al models trained on vast amounts of text data to understand and generate human language. GPT models, developed by OpenAI, are known for their text generation capabilities and are built on the Transformer architecture.

[0057] LLM chatbots, which are based on LLMs such as GPT, have a deep understanding of language, which enables them to provide more natural and contextually relevant responses. They can better understand and respond to complex queries.

[0058] An Al virtual assistant is a software application that uses Al, particularly natural language processing (NLP) and machine learning (ML), to understand and respond to user commands, either through text or voice. These assistants automate tasks, provide information, and personalize user experiences, often bridging the gap between humans and technology.

[0059] A chatbot and a virtual assistant, while both using conversational interfaces, may differ in their scope and functionality. Chatbots are generally designed for specific, often task-oriented interactions, while virtual assistants can handle a wider range of tasks, often adapting to user preferences and integrating with multiple applications.

[0060] Figure 1 of the drawings shows a system 100 for enhancing prompts provided to artificial intelligence models such as the LLM 190 illustrated in this example embodiment. The system 100 comprises a structured memory system 110 that comprises an ingestion engine 120 that receives input data 102, and is configured to (1) convert the received input data to converted data for processing, and to (2) determine and apply labels or tags (such as metadata tags), thereby creating tagged converted dat . Each tagged converted datum is associated with respective received input data. The input data 102 may relate to one or more data sources associated with the user and / or the use context, for example user chat interactions, emails, notes, or other unstructured text sources. In other words, the input data is context data that can be used to interpret user inputs, queries, user interactions, etc. In some embodiments, the system 100 further comprises a conversation interface 458 configured to receive a sequence of conversational messages from a user and to provide the conversational messages to one or more components of the system 100 for processing, storage, retrieval, and / or prompt construction.

[0061] The structured memory system 110 is external to an LLM and stores context-related data such as prior interaction data, as well as metadata. The structured memory system captures each LLM interaction and stores it with metadata such as a timestamp, topic, mode, emotional tone, a summary, and the like. The structured memory system 110 enables data labelling, data tagging, and the like (both manual and automatic), for later retrieval, and stores decision points and their context for strategic continuity.

[0062] In some embodiments, the system 100 may include a continuity engine 170 that maintains ongoing conversational threads and that links related memory units across multiple sessions. This enables retrieval of memory chains, allowing the LLM to simulate long-term reasoning or track decision-making progression. A graph structure may be optionally implemented to maintain semantic links between memory units based on co-occurrence, topic overlap, and / or narrative flow.

[0063] Referring to Figure 2 of the drawings, the structured memory system 110 comprises a memory store 130 with one or more databases, for example a metadatadatabase 132 that stores the tagged converted data, and a vector database 134 that indexes vector embeddings associated with the converted data.

[0064] In some embodiments, the memory store 130 further stores a summary memory entry 136 as an independent retrievable memory unit. The summary memory entry 136 may be a newly derived memory unit representing a compressed synthesis of a plurality of prior conversational exchanges or other related source content, rather than a mere transcript excerpt of the underlying messages. In some embodiments, the summary memory entry 136 is stored together with one or more metadata tags, for example identifying topics, entities, decisions, action items, emotional states, and / or temporal markers represented in the summarised content. The summary memory entry 136 may be indexed within the memory store 130, for example via the vector database 134 and associated metadata database 132, so as to be retrievable alongside other semantically related memory units derived from different source texts. In this manner, the summary memory entry 136 enables compressed conversational knowledge to be retained as persistent memory and later recalled without requiring reprocessing of the full underlying exchange history.

[0065] LLMs like GPT have a fixed memory limit defined by their context window. The structured memory system 110 described herein, however, uses an external vectorbased memory database that has the ability to significantly extend the functional memory of a chatbot. To the user, this creates the perception of almost unlimited memory by enabling the chatbot to recall relevant past information from a large, persistent memory store, even across different sessions or long time periods.

[0066] The system stores user interactions with content, emotion, context, and summaries. Using the vector database, the system facilitates the retrieval of relevant “memory slices” to add them to LLM prompts. This memory architecture reconstructs user intent and history in a context-aware way, providing context-aware results that are better than the results possible with simple search tools.

[0067] The process begins by converting each user interaction, such as a message, email, or file, into smaller, meaningful parts. These data portions are then transformed into vector representations using a text embedding model, and the data portions are enriched with metadata such as time, topic, emotion, decision points, etc. These data portions, or “memory units”, are stored in the vector database 134, which can hold a very large number of such entries. When the user submits a new query, the system compares the query’s vector with all stored vectors in the database 134 and retrieves only the most relevant memory units. These are selected not just by similarity, but also based on other factors such as how recent or important they are, based on stored metadata and user behaviour.

[0068] In some embodiments, the memory store is configured such that memory units are arranged according to semantic relationship among the memory units rather than according to the source document, message, file, session, or chronological sequence from which the memory units were derived. Thus, memory units extracted from different source texts may be stored proximate to one another when they relate to a common subject matter, entity, objective, event, or theme, while memory units extracted from the same source text may be stored in different regions of the memory space when they relate to different subject matters. This arrangement enables later retrieval to operate over semantically organised memory content instead of source-grouped text, so that relevant information can be assembled across multiple source texts without requiring those source texts to be re-opened or re-parsed as whole documents at query time.

[0069] As described in more detail elsewhere herein, once selected, these relevant memory units are compiled into a structured format and inserted into the prompt sent to the language model. This allows the model to respond as if it “remembers” past conversations, decisions, and topics, even though it has no built-in long-term memory. In practice, the chatbot uses only a small number of memory entries at a time (those that have been determined to matter most), but the chatbot has access to a much larger body of stored knowledge. As a result, the user experiences an interaction that feelscontinuous, context-aware, and deeply personalized, giving the impression of unlimited memory without exceeding the technical limits of the language model itself.

[0070] Referring to Figure 3 of the drawings, in some embodiments, the ingestion engine 120 comprises one or more data processing modules 121 that process the input data in order to facilitate interpretation and application of information in the input data by the system 100. In this example embodiment, the ingestion engine 120 includes a data conversion engine 122 that converts the received input data to converted data for processing, and a tagging engine 124 that processes the converted data to determine and apply metadata tags, thereby creating tagged converted data.

[0071] When the ingestion engine converts the received input data, this converting step may comprise segmenting the received input data into discrete memory units using natural language processing, so that each converted datum comprises a discrete memory unit.

[0072] In some embodiments, the ingestion engine 120 further comprises an embedding engine 126 configured to generate a respective vector representation for each discrete memory unit produced from the input data. The embedding engine 126 may apply a pretrained language embedding model, such as a transformer-based encoder, to convert each discrete memory unit into an embedding that captures semantic content of the unit in a machine-comparable form. The resulting vector representations may be stored in association with the corresponding discrete memory units and any metadata tags in the memory store 130, for example within a vector index or vector database, to facilitate semantic comparison between stored memory units and later user queries. In this manner, the embedding engine 126 enables retrieval of memory units based on semantic relatedness rather than simple keyword matching or source-document proximity.

[0073] The ingestion engine 120 may also comprise an NLP chunker and context extractor 128 that segments incoming data into granular, meaning-preserving units (memory units). The segmentation process may apply syntactic parsing, topicmodelling, entity recognition, and / or sentence boundary detection to produce logically coherent memory chunks.

[0074] The ingestion engine may comprise an embedding engine that processes the converted data to create vector embeddings associated with the converted data. For example, the data conversion engine 122 may comprise an embedding engine, which converts the unit into a vector representation using a pretrained language embedding model, such as a transformer-based encoder (e.g., BERT, SentenceTransformer, etc.).

[0075] The tagging engine 124 may comprise a metadata tagging engine that analyses the memory unit for contextual attributes, including but not limited to: a timestamp of interaction, detected entities (people, organizations, places), emotional tone, determined via classification models, topic category or tags (manually or automatically assigned), and / or decision points or action items, if present in the text.

[0076] In one embodiment, the system is designed to store, search, and recall relevant memories derived from prior user interactions, such as chat messages or other textual inputs. The architecture comprises several coordinated components that enable persistent memory handling and dynamic context enrichment during language model interaction.

[0077] The memory ingestion component processes raw text inputs by converting them into vector embeddings using a language embedding model, such as OpenAI’s 'text-embedding-ada-002'. These embeddings are stored along with corresponding metadata in a persistent storage file and indexed in a FAISS vector database to support efficient semantic retrieval. The metadata may include timestamps, tags, emotional tone, or other contextual signals useful for later filtering and ranking.

[0078] A semantic search and retrieval module performs vector-based similarity searches using FAISS with cosine similarity as the distance metric. When a user issues a query or continues a conversation, the system retrieves the top matching memory entries from the vector database. These retrieved memories are then injected into theprompt submitted to the large language model, enabling the generation of responses informed by prior relevant context.

[0079] The system also includes a web-based user interface that provides interaction endpoints for storing new inputs, conducting memory -based conversations, and browsing previously stored memory entries. For example, one route allows users to store or search entries, another supports active chat sessions, and a dedicated interface permits exploration and review of all saved memory units.

[0080] Additionally, the system features a memory summarization capability that operates during or after extended conversations. Once a predefined threshold of exchanges, such as four or more messages, has been reached, the system automatically summarizes the recent conversation and stores the summary as a new memory entry. This supports long-term continuity by transforming live dialogue into persistent, retrievable knowledge.

[0081] In some embodiments, the summarisation capability is not limited to storing a transcript excerpt, but instead creates a new derived memory unit or summary memory entry representing a plurality of prior exchanges in compressed form. The summary memory entry may be stored as an independent retrievable item in the memory store and may optionally be associated with metadata identifying one or more topics, entities, decisions, action items, emotional states, and / or temporal markers represented in the summarised conversation. By converting a run of live conversational exchanges into a newly stored summary memory entry upon satisfaction of a threshold condition, the system reduces later retrieval complexity and enables subsequent sessions to access condensed conversational knowledge without reprocessing the full underlying exchange history.

[0082] In summary, the system continuously reads, breaks down, embeds, and organizes long-form data (e.g., 50,000+ emails) into deeply structured vector memories, grouped semantically but stored for optimized retrieval, not by source proximity.

[0083] In an example embodiment, the system architecture includes an ingestion engine that receives input data (for example in the form of emails, notes, files, etc.) The ingestion engine comprises a Chunker and Context Extractor module and an Entity / Relationship Mapper that process the input data and then provide the processed data to a memory unit generator that tags the processed data (for example with semantic tags, metadata tags, etc.), and does vector embedding. The processed and tagged data is then stored in an organised vector database (e.g. a FAISS database). The data is stored based on the tags associated with the data units; for example the data may be stored based on semantic closeness, and not based on document grouping.

[0084] In some embodiments, the invention resides at least in part in the recognition that memory usefulness for later artificial-intelligence reasoning is improved when storage organisation reflects semantic affinity between memory units instead of preserving the layout of the source material from which the memory units were derived. Accordingly, the memory store may disregard original document boundaries for storage purposes and may co-locate memory units from different source texts when those memory units contribute to a common semantic context.

[0085] Compared to Google Gemini, the final product offers a series of significant advantages. Rather than relying on full document retrieval, it employs a strategy of chunking and vectorizing individual facts, which allows for efficient cross-email merging and more precise information access. Its search capability focuses on distilled memory concepts instead of the entire email blob, enabling users to find relevant details more quickly. Proximity logic is greatly enhanced by storing information near other semantically similar memories on a global scale, rather than simply grouping them based on documents or emails. For reasoning, the final product leverages only pre-embedded and tagged memory slices, which increases speed and efficiency, as opposed to Gemini’s need to parse entire emails at query time. Overall, these improvements result in instant retrieval due to pre-vectorized chunk memory and smarter, more relevant search results.

[0086] The arrangements described herein provide technical advantages over systems that preserve source-document adjacency or that later retrieve and process entire documents or transcripts. By storing memory units according to semantic relationship rather than source grouping, the system can retrieve semantically coherent information spanning multiple source texts while avoiding dependence on original document boundaries, thereby improving relevance and reducing unnecessary retrieval of unrelated text. Further, by generating a new retrievable summary memory entry when an interaction threshold is reached, the system converts extended dialogue into compact persistent knowledge that can be recalled in later sessions with reduced processing burden. These arrangements improve continuity, retrieval precision, and efficiency of context formation for artificial intelligence interactions, particularly where relevant knowledge is distributed across multiple source texts or across lengthy multi-turn conversations.

[0087] Email Example:In the following example, how the ingestion engine operates is described with reference to an email that is broken down and reorganized for storage, so that it can be used to provide context later on:Hello Joe,Glad it wasn't snowing at our warehouse so you can get it! It normally snows in March in Toronto.It was nice to meet you on Monday after all these years. I really enjoyed that burger with you.Let's see if we can get you supplying Costco! Please send through your presentation asap and I'll share it with Bob.From,David Jackman

[0088] The ingestion module chunks, tags, and groups the input data from the email as follows:Chunk (Memory Unit) Tags Group Joe hadn't met David Jackman in years ["David Jackman" A "relationship"]David Jackman enjoys burgers ["David Jackman" "food"] A Costco warehouse is in Toronto ["Costco" "location"] A It normally snows in March in Toronto ["Toronto" "weather"] B Joe wants to sell to Costco (from another ["Joe" "Costco" "goal"] A email)Joe hates snow (from another email) ["Joe" "snow" "weather"] B Joe loves snowboarding (from another email) ["Joe" "snow" "hobby"] BTable 1: Data Processing Example

[0089] The chunks are not stored adjacent to each other simply because they came from the same email. Instead, Group A memories are stored close to other memories about Costco, David Jackman, or sales strategy, even if those came from different emails. Group B memories, on the other hand, are stored near other snow-related memories, like Joe’s dislike of snow or his love of snowboarding. This ensures that semantic proximity, not temporal or source-based proximity, governs memory layout in vector space for this example.

[0090] In some embodiments, semantic organisation of the memory units is achieved by assigning each memory unit to a position, cluster, neighbourhood, index region, or other storage location based on one or more semantic indicators associated with that memory unit, such as an embedding, one or more semantic tags, one or more detected entities, one or more inferred relationships, and / or one or more topic indicators.Accordingly, the storage layout is determined by semantic affinity between memory units and not by preserving original source adjacency. This allows a later query relating to a particular topic to retrieve a set of memory units drawn from different emails,notes, chat sessions, or files that are semantically related to that topic, even where those memory units were originally separated in time, source, or document structure.

[0091] Referring again to Figure 1, the system 100 comprises a retrieval engine 140 configured to receive a user query 104, convert the received query to converted query data, query the vector index, such as vector database 134 based on the converted query data, and retrieve converted data from the metadata storage, such as metadata database 132 based on the vector database response.

[0092] The retrieval engine 140 selects relevant distinct memory units based on one or more factors, including semantic match, temporal relevance, and / or user behaviour. The retrieval engine 140 uses semantic search, like vector embeddings, to identify relevant prior interactions. In some embodiments, the retrieval engine ranks the relevance of distinct memory units based on their similarity to a current prompt, the recency of a prompt and / or memory unit, and / or user feedback. The retrieval engine 140 returns compact, context-specific snippets for prompt injection, ranking memory units by semantic relevance, recency, user behaviour, and the like.

[0093] In other words, the retrieval engine 140 is configured to rank retrieved converted data based on one or more of a similarity of the retrieved converted data to the converted query data, recency of the associated respective received input data, and user feedback.

[0094] The system 100 comprises a prompt construction module 150 that receives the retrieved tagged converted data from the metadata database via the retrieval engine 140, and creates an artificial intelligence (Al) prompt 152, e.g., an LLM prompt, based on both the received user query and the retrieved converted data. The prompt 152 is provided to an artificial intelligence model such as an LLM 190, which then uses the prompt 152 to create an LLM output 192.

[0095] The prompt construction module 150 applies a memory injection framework that dynamically compiles and formats memory context for inclusion in LLM prompts.The prompt construction module 150 assembles a custom context block before model completion, and combines the active prompt (i.e., the current user query) with memory data from the structured memory system, to create an enhanced LLM prompt. The prompt construction module 150 compiles selected memory entries into a structured prompt for LLM injection, balancing token limits with relevance scoring to optimize recall depth.

[0096] In some embodiments, the prompt construction module 150 comprises a token budgeting submodule 154 that estimates total token count and truncates or compresses lower-priority memories if a maximum prompt size is exceeded, thereby balancing token limits with relevance scoring to optimize depth of recall.

[0097] In some embodiments, the system includes a structured memory layer 110, a retrieval engine 140, a prompt construction layer 150, as well as one or more optional modules such as a memory review interface 160 and / or a learning loop 180.

[0098] The memory review user interface 160 is configured to facilitate user actions such as inputting user input in the form of annotating, modifying, deleting, and / or reorganizing stored converted data. The user actions and / or the resulting changes and / or feedback are stored in the memory store 130, and then retrieved by the retrieval engine 140 to be used by the retrieval engine. In some embodiments the retrieval engine applies the user input to further process the stored data, for example for ranking retrieved converted data for the purpose of prompt construction.

[0099] The memory review interface 160 allows users to explore, edit, and manage stored memory. It provides a timeline and filter-based dashboard for session history, supports annotating thoughts, editing summaries, pinning or discarding irrelevant memories, and tracking emotional history over time.

[0100] Optionally, the system includes a learning loop 180 that tracks which memory slices are reused or referenced by the user, adjusts future memory retrieval priorities accordingly, and learns which tone, framing, or memory patterns lead to best useroutcomes. The learning loop prioritizes memory content based on usage frequency, emotional tagging, and user feedback, adjusting memory prioritization based on historical usage and feedback.

[0101] The learning module 180 is configured to adjust a priority ranking of retrieved converted data based on historical usage. Additionally or alternatively, the learning module 180 may adjust the priority ranking based on the user input and / or user feedback, for example as received via the memory review user interface 160.

[0102] Referring to Figure 4 of the drawings, a computer-implemented method 400 comprises, at an ingestion engine, receiving 410 input data 402, and converting 420 the input data to converted data comprising discrete memory units. The method 400 includes applying 430 metadata tags to form tagged converted data associated with respective received input data, and storing 440, at a memory store 443, tagged converted data and vector embeddings associated with the converted data. In response to a user query 452, the method includes, at a retrieval engine, monitoring and determining 412, 450 a relevant data subset of the converted data based on one or more characteristics of the converted data, for example a semantic match between the user query and the converted data, temporal distance, user behaviour and / or feedback, and the like. Also, at the retrieval engine, the method includes retrieving 460 the relevant data subset from the memory store 443. The method then comprises summarising and creating 414, 470, by a prompt construction module, an artificial intelligence prompt based on both the user query 452 and the relevant data subset, and outputting 480 said prompt, for example to a memory or storage for later use, and / or submitting the prompt to an artificial intelligence model (for example to an LLM).

[0103] Optionally, the method may comprise receiving a user feedback input 456 and applying the user feedback input to determine and / or amend the relevant data subset.

[0104] In some embodiments, determining the relevant data subset may comprise determining a priority ranking of the converted data within the relevant data subset, for example based on one or more of: a similarity of the converted data to the user query,recency of the converted data, user feedback, and the like. The retrieval engine may, in some embodiments, be configured to determine, amend, and / or adjusting the priority ranking based on historical usage (e.g., as determined from the input data 402) and / or feedback (e.g., as provided via the user input 456).

[0105] Advantageously, the systems and methods described herein may be used for user memory management, user decision journals, therapy reflections, Al co-pilots with deep historical context, persistent learning environments, and character simulations.

[0106] A distinguishing feature of the systems described herein (for example, when compared to other solutions such as Gemini, ChatGPT, and Rewind) is the capacity to leverage a persistent, scalable external memory infrastructure. This capability enables the systems to reference any historical interaction, email, or content, regardless of the conversation’s length or complexity. Consequently, the novel systems and methods described are able to deliver responses that are more insightful and contextually continuous, providing users with a richer and more engaging experience.

[0107] The interface architecture described herein comprises a structured memory and is designed to manage prompt injection and formatting, learn from user engagement, and dynamically optimize prompts to ensure contextually accurate responses. This approach enhances both the flexibility and the effectiveness of the chatbot in delivering superior user outcomes. Advantageously, this approach does not rely on altering the underlying large language model (LLM), but instead provides a modular solution that can be applied to any number of artificial intelligence models.

[0108] In the claims which follow and in the preceding description, except where the context requires otherwise due to express language or necessary implication, the word “comprise” or variations such as “comprises” or “comprising” is used in an inclusive sense, i.e. to specify the presence of the stated features but not to preclude the presence or addition of further features in various embodiments.

Claims

CLAIMS:

1. A computer-implemented system (100) comprising:an ingestion engine (120) configured to receive a plurality of source texts and to segment the source texts into discrete memory units using natural language processing;an embedding engine (128) configured to generate, for each discrete memory unit, a respective vector representation;a tagging engine (124) configured to apply one or more metadata tags to each discrete memory unit;a memory store (130) comprising a vector index (134) and a metadata storage (132), the memory store (130) being configured to store the discrete memory units, the respective vector representations, and the metadata tags in a memory space organised according to semantic relationship among the discrete memory units rather than according to grouping by source text, such that semantically related discrete memory units derived from different source texts are stored closer to one another in the memory space than semantically unrelated discrete memory units derived from a same source text, independently of source-document adjacency;a retrieval engine (140) configured, in response to a user query (104), to retrieve from the memory store (130) a subset of the discrete memory units, the subset being assembled across a plurality of source texts for inclusion in a single artificial intelligence prompt; anda prompt construction module (150) configured to construct the artificial intelligence prompt using the user query (104) and the retrieved subset of discrete memory units.

2. The system (100) of claim 1, wherein the memory store (130) is configured to disregard original document boundaries when storing the discrete memory units, such that discrete memory units derived from a same source text are stored in different regions of the memory space when they are semantically unrelated.

3. The system (100) of claim 1 or claim 2, wherein the source texts comprise one or more of emails, notes, files, and chat messages.

4. The system (100) of any one of claims 1 to 3, wherein the natural language processing comprises one or more of syntactic parsing, topic modelling, entity recognition, and sentence boundary detection.

5. The system (100) of any one of claims 1 to 4, wherein the metadata tags comprise one or more of timestamps, detected entities, topic tags, emotional tone, decision points, action items, and temporal markers.

6. The system (100) of any one of claims 1 to 5, wherein the retrieval engine (140) is configured to rank candidate discrete memory units for retrieval based on one or more of semantic similarity to the user query (104), recency of the associated source text, metadata, historical usage, and user feedback.

7. The system (100) of any one of claims 1 to 6, wherein the prompt construction module (150) comprises a token budgeting submodule (154) configured to estimate a total token count for the artificial intelligence prompt and to truncate or compress lower-priority discrete memory units when a maximum prompt size is exceeded.

8. The system (100) of any one of claims 1 to 7, further comprising a memory review interface (160) configured to permit one or more of annotation, modification, deletion, and reorganisation of stored discrete memory units, wherein information representative of such user actions is stored in the memory store (130) and is usable by the retrieval engine (140) in retrieval or ranking of discrete memory units.

9. The system (100) of any one of claims 1 to 8, further comprising a learning module (180) configured to adjust a priority ranking of retrieved discrete memory units based on one or more of historical usage and user feedback.

10. The system (100) of any one of claims 1 to 9, wherein the memory store (130) is further configured to store a summary memory entry (136) as an independent retrievable memory unit in the memory space, the summary memory entry (136) being a newly derived memory unit representing a compressed synthesis of a plurality of prior conversational exchanges, stored with one or more metadata tags and retrievable alongside semantically related discrete memory units derived from other source texts.

11. A computer-implemented method (400) comprising:receiving (410), at a conversation interface (458), a sequence of messages in a conversation;monitoring (412) whether the sequence of messages has reached a predefined threshold number of exchanges;upon the threshold number of exchanges being reached, generating, by a summarisation module (414), a summary memory entry (136) comprising a newly derived memory unit representing a compressed synthesis of content of the sequence of messages, the summary memory entry (136) being distinct from a transcript of the underlying message sequence;storing (440) the summary memory entry (136) in a memory store (130; 443) as an independent retrievable memory item, the memory store (130; 443) being organised according to semantic relationship among stored memory units rather than according to source-text grouping, such that the summary memory entry (136) is retrievable alongside semantically related discrete memory units derived from other source texts; andin response to a later user query (452), retrieving (460) the summary memory entry (136) from the memory store (130; 443) based on semantic similarity between the user query (452) and the summary memory entry (136), without reprocessing the underlying message sequence.

12. The method (400) of claim 11, wherein the predefined threshold number of exchanges is four or more messages.

13. The method (400) of claim 11 or claim 12, further comprising storing (440) one or more metadata tags with the summary memory entry (136), the metadata tags comprising one or more of topic tags, detected entities, decision points, action items, emotional states, and temporal markers represented in the sequence of messages.

14. The method (400) of any one of claims 11 to 13, further comprising constructing (470) an artificial intelligence prompt using the user query (452) and the retrieved summary memory entry (136), and outputting (480) or submitting the artificial intelligence prompt to an artificial intelligence model (190).

15. The method (400) of any one of claims 11 to 14, further comprising:receiving (410) a plurality of source texts (102; 402);segmenting, at an ingestion engine (120), the source texts into discrete memory units using natural language processing;generating, by an embedding engine (128), for each discrete memory unit, a respective vector representation and one or more metadata tags; andstoring (440) the discrete memory units in the memory store (130; 443) such that semantically related discrete memory units derived from different source texts are stored closer to one another than semantically unrelated discrete memory units derived from a same source text, independently of source-document adjacency.

16. The method (400) of claim 15, wherein the memory store (130; 443) disregards original document boundaries when storing the discrete memory units.

17. The method (400) of claim 15 or claim 16, wherein retrieving (460) the summary memory entry (136) comprises assembling a retrieval subset comprising the summary memory entry (136) and discrete memory units derived from different source texts for inclusion in a single artificial intelligence prompt.

18. The method (400) of any one of claims 11 to 17, wherein candidate memory units are ranked for retrieval based on one or more of semantic similarity to the user query (452), recency, metadata, historical usage, and user feedback.

19. The method (400) of any one of claims 11 to 18, further comprising adjusting a priority ranking of retrieved memory units based on one or more of historical usage and user feedback received via a memory review interface (160).

20. The method (400) of any one of claims 11 to 19, further comprising estimating a token count for an artificial intelligence prompt constructed from retrieved memory units and truncating or compressing lower-priority memory units when a maximum prompt size is exceeded.