Database systems and methods of enhanced data management and contextual query responses

The system addresses the lack of contextually relevant responses in databases by transforming human readable data into machine readable vectors and generating hierarchical indexing, enhancing retrieval and response accuracy.

WO2026072825A1PCT designated stage Publication Date: 2026-04-02SYMBOTIC LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing database systems fail to provide contextually relevant responses to queries, often returning irrelevant or contextually disconnected information, and lack efficient hierarchical correlations between human readable data.

Method used

A system that transforms human readable data into machine readable vectors, generates hierarchical indexing through clustering and summarization, and utilizes generative AI to provide contextually relevant responses by identifying and utilizing organizational context.

Benefits of technology

Enhances data retrieval and response accuracy by preserving context and relevance, reducing computational processing, and improving response speed and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025047983_02042026_PF_FP_ABST
    Figure US2025047983_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Some embodiments provide systems to generate a contextual response to a user query comprising: a query system; an indexing system; a synthetic summary generation system; and a data store comprising human readable data; the indexing control circuit can transform the human readable data into data vectors. The summary control circuit can access the human readable data and generate first level synthetic summaries generated for data clusters of the human readable data generated based on a relative vector proximity of the data vectors. The indexing control circuit can transform the first level synthetic summaries into first level synthetic summary vectors. The summary control circuit can generate second level synthetic summaries generated for clusters of the first level synthetic summaries selected based on a relative vector proximity.
Need to check novelty before this filing date? Find Prior Art

Description

Aty. Docket No. 1127P017359-WO (PCT)DATABASE SYSTEMS AND METHODS OF ENHANCED DATA MANAGEMENT AND CONTEXTUAL QUERY RESPONSESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is a non-provisional of and claims the benefit of United States provisional patent application number 63 / 698,984, filed on September 25, 2024, the disclosure of which is incorporated herein by reference in its entirety.BACKGROUND1. Field

[0002] This invention relates generally to improved databases.2. Brief Description of Related Developments

[0003] Many systems often enable access to the content through query searching. Some of these systems can be used to ask human language questions and get artificial intelligence (Al) generated answer. To assist the Al engine, context related information can be provided. There is a need to improve responses to queries.BRIEF DESCRIPTION OF DRAWINGS

[0004] Disclosed herein are embodiments of systems, apparatuses and methods pertaining to databases and content access. This description includes drawings, wherein:

[0005] FIG. 1 illustrates a simplified block of an exemplary document management and content retrieval system, in accordance with some embodiments.

[0006] FIG. 2 illustrates a simplified block diagram representation of an exemplary data store, in accordance with some embodiments.Aty. Docket No. 1127P017359-WO (PCT)

[0007] FIG. 3 illustrates a simplified flow diagram of an exemplary indexing process that generates a hierarchical indexing relative to some or all of the human readable data, in accordance with some embodiments.

[0008] FIG. 4 illustrates a simplified graphical representation of an exemplary hierarchical indexing and correlation of data corresponding to at least a subset of the human readable data, in accordance with some embodiments.

[0009] FIG. 5 illustrates a simplified block diagram of an exemplary representation of a response process implemented to response to one or more received queries, in accordance with some embodiments.

[0010] FIG. 6 illustrates a simplified block diagram representation of an exemplary process of providing a contextually relevant response to a query, in accordance with some embodiments.

[0011] FIG. 7 illustrates a simplified flow diagram of an exemplary process of providing a contextual response to a query, in accordance with some embodiments.

[0012] FIG. 8 illustrates a simplified block diagram of an exemplary process of generating a hierarchical indexing for human readable data, in accordance with some embodiments.

[0013] FIG. 9 illustrates a simplified flow diagram of an exemplary process of generating a contextual response to a user query, in accordance with some embodiments.

[0014] FIG. 10 illustrates a simplified block diagram of an exemplary process of creating a hierarchical indexing of human readable data, in accordance with some embodiments.

[0015] FIG. 11 illustrates an exemplary system for use in implementing methods, techniques, devices, apparatuses, systems, servers, sources and providing document management and query responses, in accordance with some embodiments.

[0016] Elements in the figures are illustrated for simplicity and clarity and have not necessarily been drawn to scale. For example, the dimensions and / or relative positioning of some of the elements in the figures may be exaggerated relative to other elements to help to improve understanding of various embodiments. Also, common but well-understood elements that areAty. Docket No. 1127P017359-WO (PCT) useful or necessary in a commercially feasible embodiment are often not depicted in order to facilitate a less obstructed view of these various embodiments. Certain actions and / or steps may be described or depicted in a particular order of occurrence while those skilled in the art will understand that such specificity with respect to sequence is not actually required. The terms and expressions used herein have the ordinary technical meaning as is accorded to such terms and expressions by persons skilled in the technical field as set forth above except where different specific meanings have otherwise been set forth herein.DETAILED DESCRIPTION

[0017] The following description is not to be taken in a limiting sense, but is made merely for the purpose of describing the general principles of exemplary embodiments. Reference throughout this specification to “one embodiment,” “an embodiment,” “some embodiments”, “an implementation”, “some implementations”, “some applications”, or similar language means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. Thus, appearances of the phrases “in one embodiment,” “in an embodiment,” “in some embodiments”, “in some implementations”, and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

[0018] Generally speaking, pursuant to various embodiments, systems, apparatuses and methods are provided herein useful to provide improved databases and data management that enable enhanced and more contextually relevant responses to queries. Some embodiments provide specific databases of organized contextual information through the transformation and reduction of content that are used to define correlations between such content, and methods of automatically selecting more relevant context to supplement queries, which may be utilized with generative Al engines, in providing more contextually relevant responses to such queries. The systems provide improvements to computer implemented artificial intelligence systems through specific rules that enable an automation of contextual identification of information that can provide a more practical and relevant response to a query that could not previously be provided.

[0019] In some embodiments, databases and data management systems are provided that generate a contextual response to a user query. The system can include: a query system comprising a queryAty. Docket No. 1127P017359-WO (PCT) control circuit; an indexing system comprising an indexing control circuit configured to execute indexing code; a synthetic summary generation system comprising a summary control circuit configured to execute summaries generation code; and a data store communicatively coupled to the query control circuit and the indexing control circuit. The data store comprising human readable data and / or other relevant information. The indexing control circuit, in executing the indexing code, can be configured to transform each of the human readable data into respective data vectors assigned to the respective human readable data providing machine readable representations of the human readable data and store the data vectors in the data store. Each of the human readable data is assigned a different data vector of the data vectors. The summary control circuit, in executing the summary code, is configured to access the human readable data and generate first level synthetic summaries generated for data clusters of the human readable data. In some embodiments, each data cluster, of the data clusters, comprises one of the human readable data or a plurality of related human readable data of the human readable data. Further, each data cluster can be generated based on a relative vector proximity of the data vectors. The indexing control circuit can be further configured to transform each of the first level synthetic summaries into respective first level synthetic summary vectors assigned to the respective first level synthetic summaries providing machine readable representations of the first level synthetic summaries. Each of the first level synthetic summaries is assigned a different first level synthetic summary vector of the first level synthetic summary vectors. The summary control circuit is further configured to generate second level synthetic summaries generated for clusters of the first level synthetic summaries. The clusters for the first level synthetic summaries can be selected based on a relative vector proximity of the first level synthetic summary vectors. The second level synthetic summaries can be stored in the data store. The indexing control circuit can further be configured to transform each of the second level synthetic summaries into respective second level synthetic summary vectors assigned to the respective second level synthetic summaries providing machine readable representations of the second level synthetic summaries.

[0020] Some embodiments comprise methods for providing a contextual response to a user query, comprising: receiving a query from a user interface unit associated with a query source; accessing relevant human readable data related to the query; accessing, for each of the accessed relevant human readable data, a respective context of the relevant human readable data from a data store, the respective context comprising hierarchically dependent synthetic summaries related to theAty. Docket No. 1127P017359-WO (PCT) relevant human readable data; transferring the relevant human readable data and the context of the relevant human readable data to a generative artificial intelligence model; generating, based on the application of the generative artificial intelligence model, a contextual response to the query based on the retrieved relevant human readable data and the context of the relevant human readable data; and communicating the contextual response, from the generative artificial intelligence model, to the user interface unit associated with the query in controlling the user interface unit to present the contextual response.

[0021] It is common for a user and / or system to try and identify documents and / or content that is relevant to a particular subject, topic, issue, text, other such information, or a combination of two or more of such information. Previous systems enable a user to enter search criteria (e.g., one or more words, phrases, numbers, content, and / or other information) that is relevant to the information attempting to be identified and / or accessed. Previous search systems, however, often failed to identify the most relevant information, and typically lost the context of the documents. Further, previous systems often returned a listing of multiple different documents and / or information without context, and often many of the returned items are not on point to the desired information.

[0022] Some present embodiments provide an enhanced data management system and / or database system providing improved hierarchical content correlation and content retrieval systems and / or methods. These systems and methods can preserve the context and / or relevance of information (e.g., documents, text, images, spreadsheets, invoices, lists, and / or other such information), and in some embodiments enhances search results through a developed hierarchical correlations between the information. For example, the content retrieval systems and methods in some embodiments deliver contextual answers to inquiries using a company's data corpus to provide more accurate results that are more contextually relevant. Further, in some embodiments, the systems and methods streamline and optimize processes of creating, managing, and evolving queries of realtime data within an organization.

[0023] The system, in some embodiments, identifies and utilizes an organization's distinct terms and concepts, ensuring relevant and accurate data retrieval and answers. This feature differentiates it from generic document management solutions by providing customized responses based on theAty. Docket No. 1127P017359-WO (PCT) organization's unique context. Previous systems typically were unable to accurately provide context or the context was typically limited to context identified in the query. Alternatively, the present embodiments improve computing and data management systems, in part, through the unconventional contextual correlation between different human readable information in relation to a particular context (e.g., organization, subject matter, geographic location, etc ), which can advantageously utilize an identified unique vocabulary and / or glossary, and the synthetic summarization of contextually related clusters.

[0024] FIG. 1 illustrates a simplified block of an exemplary document management and content retrieval system 100, in accordance with some embodiments. The document management and content retrieval system 100, in some embodiments, includes one or more query systems 110, one or more indexing systems 106, one or more glossary generator systems 109, one or more content retrieval systems 116, and one or more response generator systems 114 communicatively coupled over one or more computer and / or communication networks 108. The document management and content retrieval system 100 further includes one or more data stores 104. The document management and content retrieval system 100 is further in communication with one or more user interface units 118, such as but not limited to user smartphones, tablets, laptops, computers, and / or other such systems that enable a user to submit queries and / or incorporate data to the system 100. One or more of the user interface units 118 may be part of the document management system, and / or one or more of the user interface units may be separate and distinct from the document management system. The document management system can further include and / or be in communication via the one or more networks 108 with one or more databases 112.

[0025] The document management and content retrieval system 100 greatly improves data retrieval, database technologies and data management, identification of and access to more relevant data, including human readable data, through the transformation of human readable data into unconventional contextual hierarchical correlations of different content of the data, and provides more contextually relevant responses to queries from users and / or other systems through the unconventionally managed hierarchical correlations. The improved databases, data management and contextual retrieval requires less memory than previous systems by in part establishing the correlations, and results in a faster computation times while improving data retrieval. The data management and retrieval is produced, in some embodiments, through one orAty. Docket No. 1127P017359-WO (PCT) more different machine learning operations and / or an iterative mathematical operation that begin with transformation of human readable data at least in part into respective data vectors and / or other embedded data that enable a different functional use to identify clusters of content that are contextually related, that are themselves transformed into synthetic summaries that can similarly be transformed into summary vectors used to further define additional levels of contextually related clusters. The process can repeat the embedding, clustering and synthetic summarizations in establishing a functional hierarchical indexing of the relevant data that provide a different practical function in identifying contextual correlation. Still further, the system can utilize the hierarchical indexing to identify relevant data and provide more accurate and relevant responses to queries through the enhanced databases and improved contextual indexing. The contextual processing of the content and queries results in faster computation time in correlating data, as well as faster contextual content retrieval while further providing improved quality of the resulting responses to queries as occurred in prior systems and / or process. The contextual correlation of data can be utilized, for example, by an artificial intelligence system in the generation of a more contextually relevant answers to human language questions. The contextual relevance, in part, can be enhanced through the recognition of particular vocabulary and / or a contextual relevance of terms used that are specific to a particular organization and / or subject. The present embodiments provide an innovation in database and computer technology, namely digital data processing and query responses, proving both an improvement in the functioning of computer and computer memory, and an improvement in data retrieval in providing enhanced and more relevant content that is contextually relevant to the particular query. The generation of the hierarchical contextual correlation through repeated embedding and clustering of such embedded content is not well known. The generation of hierarchical indexing of content through the sequential clustering and generation of synthetic summaries that are themselves clusters and used to generate subsequent hierarchical levels of correlation addresses specific problems in computer technology in contextually correlating content in order to provide more contextually relevant real-time responses to computer submitted queries with a reduced amount of computation processing and in a reduced amount of time.

[0026] The response generator system 114 can generate contextual responses to different user queries from multiple different local and / or remote users and / or sources through the one or more user interface units 118. The one or more data stores 104 can be implemented through one or moreAty. Docket No. 1127P017359-WO (PCT) local and / or remote computer memory and / or servers, such as RAM, ROM, flash memory, web servers, application servers, database servers, and / or any other forms of memory and / or servers capable of storing and providing access to the relevant content, indexes and / or content correlations.

[0027] In some embodiments, the indexing system 106 comprises one or more indexing control circuits or systems 126, and / or one or more summary control circuits and / or generation systems 128, each of which can be implemented through one or more processors, microprocessors, ASICs, servers, computers, other such processing systems or a combination of two or more of such processing systems. The indexing system 126 indexes human readable data into a hierarchical correlation of human readable data and synthetic summaries generated by the summary system 128.

[0028] FIG. 2 illustrates a simplified block diagram representation of an exemplary data store 104, in accordance with some embodiments. The one or more data store 104 store and / or access one or more human readable data. Such human readable data can include reports, spreadsheets, scheduling, invoices, shipping data, inventory data, product information, sales data, white papers, news reports, news papers, images, video, articles, documents, correspondences, memos, publications, other such human readable data, and typically a combination of two or more of such human readable data.

[0029] FIG. 3 illustrates a simplified flow diagram of an exemplary indexing process 300 that generates a hierarchical indexing relative to some or all of the human readable data 202, in accordance with some embodiments. The human readable data 202 may be stored in the data store 104, one or more of the databases 112 and / or other data storage systems, which may be local, remote and / or maintained by one or more third parties. Often the human readable data 202 is particular to a specific business entity or entities, technology, geographic area, language, and / or other such correlation. As such, the information has more relevance to that particular correlation (e.g., business entity and / or entities). The data can be used to identify vocabulary relevant to the correlation, definitions more relevant to the correlation, and other factors that are more relevant to the correlation and may provide more contextual significance to the respective correlation. Some embodiments include one or more glossary generator systems 109 that process the human readable data relative to a particular organization, a group of organizations, an area and / or portion of anAty. Docket No. 1127P017359-WO (PCT) organization or business, a type of organization, other such collections of data, or a combination of two or more of such collections of data. The glossary generator system can in some implementations extract glossaries and / or recognize relevant variables of an organization or collection, identify key terms within the human readable data, concepts within the human readable data, key performance indicators (KPIs) corresponding the human readable data, and can assemble one or more comprehensive glossaries for convenient reference (e.g., constructing a glossary for a retail warehouse management organization and / or system).

[0030] The document management and content retrieval system 100 utilizes the human readable data 202 and builds a hierarchical correlation between the data, as well as providing more relevant contextual information in response to queries that are relevant and / or related to the data. FIG. 4 illustrates a simplified graphical representation of an exemplary hierarchical indexing 400 and correlation of data corresponding to at least a subset of the human readable data 202, in accordance with some embodiments. Referring to FIGS. 1-4, the indexing system 106 can comprise and / or be implemented through one or more indexing control circuits, executing the indexing code, access the human readable data 202 and transform each of the human readable data into corresponding reduced, machine readable representations 204 (e.g., embeddings) of the human readable data that can be processed by models, applications and / or processing systems to understand semantic meanings. For example, some embodiments apply one or more rules to apply one or more algorithms and / or machine learning to generate respective one or more data vectors (vl , v2 ... vN) for each of the human readable data, which are assigned to the respective human readable data, providing machine readable representations 204 of the human readable data 202. In some embodiments, the vectors are a reduced state transformation representative of the context of the respective human readable data. The data vectors and / or other machine readable representations can be stored in the data store 104. Accordingly, in some embodiments, each of the human readable data is assigned at least one different respective data vector 204 or other such embedding. In some embodiments, the data vectors can be generated through one or more computer implemented applications, known algorithms, proprietary algorithms, trained machine learning models (which are retrained over time based on feedback) and / or other such methods, or a combination of two or more of such methods. For example, one or more embedding models can be applied to each of the human readable data, such as but not limited to contrastive learning model, transformer architecture, OpenAI’s Ada encoder, Alibaba’s GTE model, Google’s Universal SentenceAty. Docket No. 1127P017359-WO (PCT)Encoder, and / or other such models or combination of two or more of such models. Feedback based subsequent correlations can be used as part of the corpus data used to retrain and modify the embedding machine learning models over time to repeatedly retrain the models resulting and a transformation to different models that provide more accurate responses.

[0031] Still referring to FIGS. 1-4 and in generating and / or maintaining the index, for the different human readable documents 202 and respective embeddings (e.g., data vectors) 330, a particular document is identified and / or selected 308. For the selected document, the system evaluates one or more other human readable documents in attempts to identify nearest neighbors of the particular document. In some embodiments, further unique cluster rules are applied where one or more human readable data clusters 206 can be identified 310 and / or created based on a proximity and / or other correlation of the data vectors (v) 204. This clustering can include and / or identify one or more documents and / or other human readable data 202 that are contextually relevant to each other based at least in part on the embedding. For example, human readable data 202 directed towards the same or similar topics may be identified through one or more threshold relationships of the respective data vectors, and used to create one or more data clusters directed at least towards the shared topic. In some embodiments, one or more of the data clusters 206 can comprises one of the human readable data or a plurality of related human readable data. A level or degree of contextual correlation and / or lack of contextual correlation between human readable data can be determined, in some implementations, based on a relative proximity of the corresponding data vectors 204. In some embodiments, for each identified cluster 206 of human readable data, summary rules can be applied where one or more first level synthetic summaries (S) 208 can be generated and summary vector rules can be applied where one or more subsequent first level synthetic summary vectors (vS) 210 can be determined based on each of the first level of synthetic summaries 208. Some embodiments further determine one or more first level synthetic summary clusters 212 from the first level of synthetic summaries 208 based at least in part on the correlations and / or proximities of the first level synthetic summary vectors (vS) 210.

[0032] The process can be repeated substantially any number of times to obtain substantially any number of levels of synthetic summaries, corresponding respective level synthetic summary vectors, and respective level synthetic summary clusters. For example, the data store can store a hierarchical data index (e g., data index 400) that can comprise one or more third level syntheticAty. Docket No. 1127P017359-WO (PCT) summaries 220 generated for clusters 218 of the second level synthetic summaries 214. In some implementations, the second level synthetic summary clusters 218 can be selected based at least in part on a proximity of the second level synthetic summary vectors. The data store, in some embodiments, can further comprises respective third level synthetic summary vectors (vSSS) 222 assigned to the one or more third level synthetic summaries 220 providing respective machine readable representations of the respective third level synthetic summaries. Accordingly, the index can include and the data store can store additional higher levels of synthetic summaries, which may respectively be generated in some embodiments by applying one or more machine learning models (e.g., large language models).

[0033] In the example hierarchical indexing 400 illustrated in FIG. 4, as one nondimiting example, a first data cluster 206.1 of a first level of data clusters can be identified for a first data dl or document as just the first data dl when a threshold vector relationship of the corresponding first data vector vl is not identified with one or more of the other relevant data vectors; a second data cluster 206.2 can be identified based on a threshold correlation between a second data vector v2 corresponding to a second data d2 and a fourth data vector v4 corresponding to a fourth document d4; and a xth data cluster 206.x (e.g., third) can be identified based on a threshold correlation between a third data vector v3 corresponding to a third data d3 and an nth data vector vN corresponding to an nth document dN. It is noted that while the example in FIG. 4 shows data clusters of a single data (e.g., dl) and two data (e.g., d2 and d4), the data clusters 206 are not limited to clusters of one or two data / documents, and instead can identify clusters of one, two or more data (e.g., data cluster of documents d5, d6, d7, d8, d9, d32 and d50, not illustrated in FIG. 4). Similarly, in some embodiments, a human readable data 202 may be associated with more than a single cluster. Each data cluster may comprise and / or associate at least one of the human readable data 202 or a plurality of related human readable data. Typically, one or more vector proximity thresholds are applied in the identifying of the cluster. The one or more vector proximity thresholds may vary for different human readable data (e.g., a larger concentration of documents may result in one or more narrower thresholds being applied while one or more broader thresholds may be applied less concentrations of human readable data, different thresholds may be applied based on different queries and / or types of queries, thresholds may be varied over time based on feedback and / or machine learning from previous historic clusters and / or feedback from previous responses to queries, and / or other such variations).Aty. Docket No. 1127P017359-WO (PCT)

[0034] In some embodiments, the summary system 128 can generate 312 one or more first level of synthetic summaries 208 from each of the clusters of the human readable data. In some embodiments, the summary system 128 can apply one or more machine learning models, algorithms, applications and / or other methods to automatically generate the summaries 208. For example, some embodiments may apply one or more pretrained large language models that can be trained based on a corpus of human readable data and / or embedded human readable data, and repeatedly retrained over time based on historic summaries, clustering, data vectors, query responses and feedback of the query responses, other such feedback or a combination of two or more of such feedback over time. The first level synthetic summaries 208, in some embodiments, is generated based on the corresponding data vectors (v) 204 of the cluster. As such, the synthetic summaries are typically not a summary in a human readable language. Some embodiments, however, may generate a human readable summary based on the generated synthetic summary.

[0035] The indexing system 126 can, in some embodiments, continue generating the hierarchical indexing by further transforming each of the first level synthetic summaries 208 and generating 314 one or more first level synthetic summary vectors (vS) 210 and / or other embeddings corresponding to each of the generated first level synthetic summaries 208. The first level synthetic summary vectors 210 can be assigned to the respective first level synthetic summaries 208 as machine readable representations of the first level synthetic summaries and associated as a first level collection 331. Steps 308, 310, 312 and 314 can be repeated any number of times for each of one or more human readable data 202. In some embodiments, upon performing steps 308, 310 and 312 for a particular data (e.g., dl), that data is optionally removed 316 from subsequent consideration in the selection of a subsequent data (e.g., d2) that is clustered in step 310, generated clusters can be summarized in step 312 and vectorized in step 314, and then optionally removed in step 316 before continuing to a subsequent human readable data (e.g., d3). Some embodiments similarly may remove from subsequent consideration for clustering, those nearest neighbor documents identified in the cluster.

[0036] Still referring to FIGS. 1-4, the hierarchical indexing can continue, in some embodiments, through repeated sublevel clustering, the generation of corresponding sublevels of synthetic summaries and corresponding embeddings (e g., representative vectors). There can be substantially any number of levels, and in some embodiments, the indexing is continued for aAty. Docket No. 1127P017359-WO (PCT) number of sublevels until a single synthetic summary and corresponding embedding is established. For example, each first level embedded synthetic summary can be evaluated to identify (e.g., step 310) and / or create one or more first level summary clusters 212, based on a proximity and / or other correlation of the respective summary vectors (vS) 210. This clustering, in some embodiments, can identify one or more synthetic summaries that are contextually relevant to each other based at least in part on one or more embedding thresholds. As one non-limiting example, as illustrated in exemplary FIG. 4, a first, first level summary cluster 212.1 based on a threshold correlation between a first, first level summary vector vSl corresponding to a first, first level synthetic summary SI and a second, first level summary vector vS2 corresponding to a second, first level synthetic summary S2; and a yth, first level summary cluster 212.y (e.g., second, first level summary cluster) can be identified for just a third, first level synthetic summary S3 when a threshold vector relationship of the corresponding third, first level summary vector vS3 is not identified with one or more of the other relevant first level summary vectors. It is noted that while the example in FIG. 4 shows first level summary clusters 212 of a single synthetic summary S3 and two synthetic summaries SI and S2, the first level summary clusters 212 are not limited to clusters of one or two summaries, and instead can identify clusters of one, two or more synthetic summaries. Typically, one or more vector proximity thresholds are applied in the identifying of the summary clusters. The one or more vector proximity thresholds may vary based on one or more factors, such as but not limited to the query, the corresponding human readable data, number summaries at a corresponding level, other such factors, or a combination of two or more of such factors.

[0037] Some embodiments continue the indexing where the summary system 128 can generate 312 second level synthetic summaries 214 for the identified first level summary clusters 212 of the first level synthetic summaries 208. The first level summary clusters 212 for the first level synthetic summaries, in some embodiments, can be created based on a proximity of the first level synthetic summary vectors (vS) 210. The second level synthetic summaries (SS) 214 can further be processed in step 314 to generate embedding (e.g., vectors) of the respective second level synthetic summaries (SS) 214, and the generated second level synthetic summary vectors (vSS) 216 assigned to the respective second level synthetic summaries 214 as machine readable representations of the respective second level synthetic summaries (SS) 214, in a second level collection 332. Again, when relevant, one or more second level synthetic summary clusters 218Aty. Docket No. 1127P017359-WO (PCT) can be identified for the second level synthetic summaries (SS) 214. The hierarchical indexing can have substantially any number n of levels of human readable data and sublevel synthetic summaries and synthetic summary vectors, n being any positive integer number (e.g., with an nth level synthetic summary (S... S) 220). Some embodiments further generate an nth level synthetic summary vector (vS... S) 222 corresponding to the nth level synthetic summary. For example, referring to FIG. 4, a second level summary cluster 218 can be defined for the first, second level synthetic summary SSI and the second, second level synthetic summary SS2 based on the relevance proximity of the first, second level summary vector vSSl and the second, second level summary vector vSS2, and an nth level synthetic summary (SSS) generated based at least in part on the first, second level synthetic summary SSI and the second, second level synthetic summary SS2, and in some embodiments, one or more of the earlier level summaries and / or corresponding human readable data. Some embodiments maintain the embedded human readable data 304, synthetic summaries, the embedded synthetic summaries with the associated synthetic summary vectors in the hierarchical index 320. Further, some embodiments maintain, in the index 320 and / or a separate mapping database, a mapping of or linking between the human readable data 202, the identified clusters and the corresponding parent synthetic summaries. As such, the present embodiments enhanced computer memory systems that provide improvements in computer capabilities reducing computational processing to access and identify relevant information while reducing memory usage based in part on the hierarchal embedding and summarization, while greatly improving the computer-centric issue resulting in previous systems responses, including Al responses, to queries that typically lacked accurate contextual relevance.

[0038] The one or more data stores 104 can store the human readable data 202 that can provide information to users. Such human readable data can include reports, spreadsheets, scheduling, invoices, shipping data, inventory data, product information, sales data, white papers, news reports, news papers, images, video, articles, documents, correspondences, memos, publications, other such human readable data, and typically a combination of two or more of such human readable data. Human readable data may be generated internal by one or more internal users (e.g., reports, orders, invoices, spreadsheets, etc.), received from one or more remote users, imported into the one or more data stores 104 from one or more databases 112 and / or third party sources (e.g., servers, databases, etc.) accessible across the one or more communication networks 108. The one or more databases 112 may be local databases and / or remote databases, and some of the humanAty. Docket No. 1127P017359-WO (PCT) readable data may come from substantially any combination of publicly accessible and / or private databases. For example, a private database may be a database which only specific users have access to, such as a company specific database wherein only employees may access the human readable data stored in the company database. A public database may include substantially any database which contains human readable data accessible to others. The human readable data stored in the one or more data stores 104 may include public data accessible by, for example, public search engines. The data store 104 further stores nthlevels of vectors, nthlevels of synthetic summaries, and nthlevel data clusters generated by the document management and content retrieval system 100.

[0039] The query system 110 is configured to receive queries from the multiple different remote users and sources via the one or more user interface units 118. In some embodiments, the query system comprises one or more query receiver systems 122 and one or more query vectorization systems 124. The query receiver system 110 can comprise one or more receiver control circuit communicatively coupled with the one or more transceivers and one or more receiver computer memory storing at least code executable by the receiver control circuit. The receiver control circuit, in some embodiments, is implemented through one or more processors, microprocessors, ASICs, servers, computers, other such processing systems or a combination of two or more of such processing systems. The query receiver system 122 processes the received queries to extract relevant query information from the query and communicate the relevant query information at least to the query vectorization system 124. Upon receiving a query via the one or more user interface units 118, the query system 110 in some embodiments embeds the query with assigned machine readable vectors. After the query has been embedded with machine readable vectors, the query is sent to the retrieval system 116. After receiving the query from the query system 110, the retrieval system 116 interacts with the one or more data stores 104 via the communication network 108 to obtain data relevant to the query which has been indexed by the retrieval system 116 and stored in the one or more data stores 104.

[0040] FIG. 5 illustrates a simplified block diagram of an exemplary representation of a response process 500 implemented to response to one or more received queries 502, in accordance with some embodiments. A query 502 can be received 504 by the query system 110 that can processes the query to extract relevant information from the query. In some embodiments, the query systemAty. Docket No. 1127P017359-WO (PCT)110 generates one or more machine readable representations of the query, such as a generation of one or more query vectors through the query vectorization system 124. The query vectorization system 124, in some embodiments, comprises one or more query vectorization control circuits or systems each of which can be implemented through one or more processors, microprocessors, ASICs, servers, computers, other such processing systems or a combination of two or more of such processing systems that execute vectorization code.

[0041] The retrieval system 116, in some embodiments, comprises one or more retrieval control circuits or systems each of which can be implemented through one or more processors, microprocessors, ASICs, servers, computers, other such processing systems or a combination of two or more of such processing systems that execute retrieval code. The retrieval system can interact 508 with the one or more data stores 104 to access relevant human readable data related to the query and the context of the human readable data (e.g., the synthetic summaries). For example, a set of one or more relevant human readable data related to the query and the context of the set of the human readable data can be accessed and / or retrieved 510 relevant to the query. The context, in some embodiments, can comprise one or more (e g., a subset) of the hierarchical synthetic summaries corresponding to the set of relevant human readable data. In some embodiments, the relevant data is determined based on the proximity of the query vector to one or more of the different levels of synthetic summary vectors of the different levels of synthetic summaries. The proximity can be determined based on one or more factors and / or thresholds. When the retrieval system 116 determines that a query vector has a predefined relationship with one or more proximity thresholds to one or more synthetic summary vectors of one or more levels, which indicates a threshold relationship in terms of content, the retrieval system 116 can retrieve at least the corresponding synthetic summary, and typically some or all of the relevant parent or higher level synthetic summaries (e.g., synthetic summaries 208, 214, 220) and / or the corresponding human readable data 202 for the relevant synthetic summary. As one non-limiting example, when an exemplary query vector, transformed and generated from a query, is determined to have a threshold relationship with to a third level summary vector (vSSS) of a third level synthetic summary (SSS), the retrieval system 116 can, in some embodiments, retrieve the third level synthetic summary (SSS), one or more second level synthetic summaries (SS) 214 of one or more second level clusters 218 of the one or more second level synthetic summaries 214 from which the third level synthetic summary (SSS) is based, one or more first level synthetic summariesAty. Docket No. 1127P017359-WO (PCT)(S) 208 of one or more first level summary clusters 212 of the one or more first level synthetic summaries (S) 208 from which the one or more second level clusters 218 are based, and one or more human readable data 202 of one or more clusters 206 of human readable data from which the one or more first level synthetic summaries (S) 208 are generated. The retrieval system 116, in some embodiments, may also retrieve the corresponding second level synthetic summary vectors (vSS) 216, the first level synthetic summary vectors (vS) 210, and / or the human readable data vectors (v) 204that are determined to have relevance to the received query 502.

[0042] The retrieval system 116, in some embodiments, can transfer the identified relevant data 512 (e.g., a set of one or more human readable data 202 and context of the set of human readable data such as one or more corresponding first level and / or second level synthetic summaries (208, 214, etc.) of respective summary clusters, and in some instances the respective query) to the response generator system 114. For example, in relation to FIG. 4, a first human readable data dl and a fourth human readable data d4 may be identified as relevant to the query, and the retrieval system 116 can retrieve the first human readable data dl and the corresponding parent synthetic summaries SI, SSI, SSS; and the fourth human readable data d4 and the corresponding parent synthetic summaries S2, SSI, SSS.

[0043] The response generator system 114, in some embodiments, comprises one or more response generator control circuits or systems each of which can be implemented through one or more processors, microprocessors, ASICs, servers, computers, other such processing systems or a combination of two or more of such processing systems. The response generator system 114 can process the identified relevance data 512 to generate and send one or more responses 520 to the requesting user interface unit 118 of the user from which the query 502 is received. In some embodiments, the response is user readable data. The response generator system 114, in some embodiments, can apply one or more response algorithms, response language machine learning models 115 and / or other such processing to generate the contextual response 520. As one nonlimiting example, the response machine learning model(s) 115 can apply one or more pretrained large language models (LLM) that can be trained based on a corpus of queries, human readable data, human readable data vectors, synthetic summaries, summary vectors, and / or other relevant data, and repeatedly retrained overtime based on historic synthetic summaries, historic clustering, historic data vectors, historic query responses 520 and historic feedback of historic query responsesAty. Docket No. 1127P017359-WO (PCT)(e g., follow up queries, selection of documents and / or accessing one or more data, and / or other such feedback), other such historic data or a combination of two or more of such historic data over time. The model based response generator system 114 can employ, for example, few-shot learning based on the human readable data and their summaries to produce accurate responses. In some embodiments, the response generator system 114 generates a human readable response and / or answer to the query 502 based on the relevant data retrieved by the retrieval system 116. For example, the document management and content retrieval system 100 can apply one or more LLM- based models to process a query (e.g., generate one or more query vectors), retrieve pertinent data from the human readable data and / or multi-level synthetic summaries within a relevant corpus, the response generator system 114 can apply one or more LLM-based response models 115, and deliver a well-informed response, data and / or documents. In some embodiments, the query response 520 can be provided through a retrieval-augmented generation (RAG) process where documents and / or synthetic summaries are identified and an answer formulated through one or more large language models 115 providing a contextually relevant response based on the comprehension of an organizations terms and concepts using the organization’s data corpus and generated Novel hierarchical clustering and summarization techniques for knowledge indexing. The large language models may be implemented through one or more known proprietary and / or known trained language machine learning models such as but not limited to one or more of OpenAI’s GTP series of models, Google’s Gemini, one or more of Meta’s LLaMA family of models, other such models, or a combination of two or more of such models. The user readable response 520, which may include human readable data relevant to the response, can communicated, via a wired and / or wireless transceiver, to an intended recipient user via a respective user interface unit 118, which may or may not be the user that submitted the query 502.

[0044] FIG. 6 illustrates a simplified block diagram representation of an exemplary process 600 of providing a contextually relevant response to a query, in accordance with some embodiments. In step 602, a query is received. In step 604, the query is transformed and embedded with one or more assigned data vectors (v). In some embodiments, the query, embedded query and / or data vector can be forward, in step 606, to the retrieval system 116. In step 608, the retrieval system 116 can generate and / or access the index 320, which in some embodiments, can be generated based on organizationally relevant data 202. The index can include, in some embodiments, the relevant embedded data items 202, synthetic summaries of the data items and cluster synthetic summaries,Aty. Docket No. 1127P017359-WO (PCT) and typically further includes mappings of data items and their respective parent synthetic summaries.

[0045] In step 610, the retrieval system can obtained and / or accesses the relevant data and its context (e.g., synthetic summaries). In some embodiments, the retrieval system can retrieve a set of one or more relevant human readable data 202 related to the query and context of the relevant human readable data from the data store. The context, in some embodiments, can include at least a subset of one or more of the hierarchical synthetic summaries corresponding to the set of one or more relevant human readable data. In step 612, the relevant data and its content can be communicated to and / or identified to the response generator system 114. In step 614, the response generator system 114 can generate a user readable response 520 and / or answer. The response generator system 114, in some embodiments, can apply one or more one or more large language models to at least the relevant human readable data and the relevant contextual information (e.g., hierarchical synthetic summaries) of the relevant human readable data and / or cluster, and generate a contextual response to the query based on the retrieved relevant human readable data and the context of the relevant human readable data, communicate the contextual response to the respective user interface unit in response to the query. In step 616, the response 520 is communicated to and / or otherwise made accessible by the requesting user through the respective user interface system.

[0046] FIG. 7 illustrates a simplified flow diagram of an exemplary process 700 of providing a contextual response to a query 502, in accordance with some embodiments. In step 702, a query is received from a user interface unit 118 associated with a query source (e.g., a user). In step 704, relevant human readable data is accessed that is related to the query. Some embodiments convert the query to a query vector correspond to the query. In step 706, a respective context of the relevant human readable data is accessed from one or more data stores 104 for each of the accessed relevant human readable data. As described above, the respective context can comprise hierarchically dependent synthetic summaries (e.g., 208, 214, 220) related to the relevant human readable data. Accordingly, in some embodiments, one or more hierarchical indexes 320 of the data store 104 is accessed and / or interacted with to identify the relevant human readable data related to the query and the context of the relevant human readable data. The query vector, when generated, can be used to identify the relevant human readable data related to the query and the context of theAty. Docket No. 1127P017359-WO (PCT) relevant human readable data (e.g., based on one or more threshold correlations between the query vector and one or more data vectors generated for the different human readable data).

[0047] In step 708, the relevant human readable data and the context of the relevant human readable data can be transferred to one or more generative artificial intelligence models 115 of the response generator system 114. In step 710, a contextual response to the query is generated based on the retrieved relevant human readable data and the context of the relevant human readable data and the application of the generative artificial intelligence model. In step 712, the contextual response, from the generative artificial intelligence model, can be communicate to the user interface unit 118 associated with the query in controlling the user interface unit to present the contextual response. One or more of the above steps may be repeated substantially any number of times. For example, in some embodiments, step 708 may optionally be repeated one or more times for one or more different human readable data identified as relevant to the query. In some embodiments, the response can control the user interface unit 118 to present the response to the user. For example, in some embodiments, a user interface of the user interface unit can display the contextual response to the query. Further, in some implementations, the identified one or more relevant human readable data that relate to the query can be included in the response, and in some embodiments, is further displayed with the contextual response to the query.

[0048] FIG. 8 illustrates a simplified block diagram of an exemplary process 800 that applies a series of computer implemented rules that automate the generating of a hierarchical indexing for human readable data, in accordance with some embodiments. The indexing process 800, in some embodiments, can be utilized in cooperation with the response process 700 to provide a contextual response to a query. In step 802, at least one data vector is associated with each of the human readable data 202 stored in the data store 104. The data vectors can, in some implementations, provide machine readable representations of the respective human readable data. In step 804, first level synthetic summaries 208 can be generated for human readable data clusters 206 of the human readable data 202. Typically, each data cluster 206 comprising one of the human readable data 202 or a plurality of related human readable data 202. In defining which of the human readable data are within a defined data cluster 206, each data cluster can be selected based on proximity of the respective data vectors (e.g., having a predefined threshold relationship with one or more other data vectors).Aty. Docket No. 1127P017359-WO (PCT)

[0049] In step 806, one or more first level synthetic summary vectors 210 can be assigned to a respective one of each of the first level synthetic summaries 208 providing machine readable representation of the first level synthetic summaries. In step 808, second level synthetic summaries 214 can be generated for the first level synthetic summary clusters 212 of the first level synthetic summaries 208. In some embodiments, the association of one or more of the first level synthetic summaries 208 in defining each of the respective first level synthetic summary clusters 212 can be selected based at least in part on a relative proximity of the first level synthetic summary vectors. In some embodiments, in one or both of the steps of generating the first level synthetic summaries 204 and generating the second level synthetic summaries 208 can be generated through the application of one or more large language models.

[0050] In step 810, one or more second level synthetic summary vectors 216 can be assigned to each of the second level synthetic summaries 214 providing machine readable representations of the respective second level synthetic summaries 214. In step 812, the hierarchical index is defined and / or created based on a collection of the data vectors 204, the first level synthetic summary vectors 210, and the second level synthetic summary vectors 216. The index can, in some embodiments, further include the embedded human readable data associated with respective one or more data vectors, the embedded first level synthetic summaries associated with respective one or more first level synthetic summary vectors, and the embedded second level synthetic summaries associated with respective one or more second level synthetic summary vectors. Again, as described above, one or more indexes may include one or more additional or fewer levels, as a function of a number of human readable data 202 determined to have a threshold relevance based on one or more thresholds to the query. Typically, the levels of synthetic summaries are generated (e g., higher levels of synthetic summaries) until a single synthetic summary is generated at a particular level.

[0051] FIG. 9 illustrates a simplified flow diagram of an exemplary process 900 of generating a contextual response to a user query, in accordance with some embodiments. In step 902, human readable data 202 can be stored in one or more data stores 104. In step 904, the human readable data is transformed to generate one or more data vectors 204 that are assigned and / or associated with the respective one of the human readable data providing for machine readable representations of the human readable data. In step 906, the data vectors 204 and / or the embedded human readableAty. Docket No. 1127P017359-WO (PCT) data can be stored in the data store 104. In step 908, first level synthetic summaries 208 for data clusters 206 of the human readable data are stored in the data store 104. Typically, each data cluster 206 corresponds to and / or comprises one of the related human readable data or a plurality of the related human readable data, with each of the one or more human readable data 202 of each data cluster 206 having been selected to be associated with the respective data cluster based at least in part on a proximity of the data vectors.

[0052] In step 910, one or more first level synthetic summary vectors 210 are defined and assigned to each of the first level synthetic summaries 208 providing machine readable representations of the first level synthetic summaries. In step 912, the first level synthetic summary vectors 210 and / or embedded first level synthetic summaries are stored in the data store 104. In step 914, one or more second level synthetic summaries are generated for each of first level synthetic summary clusters 212 of the first level synthetic summaries 208. In some embodiments, the first level synthetic summary clusters for the first level synthetic summaries can be selected as part of the respective summary cluster 212 based at least in part on a threshold correlation (e.g., proximity) of the first level synthetic summary vectors 210. In some embodiments, the first level synthetic summaries, and the second level synthetic summaries are generated using one or more generative artificial intelligence models.

[0053] In step 916, one or more second level synthetic summary vectors 216 can be generated for each of one or more second level synthetic summaries 214 providing machine readable representation of the second level synthetic summaries. In step 918, the one or more second level synthetic summary vectors 216 and / or one or more embedded second level synthetic summaries 214 can be stored in the data store 104. Further, a hierarchical index can be maintained defining the hierarchical correlations between different human readable data 202, the corresponding data clusters, the corresponding synthetic summaries and the corresponding synthetic summary clusters. In some embodiments, the hierarchical index may further include the data vectors and / or the synthetic summary vectors.

[0054] FIG. 10 illustrates a simplified block diagram of an exemplary process 1000 of creating a hierarchical indexing of human readable data, in accordance with some embodiments. In step 1002, one or more human readable data 202 is converted to an embedding, which in someAty. Docket No. 1127P017359-WO (PCT) embodiments can include the generation of one or more data vectors using one or more models and / or algorithms (e.g., OpenAI ADA). In step 1004, a retrieval system is initiated typically with access to the human readable data. In step 1006, one of the human readable data 202 is selected. In step 1008, the process attempts to identify nearest neighbors of the chosen human readable data 202 in establishing a data cluster 206 of the human readable data 202. The identification of a neighbor may be identified, in some embodiments, based on a threshold relationship between the data vectors 204. In some instances one or more nearest neighbor human readable data may be identified, while in other instances for some human readable data there may not be a single nearest neighbor that satisfies a respective neighbor threshold.

[0055] In step 1010, a synthetic summary document is created based on cluster and the nearest neighbors within the cluster. In step 1012, some embodiments remove the selected human readable data, and in some implementations similarly remove the other human readable data of the nearest neighbor clusters from the original collection such that these are not subsequently considered during a particular indexing of the collection of human readable data being processed. In step 1014, the generated synthetic summary is incorporated into a next level collection (e.g., collection 331, 332, 333, 334) depending on a current level. In step 1016, each synthetic summary document is embedded. Some embodiments include step 1018 where it is determined whether levels have been processed until only a single synthetic summary document remains. Steps 1006-1018 can be repeated on a next level collection of synthetic summaries in the generation of sub-level collections until only the single synthetic summary document remains representing the hierarchical summary of all of the original human readable data 202. For example, in some embodiments, in step 1006 a synthetic summary is selected, in step 1008 a synthetic summary cluster of nearest neighboring synthetic summaries is identified, in step 1010 a next level synthetic summary is generated for the synthetic summary cluster, in step 1012 the selected synthetic summary and in some instances nearest neighbor synthetic summaries are removed from further consideration of the current level collection (e.g., collection 333), in step 1014 the current level generated synthetic summary is added to a next level collection (e.g., collection 334), in step 1016 the generated synthetic summary is embedded in the collection (e.g., collection 334), and in step 1016 it is determined whether the generated synthetic summary is a single remaining synthetic summary.Aty. Docket No. 1127P017359-WO (PCT)

[0056] In some embodiments, the query system 110 includes and / or applies one or more large language models (LLM), providing an LLM-powered system to process the received queries 502. The query system 110 can, in some implementations, generate and manage one or more query indexes by in part indexing received queries using the content (e.g., corresponding questions) of the respective queries as keys. In some embodiments, indexing values can be triplets comprising the question, sample query, and sample output. The query indexer can be based on embeddings of the keys (e.g., questions, terms, phrases, values, ranges, etc.). The indexed values and / or parameters can be retrieved based on a given received query. Again, in some instances, the values may comprise triplets. The query system 110, in some embodiments, can apply one or more k- nearest neighbors algorithms (KNN) with cosine similarity scores used to rank the retrieved triplets. Subsequently, the scores for the list of triplets may be normalized (e.g., using SoftMax). The results can be used to identify more relevant triples, and in some embodiments triplets are limited to retrieving a top-p triplets.

[0057] In some embodiments, the query system 110 can apply one or more algorithms and / or query generation machine learning models (e.g., LLM operating in a few-shot generation mode) to generate an appropriate index query that addresses the received query 502 from the user. The query system or other document management agent can leverage the indexed question-query- output triplets to respond to more complex inquiries. In some embodiments, the LLM powered query system can execute these generated queries on the system, and return one or more outputs, and feedback can be received over time (e.g., humans validations of the effectiveness of these triplets and the generated responses). The query system agent can subsequently index these validated triplets for future use in addressing increasingly complex questions. As one non-limiting example, a question can be received from a user, and the query LLM-based model can process the query, generate a database query, which can be utilized at least in part by the retrieval system 116 in retrieving more contextually relevant data that can be provide in an answers 520 the user. By utilizing the generated glossaries, the query system and retrieval system can be capable of comprehending the received queries 502 or questions within the context of an organization's specific terms and concepts. For example, a query may be received regarding an availability of a particular store (e.g., store number 052), where the concept of "availability" may vary between organizations (e.g., operating hours vs. functioning without issues). Based on a more relevant understanding as a function of at least the glossary, the system provides a more content relevantAty. Docket No. 1127P017359-WO (PCT) indexed query. The query, in some embodiments, can be augmented with relevance to the organization’s internal and / or external terms and concepts, providing more relevant context. The retrieval system can utilize the augmented query in evaluating indexed data and collections in identifying more contextually relevant information and / or answer. In some embodiments, the query system 110 is further configured to determine when further clarification is needed for a query, and can utilize the glossary in part to identify the clarification. For example, the query system 110 may processes a query, and determine further knowledge of one or more specific terms or domain expertise would be beneficial and / or needed in contextually understanding the query. A query LLM can utilize the pone or more generated glossaries to accurately comprehend the query and / or retrieve pertinent documents from the corpus. In some embodiments, the query system may utilize a clarification database query using appropriate terms from the corpus and glossaries in delivering a well-informed more contextually relevant response.

[0058] Some prior processing for document management utilize vector databases to implement retrieval -augmented generation (RAG), with each document indexed based on some metadata. The use of such prior RAG methods in organizations fail to provide context to the documents, and each single document is indexed and retrieved without context. This results in time lost to answer a question requiring contextual understanding.

[0059] However, the document management and content retrieval system 100, in some embodiments, provides an LLM-based system that preserves context, in part through contextual indexing and contextual retrieval of documents and / or corresponding information. This differentiates from prior RAG solutions by providing relevant responses (e.g., answers and / or documents) based on a particular context, which can include for example at least an the organization's unique context. The system can preserve the context of the user readable data 202 and / or other relevant documents and / or information during indexing and retrieval stages. Further, the system improves document indexing and retrieval and document-based question-answering in respective organizations, at least in part through the clustering and / or synthetic summarization of the clusters. The document management and content retrieval system 100 can deliver contextual answers using a company's data corpus, through the hierarchical clustering and summarization techniques providing knowledge indexing. Further, in some embodiments, the document management and content retrieval system 100 can streamline and optimize the process of creating,Aty. Docket No. 1127P017359-WO (PCT) managing, and evolving queries of real-time data within an organization. The system can store, manage, and generate queries for real-time data, through query indexing, query retrieval, query generation, and query evolution over time through feedback and the repeated retraining over time of the one or more machine learning models.

[0060] Further, the circuits, circuitry, systems, devices, processes, methods, techniques, functionality, services, servers, sources and the like described herein may be utilized, implemented and / or run on many different types of devices and / or systems. FIG. 11 illustrates an exemplary system 1100 that may be used for implementing any of the components, circuits, circuitry, systems, functionality, apparatuses, processes, or devices of the document management and content retrieval system 100, and / or other above or below mentioned systems or devices, or parts of such circuits, circuitry, functionality, systems, apparatuses, processes, or devices. For example, the system 1100 may be used to implement some or all of the query system 110, the indexing system 106, the glossary generator system 109, the retrieval system 116, the response generator system 114, the user interface units, and / or other such components, circuitry, functionality and / or devices. However, the use of the system 1100 or any portion thereof is certainly not required.

[0061] By way of example, the system 1100 may comprise one or more control circuits or processor modules 1112, one or more memory 1114, and one or more communication links, paths, buses or the like 1 118. Some embodiments may include one or more user interfaces 1 116, and / or one or more internal and / or external power sources or supplies 1140. The control circuit 1112 can be implemented through one or more processors, microprocessors, central processing unit, logic, local digital storage, firmware, software, and / or other control hardware and / or software, and may be used to execute or assist in executing the steps of the processes, methods, functionality and techniques described herein, and control various communications, decisions, programs, content, listings, services, interfaces, logging, reporting, etc. Further, in some embodiments, the control circuit 1112 can be part of control circuitry and / or a control system 1110, which may be implemented through one or more processors with access to one or more memory 1114 that can store instructions, code and the like that is implemented by the control circuit and / or processors to implement intended functionality. In some applications, the control circuit and / or memory may be distributed over a communications network (e.g., LAN, WAN, Internet) providing distributed and / or redundant processing and functionality. Again, the system 1100 may be used to implementAty. Docket No. 1127P017359-WO (PCT) one or more of the above or below, or parts of, components, circuits, systems, processes and the like.

[0062] The user interface 1116 can allow a user to interact with the system 1100 and receive information through the system. In some instances, the user interface 1116 includes a display 1122 and / or one or more user inputs 1124, such as buttons, touch screen, track ball, keyboard, mouse, etc., which can be part of or wired or wirelessly coupled with the system 1100. Typically, the system 1100 further includes one or more communication interfaces, ports, transceivers 1120 and the like allowing the system 1100 to communicate over a communication bus, a distributed computer and / or communication network 108 (e.g., a local area network (LAN), the Internet, wide area network (WAN), etc.), communication link 1118, other networks or communication channels with other devices and / or other such communications or combination of two or more of such communication methods. Further the transceiver 1120 can be configured for wired, wireless, optical, fiber optical cable, satellite, or other such communication configurations or combinations of two or more of such communications. Some embodiments include one or more input / output (I / O) ports 1134 that allow one or more devices to couple with the system 1100. The I / O ports can be substantially any relevant port or combinations of ports, such as but not limited to USB, Ethernet, or other such ports. The I / O interface 1134 can be configured to allow wired and / or wireless communication coupling to external components. For example, the I / O interface can provide wired communication and / or wireless communication (e.g., Wi-Fi, Bluetooth, cellular, RF, and / or other such wireless communication), and in some instances may include any known wired and / or wireless interfacing device, circuit and / or connecting device, such as but not limited to one or more transmitters, receivers, transceivers, or combination of two or more of such devices.

[0063] The system 1100 comprises an example of a control and / or processor-based system with the control circuit 1112. Again, the control circuit 1112 can be implemented through one or more processors, controllers, central processing units, logic, software and the like. Further, in some implementations the control circuit 1112 may provide multiprocessor functionality.

[0064] The memory 1114, which can be accessed by the control circuit 1112, typically includes one or more processor-readable and / or computer-readable media accessed by at least the control circuit 1112, and can include volatile and / or nonvolatile media, such as RAM, ROM, EEPROM,Aty. Docket No. 1127P017359-WO (PCT) flash memory and / or other memory technology. Further, the memory 1114 is shown as internal to the control system 1110; however, the memory 1114 can be internal, external or a combination of internal and external memory. Similarly, some or all of the memory 1114 can be internal, external or a combination of internal and external memory of the control circuit 1112. The external memory can be substantially any relevant memory such as, but not limited to, solid-state storage devices or drives, hard drive, one or more of universal serial bus (USB) stick or drive, flash memory secure digital (SD) card, other memory cards, and other such memory or combinations of two or more of such memory, and some or all of the memory may be distributed at multiple locations over the computer network 108. The memory 1114 can store code, software, executables, scripts, data, content, lists, programming, programs, log or history data, user information, customer information, product information, and the like. While FIG. 11 illustrates the various components being coupled together via a bus, it is understood that the various components may actually be coupled to the control circuit and / or one or more other components directly.

[0065] As described above, one or more machine learning models and / or generative artificial intelligence can be used. These machine learning models and / or Al can be trained and retrained over time with one or more corpuses of information and / or feedback. Still further, some embodiments utilize artificially generated training and / or re-training data. Such data can be generated to simulate one or more content, data vectors, synthetic summaries, conditions, other such information or a combination of such information. The training data can be dependent on the type of machine learning model or models employed. The machine learning models and / or modeling applications further include the trained, deep learning models that process the data. The learning models can be substantially any relevant modeling, whether custom developed or acquired by a third party. For example, in some embodiments, the trained learning models may include decision trees, XGBOOST, GRIDSEARCHCV, unsupervised learning, regression, clustering, TENSORFLOWLITE model, MOBILENETV2 model, ML KIT for FIREBASE, and substantially any other relevant modeling and supporting applications (e.g., CORE ML, VISION FRAMEWORK, CAFFE, KERAS, XGBOOST, TENSORFLOW, etc.) to implement the modeling. Additionally or alternatively, the machine learning models can comprise a neural network machine learning model, a convolutional neural network, Bayesian network learning, dynamically learned behavior based on, for example, decision tree learning, association ruleAty. Docket No. 1127P017359-WO (PCT) learning, inductive logic learning, support vector learning, cluster analysis learning, Bayesian network learning, and / or similarity and metric learning, and / or other such modeling.

[0066] Some embodiments provide systems to generate a contextual response to a user query comprising: a query system comprising a query control circuit; an indexing system comprising an indexing control circuit configured to execute indexing code; a synthetic summary generation system comprising a summary control circuit configured to execute summaries generation code; and a data store communicatively coupled to the query control circuit and the indexing control circuit, the data store comprising human readable data; wherein the indexing control circuit, in executing the indexing code, is configured to transform each of the human readable data into respective data vectors assigned to the respective human readable data providing machine readable representations of the human readable data and store the data vectors in the data store, wherein each of the human readable data is assigned a different data vector of the data vectors; wherein the summary control circuit, in executing the summary code, is configured to access the human readable data and generate first level synthetic summaries generated for data clusters of the human readable data, each data cluster, of the data clusters, comprising one of the human readable data or a plurality of related human readable data of the human readable data, wherein each data cluster is generated based on a relative vector proximity of the data vectors; wherein the indexing control circuit is further configured to transform each of the first level synthetic summaries into respective first level synthetic summary vectors assigned to the respective first level synthetic summaries providing machine readable representations of the first level synthetic summaries, wherein each of the first level synthetic summaries is assigned a different first level synthetic summary vector of the first level synthetic summary vectors; wherein the summary control circuit is further configured to generate second level synthetic summaries generated for clusters of the first level synthetic summaries, wherein the clusters for the first level synthetic summaries are selected based on a relative vector proximity of the first level synthetic summary vectors and store the second level synthetic summaries in the data store; and wherein the indexing control circuit is further configured to transform each of the second level synthetic summaries into respective second level synthetic summary vectors assigned to the respective second level synthetic summaries providing machine readable representations of the second level synthetic summaries.Aty. Docket No. 1127P017359-WO (PCT)

[0067] Some embodiments provide methods for providing a contextual response to a user query, comprising: receiving a query from a user interface unit associated with a query source; accessing relevant human readable data related to the query; accessing, for each of the accessed relevant human readable data, a respective context of the relevant human readable data from a data store, the respective context comprising hierarchically dependent synthetic summaries related to the relevant human readable data; transferring the relevant human readable data and the context of the relevant human readable data to a generative artificial intelligence model; generating, based on the application of the generative artificial intelligence model, a contextual response to the query based on the retrieved relevant human readable data and the context of the relevant human readable data; and communicating the contextual response, from the generative artificial intelligence model, to the user interface unit associated with the query in controlling the user interface unit to present the contextual response.

[0068] Some embodiments provide methods for use in generating a contextual response to a user query, the method comprising: storing human readable data in a data store; assigning data vectors to the human readable data for machine readable representations of the human readable data; storing the data vectors in the data store; storing first level synthetic summaries for data clusters of the human readable data, each data cluster comprising one of the human readable data or a plurality of related human readable data, wherein each of the one or more human readable data of each data cluster is selected to be associated with the respective data cluster based on a proximity of the data vectors; assigning first level synthetic summary vectors to the first level synthetic summaries providing machine readable representations of the first level synthetic summaries; storing the first level synthetic summary vectors in the data store; storing second level synthetic summaries for clusters of the first level synthetic summaries, wherein the clusters for the first level synthetic summaries are selected based on a proximity of the first level synthetic summary vectors; assigning second level synthetic summary vectors to the second level synthetic summaries providing machine readable representation of the second level synthetic summaries; and storing the second level synthetic summary vectors in the data store.

[0069] Some embodiments provide systems to generate a contextual response to a user query, comprising: a control circuit; and a data store communicatively coupled to the control circuit, the data store comprising: human readable data; data vectors assigned to the human readable data forAty. Docket No. 1127P017359-WO (PCT) machine readable representations of the human readable data; first level human readable synthetic summaries generated for data clusters of the human readable data, each data cluster comprising one of the human readable data or a plurality of related human readable data, wherein each data cluster is selected based on a proximity of the data vectors; first level synthetic summary vectors assigned to the first level human readable synthetic summaries for machine readable representations of the first level human readable synthetic summaries; second level human readable synthetic summaries generated for clusters of the first level human readable synthetic summaries, wherein the clusters for the first level human readable synthetic summaries are selected based on a proximity of the first level synthetic summary vectors; and second level synthetic summary vectors assigned to the second level human readable synthetic summaries for machine readable representations of the second level human readable synthetic summaries.

[0070] Further embodiments provide methods of generating a contextual response to a query, comprising: transforming, in indexing human readable data, each human readable data into respective data vectors assigned to the respective human readable data providing machine readable representations of the human readable data and storing the data vectors in a data store, wherein each of the human readable data is assigned a different data vector of the data vectors; accessing the human readable data and generating first level synthetic summaries generated for data clusters of the human readable data, each data cluster, of the data clusters, comprising one of the human readable data or a plurality of related human readable data of the human readable data, wherein each data cluster is generated based on a relative vector proximity of the data vectors; transforming each of the first level synthetic summaries into respective first level synthetic summary vectors assigned to the respective first level synthetic summaries providing machine readable representations of the first level synthetic summaries, wherein each of the first level synthetic summaries is assigned a different first level synthetic summary vector of the first level synthetic summary vectors; generating second level synthetic summaries generated for clusters of the first level synthetic summaries, wherein the clusters for the first level synthetic summaries are selected based on a relative vector proximity of the first level synthetic summary vectors and storing the second level synthetic summaries in the data store; and transforming each of the second level synthetic summaries into respective second level synthetic summary vectors assigned to the respective second level synthetic summaries providing machine readable representations of the second level synthetic summaries.Aty. Docket No. 1127P017359-WO (PCT)

[0071] Those skilled in the art will recognizethat a wide variety of other modifications, alterations, and combinations can also be made with respect to the above described embodiments without departing from the scope of the invention, and that such modifications, alterations, and combinations are to be viewed as being within the ambit of the inventive concept.

[0072] What is claimed is:

Claims

Aty. Docket No. 1127P017359-WO (PCT)Claims1. A system to generate a contextual response to a user query, the system comprising: a query system comprising a query control circuit; an indexing system comprising an indexing control circuit configured to execute indexing code; a synthetic summary generation system comprising a summary control circuit configured to execute summaries generation code; and a data store communicatively coupled to the query control circuit and the indexing control circuit, the data store comprising human readable data; wherein the indexing control circuit, in executing the indexing code, is configured to transform each of the human readable data into respective data vectors assigned to the respective human readable data providing machine readable representations of the human readable data and store the data vectors in the data store, wherein each of the human readable data is assigned a different data vector of the data vectors; wherein the summary control circuit, in executing the summary code, is configured to access the human readable data and generate first level synthetic summaries generated for data clusters of the human readable data, each data cluster, of the data clusters, comprising one of the human readable data or a plurality of related human readable data of the human readable data, wherein each data cluster is generated based on a relative vector proximity of the data vectors; wherein the indexing control circuit is further configured to transform each of the first level synthetic summaries into respective first level synthetic summary vectors assigned to the respective first level synthetic summaries providing machine readable representations of the first level synthetic summaries, wherein each of the first level synthetic summaries is assigned a different first level synthetic summary vector of the first level synthetic summary vectors; wherein the summary control circuit is further configured to generate second level synthetic summaries generated for clusters of the first level synthetic summaries, wherein the clusters forAty. Docket No. 1127P017359-WO (PCT) the first level synthetic summaries are selected based on a relative vector proximity of the first level synthetic summary vectors and store the second level synthetic summaries in the data store; and wherein the indexing control circuit is further configured to transform each of the second level synthetic summaries into respective second level synthetic summary vectors assigned to the respective second level synthetic summaries providing machine readable representations of the second level synthetic summaries.

2. The system of claim 1, wherein at least one of the first level synthetic summaries and the second level synthetic summaries is generated applying a large language model.

3. The system of claim 1, wherein the data store further comprises third level synthetic summaries generated for clusters of the second level synthetic summaries, wherein the clusters for the second level synthetic summaries are selected based on a relative vector proximity of the second level synthetic summary vectors.

4. The system of claim 3, wherein the data store further comprises third level synthetic summary vectors assigned to the third level synthetic summaries for machine readable representation of the third level synthetic summaries.

5. The system of claim 1, wherein the data store further comprises additional higher levels of synthetic summaries.

6. The system of claim 5, wherein the additional higher levels of synthetic summaries are generated applying a large language model.

7. The system of claim 1, further comprises: a response generator system comprising a response generator control circuit configured to execute response generator code; wherein the query control circuit is configured to receive a query from a user interface unit associated with a user;Aty. Docket No. 1127P017359-WO (PCT) wherein the retrieval system is configured to: access a first set of one or more relevant human readable data related to the query and context of the first set of one or more relevant human readable data from the data store; and transfer the first set of one or more relevant human readable data and the context of the relevant human readable data to the response generator system; and wherein the response generator control circuit is configured to: apply a language model to generate a contextual response to the query based on the retrieved relevant human readable data and the context of the relevant human readable data; and communicate the contextual response to a respective user interface unit associated with the query.

8. The system of claim 7, wherein the context comprising a subset of one or more hierarchical synthetic summaries comprising one of the first level synthetic summary or the second level synthetic summary related to each of the first set of one or more relevant human readable data.

9. The system of claim 7, further comprises a retrieval system comprising a retrieval control circuit, wherein the retrieval control circuit is configured to interact with an index of the data store to identify the first set of one or more relevant human readable data related to the query, wherein the index comprises: the data vectors corresponding to the human readable data; the first level synthetic summary vectors corresponding to the first level synthetic summaries; and the second level synthetic summary vectors corresponding to the second level synthetic summaries.

10. The system of claim 7, wherein the query system further comprises a query vectorization system comprising a query vectorization control circuit configured to:Aty. Docket No. 1127P017359-WO (PCT) generate a query vector comprising a machine readable representation of a context of the query; and interact with an index of the data store with the query vector to identify the first set of one or more relevant human readable data related to the query.

11. A method for providing a contextual response to a user query, the method comprising: receiving a query from a user interface unit associated with a query source; accessing relevant human readable data related to the query; accessing, for each of the accessed relevant human readable data, a respective context of the relevant human readable data from a data store, the respective context comprising hierarchically dependent synthetic summaries related to the relevant human readable data; transferring the relevant human readable data and the respective context of the relevant human readable data to a generative artificial intelligence model; generating, based on the application of the generative artificial intelligence model, a contextual response to the query based on the relevant human readable data and the respective context of the relevant human readable data; and communicating the contextual response, from the generative artificial intelligence model, to the user interface unit associated with the query in controlling the user interface unit to present the contextual response.

12. The method of claim 11, further comprising interacting with an index of the data store to identify the relevant human readable data related to the query and the respective context of the relevant human readable data.

13. The method of claim 11, further comprising creating an index of the data store comprising: assigning data vectors to human readable data stored in the data store for machine readable representation of the human readable data;Aty. Docket No. 1127P017359-WO (PCT) generating first level synthetic summaries for data clusters of the human readable data, each data cluster comprising one of the human readable data or a plurality of related human readable data, wherein each data cluster is selected based on a proximity of the data vectors; assigning first level synthetic summary vectors to the first level synthetic summaries for machine readable representation of the first level synthetic summaries; generating second level synthetic summaries for clusters of the first level synthetic summaries, wherein the clusters for the first level synthetic summaries are selected based on a proximity of the first level synthetic summary vectors; assigning second level synthetic summary vectors to the second level synthetic summaries for machine readable representation of the second level synthetic summaries; and collecting, to create the index, the data vectors, the first level synthetic summary vectors, and the second level synthetic summary vectors.

14. The method of claim 13 wherein at least one of the generating the first level synthetic summaries and the generating the second level synthetic summaries comprises applying a large language model to generate at least one of the first level synthetic summaries and the second level synthetic summaries.

15. The method of claim 13 further comprising generating additional higher levels of synthetic summaries until a single synthetic summary is generated.

16. The method of claim 11, further comprising: converting the query to a query vector correspond to the query; and interacting with an index of the data store with the query vector to identify the relevant human readable data related to the query and the context of the relevant human readable data.

17. The method of claim 11 further comprising displaying, via a user interface, the contextual response to the query.Aty. Docket No. 1127P017359-WO (PCT)18. The method of claim 17 further comprising displaying, via the user interface, the relevant human readable data related to the query along with the contextual response to the query.

19. A method for use in generating a contextual response to a user query, the method comprising: storing human readable data in a data store; assigning data vectors to the human readable data for machine readable representations of the human readable data; storing the data vectors in the data store; storing first level synthetic summaries for data clusters of the human readable data, each data cluster comprising one of the human readable data or a plurality of related human readable data, wherein each of the human readable data of each the data clusters is selected to be associated with the respective data cluster based on a proximity of the data vectors; assigning first level synthetic summary vectors to the first level synthetic summaries providing machine readable representations of the first level synthetic summaries; storing the first level synthetic summary vectors in the data store; storing second level synthetic summaries for clusters of the first level synthetic summaries, wherein the clusters for the first level synthetic summaries are selected based on a proximity of the first level synthetic summary vectors; assigning second level synthetic summary vectors to the second level synthetic summaries providing machine readable representation of the second level synthetic summaries; and storing the second level synthetic summary vectors in the data store.

20. The method of claim 19 wherein the first level synthetic summaries, and the second level synthetic summaries are generated applying a generative artificial intelligence model.