Method, apparatus, and system for embedding content from documents as machine-learning-document vectors for filtering and / or visualization
The system addresses the challenge of information overload by embedding documents as vectors, calculating similarity matrices, and organizing them into groups for efficient filtering and visualization, enhancing understanding and reducing computational overhead.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- INMTEL INC
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-07
AI Technical Summary
The rapid increase in available information due to advances in computing and internet access has led to an overwhelming amount of low-utility content, making it difficult for humans to comprehend and analyze large amounts of information efficiently, with existing technologies like search engines and large language models (LLMs) failing to effectively filter and visualize relevant insights.
A system that embeds document content as vectors, calculates similarity values, and generates a similarity matrix to filter and visualize semantic relationships among documents, using techniques such as clustering and similarity metrics to organize documents into groups, and apply filtering and visualization methods to enhance understanding.
Enables efficient processing and visualization of large document collections, reducing computational overhead and resource consumption while providing intuitive visual representations of semantic relationships, allowing for quicker insights and decision-making.
Smart Images

Figure US2024053897_07052026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: P10445PC00 PatentMETHOD, APPARATUS, AND SYSTEMFOR EMBEDDING CONTENT FROM DOCUMENTS AS MACHINE-LEARNINGDOCUMENT VECTORS FOR FILTERING AND / OR VISUALIZATIONBACKGROUND
[0001] Advances in computing, especially since the advent of the internet, have enabled easier publication and access of audiovisual and textual content (i.e., documents). This has led to an unprecedented explosion of available information. Despite the ability to access information more easily than in the pre-internet era, the information explosion does not automatically lead to a greater understanding of the world around us, primarily due to the high amount of low-utility information and limited human ability to comprehend and understand large amounts of information quickly.SOME EXAMPLE EMBODIMENTS
[0002] Therefore, there is a need for an approach for filtering and visualizing large amounts of information or documents via automated means (e.g., machine learning, large language models, etc.).
[0003] According to one embodiment, a method comprises embedding content of a plurality of documents respectively as a plurality of document vectors. The method also comprises calculating a similarity value between each pair of document vectors of the plurality of document vectors to generate a similarity matrix. In one embodiment, the method further comprises filtering the plurality of documents based on the similarity matrix, and providing the filtered plurality of documents as an output. In another embodiment, the method further comprises generating a visual representation of a plurality of semantic relationships among the plurality of documents based on the similarity matrix (in addition or as an alternate to the immediately preceding step).
[0004] According to another embodiment, an apparatus comprises at least one processor, and at least one memory including computer program code for one or more computer programs, at least one memory and the computer program code configured to, with the at least one processor, cause, at least in part, the apparatus to embed content of a plurality of documents respectively as a plurality of document vectors. The apparatus is also caused to calculate a similarity value between each pair of document vectors of the plurality of document vectors to generate a similarity matrix.Attorney Docket No.: P10445PC00 PatentIn one embodiment, the apparatus is further caused to fdter the plurality of documents based on the similarity matrix, and to provide the filtered plurality of documents as an output. In another embodiment, the apparatus is further caused to generate a visual representation of a plurality of semantic relationships among the plurality of documents based on the similarity matrix (in addition or as an alternate to the immediately preceding step).
[0005] According to another embodiment, a non-transitory computer-readable storage medium carries one or more sequences of one or more instructions which, when executed by one or more processors, cause, at least in part, an apparatus to embed content of a plurality of documents respectively as a plurality of document vectors. The apparatus is also caused to calculate a similarity value between each pair of document vectors of the plurality of document vectors to generate a similarity matrix. In one embodiment, the apparatus is further caused to filter the plurality of documents based on the similarity matrix, and to provide the filtered plurality of documents as an output. In another embodiment, the apparatus is further caused to generate a visual representation of a plurality of semantic relationships among the plurality of documents based on the similarity matrix (in addition or as an alternate to the immediately preceding step).
[0006] According to another embodiment, an apparatus comprises means for embedding content of a plurality of documents respectively as a plurality of document vectors. The apparatus also comprises means for calculating a similarity value between each pair of document vectors of the plurality of document vectors to generate a similarity matrix. In one embodiment, the apparatus further comprises means for filtering the plurality of documents based on the similarity matrix, and providing the filtered plurality of documents as an output. In another embodiment, the apparatus further comprises means for generating a visual representation of a plurality of semantic relationships among the plurality of documents based on the similarity matrix (in addition or as an alternate to the immediately preceding step).
[0007] According to another embodiment, a computer program product comprises one or more instructions which, when executed by one or more processors, cause, at least in part, an apparatus to embed content of a plurality of documents respectively as a plurality of document vectors. The apparatus is also caused to calculate a similarity value between each pair of document vectors of the plurality of document vectors to generate a similarity matrix. In one embodiment, the apparatus is further caused to filter the plurality of documents based on the similarity matrix, and to provideAttorney Docket No.: P10445PC00 Patent the filtered plurality of documents as an output. In another embodiment, the apparatus is further caused to generate a visual representation of a plurality of semantic relationships among the plurality of documents based on the similarity matrix (in addition or as an alternate to the immediately preceding step).
[0008] In addition, for various example embodiments described herein, the following is applicable: a computer program product may be provided. For example, a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to perform any one or any combination of methods (or processes) disclosed.
[0009] In addition, for various example embodiments of the invention, the following is applicable: a method comprising facilitating a processing of and / or processing (1) data and / or (2) information and / or (3) at least one signal, the (1) data and / or (2) information and / or (3) at least one signal based, at least in part, on (or derived at least in part from) any one or any combination of methods (or processes) disclosed in this application as relevant to any embodiment of the invention.
[0010] For various example embodiments of the invention, the following is also applicable: a method comprising facilitating access to at least one interface configured to allow access to at least one service, the at least one service configured to perform any one or any combination of network or service provider methods (or processes) disclosed in this application.
[0011] For various example embodiments of the invention, the following is also applicable: a method comprising facilitating creating and / or facilitating modifying (1) at least one device user interface element and / or (2) at least one device user interface functionality, the (1) at least one device user interface element and / or (2) at least one device user interface functionality based, at least in part, on data and / or information resulting from one or any combination of methods or processes disclosed in this application as relevant to any embodiment of the invention, and / or at least one signal resulting from one or any combination of methods (or processes) disclosed in this application as relevant to any embodiment of the invention.
[0012] For various example embodiments of the invention, the following is also applicable: a method comprising creating and / or modifying (1) at least one device user interface element and / or (2) at least one device user interface functionality, the (1) at least one device user interface element and / or (2) at least one device user interface functionality based at least in part on data and / orAttorney Docket No.: P10445PC00 Patent information resulting from one or any combination of methods (or processes) disclosed in this application as relevant to any embodiment of the invention, and / or at least one signal resulting from one or any combination of methods (or processes) disclosed in this application as relevant to any embodiment of the invention.
[0013] In various example embodiments, the methods (or processes) can be accomplished on the service provider side or on the mobile device side or in any shared way between service provider and mobile device with actions being performed on both sides.
[0014] For various example embodiments, the following is applicable: An apparatus comprising means for performing a method of the claims.
[0015] Still other aspects, features, and advantages of the invention are readily apparent from the following detailed description, simply by illustrating a number of particular embodiments and implementations, including the best mode contemplated for carrying out the invention. The invention is also capable of other and different embodiments, and its several details can be modified in various obvious respects, all without departing from the spirit and scope of the invention. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The embodiments of the invention are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings:
[0017] FIG. 1 is a diagram of a system capable of embedding content from documents as machine-learning- (ML-) document vectors for filtering and / or visualization, according to one example embodiment;
[0018] FIG. 2 is a flowchart of a process for embedding content from documents as ML- document vectors for filtering and / or visualization, according to one example embodiment;
[0019] FIG. 3 is a diagram of a similarity matrix, according to one example embodiment;
[0020] FIG. 4 is a flowchart of a process for grouping documents for filtering and / or visualization, according to one example embodiment;
[0021] FIG. 5 is a flowchart of a process for generating a brief, according to one example embodiment;Attorney Docket No.: P10445PC00 Patent
[0022] FIG. 6 is a diagram of example user interfaces for generating and presenting a brief, according to one example embodiment;
[0023] FIG. 7 is a flowchart of a process for generating a report, according to one example embodiment;
[0024] FIG. 8 is a diagram of example user interfaces for generating and presenting a report, according to one example embodiment;
[0025] FIG. 9 is a diagram of example user interfaces for generating and presenting an investigation, according to one example embodiment;
[0026] FIG. 10 is a diagram of hardware that can be used to implement an embodiment of the invention; and
[0027] FIG. 11 is a diagram of a chip set that can be used to implement an embodiment of the invention.DESCRIPTION OF SOME EMBODIMENTS
[0028] Examples of a method, apparatus, and computer program for embedding content from documents as machine-learning- (ML-) document vectors for filtering and / or visualization are disclosed, according to various example embodiments. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the invention. It is apparent, however, to one skilled in the art that the embodiments of the invention may be practiced without these specific details or with an equivalent arrangement. In other instances, well-known structures and devices are shown in block diagram form in order to avoid unnecessarily obscuring the embodiments of the invention.
[0029] Reference in this specification to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. The appearance of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments mutually exclusive of other embodiments. In addition, the embodiments described herein are provided by example, and as such, “one embodiment” can also be used synonymously as “one example embodiment.” Further, the terms “a” and “an” herein do not denote a limitation of quantity, but rather denote the presenceAttorney Docket No.: P10445PC00 Patent of at least one of the referenced items. Moreover, various features are described which may be exhibited by some embodiments and not by others. Similarly, various requirements are described which may be requirements for some embodiments but not for other embodiments.
[0030] FIG. 1 is a diagram of a system 100 capable of embedding content from documents as ML-document vectors for filtering and / or visualization, according to one example embodiment. As previously discussed in the background, advances in computing, especially since the advent of the internet, have enabled easier publication and access to vast amounts of audiovisual and textual content (e.g., referred to herein as documents) from a variety of publishers or document sources 101 a-lOln (also collectively referred to as document sources 101). As used herein, the term “document” refers broadly to any audiovisual or textual content accessible to an information platform 103 of the system 100. This includes, but is not limited to, written texts, images, videos, and any other forms of recorded information. It also includes texts parsed or otherwise extracted from the audiovisual content or recorded information. This technological leap has facilitated the democratization of information, allowing anyone with an internet connection to publish and consume content from virtually anywhere in the world. Consequently, users have witnessed an unprecedented explosion of available information across various platforms and media (e.g., documents, texts, information, etc. available from a services platform 105, one or more services 107a-107m - collectively referred to as services 107 - of the services platform 105, one or more content providers 109a-109k - collectively referred to as content providers 109, and / or the like available over a communication network 111).
[0031] However, despite this increased accessibility, the sheer volume of information does not inherently lead to greater understanding or enhanced knowledge. The overwhelming amount of data available often includes a significant proportion of low-utility information, which can obscure the most valuable and relevant insights. Furthermore, the human-cognitive capacity to process and understand large quantities of information remains limited. This discrepancy between the availability of information and a human’s ability to effectively analyze it has presented new challenges in the digital age.
[0032] Moreover, the rapid dissemination of content has not only led to information overload but also to the proliferation of misinformation and low-quality data. As a result, individuals and organizations often struggle to filter through the noise to identify the most credible and pertinentAttorney Docket No.: P10445PC00 Patent information. This challenge underscores the need for advanced tools and techniques that can assist in the efficient and automated sorting, filtering, and analysis of vast datasets, enabling users to extract meaningful insights and make informed decisions.
[0033] While this issue plagues all types of content, it is particularly relevant for text, which accounts for the majority of content published on the internet. Search technology partially mitigates the issue by retrieving relevant text files or documents, yet it cannot filter the information within them. Moreover, although recent advances in large language models (LLMs) (e.g., LLM 113) and / or other equivalent ML / artificial -intelligence (Al) approaches allow extraction of information from large bulks of texts or other documents, they by their very design hallucinate and return erroneous information in a way that is still not systematically understood, as well as omit essential information, especially for larger texts or documents, such as books. Although the various embodiments may refer in some instances to texts, it is contemplated that the embodiments are also applicable to documents in general including but not limited to audiovisual content or other non-text-based content.
[0034] The pitfalls of search engines and LLMs mean that in order to enhance human ability to analyze large chunks of information, an intermediate step, which fdters and / or visualizes the information from distinct text files or documents post search, is needed. Following filtering and / or visualization of texts / documents, the filtered and / or processed documents can be fed to a LLM or otherwise processed to generate outputs such as but not limited to briefs, summaries, reports, graphs, etc. of the data. It is contemplated that the step of filtering and / or visualization can also be used without a reference to search or LLMs.
[0035] Accordingly, one goal of the system 100 is to use retrieved documents to produce automated reports from a subset of information contained within them. However, the system 100 does not a priori know which documents or which information are the most important, thereby presenting a significant technical challenge to addressing this issue using automated means and at scale.
[0036] There are also significant technical challenges with respect to using LLMs. While LLMs are now commonly used to write automated reports, they suffer from several technical limitations and challenges including but not limited to:Attorney Docket No.: P10445PC00 Patent a. Limited context window: e.g., state of the art LLMs currently have approximately a maximum of 1 million tokens, or about 750,000 words, which is about 1,500 text documents. This limited context window is not sufficient for applications that seek to generate automated reports over large collections of documents. b. Cost: even if LLM context window increases, the costs of using LLMs with large context windows can be too high to be consistently deployed in production for a large number of users. These costs include both monetary and resource costs (e.g., processing power, memory, bandwidth, etc.). c. Hallucination: state-of-the-art LLMs can hallucinate, and the issue is more pronounced when texts are not coherent. d. Omission: LLMs leave out important information, especially for large texts.
[0037] Some conventional approaches to minimizing these limitations include retrieval augmented generation (RAG). For example, RAG is a method that combines information retrieval and text generation. It involves fetching relevant information from a database to support the generation of a comprehensive and accurate report or response using LLMs. However, RAG’s integration of LLMs also means that it still suffers from the same technical limitations of LLMs discussed above. In some implementations of RAG, chunking is one commonly used traditional technique to improve RAG. In chunking, only a part of each text is kept as input to LLMs. Chunking can include fixed chunking or semantic chunking. In fixed chunking, a fixed number of words from each text is taken. It is simple but inaccurate as it is not known whether the relevant information is kept. While fixed chunking is computationally and financially cheap, its low accuracy means it is not commonly used. On the other hand, semantic chunking embeds each sentence in each text as a vector, and then algorithmically determines which consecutive sentences to keep based on the similarity between them. While more accurate than fixed chunking, semantic chunking is computationally (and financially) costly to embed each sentence in each text and it is still unknown whether all the key information has been kept. Another issue with RAG is that the relative importance of texts is not addressed, so many of the issues that are associated with LLMs (e.g., hallucinations, omissions, etc.) remain unsolved even if the resulting text is below the LLM context window.Attorney Docket No.: P10445PC00 Patent
[0038] Some conventional approaches to determining the relative importance of texts are based on clustering algorithms. For example, k-means clustering historically has been popular. However, in k-means clustering, one assumes each data point (in this case, a document) belongs to a group. The optimal number of clusters is also determined somewhat subjectively, e.g., through the elbow method. This causes at least a couple of issues. Firstly, all texts are assumed to belong to a group and the number of groups is a heuristic. In reality, many texts are not semantically similar to any other texts in a group, and as a result, are not useful for determining and / or visualizing the development of ideas, concepts, and stories in these texts. Secondly, k-means clustering cannot account for any contextual properties of documents such as time, publisher / author, location, etc. Another conventional approach is based on Latent Dirichlet Allocation (LDA). However, LDA is concerned with assigning “topics” to documents, and not with looking into relationships between the documents themselves.
[0039] To address these technical limitations and challenges, the system 100 of FIG. 1 introduces a capability to filter or eliminate most documents based on their importance before they are fed to LLM 113 or other processed via automated means. Generally, there are many disparate text documents available from multiple document sources 101 over the communication network 111 (e.g., the internet). In one embodiment, the information platform 103 aggregates documents from the document sources 101 in documents database 115 for processing. Some documents aggregated into the documents database 115 may be topically connected, but in most cases, it is not known in advance of processing which documents are connected. In some embodiment, the documents can be indexed for search, and also labels or metadata (e.g., publisher, location, time, link, etc.) associated with them. The labels, for instance, can also serve as identifiers for the documents.
[0040] In one embodiment, documents are fetched from the documents database 115 for processing in response to user demand or request. For example, the fetching of the documents can be triggered by the end user’s activities (e.g., keyword search, location or time filtering, etc.). By way of example, the number of fetched texts or documents can be anywhere from tens to tens of thousands. As an example in the case of text documents, each document may have between 100 and 2,000 words, with about 500 words on average. However, these figures are provided by way of illustration and not as limitations. It is contemplated that the number of words or tokens are notAttorney Docket No.: P10445PC00 Patent limited by the ranges discussed above and can be more than 2,000 or less than 100. As such, the number of documents can easily exceed the context window of the LLM 113 or other ML / AI- based-processing pipelines. Therefore, the various embodiments of filtering described herein solve the technical challenges with associated with LLM-limited context windows, costs, hallucinations, etc.
[0041] In addition to filtering, the various embodiments of the system 100 also introduce a capability to visualize semantic similarities between texts or documents (e.g., as a function of time, publisher, location, etc.). By way of example, this advantageously enables automated mapping the development of ideas, concepts, and stories across time and space for different texts or documents. In one embodiment, the system 100 uses the similarity values (e.g., a similarity matrix representing calculated similarity values between each pair of texts or documents) and text labels (e.g., publication time, publisher, location, etc.) to represent and visualize semantic relationships between the documents.
[0042] In one embodiment, the system 100 initiates filtering and / or visualization based on analyzing the similarity between documents (e.g., documents stored in the documents database 115). The system 100, for instance, calculates the similarity between all pairs of texts or documents that are to be processed. The similarity between texts or documents can either compare the tokens within the texts directly or use embeddings. Embeddings are numerical representations (e.g., vector representations) of text that capture semantic meaning in a multi-dimensional space. By converting the content of entire documents into vectors, embeddings allow for the comparison of texts or documents based on their contextual relationships, rather than mere lexical similarity. In one embodiment, the similarity between pairs of documents can be calculated in a number of ways, such as through Jaccard similarity, cosine similarity, k-means similarity, and / or any other equivalent similarity metrics.
[0043] In one embodiment, the system 100 can optionally use text or document similarities to group the texts or documents together. Grouping documents based on their semantic similarities involves analyzing texts to identify and cluster those with similar semantic relationships (e.g., meanings, themes, context, etc.). This is achieved by using the similarity metrics such as Jaccard or cosine similarity, etc. calculated as described above, for instance, by using embeddings thatAttorney Docket No.: P10445PC00 Patent represent the texts or documents in a multi-dimensional space. Documents are then organized into groups where each group contains texts that are semantically related.
[0044] In one embodiment, the grouping enables more efficient processing, filtering, and / or visualization of large collections of documents. By clustering together documents that share thematic or contextual relationships, the system 100 reduces redundancy and isolates distinct topics or narratives. This approach allows for targeted analysis and streamlined workflows, as similar documents can be processed collectively rather than individually. Efficient grouping advantageously provides the technical benefit of minimizing computational overhead by reducing the number of comparisons needed between documents. When documents are grouped, processing algorithms can operate on clusters as units, rather than performing pairwise comparisons across the entire dataset. This not only accelerates processing times but also lowers resource consumption in terms of memory and processing power.
[0045] In terms of filtering, grouping simplifies the elimination of irrelevant or redundant content. For instance, entire clusters of documents that fall outside the scope of interest can be excluded en masse, rather than filtering document by document. This hierarchical filtering reduces the complexity of the task, making it more manageable and less time-consuming. Visualization of document collections also benefits from semantic grouping.
[0046] In one embodiment, filtering can involve removing specific groups and all the texts within them based on various filtering criteria (e.g., the total number of texts each group contains). For example, depending on the use case for the documents, groups can be removed if they contain fewer than, more than a specific number of texts, or if the number is within a range. In another embodiment, filtering includes removing the similarity values of filtered documents from the similarity matrix itself. Because the similarity matrix may initially contain similarity values for each pair of documents with the number of documents reaching into the tens of thousands, the size of the similarity matrix can be quite large (e.g., 10,000 x 10,000 for a collection of 10,000 documents). As such filtering the similarity matrix can provide technical benefits including but not limited to reduced storage size, reduced computation resources to process the matrix, reduced bandwidth to transmit the matrix, reduced visual complexity for visualizing the matrix, etc.
[0047] In addition or alternatively, the system 100 can remove or filter the texts / documents based on the similarities with other texts / documents (e.g., with or without the optional groupingAttorney Docket No.: P10445PC00 Patent process). The texts can be on one or more filtering criteria. For example, they can be removed based on a number of similarities above a threshold, strongest similarity value or rank based on the number of similarities above a threshold. It is noted that these filtering criteria are provided by way of illustration and not as limitations. It is contemplated that any equivalent filtering criteria, mechanism, or process can be used according to the various embodiments described herein. If the texts have previously been grouped together, they are removed based on calculations within the group they belong to.
[0048] In yet another embodiment, the system 100 can further filter the information within each remaining text or document (e.g., filter the contents within a document as opposed to filtering the entire document). Information or content can be filtered through any of the extractive- and abstractive- text-summarization methods (e.g., via LLM 113 or equivalent). Extractive summarization involves selecting the most critical sentences or phrases directly from the text to create a concise summary (also referred to herein as a “brief’). This method relies on identifying key information and presenting it verbatim, thereby retaining the original context and meaning. Algorithms used in extractive summarization typically rank sentences based on factors such as frequency of terms, positions within the text, and the presence of certain keywords. By compiling the highest-ranked sentences, the system generates a summary or brief that highlights the core content of the document.
[0049] On the other hand, abstractive summarization involves generating new sentences that convey the essential information of the text. This approach is more complex, as it requires understanding the content at a deeper level and then paraphrasing it using natural -language- generation techniques. Abstractive summarization systems often use advanced models (e.g., LLM 113), such as transformer-based architectures, to create coherent and contextually relevant summaries. These summaries or briefs are not mere extracts but are rephrased versions that encapsulate the main ideas and themes of the original document.
[0050] Both extractive and abstractive methods can be customized based on specific criteria or user preferences, allowing for a tailored filtering process that meets the unique needs of various use cases. For instance, extractive summarization might be preferred for technical documents where precision and retention of original wording are paramount, while abstractive summarization might be more suitable for narrative texts where a more fluid and cohesive summary is desired.Attorney Docket No.: P10445PC00 Patent
[0051] In one embodiment, visual representations, such as graphs or maps, can highlight the relationships and evolution of ideas over time and space more clearly when documents or groups of documents are organized based on their semantic relationships (e.g., via the calculated document similarities). Examples of visualizations (also referred to as visual representations) include directed graphs or maps, where documents or groups of documents are represented as nodes and their semantic relationships as connecting lines (e.g., edges), showing clusters of related documents. The direction of the graphs can indicate other contextual parameters (e.g., temporal parameters where the direction of the edge points to earlier documents or documents with more fundamental concepts / topics). This provides users with a more intuitive understanding of the data, allowing for quicker insights and decision making. It is noted that directed graphs are provided as an example of a visual representation, and it is contemplated that any other type of visualization (e.g., heat maps., cluster maps, etc.) can also be used according to the various embodiments described herein.
[0052] In one embodiment, the visualization processes of the system 100 specifically relies on the filtered similarity matrix produced by the document-processing components of the system 100, as traditional similarity measures would produce visualizations too complex for practical use. Conversely, the filtering system's effectiveness is validated and enhanced through the visualization capabilities, creating a technical interdependence between the two components.
[0053] In summary, the system 100 uses calculated similarities between each pair of documents of a collection of documents to be processed. These similarity values for each pair can then be used to filter, visualize, and / or otherwise process the collection of documents more efficiently. In one embodiment, the system 100 processes the documents database 115 to generate output data 117 comprising filtered documents 119 and / or visualization data 121. In one embodiment, the output data 117 (e.g., the filtered documents 119) can be processed using the LLM 113 (or equivalent ML / Al process) to generate briefs 123 (e.g., daily briefs summarizing tens of thousands of documents), reports 125 (e.g., on demand reports on specified topics and / or documents), etc. for presentation on client terminals 127. In addition or alternatively, the output data 1 17 (e.g., the visualization data 121) can be directly presented on the client 127 without further processing (e.g., by LLM 113).Attorney Docket No.: P10445PC00 Patent
[0054] Various embodiments of this process are described in further detail with respect to FIGs. 2-11 below.
[0055] In one embodiment, the various embodiments enable the system 100 to process large collections of documents within a time frame that would not be practical as a manual process. For example, a daily brief 123 that summarizes on average 50,000 documents per day aggregated would be well beyond manual capabilities. Consider a test case, with even a smaller 1000 text document database wherein each document has 500 words on average, and a human called X who does everything by hand, and has their own way of encoding each document into a 50 dimensional (e g., which is much fewer dimensions that a typical embedding which can have over 12,000 or more dimensions) vector based on a model they developed. An average human reads about 250 words per minute, so takes about 2,000 minutes to read these documents. Because X also has to encode each text in a vector, assume encoding takes 2s per vector dimension, meaning they take 1,666 minutes to convert texts into numbers, or 3,666 minutes to label texts in total. If they work 8 hours a day 5 days a week, X takes a week and a half to read and label 1,000 documents. This is already too slow for the daily brief, which has to be completed daily.
[0056] Now, X has to calculate similarities between all the texts. Assuming that vectors are normalized (length 1). Assume that X uses a calculator, and each multiplication takes about 10s to multiply two numbers and write the result down, and 2 more minutes to add 50 numbers on a calculator. This means X takes about 10 minutes to calculate the similarity between two texts. As the matrix of similarities is symmetric, X needs to calculate the similarity between texts about 500,000 times, meaning it takes him about 9.5 years of work to obtain a matrix of similarities. As X works 8 hours a day 5 days a week 50 weeks a year (2,000 hours per year), it would take X 42 working years to calculate the matrix of text similarities assuming he does nothing else in his work. X is average so works about 45 years, and cannot spend 8 hours a day every day using a calculator non-stop. Even if he takes a 30-minute lunch break and spends more than 10 minutes from the rest of the day not calculating, it will take him more than his entire working life to calculate a similarity matrix. If we give the same task to multiple humans, they can perform the task more quickly, but because these people have biases, these biases will be reflected in document encodings, so the document encodings will not be as precise as if one person does it. In the case of a daily brief, if it takes 83,000 hours to complete it and X does 7 hours of work a day (still optimistic), it would takeAttorney Docket No.: P10445PC00 Patent almost 12,000 people to finish working on this in time. For report writing, even though it does not have to be produced immediately, as it is used to assist decision making, a period of over 1 month would likely be deemed too long, so about 400 people would still be needed to complete this task in feasible time.
[0057] FIG. 2 is a flowchart of a process 200 for embedding content from documents as ML- document vectors for filtering and / or visualization, according to one example embodiment. In various embodiments, the information platform 103 may perform one or more portions of the process 200 and may be implemented in, for instance, a chip set including a processor and a memory as shown in FIG. 11 , hardware as shown in FIG. 10, or in circuitry, hardware, firmware, software, or in any combination thereof. As such, information platform 103 can provide means for accomplishing various parts of the process 200, as well as means for accomplishing embodiments of other processes described herein in conjunction with other components of the system 100. Although the process 200 is illustrated and described as a sequence of steps, it is contemplated that various embodiments of the process 200 may be performed in any order or combination and need not include all of the illustrated steps.
[0058] In step 201, the information platform 103 embeds content of a plurality of documents respectively as a plurality of document vectors. In one embodiment, the documents can be aggregated over a designated period of time (e.g., daily, weekly, monthly, etc.) from document sources 101 (e.g., services platform 105, services 107, content providers 109, etc.). To generate embeddings that represent the content of these documents, the system can pre-process each document to clean and normalize the text. This involves removing any irrelevant symbols, punctuation, and stop words that do not contribute to the semantic meaning of the text. Next, the text is tokenized, splitting it into smaller units like words, subwords, punctuation, etc. Once pre- processed and tokenized, the text is then transformed into numerical vectors. This transformation can utilize various ML techniques, such as but not limited to word embeddings (e.g., Word2Vec, GloVe), contextual embeddings using transformer-based models (e.g., LLMs such as BERT, GPT, etc.). Word embeddings map each word in the text to a fixed-size vector in a continuous vector space where semantically similar words are close to each other.
[0059] In the case of transformer-based models such as LLMs, the process involves encoding the entire document into a high-dimensional vector (e.g., typically 12,000 or more dimensions) byAttorney Docket No.: P10445PC00 Patent capturing the contextual relationships between all tokens in the document. These models use attention mechanisms to weigh the importance of each token relative to others, thus producing a comprehensive representation that encapsulates the meaning of the document as a whole. The resultant document vectors, now in a numerical format, can be further processed to reduce dimensionality, if necessary, for computational efficiency and storage optimization. Techniques such as Principal Component Analysis (PCA) or t-Distributed Stochastic Neighbor Embedding (t- SNE) may be employed to achieve this dimensionality reduction while preserving the core semantic relationships.
[0060] It is contemplated that the documents to be processed for embedding can be of any format or type. Depending on the format or type, the information platform 103 can apply different document-type-handling and embedding processes. Examples of different document types include but are not limited to: (a) text documents; (b) audio documents; (c) image documents; and (d) mixed media documents.
[0061] It is noted that the embedding process described above is provided by way of illustration and not as a limitation. It is contemplated that any embedding process to generate embeddings representative of an entire document can be used.
[0062] In step 203, the information platform 103 calculates a similarity value between each pair of document vectors of the plurality of document vectors to generate a similarity matrix. In other words, the information platform 103 calculates the similarity between all pairs of texts / documents using any similarity metric known in the art (e.g., Jaccard similarity, cosine similarity, etc.) between document vectors. The end result is a symmetric matrix, with values between 0 and 1.
[0063] By way of example, to calculate the similarity between document vectors, the information platform 103 employs various similarity metrics. One common method is cosine similarity, which measures the cosine of the angle between two vectors in a multi-dimensional space. This metric is invariant to the magnitude of the vectors, focusing instead on their orientation. Given two document vectors, A and B, the cosine similarity cos(0) is computed as follows:Attorney Docket No.: P10445PC00 Patent
[0064] Here, A ■ B represents the dot product of the two vectors, and ||A|| and ||B|| denote the magnitudes of A and B , respectively. The resulting value ranges from -1 to 1, with 1 indicating perfect similarity, 0 indicating no similarity, and -1 indicating perfect dissimilarity.
[0065] Another approach is the Jaccard similarity. The Jaccard similarity measures the size of — > — > the intersection divided by the size of the union of two sets. For document vectors A and B, the Jaccard similarity J can be computed as:
[0066] In this context, |A n B| represents the number of common elements between the two vectors A and B, and | / 1 U B| represents the total number of unique elements in both vectors A and B.
[0067] Additionally, techniques like Euclidean distance can be used to calculate the similarity by computing the straight-line distance between two vectors in multi-dimensional space. The - > — > formula for Euclidean distance d between two vectors A and B is given by:
[0068] Where n is the number of dimensions, and Atand B;are the components of the vectors A and B, respectively.
[0069] By leveraging these similarity metrics, the information platform 103 can effectively measure the degree of similarity between document vectors, enabling the construction of a similarity matrix that is used in subsequent steps. A similarity matrix is a structured representation used to measure and compare the degree of similarity between each pair of document vectors in the collection of documents being processed. It is constructed using various similarity metrics such as cosine similarity, Jaccard similarity, etc. as discussed above, which quantify the relationship between pairs of document vectors. Each element in the matrix corresponds to the similarity score between two documents, with values typically ranging from 0 to 1, where 1 indicates perfect similarity, 0 indicates perfect dissimilarity.Attorney Docket No.: P10445PC00 Patent
[0070] In one embodiment, the similarity matrix is a symmetric matrix. A symmetric matrix is a square matrix that is identical to its transpose, meaning that the element in the z-th row and j- th column is equal to the element in the / -th row and z-th column for all z and j. This property ensures that the matrix is symmetrical along its main diagonal, which runs from the top-left to the bottom-right corner.
[0071] FIG. 3 is a diagram of a similarity matrix 300, according to one example embodiment. As shown, similarity matrix 300 is a square matrix spanning documents Di to DN and represents all possible pairs of documents between Di to DN. The diagonal represents pairs of each document with itself with each element along the diagonal generally equal to 1 (indicating that each document is perfectly similar to itself). The similarity values Sy(e.g., computed based on a selected similarity metric such as cosine similarity) are indicated in each element other than the diagonal elements and represents the computed similarity values for each corresponding pair where z represents the document number of the first document in the pair, and j represents the document number of the second document in the pair.
[0072] In step 205, the information platform 103 optionally groups the plurality of documents into a plurality of groups based on the similarity matrix. It noted that the process of grouping the documents based on the similarity matrix, as outlined in step 205, is an optional step within the information platform 103. While grouping can enhance the filtering and visualization of document groups, providing insightful patterns and relationships among the documents, it is not a mandatory requirement for the functioning of the system. The decision to implement this step depends on the specific goals and needs of the analysis being conducted. This flexibility allows users to tailor the process according to their unique requirements and the nature of the data they are working with.
[0073] FIG. 4 is a flowchart of a process 400 for grouping documents for filtering and / or visualization, according to one example embodiment. In various embodiments, the information platform 103 may perform one or more portions of the process 400 and may be implemented in, for instance, a chip set including a processor and a memory as shown in FIG. 11, hardware as shown in FIG. 10, or in circuitry, hardware, firmware, software, or in any combination thereof. As such, information platform 103 can provide means for accomplishing various parts of the process 400, as well as means for accomplishing embodiments of other processes described herein in conjunction with other components of the system 100. Although the process 400 is illustrated andAttorney Docket No.: P10445PC00 Patent described as a sequence of steps, it is contemplated that various embodiments of the process 400 may be performed in any order or combination and need not include all of the illustrated steps.
[0074] In one embodiment, the process 400 uses the resulting similarity matrix from step 203 of process 200 to group the documents together.
[0075] In step 401, for each text or document in the collection of documents to process, the information platform 103 finds a text or document that is the most similar to it. To find a document that is most similar to another document using the similarity matrix, the following steps are typically performed. First, the information platform 103 identifies the row in the similarity matrix that corresponds to the document in question. Each row in the matrix represents one document, with columns representing its similarity scores with other documents in the collection. Next, the information platform 103 scans across the row to find the highest similarity score. This score indicates the document that is most similar to the document in question. The column index of this highest score will point to the document number that shares the greatest similarity with the target document. For instance, if document Dfs highest similarity score is with document D3, then D3 is the most similar document to Di. When looking for the maximum value of the similarity matrix, diagonal elements of the similarity matrix either can be set to 0 in advance or otherwise ignored.
[0076] In step 403, the information platform 103 checks whether the maximum similarity value for the document is below the threshold value a. In step 405, a document whose maximum value of similarity with other documents is smaller than the threshold value a is assigned to its own group in which it is the only element. In other words, for each document of the plurality of documents, the information platform 103 determines a maximum similarity value between a plurality of pairs comprising said each document and other documents of the plurality documents. The information platform 103 then assigns said each document as a sole element of its own group based on determining the that the maximum similarity value is below a threshold value. These steps are repeated until each document in document collection is processed.
[0077] In step 407, from the remaining documents (e.g., texts or documents not assigned to a group in which they are the sole element), the information platform 103 picks a document Di at random, and assigns it to the first group (Di to Gi). In step 409, the information platform 103 finds the document D2 that has the highest value of similarity with Di and assigns it to Gi. In step 411, the information platform 103 finds the document D that has the highest value of similarity withAttorney Docket No.: P10445PC00 PatentD2. In step 413, the information platform 103 determines whether D = Di. In step 415, if D = Di, the information platform 103 picks another text at random and calls it D3. Else, in step 417, the information platform 103 assigns D as D3. In step 419, the information platform 103 determines whether D3 is the most similar to any document in Gi (e.g., Di or D2). In step 421, if D3 is the most similar to a text in Gi (e g., Di or D2), the information platform 103 assigns D3 to Gi. Else, in step 423, the information platform 103 assigns D3 to G2, along with the D4, the text to which D is the most similar. In step 425, the information platform 103 repeats the previous steps for the remaining articles, starting from D4. If a document D that is being assigned has already been assigned to an existing group, all the documents from the group to which document D already belongs are assigned to the same group as well.
[0078] Returning to process 200 following the optional grouping step, in step 207, the information platform 103 filters the plurality of documents based on the similarity matrix. Filtering documents based on the similarity matrix involves the following steps.
[0079] (a) Threshold Setting: the information platform 103 sets a predetermined threshold value which dictates the minimum similarity score for documents to be considered relevant. The minimum similarity score helps to determine the relevance of documents within a collection. By setting a predetermined threshold value for similarity, the system can discern which documents are pertinent to the analysis and which are not. The threshold value acts as a benchmark for filtering; documents with similarity scores below this value are considered irrelevant and are thus excluded from further processing. This ensures that only those documents that exhibit a predetermined degree of similarity — e.g., indicative of shared themes, topics, or content — are retained for subsequent steps. Such a filtering mechanism not only streamlines the dataset but also enhances the quality and focus of the analysis by concentrating on documents that meet the established relevance criteria. This can provide technical benefits such as reducing the problems of hallucinations, omissions, etc. associated with LLMs. Thus, in one embodiment, the filtering of the plurality of documents generates a set of filtered documents with a similarity greater than a predetermined threshold value.
[0080] (b) Document Comparison: the information platform 103 compares each document's similarity scores with all other documents. Retain documents with similarity scores above theAttorney Docket No.: P10445PC00 Patent threshold and discard those below it. This achieves the same benefits as described with respect to threshold setting above.
[0081] (c) Token Reduction: In one embodiment, the number of tokens in the plurality of documents is greater than a context window of a LLM used for processing the plurality of documents. A context window in a LLM refers to the maximum amount of text that the model can process and consider at one time. It encompasses a fixed number of tokens, where a token might represent a word, part of a word, or a punctuation mark, depending on the tokenization method used by the model. The context window limits the span of text that the model can use to generate predictions or understand context, which directly impacts its performance in tasks such as text generation, translation, and summarization. When the number of tokens in a document exceeds the context window, the model cannot take into account the entire content, potentially leading to loss of coherence or context in its output. The filtering of the plurality of documents then reduces the number of tokens in the plurality of documents to below the context window of the LLM. In other words, if the number of tokens in the document collection exceeds the context window of a LLM, filter out documents to reduce the number of tokens below this limit. This provides the technical benefit of avoiding loss of coherence or context in an LLM.
[0082] In another embodiment, the filtering of the plurality of documents reduces the number of tokens in the plurality of documents to below a predetermined threshold value. In some cases, LLMs have performance drops when approaching the limits of their context windows even when context window is not exceeded. In other cases, costs (e.g., monetary and / or resource costs) can increase with larger context windows or when approaching the limits of the context window. Accordingly, to address this issue the information platform 103 can set a predetermined threshold value for the number tokens for reduction that is lower or otherwise different from the full context window of the LLM.
[0083] (d) Time-Based Filtering: Optionally, the information platform 103 can filter documents based on one or more time criteria, such as retaining only the oldest documents or those meeting specific time-related conditions. Filtering based on time criteria, such as retaining only the oldest documents, offers several technical benefits. Firstly, it allows for a historical analysis of data, which can be crucial for understanding the genesis and evolution of ideas, concepts, and narratives over time. By focusing on the earliest records, researchers and analysts can trace theAttorney Docket No.: P10445PC00 Patent development of themes from their inception, providing a clear chronological perspective. Another technical advantage is the reduction in computational load. Visualizing and analyzing large networks of documents can be computationally intensive and challenging to interpret. By limiting the dataset to older documents, the information platform 103 reduces the complexity of the visualization process, making it more manageable and easier to derive meaningful insights. This approach not only saves computational resources but also enhances the clarity and interpretability of the analysis.
[0084] Moreover, time-based filtering can help mitigate the issue of redundancy. Newer documents often reiterate or build upon the information contained in older ones. By focusing on the oldest documents, one can eliminate repetitive content, ensuring that the analysis remains concise and to the point. Ultimately, such a filtering mechanism ensures that the document collection is reduced to a more relevant and manageable subset, thereby enhancing the efficiency and effectiveness of subsequent analytical processes.
[0085] For example, the information platform 103 can time order the remaining documents. In some embodiments, the information platform 103 can keep the N-oldest ones (where N is a number that can be configured based on specific use cases). The oldest documents are kept in use cases where there is interest in analyzing the early development or emergence of ideas, concepts, topics, stories, etc. in the documents.
[0086] In one embodiment, the filtering includes removing the documents from the documents database 115 and thus from further processing. In addition or alternatively, the filtering can include removing the corresponding document entries the similarity matrix. This advantageously reduces the size of the similarity matrix to achieve a corresponding speed increase in determining similarity values for documents by excluding documents that are not likely to be relevant to specific use cases.
[0087] (e) Group-Based Filtering: If documents are grouped, the information platform 103 can filter based on groups of documents instead of individual documents. For example, the information platform 103 can rank the groups based on the number of elements they contain and retain only the top-ranked groups. In other words, in embodiments where the documents are grouped, the information platform 103 filters the plurality of groups based on a number of elements in each group of the plurality of groups. For example, the information platform 103 ranks the plurality ofAttorney Docket No.: P10445PC00 Patent groups based on a number of elements in each group of the plurality of groups, and then filters the plurality of groups based on the ranking (i.e., rank groups based on the number of documents they contain and remove groups with fewer documents).
[0088] In one embodiment, the information platform 103 can filter or remove all by the largest group or largest N number of groups (e.g., with respect to number of documents in each group). N can be set according to specific use cases if more than the largest group is to be retained. If there are two or multiple groups of the same size, the information platform 103 can keep the one with the more total characters or use any other tie-breaking criteria.
[0089] It is noted that the filtering embodiments discussed above are provided by way of illustration and not as limitations. It is contemplated that any other equivalent filtering process or mechanism can be used.
[0090] In one embodiment, the information platform 103 can also perform intra-group filtering. The information platform 103 can rank the documents in each group based, for instance, on the average value of the similarity with the other texts in that group, or any other equivalent criteria. This intra-group filtering step ensures that only the more important documents from each group remain.
[0091] These filtering steps ensure that the document collection is reduced to a more relevant and manageable subset, enhancing the efficiency and effectiveness of subsequent analysis. For example, visualization of large networks is computationally costly and more difficult to interpret. So, filtering enables the system to advantageously reduce the computational costs and complexity of such visualizations.
[0092] In step 209, the information platform 103 optionally filters information within each remaining text or document. In other words, the information platform 103 filters at a portion of the content of the plurality of documents based on one or more filtering criteria. It is contemplated that any filtering criteria can be used including but not limited to those discussed herein with respect to filtering groups or documents within groups. In one embodiment, to filter the information within each document, the information platform 103 can be a variant of an extractivetext-summarization method called TextRank. TextRank, for instance, is a graph-based ranking model for text processing including keyword and sentence extraction. More specifically, TextRank uses a graph-based ranking algorithm to determine the importance of a vertex (word or sentence)Attorney Docket No.: P10445PC00 Patent within a graph by considering global information from the entire graph. The algorithm identifies keywords by constructing a graph where vertices represent words, and edges represent cooccurrence relations within a fixed window of words. The importance of each word is determined by the TextRank score. For sentence extraction, the algorithm constructs a graph where vertices represent sentences, and edges represent similarity relations based on content overlap. The most important sentences are selected based on their TextRank scores. In one embodiment, the information platform 103 alters the TextRank algorithm by first embedding each sentence as a vector and then using similarity between the sentences as vertex weights in algorithm. It is contemplated that any other extractive algorithm (e.g., LLM-extractive summarization by prompting a LLM and returning key sentences) can be equivalently used in the various embodiments described herein. In addition, in embodiments where only LLM-extractive summarization is mentioned, it is contemplated alternate methods such as TextRank or the modified version described above can be used equivalently.
[0093] This additional filtering and the filtering previously discussed can be used to advantageously minimize LLM cost, hallucinations, and omissions.
[0094] In addition or as an alternative to filtering steps 207-209, the information platform 103 generates a visual representation of a plurality of semantic relationships among the plurality of documents based on the similarity matrix (step 211). For example, the information platform 103 can use the remaining values of the similarity matrix (e.g., remaining after filtering) and / or the properties or labels (e g., author / publisher, location, time, etc.) associated with the documents to produce visualizations. In other words, the plurality of documents are respectively associated with one or more labels. Then, the visual representation, the plurality of semantic relationships (e.g., depicted in the visual representations), or a combination thereof are based at least in part on the one or more labels.
[0095] Based on the labels, the information platform 103 can use the remaining similarity matrix to visualize relationships between documents and any of the labels (e.g., publishers, locations, time, etc. There are different ways in which the system 100 can present these semantic relationships, depending on which of the properties or labels represent the nodes (e.g., publishers, locations, etc.).Attorney Docket No.: P10445PC00 Patent
[0096] Since the entries of the similarity matrix represent the semantic similarity between the remaining documents, they are used to construct the edges of the directed graph visual representation. By way of example, there are three ways in which the information platform 103 can use the remaining similarity matrix to visualize the semantic relationships between the documents, publishers, and locations:
[0097] (1) Make no further modifications to the similarity matrix. In other words, the information platform 103 uses the similarity matrix as it exists after the previous processing and does nothing further to change the matrix. With this unmodified matrix, the information platform 103 can provide the following analysis / visualizations: a. Document level - Detection / visualization of most important documents. The “most important documents”, for instance, are those documents that have the most similarity with other remaining documents indicating that their contents are the most semantically prevalent in the remaining documents / similarity matrix. b. Publisher level - Detection / visualization of the most important publisher(s). The “most important publisher(s)”, for instance, are the publishers labeled as otherwise associated the “most important documents”. c. Location level - Detection / visualization of the most important locations. The “most important locations(s)”, for instance, are the published labeled as otherwise associated the “most important documents”.
[0098] (2) Modify the remaining portion of the similarity matrix, where for each document, only the strongest older connections are kept. It is noted that depending on the use case any other connections in time can be kept (e.g., strongest newest connections, strongest connections for a set period of time, etc.). This modification enables the information platform 103 to consider temporal parameters (e.g., via a chronology) to track the development of documents, publishers, and / or locations over time. With a modification based on keeping only the strongest oldest connections, the information platform 103 can provide the following analysis / visualizations: a. Document level - Detection / visualization of the original source of a document, and tracing of the development of the story / idea / concept in the document over time. b. Publisher level - Detection / visualization of the source publisher, and tracing the original or source publisher of a document, story, idea, concept, etc.Attorney Docket No.: P10445PC00 Patent c. Location level - Detection / visualization of the source location, and tracing the original or source location of a document, story, idea, concept, etc.
[0099] (3) Modify the remaining similarity matrix, where for each document, only the strongest connection is kept. With this modification, the information platform 103 can provide the following analysis / visualizations: a. Document level - sometimes, a story / idea / concept will have multiple sources that operate independently or are barely loosely related. This visualization allows checking for such occurrences. It also accomplishes an almost exact purpose as visualizing all the connections, but it is easier to observe, as fewer connections are present. b. Publisher level - Visualizes which publishers are the most similar to each other. This is an indication of publisher collaboration or copying. c. Location level - Visualizes which locations are the most similar to each other. This is an indication of publisher collaboration or copying.
[0100] In summary, in one embodiment, the one or more labels includes a publication time label, and wherein the plurality of semantic relationships are determined as a function of publication time. In another embodiment, the one or more labels includes a publisher label, and wherein the plurality of semantic relationships are determined as a function of publisher. In this embodiment, the visual representation indicates a similarity of one or more publishers of the plurality of documents based on the plurality of semantic relationships determined as the function of publisher. In yet another embodiment, the one or more labels includes a location label, and wherein the plurality of semantic relationships are determined as a function of location. In this embodiment, the visual representation indicates a similarity of one or more locations of the plurality of documents based on the plurality of semantic relationships determined as the function of location.
[0101] In one embodiment, the visual representation is based at least in part on a directed graph, wherein a node of the directed graph represents a document of the plurality of documents, and wherein an edge of the directed graph represents at least one relationship of the plurality of semantic relationships between two adjacent nodes of the directed graph.Attorney Docket No.: P10445PC00 Patent
[0102] In this way, the end user can visualize the relationships between disparate documents and make sense of them on the larger scale.
[0103] In step 213, a product of the various embodiments described herein is the output data 117 comprising, for instance filtered documents 119 (and / or filtered similarity matrix associated with the documents) and / or visualization data 121). In this way, the information platform 103 provides the filtered plurality of documents, the visual representations of the documents, and / or related data as an output.
[0104] By way of example, the output data 117 can be further processed using the LLM 113 to generate briefs 123 (e.g., periodic such as daily summaries of aggregated documents), reports 125 (e.g., on demand requests for data based on specific user queries / requests), and / or the like. For example, in step 215, the information platform 103 provides the remaining documents (e.g., as a set of filtered or unfiltered documents) as an input to a LLM to generate at least one automated report.
[0105] In one embodiment, the information platform 103 enables the end user to search for terms of interest as a function of time and location (categorical). Following this, the end user can visualize the similarities between the remaining articles. Visualization can be done on three levels and have three types, as described above. Edges in the document view and the publisher or location view are represented differently. In the document view, edges are thicker when the corresponding value of the similarity matrix is larger, and longer when the difference of publication times between the documents is larger. In publisher or location views, edges are thicker when the number of entries in the similarity matrix that correspond to the same publisher / location and are above the similarity threshold is larger. Edges are directed, pointing from older to newer documents. In the publisher / location view, each node can refer to itself. Any of the visualizations can be saved / exported and a report on the findings written up.
[0106] Additional description of use cases for generating briefs, reports, and investigations are provided below.
[0107] FIG. 5 is a flowchart of a process 500 for generating a brief 123, according to one example embodiment. In various embodiments, the information platform 103 may perform one or more portions of the process 500 and may be implemented in, for instance, a chip set including a processor and a memory as shown in FIG. 11, hardware as shown in FIG. 10, or in circuitry,Attorney Docket No.: P10445PC00 Patent hardware, firmware, software, or in any combination thereof As such, information platform 103 can provide means for accomplishing various parts of the process 500, as well as means for accomplishing embodiments of other processes described herein in conjunction with other components of the system 100. Although the process 500 is illustrated and described as a sequence of steps, it is contemplated that various embodiments of the process 500 may be performed in any order or combination and need not include all of the illustrated steps.
[0108] In one embodiment, the brief 123 summarizes all the key time-indexed texts from the past designated time period (e.g., 24 hours for a daily brief, 7 days for a weekly brief, etc.), either from around the world or from specific locations. The process 500 uses the steps as previously described in process 200.
[0109] In step 501, the information platform 103 queries a database (e.g., the documents database 115) for a collection of documents based on publication time, location, and / or any other search term to gather documents for designated time period.
[0110] In step 503, the information platform 103 calculates the similarities between the documents retrieved by the query of step 501 according to the various embodiments described with respect to process 200. For example, the calculation comprises embedding the contents of the documents as respective document vectors and calculating similarities between all pairs of the documents (e.g., using Jaccard similarity, cosine similarity, etc.). A similarity matrix is created to provide the similarities between each pair of document vectors.[0U1] In step 505, the information platform 103 organizes the documents into groups based on the calculated similarities as described in processes 200 and 400. In one embodiment, the grouping enables more efficient processing, filtering, and / or visualization of large collections of documents.
[0112] In 507, the information platform 103 performs group filtering to remove less important or relevant groups of documents. For example, the information platform can filter or remove all but the largest group or largest N number of groups. N can be set according to specific use cases if more than the largest group is to be retained.
[0113] In step 509, the information platform 103 can filter documents within each group to also remove less important or relevant documents from the remaining groups. For example, the information platform can rank the documents in each group based, for instance, on the averageAttorney Docket No.: P10445PC00 Patent value of the similarity with the other texts in that group, or any other equivalent criteria. This intragroup filtering step ensures that only the more important documents from each group remain.
[0114] In step 511, the information platform 103 generates briefs (e.g., LLM-extractive summaries) of the remaining documents (e.g., using the LLM 113). The process involves leveraging the advanced capabilities of a LLM to create concise briefs of extensive documents. Initially, the remaining documents are fed into the LLM, which identifies and extracts the most pertinent information, effectively summarizing each document into a shorter brief or report. This approach ensures that despite the reduction in length, the core message and vital details of the original documents are preserved.
[0115] In step 513, the information platform 103 generates briefs for entire groups that remain after filtering using the LLM 113, for instance, based on the briefs of each document within each group. The briefs for each document in the group is fed to the LLM 113 to produce a cohesive summary that captures the main ideas and critical data points from the collective group.
[0116] In step 515, the information platform 103 combines the briefs of each group into the brief 123.
[0117] FIG. 6 is a diagram of example user interfaces (UIs) 601-605 for generating and presenting a brief, according to one example embodiment. As noted, the brief 123 (e.g., daily brief) summarizes all the key time-indexed texts from the past 24 hours (or any other designated time window), either all around the world or from specific locations. Based on the process 500, once distinct texts are optionally filtered by location, the system automatically calculates all the similarities between the texts, then organizes them into groups, removes less important groups and less important texts from the groups. The remaining text is smaller than the maximum context length for the specific LLM used. Following this, the remaining documents from each group are passed to a LLM, with short reports returned as an output. Short reports are combined, and the final text is returned to the end user.
[0118] As shown, the main page UI 601 lists all the briefs by title. The end user can click on each brief or add a new one by clicking on the “Add brief’ button. If the user clicks on a brief, the page specific brief UI 603 lists all the past briefs and sorts them out by date. The user can click on a date to read a specific brief. If the user decides to add a brief, they are taken to the brief add UI 605, where they can name the brief and select the locations they are interested in.Attorney Docket No.: P10445PC00 Patent
[0119] FIG. 7 is a flowchart of a process 700 for generating a report, according to one example embodiment. In various embodiments, the information platform 103 may perform one or more portions of the process 700 and may be implemented in, for instance, a chip set including a processor and a memory as shown in FIG. 11, hardware as shown in FIG. 10, or in circuitry, hardware, firmware, software, or in any combination thereof. As such, information platform 103 can provide means for accomplishing various parts of the process 700, as well as means for accomplishing embodiments of other processes described herein in conjunction with other components of the system 100. Although the process 700 is illustrated and described as a sequence of steps, it is contemplated that various embodiments of the process 700 may be performed in any order or combination and need not include all of the illustrated steps.
[0120] In step 701, the end user can query for specific terms (or use semantic search) by time and location. Documents are returned and similarities between them are calculated (step 703), following which they are grouped together (step 705) and all but the largest group is removed (step 707). Within the remaining group, the most important documents are kept. Remaining texts are shorter than the maximum context length of the language model used. The end user can now query a LLM to look for any information contained in the remaining texts, or generate a report based on them (step 709).
[0121] FIG. 8 is a diagram of example user interfaces 801-807 for generating and presenting a report, according to one example embodiment. As shown, the main page UI 801 shows the report title, the report writeup, and a list of queries, which the user can click on. The user can also choose to edit the report, archive it, or add a query. The add query UI 803 allows the user to search for terms during a specific time period and in specific locations, and add it. In addition, the user can click on the query in the main page UI 801 and then edit or delete it. The report can be edited by clicking on the edit report button on the main page UI 801 which will take the user to the edit report UI 805. After editing in the edit report UI 805, the user can click on the write button to commit the changes to the report. To use an LLM, the user can click on the generate button on the edit report page 805 and query the LLM 113. The returning output from the LLM can be copied and used in another query or taken to the report and edited.
[0122] In one embodiment, the main page UI 801 also includes an option to visualize the similarities between documents as a function of various contextual parameters such as publisherAttorney Docket No.: P10445PC00 Patent and location as illustrated. The user can click on the visualize option to open visualization UI 807 which presents tabs for viewing relationships between documents, relationships as a function of publisher, and relationships as a function of location. In some embodiments, the relationships can be viewed as a function of any other contextual parameter such as but not limited to time. In this example, the visualization of the semantic relationships is in the form of a graph (e.g., a directed graph) where documents or clusters / groups of documents are depicted as nodes (e.g., indicated by a circle) with edges between the nodes an indication of the semantic relationships between nodes. The properties of the edges (e.g., line weights, length, etc.) can be used to indicate properties of the semantic of relationship such as but not limited to strength of the relationship, magnitude of similarity values, etc. As shown, the visualization UI 807 also includes options for selecting or filtering connections (e.g., semantic relationships) for visualization. The options include: (1) an option to visualize all connections; (2) an option to visualize only the strongest older connections (e g., top N connections older than a threshold time); and (3) an option to visualize only the strongest connections for all connections (e.g., top N connections with no time / age constraint). These temporal options enable the use to detect and / or visualize the development of the documents over time such determining original documents, original publishers, and / or original locations from which documents of certain topics developed.
[0123] FIG. 9 is a diagram of example user interfaces 901-903 for generating and presenting an investigation, according to one example embodiment. The example of FIG. 9 supplements the report of FIG. 8 to enable investigations of time separated or time ordered events. Time separated or time ordered events refer to occurrences that are organized or analyzed based on their chronological sequence. These events are often tracked to understand their progression over time and to identify patterns or causality. In one embodiment, this involves pulling textual data from databases, querying for key terms, and extracting events to write reports with the help of LLM 113. The LLM queries can be performed with reference to timeline events, comments, and the writeup to be useful.
[0124] Accordingly, as shown, main page UI 901 presents the investigation title along with queries, timeline events, and comments. The main page UI 901 includes the functionality of UIs 801-807 and adds the options to select, edit, and / or add timeline events and comments. A timeline event can be any event associated with a date / time to facilitate time ordering or generating aAttorney Docket No.: P10445PC00 Patent chronology. The user can click on a timeline event to edit the comment / timeline event, or click on the add event button on the main page UI 901 to access the timeline UI 903. The timeline UI 903 provides a description and time of the selected timeline event and options to add, edit, or delete events. The main page UI 901 also includes a listing of comments and an add comment button. Comments can be any notation that a user wants to add as part of the investigation report. The comment UI 905 is presented when a user clicks on a comment to view or the add comment button to add a new comment. The comment UI 905 displays the selected comment along with options to add, edit, or delete comments.
[0125] Returning to FIG. 1, as shown and discussed above, the system 100 includes the information platform 103 for embedding content from documents as ML-document vectors for filtering and / or visualization. In one embodiment, the information platform 103 has connectivity or access to one or more databases for storing the documents to be processed (e.g., documents database 115) according to the various embodiments described herein. In one embodiment, the information platform 103 has connectivity over a communication network 111 to the services platform 105 that provides one or more services 107 (e.g., data and / or document services). By way of example, the services 107 include, but are not limited to, social -networking services, content (e.g., audio, video, images, etc.) provisioning services, application services, storage services, contextual information determination services, location-based services, information-based services (e.g., weather, news, etc.), etc.
[0126] In one embodiment, the information platform 103 may be a platform with multiple interconnected components. The information platform 103 may also include multiple servers, intelligent-networking devices, computing devices, components, and corresponding software. In addition, it is noted that the information platform 103 may be a separate entity of the system 100, a part of the one or more services 107, a part of the services platform 105, or content providers 109.
[0127] In one embodiment, content providers 109 may provide content or data (e.g., texts, audiovisual data, documents, etc.) to the information platform 103 and / or document sources 101. The content provided may be any type of content, such as textual content, audio content, video content, image content, etc. In one embodiment, the content providers 109 may also store content associated with the information platform 103. In another embodiment, the content providers 109Attorney Docket No.: P10445PC00 Patent may manage access to a central repository of data, and offer a consistent, standard interface to data, such as the documents database 115.
[0128] In another embodiment, the communication network 111 of system 100 includes one or more networks such as a data network, a wireless network, a telephony network, or any combination thereof. It is contemplated that the data network may be any local area network (LAN), metropolitan area network (MAN), wide area network (WAN), a public data network (e.g., the Internet), short range wireless network, or any other suitable packet-switched network, such as a commercially owned, proprietary packet-switched network, e.g., a proprietary cable or fiberoptic network, and the like, or any combination thereof. In addition, the wireless network may be, for example, a cellular network and may employ various technologies including enhanced data rates for global evolution (EDGE), general packet radio service (GPRS), global system for mobile communications (GSM), Internet protocol multimedia subsystem (IMS), universal mobile telecommunications system (UMTS), etc., as well as any other suitable wireless medium, e.g., worldwide interoperability for microwave access (WiMAX), 5G New Radio Networks, Long Term Evolution (LTE) networks, code division multiple access (CDMA), wideband code division multiple access (WCDMA), wireless fidelity (Wi-Fi), wireless LAN (WLAN), Bluetooth®, Internet Protocol (IP) data casting, satellite, mobile ad-hoc network (MANET), and the like, or any combination thereof.
[0129] By way of example, the mapping platform 103, document sources 101, services platform 105, services 107, and / or content providers 109 optionally communicate with each other and other components of the system 100 using well known, new or still developing protocols. In this context, a protocol includes a set of rules defining how the network nodes within the communication network 111 interact with each other based on information sent over the communication links. The protocols are effective at different layers of operation within each node, from generating and receiving physical signals of various types, to selecting a link for transferring those signals, to the format of information indicated by those signals, to identifying which software application executing on a computer system sends or receives the information. The conceptually different layers of protocols for exchanging information over a network are described in the Open Systems Interconnection (OSI) Reference Model.Attorney Docket No.: P10445PC00 Patent
[0130] Communications between the network nodes are typically effected by exchanging discrete packets of data. Each packet typically comprises (1) header information associated with a particular protocol, and (2) payload information that follows the header information and contains information that may be processed independently of that particular protocol. In some protocols, the packet includes (3) trailer information following the payload and indicating the end of the payload information. The header includes information such as the source of the packet, its destination, the length of the payload, and other properties used by the protocol. Often, the data in the payload for the particular protocol includes a header and payload for a different protocol associated with a different, higher layer of the OSI Reference Model. The header for a particular protocol typically indicates a type for the next protocol contained in its payload. The higher layer protocol is said to be encapsulated in the lower layer protocol. The headers included in a packet traversing multiple heterogeneous networks, such as the Internet, typically include a physical (layer 1) header, a datalink (layer 2) header, an internetwork (layer 3) header, a transport (layer 4) header, and various application (layer 5, layer 6, and layer 7) headers as defined by the OSI Reference Model.
[0131] The processes described herein for embedding content from documents as ML- document vectors for filtering and / or visualization may be advantageously implemented via software, hardware (e.g., general processor, Digital Signal Processing (DSP) chip, an Application Specific Integrated Circuit (ASIC), Field Programmable Gate Arrays (FPGAs), etc.), firmware or a combination thereof. Such exemplary hardware for performing the described functions is detailed below.
[0132] Additionally, as used herein, the term ‘circuitry’ may refer to (a) hardware-only circuit implementations (for example, implementations in analog circuitry and / or digital circuitry); (b) combinations of circuits and computer program product(s) comprising software and / or firmware instructions stored on one or more computer readable memories that work together to cause an apparatus to perform one or more functions described herein; and (c) circuits, such as, for example, a microprocessor s) or a portion of a microprocessor s), that require software or firmware for operation even if the software or firmware is not physically present. This definition of ‘circuitry’ applies to all uses of this term herein, including in any claims. As a further example, as used herein, the term ‘circuitry’ also includes an implementation comprising one or more processors and / orAttorney Docket No.: P10445PC00 Patent portion(s) thereof and accompanying software and / or firmware. As another example, the term ‘circuitry’ as used herein also includes, for example, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular device, other network device, and / or other computing device.
[0133] FIG. 10 illustrates a computer system 1000 upon which an embodiment of the invention may be implemented. Computer system 1000 is programmed (e.g., via computer program code or instructions) to embed content from documents as ML-document vectors for filtering and / or visualization as described herein and includes a communication mechanism such as a bus 1010 for passing information between other internal and external components of the computer system 1000. Information (also called data) is represented as a physical expression of a measurable phenomenon, typically electric voltages, but including, in other embodiments, such phenomena as magnetic, electromagnetic, pressure, chemical, biological, molecular, atomic, sub-atomic, and quantum interactions. For example, north and south magnetic fields, or a zero and non-zero electric voltage, represent two states (0, 1) of a binary digit (bit). Other phenomena can represent digits of a higher base. A superposition of multiple simultaneous quantum states before measurement represents a quantum bit (qubit). A sequence of one or more digits constitutes digital data that is used to represent a number or code for a character. In some embodiments, information called analog data is represented by a near continuum of measurable values within a particular range.
[0134] A bus 1010 includes one or more parallel conductors of information so that information is transferred quickly among devices coupled to the bus 1010. One or more processors 1002 for processing information are coupled with the bus 1010.
[0135] A processor 1002 performs a set of operations on information as specified by computer program code related to embedding content from documents as ML-document vectors for filtering and / or visualization. The computer program code is a set of instructions or statements providing instructions for the operation of the processor and / or the computer system to perform specified functions. The code, for example, may be written in a computer programming language that is compiled into a native instruction set of the processor. The code may also be written directly using the native instruction set (e.g., machine language). The set of operations include bringing information in from the bus 1010 and placing information on the bus 1010. The set of operations also typically include comparing two or more units of information, shifting positions of units ofAttorney Docket No.: P10445PC00 Patent information, and combining two or more units of information, such as by addition or multiplication or logical operations like OR, exclusive OR (XOR), and AND. Each operation of the set of operations that can be performed by the processor is represented to the processor by information called instructions, such as an operation code of one or more digits. A sequence of operations to be executed by the processor 1002, such as a sequence of operation codes, constitute processor instructions, also called computer system instructions or, simply, computer instructions. Processors may be implemented as mechanical, electrical, magnetic, optical, chemical, or quantum components, among others, alone or in combination.
[0136] Computer system 1000 also includes a memory 1004 coupled to bus 1010. The memory 1004, such as a random access memory (RAM) or other dynamic storage device, stores information including processor instructions for embedding content from documents as ML- document vectors for filtering and / or visualization. Dynamic memory allows information stored therein to be changed by the computer system 1000. RAM allows a unit of information stored at a location called a memory address to be stored and retrieved independently of information at neighboring addresses. The memory 1004 is also used by the processor 1002 to store temporary values during execution of processor instructions. The computer system 1000 also includes a read only memory (ROM) 1006 or other static storage device coupled to the bus 1010 for storing static information, including instructions, that is not changed by the computer system 1000. Some memory is composed of volatile storage that loses the information stored thereon when power is lost. Also coupled to bus 1010 is a non-volatile (persistent) storage device 1008, such as a magnetic disk, optical disk, or flash card, for storing information, including instructions, that persists even when the computer system 1000 is turned off or otherwise loses power.
[0137] Information, including instructions for embedding content from documents as ML- document vectors for filtering and / or visualization, is provided to the bus 1010 for use by the processor from an external input device 1012, such as a keyboard containing alphanumeric keys operated by a human user, or a sensor. A sensor detects conditions in its vicinity and transforms those detections into physical expression compatible with the measurable phenomenon used to represent information in computer system 1000. Other external devices coupled to bus 1010, used primarily for interacting with humans, include a display device 1014, such as a cathode ray tube (CRT) or a liquid crystal display (LCD), or plasma screen or printer for presenting text or images,Attorney Docket No.: P10445PC00 Patent and a pointing device 1016, such as a mouse or a trackball or cursor-direction keys, or motion sensor, for controlling a position of a small-cursor image presented on the display 1014 and issuing commands associated with graphical elements presented on the display 1014. In some embodiments, for example, in embodiments in which the computer system 1000 performs all functions automatically without human input, one or more of external input device 1012, display device 1014, and pointing device 1016 is omitted.
[0138] In the illustrated embodiment, special purpose hardware, such as an application specific integrated circuit (ASIC) 1020, is coupled to bus 1010. The special purpose hardware is configured to perform operations not performed by processor 1002 quickly enough for special purposes. Examples of application specific ICs include graphics-accelerator cards for generating images for display 1014, cryptographic boards for encrypting and decrypting messages sent over a network, speech recognition, and interfaces to special external devices, such as robotic arms and medicalscanning equipment that repeatedly perform some complex sequence of operations that are more efficiently implemented in hardware.
[0139] Computer system 1000 also includes one or more instances of a communications interface 1070 coupled to bus 1010. Communication interface 1070 provides a one-way or two- way communication coupling to a variety of external devices that operate with their own processors, such as printers, scanners, and external disks. In general, the coupling is with a network link 1078 that is connected to a local network 1080 to which a variety of external devices with their own processors are connected. For example, communication interface 1070 may be a parallel port or a serial port or a universal serial bus (USB) port on a personal computer. In some embodiments, communications interface 1070 is an integrated services digital network (ISDN) card or a digital subscriber line (DSL) card or a telephone modem that provides an information communication connection to a corresponding type of telephone line. In some embodiments, a communication interface 1070 is a cable modem that converts signals on bus 1010 into signals for a communication connection over a coaxial cable or into optical signals for a communication connection over a fiber optic cable. As another example, communications interface 1070 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN, such as Ethernet. Wireless links may also be implemented. For wireless links, the communications interface 1070 sends or receives or both sends and receives electrical, acoustic,Attorney Docket No.: P10445PC00 Patent or electromagnetic signals, including infrared and optical signals, that carry information streams, such as digital data. For example, in wireless handheld devices, such as mobile telephones like cell phones, the communications interface 1070 includes a radio band electromagnetic transmitter and receiver called a radio transceiver. In certain embodiments, the communications interface 1070 enables connection to the communication network 111 for embedding content from documents as ML-document vectors for filtering and / or visualization.
[0140] The term computer-readable medium is used herein to refer to any medium that participates in providing information to processor 1002, including instructions for execution. Such a medium may take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 1008. Volatile media include, for example, dynamic memory 1004. Transmission media include, for example, coaxial cables, copper wire, fiber optic cables, and carrier waves that travel through space without wires or cables, such as acoustic waves and electromagnetic waves, including radio, optical, and infrared waves. Signals include man-made transient variations in amplitude, frequency, phase, polarization, or other physical properties transmitted through the transmission media. Common forms of computer-readable media include, for example, a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, CDRW, DVD, any other optical medium, punch cards, paper tape, optical mark sheets, any other physical medium with patterns of holes or other optically recognizable indicia, a RAM, a PROM, an EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave, or any other medium from which a computer can read.
[0141] Network link 1078 typically provides information communication using transmission media through one or more networks to other devices that use or process the information. For example, network link 1078 may provide a connection through local network 1080 to a host computer 1082 or to equipment 1084 operated by an Internet Service Provider (ISP). ISP equipment 1084 in turn provides data communication services through the public, world-wide packet-switching communication network of networks now commonly referred to as the Internet 1090.
[0142] A computer called a server host 1092 connected to the Internet hosts a process that provides a service in response to information received over the Internet. For example, server hostAttorney Docket No.: P10445PC00 Patent1092 hosts a process that provides information representing video data for presentation at display 1014. It is contemplated that the components of system can be deployed in various configurations within other computer systems, e.g., host 1082 and server 1092.
[0143] FIG. 11 illustrates a chip set 1100 upon which an embodiment of the invention may be implemented. Chip set 1100 is programmed to embed content from documents as ML-document vectors for filtering and / or visualization as described herein and includes, for instance, the processor and memory components described with respect to FIG. 10 incorporated in one or more physical packages (e.g., chips). By way of example, a physical package includes an arrangement of one or more materials, components, and / or wires on a structural assembly (e.g., a baseboard) to provide one or more characteristics such as physical strength, conservation of size, and / or limitation of electrical interaction. It is contemplated that in certain embodiments the chip set can be implemented in a single chip.
[0144] In one embodiment, the chip set 1100 includes a communication mechanism such as a bus 1101 for passing information among the components of the chip set 1100. A processor 1103 has connectivity to the bus 1101 to execute instructions and process information stored in, for example, a memory 1105. The processor 1103 may include one or more processing cores with each core configured to perform independently. A multi-core processor enables multiprocessing within a single physical package. Examples of a multi-core processor include two, four, eight, or greater numbers of processing cores. Alternatively or in addition, the processor 1103 may include one or more microprocessors configured in tandem via the bus 1101 to enable independent execution of instructions, pipelining, and multithreading. The processor 1103 may also be accompanied with one or more specialized components to perform certain processing functions and tasks such as one or more digital signal processors (DSP) 1107, or one or more applicationspecific integrated circuits (ASIC) 1109. A DSP 1107 typically is configured to process real-world signals (e.g., sound) in real time independently of the processor 1103. Similarly, an ASIC 1109 can be configured to perform specialized functions not easily performed by a general purposed processor. Other specialized components to aid in performing the inventive functions described herein include one or more field programmable gate arrays (FPGA) (not shown), one or more controllers (not shown), or one or more other special-purpose computer chips.Attorney Docket No.: P10445PC00 Patent
[0145] The processor 1103 and accompanying components have connectivity to the memory 1105 via the bus 1101. The memory 1105 includes both dynamic memory (e.g., RAM, magnetic disk, writable optical disk, etc.) and static memory (e.g., ROM, CD-ROM, etc.) for storing executable instructions that when executed perform the inventive steps described herein to embed content from documents as ML-document vectors for filtering and / or visualization. The memory 1105 also stores the data associated with or generated by the execution of the inventive steps.
Claims
Attorney Docket No.: P10445PC00 PatentCLAIMSWHAT IS CLAIMED IS:
1. A method comprising: embedding content of a plurality of documents respectively as a plurality of document vectors; calculating a similarity value between each pair of document vectors of the plurality of document vectors to generate a similarity matrix; and generating a visual representation of a plurality of semantic relationships among the plurality of documents based on the similarity matrix.
2. The method of claim 1, wherein the plurality of documents are respectively associated with one or more labels, and wherein the visual representation, the plurality of semantic relationships, or a combination thereof is based at least in part on the one or more labels.
3. The method of claim 2, wherein the one or more labels includes a publication time label, and wherein the plurality of semantic relationships are determined as a function of publication time.
4. The method of claim 3, wherein the visual representation indicates one or more documents of the plurality of documents associated based on the plurality of semantic relationships determined as the function of publication time.
5. The method according to any of claims 2-4, wherein the one or more labels includes a publisher label, and wherein the plurality of semantic relationships are determined as a function of publisher.Attorney Docket No.: P10445PC00 Patent6. The method of claim 5, wherein the visual representation indicates a similarity of one or more publishers of the plurality of documents based on the plurality of semantic relationships determined as the function of publisher.
7. The method according to any of claims 2-6, wherein the one or more labels includes a location label, and wherein the plurality of semantic relationships are determined as a function of location.
8. The method of claim 7, wherein the visual representation indicates a similarity of one or more locations of the plurality of documents based on the plurality of semantic relationships determined as the function of location.
9. The method according to any of claims 1-8, wherein the visual representation is based at least in part on a directed graph, wherein a node of the directed graph represents a document of the plurality of documents, and wherein an edge of the directed graph represents at least one relationship of the plurality of semantic relationships between two adjacent nodes of the directed graph.
10. The method according to any of claims 1-9, further comprising: filtering the plurality of documents based on the similarity matrix, wherein the visual representation is based on the filtering of the plurality of documents.
11. The method according to any of claims 1-10, further comprising: filtering at least a portion of the content of the plurality of documents based on one or more filtering criteria, wherein the visual representation is based on the filtering of the at least the portion of the content.
12. The method according to any of claims 1-11, further comprising:Attorney Docket No.: P10445PC00 Patent grouping the plurality of documents into a plurality of groups based on the similarity matrix.
13. The method of claim 12, further comprising: filtering the plurality of groups based on a number of elements in each group of the plurality of groups, wherein the visual representation is based on the filtering.
14. The method according to any of claims 12-13, further comprising: ranking the plurality of groups based on a number of elements in each group of the plurality of groups; and filtering the plurality of groups based on the ranking.
15. The method according to any of claims 1-14, further comprising: for each document of the plurality of documents, determining a maximum similarity value between a plurality of pairs comprising said each document and other documents of the plurality documents; and assigning said each document as a sole element of its own group based on determining that the maximum similarity value is below a threshold value.
16. The method according to any of claims 1-15, further comprising: filtering the plurality of documents based on one or more time criteria, wherein the visual representation is based on the filtering.
17. The method according to any of claims 1-16, wherein a number of tokens in the plurality of documents is greater than a context window of a large language model used for processing the plurality of documents.
18. An apparatus comprising: at least one processor; andAttorney Docket No.: P10445PC00 Patent at least one memory including computer program code for one or more programs, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to: embed content of a plurality of documents respectively as a plurality of document vectors; calculate a similarity value between each pair of document vectors of the plurality of document vectors to generate a similarity matrix; filter the plurality of documents based on the similarity matrix; and provide the filtered plurality of documents as an output.
19. The apparatus of claim 18, wherein the apparatus is further caused to: filter at least a portion of the content of the plurality of documents based on one or more filtering criteria, wherein the output is further based on the filtering of the at least the portion of the content.
20. The apparatus of claim 19, wherein a number of tokens in the plurality of documents is greater than a context window of a large language model used for processing the plurality of documents.
21. The apparatus of claim 20, wherein the filtering of the plurality of documents reduces the number of tokens in the plurality of documents to below the context window of the large language model.
22. The apparatus according to any of claims 20-21, wherein the filtering of the plurality of documents reduces the number of tokens in the plurality of documents to below a predetermined threshold value.
23. The apparatus according to any of claims 18-22, wherein the filtering of the plurality of documents generates a set of filtered documents with similarities greater than a predetermined threshold value.Attorney Docket No.: P10445PC00 Patent24. The apparatus of claim 23, wherein the apparatus is further caused to: provide the set of filtered documents as an input to a large language model to generate at least one automated report.
25. The apparatus according to any of claims 18-24, wherein the apparatus is further caused to: group the plurality of documents into a plurality of groups based on the similarity matrix.
26. The apparatus of claim 25, wherein the apparatus is further caused to: filter the plurality of groups based on a number of elements in each group of the plurality of groups.
27. The apparatus according to any of claims 25-26, wherein the apparatus is further caused to: rank the plurality of groups based on a number of elements in each group of the plurality of groups; and filter the plurality of groups based on the ranking.
28. The apparatus according to any of claims 18-27, wherein the apparatus is further caused to: for each document of the plurality of documents, determine a maximum similarity value between a plurality of pairs comprising said each document and other documents of the plurality documents; and assign said each document as a sole element of its own group based on determining that the maximum similarity value is below a threshold value.
29. The apparatus according to any of claims 18-28, wherein the apparatus is further caused to: filter the plurality of documents based on one or more time criteria.Attorney Docket No.: P10445PC00 Patent30. A non-transitory computer-readable storage medium, carrying one or more sequences of one or more instructions which, when executed by one or more processors, cause an apparatus to perform: embedding content of a plurality of documents respectively as a plurality of document vectors; calculating a similarity value between each pair of document vectors of the plurality of document vectors to generate a similarity matrix; and generating a visual representation of a plurality of semantic relationships among the plurality of documents based on the similarity matrix.
31. The non-transitory computer-readable storage medium of claim 30, wherein the plurality of documents are respectively associated with one or more labels, and wherein the visual representation, the plurality of semantic relationships, or a combination thereof is based at least in part on the one or more labels.
32. The non-transitory computer-readable storage medium of claim 31, wherein the one or more labels includes a publication time label, and wherein the plurality of semantic relationships are determined as a function of publication time.
33. The non-transitory computer-readable storage medium of claim 32, wherein the visual representation indicates one or more documents of the plurality of documents associated based on the plurality of semantic relationships determined as the function of publication time.
34. The non-transitory computer-readable storage medium according to any of claims 31- 32, wherein the one or more labels includes a publisher label, and wherein the plurality of semantic relationships are determined as a function of publisher.
35. The non-transitory computer-readable storage medium of claim 34, wherein the visual representation indicates a similarity of one or more publishers of the plurality of documents based on the plurality of semantic relationships determined as the function of publisher.Attorney Docket No.: P10445PC00 Patent36. The non-transitory computer-readable storage medium according to any of claims 31- 35, wherein the one or more labels includes a location label, and wherein the plurality of semantic relationships are determined as a function of location.
37. The non-transitory computer-readable storage medium of claim 36, wherein the visual representation indicates a similarity of one or more locations of the plurality of documents based on the plurality of semantic relationships determined as the function of location.
38. The non-transitory computer-readable storage medium according to any of claims 30-37, wherein the visual representation is based at least in part on a directed graph, wherein a node of the directed graph represents a document of the plurality of documents, and wherein an edge of the directed graph represents at least one relationship of the plurality of semantic relationships between two adjacent nodes of the directed graph.
39. The non-transitory computer-readable storage medium according to any of claims 30-38, wherein the apparatus is caused to further perform: filtering the plurality of documents based on the similarity matrix, wherein the visual representation is based on the filtering.
40. The non-transitory computer-readable storage medium according to any of claims SOS , wherein the apparatus is caused to further perform: filtering at least a portion of the content of the plurality of documents based on one or more filtering criteria, wherein the visual representation is based on the filtering of the at least the portion of the content.
41. The non-transitory computer-readable storage medium according to any of claims 30- 40, wherein the apparatus is caused to further perform: grouping the plurality of documents into a plurality of groups based on the similarity matrix.Attorney Docket No.: P10445PC00 Patent42. The non-transitory computer-readable storage medium of claim 41, wherein the apparatus is caused to further perform: filtering the plurality of groups based on a number of elements in each group of the plurality of groups, wherein the visual representation is based on the filtering.
43. The non-transitory computer-readable storage medium according to any of claims 41-42, wherein the apparatus is caused to further perform: ranking the plurality of groups based on a number of elements in each group of the plurality of groups; and filtering the plurality of groups based on the ranking.
44. The non-transitory computer-readable storage medium according to any of claims 30-43, wherein the apparatus is caused to further perform: for each document of the plurality of documents, determining a maximum similarity value between a plurality of pairs comprising said each document and other documents of the plurality documents; and assigning said each document as a sole element of its own group based on determining that the maximum similarity value is below a threshold value.
45. The non-transitory computer-readable storage medium according to any of claims 30-44, wherein the apparatus is caused to further perform: filtering the plurality of documents based on one or more time criteria, wherein the visual representation is based on the filtering.
46. The non-transitory computer-readable storage medium according to any of claims 30-45, wherein a number of tokens in the plurality of documents is greater than a context window of a large language model used for processing the plurality of documents.
Citation Information
Patent Citations
Systems and methods for document processing using machine learning
US20180300315A1
Methods and Systems for Identifying a Level of Similarity Between a Plurality of Data Representations
US20220091817A1
Processing System for Generating a Playlist from Candidate Files and Method for Generating a Playlist
US20220107975A1
Self-supervised document-to-document similarity system
US20220405504A1
Table information extraction and mapping to other documents
US20230065915A1